Dev & Engineering

Xberg Format-Specific Extraction Workflows

A reference skill for agents that spells out the exact extraction paths and security validation steps Xberg uses for DOCX, PDF, archives, structured text, and email formats.

66/ 100
Recommended

Generally reliable with disclosed limitations; trial as directed and keep a rollback path.

See how it was scored ↓
Works as-is in
Codex · Claude Code
Stars
★ 9.4k
Last updated
3d ago
License
MIT
pdf-extractionoffice-metadatazip-bomb-protectiondocument-parsing
+4structured-dataemail-extractionarchive-extractionrust

What does this skill do, and when should you use it?

This is one of 54 skills bundled in the xberg-io/xberg monorepo, located at .ai-rulez/skills/format-specific-extraction/SKILL.md. It contains no executable scripts; instead it is a body of format-specific workflow knowledge: for Office XML, PDF, ZIP/TAR/7z/GZIP archives, JSON/YAML/TOML/XML structured text, and EML/MSG email, it specifies which validators, parsers, and source modules to invoke. For developers maintaining or extending Xberg's extractors it is a precise roadmap; for users who only want to call the extraction API, it reads more like internal engineering documentation. The repository is MIT-licensed, but the skill assumes the reader can follow Rust code references such as extractors/pdf/mod.rs.

  • Office XML (DOCX/PPTX/ODT): runs ZipBombValidator first, unpacks word/document.xml, ppt/slides/*.xml, and content.xml, then streams through quick-xml with DepthValidator and StringGrowthValidator to pull text, tables, and metadata
  • PDF: loads via pdf_oxide::PdfDocument::from_bytes, extracts text per page, falls back to OCR when force_ocr is set or no searchable text exists, optionally extracts tables
  • Archives: validates with ZipBombValidator before any extraction, lists file metadata, extracts plaintext files only, and assembles results with build_archive_result()
  • Structured text: detects JSON/YAML/TOML/XML by MIME, parses with the format-specific library, and pretty-prints
  • Email: parses EML/MSG headers, extracts text/HTML body, and processes attachments
  • Adding a format: a seven-step checklist from MIME registration through implementing the DocumentExtractor trait to feature gating and fixture tests
Good fit
  • A Rust developer adding a new document format extractor to Xberg who needs the full steps for implementing the DocumentExtractor trait and registering it
  • An engineer auditing extraction pipeline security who wants to confirm the correct ordering of ZipBombValidator, DepthValidator, and StringGrowthValidator on Office and archive paths
  • A maintainer debugging PDF output who needs to locate the pdf_oxide OCR fallback conditions and the feature gate
  • A team wiring document-extraction knowledge into an agent so the model picks the correct workflow per format instead of guessing
Not a fit
  • End users who just want to install a library and extract documents — this is an engineering reference for Xberg contributors; regular users should follow the README's installation and Quick Start instead
  • Non-Rust developers — the skill body is written entirely as Rust module paths and APIs, with no equivalent guidance for the Python/Node or other bindings

How do you install this skill?

Before you use it
  • The skill documentation does not include explicit security/permission guidance, nor does it mention user confirmation or data-flow transparency.
  • Some capabilities (e.g., provider-hosted OCR or LLM) may rely on external services potentially inaccessible from mainland China; evaluate reachability from Chinese networks.
  • The skill document does not define a clear non-fit range or trigger boundaries, which may lead to misuse.
  • The skill itself has no versioning or changelog, maintenance responsibility is unclear, and the update path is unknown.
Before you start
Your agent needs
  • Local filesystem

The skill ships with the Agent Skills plugin in the xberg-io/xberg repo; the README documents these routes:
Claude Code

/plugin marketplace add xberg-io/xberg
/plugin install xberg@xberg

Codex CLI

/plugins add https://github.com/xberg-io/xberg

Once installed, the skill lives at .ai-rulez/skills/format-specific-extraction/SKILL.md inside the repo, one of 54 bundled skills.

How do you use this skill?

Try saying

Once installed, send your agent any of these to trigger it:

  • I want to add a new email format extractor to Xberg besides .eml — walk me through each step of the format-specific-extraction checklist and which files to touch
  • Check the archive extraction path for me: should ZipBombValidator run before or after build_archive_result()?
  • After a DOCX is parsed into a cell grid, how do I convert it to a GitHub-flavored Markdown table?
  • This PDF has no searchable text — at what point does Xberg trigger the OCR fallback?

This is a pure reference skill with no scripts or CLI operations. After installing the plugin, when a task touches format-specific extraction, security validation ordering, or source code location, the model reads the matching section of SKILL.md: the five-step Office XML flow, the pdf_oxide PDF flow, the validate-before-extract rule for archives, the StructuredExtractor for structured text, email parsing, and the seven-step new-format checklist (register MIME in EXT_TO_MIME → implement DocumentExtractor → set supported_mime_types() and priority() (default 50) → register in register_default_extractors() → feature-gate if optional → apply security validators → add tests with fixtures).

What are this skill's strengths and limitations?

Pros
  • Every workflow cites exact source locations (e.g. extractors/docx.rs, extraction/office.rs), so you can jump straight to the code
  • Security ordering is explicit: ZipBombValidator must run before any archive extraction; Office XML parsing layers DepthValidator and StringGrowthValidator against malicious input
  • The seven-step new-format checklist, including feature gating and test requirements, lowers the contribution barrier
Limitations
  • Pure knowledge document — no executable scripts or automation, so every step still requires manual developer work
  • Tightly coupled to Xberg's internal codebase layout; of little use outside this repository
  • SKILL.md carries no version information, so it is unclear which Xberg source layout it matches

How does this skill compare with similar options?

Side by side with related skills; every score comes from the same FSRS standard.

Skill FS score Stars Last updated License
Xberg Format-Specific Extraction Workflows this page 66 · Recommended ★ 9.4k 3d ago MIT
Xberg API Server & MCP Protocol Integration 52 · Use with care ★ 9.4k 3d ago MIT
Document Text Extractor ✓ Anthropic · Official 52 · Use with care ★ 422 1mo ago —
Skill Seekers Skill Builder 48 · Use with care ★ 15k 11d ago MIT
Azure Document Intelligence for Java ✓ Microsoft · Official 46 · Use with care ★ 2.7k 3mo ago MIT

Xberg is the next iteration of Kreuzberg (kreuzberg-dev/kreuzberg-v4-lts) — the same document-intelligence engine rebuilt under a fresh v1 line; this document describes the Xberg-side extractor implementation workflows.

How did FollowSkills review this skill?

FollowSkills review · FSRS-2.0
Recommended
66/ 100 5-point scale 3.3 / 5
The upstream repository has new commits since this review. The score still applies to the reviewed revision shown and may not cover the latest changes.
1Trust21 / 25 · 4.2/5

The skill is well-integrated with the Xberg document-extraction framework, which explicitly acknowledges handling potentially untrusted user documents. The repository contains a detailed security policy, dedicated input validators (ZipBombValidator, DepthValidator, StringGrowthValidator etc.), and careful default security limits, all of which strongly support safe handling of untrusted input. However, at the skill level, there is no explicit mechanism for user confirmation or transparency of data flow; the skill documentation does not cover permissions, rollback, or user confirmation. The skill's functional scope is extraction, without external network calls, but some capabilities like OCR may depend on external services. Therefore, deduction for incomplete permission and confirmation guidance at the skill level.

2Reliability9 / 20 · 2.3/5

Instructions in the skill document are consistent, and the format-specific extraction workflows are clear. The repo demonstrates extensive test and CI infrastructure (e.g., Docker CI, benchmark workflows, e2e tests). However, for the selected skill path, no tests or reproducible execution evidence are available; the skill document only references source files, with no tests or verification results. The instructions are self-consistent and well-structured, but evidence for test coverage and edge-case handling is lacking. Given static review, the maximum is capped, and this score of 9 reflects that the happy path appears plausible, but error handling and failure feedback are not detailed.

3Adaptability12 / 15 · 4.0/5

The skill is well-targeted for clear scenarios: extracting content from various specific formats (Office, PDF, archives, structured text, email), which aligns with real needs. However, the non-fit range is not clearly defined; the skill's name and description suggest a general extraction capability, but it covers only format-specific aspects and does not explicitly state when it should not be used. Additionally, documentation mentions that complex content like macros or password-protected files may not be supported, but this is not described as a skill boundary. Regarding environment fit, there is no mention of China-relevant services or reachability from mainland-China networks; some capabilities like provider-hosted OCR or LLM may depend on services potentially inaccessible from China, but the skill's core is local processing. Deduction for insufficient evidence of boundaries and trigger conditions.

4Convention12 / 15 · 4.0/5

The skill document is well-organized, with clear heading hierarchy, diagrams, and source-code references, and follows progressive disclosure. The license is explicit (MIT). Installation and dependency notes are provided (indicating Cargo feature flags). Maintenance responsibility (regular CI and versioning) exists, but the skill itself is not versioned or changelogged, and it is unclear whether the skill will be regularly updated. Known limitations are not explicitly disclosed at the skill level. Deduction for lack of versioning and explicit maintenance responsibility, and hidden assumptions such as needing to build the project first.

5Effectiveness7 / 15 · 2.3/5

Public evidence indicates that the skill can accomplish the core task of extracting content from specified formats. The repository contains extensive benchmark efforts and tests across various formats, suggesting practical effectiveness. However, for the selected skill path, there is no directly usable output example or verification result. The skill document references implementation details but does not demonstrate output for key paths or provide representative results. The benefit compared to manual effort is unclear, and static review cannot verify output quality. Hence, capped at 7 due to lack of direct, verifiable evidence from static review.

6Verifiability5 / 10 · 2.5/5

The repository contains abundant auditable primary material: detailed CI workflows, benchmark configuration, test suites, and reproducible build configurations. These provide supporting evidence for the framework's capabilities, which underpin the skill. However, the specific skill path has no dedicated tests or verification; framework coverage is diffuse. There is no independent verification or cross-source corroboration. Given static read and no tests specifically for the skill path, capped at 5 for partial evidence.

1 2 3 4 5 6

Open a dimension to read why it scored that way

Reviewed Aug 07, 2026 Reviewed revision fbdb9f18ca7c Review evidence[1][2][3][4][5][6][7][8][9][10][11][12]

Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.

See the full review method →

FAQ

Does this skill extract documents for me?
No. It is an engineering reference for developers and contains no executable scripts; actual extraction is done by Xberg itself (library, CLI, REST API, or MCP server).
Is it safe to process maliciously crafted Office or archive files?
The documented workflows run ZipBombValidator before parsing, layer DepthValidator and StringGrowthValidator into Office XML parsing, and require applying security validators to user content.
Is OCR enabled by default for PDFs?
No. The flow checks config.force_ocr or whether searchable text exists and only falls back to OCR when needed; PDF extraction is also behind the #[cfg(feature = "pdf")] feature gate.
How do I add a new format?
Follow the seven-step checklist: register the MIME type in EXT_TO_MIME in core/mime.rs, implement the DocumentExtractor trait, set supported_mime_types() and priority() (default 50), register in register_default_extractors(), optionally feature-gate, apply security validators, and add tests with fixture files.

More skills from this repository

All from xberg-io/xberg

Related skills