What does this skill do, and when should you use it?
This is one of 54 skills bundled in the xberg-io/xberg monorepo, located at .ai-rulez/skills/format-specific-extraction/SKILL.md. It contains no executable scripts; instead it is a body of format-specific workflow knowledge: for Office XML, PDF, ZIP/TAR/7z/GZIP archives, JSON/YAML/TOML/XML structured text, and EML/MSG email, it specifies which validators, parsers, and source modules to invoke. For developers maintaining or extending Xberg's extractors it is a precise roadmap; for users who only want to call the extraction API, it reads more like internal engineering documentation. The repository is MIT-licensed, but the skill assumes the reader can follow Rust code references such as extractors/pdf/mod.rs.
- Office XML (DOCX/PPTX/ODT): runs ZipBombValidator first, unpacks word/document.xml, ppt/slides/*.xml, and content.xml, then streams through quick-xml with DepthValidator and StringGrowthValidator to pull text, tables, and metadata
- PDF: loads via pdf_oxide::PdfDocument::from_bytes, extracts text per page, falls back to OCR when force_ocr is set or no searchable text exists, optionally extracts tables
- Archives: validates with ZipBombValidator before any extraction, lists file metadata, extracts plaintext files only, and assembles results with build_archive_result()
- Structured text: detects JSON/YAML/TOML/XML by MIME, parses with the format-specific library, and pretty-prints
- Email: parses EML/MSG headers, extracts text/HTML body, and processes attachments
- Adding a format: a seven-step checklist from MIME registration through implementing the DocumentExtractor trait to feature gating and fixture tests
- A Rust developer adding a new document format extractor to Xberg who needs the full steps for implementing the DocumentExtractor trait and registering it
- An engineer auditing extraction pipeline security who wants to confirm the correct ordering of ZipBombValidator, DepthValidator, and StringGrowthValidator on Office and archive paths
- A maintainer debugging PDF output who needs to locate the pdf_oxide OCR fallback conditions and the feature gate
- A team wiring document-extraction knowledge into an agent so the model picks the correct workflow per format instead of guessing
- End users who just want to install a library and extract documents — this is an engineering reference for Xberg contributors; regular users should follow the README's installation and Quick Start instead
- Non-Rust developers — the skill body is written entirely as Rust module paths and APIs, with no equivalent guidance for the Python/Node or other bindings
How do you install this skill?
- The skill documentation does not include explicit security/permission guidance, nor does it mention user confirmation or data-flow transparency.
- Some capabilities (e.g., provider-hosted OCR or LLM) may rely on external services potentially inaccessible from mainland China; evaluate reachability from Chinese networks.
- The skill document does not define a clear non-fit range or trigger boundaries, which may lead to misuse.
- The skill itself has no versioning or changelog, maintenance responsibility is unclear, and the update path is unknown.
- Local filesystem
The skill ships with the Agent Skills plugin in the xberg-io/xberg repo; the README documents these routes:
Claude Code
/plugin marketplace add xberg-io/xberg
/plugin install xberg@xbergCodex CLI
/plugins add https://github.com/xberg-io/xbergOnce installed, the skill lives at .ai-rulez/skills/format-specific-extraction/SKILL.md inside the repo, one of 54 bundled skills.
How do you use this skill?
Once installed, send your agent any of these to trigger it:
- I want to add a new email format extractor to Xberg besides .eml — walk me through each step of the format-specific-extraction checklist and which files to touch
- Check the archive extraction path for me: should ZipBombValidator run before or after build_archive_result()?
- After a DOCX is parsed into a cell grid, how do I convert it to a GitHub-flavored Markdown table?
- This PDF has no searchable text — at what point does Xberg trigger the OCR fallback?
This is a pure reference skill with no scripts or CLI operations. After installing the plugin, when a task touches format-specific extraction, security validation ordering, or source code location, the model reads the matching section of SKILL.md: the five-step Office XML flow, the pdf_oxide PDF flow, the validate-before-extract rule for archives, the StructuredExtractor for structured text, email parsing, and the seven-step new-format checklist (register MIME in EXT_TO_MIME → implement DocumentExtractor → set supported_mime_types() and priority() (default 50) → register in register_default_extractors() → feature-gate if optional → apply security validators → add tests with fixtures).
What are this skill's strengths and limitations?
- Every workflow cites exact source locations (e.g. extractors/docx.rs, extraction/office.rs), so you can jump straight to the code
- Security ordering is explicit: ZipBombValidator must run before any archive extraction; Office XML parsing layers DepthValidator and StringGrowthValidator against malicious input
- The seven-step new-format checklist, including feature gating and test requirements, lowers the contribution barrier
- Pure knowledge document — no executable scripts or automation, so every step still requires manual developer work
- Tightly coupled to Xberg's internal codebase layout; of little use outside this repository
- SKILL.md carries no version information, so it is unclear which Xberg source layout it matches
How does this skill compare with similar options?
Side by side with related skills; every score comes from the same FSRS standard.
| Skill | FS score | Stars | Last updated | License |
|---|---|---|---|---|
| Xberg Format-Specific Extraction Workflows this page | 66 · Recommended | ★ 9.4k | 3d ago | MIT |
| Xberg API Server & MCP Protocol Integration | 52 · Use with care | ★ 9.4k | 3d ago | MIT |
| Document Text Extractor ✓ Anthropic · Official | 52 · Use with care | ★ 422 | 1mo ago | — |
| Skill Seekers Skill Builder | 48 · Use with care | ★ 15k | 11d ago | MIT |
| Azure Document Intelligence for Java ✓ Microsoft · Official | 46 · Use with care | ★ 2.7k | 3mo ago | MIT |
Xberg is the next iteration of Kreuzberg (kreuzberg-dev/kreuzberg-v4-lts) — the same document-intelligence engine rebuilt under a fresh v1 line; this document describes the Xberg-side extractor implementation workflows.
How did FollowSkills review this skill?
The skill is well-integrated with the Xberg document-extraction framework, which explicitly acknowledges handling potentially untrusted user documents. The repository contains a detailed security policy, dedicated input validators (ZipBombValidator, DepthValidator, StringGrowthValidator etc.), and careful default security limits, all of which strongly support safe handling of untrusted input. However, at the skill level, there is no explicit mechanism for user confirmation or transparency of data flow; the skill documentation does not cover permissions, rollback, or user confirmation. The skill's functional scope is extraction, without external network calls, but some capabilities like OCR may depend on external services. Therefore, deduction for incomplete permission and confirmation guidance at the skill level.
Instructions in the skill document are consistent, and the format-specific extraction workflows are clear. The repo demonstrates extensive test and CI infrastructure (e.g., Docker CI, benchmark workflows, e2e tests). However, for the selected skill path, no tests or reproducible execution evidence are available; the skill document only references source files, with no tests or verification results. The instructions are self-consistent and well-structured, but evidence for test coverage and edge-case handling is lacking. Given static review, the maximum is capped, and this score of 9 reflects that the happy path appears plausible, but error handling and failure feedback are not detailed.
The skill is well-targeted for clear scenarios: extracting content from various specific formats (Office, PDF, archives, structured text, email), which aligns with real needs. However, the non-fit range is not clearly defined; the skill's name and description suggest a general extraction capability, but it covers only format-specific aspects and does not explicitly state when it should not be used. Additionally, documentation mentions that complex content like macros or password-protected files may not be supported, but this is not described as a skill boundary. Regarding environment fit, there is no mention of China-relevant services or reachability from mainland-China networks; some capabilities like provider-hosted OCR or LLM may depend on services potentially inaccessible from China, but the skill's core is local processing. Deduction for insufficient evidence of boundaries and trigger conditions.
The skill document is well-organized, with clear heading hierarchy, diagrams, and source-code references, and follows progressive disclosure. The license is explicit (MIT). Installation and dependency notes are provided (indicating Cargo feature flags). Maintenance responsibility (regular CI and versioning) exists, but the skill itself is not versioned or changelogged, and it is unclear whether the skill will be regularly updated. Known limitations are not explicitly disclosed at the skill level. Deduction for lack of versioning and explicit maintenance responsibility, and hidden assumptions such as needing to build the project first.
Public evidence indicates that the skill can accomplish the core task of extracting content from specified formats. The repository contains extensive benchmark efforts and tests across various formats, suggesting practical effectiveness. However, for the selected skill path, there is no directly usable output example or verification result. The skill document references implementation details but does not demonstrate output for key paths or provide representative results. The benefit compared to manual effort is unclear, and static review cannot verify output quality. Hence, capped at 7 due to lack of direct, verifiable evidence from static review.
The repository contains abundant auditable primary material: detailed CI workflows, benchmark configuration, test suites, and reproducible build configurations. These provide supporting evidence for the framework's capabilities, which underpin the skill. However, the specific skill path has no dedicated tests or verification; framework coverage is diffuse. There is no independent verification or cross-source corroboration. Given static read and no tests specifically for the skill path, capped at 5 for partial evidence.
Open a dimension to read why it scored that way
Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.
See the full review method →