Productivity & Collaboration ✓ Anthropic · Official

PDF Document Processing Skill

Helps Claude read, transform, create, and manage PDF documents.

45/ 100
Not recommended

Current benefit does not outweigh risk or uncertainty.

See how it was scored ↓
Works as-is in
Codex · Claude Code · Claude.ai · Claude API
Stars
★ 180k
Last updated
6d ago
pdf-extractionpdf-mergingpdf-splittingpdf-forms
+3pdf-generationocrtable-extraction

What does this skill do, and when should you use it?

This PDF skill is part of the anthropics/skills repository and targets tasks involving PDF reading, extraction, editing, and creation. It provides examples using Python libraries and command-line utilities. Covered operations include text and table extraction, merging and splitting PDFs, page rotation, watermarking, image extraction, OCR, form handling, and password protection. The skill is marked Proprietary, and the README describes the related document skills as source-available rather than open source.

Uses pypdf to read PDFs, extract text and metadata, merge or split pages, rotate pages, and encrypt files; uses pdfplumber to extract text and tables, with pandas for combining table data; uses reportlab to create single- and multi-page PDFs; uses pytesseract with pdf2image for OCR on scanned PDFs; and uses pdftotext, pdfimages, qpdf, and pdftk for command-line extraction, image export, merging, splitting, rotation, and decryption.

Good fit
  • Researchers who need text, page counts, or author metadata from PDF files.
  • Analysts who need to extract tables from PDF reports and combine them into a dataset.
  • Office users who need to merge multiple PDFs or split documents by page.
  • Document operators who need to create reports, add watermarks, rotate pages, or protect files with passwords.
  • Users who need to turn scanned PDFs into searchable text.
  • Users handling PDF forms, although the detailed workflow requires the unavailable FORMS.md file.

How do you install this skill?

Before you use it
  • Confirm all external dependencies are installed and check the case-sensitive referenced filenames before use.
  • Obtain explicit confirmation and preserve a recoverable copy before decryption, password handling, overwriting inputs, or producing sensitive outputs.
  • OCR, table extraction, form coordinates, and PDF layout fidelity are not proven by the static materials and require human review.
  • No Chinese-language or Chinese OCR/font compatibility guidance is provided, so Chinese users may need additional configuration.
Before you start
Your agent needs
  • Shell / CLI
  • Local filesystem
Install first
  • Python
  • pypdf
  • pdfplumber
  • pandas
  • reportlab
  • pytesseract
  • pdf2image
  • poppler-utils
  • qpdf
  • pdftk

In Claude Code, run /plugin marketplace add anthropics/skills, then install document-skills@anthropic-agent-skills with /plugin install document-skills@anthropic-agent-skills. The README does not document a separate procedure for copying the skills/pdf folder into other clients.

Generic route: install into Claude Code manually (macOS / Linux)
tmp="$(mktemp -d)"
git clone --depth 1 https://github.com/anthropics/skills.git "$tmp"
mkdir -p ~/.claude/skills
cp -R "$tmp/skills/pdf" ~/.claude/skills/
rm -rf "$tmp"

Generated from the source repository and skill path; it copies only this skill's folder. If the author's install steps above differ, follow those first. To scope it to one project, replace ~/.claude/skills with that project's .claude/skills.

How do you use this skill?

After installing document-skills, make a request involving a PDF, for example: Use the PDF skill to extract the form fields from path/to/some-file.pdf. The SKILL.md says to use this skill whenever a user mentions a .pdf file or asks to produce one.

What are this skill's strengths and limitations?

Pros
  • Covers a broad range of PDF reading, extraction, editing, creation, OCR, and security tasks.
  • Provides both Python and command-line examples for scripting and batch workflows.
  • Addresses text, tables, images, and metadata rather than only plain text extraction.
  • Includes a specific ReportLab warning about avoiding Unicode superscript and subscript characters to prevent font-rendering problems.
Limitations
  • Its SKILL.md lists a Proprietary license, while the README describes the document skills as source-available rather than open source.
  • The examples reference several Python libraries and command-line tools, increasing setup overhead.
  • No test suite, version requirements, or cross-platform validation are provided.
  • Advanced usage and PDF form workflows depend on REFERENCE.md and FORMS.md, whose contents are not included in the supplied material.

How does this skill compare with similar options?

Side by side with related skills; every score comes from the same FSRS standard.

Skill FS score Stars Last updated License
PDF Document Processing Skill this page ✓ Anthropic · Official 45 · Not recommended ★ 180k 6d ago —
Open Cowork PDF Skill 69 · Recommended ★ 2.2k 5d ago MIT
DeepTutor PDF Skill 61 · Recommended ★ 41k 3d ago Apache-2.0
NeMo Retriever Document Retrieval Skill ✓ NVIDIA · Official 41 · Not recommended ★ 3.5k 3d ago Apache-2.0
Image to Editable PPT 54 · Use with care ★ 2.8k 25d ago MIT

The README mentions the document-skills and example-skills plugins, but it does not provide a direct comparison with other PDF skills or tools.

How did FollowSkills review this skill?

FollowSkills review · FSRS-2.0
Not recommended
45/ 100 5-point scale 2.3 / 5
The upstream repository has new commits since this review. The score still applies to the reviewed revision shown and may not cover the latest changes.
1Trust12 / 25 · 2.4/5

The material shows local PDF-processing examples and no malware or covert exfiltration; however, it includes decryption, password protection, overwrite-style repair, and file writes without explicit user confirmation, least-privilege guidance, sensitive-data handling, rollback, or data-flow disclosure, so points are deducted.

2Reliability7 / 20 · 1.8/5

The main paths have readable Python, command-line, and form-processing examples; however, there is no test or execution evidence, dependency availability and abnormal-input handling are incomplete, failure diagnostics are thin, and SKILL.md references REFERENCE.md and FORMS.md while the supplied files use lowercase names, limiting the score.

3Adaptability9 / 15 · 3.0/5

The trigger and PDF task coverage are clear, including reading, conversion, forms, OCR, creation, and modification; however, non-fit boundaries, input/output contracts, platform requirements, and Chinese-language use are not specified, and the workflow depends on local Python, Poppler, Tesseract, and related tools.

4Convention6 / 15 · 2.0/5

The material provides an overview, quick start, layered reference, scripts, and troubleshooting, with generally clear naming; however, installation requirements, versioning, changelog, maintenance ownership, update path, and license governance are incomplete, and filename case references are inconsistent.

5Effectiveness7 / 15 · 2.3/5

The examples cover many common PDF operations and could be adapted directly; however, this is a static review with no representative output or execution verification, while complex forms, OCR, layout fidelity, and cross-tool compatibility may require substantial manual checking, so the static ceiling applies.

6Verifiability4 / 10 · 2.0/5

The source code, scripts, and dependency license notes provide some auditable evidence; however, there are no committed tests, CI coverage, real execution logs, or independent corroboration, so the conclusion is based mainly on static inspection and receives limited credit.

1 2 3 4 5 6

Open a dimension to read why it scored that way

Reviewed Jul 19, 2026 Reviewed revision fa0fa64bdc96 Review evidence[1][2][3][4][5][6][7][8][9][10][11]

Evidence confidence:Low — Mostly static review, author material or a limited demo; useful for discovery, not high-risk decisions.

See the full review method →

FAQ

Is this skill open source?
That cannot be confirmed from the supplied material. The SKILL.md says Proprietary, and the README describes the related document skills as source-available rather than open source.
What tools does it require?
The examples reference pypdf, pdfplumber, pandas, reportlab, pytesseract, and pdf2image, plus pdftotext, pdfimages, qpdf, and pdftk. No complete installation list or version requirements are provided.
Can it process scanned PDFs?
Yes. The example converts PDF pages to images with pdf2image and runs OCR on each page with pytesseract.
Does it include complete PDF form instructions?
The SKILL.md directs users to read FORMS.md for form filling, but that file's contents are not included in the supplied source.

More skills from this repository

All from anthropics/skills

Related skills