CtrlK
BlogDocsLog inGet started
Tessl Logo

pdf-extract-shell-first

PDF text extraction with tool cascade prioritizing shell pdftotext before Python fallback

60

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./benchmarks/gdpval/skills/pdf-download-extract-fallback-enhanced-e27e0c/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is highly actionable and the workflow is clearly sequenced with strong validation and feedback loops. Its main weakness is redundancy from restating the cascade as steps, a decision tree, and a full script, plus a meta Migration Notes section that adds tokens without aiding execution.

Suggestions

Collapse the redundancy: keep the step-by-step workflow and either move the full bash script to scripts/pdf-extract-cascade.sh (referenced once) or drop the ASCII decision tree, since both restate the same cascade.

Remove or trim the 'Migration Notes from pdf-download-extract-fallback' section unless lineage is genuinely needed at runtime; it is meta-context that consumes tokens.

Tighten 'Why Shell-First?' to a one-line rationale plus the bullet evidence, removing explanatory framing Claude can infer.

DimensionReasoningScore

Conciseness

The body is mostly useful but restates the same cascade three ways (step-by-step workflow, ASCII decision tree, and a full automated bash script), and the 'Migration Notes' section is meta-commentary not needed to execute, so it could be tightened.

3 / 5

Actionability

It provides copy-paste-ready curl, pdftotext, apt-get, and PyMuPDF commands plus a complete bash script, covering the common cases (URL download, missing poppler, sandbox failure) with executable code.

5 / 5

Workflow Clarity

Steps 0-4 are explicitly sequenced with validation checkpoints (response-evaluation table, 'Verify extraction quality'), feedback loops (install poppler then retry, fallback to PyMuPDF, domain-knowledge degradation), and a failure-mode reference table.

5 / 5

Progressive Disclosure

Sections are clearly headed and well-organized with no nested references, but all ~300 lines live inline in SKILL.md with no bundle files; content like the full automated script could arguably live in a scripts/ file, leaving minor organization gaps.

4 / 5

Total

17

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific about its method and tools but omits any explicit "when to use" trigger guidance, which caps completeness. It is a solid, focused description that would benefit from a Use-when clause and broader natural keywords.

Suggestions

Add a 'Use when ...' clause naming natural triggers, e.g. 'Use when extracting text from PDFs, .pdf files, or scanned documents where read_file may return binary data.'

Include user-natural synonyms and file extensions ("PDFs", ".pdf", "documents") alongside the technical term pdftotext.

Optionally enumerate the concrete actions (download, extract, verify, degrade) to lift specificity toward a 5.

DimensionReasoningScore

Specificity

The description names a concrete action ("PDF text extraction") plus a specific approach ("tool cascade prioritizing shell pdftotext before Python fallback"), naming real tools rather than vague verbs; it stops short of listing multiple distinct actions, so it sits just above the midpoint.

4 / 5

Completeness

It clearly states what the skill does (PDF text extraction via a shell-first cascade) but provides no "Use when..." clause or equivalent trigger guidance, capping completeness at 3 per the rubric.

3 / 5

Trigger Term Quality

"PDF text extraction" is a relevant keyword, but "pdftotext" is technical jargon and natural variants users say ("PDFs", ".pdf", "documents", "forms") are absent, leaving common synonyms missing.

3 / 5

Distinctiveness Conflict Risk

The shell-first cascade niche is fairly distinct from generic document skills, though it could overlap with sibling PDF-extraction skills since no explicit trigger phrases carve out its unique scope.

4 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
HKUDS/OpenSpace
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.