CtrlK
BlogDocsLog inGet started
Tessl Logo

document-ocr

Convert scanned PDFs and document images into clean Markdown using docling for layout (figures, tables, reading order) plus a vision-language OCR model. Use when a user needs high-quality OCR of scanned documents, historical literature, or photographed pages — preserving multi-column reading order, diacritics, special characters, and figures. Supports local vLLM/Ollama servers and cloud vision APIs (OpenAI, Anthropic). Assumes an OCR backend already exists.

69

Quality

87%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable skill body: copy-paste-ready commands, a clear sequenced workflow with preflight and per-page validation, and exemplary progressive disclosure via three real one-level references. The only weakness is minor verbosity in the provenance/corpus note and the lack of an explicit retry loop after page failures.

DimensionReasoningScore

Conciseness

The body is efficient — no explaining of what PDFs/Markdown/OCR are and no library tutorials — but the provenance note describing the 18-PDF historical-entomology corpus (1833–2015, 5 languages) is a mild over-share of background that could be trimmed, fitting the 'efficient; minor instances of over-explanation' anchor.

4 / 5

Actionability

Fully executable, copy-paste-ready commands throughout — 'conda env create -f environment.yml', 'python scripts/ocr_document.py --input docs/ --dry-run', the full run invocation with --backend/--host/--dpi — plus a complete actionable backend matrix and a concrete file-extension list, covering the common cases.

5 / 5

Workflow Clarity

A clear 5-step workflow (confirm backend, gather inputs, classify via --dry-run, run OCR, review) with explicit preflight validation ('fails with a clear message if the backend isn't configured') and per-page-failure handling, plus cost/safety guidance for cloud batch runs; falls just short of 5 because there is no explicit validate->fix->retry loop beyond 'iterate if needed'.

4 / 5

Progressive Disclosure

The body is a clear overview with well-signaled one-level-deep references — references/server_setup.md, references/backends.md ('See ... for the full matrix ...'), references/output_format.md ('Full schema and frontmatter fields') — all of which exist as real, non-nested files, with detail appropriately split out and scripts' internals kept out of the overview.

5 / 5

Total

18

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description that explicitly answers both what the skill does and when to use it, with concrete trigger phrases and good natural keyword coverage. It distinguishes itself via the backend-exists prerequisite, though specificity sits just below the top anchor because the core capability is one conversion operation rather than a list of distinct actions.

DimensionReasoningScore

Specificity

Concrete actions are named — 'Convert scanned PDFs and document images into clean Markdown', 'preserving multi-column reading order, diacritics, special characters, and figures', 'Supports local vLLM/Ollama servers and cloud vision APIs' — with comprehensive coverage, though the core action is a single convert-to-Markdown operation with preservation details rather than a list of distinct operations.

4 / 5

Completeness

Both 'what' ('Convert scanned PDFs and document images into clean Markdown using docling for layout ... plus a vision-language OCR model') and 'when' ('Use when a user needs high-quality OCR of scanned documents, historical literature, or photographed pages') are explicitly stated with concrete trigger phrases.

5 / 5

Trigger Term Quality

Strong natural keyword coverage — 'scanned PDFs', 'document images', 'OCR', 'scanned documents', 'historical literature', 'photographed pages' — with only a few synonyms/extensions (e.g. .pdf, image-to-text) missing, fitting the 'good coverage, a few natural terms missing' anchor.

4 / 5

Distinctiveness Conflict Risk

Clear niche (OCR of scanned documents via docling + VLM) with the 'Assumes an OCR backend already exists' constraint and explicit disambiguation from born-digital extraction giving it mostly distinct triggers, with only minor overlap risk against a generic extract-from-pdfs skill.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
brunoasm/my_claude_skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.