Convert scanned PDFs and document images into clean Markdown using docling for layout (figures, tables, reading order) plus a vision-language OCR model. Use when a user needs high-quality OCR of scanned documents, historical literature, or photographed pages — preserving multi-column reading order, diacritics, special characters, and figures. Supports local vLLM/Ollama servers and cloud vision APIs (OpenAI, Anthropic). Assumes an OCR backend already exists.
69
87%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
Turn scanned PDFs and document images into clean, structured Markdown with figures preserved. The pipeline combines:
docling supplies reliable figure/table crops; the VLM supplies high-quality text. Born-digital PDFs are detected and extracted directly (no server needed).
Adapted from Bruno de Medeiros' OntoMorphoGrapher proof-of-concept, where the
olmocr-docling backend processed an 18-PDF historical-entomology corpus
(1833–2015, 5 languages) with no catastrophic failures.
Use when:
Do not use when:
extract-from-pdfs skill, which can run downstream of this one.This skill does not start or manage any server. Standing one up is out of scope. Before running, the user must have one of:
--backend olmocr-docling);--backend ollama);--backend anthropic-docling) or OpenAI
(--backend vlm-docling pointed at the OpenAI base URL).Always remind the user of this requirement and confirm which backend they
have. For example setup commands (illustrative, not maintained), see
references/server_setup.md. The born-digital PDF fast path needs no backend.
conda env create -f environment.yml
conda activate document_ocrThen make the backend reachable (pick one):
export OCR_HOST=http://YOUR_HOST:30001 # vLLM / Ollama / OpenAI-compatible
# or
export ANTHROPIC_API_KEY=sk-ant-... # cloud Claude
# or
export OPENAI_API_KEY=sk-... # cloud OpenAI--backend | What it uses | Needs |
|---|---|---|
olmocr-docling ★ | docling figures + olmOCR-2 (vLLM) | --host |
vlm-docling | docling figures + any VLM (OpenAI-compat) | --host (+ --api-key for cloud) |
anthropic-docling | docling figures + Claude (cloud) | --api-key / ANTHROPIC_API_KEY |
ollama | full-page DeepSeek-OCR (Ollama) | --host |
docling | docling layout + per-region OCR | --host |
★ Recommended. See references/backends.md for the full matrix, why full-page
beats per-region OCR, model defaults, and tuning notes.
.png/.jpg/.jpeg/.tif/.tiff), or a
directory of them. Ask where output should go.python scripts/ocr_document.py --input docs/ --dry-runDIGITAL or NEEDS_OCR so the user knows what will hit
the server.python scripts/ocr_document.py \
--input docs/ --output-dir out/ \
--backend olmocr-docling --host "$OCR_HOST" --dpi 200out/<stem>/<stem>.md and the figures/ crops. Iterate if needed:
raise --dpi, switch --model, or adjust the caption regex / bad_words /
language list (see references/backends.md).Per-document folder with the stitched markdown (<stem>.md), a per-page JSON
cache/, cropped figures/, and rendered pages/. Full schema and frontmatter
fields: references/output_format.md.
--dry-run first and a small sample before the full
run, so cost and quality are known before committing.Pipeline and prompts adapted from Bruno de Medeiros' OntoMorphoGrapher proof-of-concept (docling + olmOCR-2 over vLLM).
5bdff99
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.