CtrlK
BlogDocsLog inGet started
Tessl Logo

ingesting-clinical-documents

Turn scanned faxes, images, and CSV/CDA exports into clean text ready for OpenMed de-identification and NER, fully on-device. Use when the user has clinical documents (image scans, photographed/faxed notes, tabular CSV/TSV exports, C-CDA XML) and needs OCR or structured intake before openmed.deidentify and openmed.analyze_text, asks about openmed.multimodal, OCR engines (Tesseract / PaddleOCR), tabular redaction, or layout and reading order. Covers the verified ocr() and redact_document() entry points and the ExtractedDocument contract. Pairs before deidentifying-clinical-text and extracting-clinical-entities.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An excellent action-oriented skill body: runnable code for every supported path, explicit pipeline ordering with sibling hand-offs, and honest error-mode documentation. The only trims available are the duplicated submodule-import and PDF/DOCX statements between the main sections and the gotchas list; validation checkpoints exist but are listed as gotchas rather than integrated steps.

DimensionReasoningScore

Conciseness

The body is efficient and assumes competence — no primer on what OCR or C-CDA is, every section carries OpenMed-specific facts ("ocr() lives in the submodule (it is intentionally not re-exported)", "PDF and DOCX have no live handler yet"). It is not a 5 because two facts are stated twice: the submodule import and the PDF/DOCX UnsupportedDocumentError both appear in their main sections and again verbatim in "Edge cases & gotchas", which could be trimmed.

4 / 5

Actionability

All code is copy-paste ready with real imports ("from openmed.multimodal.ocr import ocr", "from openmed.multimodal import redact_document"), executable calls, and expected outputs ("doc.text", "doc.spans[:3]", "redacted.manifest"). Install commands are exact ("pip install \"openmed[multimodal]\"", "pip install \"openmed[ocr-paddle]\"") and the common cases — image OCR, one-step redact_document, CSV redaction, offset-to-bbox mapping — each get a runnable example. Not a 4: no example is pseudocode and none is missing key details.

5 / 5

Workflow Clarity

The sequence is explicit and well-ordered — "This is the **first** stage. After intake, hand off to deidentifying-clinical-text then extracting-clinical-entities", plus a numbered two-step quick start and a clear one-step vs two-step decision rule ("Use redact_document when you want OpenMed to own intake **and** redaction"). Checkpoints exist ("inspect OcrWord.confidence to flag low-quality pages", "MissingDependencyError with an install hint", run NER on the **redacted** text) but they live in a gotchas list rather than being woven into the pipeline steps, which fits anchor 4 ("most checkpoints present; minor validation gaps") better than 5.

4 / 5

Progressive Disclosure

SKILL.md is a true overview: the full data contract, engine list, and tabular pipeline are split into references/multimodal-ingest.md (verified to exist), which is one level deep and clearly signaled at three points ("See [references/multimodal-ingest.md] ... for the full contract, engines, and the tabular pipeline"). Sections are well-organized (when to use, install, quick start, per-format sections, hand-off, gotchas), matching the anchor-5 pattern of a concise overview with well-signaled single-depth references. Not a 4: there are no organization gaps or inline content that belongs in the reference file.

5 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A model description: concrete capability statement, an explicit "Use when" clause rich in natural synonyms, verified entry points, and clear pipeline positioning against sibling skills. It is dense with trigger content rather than padded, and every claim is specific to the openmed.multimodal surface.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — "Turn scanned faxes, images, and CSV/CDA exports into clean text ready for OpenMed de-identification and NER", structured intake, tabular redaction, layout and reading order — and names the exact entry points ("ocr() and redact_document()") and the "ExtractedDocument contract". Coverage is comprehensive across image, tabular, and C-CDA inputs, matching the anchor for multiple specific concrete actions. It is not a 4 because there are no minor gaps: every supported input family and the core API surface is explicitly named.

5 / 5

Completeness

It explicitly answers both questions: what ("Turn scanned faxes, images, and CSV/CDA exports into clean text ready for OpenMed de-identification and NER, fully on-device") and when ("Use when the user has clinical documents ... and needs OCR or structured intake before openmed.deidentify and openmed.analyze_text"). This mirrors the anchor-5 example's structure with concrete trigger phrases; the "what" is not vague and the "when" clause is explicit, so 4 ("when could be more explicit") does not fit.

5 / 5

Trigger Term Quality

Natural trigger terms are comprehensive with synonyms: "scanned faxes", "images", "image scans", "photographed/faxed notes", "tabular CSV/TSV exports", "C-CDA XML", "OCR", "Tesseract / PaddleOCR", "tabular redaction", and "layout and reading order", plus module-level triggers ("asks about openmed.multimodal", "openmed.deidentify"). Not a 4: even the engine names and both CSV/TSV and C-CDA format variants a user would naturally say are present.

5 / 5

Distinctiveness Conflict Risk

The niche is clear — on-device clinical document intake feeding openmed.deidentify/analyze_text — and it explicitly delimits against sibling skills ("Pairs before deidentifying-clinical-text and extracting-clinical-entities"), so triggers like "C-CDA", "PaddleOCR", and "tabular redaction" are unlikely to fire the wrong skill. Not a 4: there is no meaningful overlap risk with adjacent skills since the pipeline position and module names are stated.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
maziyarpanahi/openmed
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.