CtrlK
BlogDocsLog inGet started
Tessl Logo

extracting-pii-entities

Detect PHI/PII spans in clinical text with OpenMed's extract_pii without altering the text. Use when the user wants to find names, dates, MRNs, phone numbers, addresses, SSNs, or other identifiers and get their offsets and labels (not redact them), inspect what would be removed before de-identifying, route spans to a custom redactor, normalize labels to a canonical taxonomy, or filter by confidence and language. Covers extract_pii, the PIIEntity fields, CANONICAL_LABELS / normalize_label, and how it differs from deidentify. Pairs before reidentifying-text and deidentifying-clinical-text.

74

Quality

93%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

High

Do not use without reviewing

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, highly actionable single-file skill: executable examples throughout, genuine gotchas (.confidence vs .score, language coverage, offset-ordering), and clear hand-off guidance to sibling skills. The only room for improvement is splitting peripheral detail (parameter reference, standards links) into a reference file and tightening a few prose sections.

Suggestions

Move the "Standards & references" links and the full parameter signature into a one-level-deep reference file (e.g. references/api.md) to slim the SKILL.md body and strengthen progressive disclosure.

Tighten the "Hand-off to / from OpenMed" section — the deidentify and keep_mapping call examples can be compressed to one line each since the sibling skills carry the detail.

Consider adding a short "verify coverage" step (get_default_pii_model(lang) before running) as an explicit numbered checkpoint in the workflow to close the validation gap.

DimensionReasoningScore

Conciseness

The body is code-forward and lean — the comparison table, "Signature & key parameters", and "Edge cases & gotchas" all earn their tokens. Minor trimmable padding exists (prose in "Hand-off to / from OpenMed" and the "Standards & references" section), keeping it below anchor 5's "every token earns its place".

4 / 5

Actionability

Every example is complete, executable Python with synthetic data: quick start, canonical-label normalization, and an offset-based redactor with the highest-offset-first ordering explained. Code is copy-paste ready and covers the common cases (inspect, normalize, route downstream).

5 / 5

Workflow Clarity

Decision guidance is clear ("When to use" section, extract_pii vs deidentify table) and each sub-task is an unambiguous recipe with useful checkpoints ("Use openmed.get_default_pii_model(lang) to confirm coverage", the `assert canon in CANONICAL_LABELS` check). It falls short of anchor 5 only because no explicit validate-and-retry loop is spelled out, though the non-destructive nature of the operation makes that a minor gap.

4 / 5

Progressive Disclosure

Sections are well-organized and clearly navigable, and the body is appropriately lean for a single API surface. However, the skill is ~155 lines fully inline with no bundle files — content like the full parameter signature and the standards/HIPAA links could live in a one-level-deep reference file, so it sits at anchor 4 rather than 5.

4 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: third-person voice, dense with concrete actions and natural trigger phrases, explicitly answering both what the skill does and when to use it. It also cleanly differentiates detection from the adjacent redaction/re-identification skills.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — "Detect PHI/PII spans in clinical text... without altering the text", "get their offsets and labels (not redact them)", "route spans to a custom redactor", "normalize labels to a canonical taxonomy", "filter by confidence and language" — with comprehensive coverage of the capability. It exceeds anchor 4 because there are no minor gaps; detection, preview, routing, normalization, and filtering are all explicitly named.

5 / 5

Completeness

Both "what" ("Detect PHI/PII spans in clinical text with OpenMed's extract_pii without altering the text") and "when" ("Use when the user wants to find names, dates, MRNs... or get their offsets and labels (not redact them), inspect what would be removed before de-identifying...") are explicitly and concretely stated, matching the anchor-5 example pattern.

5 / 5

Trigger Term Quality

Natural user vocabulary is comprehensively covered: "names, dates, MRNs, phone numbers, addresses, SSNs, or other identifiers", "inspect what would be removed before de-identifying", plus synonyms like "identifiers", "de-identifying", "spans", "offsets". A user needing this skill would naturally say these phrases.

5 / 5

Distinctiveness Conflict Risk

The skill occupies a clear niche (detection-without-redaction) and explicitly disambiguates from siblings: "not redact them", "how it differs from deidentify", and "Pairs before reidentifying-text and deidentifying-clinical-text". Conflict risk with redaction skills is minimal.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
maziyarpanahi/openmed
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.