CtrlK
BlogDocsLog inGet started
Tessl Logo

evaluating-with-leakage-gates

Evaluate an OpenMed de-identification or clinical NER model against the leakage-first release gates G1a through G8, which gate releases on residual PHI leakage rather than on F1. Use when the user wants to run the OpenMed eval harness on a synthetic golden set, decide whether a de-id model is RELEASABLE or QUARANTINED, enforce direct-identifier recall floors, require zero critical leakage, fit calibration thresholds, or produce a signed gate report. Trigger on "release gate", "leakage", "is this model safe to ship", "G1a", "G3", "quarantine", "recall floor", or "calibration thresholds" in an OpenMed de-id context.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, dense, practitioner-oriented body: executable quick start, a compact gate table, a sequenced fail-closed workflow, and a valuable edge-cases section. The only deductions are minor — inline version-pinned floor constants that duplicate the authoritative source, and a placeholder variable in the calibration example.

Suggestions

Move version-specific floor values (v1.6 vs v2.0 recall numbers) out of the gate table into the referenced openmed.eval.release_gates constants, keeping only the invariant rule per gate inline.

Show how calibration_samples is obtained (e.g., one line constructing or loading the held-out score/target samples) so the write_calibration_artifacts snippet is copy-paste runnable end to end.

DimensionReasoningScore

Conciseness

The body is lean and assumes competence (no explaining what HIPAA or de-id is), but the gate table duplicates version-specific floor numbers ('recall ≥ 0.990 (v1.6) / 0.995 (v2.0)') inline even while advising to confirm constants in openmed.eval.release_gates — time-sensitive constants not confined to a deprecated/old-patterns section.

4 / 5

Actionability

The quick-start path is fully executable (run_suite call with realistic metadata, ReleaseGate.evaluate, CLI with concrete flags), but the calibration snippet passes an undefined placeholder 'calibration_samples' rather than showing how it is constructed — a minor gap.

4 / 5

Workflow Clarity

The seven-step workflow is clearly sequenced with an explicit validation checkpoint ('Read the per-gate results' with gate/passed/reason/details) and a fail-closed hard stop, plus subgroup auditing so an aggregate pass can't hide an under-protected group — feedback loops are present for this batch gating operation.

5 / 5

Progressive Disclosure

No bundle files exist and none are needed: the self-contained body is organized into well-signaled sections (gates table, quick start, workflow, hand-offs, edge cases, standards) with one-level-deep pointers to the source of truth (openmed/eval/release_gates.py) and sibling skills, and no nested reference chains.

5 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is exemplary: it states concrete capabilities with named artifacts, gives an explicit 'Use when' clause plus a literal trigger-phrase list, and carves out a distinct niche from sibling skills. It is dense but every clause is substantive rather than padded.

DimensionReasoningScore

Specificity

Lists multiple concrete actions with named artifacts — run the eval harness on a synthetic golden set, decide RELEASABLE vs QUARANTINED, enforce recall floors, require zero critical leakage, fit calibration thresholds, produce a signed gate report — with no coverage gaps.

5 / 5

Completeness

Explicitly answers 'what' (evaluate against the leakage-first release gates G1a–G8) and 'when' via both a 'Use when the user wants to...' clause and a concrete 'Trigger on...' phrase list.

5 / 5

Trigger Term Quality

Triggers include natural user phrasing ('is this model safe to ship', 'release gate', 'quarantine', 'recall floor') plus synonyms and gate IDs (G1a, G3), covering the terms a user in this domain would actually say.

5 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (release gating rather than NER scorecards or CI wiring, distinguished by 'gates on residual PHI leakage rather than on F1') and scopes the generic term 'leakage' with 'in an OpenMed de-id context', minimizing overlap risk.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
maziyarpanahi/openmed
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.