CtrlK
BlogDocsLog inGet started
Tessl Logo

healthcare-eval-harness

Patient safety evaluation harness for healthcare application deployments. Automated test suites for CDSS accuracy, PHI exposure, clinical workflow integrity, and integration compliance. Blocks deployments on safety failures. Use when a healthcare deployment must be gated on patient-safety tests for CDSS accuracy, PHI exposure, and workflow integrity.

68

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-structured operational guide with executable commands and clear gating logic throughout. Its main weakness is repetition — the HIGH-gate pass-rate script appears four times across sections and the CI pipeline duplicates the category commands, inflating token cost without adding guidance.

Suggestions

State the HIGH-gate pass-rate bash block once and reference it from the CI/CD section (or vice versa) instead of repeating the mktemp/jq/bc script four times across sections 4, 5, the CI YAML, and Example 2.

Remove the duplicated '# HIGH gates — 95%+ required' comment line in the CI YAML and consolidate the two single-category Examples, which restate commands already shown.

Reconcile the '95%+ required' threshold with the CI steps that only warn below 95% — either enforce a failing exit code in CI or state explicitly in the pass/fail matrix that HIGH gates are advisory in CI.

DimensionReasoningScore

Conciseness

The body has no concept-explaining padding and assumes competence, but the HIGH-gate pass-rate bash script (mktemp / jest --json / jq / bc) is repeated nearly verbatim in sections 4, 5, the CI YAML, and Example 2, and the YAML contains a duplicated '# HIGH gates — 95%+ required' comment. Anchor 3: mostly efficient but noticeable duplication that could be tightened — not anchor 2's pervasive padding, nor anchor 4's merely trimmable text.

3 / 5

Actionability

Everything is executable: copy-paste-ready npx jest commands, complete jq/bc pass-rate computations, a full GitHub Actions YAML with error annotations, and a concrete report example with expected output. Matches anchor 5's fully executable, common-case-covering standard.

5 / 5

Workflow Clarity

Categories run 'in order' with explicit CRITICAL/HIGH thresholds, a pass/fail matrix, anti-patterns, and zero-test error handling — validation is inherent to the skill. Anchor 4 rather than 5 because the CI HIGH-gate steps only emit warnings below 95%, contradicting the '95%+ required' / matrix language, leaving a minor checkpoint inconsistency.

4 / 5

Progressive Disclosure

No bundle files exist, and the body is a self-contained, well-sectioned document (When to Use, categories, matrix, CI, anti-patterns, examples) that is easy to navigate. Anchor 4: good structure and appropriate placement, though the CI YAML section largely repeats inline content that could be split out.

4 / 5

Total

16

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states concrete capabilities, the blocking behavior, and an explicit 'Use when' clause with concrete healthcare-specific triggers. Trigger vocabulary is good, though it could add synonyms like EMR/EHR to broaden natural-user matching.

DimensionReasoningScore

Specificity

Names concrete capabilities ('Automated test suites for CDSS accuracy, PHI exposure, clinical workflow integrity, and integration compliance', 'Blocks deployments on safety failures') — several specific actions with minor gaps in coverage (how the suites execute is unspecified). Falls at anchor 4: more actions than anchor 3's 1-2, but not the fully comprehensive coverage of anchor 5.

4 / 5

Completeness

Clearly answers both: what ('Automated test suites for CDSS accuracy, PHI exposure... Blocks deployments on safety failures') and when, with an explicit 'Use when a healthcare deployment must be gated on patient-safety tests...' clause containing concrete trigger phrases. Matches anchor 5 exactly; anchor 4 would require the 'when' to be less explicit.

5 / 5

Trigger Term Quality

Includes natural phrases users would say ('patient safety', 'healthcare deployment', 'CDSS accuracy', 'PHI exposure') but misses common variations and synonyms like EMR/EHR or 'clinical decision support' spelled out. Good keyword coverage per anchor 4, short of anchor 5's comprehensive synonym/extension coverage.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (patient-safety deployment gating for healthcare applications) with distinct trigger terms (CDSS, PHI, patient-safety tests) that no general testing or deployment skill would claim. Minimal conflict risk per anchor 5.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
affaan-m/ECC
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.