CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-auditor

A comprehensive auditor for any agent skill — including Manus, OpenClaw/ClawHub, Claude, LobeHub, or custom SKILL.md-based skills. Use this skill whenever a user wants to evaluate, audit, review, score, or quality-check an agent skill before publishing, updating, or deploying. Covers two hard veto gates (structural redlines + research integrity redlines), static quality scoring across 25 criteria (ISO 25010 + OpenSSF + Agent), dynamic test input generation, multi-mode execution testing, multi-layer output evaluation with five specialized category rubrics (Evidence Insight / Protocol Design / Data Analysis / Academic Writing / Other), a Research Veto that applies to all four research categories, human eval viewer generation, actionable P0/P1/P2 optimization recommendations, and automatic skill improvement that outputs a polished, production-ready SKILL.md. Also use whenever a user says "audit my skill", "evaluate my skill", "improve my skill", or wants a corrected version after evaluation.

64

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skill-auditor/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers a well-sequenced, heavily templated audit pipeline with mandatory validation gates and a real, complete reference bundle. Its weaknesses are a phantom Step 9 (the promised improvement procedure is undefined), an unwired script dependency, and ~75 lines of duplicated schema summary plus changelog that inflate the token budget without aiding execution.

Suggestions

Resolve the 'Step 9 always runs' reference: either add an explicit Step 9 defining the procedure that produces the polished, production-ready SKILL.md (which the description promises), or delete the sentence and the improvement promise.

Wire 'scripts/evaluate_skill.py' into the workflow: give the exact invocation command (e.g., 'python scripts/evaluate_skill.py <skill_path>') in Steps 1–3 where the script's docstring says it applies, or remove it from Dependencies.

Move the JSON top-level nodes table and pre-emit checklist digest out of SKILL.md into references/report_json_schema.md (keeping only the pointer and the drift-warning), and drop or relocate the Changelog section — this trims ~75 lines that duplicate or don't serve execution.

DimensionReasoningScore

Conciseness

Mostly efficient operational content (templates, formulas, path rules), but there is unnecessary material: a full 'Changelog' section with multi-paragraph rationale about past audits that is irrelevant to executing the pipeline, and ~60 lines of JSON key rules and a pre-emit checklist that duplicate the canonical references/report_json_schema.md even while the body itself says 'the schema file is canonical'. This fits the anchor 'mostly efficient but includes some unnecessary explanation or could be tightened'; it is not a 4 because the duplication and changelog are more than minor trims, and not a 2 because the bulk of the body is instruction-bearing rather than padded concept explanation.

3 / 5

Actionability

The body is highly executable: exact output templates for every step, exact scoring formulas ('Final Score = (Static Score × 0.4) + (Execution Avg × 0.6)'), concrete file paths and fallback rules ('fall back to /tmp/eval_viewer_<skill_name>.md'), and a bash invocation pattern. It is not a 5 due to real gaps: the body asserts 'Step 9 always runs' but no Step 9 is defined anywhere (the promised 'polished, production-ready SKILL.md' improvement output has no procedure), and 'scripts/evaluate_skill.py' is listed as a dependency without a single instruction to run it.

4 / 5

Workflow Clarity

The 8-step pipeline is clearly sequenced with strong explicit checkpoints: two mandatory hard gates ('Any FAIL = immediate rejection. Do not proceed to Step 2'), a pre-emit checklist with cardinality constraints ('Each input's assertions array has 3–5 entries — count it'), and overwrite/fallback rules. Not a 5 because coherence is broken in places: the phantom 'Step 9 always runs' reference, the unwired evaluate_skill.py dependency, and the Input Validation scope-gate placed at the end of the document instead of before Step 1. Not a 3 because validation checkpoints are abundant and explicit throughout.

4 / 5

Progressive Disclosure

Structure is good: all 11 files in the Reference Files table exist in references/, links are one level deep and signaled at point of use ('→ Full criteria: references/basic_veto.md'), and a summary table maps each file to its step and gate status. Not a 5 because the body inlines a detailed digest of report_json_schema.md (the 'JSON top-level nodes' table and pre-emit checklist) that belongs in that reference file — content that should be separate is kept inline, matching 'good structure; most content appropriately placed; minor organization gaps'.

4 / 5

Total

15

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it answers what and when explicitly, quotes natural trigger phrases, and comprehensively enumerates concrete capabilities. Its only weaknesses are a few missing natural trigger variations and minor overlap risk from broad verbs like 'review'/'improve'.

DimensionReasoningScore

Specificity

The description enumerates many concrete actions covering the full workflow: 'two hard veto gates (structural redlines + research integrity redlines)', 'static quality scoring across 25 criteria (ISO 25010 + OpenSSF + Agent)', 'dynamic test input generation', 'multi-mode execution testing', 'human eval viewer generation', 'actionable P0/P1/P2 optimization recommendations', and 'outputs a polished, production-ready SKILL.md'. This matches the anchor 'lists multiple specific concrete actions; comprehensive coverage'; it is not a 4 because coverage spans every pipeline stage rather than having minor gaps.

5 / 5

Completeness

Both 'what' (the enumerated capabilities above) and 'when' are explicitly and concretely answered with 'Use this skill whenever a user wants to evaluate, audit, review, score, or quality-check an agent skill' plus a second explicit trigger sentence ('Also use whenever a user says...'). This matches the score-5 anchor with concrete trigger phrases; the score-4 anchor's 'when could be more explicit' does not apply.

5 / 5

Trigger Term Quality

Natural trigger phrases are explicitly quoted ('audit my skill', 'evaluate my skill', 'improve my skill') plus verbs 'evaluate, audit, review, score, or quality-check' and deployment context 'before publishing, updating, or deploying'. Good coverage matching the score-4 anchor; not a 5 because common variations like 'test my skill', 'check my skill', or 'review my SKILL.md file' are absent.

4 / 5

Distinctiveness Conflict Risk

The niche is clear — auditing agent skills across named ecosystems (Manus, OpenClaw/ClawHub, Claude, LobeHub) — and every trigger phrase is anchored on 'skill', so conflict risk is low. It is not a 5 because broad verbs like 'review' and 'improve' carry minor overlap risk with generic code-review or refactoring skills when the user's phrasing drops the skill-specific noun; this fits 'mostly distinct; minor overlap risk'.

4 / 5

Total

18

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (554 lines); consider splitting into references/ and linking

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 2 missing

Warning

Total

13

/

16

Passed

Repository
aipoch/medical-research-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.