CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-harness-audit

Audita un arnes agentico existente con score determinista de 35 checks: presencia, calidad (staleness por contenido, placeholders, evidencia vacia, ciclos de dependencias), comandos stack-verified, monorepo y capa de enforcement (hooks registrados que leen stdin, Stop con max_turns, alcance estructurado, ruta unica de escritura), con anti-gaming por lineas estructuradas. Usar al pedir 'auditar arnes', 'harness audit', 'score del arnes', 'por que mi agente dice done y falla'. NO para auditar skills del toolkit (usar toolkit-hardener) ni auditorias de accesibilidad.

65

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

65%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, command-driven skill body that stays lean and points to real reference files for depth, with a clear conditional procedure. The two weak spots are the Packet layer list advertising non-existent directories and validation being implied via exit code rather than an explicit re-validate step.

Suggestions

Trim the Packet block to the layers that actually ship (references/, scripts/, assets/) so the advertised navigation matches the bundle on disk.

Surface an explicit validation checkpoint in the procedure — e.g. 're-run audit_harness.py after a fix and confirm the failed check group is now green' — rather than relying solely on the exit-code coherence noted in the Contract.

Optionally note that checks-de-calidad.md is the canonical 35-check table so a reader knows where to resolve any check id referenced in passing.

DimensionReasoningScore

Conciseness

Lean and dense; assumes Claude's competence, explains no basic concepts, and tags provenance ([CÓDIGO], [EXPLICIT], [SUPUESTO]) instead of padding. Every line carries load (e.g. "Mide la calidad real de un arnes, no su presencia").

5 / 5

Actionability

Copy-paste-ready commands with concrete flags — "Corre ${CLAUDE_PLUGIN_ROOT}/skills/agent-harness-creator/scripts/audit_harness.py <dir> --json" and "--require H1-hooks-registered,H2-hooks-read-stdin,..." — plus explicit chaining outputs (--emit-fix) covering the common cases.

5 / 5

Workflow Clarity

A numbered procedure with clear conditional branching (fallos mecánicos → repair; drift docs-vs-realidad → diagnose) and an acceptance criterion tying exit code to the threshold. It is a read-only audit so the destructive-validation cap does not apply, but validation is delegated to the script's exit code rather than surfaced as an explicit re-check loop.

4 / 5

Progressive Disclosure

Well-structured body with inline one-level-deep references to real files (references/checks-de-calidad.md, references/anti-gaming.md) and a Packet layer map. However the Packet block advertises knowledge/, prompts/, examples/, and agents/ layers that are not present in the actual bundle (only references/, scripts/, assets/ exist), a minor navigation mismatch.

4 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: concrete capabilities, natural bilingual trigger phrases, explicit what/when, and clear boundary guidance against adjacent skills. Third-person voice is maintained throughout with no over-claims.

DimensionReasoningScore

Specificity

Lists multiple concrete actions across the full audit surface — 'presencia, calidad (staleness por contenido, placeholders, evidencia vacia, ciclos de dependencias), comandos stack-verified, monorepo y capa de enforcement (hooks registrados que leen stdin, Stop con max_turns, alcance estructurado, ruta unica de escritura)' — with comprehensive, specific coverage rather than generic verbs.

5 / 5

Completeness

Explicitly answers both 'what' (deterministic 35-check harness audit with named check groups and anti-gaming) and 'when' ("Usar al pedir ...") with concrete trigger phrases, plus negative scope guidance.

5 / 5

Trigger Term Quality

Provides natural user phrasings in both languages — "auditar arnes", "harness audit", "score del arnes", "por que mi agente dice done y falla" — covering synonyms and the exact words a user would say when needing this skill.

5 / 5

Distinctiveness Conflict Risk

Clear niche (agent-harness auditing) with explicit disambiguation — "NO para auditar skills del toolkit (usar toolkit-hardener) ni auditorias de accesibilidad" — minimizing overlap with related skills.

5 / 5

Total

20

/

20

Passed

Validation

68%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 11 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

11

/

16

Passed

Repository
JaviMontano/claude-plugins
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.