CtrlK
BlogDocsLog inGet started
Tessl Logo

spec-driven-eval

Scores how completely an implementation fulfills a PRD/spec, case by case, and produces a single comparable final grade. Invoke only when explicitly named (e.g. run spec-driven-eval); do not auto-trigger. Use when benchmarking spec-driven implementations, grading acceptance criteria, evaluating whether a feature was 100% implemented, comparing multiple implementations of the same PRD, or auditing implementation and test coverage (unit and e2e) against product requirements. Do NOT use for planning or building features (use tlc-spec-driven), writing PRDs, or general code review unrelated to a spec.

73

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an exceptionally rigorous, actionable methodology with a well-sequenced 14-step process, embedded validation checkpoints, and correctly-signaled one-level-deep references. Its main weakness is redundancy: core rules are restated nearly verbatim in the step rules and anti-patterns sections, and a few dense run-on paragraphs (notably Engineering Gates G) could be tightened or moved to a reference file.

Suggestions

Cut the anti-patterns section down to genuinely new failure modes (or convert it to a cross-reference table pointing back to the core/step rules) — at least 10 of its 15 bullets restate rules already stated verbatim in Core rules, Reproducibility rules, or Step rules.

Break the single-paragraph Engineering Gates G bullet into a short definition plus a small table (gate / verdict rules / evidence required), moving the probe-before-not-run and pre-existing-failure carve-out detail into that structure.

Consider moving the 10-category Elicitation rubric table into references/reference.md alongside the report template and calibration anchors, keeping only the E_recall/E_precision/E_justified formulas and a pointer in SKILL.md — it is applied once per evaluation (Step 5) rather than continuously.

DimensionReasoningScore

Conciseness

There is no filler — no explanations of concepts Claude already knows — but the same rules are restated across four sections: e.g. read-only evaluation appears in Core rule 5, Step 11, and the anti-patterns; 'compute, don't calculate' in Reproducibility rule 4, Step 13, and anti-patterns; search-before-zero in Core rule 2, Step 8, and anti-patterns. This fits 'Mostly efficient but... could be tightened' rather than 4, since the ~15-bullet anti-patterns section largely duplicates rules already stated verbatim, and the Engineering Gates paragraph is a single ~350-word run-on that could be condensed.

3 / 5

Actionability

Guidance is fully executable: exact git commands for the diff surface, exact formulas (I, T, AC_score, Story_score, Final), pinned weights tables, a mandated compute mechanism ('executed by a script (e.g. node -e / python3 -c)... with the script's output pasted into the report'), and a precise report filename format with the timestamp command (date -u +%Y%m%dT%H%M%SZ). It matches 'Fully executable; copy-paste ready code or commands' — the only unpasted artifact (the roll-up script) is deliberately parameterized per benchmark, which is a justified flexibility.

5 / 5

Workflow Clarity

The 14-step process checklist is clearly sequenced with explicit validation checkpoints and feedback loops: search-before-zero for UNMET checks, probe-before-not-run for gates, the k=3 majority-vote disagreement handling, and the calibration loop ('If verdicts disagree on more than ~20% of checks... sharpen the ambiguous checks... and re-run'). This matches the anchor 'Clear sequence with explicit validation steps; feedback loops for error recovery; checklists for complex processes'.

5 / 5

Progressive Disclosure

References are real (references/quickstart.md and references/reference.md both exist), one level deep (reference.md points to no further bundled .md files), and well-signaled with purpose statements ('Read it when you need the operational how-to; the rest of this file is the scoring methodology'; reference.md holds 'worked checklist anchors... report template... worked example'). However, the ~330-line body retains the full scoring model plus the entire 10-category Elicitation rubric inline, so it is more than an overview — 'Good structure; most content is appropriately placed... minor organization gaps' fits better than the 5 anchor's 'content appropriately split'.

4 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person voice, concrete capability list, comprehensive natural-language trigger phrases, and explicit negative guidance that separates it from sibling planning/building skills. The 'invoke only when explicitly named' clause is operationally consistent with the frontmatter's disable-model-invocation flag rather than padding.

DimensionReasoningScore

Specificity

The description lists multiple concrete, domain-specific actions — "Scores how completely an implementation fulfills a PRD/spec, case by case", "grading acceptance criteria", "auditing implementation and test coverage (unit and e2e)", "comparing multiple implementations of the same PRD" — with comprehensive coverage of the skill's capabilities. It matches the anchor 'Lists multiple specific concrete actions; comprehensive coverage' and does not fall to 4, since no meaningful capability of the skill is left unnamed.

5 / 5

Completeness

It explicitly answers both questions: the 'what' ("Scores how completely an implementation fulfills a PRD/spec, case by case, and produces a single comparable final grade") and the 'when' ("Use when benchmarking... grading... evaluating... comparing... auditing...") with concrete trigger phrases. It also adds explicit negative guidance ("Do NOT use for planning or building features... writing PRDs, or general code review unrelated to a spec"), which exceeds the score-5 anchor rather than falling short of it.

5 / 5

Trigger Term Quality

Natural user phrasings are comprehensively covered with synonyms: "benchmarking spec-driven implementations", "grading acceptance criteria", "evaluating whether a feature was 100% implemented", "comparing multiple implementations of the same PRD", "auditing implementation and test coverage (unit and e2e)", plus the PRD/spec synonym pair. It fits the anchor 'Comprehensive coverage of natural terms including synonyms'; a 4 would require identifiable missing natural terms, and none are apparent for this domain.

5 / 5

Distinctiveness Conflict Risk

It occupies a clear niche (spec-driven implementation evaluation/benchmarking) and explicitly disambiguates the nearest competing skills — "Do NOT use for planning or building features (use tlc-spec-driven), writing PRDs, or general code review unrelated to a spec" — and further constrains invocation to "explicitly named (e.g. run spec-driven-eval)". This matches 'Clear niche with distinct triggers; minimal conflict risk'.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
tech-leads-club/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.