CtrlK
BlogDocsLog inGet started
Tessl Logo

review-spec

Deep adversarial specification review using the 8-lens review constitution. Evaluates design documents, plans, and specs for ambiguity, completeness, feasibility, security, and MockServer-specific concerns. Loaded by review-cheap and review-final agents.

61

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.opencode/skills/review-spec/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-sequenced, highly actionable adversarial review procedure with explicit validation checkpoints, a finding format, and a strict verdict rule. Its main weakness is progressive disclosure: a large inlined principle-ID index duplicates the constitution file it points to, and the referenced files (e.g., review-constitution.md) live outside the skill bundle. Trimming the lens index to priorities and keeping detail in the constitution would tighten it further.

Suggestions

Reduce the 'Lens Priority for Spec Reviews' section to the priority ordering and rationale, leaving full principle-ID enumeration to the constitution file it already directs the reader to load.

Move the lens-by-lens ID checklists into a reference file in the skill bundle (e.g., references/lens-index.md) so SKILL.md stays an overview with one-level-deep references.

Ship the referenced files (review-constitution.md, documentation-style.md, spec-template.md) inside the skill bundle or note that they are repo-external, so the skill is self-contained when distributed.

DimensionReasoningScore

Conciseness

The body is dense, directive prose with no concept explanations Claude already knows ('The spec is wrong until proven right', 'This prevents anchoring bias' are brief purposeful rationales). Not 5 because the ~60-line lens-priority ID index partially duplicates the constitution that Step 1 already requires reading in full, and could be trimmed to priorities only.

4 / 5

Actionability

Guidance is copy-paste ready for an instruction skill: an exact finding-format block ('[PRINCIPLE-ID] Severity: ... Location: ... Evidence: ...'), an exact output-structure markdown template, a binary verdict rule (PASS/BLOCK with explicit anti-hedging), and a concrete sampling rule ('minimum 3 or 20%, whichever is larger'). Everything needed to execute is specified; not 4 because there are no material gaps.

5 / 5

Workflow Clarity

Eight clearly sequenced steps with explicit validation checkpoints: an anchoring-bias countermeasure (build an independent model before reading the spec), codebase verification of claims with a failure escalation rule ('If ANY verification fails, flag ALL unverified claims as suspect'), a pre-verdict completeness checklist, and a binary verdict. This matches the feedback-loop/checklist anchor exactly; not 4 because checkpoints are explicit at every stage.

5 / 5

Progressive Disclosure

References are clearly signaled (read '.opencode/rules/review-constitution.md' in full; '.opencode/rules/documentation-style.md'; '.opencode/skills/ideate/spec-template.md'), but the bulk of the body is an inlined index of principle IDs (AMB-01..CPX-07) that duplicates detail belonging in the constitution file. Not 4 because roughly a third of the body is inlined material that belongs in the referenced file; not 2 because structure and navigation are otherwise sound and references are explicit, not buried.

3 / 5

Total

17

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description answers 'what' concretely and in proper third person, but lacks any 'when to use' trigger guidance, which caps its usefulness for invocation. It is specific and reasonably distinctive within a spec-review niche. Adding a 'Use when reviewing specs, design docs, or plans...' clause would resolve the main gap.

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user asks to review a spec, design document, or plan, or before implementing a proposed design.'

Include common user phrasings/synonyms such as 'design review', 'spec review', or 'review this plan' to improve trigger term coverage.

Drop internal plumbing details ('Loaded by review-cheap and review-final agents') in favor of user-facing invocation guidance.

DimensionReasoningScore

Specificity

Phrases like 'Deep adversarial specification review', 'Evaluates design documents, plans, and specs for ambiguity, completeness, feasibility, security, and MockServer-specific concerns' name the domain plus several concrete evaluation dimensions. Not 5 because the actions are evaluation criteria rather than a comprehensive list of concrete operations; not 3 because coverage goes beyond 1-2 actions.

4 / 5

Completeness

The 'what' is clear (adversarial spec review across named quality dimensions), but there is no 'Use when...' or equivalent trigger clause; 'Loaded by review-cheap and review-final agents' describes internal orchestration, not when Claude should invoke it. Per the guideline, a missing explicit trigger guidance caps completeness at 3; it is not 2 because the 'what' is fully explicit.

3 / 5

Trigger Term Quality

Relevant keywords exist ('specification review', 'design documents', 'plans', 'specs', 'ambiguity', 'security'), which a user might plausibly say, but common variations/synonyms ('review my spec', 'check this design doc', 'RFC') are missing. Not 4 because natural phrasing coverage is thin; not 2 because several genuinely natural domain terms are present.

3 / 5

Distinctiveness Conflict Risk

The niche is fairly distinct ('adversarial specification review', '8-lens review constitution', 'MockServer-specific concerns'), though the generic word 'review' and evaluation dimensions like 'security' create minor overlap risk with general code-review or security skills. Not 5 due to that overlap risk; not 3 because the framing is clearly specialized to spec/design documents.

4 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mock-server/mockserver-monorepo
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.