CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-staged-review

Use when a PR or feature needs both specification and code-quality review

54

Quality

61%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/skill-staged-review/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-sequenced, highly actionable two-stage pipeline with explicit gates, status labels, decision tables, and executable bash throughout — its workflow clarity is exemplary. Its weaknesses are structural: everything lives in one long file with no reference/script split, and a couple of sections (Bottom Line, PR posting) add tokens or depend on undefined variables like ${COMBINED_REPORT}.

Suggestions

Move the self-contained stub-detection and PR-posting bash blocks into scripts/ files (e.g. scripts/stub-detect.sh, scripts/post-review.sh) and reference them from SKILL.md, reducing the body to an overview plus orchestration guidance.

Define or populate ${COMBINED_REPORT} in the PR-posting snippet, and replace 'Wait for external reviews to complete' with a concrete wait/poll command so the guidance is fully copy-paste executable.

Cut 'The Bottom Line' and fold any non-redundant content into the intro to trim tokens without losing information.

DimensionReasoningScore

Conciseness

The body is dense with executable bash and tables and does not explain concepts Claude already knows, but has minor padding: 'The Bottom Line' restates the pipeline in three redundant lines, and the multi-LLM section includes justification prose ('A Claude-only review pipeline misses what external models catch — Codex excels at...'). This fits 'Efficient; minor instances of over-explanation that could be trimmed' rather than 5, where every token earns its place.

4 / 5

Actionability

Concrete, mostly copy-paste-ready bash is given for loading the intent contract, stub detection, dispatching external providers, and posting to a PR, plus explicit status labels and report templates. Minor gaps keep it at 4 rather than 5: ${COMBINED_REPORT} is referenced but never defined, and the external-provider step says 'Wait for external reviews to complete' without a wait/poll command.

4 / 5

Workflow Clarity

The two-stage sequence is explicit and gated: Stage 1 validates success criteria and boundaries with PASS/FAIL/PARTIAL and RESPECTED/VIOLATED statuses, a decision table routes outcomes (proceed / ask user fix-or-override), and the Error Handling table plus validation-gate frontmatter close the loop. This matches 'Clear sequence with explicit validation steps; feedback loops for error recovery' — the gate is an explicit checkpoint and failure paths route to fix-or-override.

5 / 5

Progressive Disclosure

The single SKILL.md is ~320 lines with no bundle files (references/, scripts/, assets/ are absent) and no pointers to separate material; sizable self-contained scripts (the ~55-line stub-detection block, the PR-posting block) are inlined where scripts/ files would fit. This matches 'Some structure but could be better organized; content that should be separate is inline' — the section structure itself is good, but nothing is split out, keeping it below anchor 4.

3 / 5

Total

16

/

20

Passed

Description

45%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a valid trigger sentence with a clear 'Use when' clause, but it is condition-only: it tells Claude when to activate without stating what the skill does. It also omits the skill's own natural trigger phrases ('staged review', 'full review'), limiting discoverability and distinctiveness against sibling review skills. Overall it reads as a deliberately minimal one-line description that leans below its potential.

Suggestions

State the 'what' explicitly, e.g. 'Runs a two-stage review pipeline: validates the implementation against the intent contract's success criteria and boundaries, then runs stub detection and a full code-quality review.'

Include the skill's natural trigger phrases and synonyms in the description itself ('staged review', 'full review', 'review against spec', 'spec compliance') so users' phrasing matches without relying on the frontmatter trigger field.

Add a disambiguating clause distinguishing it from a quick code review (e.g. 'use quick-code-review for code quality only') to reduce overlap risk with sibling review skills.

DimensionReasoningScore

Specificity

The description only states a condition ('Use when a PR or feature needs both specification and code-quality review') and names the review domain, but lists no concrete actions the skill performs — it never says it runs a two-stage pipeline, validates success criteria, or detects stubs. This matches the anchor 'Names the domain but actions are minimal or generic' rather than score 3, which requires 1-2 explicitly named concrete actions.

2 / 5

Completeness

The 'when' is explicit ('Use when a PR or feature needs...') and a 'what' is only weakly implied — that Claude performs specification and code-quality review — but the description never states what the skill actually does. This sits between anchor 2 (only 'when' present, no 'what') and anchor 4 (both present with the 'when' explicit), so it scores the midpoint; it is not a 4 because the 'what' must be inferred rather than stated.

3 / 5

Trigger Term Quality

'PR', 'feature', 'specification', and 'code-quality review' are relevant keywords a user might say, but common natural variations like 'PR review', 'full review', 'staged review', 'review against spec', or 'spec compliance' are absent from the description. This fits 'Some relevant keywords but missing common variations or synonyms'; it is not a 4 because several natural phrases users would actually say are missing.

3 / 5

Distinctiveness Conflict Risk

'Both specification and code-quality review' carves out a somewhat specific niche, but the phrasing 'a PR or feature needs... review' overlaps heavily with a plain code-review or PR-review skill, and the description omits the distinguishing trigger phrases ('staged review', 'two-stage review') that appear only in the frontmatter trigger field. This matches 'Somewhat specific but could still overlap with similar skills' rather than 4, which would require clearly distinct triggers.

3 / 5

Total

11

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
nyldn/claude-octopus
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.