CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-code-review

Expert multi-AI code review with inline PR comments — use for thorough quality and security analysis

53

Quality

59%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./skills/skill-code-review/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

60%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly concrete with well-sequenced workflows and thoughtful failure handling, but it is bloated by compliance rhetoric and inlines large bash blocks that belong in script files. Its cross-references all point to files absent from the bundle, which undermines navigation and progressive disclosure.

Suggestions

Move the ~90-line stub-detection loop into a scripts/stub-detection.sh file and reference it with a one-line invocation, pulling the detailed patterns into references/stub-detection.md.

Fix or remove dangling references (skills/blocks/codex-host-adapter.md, agents/personas/code-reviewer.md, .claude/references/stub-detection.md) — either ship these files in the bundle or describe their content inline.

Trim the MANDATORY COMPLIANCE rhetoric, the capabilities bullet list, and the version-pinned heading down to essential routing guidance to reduce token cost.

DimensionReasoningScore

Conciseness

The bulk of the body is executable bash that earns its place, but sections like "MANDATORY COMPLIANCE — DO NOT SKIP" (rhetorical prohibitions), the "CLAUDE OCTOPUS ACTIVATED" mascot line, the Capabilities bullet list, and the version-pinned "(v8.44.0)" heading are padding Claude does not need. This matches 'mostly efficient but includes some unnecessary explanation or could be tightened'; not 2 because the padding is confined to a few sections rather than pervading the document.

3 / 5

Actionability

Concrete, mostly copy-paste-ready bash is provided for every phase (orchestrate.sh invocations, stub-detection greps, gh PR detection and posting with a credential-gated helper), matching 'mostly executable guidance; concrete code or commands with minor gaps'. Not 5 because the code depends on external scripts outside the bundle (${HOME}/.claude-octopus/...) that may not exist, REVIEW_SYNTHESIS is referenced but never constructed, and some greps (e.g. "\s" in grep -E) will not behave as intended.

4 / 5

Workflow Clarity

Sequences are explicit and numbered (stub detection Steps 1-3, PR posting Steps 1-2, quick vs full mode selection) with validation checkpoints and fallbacks: gh CLI availability check, no-PR branch, failed-post recovery path, and ask-first gating for standalone invocation. Matches 'clear sequence with most checkpoints present'; not 5 because the top-level pipeline order (grasp → tangle → synthesis → completeness → PR post) is scattered across sections and error-recovery for a missing orchestrate.sh is not addressed.

4 / 5

Progressive Disclosure

Roughly 90 lines of stub-detection bash are inlined in SKILL.md where a scripts/ helper clearly belongs, and every cross-file reference points outside the bundle — skills/blocks/codex-host-adapter.md, agents/personas/code-reviewer.md, .claude/references/stub-detection.md, and plugin scripts — none of which exist in the skill's references/, scripts/, or assets/ directories. This matches 'content that clearly belongs in separate files is inlined; or references are buried'. Not 3 because the references are not merely unclearly signaled — they are dangling, and no bundle structure exists to offload the inline bulk.

2 / 5

Total

13

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description communicates a clear niche (multi-provider review with inline PR posting) and includes both a what and a when clause, but the trigger guidance is generic and lacks the natural phrases and variations users would actually say. It sits slightly above the midpoint of the rubric.

Suggestions

Add concrete trigger phrases users would naturally say, e.g. "Use when the user asks to review a PR, audit changes for security issues, or run a pre-commit quality check".

Include common synonyms and variations (pull request, code audit, security review, staged changes) to improve trigger term coverage.

Name one or two more concrete capabilities (e.g. OWASP vulnerability detection, TDD evidence review) to sharpen the 'what' clause.

DimensionReasoningScore

Specificity

"Expert multi-AI code review with inline PR comments" names the domain and 1-2 concrete actions (multi-AI review, inline PR comments), matching the 'names domain and 1-2 concrete actions' anchor. It does not reach 4 because the action list is thin — no mention of specific analysis types like vulnerability detection or test review in the description itself.

3 / 5

Completeness

The 'what' is clear ("multi-AI code review with inline PR comments") and a trigger clause exists ("use for thorough quality and security analysis"), so both what and when are present per the anchor-4 example. Not 5 because the 'when' is generic — no concrete trigger phrases like "use when the user asks to review a PR or staged changes"; not 3 because an explicit 'use for' clause is present.

4 / 5

Trigger Term Quality

Terms "code review", "PR", "quality", and "security analysis" are relevant keywords, but common natural phrases users would say ("review my changes", "pull request", "audit this PR", "pre-commit check") are missing. Not 4 because several natural variations are absent; not 2 because the present keywords are genuinely on-domain rather than generic.

3 / 5

Distinctiveness Conflict Risk

"Code review" is a crowded namespace with high overlap risk against built-in and common review skills, but the "multi-AI" and "inline PR comments" qualifiers give it some distinct identity — matching the 'somewhat specific but could still overlap' anchor. Not 4 because the differentiators are jargon-y rather than distinct trigger terms.

3 / 5

Total

13

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
nyldn/claude-octopus
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.