CtrlK
BlogDocsLog inGet started
Tessl Logo

code-review

Performs a comprehensive, multi-step code review of pull requests or local code changes, using iterative refinement (generation, critique, synthesis) to ensure high-quality, actionable feedback. Use when you need to review code changes thoroughly.

60

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./.agents/agents/reidbaker-agent/skills/code-review/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, mostly actionable review workflow with good reference discipline and a genuine self-critique feedback loop. Its main costs are persona/intro padding and an orphaned 8KB reference (reviewing_tests.md) that nothing routes to, plus missing example output and error-recovery handling.

Suggestions

Delete the persona paragraph ('You are an expert Senior Software Engineer...') and fold the intro's pitfall list into Core Principles to remove redundant tokens.

Reference reviewing_tests.md from Step 2 or Step 3 (e.g. 'For test files, use the guidelines in [reviewing_tests.md](references/reviewing_tests.md)') so the bundled content is discoverable.

Add one example of a well-formed review comment (File/Line/Severity/Body/Suggestion) and a brief fallback for edge cases like an empty diff or a diff too large to review in one pass.

DimensionReasoningScore

Conciseness

The workflow steps are tight, but the persona paragraph ('You are an expert Senior Software Engineer... You are meticulous, collaborative') and the intro's restatement of the Core Principles ('avoiding common pitfalls of AI-generated reviews') are padding Claude does not need. This matches 'mostly efficient but includes some unnecessary explanation'.

3 / 5

Actionability

Provides concrete commands (`gh pr view`, `gh pr diff`, `git diff --staged`, `git log -p`), an explicit output schema with enumerated severity levels, and a checkable anchoring rule ('comments are only on lines that begin with + or -'). Falls short of fully executable: no example review comment, and 'code suggestions are compilable' is asserted without any mechanism to verify it.

4 / 5

Workflow Clarity

Five clearly sequenced steps with a built-in feedback loop (Step 4 critiques Step 3's output before synthesis) and dedup/severity prioritization at the end. Not a 5: no error-recovery checkpoints such as what to do when the diff is empty or when a comment fails the filtering rules.

4 / 5

Progressive Disclosure

The body is a lean overview and review_criteria.md, critique_rules.md, and splitting_reviews.md are real, one level deep, and signaled with markdown links at the point of need; scripts/split_diff.py is correctly second-level via splitting_reviews.md. The gap: references/reviewing_tests.md (8KB) is never referenced from SKILL.md or any other file, leaving it undiscoverable despite Step 2 listing 'test files corresponding to changed files' as review context.

4 / 5

Total

15

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A solid third-person description with an explicit what and when and mostly natural trigger terms. Its main weaknesses are buzzword padding and a when-clause that lacks concrete trigger phrases or output specifics.

Suggestions

Add concrete output specifics to the what-clause, e.g. 'produces severity-rated findings and prioritized recommendations' instead of the vague 'high-quality, actionable feedback'.

Expand the when-clause with concrete triggers users would actually say, e.g. 'Use when asked to review a PR, a diff, staged changes, or a branch before merging'.

Trim buzzwords like 'comprehensive' and 'ensure high-quality' — they add no trigger value and dilute distinctiveness from general review skills.

DimensionReasoningScore

Specificity

Names concrete actions ('multi-step code review of pull requests or local code changes', 'iterative refinement (generation, critique, synthesis)') but describes the method rather than outputs — no mention of severity-rated findings or recommendations, so it falls short of the comprehensive-coverage anchor.

4 / 5

Completeness

Both what ('Performs a comprehensive, multi-step code review...') and when ('Use when you need to review code changes thoroughly') are present, but the when-clause is generic — 'thoroughly' adds no trigger specificity, so it does not reach the explicit concrete-trigger anchor.

4 / 5

Trigger Term Quality

'code review', 'pull requests', and 'local code changes' are natural user phrases, but common variations like 'review my diff', 'review staged changes', or 'PR review' are missing, matching the good-but-incomplete keyword anchor.

4 / 5

Distinctiveness Conflict Risk

'code review of pull requests or local code changes' carves a clear niche, but buzzwords ('comprehensive', 'high-quality, actionable') and a broad trigger clause create minor overlap risk with general review-related skills.

4 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 2 suspicious

Warning

Total

15

/

16

Passed

Repository
flutter/agent-plugins
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.