CtrlK
BlogDocsLog inGet started
Tessl Logo

correctness-reviewer

Always-on code-review persona. Reviews code for logic errors, edge cases, state management bugs, error propagation failures, and intent-vs-implementation mismatches.

63

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/correctness-reviewer/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, high-signal persona definition with concrete bug taxonomy, calibrated confidence rules, and explicit exclusions. The one real defect is the dangling reference to a 'findings schema' that is neither inlined nor shipped as a bundle file.

Suggestions

Inline the findings schema fields (e.g. what each entry in 'findings', 'residual_risks', 'testing_gaps' should contain) or ship it as a bundled reference file.

Add a brief explicit sequence for performing a review (read diff -> trace paths -> calibrate confidence -> emit JSON) to anchor the workflow.

State where the skill should look for the diff or code under review (e.g. current working tree, PR diff) to remove ambiguity at invocation.

DimensionReasoningScore

Conciseness

Lean and efficient: every section delivers concrete failure patterns or decision rules ('Only flag missing checks when the null/undefined can actually occur') with zero padding and no explanation of concepts Claude already knows.

5 / 5

Actionability

Concrete, executable guidance including numeric confidence bands ('high (0.80+)', 'moderate (0.60-0.79)') and an exact JSON output block, but 'Return your findings as JSON matching the findings schema' references a schema that is never defined or bundled.

4 / 5

Workflow Clarity

Scope rules and the confidence-calibration checkpoint are clear for a single-purpose skill, but there is no explicit review sequence and the output step depends on an undefined external schema, leaving a minor gap.

4 / 5

Progressive Disclosure

Well-organized sections ('What you're hunting for', 'Confidence calibration', 'What you don't flag', 'Output format') appropriate for a single-file skill, but the body points to an external 'findings schema' that does not exist in the bundle.

4 / 5

Total

17

/

20

Passed

Description

71%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, well-scoped description that clearly names its capabilities, held back by the absence of any 'Use when...' trigger guidance and some missing natural synonyms. Adding an explicit trigger clause would lift it to the top band.

Suggestions

Add a 'Use when...' clause, e.g. 'Use when reviewing diffs or PRs for correctness bugs before merge.'

Include natural user phrasings like 'find bugs', 'review my diff', or 'pull request review' as trigger terms.

Sharpen distinctiveness by naming the boundary with sibling reviewers (e.g. 'correctness only; not performance or style').

DimensionReasoningScore

Specificity

Lists five concrete, specific review capabilities ('logic errors, edge cases, state management bugs, error propagation failures, and intent-vs-implementation mismatches'), which is comprehensive coverage for the code-review domain with no generic filler.

5 / 5

Completeness

Has a clear 'what' (five named bug categories) but no 'when' clause; 'Always-on' only weakly implies trigger context, so per the judging guideline completeness is capped at 3.

3 / 5

Trigger Term Quality

Good natural keywords ('code-review', 'logic errors', 'edge cases', 'state management') but misses common user phrasings like 'find bugs', 'review my diff', or 'pull request', so coverage is not comprehensive.

4 / 5

Distinctiveness Conflict Risk

The enumerated bug taxonomy ('error propagation failures', 'intent-vs-implementation mismatches') makes it mostly distinct, but the generic 'code-review persona' framing overlaps with sibling reviewer skills such as a performance reviewer.

4 / 5

Total

16

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
udecode/plate
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.