CtrlK
BlogDocsLog inGet started
Tessl Logo

factory-review

Adversarially review one ready-for-review ticket task in a read-only Pi subagent using a model different from the worker, then return a structured markdown findings list and a verdict. Use only for the review stage; it never edits code or applies its own findings.

71

Quality

89%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

96%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplary lean skill body: it encodes session isolation as hard, testable gates, sequences the workflow unambiguously, and delegates methodology and the quality lens to appropriately split reference files. The only structural weakness is that the verdict/severity contract lives two references deep and is duplicated across two files.

DimensionReasoningScore

Conciseness

The body is ~24 lines with zero padding: 'Remain read-only. Never edit code, tests, plans, or state.' and 'Never bash, edit, or write.' assume Claude's competence and add only non-obvious protocol. Every token earns its place, matching anchor 5; there are no unnecessary explanations that would drop it to 4.

5 / 5

Actionability

Directives are exact and executable: the tool allowlist 'tools=read,grep,find,ls', the prohibition 'Never `bash`, `edit`, or `write`', the output contract 'Return only `factory.review.v4`', and the termination step 'After the template, call `subagent_done` and emit nothing else.' This fully matches anchor 5 for an instruction-only skill — guidance is specific and unambiguous, covering the common cases (model mismatch, mutation tools present, missing packet).

5 / 5

Workflow Clarity

Three clearly sequenced phases (isolate session → review the task → return the verdict), each with explicit validation checkpoints and named failure paths: 'If the model matches the work model, return `blocked`', 'If mutation tools are present, return `blocked`', 'If the packet or diff file is missing... return `blocked` and name the missing input.' Every gate has a defined recovery/report action, matching anchor 5; no validation gaps exist that would justify anchor 4.

5 / 5

Progressive Disclosure

The body is a clean overview with well-signaled, purposeful references ('Read [references/reviewer-prompt.md] and follow it', 'After behavior and acceptance criteria, apply [references/code-quality.md]'), and both referenced files exist with substantive content. However, review-contract.md is only reachable two hops away via reviewer-prompt.md and is never mentioned in the body, and verdict/severity definitions are duplicated between reviewer-prompt.md and review-contract.md — minor organization gaps that keep it at anchor 4 rather than the fully one-level-deep, cleanly split structure of anchor 5.

4 / 5

Total

19

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states concrete, distinctive capabilities in third person, includes an explicit use-when clause, and clearly separates this read-only review-stage skill from editing or worker skills. The main room for improvement is a slightly richer set of trigger phrases and a fuller enumeration of what the review produces (verdicts, severities).

Suggestions

Add one or two natural trigger variations to the when-clause, e.g. 'Use when a ticket task reaches ready-for-review or the supervisor asks for an adversarial review of the worker's implementation'.

Briefly surface the output contract in the description — e.g. mention that the verdict is approve / changes_requested / blocked — so the what-clause covers the full result surface.

DimensionReasoningScore

Specificity

Lists several concrete mechanisms — 'read-only Pi subagent', 'model different from the worker', 'return a structured markdown findings list and a verdict', 'never edits code or applies its own findings' — with minor gaps in coverage (verdict types and severity levels are not surfaced). It exceeds anchor 3 (only 1-2 actions) but is not the comprehensive coverage of anchor 5.

4 / 5

Completeness

Both parts are explicit: the 'what' ('Adversarially review one ready-for-review ticket task... return a structured markdown findings list and a verdict') and the 'when' ('Use only for the review stage'). It falls short of anchor 5 because the trigger guidance is a single clause rather than multiple concrete trigger phrases, so the 'when' could be more specific; it is well above anchor 3 since the 'when' is explicit, not implied.

4 / 5

Trigger Term Quality

Natural terms a user in this workflow would say are present: 'review', 'ready-for-review ticket task', 'review stage', 'findings', 'verdict'. Coverage is good but misses common variations like 'check the implementation' or 'verify the worker's output', placing it between anchors 3 and 4, noticeably above the midpoint.

4 / 5

Distinctiveness Conflict Risk

The description carves a clear niche with distinct triggers — 'read-only Pi subagent', 'review stage', 'model different from the worker', 'never edits code or applies its own findings' — making confusion with general code-review or editing skills very unlikely. This matches anchor 5 (clear niche, minimal conflict risk); anchor 4 would require some notable overlap with a closely related skill, which is absent.

5 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
geut/factory-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.