CtrlK
BlogDocsLog inGet started
Tessl Logo

evaluate-pr

Evaluate a PR the harness produced — walk the change, run the system together, and build a firm understanding before you merge it. Use after the harness hands a converged PR to you for review (the "Ready for your review" comment), or any time you want to deeply review an agent-authored PR. The human-attentive skill at the back of the chain; the mirror of /intent. Outcomes — merge, close, or fix-it-yourself-and-push - no handing work back to the loop.

68

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

The canonical home for this skill is evaluate-pr in tdg-ninja/context-specs-factory-ai

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-engineered skill body: precise step sequencing, executable commands, real feedback loops, and exemplary use of one-level-deep reference files. The only notable weakness is redundancy — the same hard rules are restated in the philosophy, contract, and hard-nevers sections.

Suggestions

Consolidate the triple statement of the memory rule (E7, the "Invocation & output contract" bullet, and the first "Hard never"): state it once as a Hard never and have E7 and the contract reference it in a single line.

Similarly fold the repeated "never hand work back to the loop" rule (E9, Step 7 closing, and the Hard never) into a single authoritative location, keeping only one-line cross-references elsewhere.

The E7 paragraph's deep dive into /learn's P7 semantics belongs in references/verdict.md (where the merge → /learn path is already explained); replacing it with a one-line pointer would cut ~10 lines without losing the rule.

DimensionReasoningScore

Conciseness

The body carries dense, genuinely novel workflow knowledge with no padding about concepts Claude already knows, but the same invariants are restated two to three times: E7's memory rules appear in the philosophy section, the "Invocation & output contract", and "Hard nevers"; E3, E8, and E9 are similarly repeated. This fits anchor 3 ("mostly efficient but... could be tightened") — more than minor trimmable redundancy, but not the majority-padded anchor 2.

3 / 5

Actionability

Fully executable guidance throughout: exact commands ("git checkout --detach origin/feature/<f>", "git push origin HEAD:feature/<f>", "gh pr view <pr> --json reviews,comments", "git checkout main"), an ordered reading list of concrete paths (prds/<f>/prd.md, run-prd-test.sh, specs/<f>/mainspec.md), and copy-paste verdict commands in references/verdict.md. As an instruction-only skill this matches the anchor-5 standard.

5 / 5

Workflow Clarity

Steps 0–8 are clearly sequenced with explicit validation checkpoints (clean-tree preflight, the Step 6 understanding gate, human go-ahead required before merge/close) and a genuine feedback loop in the fix-then-merge path (push → reviewer re-runs → wait for HARNESS_REVIEW_CLEAN → merge). The "Hard nevers" section functions as a checklist. This matches anchor 5.

5 / 5

Progressive Disclosure

The body is a clear overview that moves bulk detail into two well-signaled, one-level-deep references (references/walkthrough.md and references/verdict.md — both present in the bundle and delivering what their intros promise), each introduced with its scope and hackable seams, with no nested reference chains or orphan pointers. Matches anchor 5.

5 / 5

Total

18

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit and concrete what/when structure and a well-defined niche. Its main weaknesses are the second-person voice (explicitly penalized by the rubric) and some space spent on ecosystem jargon rather than natural trigger synonyms.

Suggestions

Rewrite the description in third person to avoid the rubric's second-person penalty — e.g. "Evaluates a PR the harness produced: walks the change, runs the system, and builds firm understanding before merge. Use after the harness posts the 'Ready for your review' comment, or whenever an agent-authored PR needs deep human review."

Replace chain-internal jargon ("the human-attentive skill at the back of the chain", "the mirror of /intent") with natural trigger synonyms users would actually say, such as "review", "code review", "approve", or "sign off".

Trim the trailing meta-commentary ("Outcomes — merge, close, or fix-it-yourself-and-push - no handing work back to the loop") to free tokens for concrete capability verbs.

DimensionReasoningScore

Specificity

The description lists many concrete actions ("walk the change, run the system together", "merge, close, or fix-it-yourself-and-push"), which on its own matches anchor 4. However, the rubric guideline penalizes second-person voice by one point, and the description repeatedly addresses the user directly ("before you merge it", "any time you want to deeply review an agent-authored PR"), so the score drops to 3.

3 / 5

Completeness

Clearly and explicitly answers both questions: what ("Evaluate a PR the harness produced — walk the change, run the system together, and build a firm understanding before you merge it") and when ("Use after the harness hands a converged PR to you for review (the 'Ready for your review' comment), or any time you want to deeply review an agent-authored PR") with concrete trigger phrases. This matches anchor 5 exactly.

5 / 5

Trigger Term Quality

Good natural-term coverage: "Evaluate a PR", "review", "merge", "agent-authored PR", and the concrete "Ready for your review" handoff phrase. It falls short of anchor 5 because common variations like "code review" or "approve" are absent and some space is spent on internal jargon ("the mirror of /intent", "human-attentive skill at the back of the chain").

4 / 5

Distinctiveness Conflict Risk

A clear niche with distinct triggers (harness-produced PR, the "Ready for your review" comment, mirror of /intent), but the broadening clause "or any time you want to deeply review an agent-authored PR" opens minor overlap with any generic PR-review skill, fitting anchor 4 rather than the minimal-conflict anchor 5.

4 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
tdg-ninja/context-specs-claude-code
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.