CtrlK
BlogDocsLog inGet started
Tessl Logo

evaluate-pr

Evaluate a PR the harness produced — walk the change, run the system together, and build a firm understanding before you merge it. Use after the harness hands a converged PR to you for review (the "Ready for your review" comment), or any time you want to deeply review an agent-authored PR. The human-attentive skill at the back of the chain; the mirror of /intent. Outcomes — merge, close, or fix-it-yourself-and-push - no handing work back to the loop.

67

Quality

84%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-engineered process skill: explicit step sequence with real validation gates and destructive-operation safeguards, copy-ready git/gh commands, and clean two-file progressive disclosure. Its one real weakness is token efficiency — the E7 memory-path essay and the Hard-nevers section restate the philosophy section at noticeable length.

Suggestions

Compress E7 to its actionable rule (never write memory autonomously; capture human-recognized patterns with them and commit on the feature branch) and cut the /learn-P7 ground-truth mechanics, which belong in the /learn skill, not here.

Replace the full-sentence Hard nevers that merely restate E3/E7/E8/E9 with one-line pointers (e.g. "Never rehash bot findings (E3)"), keeping the list as a compact checklist.

Add a concrete command for finding the handoff PR in Step 0 (e.g. a gh pr list/--search invocation for the "Ready for your review" comment) so every step is executable as written.

DimensionReasoningScore

Conciseness

Nothing explains concepts Claude already knows, but E7 ("Memory written here is the human's call, and it's authoritative...") sprawls across two long paragraphs of /learn mechanics, and "Hard nevers" restates E3, E7, E8, and E9 nearly verbatim with cross-references. The duplication is more than the minor trimming anchor 4 allows; fits anchor 3's "could be tightened".

3 / 5

Actionability

Concrete, executable guidance throughout: "git checkout --detach origin/feature/<f>", "gh pr diff <pr>", "git push origin HEAD:feature/<f>", plus exact file paths ("prds/<f>/prd.md", "run-prd-test.sh"). Minor gaps — "find the open PR carrying the harness's 'Ready for your review' handoff comment" and the merge/close mechanics have no command (delegated to verdict.md). Anchor 4.

4 / 5

Workflow Clarity

Steps 0–8 are clearly sequenced with explicit validation checkpoints: clean-tree preflight, the Step 6 gate ("Do you feel you understand this change?"), explicit human go-ahead required for the destructive merge/close, a fix-push-review feedback loop ("once it's clean again you merge"), idempotent re-run handling, and a Hard-nevers checklist. Matches anchor 5.

5 / 5

Progressive Disclosure

The body is an overview-plus-guided-flow that splits depth into two real, one-level-deep references ("references/walkthrough.md", "references/verdict.md"), each clearly signaled with a content description and hackable-seam annotation; the reference files themselves nest no further. Matches anchor 5's well-signaled, easy-navigation structure.

5 / 5

Total

17

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that explicitly states both capabilities and trigger conditions with a concrete handoff signal (the "Ready for your review" comment). Its main weaknesses are chain-ecosystem positioning that displaces capability wording and slightly thin synonym coverage for PR-review language.

DimensionReasoningScore

Specificity

Lists several concrete actions — "walk the change, run the system together", "merge, close, or fix-it-yourself-and-push" — but dilutes them with positioning language like "the human-attentive skill at the back of the chain; the mirror of /intent" that names no capability. Not anchor 5's comprehensive action coverage; clearly above anchor 3's 1-2 actions.

4 / 5

Completeness

Explicitly answers what ("Evaluate a PR the harness produced — walk the change, run the system together, and build a firm understanding before you merge it") and when with a concrete trigger phrase ("Use after the harness hands a converged PR to you for review (the 'Ready for your review' comment), or any time you want to deeply review an agent-authored PR"). Matches the anchor-5 example structure.

5 / 5

Trigger Term Quality

Natural triggers are present — "deeply review an agent-authored PR", "Ready for your review" comment, "evaluate a PR" — but common variations like "pull request", "review the diff", or "code review" are missing. Good coverage per anchor 4, below anchor 5's comprehensive synonym coverage.

4 / 5

Distinctiveness Conflict Risk

The harness-ecosystem framing ("mirror of /intent", "the human-attentive skill at the back of the chain") carves a clear niche, but "any time you want to deeply review an agent-authored PR" overlaps a generic PR-review skill. Mostly distinct with minor overlap risk — anchor 4.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
tdg-ninja/context-specs-factory-ai
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.