CtrlK
BlogDocsLog inGet started
Tessl Logo

the-judge

Evidence-first pull request judge that reviews a PR and posts one consolidated GitHub review with inline comments via the gh CLI. Runs the repo's own deterministic checks first, researches current official docs before any claim about external libraries or APIs, then reviews correctness, security, structural quality (code judo, spaghetti growth, file-size limits), and AI slop including useless code comments. Every finding must carry evidence, every comment passes a deterministic noise gate before posting, and the verdict (APPROVE, COMMENT, REQUEST_CHANGES) is weighed by findings. Use when asked to review a PR, judge this PR, review this branch or diff before merge, run the-judge, "revise esse PR", "faca o code review", or "julgue esse PR". Do NOT use for reviewing prose or documents, fixing CI failures, resolving merge conflicts, writing the fix itself, or responding to review comments (use gh-address-comments).

77

Quality

96%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong operational skill body: fully executable commands and schemas, a clearly sequenced workflow with deterministic validation gates and error-recovery loops, and correct use of one-level-deep references. The only weakness is minor redundancy and rhetorical padding that could be tightened without losing meaning.

DimensionReasoningScore

Conciseness

The body never explains concepts Claude already knows and mostly earns its tokens with rules, commands, and tables. Minor over-explanation remains: a few rules stated twice (the gate appears in Non-Negotiable #5, Step 6, and the Convergence Contract; the noise budget repeats in the severity table) and rhetorical flourishes like 'Training data is a rumor; the changelog is a source' and 'arguing with it is arguing with a regex' could be trimmed. This is efficient with minor trimmable instances, not the lean every-token-earns-its-place level.

4 / 5

Actionability

Guidance is fully executable: exact gh CLI commands ('gh pr view --json number,title,body,author,url...', 'gh api user --jq .login'), concrete script invocations ('python3 scripts/review_gate.py findings.json'), a complete findings.json example with schema fields explained, and worked examples covering routine review, blockers, unevidenced claims, and round-2 re-reviews. Nothing is pseudocode.

5 / 5

Workflow Clarity

Steps 0-7 are clearly sequenced with explicit validation checkpoints and feedback loops: Step 4 verification kills un-evidenced candidates, Step 6 mandates 'Fix every reported violation and re-run until exit code 0', the deterministic ladder runs before any judgment, and a Troubleshooting section maps errors to solutions (422, gh auth, no PR, gate loops). The outward-facing post step is gated behind deterministic validation.

5 / 5

Progressive Disclosure

Structure follows progressive disclosure well: both referenced files exist and are one level deep, clearly signaled at exactly the point of use ('Read `references/review-standards.md` now', 'Read `references/comment-voice.md` now'), and the three scripts are referenced by path and present in scripts/. Detailed standards and comment voice are correctly externalized while the body keeps the overview, workflow, and contract.

5 / 5

Total

19

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: comprehensive concrete capabilities, natural trigger phrases in two languages, an explicit Use-when clause, and a Do-NOT-use boundary that sharply reduces conflict risk with sibling skills. No penalties apply under the rubric guidelines.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions: 'reviews a PR and posts one consolidated GitHub review with inline comments via the gh CLI', 'Runs the repo's own deterministic checks first, researches current official docs', and enumerates review dimensions (correctness, security, structural quality, AI slop) plus the verdict mechanism. Coverage is comprehensive with no vague filler. Third-person voice is used throughout ('reviews', 'runs', 'researches').

5 / 5

Completeness

It explicitly answers what the skill does (detailed multi-sentence capability list) and when to use it ('Use when asked to review a PR, judge this PR, review this branch or diff before merge...') with concrete trigger phrases, and adds a 'Do NOT use for...' exclusion clause. Both what and when are explicit and complete.

5 / 5

Trigger Term Quality

Natural trigger phrases users would actually say are covered including synonyms and other languages: 'review a PR', 'judge this PR', 'review this branch or diff before merge', 'run the-judge', plus 'revise esse PR', 'faca o code review', 'julgue esse PR'. Both English and Portuguese variants of the same intent are present.

5 / 5

Distinctiveness Conflict Risk

The niche is clear (evidence-first PR review via gh CLI) and the explicit negative scope ('Do NOT use for reviewing prose or documents, fixing CI failures, resolving merge conflicts, writing the fix itself, or responding to review comments (use gh-address-comments)') minimizes overlap with adjacent review/fix skills by naming the boundary.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
tech-leads-club/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.