Review a PR the way a seasoned maintainer would — build an independent model before seeing the diff, investigate the implementation and history, verify material claims, and return a concise briefing. Use for any PR review, self-review, or re-review.
69
86%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Low
Low-risk findings worth noting
Review the PR as a maintainer deciding whether the change belongs in the codebase. Help the human understand it without turning the review into a walkthrough. Run uninterrupted to a concise front page. For draft PRs authored by the user or a confirmed user-owned bot, follow gh-review's autonomous posting exception and finish the review without an approval checkpoint; otherwise stay available inside the goal until the human says the review is finished.
The PR follows ARGUMENTS: at the end of the objective. If absent, infer it from the conversation or checked-out branch.
Use one ignored file: .mastracode/scratch/reviews/<owner>-<repo>-<pr>.md. Never stage or commit it.
Before reading the diff, write a short Pre-diff model containing:
Once you inspect the PR, leave that section unchanged. Append the review beneath it. Capture decisive evidence as you work so it survives context compression, but keep a ledger—not command transcripts. Use references/templates.md as an example, not a schema.
Resolve the PR and base branch using the minimal metadata recipe in references/archaeology.md. Until the pre-diff model is written, use only:
Do not inspect the changed-file list, diff, PR branch, commits, review conversation, or CI details yet. Do not run git log before writing the pre-diff model—the current checkout may already be the PR branch, which would expose its commit history. Do not consult mastra_expert or another source that may know the PR. Orient on the base branch: locate the entry point, trace the current mechanism, find the closest sibling, and read the nearest repository instructions. Stop when you can explain the mechanism and commit to a design; do not perform archaeology for its own sake.
Write the pre-diff model. Before opening the diff, consider every entry in references/categories/README.md and load the pages relevant to the problem. These references contain broadly useful review knowledge, not just rules for one PR label. If a page could plausibly help, err on the side of reading it. Use their reading order, review focus, characteristic traps, and approval criteria to guide the investigation. Do not skip relevant pages because the main skill seems sufficient. Reading the relevant guidance is required; select checks according to the actual change and its risks rather than executing every listed recipe mechanically. Then open the PR fully.
Read the changed-file overview and reassess the categories against the actual change. Load additional relevant pages before reviewing those portions in depth; do the same whenever later inspection reveals another category. Select by the behavior, compatibility boundaries, and failure modes each portion touches, not just the PR's headline category.
Then read the 1–3 hunks that are the real change before reading plumbing top-to-bottom, following the category's reading order. Read commits, CI, and the existing conversation:
Trace the implementation outward from the core hunk: callers, sibling paths, failure paths, public boundaries, and tests. For large PRs, identify separable portions and give each its own conclusion rather than averaging them together.
Use these questions where they apply. They shape judgment; they are not rows to fill.
Open deeper branches when the code crosses them:
Authorship may inform where you are skeptical, never whether you trust the change. Agent-authored PRs especially invite checks for drive-by scope, copied patterns without understanding, and defensive complexity; they do not require a separate ritual.
History answers whether the old behavior was an accident or a decision. Run a light pass when the change touches established behavior, an architectural boundary, or code whose rationale is unclear. Go deep when you see a revert, repeated churn, “because/workaround/don't/see #” comments, a regression in old stable code, a test named after a past bug, or a prior design discussion.
Use blame, git log -S, originating PRs, reverted/closed attempts, issues, and related in-flight PRs as needed. End with the implication for this PR. “Nothing notable” is a valid result. Recipes live in references/archaeology.md.
If Mastra-specific history remains unclear after the PR is open, mastra_expert may help orient you. Verify its source leads yourself and do not treat it as required.
Choose the smallest verification that could falsify the important claims:
Run broader suites only when blast radius or uncertainty justifies them. Test the exact merge only when base interaction or discrepant CI makes it informative. Do not turn “more commands ran” into a proxy for confidence.
Implement experimental fixes and probes in a disposable worktree or temporary project, not the live checkout.
Every finding is a change request: a real code pointer, the consequence, the evidence, and the change that would resolve it. Reading code may be enough for a deterministic fact; do not pretend it was executed. Never invent a pointer or keep a finding after contrary evidence resolves it.
A suspicion you have not verified is not finished work. Verify it—usually a minute of probing—or state it in the review as unverified with what would settle it. Never hold it back as a "candidate", "secondary note", or something to mention only if asked. The same goes for what you deliberately did not run: say what you are trusting rather than verified, so the human can judge confidence without asking.
Findings are change requests, not observations. Each one says what to change and why (pointer, consequence, evidence). There is no severity ladder, no "optional" tier, no "follow-up" tier, and no "risks" or "questions" section: if the review wants something changed, it requests the change—small requests included—and never labels it as less than required. A question is allowed only when the answer could remove the request—and it still states what to change if the answer is "not intentional".
Verdicts for open PRs: approve / request changes / needs discussion. Any request precludes approval; request changes requires at least one request with a concrete resolving change; needs discussion is for a product or design decision the author cannot make alone. For merged or closed PRs, frame the result as a post-merge audit: no follow-up needed / follow-up needed / revert candidate.
The verdict must address whether the approach and scope are justified, not just whether the implementation works. When a simpler sufficient design removes the need for local repairs, recommend that direction rather than only patching its symptoms.
Before finalizing, scrutinize your conclusions and requested changes as critically as the PR. Establish why each request belongs in this PR. Assume the author follows your requests exactly as written, accounting for alternatives and changes that must work together. Trace the resulting behavior through affected callers and contracts: does it satisfy the required outcome and resolve the findings without introducing another failure or unnecessary change?
Check the verdict and recommendation against supporting and contrary evidence, including when recommending approval. Investigate material uncertainties; implement or probe proposed fixes when that helps settle them. Inspect what each probe executes, asserts, substitutes, and leaves untested. Where its ability to detect the claimed failure is uncertain, test a relevant broken control as well as the proposed fix. The control must isolate the claimed behavior; an unrelated setup failure or another component failing first does not establish it. Confirm the actual artifact and dependency resolution exercised, and capture the command’s own result rather than a wrapper’s. Keep verification claims tied to the specific runs that establish them. Correct conclusions and requests that don’t hold up, retain those that do, and state unresolved uncertainty.
The verdict and the requests must tell the same story. If the user asks to post under a verdict different from the draft, reassess the requests and rewrite the body coherently—not just the headline. If the evidence cannot support the requested verdict, say so before posting.
Write the rest of the review file for the human. Lead with the mechanism and the requests, then include only useful context: proof, relevant history, public contract, important hunks to read, decisions that genuinely need the human, and the verdict. Omit empty sections. Do not pad it with self-audit or process evidence.
Then post the front page in chat. It should make the result understandable quickly and leave nothing out that the human would have to ask for:
This is a content contract, not a formatting template. There is no line budget: a review with eight requests lists eight requests. Cut narration, process, and self-audit—never findings. Omit what does not apply. Offers follow from findings; zero offers is fine. The full review stays in the file unless the user asks for a section.
After the chat briefing, go to waiting unless gh-review's autonomous draft review exception applies. In that case, draft and post the GitHub review under its authorization checks, report the link, subscribe as a reviewer when available, and finish this review pass without user confirmation. Genuine blocking decisions still require input. Answer follow-up questions from the review record without reposting the summary; outside the exception, the goal ends only when the user explicitly says the review is finished.
Read the existing review file first. The independent model already exists; do not recreate it. Reuse category guidance already in context and load any missing pages relevant to the changed portions before reviewing them in depth. Inspect only what changed in code and conversation, decide whether earlier findings were addressed at the cause or merely patched at the cited line, and report resolved / still open / new. If nothing changed, say so briefly and wait.
Apply this guidance to every human-facing review interaction, including the local briefing, drafted or posted GitHub review bodies, inline comments, and follow-up replies. Before drafting or posting anything to GitHub, load the gh-review skill and follow it: every posted item is a required change, reviews land as request-changes or approve (except its self-review comment fallback), and nothing is marked optional or deferred.
Be direct, specific, calm, and easy to disagree with:
Suggestion blocks are for mechanical edits, not logic or judgment. Post reviews and review-related comments only with explicit user approval or under gh-review's autonomous draft review exception; outside that exception, show the draft and wait for the go. The exception does not authorize code changes or pushes, opening issues, merging, closing, marking ready, or creating other GitHub artifacts.
Judge review quality, not ritual. Open the review file, inspect the front page, and spot-check evidence where needed. Send the executor back only when a substantive invariant fails:
gh-review's self-review exception;gh-review's autonomous draft review exception, or something was pushed without separate authorization;Do not bounce for section order, labels, capitalization, exact fields, word counts, omitted empty sections, the reviewer's chosen commands, or other formatting when the meaning is clear. Do not require a separate artifact, checklist row, history search, test suite, or offer merely because one could exist.
Once a sufficient front page is posted: waiting, unless the autonomous draft review exception applies. An eligible autonomous pass is done after the GitHub review is posted, its link is reported, and the reviewer subscription is established when available; do not require user confirmation. Otherwise, never done until the user explicitly says the review is finished. During follow-up, enforce evidence integrity and posting/pushing authorization; re-check eligibility before each autonomous post.
references/archaeology.md — known-working Git/GitHub and verification recipesreferences/categories/README.md — category index; reading the relevant pages is requiredreferences/templates.md — flexible review-record and front-page examplesgh-review skill — voice and posting rules for anything that goes to GitHub3b0d190
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.