github.com/mthines/agent-skills
| Skill | Added | Review |
|---|---|---|
pr-review skills/quality/pr-review/SKILL.md One-shot read-only review of a GitHub PR — dispatches the `pr-reviewer` agent and reports its verdict and findings without touching your code. The short entry point for "review this PR" when you do not want an apply-and-converge loop. Also writes maintainer relevance rules via `/pr-review remember <fact>`. Invoke with /pr-review. | 69 69 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 39b3f44 | |
resolve-conflicts skills/delivery/resolve-conflicts/SKILL.md Analyze and resolve Git merge/rebase conflicts intelligently, showing diffs and asking clarifying questions when needed. Invoke with /resolve-conflicts. | 64 64 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 39b3f44 | |
review-branch skills/quality/review-branch/SKILL.md Converges a LOCAL branch with no pull request through a bounded review-apply-simplify loop. Dispatches the branch-reviewer agent, which runs the same impact graph, finders, verifier, and confidence/severity gates as pr-reviewer, but carries findings in a local findings.jsonl instead of GitHub review threads — so a run makes zero GitHub API calls, posts nothing, and works on an un-pushed branch or offline. Stops when every finding is applied or honestly declined AND the repo's own fast checks are green; a finding it can neither fix nor honestly decline stays flagged and visible, never quietly resolved. Use it to iterate before opening a PR, or when the comment round trip on your own PR is pure overhead. Once a PR exists use /review-loop instead — threads are the right bus when other people read them. Invoke with /review-branch [--base <ref>] [--cap N] [--effort high] [--no-simplify] [--no-checks] [--report] [--include-untracked]. | 69 69 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 39b3f44 | |
review-loop skills/quality/review-loop/SKILL.md Bounded review-apply-resolve convergence loop for a GitHub PR, drafts included. Runs up to N=5 iterations of pr-reviewer → implement-suggestion (--resolve-all) → polish simplify, converging until every review thread is resolved through a fix OR a reply, so the PR ends with zero open threads and only genuine human-judgment flags left open. Convergence also means CI is not red: each push is check-read and a red mechanical failure delegated to ci-auto-fix (--no-ci). On convergence it refreshes the PR description (--no-refresh) and, on a UI PR, runs ui-verify against the live preview once, report-only (--no-preview-run). When the last review still stands (unmoved head, open unreplied threads), iteration 1 applies before re-reviewing. --merge squash-merges on a clean convergence, an approving verdict, and green CI. Run it at the TOP LEVEL of a session holding a sub-agent dispatch tool, never nested in a sub-agent. Use after opening a draft PR to converge it before undrafting. Triggers on "/review-loop". | 61 61 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 39b3f44 | |
rum-tracking skills/analysis/rum-tracking/SKILL.md Guides product analytics and RUM (Real User Monitoring) event tracking in web (React/Next.js) and mobile (React Native/Expo) apps. Decides what user interactions are valuable to capture, what's noise, what's PII to avoid, and how to implement, audit, update, and remove tracking code cleanly. Covers event naming, property schemas, tracking plans, GDPR/CCPA/DPDPA compliance, OpenTelemetry semantic conventions for browser and mobile RUM, and platforms (PostHog, Segment, Mixpanel, Amplitude, Datadog RUM, Sentry, OTel, Dash0). Modes: guide (default), implement, audit, remove, plan. Triggers on "track this event", "add analytics", "what should I track", "is this PII", "tracking plan", "remove tracking", "audit analytics", "/rum-tracking". | 67 67 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 39b3f44 | |
screen-recorder skills/analysis/screen-recorder/SKILL.md Records short videos of specific page sections using Playwright's `recordVideo` API, plays scripted interactions (hover, click, focus, scroll, keypress), crops the output to a target element via `ffmpeg`, and saves a `.webm` (or `.mp4` / `.gif`) artifact to `.agent/recordings/`. Use when a still screenshot cannot prove the change — verifying a View Transition, a Motion `layout` morph, a hover stagger, a scroll-driven timeline, an `@starting-style` entry, or any multi-frame interaction. Called by the `animations` skill to validate a generated animation, by the `ux` skill to capture an interaction the `pr-reviewer` cannot read from code, and by the `pr-reviewer` agent to produce a local clip it can inspect on motion-heavy diffs. Triggers on "record this interaction", "capture this animation", "video of this section", "validate the transition visually", "screen recording", "/screen-recorder". | 72 72 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 39b3f44 | |
severity skills/quality/severity/SKILL.md Rates how severe a finding or bug is if it is real — the blast-radius axis, complementary to confidence's is-it-real axis. Emits a lowercase tier (critical / high / medium / low) from a fixed axis rubric, then applies one deterministic, executable path floor (auth / billing / migration / infra / secrets paths) plus heuristic escalators (data-loss, security, concurrency shapes) — both scoped to reachable production code, never test / fixture / generated paths. Callers own how they gate on the tier and map it to their own blocking rules; the skill stays policy-free. Use to triage a review finding or a bug. Triggers on "how severe", "rate severity", "severity check", "triage this finding", "how bad is this", "/severity". | 73 73 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 39b3f44 | |
storybook skills/design/storybook/SKILL.md Scaffolds, audits, and tests Storybook stories for React (web) and React Native / Expo (native) component libraries. Generates three artefacts in two files per invocation: a visual regression `*.stories.tsx` file holding a `Default` story (variants grouped into one snapshot) and a `Playground` story (interactive `args` / `argTypes`), plus a sibling `*.test.stories.tsx` interaction test under a `/Tests` namespace. Supports an opt-in, per-pathname auth flow whose credentials live in the OS keychain, never in the repo. The iteration loop drives the Playwright CLI against the running Storybook URL; visual evidence delegates to the `pr-reviewer` agent and the `screen-recorder` skill. An opt-in `--validate` phase drives Playwright adversarially against the rendered story, then fixes what it breaks behind a confidence gate. Triggers on "scaffold stories", "add storybook", "story for this component", "interaction test for this story", "validate this story", "/storybook". | 72 72 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 39b3f44 | |
tdd skills/quality/tdd/SKILL.md Enforces strict Test-Driven Development with RED-GREEN-REFACTOR cycles. Writes one failing test at a time, implements minimal code to pass, then refactors. Delegates to the `test-provenance-guard` skill during REFACTOR to detect tests-by-construction (static + mutation checks). Pairs with the `code-quality` skill: invokes `Skill('code-quality')` during the REFACTOR phase to apply the full code-quality rule set against the GREEN output, and cites refactor recipes (R1–R20) by ID when reporting changes. Triggers on: "tdd", "write tests", "test this", "add test coverage", "test driven", "red green refactor", "/tdd". | 71 71 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 39b3f44 | |
test-provenance-guard skills/quality/test-provenance-guard/SKILL.md Detects tests that pass by construction — tests that define a private copy of the function under test instead of importing the production module — and self-heals by extracting the inline logic to an exported function, updating production callers, and rewriting the test to import the export. Two checks: (1) static — the test file must import the SUT and must not shadow its exported names; (2) mutation — blanking the production function body re-runs the test and expects failure. Runs autonomously inside autonomous-workflow Phase 4 and as a slash command for human-driven PR review. Use when adding new tests for existing or refactored code, when CI is green but you are unsure whether the tests actually exercise production, or when reviewing a PR for tests-by-construction. Triggers on "test provenance", "tests by construction", "verify tests cover real code", "tests duplicate logic", "mutation sanity check", "are these tests fake", "/test-provenance-guard". | 65 65 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 39b3f44 | |
ui-verify skills/testing/ui-verify/SKILL.md Makes a UI pull request autonomously verifiable. `author` writes a Markdown intent spec — the happy path plus out-of-bounds cases from a PR-scoped brainstorm — into the PR description as a collapsed, machine-findable block (delegated to by `create-pr` on UI diffs). `run` resolves the PR's live preview URL via the GitHub deployments API and runs the spec there through `aw-tester`, which may adapt the route but checks every expected outcome with evidence, and reports a verdict plus full-page screenshots (`--no-screenshots` opts out). It then tries to break the same change — hostile input, double submits, failed requests, small viewports — and documents every probe with screenshots, never changing the verdict (`--no-adversarial` opts out). `verify` is author-if-needed, then run. `setup` delegates to `aw-setup --target preview`. A LoreKit loop feeds runner friction into the next spec. Web only. Triggers on "write a preview spec", "verify this PR's preview", "try to break this PR's preview", "/ui-verify". | 66 66 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 39b3f44 | |
ux skills/design/ux/SKILL.md Reviews UX, accessibility, and microcopy for web and React Native (Expo) applications. Analyzes UI code against established UX principles, WCAG 2.2 accessibility standards, platform guidelines (Apple HIG, Material Design 3), and UX writing best practices. Triggers on: "ux review", "review ux", "check accessibility", "improve the UI", "ux audit", "review this component", "is this accessible", "check usability", "ux feedback", "review the design", "improve usability", "check contrast", "review navigation", "ux writing", "improve copy", "review microcopy", "make this more intuitive", "ux best practices", "/ux". | 68 68 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 39b3f44 | |
verify-behavior skills/quality/verify-behavior/SKILL.md Owns a cheapest-first three-tier verification ladder — Tier 1 syntactic (grep / ast-grep / read), Tier 2 semantic-no-execution (typecheck / build / lint), Tier 3 execution (run the covering test, or a minimal synthesized repro) — and reports the result as an evidence receipt (confirms / contradicts / ambiguous / null). It never scores; `confidence(code)` owns the number. Two consumer shapes: claim-verification (read-only, feeds `confidence(code)`) and change-verification (post-apply green/red gate). Called by `verification-receipt.md` (pr-reviewer Tier 2/3), `bug-fix-verifier`, `feature-pr-verifier`, and the `aw-executor` Phase 4 checks loop. Use when a finding or a change needs executed proof, not just a plausible-sounding claim. Triggers on "verify this claim", "does this actually happen at runtime", "prove this behavior", "run this to confirm", "/verify-behavior". | 64 64 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 39b3f44 | |
video-analyser skills/analysis/video-analyser/SKILL.md Analyse a video file — primarily a screen recording of a bug — to extract errors, UI state, and reproduction steps. Resolves input from a Linear ticket URL, a local file path, or a direct video URL. Extracts keyframes with ffmpeg, runs optional Tesseract OCR and Whisper audio transcription, then delivers structured findings. Trigger phrases: "analyse this video", "analyze this recording", "what does this video show", "extract bugs from this recording", "analyse this screen recording", "investigate this mp4", "investigate this mov", "analyse this clip", "look at this screen capture", "what is happening in this video", "analyse this screen capture", "video-analyser", "/video-analyser". | 73 73 Impact — No eval scenarios have been run Securityby Low Low-risk findings worth noting Version: 39b3f44 | |
visual-design skills/design/visual-design/SKILL.md Guides and reviews the visual design and brand identity of UI components for web and React Native — color systems, typography pairing, visual hierarchy, signature details, and named style directions (minimal, swiss, editorial, brutalist, neo-brutalist, glass, soft-UI, terminal, playful, retro). Owns the generative, brand-aware side; defers WCAG contrast math, size minimums, and dark-mode mechanics back to /ux. Modes: `guide` (default — build a component from scratch), `review` (audit existing visuals against direction), `direction` (propose a style direction for a new product or feature). Triggers on "visual design", "make this look good", "brand identity", "style direction", "improve the visuals", "review the look", "does this look generic", "/visual-design". | 69 69 Impact — No eval scenarios have been run Securityby Passed No findings from the security scan Version: 39b3f44 |