CtrlK
BlogDocsLog inGet started
Tessl Logo

triaging-visual-review-runs

Inspects PostHog Visual Review (VR) runs that gate PR merges with screenshot regression checks. Use when the user mentions "visual review", "VR", "snapshot diff", "screenshot test", "storybook regression", "playwright snapshot", asks why a PR is blocked or what changed visually, wants to triage the VR backlog, decide whether a snapshot diff is real vs flaky, or check whether a story has been changing across runs. Also invoke when a PR has a failing `visual-review` status check, when a PR comment mentions "Visual review", or when the user is on a branch with an open VR run.

76

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a high-quality, dense, and highly actionable skill document with strong workflows and an explicit safety gate around the one irreversible action. Its only weak spots are minor intro padding and a monolithic structure that could split the tool catalog and vocabulary into reference files.

Suggestions

Tighten the opening two paragraphs: drop the meta framing ('This skill teaches an agent how to answer the questions a human reviewer would actually ask') and lead directly with what the skill does and its tool categories.

Consider moving the full tool tables and 'Vocabulary cheat sheet' into a references/ file (e.g. TOOLS.md), keeping SKILL.md as an overview that links to it — this would raise progressive disclosure by giving the body one-level-deep, clearly signaled references.

The finalize-gate content is excellent; mirror that explicit 'verify -> present -> wait -> confirm' pattern more briefly in the triage and status workflows so the validation discipline is consistent across all workflows.

DimensionReasoningScore

Conciseness

Largely efficient and assumes Claude's competence (no explanations of CI, screenshots, or libraries), but the intro framing — 'This skill teaches an agent how to answer the questions a human reviewer would actually ask' — and a few setup sentences add mild padding that could be trimmed.

4 / 5

Actionability

Provides concrete tool tables, exact call shapes with required/optional parameters, executable curl snippets for pulling PNGs, and specific workflow steps with named tool calls — copy-paste-ready guidance covering the common cases.

5 / 5

Workflow Clarity

Multiple clearly sequenced numbered workflows with explicit validation checkpoints and feedback loops, notably the 'finalize gate' pre-conditions and error-recovery paths for '409 stale_run' / '409 sha_mismatch', and the three-signal diff-verdict matrix.

5 / 5

Progressive Disclosure

Well-organized into clearly signaled sections (When this skill applies, Tools, Vocabulary cheat sheet, Workflows, Output expectations, What NOT to do) with no nested or broken references, but it is monolithic — the tool catalog and vocabulary cheat sheet are inline rather than split into one-level-deep reference files, so it stops short of the anchor for clean external reference navigation.

4 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is exemplary: it pairs a concrete statement of what the skill does with an unusually thorough set of natural-language triggers and operational status-check signals. It is specific, complete, and clearly distinct from neighboring skills.

DimensionReasoningScore

Specificity

Lists multiple specific concrete actions — 'Inspects PostHog Visual Review (VR) runs that gate PR merges with screenshot regression checks', 'triage the VR backlog', 'decide whether a snapshot diff is real vs flaky', 'check whether a story has been changing across runs' — with comprehensive coverage of the skill's purpose.

5 / 5

Completeness

Explicitly answers both 'what' (inspects VR runs gating PR merges via screenshot regression) and 'when' (extensive 'Use when...' and 'Also invoke when...' clauses with concrete trigger phrases), matching the anchor for clearly answering both.

5 / 5

Trigger Term Quality

Comprehensive natural-language triggers including 'visual review', 'VR', 'snapshot diff', 'screenshot test', 'storybook regression', 'playwright snapshot', plus operational signals like 'failing visual-review status check' and 'PR comment mentions Visual review'.

5 / 5

Distinctiveness Conflict Risk

Narrow PostHog-specific niche with distinct triggers ('visual review', 'VR', 'visual-review status check') makes conflict with other skills minimal.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
PostHog/posthog
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.