CtrlK
BlogDocsLog inGet started
Tessl Logo

fixing-flaky-tests

Guides an agent through reproducing, root-causing, fixing, and validating flaky tests in the PostHog monorepo. Use when a test fails intermittently in CI but passes on rerun or locally, when `hogli ci:insights` or the debugging-ci-failures skill classifies a failure as a flaky test, when given a GitHub Actions URL for a flaky job, when asked to check Trunk Flaky Tests for a test, PR, or master, or when asked to deflake, stabilize, or fix a flaky Jest, pytest, or Playwright test. Core discipline: reproduce locally before changing anything, fix the root cause (never mask it with sleeps, retries, or bigger timeouts), and prove the fix with an N-run validation loop sized to the observed failure rate. Stabilizing is not the only valid outcome — the skill also gates whether the test should exist, so deleting a test that catches nothing real, or re-leveling one that flakes because of the level it runs at, are first-class endings.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

—

The risk profile of this skill

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a strong, executable runbook with a clear multi-step workflow, validation feedback loops, and careful gating of destructive outcomes. It is tight and domain-specific, with only minor room to split inlined reference material or trim a few asides.

Suggestions

Consider moving the large symptom→cause table and the per-stack PostHog-specific patterns (frontend/backend) into a reference file under references/, keeping SKILL.md as a tighter overview that links out.

Trim a few justificatory asides (e.g. restatements of why a masking move is wrong) where the table already conveys the rule, to tighten conciseness.

DimensionReasoningScore

Conciseness

The body is dense with PostHog-specific operational knowledge Claude does not already have (hogli, Trunk, kea/MSW internals) and mostly assumes competence, but a few justificatory asides and explanatory clauses could be trimmed without losing clarity.

4 / 5

Actionability

Provides copy-paste-ready, executable guidance throughout — concrete `gh run list`/`gh api` queries, a complete N-run bash loop harness, `git bisect run` invocation, and `pnpm --filter` jest commands — covering the common cases fully.

5 / 5

Workflow Clarity

Eight clearly numbered steps with an explicit escalation ladder, validation checkpoints ("Stop at the first level that reproduces it", "Any failure in the loop → back to step 4"), and a destructive-operation gate requiring explicit user approval before deletion.

5 / 5

Progressive Disclosure

Well-organized into clearly headed sections with one-level-deep, clearly signaled references to sibling skills and internal docs (no nested-reference problem), but the body inlines sizable pattern/cause tables that could arguably live in separate reference files; no bundle files are present to offload them.

4 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is exemplary: third-person, concrete, and dense with natural trigger phrases, explicitly answering both what the skill does and when to use it while staking out a distinct niche. No vague fluff or over-claims.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — "reproducing, root-causing, fixing, and validating flaky tests" plus "deleting a test that catches nothing real, or re-leveling one" — giving comprehensive coverage rather than vague verbs.

5 / 5

Completeness

Explicitly answers both — the "what" (reproduce/root-cause/fix/validate flaky tests in the PostHog monorepo) and the "when" via a concrete "Use when..." clause enumerating trigger scenarios.

5 / 5

Trigger Term Quality

Covers natural user phrases and synonyms densely: "fails intermittently in CI but passes on rerun or locally", "deflake, stabilize, or fix a flaky Jest, pytest, or Playwright test", plus tool-specific triggers (hogli ci:insights, GitHub Actions URL, Trunk Flaky Tests).

5 / 5

Distinctiveness Conflict Risk

Occupies a clear PostHog-specific flaky-test niche with distinct tool triggers, and explicitly distinguishes itself from the sibling debugging-ci-failures skill, minimizing wrong-skill activation.

5 / 5

Total

20

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 suspicious

Warning

Total

15

/

16

Passed

Repository
PostHog/posthog
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.