CtrlK
BlogDocsLog inGet started
Tessl Logo

deflake

Finds flaky tests on the main branch by analyzing GitHub Actions failures, ranks them by frequency, and enters parallel plan mode to design deflake strategies. Use when you want to find and fix the flakiest tests.

66

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable workflow skill with excellent sequencing and user-approval checkpoints. The one notable gap is that the referenced collection script is not part of the bundle, leaving Phase 1 non-executable as shipped.

Suggestions

Ship `collect-flakes.py` in a `scripts/` directory within the skill bundle and reference it with a skill-relative path so Phase 1 is executable as distributed.

Trim the parenthetical examples in the root-cause checklist (e.g., "hardcoded sleeps, tight timeouts") to save tokens where the category name alone suffices.

DimensionReasoningScore

Conciseness

The body is procedural and assumes competence (e.g., "If CI log formats change over time, update the script directly") without tutoring Claude on known concepts. Minor trimmable spots exist, such as the parenthetical glosses in the root-cause checklist ("hardcoded sleeps, tight timeouts"), placing it at anchor 4 rather than the fully lean anchor 5.

4 / 5

Actionability

Provides a concrete command (`python3 .claude/skills/deflake/collect-flakes.py`), a six-step agent task list, and copy-shaped report/plan templates — mostly executable. It stops short of anchor 5 because the single command depends on a script that is not present in the skill bundle, so the core collection step cannot be run as shipped here.

4 / 5

Workflow Clarity

Three phases are clearly sequenced with explicit checkpoints and feedback loops: a hard stop if `--report` was passed, "Wait for all agents to complete, then consolidate findings", "Present all plans and wait for user feedback", and "Do NOT enter plan mode or start implementing until the user approves". This matches anchor 5's explicit validation steps and error-recovery gating, and exceeds anchor 4's minor-gaps profile.

5 / 5

Progressive Disclosure

Sections are well organized (Arguments, three phases, Deflake Principles) with one-level-deep structure and no nested references. Scored against the actual bundle — no references/, scripts/, or assets/ directories exist, yet the body references `collect-flakes.py` at a non-standard path, a minor organization gap that keeps it below anchor 5.

4 / 5

Total

17

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete third-person actions, an explicit "Use when" trigger, and a well-delineated niche. The main improvement would be broadening trigger coverage with synonyms (intermittent failures, CI flakiness) and hinting at the report-only mode.

DimensionReasoningScore

Specificity

"Finds flaky tests on the main branch by analyzing GitHub Actions failures, ranks them by frequency, and enters parallel plan mode to design deflake strategies" lists three concrete, mechanized actions in third person, but coverage has minor gaps (no mention of the report-only mode or flake/bug/infra categorization). It is above anchor 3 (only 1-2 actions) but does not match anchor 5's comprehensive coverage.

4 / 5

Completeness

Explicitly answers both "what" (finds, ranks, plans fixes for flaky tests via GitHub Actions analysis) and "when" ("Use when you want to find and fix the flakiest tests"), so above anchor 3 where "when" is only implied. It falls short of anchor 5 because the trigger clause is a single condition rather than multiple concrete trigger phrases covering related scenarios.

4 / 5

Trigger Term Quality

"flaky tests", "flakiest tests", and "GitHub Actions failures" are natural terms a user would say, but common synonyms like "intermittent test failures", "CI flakiness", or "test reliability" are missing. Good coverage short of the comprehensive anchor-5 example (which includes synonyms and extensions).

4 / 5

Distinctiveness Conflict Risk

"flaky tests", "GitHub Actions failures", and "deflake" carve out a clear niche with distinct triggers; minimal overlap risk with general testing or CI skills. Matches anchor 5 clearly and is more specific than the anchor-4 example's minor overlap.

5 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
stacklok/toolhive
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.