CtrlK
BlogDocsLog inGet started
Tessl Logo

test-flakiness

Find flaky tests from CI logs — aggregates pass rates, spots intermittent failures, recommends quarantine. After multiple runs.

64

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This is a tightly written, highly operational skill: concrete commands, numeric thresholds, an engine-specific cause/fix table, and unusually disciplined sourcing (verifying skip annotations against the installed gdUnit source, flagging unverifiable claims as NOT SOURCEABLE). The few weaknesses are minor: a small amount of introductory flakiness philosophy, no error-recovery loop for unparseable logs, and a single-file layout that could offload engine details to reference files.

Suggestions

Trim the opening definition of what a flaky test is and the "worse than no tests" rationale — Claude already knows this — to gain token efficiency.

Add a small error-recovery branch for logs that fail to parse (e.g., unrecognized XML schema → report the format found and ask the user) to close the workflow-clarity feedback-loop gap.

Move the engine-specific parsing and skip-mechanism details (Godot/Unity/Unreal sections) into a references/ file, keeping SKILL.md as a tighter overview.

DimensionReasoningScore

Conciseness

The body is dense and operational — thresholds, grep patterns, and an engine-specific cause table with almost no filler. The main over-explanations are the opening definition of what a flaky test is ("A flaky test is one that sometimes passes and sometimes fails...") and its "worse than no tests" rationale, plus a few motivational lines in the Collaborative Protocol — minor trimmable material, not padding.

4 / 5

Actionability

Guidance is fully executable: exact grep patterns ("<testcase name=", "Result={Success}"), numeric flakiness tiers (>25% / 5–25% / 1–5%), engine-specific skip syntax with source-file citations, verbatim user prompts, and a copy-paste report template. The honest "NOT SOURCEABLE" markers that instruct asking the user instead of inventing attributes make the guidance exceptionally safe to execute.

5 / 5

Workflow Clarity

A clear numbered sequence (parse arguments → locate data → parse results → identify → recommend → report → update suite) with real checkpoints: the under-3-runs statistical guard, the stop-and-ask branch when no logs exist, and ask-before-write gates on both file writes. It falls short of the top anchor because there is no validate-and-retry feedback loop (e.g., what to do when a log format fails to parse), though the operations are append-only and user-approved, so the destructive-operation cap does not apply.

4 / 5

Progressive Disclosure

No bundle files exist (references/, scripts/, assets/ are absent), so everything lives in a single ~240-line SKILL.md. Section structure is clean and numbered, and external details are correctly pushed to clearly signaled one-level-deep project docs (e.g., "docs/engine-reference/unreal/current-best-practices.md, 'Command Line'"). At this length, the engine-specific parsing and skip-mechanism material is a plausible candidate for a references/ split, which keeps it just below the top anchor.

4 / 5

Total

17

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description with concrete actions and genuinely natural trigger terms ("flaky", "intermittent failures", "CI logs", "quarantine"). Its main weakness is the terse "After multiple runs" trigger clause, which reads as a precondition rather than explicit "use when..." guidance, and coverage that understates the skill's full capability set.

Suggestions

Expand the when-clause into explicit trigger guidance, e.g. "Use when the user mentions flaky, intermittent, or randomly failing tests, or when CI failures keep reappearing without code changes."

Add missing natural trigger variations such as "unstable tests", "tests failing randomly", and "reruns" to broaden keyword coverage.

Mention one or two more capabilities the skill actually performs (cause diagnosis, regression-suite quarantine updates) to close the coverage gap in specificity.

DimensionReasoningScore

Specificity

The description lists several concrete actions — "aggregates pass rates, spots intermittent failures, recommends quarantine" — grounded in a clear domain ("flaky tests from CI logs"), but it omits other capabilities the skill actually performs (cause classification, report generation, regression-suite updates), so coverage has minor gaps rather than being comprehensive.

4 / 5

Completeness

The "what" is clear (find flaky tests from CI logs, aggregate pass rates, spot intermittent failures, recommend quarantine) and an explicit "when" clause exists ("After multiple runs"), so it is not merely implied. However, the when-clause is terse — it states a precondition rather than concrete trigger phrases like "use when the user mentions flaky tests or intermittent CI failures" — so it does not reach the fully explicit anchor.

4 / 5

Trigger Term Quality

Natural trigger terms users would actually say are present: "flaky tests", "CI logs", "intermittent failures", "quarantine", "pass rates". A few natural variations are missing (e.g., "unstable tests", "tests failing randomly", "reruns"), keeping it just below comprehensive synonym coverage.

4 / 5

Distinctiveness Conflict Risk

The niche is distinct: flakiness detection, pass-rate aggregation, and quarantine recommendation are unlikely to collide with unrelated skills. There is minor overlap risk with closely related testing skills (the body references a sibling `/regression-suite` skill covering the same quarantine registry), which keeps it below the clear-niche anchor.

4 / 5

Total

16

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
Donchitos/Claude-Code-Game-Studios
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.