CtrlK
BlogDocsLog inGet started
Tessl Logo

repro-flaky-tests

Reproduce and investigate flakiness in a test.

56

Quality

64%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/repro-flaky-tests/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a well-structured, actionable workflow with concrete commands, validation checkpoints, and feedback loops; its main weaknesses are minor redundancy, unfilled placeholders, and prose-style rather than checklist-style sequencing.

Suggestions

Tighten redundancy between the 'Next steps' and 'Results' sections (e.g. state the local-vs-bot verification preference once) and fix typos like 'Iff' and 'initially attempt'.

Give '<count>' a concrete default or range so the repeat command is more copy-paste ready.

Consider converting the fix-test and verification steps into a short numbered checklist with explicit retry/abort conditions for tighter workflow clarity.

DimensionReasoningScore

Conciseness

The body is mostly lean and task-focused with concrete commands and examples and little concept padding; minor redundancy (local-vs-bot preference restated across 'Next steps' and 'Results') and a couple of typos ('Iff', 'initially attempt') keep it just short of fully lean.

4 / 5

Actionability

It provides concrete, near-copy-paste commands with specific bots, flags, branch names, and table formats, but unfilled placeholders like '<count>' and required substitutions prevent a fully copy-paste-ready 5.

4 / 5

Workflow Clarity

There is a clear sub-agent sequence with explicit validation checkpoints and a feedback loop ('repeat the verification cycle', 'at least one test ran to completion'), though the steps are prose-style rather than a tight numbered checklist with full error-recovery guidance.

4 / 5

Progressive Disclosure

The document is a single well-sectioned file with no nested or buried references and no bundle files to misorganize; it stays at one level deep, just short of 5 because it exceeds ~50 lines and a couple of sections could be trimmed.

4 / 5

Total

16

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific to a clear niche but is terse: it states what the skill does without an explicit 'Use when...' trigger clause or richer synonyms, which limits trigger quality and completeness.

Suggestions

Add an explicit 'Use when...' clause, e.g. 'Use when a test is flaky, intermittently failing, or the user reports an intermittent test failure.'

Expand trigger terms with natural synonyms such as 'flaky test', 'intermittent test failure', and 'deflake'.

Optionally list a couple more concrete actions (e.g. run local repeats, drive stressor bots, apply and verify a fix) to lift specificity toward 4-5.

DimensionReasoningScore

Specificity

The phrase 'Reproduce and investigate flakiness in a test' names the domain and two concrete actions, matching the anchor for naming a domain plus 1-2 actions without comprehensive coverage.

3 / 5

Completeness

It gives a clear 'what' but no 'Use when...' trigger clause, which per the guidelines caps completeness at 3 since 'when' is only weakly implied.

3 / 5

Trigger Term Quality

'Reproduce', 'investigate', 'flakiness', and 'test' are relevant natural terms, but common synonyms like 'flaky test', 'intermittent failure', or 'failing test' are missing, so coverage is partial rather than good.

3 / 5

Distinctiveness Conflict Risk

The flaky-test reproduction niche is mostly distinct with only minor overlap risk against related testing skills, though the brief trigger set keeps it just short of a clear-no-conflict 5.

4 / 5

Total

13

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
ChromeDevTools/devtools-frontend
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.