CtrlK
BlogDocsLog inGet started
Tessl Logo

tidb-failpoint-test-runner

Use when running TiDB package tests and deciding whether failpoint enable/disable is required before and after the test command.

82

0.96x
Quality

79%

Does it follow best practices?

Impact

96%

0.96x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/tidb-failpoint-test-runner/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A genuinely lean, well-structured instruction-only skill: it carries only non-obvious TiDB-specific knowledge and delegates canonical detail to a single named document. Its weaknesses are that no executable command appears in the skill itself (everything actionable lives in an external, non-bundled file), there is no outcome-validation step, and the external dependency is unverifiable from within the skill.

Suggestions

Inline the two actual command templates (failpoint-enabled run and plain unit-test run) or copy them into a bundled references/ file so the skill is executable without the external docs/agents/testing-flow.md.

Add a post-run validation step (e.g., check test output for pass/fail and re-check failpoint state is restored) to strengthen the workflow's feedback loop.

State the failpoint decision criteria (or at least a summary of the checks) in the skill so the decision in step 1 does not depend entirely on an out-of-bundle document.

DimensionReasoningScore

Conciseness

The ~20-line body is lean: every line is a non-obvious fact or an instruction ("`-tags=intest,deadlock` does not enable failpoints" is repo-specific knowledge Claude would not have). There is no padding, no explanation of concepts Claude already knows, and no verbosity. Matches the "lean and efficient; every token earns its place" anchor.

5 / 5

Actionability

Some concrete guidance exists — a named decision doc with section anchors, the "`-tags=intest,deadlock`" fact, and the "`-run <TestName>`" targeting flag — but the actual executable commands (the failpoint-enabled run and the plain unit-test run) are entirely delegated to an external file, so nothing in the skill itself is copy-paste executable. Not a 4 because the core execution detail is missing rather than a minor gap; not a 2 because the pointer is specific (named sections and command-set labels) rather than a high-level hint.

3 / 5

Workflow Clarity

The four numbered steps form a clear sequence — decide via the testing-flow doc, run the matching command set, keep runs targeted, record evidence — and step 4 provides a lightweight checkpoint (reporting decision evidence and the exact command). Not a 5 because there is no validation of test outcomes or error-recovery loop; not a 3 because the sequence is explicit and coherent with the evidence-recording checkpoint present.

4 / 5

Progressive Disclosure

Structure is good: Overview and Workflow sections, with detail deferred one level deep via clearly signaled pointers to "docs/agents/testing-flow.md" (named with specific section anchors). Not a 5 because the referenced file is external to the skill bundle (no references/ directory exists), the same doc path is repeated three times, and the dependency makes the skill unusable if that file moves; not a 3 because what structure exists is clean, minimal, and clearly signaled.

4 / 5

Total

16

/

20

Passed

Description

73%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A focused, third-person, appropriately concise description with an explicit "Use when..." trigger and domain-specific vocabulary. Its main limitations are a thin statement of what the skill actually does (only the decision step, not the command execution it governs) and missing common trigger variations like "unit tests" or "go test".

Suggestions

Expand the "what" clause to summarize the full capability, e.g. "Runs the correct failpoint-enabled or plain test command for TiDB packages..." rather than only describing the enable/disable decision.

Add natural trigger variations such as "unit tests", "go test", or "pkg/..." to broaden keyword coverage.

Optionally mention the outcome (recording decision evidence and the exact command) so the description reflects the complete workflow.

DimensionReasoningScore

Specificity

"Use when running TiDB package tests and deciding whether failpoint enable/disable is required before and after the test command" names the domain (TiDB package tests, failpoint enable/disable) and two concrete actions (running package tests, deciding failpoint enablement around the test command), but coverage is not comprehensive — the actual command execution variants, targeting, and evidence recording from the body are absent. Not a 4 because it lists only two actions rather than several; not a 2 because the actions named are concrete rather than generic.

3 / 5

Completeness

Both parts are present: the "when" is explicit ("Use when running TiDB package tests...") and the "what" states the decision the skill governs ("deciding whether failpoint enable/disable is required before and after the test command"). Not a 5 because the "what" leans toward restating the trigger rather than summarizing the full capability (e.g., selecting and running the correct failpoint-enabled or plain test command); not a 3 because the "when" clause is explicit, not weakly implied.

4 / 5

Trigger Term Quality

Natural terms a user would say are present — "TiDB", "package tests", "failpoint enable/disable", "test command" — which are exactly the phrases someone working in this repo would use. Not a 5 because common variations like "unit tests", "go test", or package-path patterns ("pkg/...") are missing from the description; not a 3 because the core vocabulary of the niche is well covered rather than partially covered.

4 / 5

Distinctiveness Conflict Risk

The description occupies a clear niche — TiDB failpoint-aware test execution — with distinctive triggers ("failpoint enable/disable", "TiDB package tests") that no generic testing or build skill would claim. Conflict risk is minimal.

5 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
pingcap/tidb
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.