CtrlK
BlogDocsLog inGet started
Tessl Logo

tidb-test-diff-triage

Investigate unexpected TiDB plan or test-result diffs unexplained by the change, including merge and environment effects.

74

1.16x
Quality

83%

Does it follow best practices?

Impact

50%

1.16x

2 of 3 eval scenarios. Add 1 more for a full score.

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A lean, well-structured triage workflow with concrete commands and explicit validation gates that prevent premature testdata updates. The only notable gap is actionability detail: the bisect example lacks a 'git bisect run' command and the failpoint check depends on an external doc not bundled with the skill.

DimensionReasoningScore

Conciseness

Lean and efficient throughout: every line carries non-obvious domain knowledge (e.g., '-tags=intest,deadlock does not enable failpoints', 'add -count=1 ... for reproducibility') with no padding or explanation of concepts Claude already knows. Matches anchor 5; not 4 because nothing needs trimming.

5 / 5

Actionability

Concrete, executable commands are present ('-run <TestName> -count=1', 'git bisect start/bad/good'), but the bisect block omits 'git bisect run' with the test command, and the Rule 1 check defers to 'docs/agents/testing-flow.md', which is not part of this bundle. Anchor 4 ('mostly executable, minor gaps'); not 5 because those gaps prevent full copy-paste readiness.

4 / 5

Workflow Clarity

Clear sequenced rules with explicit validation gates and feedback loops: 'If the diff disappears after failpoint enable, classify it as an environment/setup issue', 'Identify first bad commit before updating expected outputs', 'Do not record/update testdata before root cause is identified', plus a gated condition list in Rule 3. Matches anchor 5.

5 / 5

Progressive Disclosure

The skill is under 50 lines, needs no bundle files, and is organized into well-labeled sections (Trigger, Rules 1-3, Output format) with a single clearly signaled one-level external reference. Per the simple-skill guideline this matches anchor 5.

5 / 5

Total

19

/

20

Passed

Description

65%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A distinctive, domain-specific description with good natural keywords, but it states only what the skill does without any explicit when-to-use trigger clause, which caps completeness. Adding a 'Use when...' sentence covering natural trigger phrases (e.g., test diff after merge/rebase, flaky full-suite runs) would raise both completeness and trigger coverage.

Suggestions

Append an explicit trigger clause, e.g. 'Use when a plan or test diff is unexplained by the PR's touched code, when single vs full-suite runs disagree, or when planner/executor testdata changes unexpectedly after a merge/rebase.'

Name additional natural trigger synonyms such as 'flaky test', 'test failure', 'unexpected testdata change', and 'rebase' so users' phrasing matches the description.

Consider naming the concrete actions the investigation covers (failpoint setup check, bisecting the merged range, syncing stale expectations) to move specificity from one verb to a fuller capability list.

DimensionReasoningScore

Specificity

Names the domain and one concrete action ('Investigate unexpected TiDB plan or test-result diffs') with qualified scope, but lists only a single verb rather than several specific actions. Anchor 3 fits best; not 4 because coverage of distinct actions is not comprehensive.

3 / 5

Completeness

The 'what' is clear ('Investigate unexpected TiDB plan or test-result diffs unexplained by the change, including merge and environment effects'), but there is no 'Use when...' clause or equivalent explicit trigger guidance anywhere in the description. Per the judging guidelines a missing explicit trigger clause caps completeness at 3; it is not 2 because the 'what' is concrete rather than vague.

3 / 5

Trigger Term Quality

'TiDB plan', 'test-result diffs', 'unexplained by the change', 'merge', 'environment effects' are the natural phrases a user would say for this niche. Falls between anchors 3 and 4: core coverage is good but common variations like 'flaky test', 'test failure', 'rebase', or 'expected output changed' are missing.

4 / 5

Distinctiveness Conflict Risk

'TiDB plan or test-result diffs unexplained by the change' carves out a clear niche with distinct triggers and minimal conflict risk with other skills.

5 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
pingcap/tidb
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.