CtrlK
BlogDocsLog inGet started
Tessl Logo

fix-random-ci-test-failure

Investigate and fix flaky/random CI test failures in dotnet/macios. Trigger on GitHub issues describing intermittent test failures, CI postmortem issues, or when asked to fix a flaky test. Analyzes test code, identifies root causes (shared state, environment dependencies, race conditions), and applies fixes.

68

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-crafted, actionable skill body: concrete commands, repo-specific file references, and a clear workflow with a fallback branch for unclear root causes. Its main weaknesses are mild — repeated keychain guidance, no local run-to-verify step before the PR, and no before/after code example demonstrating an applied fix.

Suggestions

Add a validation step before section 6 (Create the PR): build and run the affected test locally (or note why it may not be runnable, e.g., simulator-only tests) to confirm the fix, creating a feedback loop for error recovery.

Include one short before/after code example of a real fix from this repo (e.g., replacing a hardcoded keychain identifier with a per-process unique one in KeyChainTest.cs style) to make the fix patterns copy-paste ready.

Consolidate the keychain/LAContext guidance that currently appears in sections 3, 4, and "Key Patterns in This Repo" — e.g., keep the symptom/fix pairs in section 3 and reduce later sections to one-line pointers.

DimensionReasoningScore

Conciseness

The body is lean and imperative with no padding or explanation of concepts Claude already knows, fitting anchor 4's "minor instances that could be trimmed". Keychain/LAContext guidance is repeated across section 3, section 4, and "Key Patterns in This Repo", which is redundant rather than anchor-5's every-token-earns-its-place.

4 / 5

Actionability

Provides concrete executable guidance — `grep -r "TestMethodName" tests/`, `Process.GetCurrentProcess ().Id`, `{bundleId}-{testType}-{pid}`, `TestContext.Out.WriteLine`, and precise PR conventions (Fixes #NNNN vs Ref #NNNN, copilot label). Falls short of anchor 5 only in lacking a copy-paste-ready before/after code example of an applied fix.

4 / 5

Workflow Clarity

A clear six-step sequence with explicit checkpoints (verifying failures span unrelated PRs to confirm flakiness) and a decision branch for unresolved causes (section 5 "If the Fix is Unclear"). The gap versus anchor 5 is a missing validation step — no instruction to run or verify the fixed test locally before creating the PR.

4 / 5

Progressive Disclosure

No bundle files exist, and this ~100-line single-file skill is well organized with clear section headers, so structure is appropriate. It is not anchor 5 because the simple-skill exception targets under-50-line skills, and the root-cause catalog / repo key patterns could reasonably live in a separate reference file.

4 / 5

Total

16

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete third-person capabilities, an explicit trigger clause, and a tightly-scoped niche in dotnet/macios. Only minor gaps — a few natural synonyms and unmentioned workflow steps like diagnostics/PR creation keep it just below perfect on specificity and trigger coverage.

DimensionReasoningScore

Specificity

Lists several concrete third-person actions ("Investigate and fix", "Analyzes test code", "identifies root causes (shared state, environment dependencies, race conditions)", "applies fixes") with enumerated root-cause categories. Not anchor 5 because parts of the workflow's coverage (diagnostics and PR creation) go unmentioned — a minor gap.

4 / 5

Completeness

Explicitly answers both what ("Analyzes test code, identifies root causes... and applies fixes") and when ("Trigger on GitHub issues describing intermittent test failures, CI postmortem issues, or when asked to fix a flaky test") with concrete trigger phrases — a direct match for the anchor-5 example pattern.

5 / 5

Trigger Term Quality

Natural phrases users would say are well covered: "flaky/random CI test failures", "intermittent test failures", "CI postmortem issues", "asked to fix a flaky test". A few common synonyms (e.g., "unstable test", "randomly failing test") are missing, so it fits anchor 4 rather than 5.

4 / 5

Distinctiveness Conflict Risk

Scoped narrowly to "dotnet/macios" flaky CI test failures with repo-specific markers (the ci-postmortem label), giving it a clear niche and minimal overlap risk with other skills.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
dotnet/macios
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.