CtrlK
BlogDocsLog inGet started
Tessl Logo

test-triage

Triage failing macOS tests across Xcode and SwiftPM workflows. Use when asked to run macOS tests, narrow failing scopes, explain assertion or crash failures, or separate real test regressions from setup and environment problems.

71

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is test-triage in openai/plugins

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, token-efficient triage workflow with a concrete failure taxonomy and sensible guardrails against overstating flakes as regressions. The main gap is actionability detail: harness detection and focused-rerun steps describe what to do without giving runnable command forms, and there is no explicit fix-verification step.

Suggestions

Add runnable example commands for the rerun and scope-narrowing steps, e.g. `xcodebuild test -only-testing:TestTargetTests/TestClass/testMethod` and `swift test --filter TestClass.testMethod`.

Make harness detection concrete: detect the harness by the presence of `Package.swift` (SwiftPM) vs. `.xcodeproj`/`.xcworkspace` (Xcode) instead of only naming the two commands.

Close the workflow loop with an explicit verification checkpoint after a fix, e.g. rerun the originally failing test and confirm it passes before reporting the regression as resolved.

DimensionReasoningScore

Conciseness

The body is lean and assumes Claude's competence: no concept explanations, no tool tutorials, just directives like "Use `xcodebuild test` for Xcode-based projects" and "Distinguish compilation failures from test execution failures". Every section earns its tokens, matching the top anchor rather than the 4 anchor ("minor instances of over-explanation").

5 / 5

Actionability

Concrete commands are named (`xcodebuild test`, `swift test`) and the six-category failure taxonomy plus output checklist are specific, but the rerun and scope-narrowing steps lack executable specifics (no example flags like `-only-testing:` or `--filter`, no example commands). This sits between the 3 anchor (missing key details) and the 5 anchor (copy-paste ready commands) but noticeably above the midpoint.

4 / 5

Workflow Clarity

A clear five-step sequence (detect harness → narrow scope → classify → rerun → summarize) with a built-in decision loop ("Use focused reruns when a specific case fails... without new information") and an output checklist. It misses a 5 because there is no explicit verify-the-fix checkpoint; it exceeds 3 because the sequence and checkpoints are explicit rather than implicit, and no destructive/batch validation cap applies.

4 / 5

Progressive Disclosure

The skill is under 50 lines, needs no external references (none exist in the bundle), and is organized into four clearly labeled sections (Quick Start, Workflow, Guardrails, Output Expectations) with nothing inlined that belongs in a separate file — matching the simple-skill exception for a top score.

5 / 5

Total

18

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly states a specific niche, enumerates concrete triage actions, and provides an explicit multi-clause "Use when" trigger. It is concise with no padding or over-claims; the only room for improvement is adding a few more trigger synonyms (e.g., flaky tests, CI failures).

DimensionReasoningScore

Specificity

The description lists several specific actions — "Triage failing macOS tests across Xcode and SwiftPM workflows", "run macOS tests, narrow failing scopes, explain assertion or crash failures, or separate real test regressions from setup and environment problems" — with only minor coverage gaps (e.g., rerun strategy, log parsing). It falls below 5, which demands comprehensive multi-action coverage, and above 3, which expects only 1-2 concrete actions.

4 / 5

Completeness

It explicitly answers what ("Triage failing macOS tests across Xcode and SwiftPM workflows") and when ("Use when asked to run macOS tests, narrow failing scopes, explain assertion or crash failures, or separate real test regressions from setup and environment problems") with multiple concrete trigger phrases, matching the top anchor exactly.

5 / 5

Trigger Term Quality

Natural phrases users would say are present: "run macOS tests", "narrow failing scopes", "explain assertion or crash failures", "Xcode", "SwiftPM", "test regressions", "setup and environment problems". A few common variations are missing ("flaky tests", "unit tests", "CI failures"), placing it between the 3 and 5 anchors but noticeably above the midpoint.

4 / 5

Distinctiveness Conflict Risk

The niche is clearly delimited — macOS test triage across two named toolchains (Xcode and SwiftPM) — and the trigger phrases are specific to test-failure diagnosis, giving minimal conflict risk with other skills. It fits the 5 anchor better than the 4 anchor, which presumes overlap with closely related skills.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
robinebers/openusage
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.