CtrlK
BlogDocsLog inGet started
Tessl Logo

test-triage

Triage macOS tests across Xcode and SwiftPM. Use when narrowing failures, explaining assertions or crashes, or separating setup from regressions.

68

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, well-structured instruction-only skill: concrete commands, an explicit failure taxonomy, and a coherent five-step workflow with decision rules. The main gaps are the absence of copy-paste filter examples for narrowing test scope, a redundant Quick Start that restates the description, and no explicit flake-confirmation checkpoint.

Suggestions

Add concrete filter invocations to the 'Narrow the scope' step, e.g. `xcodebuild test -only-testing:TestTarget/TestClass/testMethod` and `swift test --filter TestClass/testMethod`, so the narrowing instruction is copy-paste ready.

Trim the Quick Start section, which restates the frontmatter description, or replace it with a one-line pointer into the Workflow steps.

Add an explicit flake-confirmation checkpoint in step 4, e.g. 'rerun an isolated failure twice before marking it as a flake,' to close the validation gap before classifying async timing issues.

DimensionReasoningScore

Conciseness

Lean and efficient with no concept explanations and no padding; every section instructs rather than describes. Not 5 because the Quick Start sentence ('run the smallest meaningful test scope first, classify failures precisely, and avoid treating every test failure like a product bug') substantially restates the frontmatter description and could be trimmed.

4 / 5

Actionability

Names concrete commands ('xcodebuild test', 'swift test'), a six-category failure taxonomy, and decision rules for scope narrowing and reruns. Not 5 because the narrowing step lacks copy-paste-ready filter invocations (e.g., '-only-testing:TestSuite/testMethod', 'swift test --filter'); not 3 because the guidance given is executable and specific, not pseudocode.

4 / 5

Workflow Clarity

Clear five-step sequence with decision points (use user filter vs. smallest likely failing target; focused vs. full-suite rerun) and a classify-then-rerun feedback loop. Not 5 because the 'mark likely flakes as such' guardrail has no explicit confirmation checkpoint (e.g., rerun twice before calling a flake); not 3 because checkpoints are mostly present and the sequence is coherent.

4 / 5

Progressive Disclosure

Under 50 lines, single-purpose, no external references needed, and content is organized into clear headed sections (Workflow, Guardrails, Output Expectations) — per the rubric's simple-skill note this earns full marks with no bundle files required.

5 / 5

Total

17

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that pairs a clear, domain-specific 'what' with an explicit, multi-trigger 'when' clause. Keyword coverage is good though missing a few natural synonyms (e.g., 'flaky', 'swift test', 'XCTest'), and a couple of the skill's actual capabilities (focused reruns, output expectations) are not surfaced.

DimensionReasoningScore

Specificity

Names the domain and several concrete actions ('Triage macOS tests', 'narrowing failures', 'explaining assertions or crashes', 'separating setup from regressions'). Not 5 because coverage has minor gaps — rerun strategy and harness detection from the body are not surfaced; not 3 because it lists well more than 1-2 concrete actions.

4 / 5

Completeness

Clearly answers both: what ('Triage macOS tests across Xcode and SwiftPM') and when ('Use when narrowing failures, explaining assertions or crashes, or separating setup from regressions') with concrete trigger phrases. Not 4 because the 'when' clause is already explicit and specific, matching the anchor-5 example's structure.

5 / 5

Trigger Term Quality

Good keyword coverage with natural user phrases: 'macOS tests', 'Xcode', 'SwiftPM', 'failures', 'assertions', 'crashes', 'regressions'. Not 5 because common variations like 'swift test', 'XCTest', 'unit tests', or 'flaky' are missing; not 3 because coverage goes well beyond 'some relevant keywords'.

4 / 5

Distinctiveness Conflict Risk

Clear niche (macOS test triage across both Xcode and SwiftPM) with distinct triggers ('assertions', 'crashes', 'regressions'); minimal conflict risk with other skills. Third-person imperative voice, no person shifts.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
openai/plugins
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.