CtrlK
BlogDocsLog inGet started
Tessl Logo

pipeline-test-triage

Triage test failures in an Azure DevOps pipeline run for Fluid Framework. Use when the user shares an ADO build/pipeline URL or build ID and wants failures analyzed, when they ask "why did this pipeline fail", "summarize the test failures", "is this test flaky or a real bug", "compare these runs", or wants bugs filed for genuine failures. Covers the Real Service End to End Tests pipeline as the worked example but applies to any FF test pipeline.

72

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an outstanding operational document — dense, executable, and full of verified gotchas (api-version previews, resource GUIDs, hidden retried attempts) with a clearly sequenced tiered workflow and explicit verdict criteria. Its main structural weakness is that it is a monolithic single file: REST recipes, auth setup, bug-filing templates, and pipeline-specific appendices that would work better as one-level-deep reference files are all inlined, and a couple of warnings are repeated.

Suggestions

Move the Tier-3 deep material — the REST auth token block, "Timeline → deep link recipe", the retried-attempts recipe, and the one-time "Enabling REST auth" setup — into a references/ file (e.g. references/rest-recipes.md), keeping only the endpoint list and key gotchas inline with a clearly signaled pointer.

Move the bug-filing gotchas and the Description template into references/filing-bugs.md, leaving just the create/update command essentials and the "confirm scope first" checkpoint in SKILL.md.

Deduplicate the green-build/retried-attempt warning, which appears in the taxonomy section, the Tier-3 intro, and the retry-recipe heading — state it once as a callout and cross-reference it from the other two spots.

DimensionReasoningScore

Conciseness

Nearly every token is non-obvious operational knowledge (api-version gotchas, resource GUIDs, log-size heuristics) rather than concepts Claude already knows, so it is well past the 3 anchor. Minor trim opportunities keep it off the 5 anchor: the tier-escalation rationale is stated both in the table and again in the paragraph after it ("Descend a tier only when the one above can't answer the question" vs. the table's 'Use when' column), and the green-build/retried-attempt warning appears three times ("Never conclude 'no failures' from a green status alone", "⚠️ Green builds can still hide failures", "Always check for retries before declaring a build clean").

4 / 5

Actionability

Fully executable throughout: exact REST URLs with api-version strings ("GET https://vstmr.dev.azure.com/fluidframework/internal/_apis/testresults/resultsbybuild?buildId={ID}&outcomes=Failed&api-version=7.1-preview.1"), ready MCP tool invocations ("ado-pipelines_build action=get_status buildId={ID} project=internal"), runnable PowerShell ("$tok = az account get-access-token --resource 499b84ac-1321-427f-aa17-267ca6975798 --query accessToken -o tsv"), a real helper script ("scripts/Parse-AdoTestLog.ps1 -Path <temp-file-from-get_content> [-Context 30]" — verified present in the bundle and matching its body description), and a copy-paste bug Description template. Specific examples cover the common cases.

5 / 5

Workflow Clarity

A clearly sequenced tiered workflow with an explicit descent criterion ("Descend a tier only when the one above can't answer the question"), numbered steps within each tier, explicit verdict rules ("Deterministic if the same test + same assertion/stack fails in every run... Flaky if it appears in some runs and not others"), and genuine feedback/recovery loops ("If auth is missing you'll land on a login page — pause and let the user sign in, then continue"; "If the history view is awkward to read, drop to Tier 2/3"), plus pre-declaration checkpoints ("Never conclude 'no failures' from a green status alone"; "Confirm scope with the user first" before filing bugs). Matches the 5 anchor's sequence + validation + recovery criteria; the operations are read-then-file so the destructive-cap is not triggered.

5 / 5

Progressive Disclosure

Section structure and headers are excellent, but the ~300-line body inlines substantial reference material that belongs in bundle files: the full Tier-3 REST auth block and token recipes, the "Timeline → deep link recipe", the "Finding failures hidden in retried attempts" recipe, the one-time "Enabling REST auth" setup, the bug-filing gotchas + template, and the "Reference specifics — Real Service End to End Tests" appendix all live in the always-loaded SKILL.md, with no references/ directory in the bundle (only the one script). This matches the 3 anchor ('content that should be separate is inline'); it is above 2 ('minimal structure') because organization is strong, and below 4 because one-time/setup material and deep recipes are not split out.

3 / 5

Total

17

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a model example: it states a concrete capability, gives an explicit 'Use when' clause with verbatim user-phrasing triggers, and carves out a distinct niche (ADO/Fluid Framework pipeline test triage). The only minor gap is that secondary capabilities (run comparison, bug filing) ride inside the trigger clause rather than being listed as actions, leaving specificity just short of comprehensive.

DimensionReasoningScore

Specificity

Concrete actions are named — "Triage test failures in an Azure DevOps pipeline run", "wants failures analyzed", "compare these runs", "bugs filed for genuine failures" — several specific capabilities beyond generic language. It falls short of the 5 anchor because the secondary actions (compare runs, file bugs) are carried inside the trigger clause rather than enumerated as capabilities, and the classification/verdict activity is only implied by 'triage'; it clearly exceeds the 3 anchor ('1-2 concrete actions') on coverage.

4 / 5

Completeness

Explicitly answers both questions: what ("Triage test failures in an Azure DevOps pipeline run for Fluid Framework") and when ("Use when the user shares an ADO build/pipeline URL or build ID... when they ask...") with concrete trigger phrases — matching the 5 anchor exactly. A 4 would require the 'when' to be merely present but less explicit, which is not the case here.

5 / 5

Trigger Term Quality

Quotes natural phrases users would actually say — "why did this pipeline fail", "summarize the test failures", "is this test flaky or a real bug", "compare these runs" — plus artifact-level triggers ("ADO build/pipeline URL", "build ID") and full-name/synonym coverage ("Azure DevOps", "ADO"). Comprehensive including synonyms; no common variation is missing.

5 / 5

Distinctiveness Conflict Risk

A clear niche with distinct triggers: Azure DevOps pipeline test-failure triage for Fluid Framework, even scoping the worked example ("Covers the Real Service End to End Tests pipeline"). Minimal overlap risk with generic testing or CI skills; matches the 5 anchor ('clear niche with distinct triggers').

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 2 missing, 1 deeper-than-1-level

Warning

Total

15

/

16

Passed

Repository
microsoft/FluidFramework
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.