CtrlK
BlogDocsLog inGet started
Tessl Logo

pipeline-test-triage

Triage test failures in an Azure DevOps pipeline run for Fluid Framework. Use when the user shares an ADO build/pipeline URL or build ID and wants failures analyzed, when they ask "why did this pipeline fail", "summarize the test failures", "is this test flaky or a real bug", "compare these runs", or wants bugs filed for genuine failures. Covers the Real Service End to End Tests pipeline as the worked example but applies to any FF test pipeline.

76

Quality

96%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

92%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a dense, highly actionable operations runbook with excellent conciseness, actionability, and workflow clarity — every section gives executable commands, explicit sequencing, and validation/feedback loops. Its only weakness is progressive disclosure: despite one clean external reference, the bulk of the advanced REST and bug-filing detail is inline rather than off-loaded to one-level reference files.

Suggestions

Extract the detailed Tier 3 REST/timeline and retried-attempt recipes into a reference file (e.g. references/ado-rest-recipes.md) and signal it from the body, so the SKILL.md overview stays lighter while preserving the deep guidance.

Move the full bug Description template and ADO work-item gotchas into references/bug-template.md, linked from the Filing bugs section, to reduce inline bulk.

Consolidate the "green builds hide failures" guidance into a single clearly labeled callout to remove the small amount of duplication across Tier 1 and Tier 3.

DimensionReasoningScore

Conciseness

The body assumes Claude's competence (no explaining what ADO/mocha/PDF is) and packs only operational specifics — endpoints, GUIDs, regexes, gotchas — with every token earning its place; minor re-emphasis of the green-builds-hide-failures gotcha in context is not padding.

3 / 3

Actionability

Provides fully executable guidance throughout — exact MCP tool calls with params, exact REST URLs with api-versions, copy-paste PowerShell token recipes, a verified helper script invocation, and a ready-to-paste bug Description template — all copy-paste ready with no pseudocode.

3 / 3

Workflow Clarity

Multi-step processes are explicitly sequenced with validation checkpoints and feedback loops: tier escalation, Tier 2 Steps 1–4, the deterministic/flaky decision criteria, the numbered timeline→deep-link recipe, and a confirm-scope gate before any bug write.

3 / 3

Progressive Disclosure

Header organization and the single one-level-deep, clearly signaled reference (scripts/Parse-AdoTestLog.ps1, verified present) are good, but most advanced material (REST/timeline/retried-attempt recipes, full bug template) lives inline rather than split into separate reference files.

2 / 3

Total

11

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is exemplary: it states concrete capabilities in third person, includes verbatim natural-language trigger phrases, explicitly answers both what and when, and occupies a clearly distinct niche unlikely to conflict with other skills. No improvements needed.

DimensionReasoningScore

Specificity

Enumerates multiple concrete actions ("Triage test failures", "wants failures analyzed", "compare these runs", "wants bugs filed") in third person, matching the score-3 anchor that lists several specific concrete actions.

3 / 3

Completeness

Answers both what (triage/classify/file bugs) and when via an explicit "Use when the user shares an ADO build/pipeline URL..." clause enumerating trigger conditions, satisfying the score-3 anchor for explicit what-and-when triggers.

3 / 3

Trigger Term Quality

Quotes verbatim natural user phrasings ("why did this pipeline fail", "summarize the test failures", "is this test flaky or a real bug", "compare these runs") plus concrete artifacts (ADO URL/build ID), giving good coverage of terms users would actually say.

3 / 3

Distinctiveness Conflict Risk

A tightly scoped niche (Azure DevOps pipeline runs for Fluid Framework, flaky-vs-deterministic test triage, ADO bug filing) with distinct triggers makes it unlikely to fire for the wrong skill.

3 / 3

Total

12

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 2 missing, 1 deeper-than-1-level

Warning

Total

15

/

16

Passed

Repository
microsoft/FluidFramework
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.