CtrlK
BlogDocsLog inGet started
Tessl Logo

macios-ci-postmortem

Post-mortem analysis of CI failures across recent PRs in dotnet/macios. Identifies flaky tests, infrastructure issues, and shared regressions by analyzing builds from the last week. Files or updates GitHub issues for failures unrelated to any specific PR. Use when asked to "find flaky tests", "CI post-mortem", "what's been failing in CI", or "file issues for flaky failures".

73

Quality

92%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A dense, highly actionable, well-sequenced operational playbook: executable commands and parsers, explicit validation checkpoints, and a strong confirmation gate before mutating GitHub state. Its main weakness is progressive disclosure — nearly everything lives inline in a ~740-line SKILL.md despite having a references directory available.

Suggestions

Move the full GitHub issue/comment body templates (Step 4.5) into references/issue-templates.md and keep only field-by-field guidance inline, cutting a large block from the always-loaded context.

Extract the build-failure deep-dive (Step 2.6a: error-code patterns, binlog collection/attachment) and the Windows-integration bot-correlation procedure (Step 3.3c) into separate reference files, signaling them from the main workflow.

Fix the duplicate step numbering (two sections labeled 'Step 3.5') and drop the trivial Step 3.1 SQL predicate so phase numbering and queries stay tight.

DimensionReasoningScore

Conciseness

Nearly every token is domain-specific operational knowledge Claude would not know (TestSummary vs HtmlReport artifact strategy, '-clean.xml' skipping, first-failed-step root-cause rule, deep-link j=/t= URL format), with almost no explanation of concepts Claude already knows. It is not a 5 because there are minor tightening opportunities: duplicate step numbering ('Step 3.5: File one issue per test' and 'Step 3.5: Produce classification summary'), and the Step 3.1 SQL ('HAVING COUNT(DISTINCT build_id) > 0') is trivially true and adds little.

4 / 5

Actionability

The body provides copy-paste-ready az/gh CLI commands, executable Python parsers, SQL schemas and queries, concrete artifact-name patterns, exact deep-link URL formats, and a complete fill-in issue body template — covering the common cases end to end. This matches the 5 anchor ('fully executable; copy-paste ready code or commands'); the 4 anchor would require minor gaps, and the few placeholder fragments (e.g. '<buildId>') are standard parameterization, not missing detail.

5 / 5

Workflow Clarity

Four phases are explicitly sequenced with numbered steps, validation checkpoints throughout (TestSummary first-pass filter before expensive HtmlReport downloads, cross-referencing commit-SHA groups for flake detection, first-failed-step root-cause verification, reopen-decision rules with a 2-week grace window), and a hard user-confirmation gate before any batch issue action ('Never file or modify issues without user confirmation'). Despite being a batch operation it has explicit validation and feedback loops, matching the 5 anchor.

5 / 5

Progressive Disclosure

The single reference ('references/azure-devops-cli.md') is clearly signaled in a References section and is one level deep (no nested references), but the ~740-line body keeps nearly all detail inlined. Material that clearly belongs in separate reference files — the full GitHub issue body template (Step 4.5), the build-failure deep-dive (Step 2.6a), and the Windows-integration bot-correlation procedure (Step 3.3c) — is inline in SKILL.md. This sits between the 2 anchor ('content that clearly belongs in separate files is inlined') and the 4 anchor ('most content is appropriately placed'): structure and signaling are good, but the split is not.

3 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

Excellent description: it states concrete capabilities in third person, scopes them to a specific repository and workflow, and provides an explicit 'Use when' clause with natural trigger phrases including question-form variants. No fluff, over-claims, or verbosity.

DimensionReasoningScore

Specificity

The description lists multiple concrete, specific actions — 'Identifies flaky tests, infrastructure issues, and shared regressions by analyzing builds from the last week. Files or updates GitHub issues for failures unrelated to any specific PR' — grounded in a named domain (dotnet/macios CI). Nothing is vague or padded, so it matches the top anchor rather than the 4 anchor (which expects minor coverage gaps).

5 / 5

Completeness

Both 'what' (identify flaky tests, infrastructure issues, shared regressions, file/update GitHub issues) and 'when' (explicit 'Use when asked to...' with concrete trigger phrases) are clearly and explicitly stated, matching the 5 anchor exactly. A 4 would require the 'when' to be less explicit — it is not.

5 / 5

Trigger Term Quality

The 'Use when' clause enumerates natural user phrasings including synonyms and question forms: 'find flaky tests', 'CI post-mortem', 'what's been failing in CI', 'file issues for flaky failures'. This comprehensively covers how a user would actually phrase the request, matching the 5 anchor ('natural terms including synonyms') rather than 4 ('a few natural terms missing').

5 / 5

Distinctiveness Conflict Risk

The description names a clear niche (CI failure post-mortems for the dotnet/macios repository) with distinct triggers ('CI post-mortem', 'find flaky tests' in this repo context). It is third person, non-generic, and unlikely to fire for unrelated skills, matching the 5 anchor 'clear niche with distinct triggers'.

5 / 5

Total

20

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (742 lines); consider splitting into references/ and linking

Warning

relative_links

Relative link issues: 2 missing

Warning

Total

14

/

16

Passed

Repository
dotnet/macios
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.