CtrlK
BlogDocsLog inGet started
Tessl Logo

diagnosing-ci-and-merge-bottlenecks

Diagnoses CI and pull-request pipeline health for a GitHub repo using the engineering analytics MCP tools — pull-requests (PR list with CI status), workflow-health (per-workflow CI trends), and pr-lifecycle (a single PR's timeline). Use when asked whether CI is getting faster or slower, which GitHub Actions workflow is the slow or flaky long-pole, how long PRs take from open to merge, how an author's merge time compares to the cohort, which open PRs have failing or pending CI, or where a specific pull request is stuck. Triggers on "engineering analytics", "is CI getting slower", "slow workflow", "flaky CI", "time to merge", "cycle time", "PR throughput", "failing checks", "where is PR <n> stuck", "CI long pole", "what's holding up this PR". For a verdict on one specific CI failure (whose fault, which commit) use investigating-ci-failures; to save these numbers as insights use turning-engineering-analytics-into-insights.

76

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

—

The risk profile of this skill

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A dense, well-structured analytics skill: concrete field-level guidance, an explicit investigative chain with validation guards, and clear output expectations. The only weak spots are minor density in the flaky-tests/DORA material and the absence of any reference-file split that could carry the deeper per-tool detail.

Suggestions

Tighten the engineering-analytics-flaky-tests bullet and the trailing DORA paragraph — both are dense walls of text; pull the recovery/quarantine definitions and the DORA environment-default rules into a short reference file or a tighter bullet list.

Consider externalizing per-tool deep semantics (e.g. the flaky-tests classification rules, DORA deploy-status proxies) into a references/ file referenced one level deep from 'The tools', leaving SKILL.md as a leaner overview + decision table.

Add a one-line 'How to verify your numbers' checkpoint summarizing the scattered guards (null-before-compare, non-null aggregation count, truncation narrowing) so the validation loop is as discoverable as the tool chain.

DimensionReasoningScore

Conciseness

Assumes Claude's intelligence (no 'what is a PR/CI' padding) and every caveat conveys non-obvious snapshot-data semantics, but the long flaky-tests bullet and the appended DORA paragraph are dense and could be trimmed — efficient with minor over-density rather than fully lean.

4 / 5

Actionability

Highly concrete instruction-only guidance: exact field names (ready_to_merge_seconds, ci.failing, p50_seconds), exact filters (state = open, not is_draft, not author.is_bot), a question→tool→how decision table, and a worked high-value chain — the instruction-skill equivalent of copy-paste-ready.

5 / 5

Workflow Clarity

Tools are read-only so the destructive/batch cap does not apply; the high-value chain explicitly sequences workflow-health → pull-requests → pr-lifecycle with validation guards ('guard for null before comparing', 'aggregate only over non-null rows', 'narrow until the real set fits') and a 'What NOT to do' checklist.

5 / 5

Progressive Disclosure

No bundle files exist; the skill is a single self-contained file with clear section headers and a decision table for navigation, but there is no one-level-deep reference split and the trailing DORA paragraph feels appended rather than integrated.

4 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplar description: third-person voice, concrete named tools with per-tool actions, an explicit 'Use when' clause, a rich natural-language trigger list, and explicit boundary routing to adjacent skills. All four dimensions land at the top anchor.

DimensionReasoningScore

Specificity

Names the domain (CI/PR pipeline health) and three concrete tools each tied to a specific action — 'PR list with CI status', 'per-workflow CI trends', 'a single PR's timeline' — giving comprehensive, not merely minor-gap, coverage.

5 / 5

Completeness

Clearly states what the skill does (diagnoses pipeline health via named MCP tools) and explicitly when to use it ('Use when asked whether…') with concrete trigger phrases.

5 / 5

Trigger Term Quality

An explicit 'Triggers on' list covers many natural user phrases plus synonyms and variations ('is CI getting slower', 'flaky CI', 'time to merge', 'cycle time', 'CI long pole', 'what's holding up this PR').

5 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (aggregate CI/PR pipeline analytics) and explicitly routes away to sibling skills ('use investigating-ci-failures', 'use turning-engineering-analytics-into-insights'), minimizing conflict risk.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
PostHog/posthog
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.