CtrlK
BlogDocsLog inGet started
Tessl Logo

debugging-ci-failures

Debugs failing GitHub Actions CI runs for PostHog PRs, commits, and branches, and answers broad CI-health questions ("is CI red?", "is master green today?", "what's broken right now?"). Use when the user asks why CI is red, asks for the current CI or master status, or mentions a failing check, GitHub Actions run, Depot runner, workflow, job, shard, merge queue kick, flaky test, lint failure, typecheck failure, snapshot diff, migration check, generated types drift, or skills build failure. Interactive runs start with the `hogli ci:insights` digest (cross-run CI history from engineering analytics), then use read-only inspection, failure classification, the smallest local reproduction with hogli, and safe reporting without rerunning CI or posting to GitHub. Running unattended as the "Master-red diagnosis" workflow: see references/master-red-incident.md for its sandbox-compatible first step.

73

Quality

92%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

—

The risk profile of this skill

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-sequenced triage skill with concrete commands, validation checkpoints, and safety gating throughout. Its length is mostly justified but slightly trimmable, and one signaled reference points to a bundle file that is absent.

Suggestions

Add the missing references/master-red-incident.md (or remove the dangling reference) so the unattended workflow's first step is actually reachable from the bundle.

Tighten the merge-queue bisection explanation and consolidate the repeated platform-outage reminders into the single 'Rule out a platform outage first' section to reduce length.

Consider moving the classification and base-rate interpretation tables into a one-level-deep reference file so SKILL.md stays a lean overview.

DimensionReasoningScore

Conciseness

Information-dense with no conceptual padding and assumes Claude's competence, but the ~345-line body has passages (merge-queue bisection sub-bullets, repeated platform-outage reminders, base-rate interpretations) that could be trimmed.

4 / 5

Actionability

Fully executable: copy-ready hogli/gh commands, jq filters, exact log search terms (FAIL, Error, ##[error]), named MCP tools, and per-class "First action" / repro tables with exact commands covering the common cases.

5 / 5

Workflow Clarity

Clear numbered sequence (outage check → safety → insights → find run → read report → classify → base rate → reproduce → report) with explicit validation checkpoints ("confirm pr_only from the current run", "potentially_resolved is a hint, not a conclusion") and a Safety rules gate around irreversible actions.

5 / 5

Progressive Disclosure

Good section structure with clearly signaled one-level-deep references, but the referenced references/master-red-incident.md does not exist in the bundle (no references/ directory) and substantial detail (classification, base-rate, repro tables) is inlined rather than split out.

4 / 5

Total

18

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description that concretely states capabilities, provides exhaustive natural-language triggers, and explicitly covers both what and when. Only minor distinctiveness risk against closely-related sibling CI skills keeps it from a clean sweep.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — "Debugs failing GitHub Actions CI runs", "failure classification", "the smallest local reproduction with hogli", "safe reporting without rerunning CI or posting to GitHub" — with comprehensive coverage and no real gaps.

5 / 5

Completeness

Explicitly answers both what (debug/classify/reproduce/report failing CI runs and CI-health questions) and when ("Use when the user asks why CI is red… or mentions a failing check…") with concrete trigger phrases.

5 / 5

Trigger Term Quality

Quotes verbatim natural user phrasings ("is CI red?", "is master green today?", "what's broken right now?") alongside a broad synonym set (failing check, Depot runner, merge queue kick, flaky test, lint/typecheck failure, snapshot diff, migration check, generated types drift, skills build failure).

5 / 5

Distinctiveness Conflict Risk

Clear PostHog-specific niche (hogli digest, Depot runners) with distinct triggers, but documented overlap with sibling skills (fixing-flaky-tests, investigating-ci-failures) it explicitly hands off to constitutes minor overlap risk.

4 / 5

Total

19

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 missing, 1 suspicious

Warning

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

14

/

16

Passed

Repository
PostHog/posthog
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.