CtrlK
BlogDocsLog inGet started
Tessl Logo

pipeline-investigation

Investigates Buildkite pipeline failures to find root causes. Returns structured JSON to the parent for formatting. Triggers when users ask about failing pipelines, build errors, or need help debugging CI/CD issues. Accepts Buildkite build URLs or build numbers and performs deep investigation.

67

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-sequenced investigation runbook with strong verification checkpoints (flaky-vs-real, competing hypotheses, already-pushed fixes). Its weaknesses are a monolithic single-file layout with no progressive disclosure and a few token-efficiency lapses (repeated token boilerplate, auto-confirmed destructive commands without validation).

Suggestions

Move the error-pattern classification table, the ASG/terraform scheduling deep-dive, and the full output JSON schema into reference files (e.g., references/error-patterns.md, references/scheduling.md) with clearly signaled one-level-deep links from SKILL.md.

Define the token-extraction helper once (e.g., 'TOKEN=$(bk auth token)' with a guard) and reference it in subsequent curl blocks instead of repeating the boilerplate verbatim in ~8 places.

Add a validation checkpoint before destructive operations: require confirmation of build state (e.g., stuck > N minutes or blocking queued builds) before 'bk build cancel', and drop '-y' so the destructive path is deliberate.

DimensionReasoningScore

Conciseness

The body is overwhelmingly operational (commands, tables, a JSON schema) but repeats 'TOKEN=$(bk auth token)' verbatim in ~8 curl blocks, re-uses the Step 2 command in Step 5, and includes trimmable asides like 'This extracts the OAuth token from the CLI's keychain storage' — efficient with minor over-explanation, matching anchor 4 rather than the fully lean anchor 5.

4 / 5

Actionability

Every step gives copy-paste-ready bk/curl/git/jq commands with explicit placeholders, a URL parse pattern, an error-pattern-to-category table, and a complete JSON output schema; concrete commands cover the common investigation cases, matching the fully-executable anchor 5.

5 / 5

Workflow Clarity

A clear 13-step sequence with genuine validation checkpoints (auth preflight, Step 10 flaky-vs-real re-run, Step 11 competing-hypotheses check, Step 13 already-fixed check), but destructive/impactful operations 'bk build cancel' and 'bk build rebuild' use '-y' auto-confirm with no pre-cancel validation checkpoint, leaving a minor validation gap at anchor 4 rather than 5.

4 / 5

Progressive Disclosure

No bundle files exist and the ~330-line body is entirely inline: content that naturally belongs in one-level-deep references (the error-classification table, the ASG/terraform scheduling deep-dive, the full output JSON schema) is embedded in SKILL.md despite good section headers — matching anchor 3 ('content that should be separate is inline') rather than anchor 4's appropriately split structure.

3 / 5

Total

16

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description in third-person voice with explicit what/when structure and a well-anchored niche. Its main weaknesses are the generic phrase 'performs deep investigation' and a trigger set that misses common colloquial variations for CI failures.

Suggestions

Drop the redundant 'performs deep investigation' clause and replace it with a concrete action, e.g., 'retrieves job logs, artifacts, and agent state to pinpoint root causes'.

Broaden trigger terms with natural user phrasings such as 'broken build', 'red pipeline', or 'CI is failing' alongside the existing 'failing pipelines' and 'build errors'.

Narrow 'debugging CI/CD issues' to 'debugging Buildkite CI failures' to reduce overlap risk with generic CI/CD or GitHub Actions skills.

DimensionReasoningScore

Specificity

Concrete actions like 'Investigates Buildkite pipeline failures to find root causes', 'Returns structured JSON to the parent', and 'Accepts Buildkite build URLs or build numbers' are specific, but 'performs deep investigation' is generic filler, keeping it below the comprehensive anchor 5.

4 / 5

Completeness

Explicitly answers both 'what' (investigates failures to find root causes, returns structured JSON) and 'when' ('Triggers when users ask about failing pipelines, build errors, or need help debugging CI/CD issues') with concrete trigger phrases, matching the anchor 5 example exactly; anchor 4 would require a less explicit 'when' clause.

5 / 5

Trigger Term Quality

Natural trigger terms 'failing pipelines', 'build errors', and 'debugging CI/CD issues' match what users would say, but common variations like 'build broke', 'red pipeline', or 'CI failing' are missing, so it sits between anchors 3 and 5 rather than at 5.

4 / 5

Distinctiveness Conflict Risk

'Buildkite' and 'Buildkite build URLs' establish a clear niche with minimal conflict, but the broad phrase 'need help debugging CI/CD issues' could overlap with generic CI skills (e.g., GitHub Actions debugging), placing it between anchors 4 and 5.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mock-server/mockserver-monorepo
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.