CtrlK
BlogDocsLog inGet started
Tessl Logo

debug-pipeline

Diagnose Harness pipeline executions via MCP. Analyzes any execution (failed or successful) to produce structured reports with stage/step breakdown, timing, bottlenecks, failure details, chained pipeline drill-down, and execution logs. Use when asked to debug a pipeline, investigate a failure, find out why a build failed, analyze pipeline errors, check execution logs, review execution performance, or find bottlenecks. Trigger phrases: debug pipeline, pipeline failed, why did my build fail, analyze failure, pipeline error, execution logs, fix pipeline, execution bottleneck, slow pipeline.

72

Quality

89%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, highly actionable skill body: concrete MCP tool calls with documented defaults, exact error-string matching, a response template, and a clear preferred-to-fallback flow. Weaknesses are mild — some generic coaching in Performance Notes, implicit rather than explicit conditions for the fallback steps, and a long single file that inlines reference-style material.

Suggestions

Remove or drastically trim the 'Performance Notes' section — 'take your time', 'quality matters more than speed' is generic coaching that assumes Claude would otherwise be hasty, and it adds tokens without actionable content.

Make the fallback flow explicit: state when to skip harness_diagnose and use harness_list/harness_get directly (e.g., 'if diagnose returns no execution or you need to browse before choosing one, run Step 3').

Move the Analysis Framework error catalog and Troubleshooting sections into a references/ file (e.g., references/error-patterns.md) and link to it, keeping SKILL.md as a leaner overview of the diagnostic workflow.

DimensionReasoningScore

Conciseness

The body is dense and mostly earns its tokens — tool-call blocks, a parameter table, and a terse error catalog — but 'Performance Notes' ('Take your time analyzing logs thoroughly', 'Quality of diagnosis is more important than speed') is generic coaching Claude doesn't need, and Steps 4–6 restate parameters already covered by the diagnose table. This fits the score-4 anchor ('efficient; minor instances of over-explanation that could be trimmed') rather than the 5 anchor where every token earns its place.

4 / 5

Actionability

Every step is a concrete, copy-paste-ready MCP call with named parameters, the parameter table documents defaults ('log_snippet_lines: 120, max_failed_steps: 5'), the analysis framework matches exact error strings ('No delegate available', 'ImagePullBackOff'), and a response template is provided. This is fully executable guidance covering the common cases — the score-5 anchor, not the 4 anchor's 'minor gaps'.

5 / 5

Workflow Clarity

The sequence is clear with a preferred path (Step 1/1b), a health overview, and labeled fallbacks, plus a Troubleshooting section covering recovery when logs or executions can't be found. It falls short of the score-5 anchor because there are no explicit validation checkpoints or decision rules for when to switch from harness_diagnose to the fallback steps (Steps 3–6) — conditions are only weakly implied by '(if needed)'.

4 / 5

Progressive Disclosure

No bundle files exist (no references/, scripts/, or assets/), and the single ~190-line body is well-sectioned with headers and a table, fitting the score-4 anchor ('good structure; most content is appropriately placed'). It is not a 5 because the error-pattern catalog and troubleshooting content are inlined in one long file where a references file would keep SKILL.md a leaner overview; it is above the 3 anchor since structure and navigation are genuinely good.

4 / 5

Total

17

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An excellent description: it names the domain and concrete capabilities, explicitly states when to use it with an extensive list of natural trigger phrases, and stays in third person without padding. The only weakness is that many trigger terms are generic pipeline-debugging vocabulary rather than Harness-specific, creating minor overlap risk with other CI-debugging skills.

DimensionReasoningScore

Specificity

The description lists multiple concrete capabilities — 'structured reports with stage/step breakdown, timing, bottlenecks, failure details, chained pipeline drill-down, and execution logs' — giving comprehensive coverage of what the skill does. It clearly matches the score-5 anchor ('lists multiple specific concrete actions; comprehensive coverage'), not the 4 anchor which expects minor coverage gaps.

5 / 5

Completeness

It explicitly answers both 'what' ('Analyzes any execution (failed or successful) to produce structured reports...') and 'when' ('Use when asked to debug a pipeline... or find bottlenecks') with concrete trigger phrases. This matches the score-5 anchor exactly; the score-4 anchor requires the 'when' to be less explicit, which is not the case here.

5 / 5

Trigger Term Quality

Natural user phrasings are comprehensively covered both in the 'Use when' clause ('debug a pipeline, investigate a failure, find out why a build failed, check execution logs, review execution performance, or find bottlenecks') and in an explicit trigger list ('debug pipeline, pipeline failed, why did my build fail, fix pipeline, execution bottleneck, slow pipeline'). Coverage includes synonyms and phrasings a user would naturally say, matching the score-5 anchor rather than the 4 anchor's 'a few natural terms missing'.

5 / 5

Distinctiveness Conflict Risk

The niche is clear (Harness pipeline diagnostics via MCP), but several trigger phrases — 'pipeline failed', 'fix pipeline', 'slow pipeline' — are generic CI terms that would also naturally fire on other pipeline/build-debugging skills (e.g., GitHub Actions or Jenkins debugging). This fits the score-4 anchor ('mostly distinct; minor overlap risk with closely related skills') better than the score-5 anchor's 'minimal conflict risk'.

4 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
harness/harness-ai
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.