CtrlK
BlogDocsLog inGet started
Tessl Logo

debugging-dags

Comprehensive DAG failure diagnosis and root-cause analysis with structured investigation and prevention recommendations. Use when deep failure investigation is needed, a DAG fails to import/parse or 'airflow dags list' errors on a file; a task or run is failing and must be diagnosed and fixed; requests like 'why did X fail', 'my dag keeps failing — find and fix it', or fixing a broken DAG so it loads cleanly. For simple 'why did it fail / show logs', the airflow skill handles it directly.

72

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A high-quality diagnostic skill body: every step is backed by concrete, executable commands, the multi-step workflow is clearly sequenced, and it teaches non-obvious material (dependency drift detection via image diffs and index lookups) without padding. The main improvement levers are moving the package-drift deep dive to a reference file and adding an explicit verify-after-fix checkpoint.

Suggestions

Add an explicit verification step after the fix (e.g., 'run `af tasks logs` / re-check `af dags stats` on the rerun to confirm the task goes green before reporting resolution') to close the workflow's feedback loop.

Move the 'Package version changes' deep dive (image diff, venv-operator specifics, index lookup) into a references/ file and link to it, keeping SKILL.md as a leaner overview with a one-line pointer.

Tighten the 'Prevention' section by replacing the trailing question-mark bullets ('Add data quality checks?') with declarative recommendations keyed to the failure categories defined in Step 2.

DimensionReasoningScore

Conciseness

The body is command-first and assumes Airflow knowledge (no 'what is a DAG' padding), with the package-drift section adding genuinely non-obvious material. Minor trimmable bits — the trailing question-mark bullets in 'Prevention' and lines like 'Gather additional context to understand WHY this happened' — keep it at anchor 4 rather than 5's 'every token earns its place'.

4 / 5

Actionability

Fully executable throughout: parameterized CLI commands (`af runs diagnose <dag_id> <dag_run_id>`, `af tasks logs ...`), copy-paste-ready `docker run ... pip freeze` diffs, a concrete `curl | jq` PyPI query, and specific remediation patterns (pin specifiers, image SHA). This matches anchor 5's 'copy-paste ready commands covering common cases'.

5 / 5

Workflow Clarity

A clear four-step sequence (identify failure → get error details → check context → structured output) with a failure-taxonomy categorization checkpoint and a decision tree for run_id/DAG presence. It falls short of anchor 5 because there is no explicit validate-the-fix loop (e.g., 'rerun and confirm the task goes green' before reporting), though it is well above anchor 3's implicit checkpoints.

4 / 5

Progressive Disclosure

Well-organized sections with an in-file anchor link ('See [Package version changes](#package-version-changes) below') that defers the deep-dive material; the body is coherent as a single overview. It is not 5 because the ~30-line package-drift deep dive is a natural candidate for a references/ file (no bundle exists), which would make SKILL.md a leaner overview; it is clearly not 3, since structure and navigation are good.

4 / 5

Total

17

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states concrete capabilities, includes natural trigger phrasing users would actually say, explicitly covers both what and when, and even disambiguates against a sibling skill. The only minor weakness is that the capability list is process-flavored ('diagnosis', 'root-cause analysis') rather than enumerating concrete operations, which keeps specificity at 4.

DimensionReasoningScore

Specificity

Names the domain (Airflow DAG failures) and several distinct actions — 'DAG failure diagnosis and root-cause analysis', 'structured investigation', 'prevention recommendations' — but the actions are process-level rather than fully concrete operations. Fits anchor 4 ('lists several specific actions; minor gaps') rather than 5, and is well above anchor 3's '1-2 concrete actions'.

4 / 5

Completeness

Explicitly answers both what ('Comprehensive DAG failure diagnosis and root-cause analysis with structured investigation and prevention recommendations') and when ('Use when deep failure investigation is needed... requests like...') with concrete trigger phrases — the anchor 5 example pattern. Not 4, since the 'when' is fully explicit rather than improvable.

5 / 5

Trigger Term Quality

Comprehensive natural-language coverage: 'why did X fail', 'my dag keeps failing — find and fix it', 'DAG fails to import/parse', "'airflow dags list' errors on a file", plus symptom variations. Matches anchor 5 (comprehensive natural terms including synonyms); nothing common is missing.

5 / 5

Distinctiveness Conflict Risk

Clear niche (deep DAG failure investigation) with an explicit boundary clause — 'For simple "why did it fail / show logs", the airflow skill handles it directly' — minimizing overlap with the general airflow skill. Distinct triggers (import errors, 'find and fix it') match anchor 5.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
astronomer/agents
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.