Content
82%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A high-quality diagnostic skill body: every step is backed by concrete, executable commands, the multi-step workflow is clearly sequenced, and it teaches non-obvious material (dependency drift detection via image diffs and index lookups) without padding. The main improvement levers are moving the package-drift deep dive to a reference file and adding an explicit verify-after-fix checkpoint.
Suggestions
Add an explicit verification step after the fix (e.g., 'run `af tasks logs` / re-check `af dags stats` on the rerun to confirm the task goes green before reporting resolution') to close the workflow's feedback loop.
Move the 'Package version changes' deep dive (image diff, venv-operator specifics, index lookup) into a references/ file and link to it, keeping SKILL.md as a leaner overview with a one-line pointer.
Tighten the 'Prevention' section by replacing the trailing question-mark bullets ('Add data quality checks?') with declarative recommendations keyed to the failure categories defined in Step 2.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is command-first and assumes Airflow knowledge (no 'what is a DAG' padding), with the package-drift section adding genuinely non-obvious material. Minor trimmable bits — the trailing question-mark bullets in 'Prevention' and lines like 'Gather additional context to understand WHY this happened' — keep it at anchor 4 rather than 5's 'every token earns its place'. | 4 / 5 |
Actionability | Fully executable throughout: parameterized CLI commands (`af runs diagnose <dag_id> <dag_run_id>`, `af tasks logs ...`), copy-paste-ready `docker run ... pip freeze` diffs, a concrete `curl | jq` PyPI query, and specific remediation patterns (pin specifiers, image SHA). This matches anchor 5's 'copy-paste ready commands covering common cases'. | 5 / 5 |
Workflow Clarity | A clear four-step sequence (identify failure → get error details → check context → structured output) with a failure-taxonomy categorization checkpoint and a decision tree for run_id/DAG presence. It falls short of anchor 5 because there is no explicit validate-the-fix loop (e.g., 'rerun and confirm the task goes green' before reporting), though it is well above anchor 3's implicit checkpoints. | 4 / 5 |
Progressive Disclosure | Well-organized sections with an in-file anchor link ('See [Package version changes](#package-version-changes) below') that defers the deep-dive material; the body is coherent as a single overview. It is not 5 because the ~30-line package-drift deep dive is a natural candidate for a references/ file (no bundle exists), which would make SKILL.md a leaner overview; it is clearly not 3, since structure and navigation are good. | 4 / 5 |
Total | 17 / 20 Passed |