Content
63%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A well-structured, expert-level analysis skill with a clear sequenced workflow, severity framework, and a concrete output template, though it is instruction-only with no code examples. Its weaknesses are redundancy between the two check-catalog sections and a monolithic single-file layout with no progressive disclosure of reference material.
Suggestions
Consolidate 'Core Checks' and 'Specific Check Guidance' into a single deduplicated section — duplicates/keys, missingness, freshness/schema drift, and distribution shifts are each covered twice, and merging them would cut significant tokens.
Move the check catalogs ('Core Checks', 'Specific Check Guidance', 'Automated Test Guidance') into a references/ file (e.g., references/checks.md) and keep SKILL.md as a workflow overview, so the detailed catalogs load only when needed.
Add one or two short executable examples (e.g., a snippet computing null/duplicate rates by segment, or a temporal trend query) to raise actionability from concrete-but-abstract guidance to copy-paste-ready.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is dense domain guidance with no library tutorials or concept padding, but it is noticeably redundant: 'Core Checks' and 'Specific Check Guidance' cover the same ground twice (duplicates/keys, missingness, freshness/schema drift, distribution shifts appear in both sections), and the five shape-specific check lists could be tightened. This fits the 'mostly efficient but could be tightened' anchor — not 4 because the duplicated check catalogs are more than 'minor instances' of excess, and not 2 because there is no explanatory filler or concept teaching. | 3 / 5 |
Actionability | Although it is instruction-only with no code, the guidance is concretely executable: named checks per category, specifics like sentinel values "'', 'unknown', 'n/a', 0, or -1", robust-method defaults ('quantiles, MAD, or IQR before defaulting to z-scores'), a concrete cross-field rule example ('is_cancelled = false with a non-null cancelled_at'), and a structured output template with per-finding fields. Minor gap versus the 5 anchor: no example SQL/Python snippets or thresholds for the checks, so it is not fully copy-paste ready. | 4 / 5 |
Workflow Clarity | The 8-step workflow is clearly sequenced (context → path → profile → core checks → shape-specific checks → temporal → risks → fixes) with embedded checkpoints ('Confirm grain before interpreting anomalies'), a severity classification scheme, and a 7-part report structure with a per-finding template. Minor validation gaps keep it below 5: there are no explicit feedback loops (e.g., re-check after a suspected cause is confirmed) and checkpoint language is suggestive rather than mandated — but the sequence itself is coherent and gap-free, so it sits above the 3 anchor. | 4 / 5 |
Progressive Disclosure | The file is well-organized with clear headers and clearly signaled cross-skill references ($build-report, $jupyter-notebooks, $validate-data, $design-kpis), but everything lives in a single ~160-line SKILL.md with no reference files: the check catalogs ('Core Checks', 'Specific Check Guidance', 'Automated Test Guidance') are reference-like material that could be split out, and the >50-line simple-skill exception does not apply. This matches the 'some structure but content that should be separate is inline' anchor rather than the 4 anchor, since the bulk of the standards material is inlined rather than 'mostly appropriately placed' in separate files. | 3 / 5 |
Total | 14 / 20 Passed |