Content
71%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body delivers a highly actionable, well-sequenced churn-analysis workflow with exact signal thresholds, a numeric scoring model, and complete output templates. Weaknesses are redundancy between 'When to Use' and 'Trigger Phrases' plus a low-value 'Cost' section, and a monolithic structure that inlines large templates and references a non-existent script instead of splitting them into bundle files.
Suggestions
Deduplicate the 'Trigger Phrases' section against 'When to Use' and trim or remove the 'Cost'/'Tools Required' filler to cut tokens without losing guidance.
Move the Phase 4 output template and the save play template into a references/ file (e.g., references/report-template.md) and link to them one level deep, keeping SKILL.md as an overview.
Either provide scripts/run_skill.py or remove/adjust the cron scheduling command that references it, and add a brief intake-validation step (check data completeness before scoring) to strengthen the workflow.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The signal tables, scoring model, and templates are tight, but there is duplication and padding: the "When to Use" phrases are repeated nearly verbatim in "Trigger Phrases" ("Which customers are at risk?" / "Which customers are at risk?"), and the "Cost" table ("All signal analysis | Free (LLM reasoning)") adds little. This matches score 3 (mostly efficient but could be tightened), below 4 because whole redundant sections could be removed. | 3 / 5 |
Actionability | Quotes: ">2x their average in last 30 days", "Critical signal = 25 points each", "Red | 70-100 | Critical risk... This week", plus a fill-in save play template and a complete output-format skeleton. The guidance is copy-paste concrete for an analysis skill: exact thresholds, exact weights, exact templates covering the common cases — matching the score-5 anchor (per the rubric note that instruction-only skills need not contain code). | 5 / 5 |
Workflow Clarity | Phases 0-4 (Intake → Signal Extraction → Risk Scoring → Save Play Generation → Output) form a clear, well-sequenced pipeline, and Phase 1C is conditioned on "if data available". It sits at 4 rather than 5 because explicit validation checkpoints are absent — e.g., no step to verify intake data completeness or flag accounts with missing sources before scoring — though this read-only analysis is neither destructive nor risky enough to trigger the cap at 3. | 4 / 5 |
Progressive Disclosure | Sections are clearly organized, but the skill is a ~260-line monolith: the 77-line Phase 4 output template and the save play template are inlined content that belongs in a reference file, and no references/ or scripts/ files exist. The scheduling section references `run_skill.py`, which is not present in any bundle. This matches score 3 (some structure, but content that should be separate is inline); it is not 2 because headers make navigation easy, and not 4 because no appropriate split exists at all. | 3 / 5 |
Total | 15 / 20 Passed |