Content
96%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
An exemplary process skill body: an unambiguous step spine, explicit validation and safety gates, a well-defined failure-delegation and retest loop, and correct offloading of exact commands to real one-level-deep references. The only deductions are minor: some retest-hygiene and AF2/AF3 readiness detail is duplicated between the steps and the Troubleshooting table, making the main file slightly longer than a pure overview.
Suggestions
Deduplicate the retest-hygiene rules stated in Step 2 and again in Step 5b, and the AF2/AF3 is_active/is_stale readiness logic stated in Step 1 and again in the Troubleshooting table — state each once and cross-reference it.
Consider moving the detailed AF2/AF3 field semantics (last_parsed_time polling window, RestApiClientException-on-POST diagnosis) into references/provisioned-testing.md, keeping Step 1 to the decision rule, to shorten the main file.
Tighten the Step 3 timeout prose by collapsing the three flavor cases into the existing table format alongside the poll-interval table.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is dense operational guidance with no concept explanations Claude already knows — it stays on non-obvious specifics like "On AF3 (REST API v2) the is_active field does not exist — require is_stale: false instead". Minor trimmable redundancy keeps it below the lean anchor: retest hygiene is stated in Step 2 ("On a retest, all run identifiers must be fresh...") and restated in 5b, and the Troubleshooting rows re-explain Step 1 AF2/AF3 readiness logic already covered inline. Not 5 because of that duplication; not 3 because there is no padded or over-explanatory material. | 4 / 5 |
Actionability | Concrete, executable specifics throughout: exact CLI calls ("aws mwaa get-environment", "get-workflow --workflow-arn"), endpoints ("the /dags collection endpoint", "PATCH /dags/<dag_id> {\"is_paused\": false}"), field names ("last_parsed_time", "has_import_errors", "is_stale"), numeric thresholds ("default to 300s", "15s intervals", "default 3"), and a copy-paste report template, with the exact command sequences correctly carried by the two real reference files. It covers the common Provisioned and Serverless cases the way the fully-executable anchor requires; it is not 4 because there is no pseudocode or missing key detail. | 5 / 5 |
Workflow Clarity | Steps 0–6 are explicitly sequenced with validation checkpoints at every risky point: a mandatory readiness check before triggering ("Do not trigger blind"), a confirmation safety gate with prod re-approval at every state-changing step, a freshness gate before adopting prior runs, terminal-state confirmation before re-triggering, and an explicit debug → fix → retest feedback loop with a progress-based cap and enumerated stop conditions. This is the clear-sequence-with-explicit-validation-and-recovery-loops anchor; it exceeds 4 because validation is not merely present but explicitly mandatory and repeated at each loop iteration. | 5 / 5 |
Progressive Disclosure | Decision logic lives inline while exact commands are properly split into two real, well-signaled, one-level-deep references (references/provisioned-testing.md and references/serverless-testing.md, both present in the bundle, each summarized in the References section and linked contextually in Step 1). Not 5 because the SKILL.md body itself runs ~280 lines and carries some content that could equally live in the references (the full AF2/AF3 readiness field semantics and the RestApiClientException rows duplicated between Step 1 and the Troubleshooting table), leaving minor organization gaps; not 3 because the split is clean, clearly signaled, and easy to navigate. | 4 / 5 |
Total | 18 / 20 Passed |