Content
82%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
This is a tightly written, highly operational skill: concrete commands, numeric thresholds, an engine-specific cause/fix table, and unusually disciplined sourcing (verifying skip annotations against the installed gdUnit source, flagging unverifiable claims as NOT SOURCEABLE). The few weaknesses are minor: a small amount of introductory flakiness philosophy, no error-recovery loop for unparseable logs, and a single-file layout that could offload engine details to reference files.
Suggestions
Trim the opening definition of what a flaky test is and the "worse than no tests" rationale — Claude already knows this — to gain token efficiency.
Add a small error-recovery branch for logs that fail to parse (e.g., unrecognized XML schema → report the format found and ask the user) to close the workflow-clarity feedback-loop gap.
Move the engine-specific parsing and skip-mechanism details (Godot/Unity/Unreal sections) into a references/ file, keeping SKILL.md as a tighter overview.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is dense and operational — thresholds, grep patterns, and an engine-specific cause table with almost no filler. The main over-explanations are the opening definition of what a flaky test is ("A flaky test is one that sometimes passes and sometimes fails...") and its "worse than no tests" rationale, plus a few motivational lines in the Collaborative Protocol — minor trimmable material, not padding. | 4 / 5 |
Actionability | Guidance is fully executable: exact grep patterns ("<testcase name=", "Result={Success}"), numeric flakiness tiers (>25% / 5–25% / 1–5%), engine-specific skip syntax with source-file citations, verbatim user prompts, and a copy-paste report template. The honest "NOT SOURCEABLE" markers that instruct asking the user instead of inventing attributes make the guidance exceptionally safe to execute. | 5 / 5 |
Workflow Clarity | A clear numbered sequence (parse arguments → locate data → parse results → identify → recommend → report → update suite) with real checkpoints: the under-3-runs statistical guard, the stop-and-ask branch when no logs exist, and ask-before-write gates on both file writes. It falls short of the top anchor because there is no validate-and-retry feedback loop (e.g., what to do when a log format fails to parse), though the operations are append-only and user-approved, so the destructive-operation cap does not apply. | 4 / 5 |
Progressive Disclosure | No bundle files exist (references/, scripts/, assets/ are absent), so everything lives in a single ~240-line SKILL.md. Section structure is clean and numbered, and external details are correctly pushed to clearly signaled one-level-deep project docs (e.g., "docs/engine-reference/unreal/current-best-practices.md, 'Command Line'"). At this length, the engine-specific parsing and skip-mechanism material is a plausible candidate for a references/ split, which keeps it just below the top anchor. | 4 / 5 |
Total | 17 / 20 Passed |