Content
90%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A dense, expert-level pattern reference: executable code for every technique, a symptom→fix pitfalls table, and genuine validation checkpoints (val/test discipline, dual-metric reporting). The only weaknesses are the absence of an explicit ordered workflow with error-recovery loops and a monolithic single-file layout that could offload reference detail to keep the body leaner.
Suggestions
Add a short numbered workflow section (data layout → split → training loop with noise → loss selection → val-based checkpointing → one-shot test evaluation) so the topical sections read as an explicit pipeline with feedback loops (e.g., 'if rollout diverges after ~10 steps, re-check noise σ and layout', expanding the pitfalls row into a retry step).
Move the full loss-function implementations and the two nRMSE variants into a references/ file (e.g., references/losses.md, references/metrics.md), keeping one-line summaries and usage weights in SKILL.md, to bring the body closer to a lean overview with one-level-deep references.
Fix the small code smell of two consecutive NOISE_STD assignments (1e-3 then 5e-3) by making the value a single configurable constant or a commented choice, so the snippet is unambiguous when copied verbatim.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is lean and code-forward: every section delivers a pattern as executable Python plus a one-line rationale ("This prevents the model from relying on artificially clean inputs"), and it assumes Claude's competence by never explaining what FNO, PDEs, or teacher forcing are. The only near-redundancy is the short noise-level bullet list echoing the code comments, which is minor and matches the level-5 anchor; level 4 would require trimmable over-explanation, which is not present. | 5 / 5 |
Actionability | Guidance is copy-paste ready throughout: a complete rollout training loop with tensor shapes ("inp = torch.cat([inp[:, :, 1:, :], pred], dim=-2)"), working h1_loss / frequency_loss / boundary-weighting implementations, both PDEBench-style nRMSE functions, a concrete 8000/1000/1000 split, and a pitfalls table with symptom→fix pairs. This matches the level-5 anchor (fully executable, covering common cases) rather than level 4's 'minor gaps'. | 5 / 5 |
Workflow Clarity | Sections follow a coherent training-pipeline order (data layout → rollout training → noise → losses → normalization → metric → split → pitfalls) with real checkpoints: validation split for "Checkpoint selection", test "evaluated ONCE", and "Always compute BOTH and verify you beat baselines under both". It falls short of level 5 because there is no explicit step-by-step sequence with feedback/recovery loops (e.g., what to do when rollout diverges beyond the one-line pitfalls table), but it is clearly above level 3, where checkpoints would be merely implicit. | 4 / 5 |
Progressive Disclosure | The skill is a single self-contained file with well-organized, clearly headed sections and no nested or chained references — nothing is buried and navigation is easy. It scores 4 rather than 5 because at ~190 lines some reference-grade detail (full loss-function variants, nRMSE discussion) is inlined in SKILL.md where a one-level-deep references/ file would keep the overview leaner; it is well above level 3, since the structure is good and no content is misplaced or hidden. | 4 / 5 |
Total | 18 / 20 Passed |