Content
57%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A well-organized, mostly lean instruction skill whose core defect is a broken bundle contract: it delegates room prioritization to a script that is absent and never surfaces the reference files that do exist. The search loop itself is clearly specified and the example is concrete, so the skill is usable if the prioritization gap is fixed.
Suggestions
Either add the referenced prioritize_rooms.py to the bundle (e.g., under scripts/) or remove the reference and inline an explicit default room priority list so 'Plan Search Order' is executable as written.
Link the existing reference files from the body (e.g., 'See [action_guide.md](references/action_guide.md) for action semantics and [example_trajectory_breakdown.md](references/example_trajectory_breakdown.md) for a worked trajectory') so they are discoverable rather than orphaned.
Trim the 'Notes & Best Practices' section of items that duplicate 'Core Process' and 'Key Actions' to tighten token efficiency.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is efficient: it does not explain concepts Claude already knows, and the ScienceWorld-specific details (observation format, ambiguous object names) are genuinely additive. It is not a 5 because the 'Notes & Best Practices' section repeats guidance already given in 'Core Process' and 'Key Actions' (e.g., 'look around' after teleporting, using 'examine' for ambiguity). | 4 / 5 |
Actionability | Concrete commands are given ('teleport to <ROOM_NAME>', 'look around', 'examine <OBJECT>') and the example is worked end-to-end, but key details are missing: step 2 says 'Use the provided prioritize_rooms.py script' yet no such script exists in the bundle, and the promised 'default priority list based on common object locations' is never provided. This incomplete-but-concrete state matches the score-3 anchor rather than the mostly-executable anchor at 4. | 3 / 5 |
Workflow Clarity | The search loop is clearly sequenced with an explicit stop condition and a failure report ('If the object is not found after searching all candidate rooms, report this failure'), and identity-confirmation via 'examine' acts as a checkpoint. However, the 'Plan Search Order' step is unexecutable as written because it depends on a nonexistent script and an unstated default room list, leaving a material gap that the score-4 anchor's 'minor validation gaps' does not cover. | 3 / 5 |
Progressive Disclosure | The body itself is short and well-sectioned, but judged against the actual bundle: the three files in references/ (action_guide.md, action_primer.md, example_trajectory_breakdown.md) are never referenced from the body, so relevant existing material is undiscoverable, while the one file the body does reference (prioritize_rooms.py) does not exist. This fits 'some structure but could be better organized; references present but not clearly signaled' rather than the good-structure anchor at 4. | 3 / 5 |
Total | 13 / 20 Passed |