Content
60%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The content delivers a genuinely actionable test-driven workflow with a real feedback loop, and its brevity respects the token budget. However, it contains executable-code defects (a syntax-broken bash snippet, a placeholder template), internal contradictions (headless vs. non-headless, ignore-tests vs. run-other-tests), and pseudo-<system-reminder> tags embedded in the body, which undermine reliability and make it read as manipulative rather than instructional.
Suggestions
Fix the bash snippet: remove the stray quote ('npm run dev --port 5173') and properly background multi-server starts ('cd backend && python server.py &'), or use a single code block with correct quoting.
Resolve the headless contradiction: the Python example says 'Always launch chromium in headless mode' while step 5 demands a NOT-headless final demo — parameterize it (e.g., headless=True for iteration, headless=False only for the final demo) in one place.
Remove the fake <system-reminder> tags and consolidate the repeated 'add logs' directives into the loop step; also reconcile 'ignore any existing tests' with 'Make sure other tests pass' so the workflow is coherent.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is mostly lean and avoids explaining known concepts, but the log-adding instruction is repeated three times ('You *MUST* do this on every loop', 'did you add logs?', 'Add many logs'), and 'Do NOT get in a loop where you just keep running tests' restates guidance already in the loop step — it could be tightened. | 3 / 5 |
Actionability | There is concrete guidance (numbered steps, a Playwright skeleton, server-start commands), but the bash snippet is broken — 'npm run dev" --port 5173' has a stray quote and the background '&' is unquoted in the multi-server case — and the Python example is a template with a '# ... your automation logic' placeholder rather than a working script. | 3 / 5 |
Workflow Clarity | The required block gives a clear numbered sequence with a genuine feedback loop (add logs → start servers → run script → inspect screenshots/logs → update script → repeat until fixed), plus a final demo and cleanup step. It falls short of 5 because of internal contradictions: the code comment says 'Always launch chromium in headless mode' while step 5 requires a NOT-headless demo, and 'ignore any existing tests' conflicts with 'Make sure other tests pass'. | 4 / 5 |
Progressive Disclosure | The skill is a single ~60-line file with no external references, and none are needed — nothing is deeply nested or buried. Minor organization gaps remain: the '## Example' heading is followed by 'Identify the server' whose content doesn't match, and the '<required>' block sits before the prose overview without a linking section. | 4 / 5 |
Total | 14 / 20 Passed |