Content
82%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
A strong, dense body: executable code, real API gotchas, and honest statistical caveats (disproportionality ≠ causality, small-count instability, OpenFDA non-deduplication). The main deductions are the duplicated metric formulas between prose and code, a duplicated step number in the workflow, and no error-recovery guidance for the cell-sum validation.
Suggestions
State the PRR/ROR/IC formulas once — either keep the formulas in 'The 2x2 table' as math and let the Python block implement them, or drop the prose restatement — to trim the duplicated explanation.
Fix the workflow numbering: 'Resolve the drug field' and 'Use .exact' are both numbered '2', so the six steps mis-number as 1, 2, 2, 3, 4, 5, 6.
Add one line of recovery guidance after 'Verify a + b + c + d == N' (e.g. what a mismatch implies about the background population choice) to turn the checkpoint into a feedback loop.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Largely efficient — the OpenFDA gotchas (404-as-empty, .exact, rate limits, MedDRA licensing) are exactly what Claude does not already know. But the PRR/ROR/IC formulas appear twice (prose in 'The 2x2 table' plus the Python implementations), and the opening paragraph re-explains what an SDR is; both could be trimmed. Efficient with minor over-explanation, which is the level-4 anchor, not level 5 where every token earns its place. | 4 / 5 |
Actionability | Copy-paste-ready executable code covering the common case end-to-end: OpenFDA count queries with 404 handling, 2x2 construction for a concrete example (warfarin x gastrointestinal haemorrhage), and PRR/ROR/CI/IC implementations. The one hand-off (EBGM to the openEBGM/PhViD packages) is explicitly justified — 'the shrinkage prior is the whole point and easy to get wrong' — which the rubric's justification carve-out covers. | 5 / 5 |
Workflow Clarity | The six-step Workflow is clearly sequenced with an explicit validation checkpoint ('Verify a + b + c + d == N') and a triage-not-verdict framing. It falls short of 5 because there is no error-recovery loop around the validation (what to do when the cells don't sum to N), and the step list has a numbering glitch — two consecutive steps are both numbered '2' ('Resolve the drug field' and 'Use .exact'), which disrupts the sequence. | 4 / 5 |
Progressive Disclosure | Well-organized with clear headers (When to use, Quick start, Workflow, Hand-off, Edge cases, Standards) and one-level-deep navigation — no buried or nested references. No bundle files exist, and at ~190 lines the body is self-contained but borderline: the quick-start code block and the standards list could live in reference files to keep the overview lighter, which keeps it at 'good structure, minor organization gaps' rather than a model split. | 4 / 5 |
Total | 17 / 20 Passed |