Content
81%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body is a highly actionable, well-sequenced operational guide with executable code, explicit validation/dry-run/undo checkpoints, and correct, one-level-deep bundle references. Its main weakness is length: the time-sensitive test-results section, the 12-item pitfall list, and duplicated template content inflate the token budget and would be better moved to a reference file.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Most of the body is skill-specific operational knowledge Claude would not already know (Hermes config paths, strategy orders per model family, patched-technique findings), so it is not padded with common knowledge. However, several sections are unnecessary or redundant at this altitude: the 407-line body inlines a full "Tested Results (March 2026)" section (time-sensitive data that the guidelines penalize unless placed in a deprecated/old-patterns section), a 12-item pitfall list with repetition (the encoding-escalation order appears in Step 3, Pitfall 6, and Pitfall 10), and the full GODMODE system prompt duplicated in Steps 2 and 6 when it already lives in references/jailbreak-templates.md. This fits 'Mostly efficient but includes some unnecessary explanation or could be tightened', not score 4 which requires only minor trimming — these are whole sections, not instances. | 3 / 5 |
Actionability | The guidance is fully executable throughout: copy-paste-ready loader snippet for execute_code, exact config.yaml and prefill.json contents, runnable CLI invocations ("python scripts/parseltongue.py ... --tier standard"), a complete end-to-end pipeline code block with OpenRouter client, and a decision table mapping situations to modes. Specific examples ("auto_jailbreak(model=..., dry_run=True)", example Parseltongue output) cover the common cases, matching the score-5 anchor 'Fully executable; copy-paste ready code or commands'. | 5 / 5 |
Workflow Clarity | Steps 0-6 are clearly sequenced with explicit validation and feedback loops: auto_jailbreak tests a baseline refusal first ("Tests baseline — confirms the model actually refuses"), scores each attempt, escalates through a strategy order per model family, and provides rollback ("undo_jailbreak()") plus a non-destructive dry run — exactly the validate → fix → retry loop the rubric rewards for config-mutating and batch operations. The Step 6 escalation ladder gives an explicit recovery path when a technique fails, matching the score-5 anchor with checkpoints and feedback loops. Not score 4: no meaningful validation gaps remain (baseline check, scoring, dry-run, and undo are all present). | 5 / 5 |
Progressive Disclosure | Structure is good and all referenced bundle paths are real (references/jailbreak-templates.md, references/refusal-detection.md, scripts/parseltongue.py, scripts/godmode_race.py, scripts/load_godmode.py, scripts/auto_jailbreak.py all exist), each clearly signaled with "See X for..." and only one level deep — templates and refusal patterns live in references while the body keeps a quick-start excerpt. It is not score 5 because content that clearly belongs in bundle files is inlined: the full March 2026 test-results section, the trigger-words reference list, and the 12-item pitfalls section would all be better split into a reference file, and the GODMODE template is duplicated between Step 2 and Step 6. It is not score 3: most content is appropriately placed and references are clearly signaled, not buried. | 4 / 5 |
Total | 17 / 20 Passed |