CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-debate

Structured multi-provider AI debates between Claude and available advisors — use for critical decisions

51

Quality

57%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/skill-debate/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

52%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable skill with a well-sequenced, checkpoint-validated debate workflow, undermined by severe duplication and padding (~700 lines with at least four verbatim-repeated sections) and broken progressive disclosure: multiple referenced bundle files do not exist while detail content is inlined monolithically. Consolidating duplicates and either shipping the referenced files or moving detail out of SKILL.md would substantially improve it.

Suggestions

Deduplicate the repeated sections — banner template, Quality Gates table/metrics, cost tracking, and export instructions each appear twice; keep one canonical copy and cross-reference it.

Ship the referenced bundle files (skills/blocks/codex-host-adapter.md, skills/blocks/architecture-simplification.md, scripts/lib/dispatch.sh) or remove the references — as written they point to nonexistent paths, and no references/ or scripts/ directories exist.

Move the flag reference, quality-gate metrics, cost tables, and export/integration detail into a separate reference file, keeping SKILL.md as a lean overview of the 7-step workflow; this would also fix the ~700-line token bloat.

DimensionReasoningScore

Conciseness

The ~700-line body is heavily padded and duplicated: the banner template appears twice (lines 29-40 and 271-281), the Quality Gates table is repeated verbatim (lines 228-240 and 673-680), cost tracking and export sections each appear twice, and sections like 'WHY Sonnet and not just more Opus?' explain model characteristics Claude already knows. This matches 'noticeably verbose; several unnecessary explanations or padded sections'. It is above score 1 because the bulk is still operational instruction rather than tutorial-style conceptual explanation.

2 / 5

Actionability

Mostly executable: exact bash commands (orchestrate.sh spawn with quoted paths, check-providers.sh, mkdir/heredoc for context.md and state.json), a concrete evaluate_response_quality shell function, and explicit AskUserQuestion payloads with options. It falls short of 5 because key variables (USER_GOAL, USER_PRIORITY, USER_CONTEXT, QUESTION) are never populated, the round-N loop only shows round-1 strings, and '# Re-prompt for more detail' is a placeholder comment rather than code — concrete guidance with minor gaps, i.e. anchor 4.

4 / 5

Workflow Clarity

Steps 1-7.5 are clearly sequenced with explicit validation checkpoints: provider availability check with portable grep rules, a two-provider minimum guard with exit on failure, quality-gate thresholds (>=75 proceed, <50 re-prompt), and a pre-application approval gate for the deliverable. It does not reach 5 because of internal contradictions — two different banner rosters (fixed participants vs. runtime-selected), a duplicated Antigravity entry with conflicting emoji, and Step 5's prompt hardcoding 'round 1' — which introduce minor ambiguity rather than missing checkpoints, placing it at anchor 4.

4 / 5

Progressive Disclosure

The bundle contains no references/, scripts/, or assets/ directories, yet the body points to 'skills/blocks/codex-host-adapter.md', 'skills/blocks/architecture-simplification.md', and 'scripts/lib/dispatch.sh' as if they were loadable files — references to nonexistent paths. Meanwhile ~700 lines of content that clearly belongs in separate files (quality gates, cost tables, flag reference, export/integration sections) is inlined in SKILL.md. This matches 'minimal structure; content that clearly belongs in separate files is inlined; or references are buried'. It is above score 1 because section headers do provide navigable structure.

2 / 5

Total

12

/

20

Passed

Description

62%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A concise, correctly third-person description that answers both 'what' and 'when', with a distinct multi-provider-debate identity. Its weaknesses are the single-action 'what', a thin set of natural trigger phrases, and a generic 'when' clause that could overlap other decision-support skills.

Suggestions

List 2-3 concrete actions to raise specificity, e.g. 'Runs structured multi-provider AI debates, collects round-by-round positions, and synthesizes a recommendation with next steps'.

Expand trigger terms with natural variations users would actually say, such as 'second opinion', 'compare models', 'adversarial review', or 'multi-model debate'.

Make the 'when' clause more specific, e.g. 'Use when making architecture or design decisions, weighing trade-offs, or when the user wants multiple AI perspectives on a question'.

DimensionReasoningScore

Specificity

The description names the domain and a single concrete action — "Structured multi-provider AI debates between Claude and available advisors" — but gives no further concrete actions (rounds, synthesis, deliverable generation), matching the anchor 'names domain and 1-2 concrete actions, but not comprehensive'. It is above score 2 because the debate mechanism is concrete rather than generic ('Processes PDF files'), and below score 4 because no additional specific capabilities are listed.

3 / 5

Completeness

Both parts are present: what it does ("Structured multi-provider AI debates between Claude and available advisors") and when to use it ("use for critical decisions"). The 'when' clause is explicit but generic — 'critical decisions' does not specify which decision types or contexts — matching anchor 4 ('when' could be more explicit or specific) rather than 5, and clearly above anchor 3 where 'when' would be absent or only implied.

4 / 5

Trigger Term Quality

"debates" and "critical decisions" are natural user phrases, but coverage stops there — missing common variations like 'second opinion', 'compare AI models', 'adversarial review', or 'multi-model review'. This fits 'some relevant keywords but missing common variations or synonyms'; it is not score 4 because 'multi-provider' and 'advisors' lean toward jargon rather than terms a user would naturally say.

3 / 5

Distinctiveness Conflict Risk

The multi-provider debate framing carves a clear niche distinct from general decision-support or review skills, leaving only minor overlap risk — 'use for critical decisions' alone is broad enough to collide with other decision-making skills. Mostly distinct with minor overlap risk fits anchor 4; it falls short of 5 because the trigger clause is not narrowly tied to the debate context.

4 / 5

Total

14

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (708 lines); consider splitting into references/ and linking

Warning

referenced_paths_exist

Referenced path issues: 1 missing, 1 deeper-than-1-level

Warning

Total

14

/

16

Passed

Repository
nyldn/claude-octopus
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.