CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-debate

Structured multi-provider AI debates between Claude and available advisors — use for critical decisions

48

Quality

53%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/skill-debate/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

45%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers a genuinely clear, well-sequenced debate workflow with concrete dispatch commands and real validation checkpoints, which is its strength. However, it is heavily padded with duplicated sections and admonitions, its supporting code snippets mix executable commands with buggy or placeholder templates, and it makes no use of progressive disclosure — everything lives inline in one long file.

Suggestions

Deduplicate the repeated content: the octopus banner, provider-status block, and quality-gates table each appear twice; keep one canonical copy and cut the redundant 'MANDATORY' preamble sections.

Move stable reference material (flag reference, cost tables, export/integration docs, worked examples, the interface-design-debate rules) into references/ files and keep SKILL.md as a lean overview pointing to them.

Fix or complete the executable snippets: define how ${USER_GOAL}/${CONTEXT}/${MAX_WORDS} are populated, replace the Agent(...) pseudocode with the real tool-call form, and correct the 'grep -c ... || echo 0' bug in evaluate_response_quality (use 'grep -c ... || true' or capture exit status).

DimensionReasoningScore

Conciseness

The ~710-line body is noticeably verbose with real duplication — the octopus banner/provider-status block appears twice, the quality-gates table is repeated verbatim in two sections, and flags are documented in both the body table and examples — plus ASCII art diagrams, cost tables, attribution, and repeated 'MANDATORY/PROHIBITED' admonition sections that add tokens without adding information.

2 / 5

Actionability

The central dispatch pattern is exact and executable ('orchestrate.sh spawn "$advisor" "$prompt"', check-providers.sh, the fleet-builder sourcing), but many supporting snippets are incomplete: templates reference undefined variables (${USER_GOAL}, ${CONTEXT}, ${MAX_WORDS}), the Agent(...) call is pseudocode, and evaluate_response_quality has a real bug ('grep -c ... || echo 0' yields a two-line '0\n0' that breaks the (( )) arithmetic), which is more than the minor gaps of anchor 4.

3 / 5

Workflow Clarity

Steps 1-7.5 are clearly sequenced with genuine checkpoints: a mandatory provider-availability check before the banner, flag validation with explicit errors (--rounds 0/11+), a two-provider minimum enforcement after dispatch, and per-response quality gates with re-prompt thresholds. It falls short of anchor 5 because the low-quality re-prompt is a stub comment ('# Re-prompt for more detail') with no actual retry loop and rounds 2+ handling is only sketched.

4 / 5

Progressive Disclosure

The skill is a monolithic single file with no references/, scripts/, or assets/ directories in the bundle, while the body inlines large blocks that belong in separate files (flag tables, cost-tracking data, export/integration docs, worked examples) and points to paths outside the skill ('skills/blocks/architecture-simplification.md', '~/.claude-octopus/plugin/scripts/...') that are not part of this bundle — minimal reference structure, matching anchor 2.

2 / 5

Total

11

/

20

Passed

Description

62%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is compact, grammatical, and third-person, covering both what the skill does and when to use it, with a reasonably distinct multi-provider-debate niche. Its main weaknesses are thin action detail and a broad 'critical decisions' trigger with few natural keyword variations.

Suggestions

Enumerate 2-3 concrete actions in the description, e.g. '...dispatches advisor models in structured rounds and synthesizes their positions into a recommendation'.

Sharpen the when-clause with explicit trigger phrases, e.g. 'Use when the user asks for a debate, second opinions from multiple AI models, or a decision review'.

Add natural synonyms users would say ('AI debate', 'compare models', 'multi-model review') to improve trigger-term coverage.

DimensionReasoningScore

Specificity

Names the domain ('multi-provider AI debates between Claude and available advisors') and the core action of running structured debates, but does not enumerate what the skill concretely does (rounds, advisor dispatch, synthesis, deliverables) — several specific actions are missing, so it sits at anchor 3 rather than 4.

3 / 5

Completeness

It answers both parts — what ('Structured multi-provider AI debates between Claude and available advisors') and when ('use for critical decisions') — but the when-clause is broad and lacks concrete trigger phrases, matching anchor 4 ('when' could be more explicit) rather than 5; it is above anchor 3 because the when is explicitly stated, not merely implied.

4 / 5

Trigger Term Quality

'debates', 'multi-provider', and 'critical decisions' are relevant keywords users might say, but common natural variations ('second opinion', 'compare models', 'have the AIs review X') and synonyms are absent, matching anchor 3 rather than the good coverage of anchor 4.

3 / 5

Distinctiveness Conflict Risk

The multi-provider debate niche is clearly distinguishable from generic review or research skills, though 'critical decisions' as a trigger is broad enough to create minor overlap risk with general decision-support skills — anchor 4 rather than 5.

4 / 5

Total

14

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (737 lines); consider splitting into references/ and linking

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 1 missing, 1 deeper-than-1-level

Warning

Total

13

/

16

Passed

Repository
nyldn/claude-octopus
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.