CtrlK
BlogDocsLog inGet started
Tessl Logo

debate-kickoff

Starter: frame a decision as a multi-model debate — picks sides, seats providers, and launches /octo:debate with a well-formed motion

60

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/octopus-starter-pack/debate-kickoff/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, well-structured instruction-only skill: the workflow is clearly sequenced with a pre-flight provider check, an explicit fallback branch, and failure-handling guardrails, all in a lean body with no padding. The main gaps are unspecified invocation details for /octo:debate and a referenced helper script that is not part of the bundle.

Suggestions

Show the exact /octo:debate invocation shape (motion plus side-assignment arguments) so the launch step is copy-paste ready.

Ship `scripts/helpers/check-providers.sh` in the bundle (or inline a one-line check command) so the referenced path resolves.

Briefly define what a completed 'round' looks like in step 5 so the synthesis step has a concrete checkpoint.

DimensionReasoningScore

Conciseness

The 26-line body is lean and assumes Claude's competence: no concept explanations, no padding, and terse imperatives like "Confirm silently from context; do not interrogate the user" and "Never fabricate a provider's position." Every token carries instruction, matching the 'lean and efficient' anchor — there is nothing an intelligent reader would need trimmed.

5 / 5

Actionability

Guidance is mostly executable for an instruction-only skill: a concrete command to run (`${CLAUDE_PLUGIN_ROOT}/scripts/helpers/check-providers.sh`), a specific slash command to invoke with defined arguments, and concrete guardrails ('state the expected cost band from the CLAUDE.md cost table'). It falls short of 5 because the invocation format for /octo:debate, the composition of the 'standard Octopus banner', and what 'rounds' mean are left unspecified, and the referenced helper script is not present in this bundle to verify.

4 / 5

Workflow Clarity

The five steps are clearly sequenced with a pre-flight check (check-providers.sh before launch), an explicit conditional branch ('if only Claude is available... offer a single-model pro/con instead'), and error recovery ('If a seat fails, report the failure and continue with the remaining seats'). It is not 5 because the synthesis step's convergence check is a one-line summary without validation of round completion or output shape, leaving minor checkpoint gaps.

4 / 5

Progressive Disclosure

The skill is under 50 lines with well-organized sections (When to use / Steps / Guardrails) and no inlined content that belongs in a separate file, so it nearly earns the simple-skill 5. It drops to 4 because the sole external reference, `scripts/helpers/check-providers.sh`, is clearly signaled one level deep but is not present in the bundle (no scripts/ directory exists), so the reference cannot be resolved as structured.

4 / 5

Total

17

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise, concrete, and clearly distinguishes the skill via the /octo:debate command and multi-model debate framing, but it omits any explicit 'when to use this' trigger guidance. Adding a 'Use when...' clause with natural user phrasings would lift both completeness and trigger-term quality.

Suggestions

Append an explicit trigger clause such as 'Use when the user has a two-sided technical decision (e.g. "Redis or Memcached?", "monorepo vs polyrepo") and wants opposing perspectives from different model families.'

Include natural user phrasings users would actually say — 'pros and cons', 'second opinion', 'X vs Y', 'adversarial review' — to improve trigger-term coverage.

Mention the synthesis output ('summarize where models converged and give one recommendation') so the 'what' covers the skill end-to-end.

DimensionReasoningScore

Specificity

The description lists several concrete actions — "frame a decision as a multi-model debate", "picks sides, seats providers", "launches /octo:debate with a well-formed motion" — which names the domain and multiple specific capabilities. It is not 5 because coverage is incomplete (the synthesis/recommendation step and the cross-lab pairing rationale are absent), but it clearly exceeds the 3 anchor's '1-2 concrete actions'.

4 / 5

Completeness

The 'what' is clear (frame a debate, assign sides, seat providers, launch /octo:debate), but there is no 'Use when...' clause or equivalent trigger guidance anywhere in the description — only the bare 'When to use' exists in the body, which does not count here. Per the judging guideline, a missing explicit trigger clause caps completeness at 3; it is not 2 because the 'what' half is explicit and concrete rather than vague.

3 / 5

Trigger Term Quality

Relevant keywords exist ("decision", "debate", "multi-model", "providers", "/octo:debate") but common natural phrasings users would actually say — "second opinion", "pros and cons", "adversarial review", "X vs Y", "which should we pick" — are missing, and there are no synonyms or variations. This matches the 'some relevant keywords but missing common variations or synonyms' anchor rather than 4, since more than a few natural terms are absent.

3 / 5

Distinctiveness Conflict Risk

The niche is fairly distinct — "multi-model debate", "seats providers", and the specific command "/octo:debate" are unlikely to trigger for unrelated skills. It is not 5 because generic words like "decision" and "picks sides" could overlap with other decision-support or planning skills, giving minor overlap risk with closely related skills.

4 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
nyldn/claude-octopus
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.