CtrlK
BlogDocsLog inGet started
Tessl Logo

octopus-research

Thorough research across multiple sources — use for complex topics needing broad synthesis

53

Quality

58%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/octopus-research/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a strongly sequenced, largely executable workflow: five enforced steps, concrete shell commands, a validation gate, and per-step error handling — the coercive framing is repetitive but the underlying instructions are real and runnable. Its main defects are token-wasteful duplication of the same prohibition, a banner template with a duplicated provider line, a placeholder-laden task snippet, and two references to files that are absent from the bundle, which leaves a reader directed to documentation that does not exist.

Suggestions

State the no-fallback rule once (e.g., in the HARD-GATE block) and have Steps 3-4 and the Error Handling section reference it instead of restating it — this alone would remove roughly a dozen redundant lines.

Fix the banner template's duplicate Antigravity entry (two lines with different emoji both rendering ${agy_status}) and drop the hardcoded cost/time estimates or mark them as illustrative.

Resolve the dangling references: either ship 'skills/blocks/codex-host-adapter.md' and 'skill-security-framing.md' in the bundle, or inline their essential guidance (e.g., the URL-validation rules) so no step points at a missing file.

Replace the TaskCreate/TaskUpdate 'taskId: "..."' placeholders with concrete guidance on capturing the returned task ID, making that section copy-paste executable like the rest.

DimensionReasoningScore

Conciseness

The body is mostly dense, actionable instruction rather than explanation of known concepts, but the same enforcement constraint ('do NOT research directly / do not substitute') is repeated four times — in the HARD-GATE block, the Step 3 prohibition list, Step 4, and the Error Handling section — and the banner template duplicates the Antigravity line twice ('🟡 Antigravity CLI' and '🧭 Antigravity CLI' both rendering ${agy_status}). It is above the verbose/padded 2-anchor but clearly could be tightened to earn 4.

3 / 5

Actionability

Guidance is largely executable: complete bash snippets for provider detection, the orchestrate.sh invocation with concrete flags, a ready-to-use AskUserQuestion payload, and a runnable validation script. It falls short of 5 because the TaskCreate/TaskUpdate snippet leaves 'taskId: "..."' as placeholders, and no example shows how the ${depth_choice} placeholders in the banner are populated versus displayed.

4 / 5

Workflow Clarity

Five mandatory steps are explicitly sequenced with a hard ordering ('DO NOT PROCEED TO STEP 2 until...'), an explicit validation gate in Step 4, and a per-step error-handling section. Not 5 because the validation is a fragile post-hoc check (find -mmin -10 on a results directory) with no validate→fix→retry loop for orchestrate.sh failures, and the banner template contains a copy-paste error (the duplicate Antigravity line).

4 / 5

Progressive Disclosure

Section structure is clear and the operational core is appropriately inlined, but both file references are dangling: 'skills/blocks/codex-host-adapter.md' and 'skill-security-framing.md' do not exist in the bundle (no references/, scripts/, or assets/ directories are present). This matches the 3-anchor (structure present, references present but not backed by an organized bundle) rather than 4, which requires well-signaled references that actually resolve.

3 / 5

Total

14

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description has the right two-part shape (what + explicit 'use for' when) with relevant trigger keywords, but the 'what' is generic: it names the research domain without a single concrete capability, and it omits the skill's actual differentiator (multi-provider/multi-AI orchestration). It reads as a competent but under-specified trigger description.

Suggestions

Name the distinctive concrete actions, e.g. 'Orchestrates deep research across multiple AI providers (Codex, Antigravity, Claude) and synthesizes their findings into a single report' — this both raises specificity and sharply reduces conflict risk with generic research/web-search skills.

Add natural trigger phrases users would actually say, such as 'deep research', 'investigate this topic', or 'literature review', to lift trigger-term coverage.

Make the 'when' clause more concrete: 'use when the user asks for deep research, a multi-perspective investigation, or synthesis of a complex topic across sources'.

DimensionReasoningScore

Specificity

The phrase 'Thorough research across multiple sources' names the domain but the actions are generic — 'research' and 'synthesis' restate the category without any concrete capability (no mention of provider orchestration, synthesis file generation, or output formats). It sits above 1 (a domain is named) but below 3, which requires 1-2 concrete actions like 'extracts text from PDFs'.

2 / 5

Completeness

Both parts are explicit: 'Thorough research across multiple sources' answers what, and 'use for complex topics needing broad synthesis' answers when. It is not 5 because the 'when' clause lacks concrete trigger phrases (compare the 5-anchor's 'when the user mentions PDFs, forms, or document extraction'), and not 3 because a clear 'what' plus an explicit 'use for' clause is present.

4 / 5

Trigger Term Quality

'research', 'multiple sources', 'complex topics', and 'broad synthesis' are relevant keywords a user might use, but common variations and synonyms are missing ('deep research', 'investigate', 'literature review', 'multi-provider'). Matches the 'some relevant keywords but missing common variations' anchor rather than 4, which expects good natural-term coverage.

3 / 5

Distinctiveness Conflict Risk

'Thorough research across multiple sources' and 'complex topics needing broad synthesis' carve out a recognizable niche (multi-source synthesis research) that is distinct from single-source or extraction skills, though it would still compete with any built-in deep-research or web-search skill since its unique multi-provider mechanic is never mentioned. Between 3 (overlap risk with similar skills) and the 5-anchor's clearly distinct trigger set — leaning 4.

4 / 5

Total

13

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
nyldn/claude-octopus
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.