CtrlK
BlogDocsLog inGet started
Tessl Logo

model-cost-compare

Starter: compare model costs for a described task — maps task shape to the cheapest adequate seat and shows the price spread

58

Quality

66%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/octopus-starter-pack/model-cost-compare/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a lean, well-sequenced workflow with concrete pricing data and useful guardrails, weakened mainly by the duplicated inline price roster, unquarantined version-sensitive details, and a reference to a file that is not part of the skill bundle. Moving the roster to a reference file and adding a worked cost example would lift both conciseness and actionability.

Suggestions

Move the model price roster out of step 3 into a one-level-deep reference file (e.g. references/pricing.md) and point to the CLAUDE.md cost table instead of duplicating every rate inline.

Isolate version-sensitive details (Opus 5.5 vs Opus 5 defaulting, Claude Code v2.1.280+) into a clearly labeled 'current defaults' or 'old patterns' section so they can be maintained without cluttering the steps.

Add a short worked example of the cost arithmetic (input tokens × input rate + output tokens × output rate, including the >272K Astra 2x/1.5x case) to make the pricing step copy-paste executable.

DimensionReasoningScore

Conciseness

Mostly efficient and free of concept over-explanation, but step 3 restates every price inline despite pointing to the CLAUDE.md cost table, and version-conditional details ('Opus 5.5', 'Claude Code v2.1.280+') are time-sensitive information not placed in an old-patterns/deprecated section — it could be tightened.

3 / 5

Actionability

Gives concrete, executable guidance: named seats with $/MTok rates, an estimation heuristic ('files touched × average size; state the assumption'), a specific escalation list, and a three-row deliverable. Minor gaps — no worked arithmetic example of the cost computation and the >272K Astra multiplier is stated but never exemplified.

4 / 5

Workflow Clarity

Six clearly sequenced steps with explicit checkpoints ('state the assumption', 'defend it in two sentences', 'Check risk surfaces', the >$1 pre-dispatch guardrail); not a batch/destructive skill, but there are no feedback loops (e.g., re-estimate when assumptions prove wrong), so minor validation gaps remain.

4 / 5

Progressive Disclosure

Under 50 lines with well-organized sections (When to use / Steps / Guardrails), fitting the simple-skill pattern, but the inline price roster is data that belongs in a one-level-deep reference file and 'skills/blocks/fable5-prompting.md' is a dangling external reference outside this bundle.

4 / 5

Total

15

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, action-oriented, and clearly differentiated, but it lacks an explicit 'Use when...' trigger clause, leaving its activation conditions only weakly implied. Adding natural trigger phrases and common synonyms (pricing, which model, budget) would raise both completeness and trigger-term quality.

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user asks which model to use, what a task will cost, or wants the cheapest adequate model for a job.'

Include common user phrasings as trigger terms — 'model pricing', 'which model', 'budget', 'expensive' — to broaden natural keyword coverage.

Briefly name the task buckets handled (refactoring, code review, long-context analysis, web research) to close the coverage gap in specificity.

DimensionReasoningScore

Specificity

Names a concrete domain and three specific actions ('compare model costs', 'maps task shape to the cheapest adequate seat', 'shows the price spread'), but does not enumerate the task types it handles, leaving minor gaps in coverage — anchor 4, not 5.

4 / 5

Completeness

Has a clear 'what' but the 'when' is only weakly implied by 'for a described task'; there is no explicit 'Use when...' clause or equivalent trigger guidance, which caps completeness at 3.

3 / 5

Trigger Term Quality

Includes natural terms a user would say ('model costs', 'cheapest', 'price spread'), but misses common variations like 'pricing', 'which model should I use', or 'budget' — good coverage with a few natural terms missing.

4 / 5

Distinctiveness Conflict Risk

The niche (model cost comparison, cheapest adequate seat, price spread) is well-differentiated with only minor overlap risk against a generic model-selection skill; not anchor 5 since explicit trigger phrases are absent.

4 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
nyldn/claude-octopus
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.