CtrlK
BlogDocsLog inGet started
Tessl Logo

flow-deliver

Multi-AI validation, scoring, and review using available external providers (Double Diamond Deliver phase)

48

Quality

53%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/flow-deliver/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

56%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers genuinely executable, well-gated workflow instructions with concrete commands and error paths, but at roughly three times the necessary length: the same steps, banners, and report formats are restated in three redundant sections with conflicting step numbering. Splitting the example report, format template, and subtype supplements into reference files and collapsing the duplication would materially improve it.

Suggestions

Collapse the duplicated "Deliver Workflow" and "Implementation Instructions" sections into the single EXECUTION CONTRACT sequence, keeping one authoritative step numbering and one banner template — this alone would remove hundreds of lines.

Move the ~100-line Example 1 validation report, the Validation Report Format template, and the dev-subtype supplement table into a references/ file (e.g. references/report-format.md, references/subtype-supplements.md) linked with a one-line pointer.

Fix the cross-shell variable bug by re-deriving $VALIDATION_FILE at the top of Step 6's bash block (or merging Steps 5-6 into one invocation), and consolidate version notes (v2.1.16+, v8.44.0, v8.49.0) into a single compatibility section.

DimensionReasoningScore

Conciseness

The ~885-line body states the same workflow three times (the EXECUTION CONTRACT steps 1-7, the later "Deliver Workflow" section repeating steps 1-2 with near-identical banner templates, and "Implementation Instructions" restating the sequence again), and the provider banner appears in four variants. Version-specific details ("Claude Code v2.1.16+", "v7.16.0", "v8.44.0", "v8.49.0") are scattered inline rather than isolated. This matches anchor 2 (noticeably verbose, several unnecessary/padded sections); it is not 1 because the body mostly avoids explaining concepts Claude already knows and the specifics it does carry are genuine instructions.

2 / 5

Actionability

Guidance is highly concrete: copy-paste bash invocations of orchestrate.sh/state-manager.sh, a subtype table with exact validation supplements to append, a full report format template, and a complete PR-posting script. It is not 5 because of small executability gaps: $VALIDATION_FILE is assigned in Step 5's shell but consumed in Step 6's separate shell invocation (variables don't persist between Bash calls), and the orchestrate example embeds a literal \n\n inside double quotes which bash will not expand to newlines.

4 / 5

Workflow Clarity

The execution contract gives an explicit numbered sequence with mandatory blocking gates, a verify-file-exists validation checkpoint in Step 5, per-step error handling, a "Validation Checklist" section, and explicit no-fallback rules. It is not 5 because the three redundant restatements use conflicting step numberings (the contract's Step 3 is orchestrate.sh, but "How It Works" labels orchestrate.sh as Step 1 and "Implementation Instructions" as Step 3), which creates genuine ambiguity about which numbering is authoritative.

4 / 5

Progressive Disclosure

The bundle contains no references/, scripts/, or assets/ files, and everything lives in one monolithic SKILL.md: a ~100-line fully worked example validation report, the complete report-format template, the subtype supplement table, and repeated banner templates — content that clearly belongs in separate reference files. External skill references (skill-doc-sync, skill-ship) are one level deep and clearly signaled. This matches anchor 3 (some structure, but content that should be separate is inline); not 4 because well over a third of the body is template/example material inlined rather than split out.

3 / 5

Total

13

/

20

Passed

Description

50%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description communicates a clear niche (multi-provider validation/review) but is a single clause with no use-when guidance, near-synonymous action verbs, and thin trigger-term coverage. It sits squarely at the midpoint of the rubric on all four dimensions.

Suggestions

Append an explicit 'Use when...' clause with concrete trigger phrases, e.g. "Use when the user asks to review, validate, score, quality-check, or verify an implementation or document before shipping."

Broaden trigger terms with natural synonyms users actually say — "check if X works", "verify the implementation of Y", "quality check", "audit" — instead of the near-synonyms validation/scoring/review.

State what gets validated (code and documents) so the description is distinguishable from generic single-model review skills.

DimensionReasoningScore

Specificity

The description names the domain and three actions — "validation, scoring, and review using available external providers" — but the actions are near-synonyms and it never says what is validated (code, documents, etc.). It matches anchor 3 (names domain and 1-2 concrete actions, not comprehensive); it is not 4 because it does not list several genuinely distinct actions like the 'extracts text, fills forms, converts pages' example, and not 2 because it does give multiple named actions plus the operating mechanism.

3 / 5

Completeness

The 'what' is clearly stated (multi-AI validation/scoring/review via external providers) but there is no 'when' — no "Use when..." clause or equivalent trigger guidance anywhere in the description, which per the judging guidelines caps completeness at 3. It is not 4 because the 'when' is entirely absent rather than merely imprecise, and not 2 because the 'what' is clear, not vague.

3 / 5

Trigger Term Quality

"validation, scoring, and review" are natural phrases users would say, but common variations and synonyms are missing — "check", "verify", "test", "quality check", "audit" — several of which the skill's own trigger field lists (e.g. "quality check for X", "verify the implementation of Y"). This matches anchor 3 (some relevant keywords, missing common variations); not 4 because coverage of natural phrasings is thin, not 5 because no synonyms or file extensions appear.

3 / 5

Distinctiveness Conflict Risk

"Multi-AI" and "(Double Diamond Deliver phase)" signal a niche mechanism, but the headline terms "validation" and "review" are extremely broad and would overlap generic review/quality-check skills. This matches anchor 3 (somewhat specific but could still overlap with similar skills); not 4 because requests like "review X" or "validate Y" would plausibly match many competing skills.

3 / 5

Total

12

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (926 lines); consider splitting into references/ and linking

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
nyldn/claude-octopus
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.