CtrlK
BlogDocsLog inGet started
Tessl Logo

flow-deliver

Multi-AI validation, scoring, and review using available external providers (Double Diamond Deliver phase)

48

Quality

53%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/flow-deliver/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

52%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers genuinely concrete, executable orchestration commands and a well-gated step sequence, but it is roughly 3-4x longer than needed: the workflow is specified twice with conflicting numbering, and large example/report content is inlined rather than split into reference files. The skill would be dramatically stronger at ~200 lines with the worked example and supplements moved to references.

Suggestions

Delete the duplicate "Deliver Workflow - Deliver Phase" section (and its conflicting 5-step 'Implementation Instructions' restatement) — keep the EXECUTION CONTRACT as the single source of truth for the sequence.

Move the ~120-line Example 1 validation report and the PR-posting bash block into reference files (e.g. references/example-report.md, references/pr-posting.md) linked one level deep from the body.

Trim the three repeated banner templates to one parameterized template and cut the motivational/repeated 'MANDATORY / DO NOT PROCEED' boilerplate to a single statement of the gating rule.

DimensionReasoningScore

Conciseness

The body is ~875 lines with heavy padding and outright duplication: context detection and banner instructions appear twice (the EXECUTION CONTRACT Steps 1-2 and again under "Deliver Workflow - Deliver Phase"), banner templates are printed three times, and Example 1 embeds a ~120-line fully worked validation report. This is noticeably below the midpoint — several padded, removable sections — though it stops short of score 1, which is reserved for extensively explaining concepts Claude already knows rather than duplication.

2 / 5

Actionability

Most guidance is executable: exact script invocations ("${HOME}/.claude-octopus/plugin/scripts/orchestrate.sh deliver \"<user's validation request>\""), a runnable provider check, a complete validation-gate block that locates the results file, and a full PR-posting bash block. It falls short of 5 because some blocks are illustrative rather than executable (the spinner-progress prose, the fill-in report templates) and variables like $VALIDATION_FILE must be threaded across separate blocks; it is well above 3 since no pseudocode stands in for real commands.

4 / 5

Workflow Clarity

The EXECUTION CONTRACT lays out a clearly sequenced 7-step process with explicit validation checkpoints (Step 5 verifies the validation file exists, failure branches are enumerated in Error Handling, and a Validation Checklist closes the loop). It misses 5 because the duplicated second workflow section restates the steps with conflicting numbering ("Implementation Instructions" lists 5 steps vs. the contract's 7), which muddies the authoritative sequence rather than sharpening it.

4 / 5

Progressive Disclosure

No bundle files exist (references/, scripts/, assets/ are absent), and everything is inlined in one ~875-line SKILL.md — including the ~120-line example validation report, the dev-subtype supplement table, and the PR-posting script block, all of which clearly belong in separate reference files. There are section headers, so it rises above the no-structure example of score 2's low end, but the inlining of large self-contained chunks with no one-level-deep reference files fits 'content that clearly belongs in separate files is inlined' better than score 3's 'some structure, could be better organized'.

2 / 5

Total

12

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly states what the skill does within its plugin niche, but omits any 'use when' trigger guidance and leaves the validation target (code vs. documents) unstated. Trigger terms are relevant but thin on natural variations. Adding an explicit 'Use when...' clause and naming concrete targets would lift the two heaviest dimensions.

Suggestions

Add an explicit trigger clause, e.g. "Use when the user asks to validate, review, score, or audit completed work (code or documents) before shipping, or mentions the Deliver phase."

Name the concrete targets of validation (code implementations, PRs, documents) so the 'what' answers what is being validated, not just the verbs.

Include natural user phrasings/synonyms such as "code review", "quality check", "pre-ship validation", and "audit" to broaden trigger-term coverage.

DimensionReasoningScore

Specificity

The description names the domain ("Multi-AI validation") and a few actions ("validation, scoring, and review using available external providers"), but the actions are generic verbs with no stated objects — it never says what is validated or reviewed (code, documents, PRs). It matches 'names domain and 1-2 concrete actions, but not comprehensive' rather than score 4, whose example lists several specific actions with concrete targets; it is above score 2 because more than one action and the provider mechanism are named.

3 / 5

Completeness

The 'what' is stated ("Multi-AI validation, scoring, and review using available external providers") but the 'when' is missing — there is no 'Use when...' clause or equivalent trigger guidance; "Double Diamond Deliver phase" is plugin jargon, not an explicit trigger. Per the judging guideline, a missing 'Use when...' clause caps completeness at 3, which this squarely fits.

3 / 5

Trigger Term Quality

It contains relevant natural terms users would say — "validation", "review", "scoring" — but misses common variations and synonyms ("code review", "audit", "quality check", "validate implementation", file types). Not score 4: coverage lacks the broader natural phrasing users actually use; not score 2: more than one or two relevant keywords are present alongside the domain.

3 / 5

Distinctiveness Conflict Risk

"Multi-AI validation ... using available external providers (Double Diamond Deliver phase)" carves out a fairly distinct niche (multi-provider orchestration within this plugin's Double Diamond flow), with mostly minor overlap risk against generic single-model review skills. It is below score 5 because broad words like "review" and "validation" with no 'use when' scoping could still collide with plain review requests; it is above score 3 because the multi-AI/external-provider framing is more distinctive than a generic 'works with document files' style description.

4 / 5

Total

13

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (876 lines); consider splitting into references/ and linking

Warning

Total

15

/

16

Passed

Repository
nyldn/claude-octopus
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.