CtrlK
BlogDocsLog inGet started
Tessl Logo

flow-deliver

Multi-AI validation, scoring, and review using available external providers (Double Diamond Deliver phase)

40

Quality

38%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/flow-deliver/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

27%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is actionable and richly gated with validation steps, but it is severely verbose, structurally duplicated with conflicting numbering, and monolithic with no bundle-based progressive disclosure. The unresolved template placeholders suggest an incomplete generation pass.

Suggestions

Collapse the three restatements of the workflow (EXECUTION CONTRACT, 'Deliver Workflow', 'Implementation Instructions') into one canonical numbered sequence; remove the duplicated banners.

Move the 130-line example validation report and the dev-subtype validation-supplement table into separate reference files under references/ and link to them one level deep, instead of inlining everything.

Resolve the template placeholders ({{PREAMBLE}}, {{VISUAL_INDICATORS}}, {{QUALITY_GATES}}) into real content or delete them, and trim emoji decoration and inline version notes (v7.16.0+, v8.44.0, v8.49.0) to cut token weight.

DimensionReasoningScore

Conciseness

The ~855-line body is heavily padded: the EXECUTION CONTRACT (Steps 1-7) is restated again as 'Deliver Workflow' (Steps 1-5) and again as 'Implementation Instructions', banners are shown 3+ times, and a 130-line illustrative report with 'XX/100' placeholders is inlined; far more tokens than the task warrants.

1 / 3

Actionability

It provides concrete, executable bash commands (orchestrate.sh, state-manager.sh, gh pr comment), but unresolved template placeholders ({{PREAMBLE}}, {{VISUAL_INDICATORS}}, {{QUALITY_GATES}}, ${CLAUDE_SESSION_ID}) and illustrative 'XX/100' report content keep it short of copy-paste ready.

2 / 3

Workflow Clarity

Validation checkpoints and feedback loops are explicit and strong (Step 5 file-existence gate, error-handling section, 'DO NOT PROCEED' gates), but three conflicting step-numbering schemes for the same sequence make the actual workflow hard to follow, capping clarity at 2.

2 / 3

Progressive Disclosure

No bundle files exist (references/, scripts/, assets/ absent) and the skill is a single monolithic 855-line document with unresolved {{...}} placeholders pointing to content that is not present; there is no file splitting or one-level-deep reference structure.

1 / 3

Total

6

/

12

Passed

Description

50%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description states a clear purpose and domain but omits any explicit 'Use when...' trigger guidance and leans on process jargon ('Double Diamond Deliver phase') over natural user phrasing. It is adequate but generic, scoring 2 across all dimensions.

Suggestions

Add an explicit 'Use when ...' clause naming natural triggers (e.g., 'Use when the user asks for code review, PR validation, a security audit, or a ship-readiness check').

Replace jargon ('Double Diamond Deliver phase', 'external providers') with concrete user-facing terms ('code review', 'validate endpoints', 'multi-AI review') to improve trigger-term quality and distinctiveness.

Tighten the action list to specific concrete operations (e.g., 'runs Codex, Gemini, and Claude in parallel and synthesizes a scored validation report') to lift specificity from 2 to 3.

DimensionReasoningScore

Specificity

It names the domain ('Multi-AI validation, scoring, and review') and several actions, but the actions ('validation, scoring, and review') are generic rather than the concrete, granular operations the top anchor calls for; it does not reach 'multiple specific concrete actions'.

2 / 3

Completeness

It clearly states what the skill does, but there is no explicit 'Use when...' trigger clause; the 'when' is only implied by the phase name, which the rubric caps at 2.

2 / 3

Trigger Term Quality

Terms like 'validation', 'scoring', and 'review' are user-relevant, but 'Double Diamond Deliver phase' and 'external providers' are jargon, and common variations users actually say ('code review', 'PR validation', 'security audit', 'ship-ready check') are absent.

2 / 3

Distinctiveness Conflict Risk

The 'Double Diamond Deliver phase' framing gives it a niche, but 'validation, scoring, and review' is broad enough to overlap with other review/quality skills, and there are no explicit distinguishing triggers.

2 / 3

Total

8

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (856 lines); consider splitting into references/ and linking

Warning

Total

15

/

16

Passed

Repository
nyldn/claude-octopus
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.