CtrlK
BlogDocsLog inGet started
Tessl Logo

delivery-flow

Carries a code change through to a tested, verified and reviewed local change, loading the stage skill for each step: planning, test-first implementation, debugging, verification and review. Use when a request asks for code to change, whether it arrives as an approved ticket or a one-line message: a bug fix, a failing test or crash, a small feature or flag, a behavior change, or review comments to fix. Not for questions about the code, for merging or deleting a finished branch, or for deciding which review comments to accept without changing code. Use it whenever the work will change code, even when the user does not mention tests, review or a pull request.

76

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, well-structured router body: zero padding, an explicit stage sequence with a real review-feedback loop, and clean one-level skill references. The only meaningful gap is that a few routing boundaries rest on undefined judgment terms rather than concrete criteria or examples.

Suggestions

Define or give a one-line example for the borderline gating terms — 'multiple meaningful implementation units', 'a missing decision that materially changes scope, behavior, or safety' — so the routing decision is mechanical rather than judgmental.

Add a brief tie-breaking rule for when a task is both a bug fix and a small feature, since steps 2 and 3 (TDD vs systematic-debugging) can both claim the same request.

Give one concrete example of 'valid findings' vs. dismissible findings in step 5 to make the review-fix feedback loop's exit condition unambiguous.

DimensionReasoningScore

Conciseness

Every line is directive and earns its place — no concept is explained that Claude doesn't already know; the only near-explanatory line ('Mentioning a skill or paraphrasing it is not equivalent to loading it') is a real behavioral correction, not padding. The 25-line body matches the lean anchor and cannot be a 4 because there are no over-explanatory passages to trim.

5 / 5

Actionability

Guidance is concrete and executable for an instruction-only router: each numbered rule names the exact skill to load and its trigger condition ('Load `finishing-a-development-branch` only when the user explicitly asks to integrate, push, or create a pull request'), and the routing contract gives checkable rules ('Never push, merge, open a pull request, or delete a worktree without explicit authorization'). It falls short of the fully-executable 5 anchor because several gating conditions remain judgment calls — 'when the task spans multiple meaningful implementation units', 'a missing decision that materially changes scope, behavior, or safety', 'If it returns valid findings' — with no worked example or tie-breaking rule for borderline cases.

4 / 5

Workflow Clarity

The six numbered stages give an explicit sequence with load conditions, step 4 places verification before any completion claim, and step 5 contains a genuine feedback loop ('If it returns valid findings, load `receiving-code-review`, address them, and return to verification'). This matches the top anchor's explicit-validation-plus-error-recovery-loop pattern; it is not a 4 because no checkpoint in the routed sequence is implicit.

5 / 5

Progressive Disclosure

No bundle files exist (references/, scripts/, assets/ are absent), and the named skills (`writing-plans`, `systematic-debugging`, etc.) are installed stage skills this router dispatches to, not nested bundle references — so there is exactly one level of indirection, clearly signaled under 'Required stage loading'. Per the simple-skill guidance, a well-organized 25-line skill with no need for external files scores 5 on section organization alone; the two headers plus numbered list make navigation trivial.

5 / 5

Total

19

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary router description: third-person, concrete about both capability and trigger, with explicit positive and negative triggering guidance. The only residual risk is the inherent breadth of 'any code change' overlapping with the specialist skills it loads.

DimensionReasoningScore

Specificity

The description enumerates concrete, sequenced capabilities — 'planning, test-first implementation, debugging, verification and review' — and a concrete end state ('a tested, verified and reviewed local change'), giving comprehensive coverage of what the skill does. It is not below 5 because no stage of the routed work is left abstract; nothing above 5 exists on the scale.

5 / 5

Completeness

Both 'what' (first sentence: carries a code change through named stages to a reviewed local change) and 'when' are explicit, with concrete trigger phrases ('Use when a request asks for code to change') and reinforced triggering ('Use it whenever the work will change code'). It exceeds the 4 anchor because the 'when' clause is specific rather than generic, and it adds negative scope ('Not for questions about the code, for merging or deleting a finished branch') that sharpens both answers.

5 / 5

Trigger Term Quality

It covers natural user phrasings with synonyms across request shapes: 'a bug fix, a failing test or crash, a small feature or flag, a behavior change, or review comments to fix', plus 'approved ticket or a one-line message' and 'even when the user does not mention tests, review or a pull request'. A user asking for any of these would naturally surface the skill, matching the comprehensive-synonyms anchor; it is not a 4 because no common variation of a code-change request is obviously missing.

5 / 5

Distinctiveness Conflict Risk

The routed stage list (planning, TDD, debugging, verification, review) carves out a distinct orchestrator niche, and the explicit exclusions ('Not for questions about the code, for merging or deleting a finished branch, or for deciding which review comments to accept without changing code') remove the most likely mis-triggers. It is not a 5 because the domain 'any request that asks for code to change' is inherently broad and could overlap with the individual stage skills it routes to or a general coding skill.

4 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
tesslio/tessl-eval-demo
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.