CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-factory

Run a full build-and-ship pipeline from a spec — use for hands-off project generation

56

Quality

63%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/skill-factory/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

70%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strongly sequenced, highly actionable orchestration skill with explicit validation gates and feedback loops. Its main weaknesses are redundancy (duplicated prohibition/error-handling text), template placeholders instead of worked examples, and inlined script/banner content plus a reference to a non-existent file where progressive disclosure into bundle files is expected.

Suggestions

Consolidate the <HARD-GATE> block and the 'Prohibited Actions' section into one list, and drop the 'Error Handling (by step)' section that restates rules already stated inline in each step.

Move the Step 4.5 scenario-coverage script and the provider banner into scripts/ or references/ files referenced one level deep, keeping SKILL.md as a lean overview.

Fix or remove the reference to 'skills/blocks/codex-host-adapter.md' (no such file in the bundle) and remove the pinned 'v8.25.0' version from the footer, or place version notes in a dedicated section.

DimensionReasoningScore

Conciseness

The body is largely efficient with compact bash blocks, but the 'Prohibited Actions' section nearly duplicates the <HARD-GATE> block verbatim, 'Error Handling (by step)' restates rules already given inline, and the footer pins a version number ('Claude Octopus v8.25.0'), a time-sensitive detail that penalizes conciseness outside any 'old patterns' section.

3 / 5

Actionability

Concrete executable commands appear throughout ('orchestrate.sh factory --spec', 'state-manager.sh record_decision', the find-based report/score discovery), making the guidance mostly copy-paste ready. It falls short of 5 because the banner and result-presentation blocks are unfilled templates with <spec_path>/<status>/<verdict> placeholders rather than a worked example.

4 / 5

Workflow Clarity

The 8-step sequence is explicitly ordered with mandatory validation gates (spec-file check in Step 3, factory-report verification in Step 6), per-step error handling, and explicit feedback loops (FAIL verdict -> refine spec, re-run with --max-retries 2). This matches the anchor for clear sequence, explicit validation, and error-recovery loops; no destructive/batch cap applies since validation is present.

5 / 5

Progressive Disclosure

Sections are clearly headed, but no bundle exists: the body references 'skills/blocks/codex-host-adapter.md', a path not present in this skill, and long content that belongs in separate files (the Step 4.5 cross-provider QA script, the banner templates) is inlined in SKILL.md. This fits 'some structure but could be better organized' rather than the well-split anchor 4.

3 / 5

Total

15

/

20

Passed

Description

57%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is short, honest, and includes an explicit trigger clause, giving it adequate completeness and reasonable distinctiveness. Its weakness is low specificity — it abstracts the whole pipeline into one generic action — and sparse trigger-term coverage with no synonyms or spec-format keywords.

Suggestions

Add 2-3 concrete pipeline actions to the 'what' clause, e.g. 'Parses an NLSpec, generates test scenarios, runs multi-provider holdout evaluation, and scores results against a satisfaction target'.

Strengthen the 'when' clause with natural trigger phrases such as 'Use when the user provides an NLSpec or asks to build/generate a project hands-off from a spec'.

Include the spec-format keyword (NLSpec) users would actually mention, since it is the skill's primary input yet absent from the description.

DimensionReasoningScore

Specificity

'Run a full build-and-ship pipeline from a spec' names the domain but gives only a single generic action; none of the concrete pipeline capabilities (scenario generation, multi-provider holdout evaluation, satisfaction scoring, verdict reporting) are stated. This matches 'Names the domain but actions are minimal or generic' and is above the purely vague anchor 1.

2 / 5

Completeness

Both parts are explicit: the 'what' ('Run a full build-and-ship pipeline from a spec') and a 'use for...' trigger clause ('use for hands-off project generation'), so the missing-trigger cap of 3 does not apply. It is not 5 because the 'when' clause is thin and lacks concrete trigger phrases like 'Use when the user provides an NLSpec or asks to generate a project hands-off'.

4 / 5

Trigger Term Quality

The description includes some relevant keywords ('spec', 'build-and-ship', 'hands-off project generation') but misses natural variations users would say, such as 'NLSpec', 'generate/scaffold a project', or 'build from a spec'. Coverage is partial rather than good, placing it at anchor 3, not 4.

3 / 5

Distinctiveness Conflict Risk

'Build-and-ship pipeline from a spec' carves a fairly distinct niche with minor overlap risk against generic scaffolding or code-generation skills. It is not 5 because 'hands-off project generation' still overlaps broad project-scaffolding skills without sharply distinct triggers.

4 / 5

Total

13

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
nyldn/claude-octopus
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.