CtrlK
BlogDocsLog inGet started
Tessl Logo

cli-commands

MUST use when using the CLI, including debugging job failures and inspecting run history via `wmill job`.

52

Quality

66%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./system_prompts/auto-generated/skills/cli-commands/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

65%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exceptionally thorough and executable CLI reference — every command, flag, and non-obvious gotcha is documented in copy-paste-ready form. Its weaknesses are architectural: the whole reference is inlined in SKILL.md with no progressive disclosure or split files, destructive sync/instance push operations lack validate-first workflow guidance, and the sync pull/push flag duplication plus the repeated object-storage listing could be tightened.

Suggestions

Split the per-command reference into references/ files (e.g. references/commands.md or per-domain files) and keep SKILL.md as a concise overview plus the high-value non-obvious guidance (global options, key concepts, decision guides), loading detail only when needed.

Add a validate-then-apply workflow for destructive operations: e.g. 'wmill sync push --dry-run (or --lint) first, review the diff, then push --yes' — the flags exist but no workflow prescribes them.

Deduplicate the ~35 identical flags shared by `sync pull` and `sync push` (state shared flags once, e.g. 'both directions accept: --skip-*, --include-*, --keep-deleted...'), and have the Object Storage section link to its subcommands in the reference instead of re-listing them.

DimensionReasoningScore

Conciseness

Line-for-line the reference is dense and unpadded — flags, defaults, and gotchas Claude cannot know elsewhere (e.g. the $res/$var bare-string rule, the fork-parent resolution notes). However, `sync pull` and `sync push` duplicate ~35 near-identical flags verbatim, and the trailing Object Storage section re-documents subcommands already listed in the command reference. Fits 'mostly efficient but could be tightened'; not a 4 given the sizeable duplicated blocks.

3 / 5

Actionability

Complete, copy-paste-ready command grammar: every subcommand with argument types, flags, defaults, and exact value formats (e.g. 'SCRIPT:PARAM=VALUE', '--upload SCRIPT[:PARAM]=SOURCE', '@<filename>' stdin conventions, port/behavior defaults). A user can construct any invocation directly from the text.

5 / 5

Workflow Clarity

This is a lookup reference, so sequence is mostly per-command, and the Object Storage section adds a genuinely useful 'Choosing a subcommand' decision guide. But destructive/batch operations exist (`sync push` 'overrides any remote versions', `instance push` 'overwrite remote', `workspace merge` deploys, `object-storage delete`) with no prescribed validate-then-apply loop even though --dry-run/--diff flags exist — the rubric's cap for destructive/batch operations without validation applies. Not a 4: no checkpoint guidance anywhere for the overwriting commands.

3 / 5

Progressive Disclosure

No bundle files exist (no references/, scripts/, or assets/), so the entire ~800-line command reference is inlined in SKILL.md rather than split into reference files behind a concise overview — 'content that should be separate is inline'. It stays at 3 rather than 2 because internal structure is good: per-command headers, consistent formatting, and the closing Key-concepts/Choosing-a-subcommand section model the right pattern; navigation is easy even though nothing is offloaded.

3 / 5

Total

14

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description has a strong, explicit trigger clause but undersells the skill: it reads as a job-debugging trigger rather than the comprehensive wmill command reference the body actually is. Both the capability list and the trigger terms cover only a sliver of the skill's surface.

Suggestions

State the 'what' explicitly, e.g. 'Reference for the wmill CLI: managing scripts, flows, apps, datatables, resources, sync, workspaces, and jobs' so the description reflects the skill's full scope.

Broaden trigger terms beyond job debugging to the natural phrases users say for the other commands: 'sync', 'push or deploy a script/flow/app', 'run a flow or script', 'workspace', 'schedule', 'token'.

Tighten the trigger so 'the CLI' clearly means the wmill CLI (e.g. 'MUST use when running wmill CLI commands') to reduce false triggers on unrelated CLI work.

DimensionReasoningScore

Specificity

Names the domain (the `wmill` CLI) and 1-2 concrete actions ("debugging job failures", "inspect run history via `wmill job`"), but omits nearly all of what the skill actually covers (scripts, flows, apps, sync, datatables, workspaces, tokens). Not a 4: the listed actions are a small fraction of the skill's scope, so coverage is not 'several specific actions with minor gaps'.

3 / 5

Completeness

The 'when' is explicit and strong ("MUST use when using the CLI"), but the 'what' is only weakly implied: the description never states the skill is a reference for wmill commands covering scripts, flows, apps, and sync. Not a 4 because the 'what' is thinner than the 'when'; not a 2 because the job-related examples do gesture at supported capabilities.

3 / 5

Trigger Term Quality

Contains well-anchored terms a user would naturally say for job work ("wmill", "job failures", "run history", "debugging"), but misses the natural phrases for the rest of the skill's surface: sync, push/deploy, script, flow, workspace, schedule, token. Fits 'some relevant keywords but missing common variations or synonyms' rather than the good coverage of a 4.

3 / 5

Distinctiveness Conflict Risk

'`wmill job`' and 'the CLI' in a Windmill context anchor this to a clear niche with minimal conflict risk. Not a 5: 'MUST use when using the CLI' is broad enough that generic CLI work in the environment could trip the trigger even when unrelated to wmill.

4 / 5

Total

13

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (867 lines); consider splitting into references/ and linking

Warning

referenced_paths_exist

Referenced path issues: 2 missing, 2 deeper-than-1-level

Warning

Total

14

/

16

Passed

Repository
windmill-labs/windmill
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.