CtrlK
BlogDocsLog inGet started
Tessl Logo

cli-anything

Use when the user wants Codex to build, refine, test, validate, or list CLI-Anything harnesses for GUI applications or source repositories. Adapts the full CLI-Anything methodology to Codex without changing the generated Python harness format.

86

1.55x
Quality

81%

Does it follow best practices?

Impact

95%

1.55x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-engineered content body that keeps SKILL.md as a navigation layer over a one-level-deep reference bundle, with concrete packaging/testing requirements and a graceful condensed fallback. Remaining gaps are small: a slightly redundant opening line, no sample harness invocation, and error-recovery loops that live only in the referenced specs.

DimensionReasoningScore

Conciseness

The body is dense and assumes Claude's competence — no concept explanations or filler — and the ~70-line condensed fallback is explicitly gated to when the full methodology is unavailable. Minor over-explanation remains: the opening line "Use this skill when the user wants Codex to act like the CLI-Anything builder" duplicates the frontmatter description, and the path-remapping table could be tightened. Anchor 4, not 5.

4 / 5

Actionability

Concrete guidance throughout: an exact generated-harness directory tree, specific packaging directives ("find namespace_packages(include=[\"cli_anything.*\"])", "console_scripts"), test requirements ("_resolve_cli()", "CLI_ANYTHING_FORCE_INSTALLED=1"), and real flags (--json, --dry-run). Not 5: no sample harness invocation is shown and items like "Copy and use the unified ReplSkin" name a resource without showing its use.

4 / 5

Workflow Clarity

The methodology-read step is a clearly ordered 6-step chain with explicit fallbacks, and every mode follows a 'Read <spec>, then do X' pattern with checkpoints such as "update TEST.md only with passing results" and a dedicated validation checklist mode. Not 5: error-recovery loops (validate → fix → re-run) are delegated to the referenced specs rather than stated in the body.

4 / 5

Progressive Disclosure

The body is a genuine overview: a Resource Map table signals every one-level-deep reference with its purpose, a remapping table resolves vendored paths, and the condensed rules keep the skill actionable even when references are unavailable (they are vendored at install time by scripts/install.sh). This matches anchor 5 — clear overview, well-signaled single-level references, easy navigation.

5 / 5

Total

17

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit and well-formed 'Use when...' trigger, five enumerated concrete actions, and a clear niche. Main weaknesses are reliance on the branded 'CLI-Anything' term to carry meaning and the absence of common synonyms like 'command-line' or 'CLI tool' in the trigger vocabulary.

DimensionReasoningScore

Specificity

Enumerates five concrete actions ("build, refine, test, validate, or list") with target scope ("GUI applications or source repositories") and a concrete constraint ("without changing the generated Python harness format"). Falls short of anchor 5 because the central deliverable — a 'CLI-Anything harness' — is characterized only by brand jargon, leaving minor coverage gaps.

4 / 5

Completeness

Explicitly answers both questions: the 'what' (build/refine/test/validate/list harnesses, adapting the methodology) and the 'when' via the concrete trigger clause "Use when the user wants Codex to build, refine, test, validate, or list CLI-Anything harnesses". This mirrors the anchor-5 exemplar pattern; the when-clause is fully explicit, so 4 does not fit.

5 / 5

Trigger Term Quality

Terms users would naturally say are present ("build", "test", "validate", "list", "GUI applications", "source repositories") and the branded "CLI-Anything" anchors recognition. Not 5: common synonyms such as "command-line interface" or "CLI tool" are missing, so a few natural trigger terms are absent.

4 / 5

Distinctiveness Conflict Risk

The "CLI-Anything harness" niche with its distinct triggers is mostly distinguishable, but broad verbs like "build", "test", and "validate" leave minor overlap risk with general build/test tooling skills. Anchor 4 ('mostly distinct; minor overlap risk with closely related skills') fits better than 5.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 26 missing, 14 deeper-than-1-level

Warning

Total

15

/

16

Passed

Repository
HKUDS/CLI-Anything
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.