CtrlK
BlogDocsLog inGet started
Tessl Logo

debug-cli

Use when users need to debug, modify, or extend the code-forge application's CLI commands, argument parsing, or CLI behavior. This includes adding new commands, fixing CLI bugs, updating command options, or troubleshooting CLI-related issues.

65

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable with executable commands throughout and a clear debug workflow including clone-before-reproduce safety practices. Its weaknesses are token efficiency (heavy duplication across six overlapping sections) and progressive disclosure (an existing smoke-test script is never referenced, and reference-worthy content is inlined).

Suggestions

Collapse the duplicated guidance: 'Core Principles', 'Common Testing Patterns', 'Integration with Development Workflow', 'Quick Reference', and 'Tips' repeat the same build/--help/-p/clone commands — consolidate into the single 'Workflow' section plus one short quick-reference block.

Reference the existing bundle: add a pointer to `scripts/test_cli.sh` (e.g. 'Smoke test: run `./scripts/test_cli.sh` after CLI changes') so the provided script is discoverable from SKILL.md.

Make verification steps concrete per workflow: replace 'Check output matches expectations' with explicit checks like `echo $?` for exit codes and a specific expected-output comparison in the 'Debugging a Bug Report' sequence.

DimensionReasoningScore

Conciseness

The same commands (cargo build, --help, -p, conversation dump/clone) are repeated across 'Core Principles', 'Workflow', 'Common Testing Patterns', 'Integration with Development Workflow', 'Quick Reference', and 'Tips' — several padded, redundant sections. It does not reach anchor 3 ('mostly efficient') because the duplication is pervasive rather than incidental, though it avoids explaining concepts Claude already knows, so it stays above anchor 1.

2 / 5

Actionability

Every section gives copy-paste-ready executable commands with real examples (e.g. `./target/debug/forge -p "create a hello world rust program"`, `forge conversation dump --html <id>`, `cat 2025-11-23_12-28-52-dump.json | jq '.messages[] | {role, content}'`) and clearly marked placeholders covering the common debug cases. It fully matches anchor 5 rather than anchor 4's 'minor gaps'.

5 / 5

Workflow Clarity

Sequences are clear and ordered (build → docs → test → verify; clone → build → reproduce → iterate) with feedback loops ('Iterate: Repeat until verified', 'Keep cloning the source until the fix is verified'). It sits below anchor 5 because validation checkpoints are described generically ('Check output matches expectations') rather than as explicit per-step commands — the exit-code check (`echo $?`) appears only in Tips.

4 / 5

Progressive Disclosure

Sections are well organized with clear headers, but the bundle's actual script (scripts/test_cli.sh, documented in scripts/README.md) is never referenced or linked from the body, and ~200 lines of quick-reference/tips content that could live in a references file are fully inlined. This matches anchor 3 ('some structure... content that should be separate is inline') better than anchor 4, since the one bundled resource is unsignaled.

3 / 5

Total

14

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is strong: it names a specific application and task domain, enumerates concrete capabilities, opens with an explicit 'Use when' trigger, and stays concise without fluff or over-claims. Only minor keyword synonyms (e.g. 'command-line', 'flags') are absent.

DimensionReasoningScore

Specificity

Quotes like "debug, modify, or extend the code-forge application's CLI commands, argument parsing, or CLI behavior" and "adding new commands, fixing CLI bugs, updating command options, or troubleshooting" list multiple concrete actions with comprehensive coverage of the CLI-development niche. It clearly exceeds anchor 4 ('minor gaps in coverage') since every common CLI task variation is enumerated.

5 / 5

Completeness

It explicitly answers 'when' with a leading "Use when users need to..." clause containing concrete trigger conditions, and answers 'what' by naming the specific actions covered (debug/modify/extend CLI commands, argument parsing, CLI behavior). It does not fit anchor 4, whose 'when' is weaker or less specific than this one.

5 / 5

Trigger Term Quality

Natural trigger phrases are present: "debug", "fixing CLI bugs", "adding new commands", "updating command options", "troubleshooting", "argument parsing". A few natural variations users might say are missing (e.g. "command-line", "flags", "--help"), so it falls just below the comprehensive-synonym coverage of anchor 5.

4 / 5

Distinctiveness Conflict Risk

The scope is pinned to a named application ("the code-forge application's CLI"), giving it a clear niche with distinct triggers and minimal conflict risk against generic debugging or general CLI skills. It is far more targeted than anchor 4's example, which spans multiple file formats.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
tailcallhq/forgecode
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.