CtrlK
BlogDocsLog inGet started
Tessl Logo

crabbox

Use the Crabbox wrapper for OpenClaw remote validation across Linux, macOS, Windows, and WSL2, including delegated Blacksmith Testbox proof. Report the actual provider and id.

56

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Medium

Suggest reviewing before use

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/crabbox/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is exceptionally actionable with concrete, executable commands and well-checkpointed workflows, but it is a ~700-line monolith with significant duplicated content and no progressive disclosure into reference files.

Suggestions

Split the monolith into one-level-deep reference files (e.g. references/observability-flags.md, references/macos-windows.md, references/webvnc.md, references/e2e-verification.md) and keep SKILL.md as a concise overview with clearly signaled links.

Deduplicate repeated content: merge the two Blacksmith footgun lists, consolidate the WebVNC guidance into one section, merge 'If Crabbox Fails' with 'Failure Triage', and factor the repeated CI=1 NODE_OPTIONS=... env prefix into a single example with a note to reuse it.

Trim sections that restate the same instruction in multiple places (e.g. 'Always report the actual provider and id' appears in First Checks and elsewhere) down to one authoritative statement.

DimensionReasoningScore

Conciseness

At ~700 lines the body is noticeably padded with duplication: Blacksmith footguns appear twice ('Run from repo root. The CLI syncs the current directory' and 'Raw commit SHAs are not reliable warmup --ref refs' recur in both 'Raw Blacksmith footguns' and 'Important Blacksmith footguns'), the identical CI=1 NODE_OPTIONS=... env prefix is repeated across ~6 command blocks, WebVNC guidance is duplicated between 'Interactive Desktop And WebVNC' and '### Interactive Desktop / WebVNC', and 'If Crabbox Fails' overlaps 'Failure Triage'. The material is non-obvious operational knowledge, but the redundancy goes beyond 'could be tightened', placing it below the 3 anchor.

2 / 5

Actionability

Fully executable throughout: copy-paste-ready commands with concrete flags (--provider blacksmith-testbox, --blacksmith-workflow .github/workflows/ci-check-testbox.yml), named JSON fields (leaseId, syncDelegated, commandPhases), expected id shapes (cbx_/tbx_), and exact validation commands covering the common cases. Placeholders like <cbx_id-or-slug> are appropriately scoped.

5 / 5

Workflow Clarity

Clear sequenced workflows with checkpoints: the numbered 'Efficient flow' (reproduce pre-fix symptom, patch locally, run one E2E command, record proof, ask for missing fields), 'First Checks' gating remote work, explicit cleanup verification, and feedback loops in triage ('rerun with --debug', 'follow next_action= hints'). Falls short of 5 because triage guidance is split across two overlapping sections with no single master decision path.

4 / 5

Progressive Disclosure

No bundle files exist (references/, scripts/, assets/ are all absent), so everything lives in one ~700-line monolith; content that clearly belongs in separate files is inlined (the 'Observability Flags' section is a long flag reference; the macOS/Windows, WebVNC, and E2E playbook sections each merit their own file). Consistent section headers keep it above the 'minimal structure' anchor 2, matching anchor 3.

3 / 5

Total

14

/

20

Passed

Description

65%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is distinctive and domain-specific with good trigger terms, but it omits any explicit 'Use when...' trigger guidance and lists only one genuinely concrete action, capping both completeness and specificity.

Suggestions

Add a 'Use when...' clause naming the situations that call for the skill (e.g. broad pnpm gates, CI-parity checks, Docker/E2E/package lanes, cross-platform proof), which would raise completeness from 3.

Enumerate 2-3 more concrete actions (e.g. 'warm reusable boxes, run CI-parity checks, capture logs and artifacts, stop leases') instead of the single generic 'remote validation'.

DimensionReasoningScore

Specificity

Names the domain ('Crabbox wrapper for OpenClaw remote validation across Linux, macOS, Windows, and WSL2') with one concrete action ('Report the actual provider and id'), but 'remote validation' is a single generic action rather than a list of specific capabilities, matching the '1-2 concrete actions, not comprehensive' anchor. It is not a 4 because no several specific actions are enumerated.

3 / 5

Completeness

The 'what' is clear (wrapper for remote validation across platforms including delegated Testbox proof plus a reporting rule), but there is no 'Use when...' clause or equivalent explicit trigger guidance, which caps completeness at 3 per the judging guidelines. Not a 4 because the 'when' is entirely absent rather than merely imprecise.

3 / 5

Trigger Term Quality

Natural terms a user needing this skill would say are present: 'Crabbox', 'OpenClaw', 'remote validation', 'Blacksmith Testbox', 'WSL2'. Falls short of comprehensive 5-level coverage because common phrasings like 'run tests remotely', 'CI proof', or 'warmup' variants are missing.

4 / 5

Distinctiveness Conflict Risk

Highly unique named tools ('Crabbox', 'OpenClaw', 'Blacksmith Testbox') define a clear niche with distinct triggers and minimal overlap risk with any other skill.

5 / 5

Total

15

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (712 lines); consider splitting into references/ and linking

Warning

referenced_paths_exist

Referenced path issues: 3 missing

Warning

Total

14

/

16

Passed

Repository
openclaw/gogcli
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.