CtrlK
BlogDocsLog inGet started
Tessl Logo

crabbox

Crabbox/Testbox remote proof: portable provider routing, untrusted isolation, Linux/macOS/Windows/WSL2, live E2E, diagnostics, cleanup.

65

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is crabbox in openclaw/openclaw

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exceptionally lean, command-dense skill with concrete executable guidance, real validation checkpoints, and error-recovery loops. The main structural weakness is that it is a monolithic single file: provider-specific and cross-OS detail would be better split into one-level-deep reference files.

Suggestions

Split provider-specific lanes (Untrusted AWS, Desktop/Cross-OS) into references/ files (e.g. references/untrusted-aws.md, references/desktop-cross-os.md) and signal them from short sections in SKILL.md.

State the top-level order once explicitly (Route → Preflight → Run → Record → Cleanup) so the pipeline is not reconstructed from rule fragments.

DimensionReasoningScore

Conciseness

Telegraphic rule style ("No speculative warmup. Acquire when first heavy command ready. Reuse id. Stop.") with zero padding and no explanation of concepts Claude already knows; every token carries operational content. Lean throughout a dense 365-line skill.

5 / 5

Actionability

Copy-paste shell with exact flags throughout, including the full untrusted-AWS sequence ("env -u CRABBOX_AWS_INSTANCE_PROFILE ... warmup --provider aws --network public ... --fresh-pr <owner/repo#number> --no-hydrate"). Placeholders like <check-command> are explicitly governed by the Repository Contract section, a justified flexibility rather than pseudocode.

5 / 5

Workflow Clarity

Clear routing rules, preflight verification, explicit validation checkpoints (jq -e guards, post-warmup inspect, status --id --wait) and recovery loops ("retry --full-resync once. Still bad: fresh lease"). Below 5 because the top-level pipeline must be assembled from rule fragments across sections rather than one explicit ordered sequence with numbered checkpoints.

4 / 5

Progressive Disclosure

Section headers are clear, but the whole skill is one inlined file: per-provider detail (Untrusted AWS, Desktop/Cross-OS) that belongs in separate reference files sits inline in a 365-line body, and no bundle files exist to offload it. Anchor 3: structure present, but content that should be separate is inline.

3 / 5

Total

17

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A dense, tool-specific description that names real capabilities and is highly distinctive, but it is written as a keyword fragment list with no explicit trigger guidance. Adding a "Use when..." clause would lift the weakest dimension.

Suggestions

Append an explicit trigger clause, e.g. "Use when the user asks to run validation on a remote/clean machine, mentions Crabbox/Testbox, or needs proof on Windows/WSL2/macOS."

Rewrite the capability list as third-person verbs ("Routes work across trusted and untrusted providers, isolates untrusted PRs, runs live E2E checks") so the what reads as actions rather than labels.

Add natural synonyms users would actually say — "clean machine", "remote CI", "AWS run" — to improve trigger term coverage.

DimensionReasoningScore

Specificity

Names several concrete capabilities — "portable provider routing", "untrusted isolation", "live E2E", "diagnostics, cleanup" — covering the domain's main actions. Falls short of 5 because the telegraphic noun-fragment style leaves the actions underdeveloped rather than comprehensively described.

4 / 5

Completeness

The "what" is stated (remote proof, routing, isolation, diagnostics, cleanup) but there is no "Use when..." clause or equivalent trigger guidance, which caps completeness at 3 per the rubric guideline. Not 2 because the "what" is substantive and multi-part.

3 / 5

Trigger Term Quality

Exact tool names ("Crabbox/Testbox") plus OS targets ("Linux/macOS/Windows/WSL2") and "E2E" give good keyword coverage. Missing a few natural phrasings a user would say, such as "clean machine", "remote CI", or "AWS test run".

4 / 5

Distinctiveness Conflict Risk

"Crabbox/Testbox" are unique product names no other skill would claim, giving a clear niche with distinct triggers and minimal conflict risk.

5 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
openclaw/acpx
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.