CtrlK
BlogDocsLog inGet started
Tessl Logo

test-all

Parallel test orchestrator. Runs all 9 test suites concurrently via Task sub-agents and the iwsdk CLI. Handles build, example setup, dev servers, agent launch, polling, retries, and result aggregation.

56

Quality

64%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/test-all/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

70%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong orchestrator document: fully phased workflow with explicit validation, retries, timeouts, and troubleshooting — the workflow clarity is exemplary for a batch operation. The weaknesses are repetition that inflates token cost, a few questionable Task-tool parameters, and a monolithic structure with no progressive disclosure into reference files.

Suggestions

State the dynamic-port design decision once (Phase 3 is the natural home) and cut the repeats in the Test Map note and 'Key Design Decisions'; move the "Boolean values must be JSON booleans" note to the test sub-skills where ecs_set_component is actually used.

Move the sub-agent prompt template, Troubleshooting, and Key Design Decisions sections into a references/ file (e.g. references/sub-agent-template.md) linked one level deep, keeping SKILL.md as a lean phase-by-phase overview.

Correct the Task launch parameters (`subagent_type: "Bash"` is not a standard subagent type) and replace the illustrative `agents = {...}` block with the concrete fields to record for each agent.

DimensionReasoningScore

Conciseness

The body is mostly efficient (command blocks, a compact test map table, direct instructions), but includes unnecessary repetition: the ports-are-not-pre-assigned caveat appears three times (Test Map note, Phase 3, and "Runtime-first server discovery"), the sub-agents-read-skill-files point is stated twice, and the "Boolean values must be JSON booleans" design note belongs in the test sub-skills, not this orchestrator. Not 2 because most sections are tight and earn their tokens.

3 / 5

Actionability

Every phase gives concrete, runnable commands (node -v, pnpm build:tgz, node scripts/test-prep.mjs clone/install, node scripts/test-servers.mjs start/ports/stop), a fill-in sub-agent prompt template, specific timeouts (60s servers, 20min hard cap), and log paths. It falls short of fully executable because `subagent_type: "Bash"` and `mode: "bypassPermissions"` are questionable/non-standard Task parameters and the `agents = {...}` tracking block is illustrative pseudocode rather than a runnable artifact.

4 / 5

Workflow Clarity

Phases 1–7 are clearly sequenced with explicit validation checkpoints: "Stop on failure" on prerequisites/build, a 60-second server-readiness check with exit code and log paths, retry policy distinguishing transient from assertion failures with a max-1-retry limit, a 20-minute hard timeout with a defined fallback, and a troubleshooting section with recovery loops. This matches the top anchor (clear sequence, explicit validation, feedback loops for error recovery) — important given this is a batch operation across 9 agents.

5 / 5

Progressive Disclosure

Section headers and phase structure give good navigation, but this is a ~250-line monolithic file with no bundle files at all (no references/, scripts/, or assets/): the sub-agent prompt template, troubleshooting scenarios, and "Key Design Decisions" rationale are inlined where a well-signaled reference file would reduce context load. It matches anchor 3 (some structure, content that should be separate is inline) rather than 4, which would require most content appropriately split.

3 / 5

Total

15

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description communicates a concrete, well-scoped 'what' with specific lifecycle actions, but entirely lacks a 'when to use' trigger clause and natural user phrasing. It also contains a minor internal inconsistency ('polling') versus the body's no-polling design. Adding an explicit "Use when..." clause with natural trigger phrases would resolve the two lowest-scoring dimensions at once.

Suggestions

Append an explicit trigger clause, e.g. "Use when the user asks to run all tests, the full IWSDK test suite, or a complete regression run" — this addresses both the missing 'when' (completeness) and natural trigger terms (trigger_term_quality).

Replace 'polling' in the action list with 'completion notifications' to match the body's Phase 5 design ("Do NOT poll"), removing the description/body contradiction.

Add one or two natural synonyms ("run all tests", "full test run") so the description matches how users actually phrase the request, and to better distinguish it from the individual test-* skills.

DimensionReasoningScore

Specificity

The description lists several concrete lifecycle actions ("build, example setup, dev servers, agent launch, polling, retries, and result aggregation"), matching the 'several specific actions; minor gaps' anchor. It falls short of 5 because the actions are lifecycle stage names rather than fully concrete capabilities, and 'polling' contradicts the body's explicit "Do NOT poll, do NOT run bash commands to check output files" instruction.

4 / 5

Completeness

There is a clear 'what' ("Runs all 9 test suites concurrently via Task sub-agents and the iwsdk CLI"), but no "Use when..." clause or equivalent explicit trigger guidance exists. Per the judging guidelines, a missing 'Use when' clause caps completeness at 3; not lower because the 'what' is concrete and unambiguous.

3 / 5

Trigger Term Quality

Relevant keywords are present ("test suites", "parallel", "orchestrator", "iwsdk"), but common natural variations a user would actually say ("run all tests", "full test run", "regression run", "run the test suite") are missing. Not 4 because keyword coverage is domain-specific jargon rather than natural user phrasing.

3 / 5

Distinctiveness Conflict Risk

The niche is fairly distinct ("9 test suites", "iwsdk CLI", "Task sub-agents"), but the body reveals sibling skills (test-interactions/SKILL.md, test-ecs-core/SKILL.md, etc.), so a generic "run the tests" request could overlap with those individual test skills — minor overlap risk with closely related skills, matching anchor 4 rather than 5.

4 / 5

Total

14

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 10 missing

Warning

Total

14

/

16

Passed

Repository
facebook/immersive-web-sdk
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.