CtrlK
BlogDocsLog inGet started
Tessl Logo

build-test

Run the project's build / typecheck / lint / test commands and emit the build.passing + tests.passing signals devloop convergence reads.

61

Quality

72%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./plugins/_official/atoms/build-test/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a lean, well-sectioned spec with concrete commands, an exact output contract, and a genuine convergence feedback loop. Weaknesses are minor: unresolvable spec citations, an incomplete inline report schema, and a lint capability named in the description that the body never defines. No external references are needed, so progressive disclosure is fully appropriate.

Suggestions

Drop or contextualize the "Spec §20.2 / §22.4" and "critique-theater" framing — an executor gains nothing from references to sections it cannot read.

Flesh out the report schema (types for durationMs, commandsRun, failures) as a real JSON block instead of a one-line comment, and either define lint handling or remove it.

Add one explicit failure-path step (e.g. "On failing tests: read `failures[]`, return control to patch-edit rather than editing test files") to close the validation loop.

DimensionReasoningScore

Conciseness

The body is compact and information-dense — a framework command table, a terse output tree, and short signal definitions with no padding or explanations of concepts Claude already knows, matching "Efficient; minor instances of over-explanation that could be trimmed". Not a 5: the opening "Spec §20.2 / §22.4" citation and the "critique-theater" framing add context tokens that earn little for an executor.

4 / 5

Actionability

Concrete, executable specifics are present: `pnpm typecheck` / `pnpm test` defaults, the `od plugin run --input testCommand='pnpm test'` override, the exact JSON report shape, and a copy-paste devloop wiring block — matching "Mostly executable guidance; concrete code or commands with minor gaps". Not a 5 because the full report schema is sketched inline in a comment (e.g. failures: [...]) and lint — named in the description — has no command or signal defined.

4 / 5

Workflow Clarity

Sequencing and validation are explicit: the atom "always runs after `patch-edit`, and only when `plan.steps`'s current step is in `completed` state", and the devloop "repeat": true with "until": "(build.passing && tests.passing) || iterations >= 8" is a real feedback loop with a bounded retry — matching "Clear sequence with most checkpoints present; minor validation gaps". Not a 5: failure handling (what happens on failing tests, who reads `failures`) is described structurally but not as an explicit recovery sequence.

4 / 5

Progressive Disclosure

This is a simple, single-purpose skill (~70-line body) with well-organized sections (Inputs, Default commands, Output, Convergence, Anti-patterns, Status), no bundle files, and no inlined bulk that belongs in separate files — per the scoring notes, such skills can score 5 with just well-organized sections. References that do appear (`code/index.json`, `plan.md`, the atom implementation path) are one level deep and clearly signaled.

5 / 5

Total

17

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description states concrete capabilities in third person with good natural keywords (build, typecheck, lint, test). Its main weakness is the complete absence of a "Use when..." trigger clause, which caps completeness and limits discoverability. Adding explicit usage triggers and resolving the lint mismatch would lift it substantially.

Suggestions

Append a trigger clause such as "Use when the user asks to verify the build, run tests, or check whether a code migration still passes typecheck/tests."

Reconcile the mention of "lint" with the rest of the skill — the command table and output signals only cover typecheck and test, so either add a lint signal or drop the word.

Include natural synonyms users might say (e.g. CI, compile, test suite) to broaden trigger-term coverage beyond the devloop jargon.

DimensionReasoningScore

Specificity

Quotes: "Run the project's build / typecheck / lint / test commands" and "emit the build.passing + tests.passing signals devloop convergence reads" — it names several concrete actions (run commands, emit two named signals), matching the anchor "Lists several specific actions; minor gaps in coverage". It falls short of 5 because coverage is incomplete (lint is named but never reflected in the outputs, and the signals' semantics live in the body, not the description).

4 / 5

Completeness

The "what" is clear ("Run the project's build / typecheck / lint / test commands and emit the build.passing + tests.passing signals"), but there is no "Use when..." clause or equivalent trigger guidance, which per the judging guidelines caps completeness at 3 ("clear 'what' but 'when' is missing or only weakly implied"). It is not a 2 because the what is concrete, and not a 4 because no when-usage is stated at all.

3 / 5

Trigger Term Quality

Natural terms users would say include "build", "typecheck", "lint", and "test", giving good keyword coverage akin to the anchor example "PDF files, forms, document extraction". It misses a 5 because it offers no synonyms or variations (e.g. CI, compile, run tests) and leans on ecosystem jargon like "devloop convergence".

4 / 5

Distinctiveness Conflict Risk

The pairing of "build.passing + tests.passing signals devloop convergence reads" carves a clear niche tied to this devloop/atom ecosystem, matching "Mostly distinct; minor overlap risk with closely related skills". Minor overlap remains with generic build/test-runner skills, and "build-test" as a name is a common phrase.

4 / 5

Total

15

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
nexu-io/open-design
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.