CtrlK
BlogDocsLog inGet started
Tessl Logo

builder-smoke-test

Smoke test the Agent Builder feature branch end-to-end against a hermetic project scaffolded by the skill (linked to the current worktree). Covers workspace reconciliation, stored agents/skills CRUD, ownership, visibility, stars, registry/library Copy flow, picker allowlists, model policy, RBAC role gating, role impersonation UI, builder defaults, infrastructure diagnostics, channels, and Studio + Agent Builder UI. Trigger when validating the agent-builder feature branch, PRs that touch packages/server, packages/playground, packages/playground-ui agent-builder routes, or builder EE code paths.

72

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A high-quality orchestration body for a complex QA skill: fully executable commands, a canonical ordering with rationale, dense validation feedback loops, and a well-signaled one-level-deep reference bundle that all exists on disk. The main improvement areas are minor: some duplicated content (env-override behavior, cookie warnings) and a few multi-page operational procedures inlined in the overview that could be pushed into the per-section references.

Suggestions

Deduplicate the 'mastra dev overwrites process.env from .env' behavior, which is explained in 'How mastra dev reads env (important)' and repeated nearly verbatim in 'Known rough edges' — keep one authoritative statement and cross-link it.

Move the detailed 'Extracting the session cookie for curl (auth on)' procedure and the env-var resolution ladder into references/auth.md and references/setup.md respectively, keeping a one-line pointer plus the failure-mode warning in SKILL.md so the overview stays lean.

DimensionReasoningScore

Conciseness

Nearly every section carries non-obvious, project-specific operational knowledge (env-resolution ladder, error-code remediation table, .env ownership policy, port-bump behavior) — no padding explaining concepts Claude already knows. It misses a 5 because of duplication: the .env/process.env override behavior appears in both 'How mastra dev reads env' and 'Known rough edges', the cookie-extraction warning appears in the execution flow and again as a full section, and the scope/section mappings are effectively tabulated twice.

4 / 5

Actionability

Fully executable throughout: exact invocations ('bash .claude/skills/builder-smoke-test/scripts/preflight.sh --expect off --openai-key "$OPENAI_API_KEY"'), copy-paste curl patterns, a per-error-code remediation table, a concrete cookie-extraction procedure with a 404 troubleshooting loop, and a filled-in report template. Commands and expected outputs are specific enough to run verbatim.

5 / 5

Workflow Clarity

The multi-step process is exceptionally sequenced: a numbered execution flow, a canonical section order with rationale, preflight gating before every later section, an error-code table defining fix-and-retry loops (e.g. 'scaffold-failed' → re-run with --no-reuse, inspect output), a role-mismatch hard stop, and a 'Verify before filing' checklist that mandates re-confirmation before reporting issues. Validation checkpoints and feedback loops are explicit and everywhere, far beyond the anchor-4 'minor validation gaps' level.

5 / 5

Progressive Disclosure

Excellent bundle structure: a 15-row section table maps each test area to a real one-level-deep reference file (all 15 references/*.md and 4 scripts/*.sh verified to exist), plus explicit 'Required vs optional reference tiers' and a References index. It falls just short of anchor 5 because the SKILL.md body itself inlines sizable operational how-to content (the ~20-line cookie-extraction procedure, the env-var resolution ladder, and the error-code table) that could live in references/setup.md or references/auth.md to keep the overview leaner.

4 / 5

Total

18

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states precisely what the skill does, enumerates the tested surface concretely, and gives an explicit 'Trigger when' clause with repo-specific conditions that distinguish it from the release-time smoke test. The only minor gap is a lack of synonym phrasings for the QA/validate intent.

DimensionReasoningScore

Specificity

The description enumerates many concrete capabilities — "workspace reconciliation, stored agents/skills CRUD, ownership, visibility, stars, registry/library Copy flow, picker allowlists, model policy, RBAC role gating, role impersonation UI, builder defaults, infrastructure diagnostics, channels" — comprehensive and specific rather than generic. Anchor 4 is ruled out because coverage is not merely 'several actions with minor gaps' but a full enumeration of the tested surface.

5 / 5

Completeness

Both halves are explicit: the 'what' ("Smoke test the Agent Builder feature branch end-to-end against a hermetic project scaffolded by the skill... Covers [full surface list]") and the 'when' ("Trigger when validating the agent-builder feature branch, PRs that touch..."). Matches the anchor-5 example pattern of concrete what + explicit trigger clause.

5 / 5

Trigger Term Quality

Triggers include natural developer phrasing: "Trigger when validating the agent-builder feature branch, PRs that touch packages/server, packages/playground, packages/playground-ui agent-builder routes, or builder EE code paths" — good coverage of the phrases a repo developer would say. Not a 5 because it lacks synonyms for the QA intent (e.g. 'QA the branch', 'regression-test', 'end-to-end test the PR') beyond the smoke/validate family; not a 3 because the terms present are the natural ones for this audience, not jargon-only.

4 / 5

Distinctiveness Conflict Risk

It carves a clear niche ("Smoke test the Agent Builder feature branch end-to-end against a hermetic project") with repo-specific triggers (named packages and routes), and even distinguishes itself from the adjacent skill ("for branch QA — it complements the release-time mastra-smoke-test"). Minimal conflict risk; nothing generic that would fire for unrelated skills.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (519 lines); consider splitting into references/ and linking

Warning

Total

15

/

16

Passed

Repository
mastra-ai/mastra
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.