CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/smoke-suite-gate

Build-an-X workflow for a critical-path smoke suite that runs in <5 minutes - picks the 5-15 highest-business-value journeys (login, hero flow, checkout, payment, primary read), implements as fast E2E or API tests, gates per-deploy, retries on transient failures with quarantine. Use as the canary-precursor or per-deploy verification gate; the team's "if this fails, the build can't proceed" floor.

68

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Overview
Quality
Evals
Security
Files

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable body: concrete budgets, executable Playwright and CI examples, an explicit retry/quarantine failure-handling loop, and useful anti-pattern and limitations sections. All four dimensions land at 4 rather than 5 due to small trims available, the missing deploy-ephemeral.sh and seeding guidance, a quarantine step that blocks rather than quarantines, and a ~200-line single-file layout that could use reference files.

Suggestions

Provide or scaffold 'scripts/deploy-ephemeral.sh' and the test-account/SKU seeding steps it depends on, so the CI example is fully executable end-to-end.

Make the 'Quarantine repeat failure' step actually quarantine (e.g., move the spec to a quarantine dir or label the run) instead of only exiting 1, matching the Step 4 prose.

Move the CI workflow YAML and the full Playwright spec into references/ files, keeping SKILL.md as the lean overview with well-signaled one-level-deep pointers.

DimensionReasoningScore

Conciseness

Dense tables, concrete budgets, and no explaining of concepts Claude already knows; a few phrases repeat earlier points ('Per stage, smoke acts as the "is this deploy worth proceeding with" gate' restates the Overview) and some Limitations prose could be tightened — efficient with minor trims, not anchor-5 lean.

4 / 5

Actionability

Copy-paste-ready Playwright spec, full CI workflow YAML, and concrete budgets (30-60s per test, '--retries=2', 'timeout-minutes: 10'), but 'scripts/deploy-ephemeral.sh' is invoked without being provided and secrets/test-data seeding is never set up — minor gaps versus anchor 5.

4 / 5

Workflow Clarity

Steps 1-6 are clearly sequenced with an explicit error-recovery loop (retry → real regression blocks deploy vs. flaky test quarantined via 'flaky-test-quarantine') and a hard timeout cap; however the 'Quarantine repeat failure' CI step only exits 1 rather than actually quarantining, a minor validation gap versus anchor 5.

4 / 5

Progressive Disclosure

No bundle files exist; the single SKILL.md is well-organized with clear section headers, a table-driven structure, and a References section for related skills. At ~200 lines it could offload the CI workflow and Playwright spec to references/ files, so it sits at 'good structure, minor organization gaps' rather than anchor 5.

4 / 5

Total

16

/

20

Passed

Description

91%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states what the skill builds and how (journey selection, E2E/API implementation, per-deploy gating, retry/quarantine) and gives explicit 'Use as...' trigger guidance, with concrete domain keywords throughout. The only weaknesses are the vague 'Build-an-X workflow' opener and slight overlap risk with general E2E testing skills.

DimensionReasoningScore

Specificity

Lists several concrete actions — 'picks the 5-15 highest-business-value journeys', 'implements as fast E2E or API tests', 'gates per-deploy, retries on transient failures with quarantine' — but the opener 'Build-an-X workflow' is template-like and unexplained, a minor gap versus the comprehensive-anchor 5.

4 / 5

Completeness

Explicitly answers both: what it does ('picks... 5-15 journeys', 'implements as fast E2E or API tests', 'gates per-deploy', 'retries... with quarantine') and when to use it ('Use as the canary-precursor or per-deploy verification gate') with concrete trigger phrases.

5 / 5

Trigger Term Quality

Covers the natural terms users would say in this domain — 'smoke suite', 'E2E', 'API tests', 'canary', 'per-deploy verification gate', plus journey names like 'login, hero flow, checkout, payment' — comprehensive coverage including synonyms.

5 / 5

Distinctiveness Conflict Risk

Clear niche — the fast, narrow, gating subset framed as 'the team's "if this fails, the build can't proceed" floor' — but 'implements as fast E2E or API tests' leaves minor overlap risk with a general E2E-testing skill.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Reviewed

Table of Contents