CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/e2e-suite-budget

Caps E2E suite size by computing per-test ROI - (regressions caught × value) ÷ (runtime × flake rate × maintenance) - then ranks every end-to-end test and recommends which bottom-decile ones to retire, move to a lower layer, or fix. Use when CI is slow or E2E-dominated, flaky failures are rising, or quarterly to keep suite size within maintenance capacity. For strategic unit:service:UI layer ratios use test-pyramid-balancer, for the minimal per-deploy critical-path gate use smoke-suite-gate, and for quarantining flaky tests use flaky-test-quarantine; this prunes low-signal tests by ROI.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Overview
Quality
Evals
Security
Files

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, well-structured workflow body: executable scoring code, explicit input validation with feedback loops, concrete output and config templates, a worked example, and a properly delegated anti-patterns reference file. The only deductions are minor verbosity — a duplicated step summary, an educational test-pyramid quote, and a slightly malformed References section.

DimensionReasoningScore

Conciseness

The body is largely lean — formula, executable code, output templates, and a worked example with no filler. Minor over-explanation keeps it off anchor 5: the References section quotes Cohn explaining the test pyramid ("brittle, expensive to write, and time consuming to run"), a concept Claude already knows; the "How to use" list restates all seven steps that are then detailed in full; and "Higher ROI = more value per cost" is trivial. It is clearly above anchor 3, which would require noticeable unnecessary explanation.

4 / 5

Actionability

Step 3 contains complete, executable Python implementing the Step 2 formula (argument-based JSON inputs, median normalization, sorted output), Step 6 gives a copy-paste `e2e-budget.yml`, Step 4 shows a concrete output template, and the worked example computes real ROI values. This matches 'fully executable; copy-paste ready code or commands; specific examples cover the common cases'.

5 / 5

Workflow Clarity

A clear 7-step sequence with a dedicated validation checkpoint (Step 1a: flag missing runtime as SKIP, flag null flake_rate as NEEDS MANUAL REVIEW and exclude, cross-check test-ID sets, warn when >50% of tests score 0.0 before generating recommendations) — an explicit validate-before-recommend feedback loop for a batch operation. Decisions are human-gated in Step 5 ("The team picks the appropriate class per test; the skill recommends"), which correctly guards against auto-retiring. Matches anchor 5's explicit validation and error-recovery loop.

5 / 5

Progressive Disclosure

The SKILL.md is a well-organized overview-plus-workflow with a single, clearly signaled one-level-deep reference: "See [references/anti-patterns-and-limitations.md](references/anti-patterns-and-limitations.md) for common failure modes... and known constraints..." — the referenced file exists in the bundle and contains exactly what the link promises, with no further nesting. Formula, code, and templates are appropriately inline for a single-workflow skill rather than over-split.

5 / 5

Total

19

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An excellent description: concrete actions with the formula inline, an explicit multi-trigger 'Use when' clause, and explicit boundary routing to three sibling skills. The only soft spot is trigger synonym coverage, which is good but not exhaustive of the phrasings a user might naturally use.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions with the exact mechanism inline: "Caps E2E suite size by computing per-test ROI", "ranks every end-to-end test and recommends which bottom-decile ones to retire, move to a lower layer, or fix", plus the full formula. Coverage is comprehensive; anchor 4's 'minor gaps in coverage' does not apply since even the formula and decision classes are stated.

5 / 5

Completeness

Both what and when are explicit: the what is the ROI computation, ranking, and bottom-decile recommendations with the formula spelled out; the when is a full "Use when..." clause with three concrete triggers plus a quarterly cadence. Clearly matches the anchor-5 example structure rather than anchor 4, where the 'when' could be more specific.

5 / 5

Trigger Term Quality

Natural trigger phrases are strong: "Use when CI is slow or E2E-dominated, flaky failures are rising, or quarterly to keep suite size within maintenance capacity" maps well to what a user would say (slow CI, flaky tests, too many E2E tests). It stops short of anchor 5's comprehensive synonym coverage (e.g., "test suite too slow", "E2E tests take forever", "trim/prune the suite") — a few common user phrasings are missing, so it sits at 'good coverage, a few natural terms missing'.

4 / 5

Distinctiveness Conflict Risk

It has a clear niche (pruning low-signal tests by ROI) and explicitly routes adjacent needs to siblings: "For strategic unit:service:UI layer ratios use test-pyramid-balancer, for the minimal per-deploy critical-path gate use smoke-suite-gate, and for quarantining flaky tests use flaky-test-quarantine; this prunes low-signal tests by ROI". This explicit boundary-setting minimizes conflict risk beyond anchor 4's 'minor overlap risk'.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Reviewed

Table of Contents