CtrlK
BlogDocsLog inGet started
Tessl Logo

integration-e2e-testing

Integration and E2E test design principles, ROI calculation, test skeleton specification, and review criteria. Use when designing integration tests, E2E tests, or reviewing test quality.

68

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is integration-e2e-testing in shinpr/claude-code-workflows

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-engineered, dense ruleset: exact formulas, numeric budgets and thresholds, worked examples, and a real one-level-deep reference make it highly actionable and verifiable. The main weakness is readability of the ROI unknown-value handling prose, which is overwritten relative to the crisp tables surrounding it, plus the absence of an explicit end-to-end workflow ordering.

Suggestions

Rewrite the ROI unknown-value paragraphs (Unknown-Value Ordering and the preceding paragraph) as a short numbered procedure or table; the current run-on sentences bury the decision rules that the rest of the document expresses cleanly.

Add a brief ordered workflow at the top (e.g., Design Doc ACs → candidate extraction → ROI scoring → lane/budget selection → skeleton → implementation → review) so the section sequence reads as an explicit process rather than a reference rulebook.

DimensionReasoningScore

Conciseness

The body is dense and assumes Claude's competence — no space is spent explaining what E2E testing or AAA structure is, and nearly every table (budgets, ROI scales, thresholds) carries project-specific rules Claude cannot know. However, the ROI unknown-value prose (e.g., 'When ROI can change candidate ranking, a lane threshold, or budget selection, return the exact missing product input and its decision effect when `test_value_context` has not yet supplied it') is convoluted and could be tightened, fitting anchor 4 rather than anchor 5's 'every token earns its place'. It is clearly above anchor 3, which requires unnecessary explanation.

4 / 5

Actionability

Guidance is fully concrete and executable: an exact ROI formula ('ROI Score = Business Value × User Frequency + Legal Requirement × 10 + Defect Detection'), exact numeric lane thresholds (ROI ≥ 20 / ROI > 50), a copy-paste-ready annotation template, an eight-row worked example table with computed scores and selection outcomes, and specific file-naming patterns. This matches anchor 5's 'specific examples cover the common cases'; per the scoring notes, absence of code is not penalized in an instruction-only skill when the guidance is this actionable.

5 / 5

Workflow Clarity

The sections sequence coherently (test types and budgets → ROI ranking → journey definition → skeleton spec → review criteria), and the Review Criteria tables supply explicit checkpoints for verifying both skeletons and implementations. It falls short of anchor 5 because there is no explicit ordered procedure from Design Doc to selected/implemented tests and no validate-and-recover feedback loop; it exceeds anchor 3 because checkpoints are explicit rather than implicit. No destructive/batch-operation cap applies.

4 / 5

Progressive Disclosure

The bundle structure is sound: the single reference (references/e2e-design.md) exists, is clearly signaled at the top ('See [references/e2e-design.md]... for UI Spec-driven E2E test candidate selection and browser test architecture'), and is one level deep. This fits anchor 4-5; it falls short of anchor 5 only because the ~200-line body inlines detailed material (ROI example table, EARS mapping, naming conventions) that could arguably live in references, though the core ruleset legitimately belongs in SKILL.md. It clearly exceeds anchor 3, whose references are poorly signaled.

4 / 5

Total

17

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that explicitly answers both what the skill does and when to use it, with concrete, third-person capability statements and natural trigger terms. Its only weaknesses are minor: a few capability gaps in the 'what' list and missing common synonyms like 'end-to-end' in the trigger terms.

DimensionReasoningScore

Specificity

The description lists four specific capabilities — 'test design principles, ROI calculation, test skeleton specification, and review criteria' — which are concrete and map to real body sections, but coverage has minor gaps (lane budgets, journey definition, naming conventions are not named). This fits anchor 4 ('several specific actions; minor gaps') better than anchor 5, which demands comprehensive coverage.

4 / 5

Completeness

Both parts are explicit: the what ('Integration and E2E test design principles, ROI calculation, test skeleton specification, and review criteria') and the when ('Use when designing integration tests, E2E tests, or reviewing test quality') with concrete trigger phrases — a direct match for anchor 5. It clearly exceeds anchor 4, where the 'when' is only partially explicit.

5 / 5

Trigger Term Quality

'designing integration tests, E2E tests, or reviewing test quality' provides good natural trigger phrases a user would actually say. It misses common variations like 'end-to-end tests' (spelled out), 'test coverage', or 'test plan', so it fits anchor 4 rather than the comprehensive synonym coverage of anchor 5.

4 / 5

Distinctiveness Conflict Risk

The integration/E2E test-design niche is clear and specific, but 'reviewing test quality' is broad enough to overlap with general code-review or unit-testing skills, giving minor overlap risk — anchor 4 rather than anchor 5's 'clear niche with minimal conflict risk'. It is well above anchor 3, since the domain is distinctly scoped.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
shinpr/claude-code-workflows
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.