CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/test-framework-architecture-audit

Audits an existing test automation framework across eight architecture-tier axes and bands each one PASS, WARN, or FAIL: page-object coverage and purity, base-class inheritance depth, fixture scope and coupling, helper sprawl, naming-convention drift, retry and wait consistency, documented-versus-actual convention drift, and CI integration health. Carries the numeric cut behind every band and labels which cuts are practitioner conventions rather than published standards. Measures the framework's own structure (page objects, base classes, fixtures, helpers, conventions), not the suite's tier mix or flake rate, and not the design of a framework that does not exist yet. Use when a test framework has grown for a release or more without structural review, before a major refactor, or when a team suspects its written test conventions no longer match what the code actually does.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Overview
Quality
Evals
Security
Files

Quality

Content

85%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A dense, highly actionable audit methodology: concrete thresholds, per-axis band tables, a baseline gate, an explicit rollup and fix-order, and a complete output template, with the worked example correctly offloaded to a real one-level reference. Its only weakness is length and some redundancy (the Sources index, repeated framing) that pull conciseness below the top anchor.

Suggestions

Remove or shrink the 'Sources' index section — every citation already appears inline at the claim it supports, so the index repeats information without adding value.

Consolidate the repeated 'is a practitioner convention, not a published number' framing into a single statement near the rollup rather than restating it verbatim under each axis.

Tighten the 'What this owns, and what it does not' boundary table; the distinctiveness point is already made in the description, so the body version can be shortened to the suite-vs-framework distinction alone.

DimensionReasoningScore

Conciseness

Most lines earn their place as concrete thresholds and citations, but at ~440 lines the body is long and carries redundancy — the closing 'Sources' index restates the inline citations, and the ownership/boundary and repeated 'convention of this audit' framing could be tightened, placing it at 'mostly efficient but could be tighter' rather than 'every token earns its place'.

2 / 3

Actionability

Concrete and copy-ready throughout: explicit numeric cuts (90%/70%, depth 2/3/4, 1:10 ratio, 20% adoption, 80%/50%, retry >2), per-axis band tables, and a full Markdown output template with exact field placeholders — the top anchor for actionable instruction guidance.

3 / 3

Workflow Clarity

The process is clearly sequenced — a 'Before scoring' baseline gate (A1 pattern, A7 conventions doc), per-axis measurement, an explicit rollup rule (any FAIL→FAIL, else any WARN→WARN, else PASS), a blast-radius fix order, and a fixed output template — with checkpoints; the read-only audit is not destructive so the feedback-loop cap does not apply.

3 / 3

Progressive Disclosure

The body keeps the methodology inline and signals one clearly-named one-level-deep reference for the worked example and anti-pattern table ('A full worked audit... in references/audit-detail.md'), and that file exists in references/, matching the 'clear overview with well-signaled one-level-deep references' anchor.

3 / 3

Total

11

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description that names concrete actions, all eight audit axes, and the convention-vs-standard distinction, then closes with explicit 'Use when' triggers and explicit non-overlap with neighbouring skills. All four dimensions land at the top anchor.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'bands each one PASS, WARN, or FAIL', names all eight axes, 'carries the numeric cut behind every band', and 'labels which cuts are practitioner conventions' — far above the 'names domain and some actions' anchor at 2.

3 / 3

Completeness

Explicitly answers both what ('Audits an existing test automation framework across eight... axes') and when ('Use when a test framework has grown...'), satisfying the top anchor with an explicit trigger clause.

3 / 3

Trigger Term Quality

The 'Use when' clause gives natural triggers a practitioner would voice — 'grown for a release or more without structural review', 'before a major refactor', 'suspects its written test conventions no longer match what the code actually does' — matching the good-coverage anchor; not just technical jargon.

3 / 3

Distinctiveness Conflict Risk

It carves a clear niche ('Measures the framework's own structure... not the suite's tier mix or flake rate, and not the design of a framework that does not exist yet'), explicitly disambiguating from suite-health and design skills, so it is unlikely to trigger for the wrong skill.

3 / 3

Total

12

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Reviewed

Table of Contents