CtrlK
BlogDocsLog inGet started
Tessl Logo

e2e-template-testing

End-to-end validation of coordinator and agent template changes

53

Quality

61%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.squad-templates/skills/e2e-template-testing/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

70%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a genuinely expert, highly actionable E2E playbook with excellent sequencing and validation gates — the strongest part of the skill. Its weaknesses are structural: it is a monolith with no progressive disclosure to reference files, and it carries notable duplication (two divergent copies of the tracking-comment body, repeated legends) that inflates token cost without adding information.

Suggestions

Move the Progress Reporting lifecycle (initial comment, update patterns, Windows-safe posting, final verdict replacement) into a references/progress-reporting.md file and keep a 10-line summary plus a pointer in SKILL.md.

De-duplicate the initial tracking comment body: Step 0 and the Progress Reporting section show it in two different emoji formats (shortcodes vs unicode); keep one canonical version.

Consolidate the repeated status legend and the Progressive Verdicting rules that are restated in Anti-Patterns into single referenced sections to cut token cost.

DimensionReasoningScore

Conciseness

The body is mostly dense, non-obvious operational knowledge (the ~15-minute connection budget, PS 5.1 BOM corruption, stale-SDK shadowing) that earns its tokens, but it could be tightened: the initial tracking comment body appears twice in two inconsistent formats (GitHub shortcodes in Step 0, unicode emoji in the bash snippet), the status legend is repeated three times, and Progressive Verdicting rules are restated verbatim in Anti-Patterns. It is not 4 because this duplication is a recurring pattern, not a minor instance.

3 / 5

Actionability

Guidance is mostly copy-paste executable: build/link commands, test-repo creation, session invocation, gh api PATCH calls, and a verdict template with concrete checks. It falls short of 5 because a few spots leave the executor to fill in mechanics — the Progressive Verdicting snippet contains '# ...rebuild the full comment body...' and the fallback 'evidence/git-log.txt' capture is only implied — and placeholder prompts ('prompt A', '[scenario name]') are left to be improvised.

4 / 5

Workflow Clarity

Steps 0-6 are explicitly sequenced with fast-fail gates and named error codes (BUILD_FAILED, LINK_FAILED, CLI_NOT_FOUND), per-step verification (squad version preview suffix, .squad/ file checks, session-log greps), a verdict checklist, and error-recovery feedback loops (build failure -> npm install -> retry; connection drop -> resume from last PATCHed comment). It is not 4 because validation checkpoints are present at every stage rather than having minor gaps.

5 / 5

Progressive Disclosure

Sections are well-labeled, but the skill is a ~550-line monolith with no bundle files (references/, scripts/, assets/ do not exist); the Progress Reporting comment-lifecycle (~150 lines of update patterns and PowerShell posting code), PII examples, and the Windows-handling detail clearly belong in a separate reference file per the anchor 'content that should be separate is inline'. It is not 2 because headers and cross-links do provide real structure and navigation, but not 4 because nothing is split out and the sole external reference points outside the bundle (../../../CONTRIBUTING.md).

3 / 5

Total

15

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description states a clear, niche 'what' but omits any 'when to use' guidance, which both caps completeness and weakens trigger-term coverage. It is concise and distinct within its project, but a user asking for help after editing a template would not reliably surface it.

Suggestions

Append an explicit trigger clause, e.g. 'Use when you change files in .squad-templates/ (squad.agent.md, agent charters, notes-protocol.md) or the init scaffolding that writes them.'

Add 1-2 concrete capability phrases so the 'what' is comprehensive, e.g. 'Builds the local CLI, runs real squad sessions in disposable test repos, and verifies coordinator behavior and state persistence.'

Include natural synonyms users would actually say — 'coordinator prompt', 'agent charter', 'squad templates' — to strengthen trigger-term coverage.

DimensionReasoningScore

Specificity

The description names its domain ('coordinator and agent template changes') and a single concrete action ('End-to-end validation'), matching the 'names domain and 1-2 concrete actions, but not comprehensive' anchor. It does not reach 4 because it never lists the underlying actions (build the CLI, run sessions, verify state), and it is not a 2 since the domain is concrete rather than generic.

3 / 5

Completeness

It answers 'what' clearly ('End-to-end validation of coordinator and agent template changes') but contains no 'Use when...' clause or equivalent trigger guidance, which caps completeness at 3 per the judging guidelines. Score 4 is ruled out because 'when' is entirely absent, not merely implicit.

3 / 5

Trigger Term Quality

Terms like 'end-to-end', 'validation', and 'template changes' are relevant but lean on project-internal jargon; the natural phrases a user would say ('I changed the coordinator prompt', 'squad.agent.md', 'charter') are missing, matching 'some relevant keywords but missing common variations or synonyms'. It is above 2 because the keywords are not purely generic, but below 4 because coverage of user-spoken variants is thin.

3 / 5

Distinctiveness Conflict Risk

'Coordinator and agent template changes' carves out a clear niche within the squad project with distinct triggers, giving minor overlap risk only with closely related testing skills. It is not 5 because 'end-to-end validation' alone is broadly generic and could overlap with general E2E testing skills.

4 / 5

Total

13

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (559 lines); consider splitting into references/ and linking

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 1 missing, 1 suspicious

Warning

Total

13

/

16

Passed

Repository
bradygaster/squad
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.