CtrlK
BlogDocsLog inGet started
Tessl Logo

soak-test

Soak test protocol for extended play — what to observe and log for slow leaks, fatigue, late-appearing edge cases.

57

Quality

72%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/soak-test/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a strong, highly executable skill: a fully specified protocol generator with per-duration checkpoint schedules, engine-specific monitoring steps, a complete output template, approval gates, and verdict rules including edge cases like aborted runs. The main weakness is two verbose provenance blockquotes that explain past-edit history instead of just stating the operational rule.

DimensionReasoningScore

Conciseness

The bulk of the body is tight, actionable checklists and tables, but the two long provenance blockquotes (e.g. 'This replaces a note asserting the return units of Performance.get_monitor … the safest fix was to stop needing it', and 'The previous threshold was an absolute "> 50MB", which silently assumes both the unit and a project scale') narrate the history of past edits rather than instruct the reader, and the operational rule in each could be stated in one line. This matches anchor 3 ('Mostly efficient but includes some unnecessary explanation or could be tightened') — not 2 because the padding is confined to two blockquotes, not several padded sections.

3 / 5

Actionability

Everything is executable: per-duration checkpoint schedules ('30m soak: T+0, T+10, T+20, T+30'), engine-specific commands ("Use `stat memory` console command at each checkpoint", 'Open Memory Profiler (Window → Analysis → Memory Profiler)'), a complete copy-paste protocol template with tables and checklists, the exact output path 'production/qa/soak-test-[date]-[duration].md', and a verbatim post-write message. This matches anchor 5 ('Fully executable; copy-paste ready … specific examples cover the common cases').

5 / 5

Workflow Clarity

Six numbered phases run in a clear sequence with explicit validation checkpoints: the approval gate ('May I write this soak test protocol to …?' / 'Write only after approval'), the NOT ASSESSED verdict rules covering aborted runs and missing instrumentation, and a feedback loop ('If the verdict is FAIL, run `/smoke-check` again after fixing the issues'). This matches anchor 5 ('Clear sequence with explicit validation steps; feedback loops for error recovery; checklists for complex processes').

5 / 5

Progressive Disclosure

A single well-organized file with clear phase sections and no bundle files; no nested or buried references. The ~130-line protocol template is the skill's own deliverable so keeping it inline is defensible, though it could be split into a reference file to slim the overview. This matches anchor 4 ('Good structure; most content is appropriately placed … minor organization gaps') rather than 5, which would expect the large template to live one level deep.

4 / 5

Total

17

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly identifies the niche (soak/endurance testing) and what the skill produces, but it omits any 'when to use' trigger guidance and lacks common synonym coverage (endurance test, memory leak, performance drift). It sits at the mid-range: specific enough to be found by someone who already says 'soak test', weak at capturing users who describe the need in other words.

Suggestions

Add an explicit trigger clause, e.g. 'Use when extended play needs to be formally tracked, before /gate-check release, or after fixing a memory or stability issue' — this currently only exists in the body, not the description.

Include natural synonyms users would say: 'endurance test', 'memory leak', 'performance drift', 'long play session' alongside the existing 'soak test' and 'extended play'.

State the concrete action the skill takes (generates the observation protocol document at production/qa/soak-test-[date]-[duration].md) rather than only describing what will be observed.

DimensionReasoningScore

Specificity

The description names the domain ("Soak test protocol for extended play") and the observables ("what to observe and log for slow leaks, fatigue, late-appearing edge cases"), which matches the anchor 'Names domain and 1-2 concrete actions, but not comprehensive'. It does not list several concrete actions (e.g., generate the protocol, define timed checkpoints, write the QA document), so it is not a 4; the observables are specific rather than minimal or generic, so it is above a 2.

3 / 5

Completeness

The 'what' is clear — a soak test protocol defining what to observe and log — but there is no 'when' clause at all (no 'Use when...' or equivalent trigger guidance). Per the judging guideline, a missing 'Use when' clause caps completeness at 3, matching the anchor 'Has a clear what but when is missing or only weakly implied'.

3 / 5

Trigger Term Quality

It contains the natural core term "soak test" plus "extended play", "slow leaks", and "fatigue", but misses common variations and synonyms a user would say — "endurance test", "memory leak", "performance drift", "long play session". Anchor 3 ('Some relevant keywords but missing common variations or synonyms') fits; keyword coverage is not broad enough for 4.

3 / 5

Distinctiveness Conflict Risk

Soak/endurance testing with its specific triggers (slow leaks, fatigue, late-appearing edge cases) is a clear niche distinct from broad QA skills, with only minor overlap risk against closely related playtest/smoke-check skills. This matches anchor 4 ('Mostly distinct; minor overlap risk with closely related skills') — not 5 because neighboring short-session QA skills cover adjacent territory.

4 / 5

Total

13

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
Donchitos/Claude-Code-Game-Studios
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.