CtrlK
BlogDocsLog inGet started
Tessl Logo

testing

Skill validation framework PLUS daily test-suite health and regression intelligence. Validates skill conformance (frontmatter, manifest coverage, resolver coverage). Runs the project test suite in tiered phases (unit / evals / integration / system health), classifies failures, and produces a regression-aware report.

53

Quality

61%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/testing/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A genuinely actionable, well-sequenced protocol body with executable commands, a failure-classification scheme, and explicit safety boundaries — its workflow quality is high. The main costs are self-admitted conformance-test stub sections at the tail, loose non-executable health/eval steps, and a monolithic ~235-line layout that keeps everything inline despite clear opportunities to split Mode 2 detail into reference files.

Suggestions

Delete or merge the trailing '## Contract' and '## Output Format' stub sections, which exist solely to satisfy test/skills-conformance.test.ts and add no behavioral guidance.

Make the loose steps executable: give a concrete evals invocation instead of '# Adapt to the project's eval config', and specify actual disk/memory/CPU/liveness commands (e.g. df -h, free -m, the gbrain doctor call is already good).

Move version/time-sensitive details ('v0.25.1 extension', dated state JSON example) out of the body and split Mode 2's state schema and report templates into a reference file to reduce the monolithic inline bulk.

DimensionReasoningScore

Conciseness

Mostly efficient, but there is measurable padding: the trailing '## Contract' and '## Output Format' sections are self-admitted stubs ('this section exists for the conformance test', 'The literal section header here exists for the conformance test'), the 'Two modes' intro repeats the frontmatter, and the version strings ('v0.25.1 extension') are time-sensitive details that penalize conciseness. Not 2 because the bulk of the body is dense, useful protocol content rather than explanation of concepts Claude already knows.

3 / 5

Actionability

Mostly executable guidance: concrete commands ('bun test test/skills-conformance.test.ts ...', 'git log --oneline --since="24 hours ago"', 'gbrain doctor --fast --json'), a copy-paste report format, a failure-classification table with actions, and explicit do/don't auto-fix rules. Not 5 because steps 2 and 3 of the daily protocol are loose ('# Adapt to the project's eval config', 'Disk / memory / CPU' with no commands) and no concrete example of running a health check is given.

4 / 5

Workflow Clarity

The daily protocol is a clear numbered sequence (1–7) with validation checkpoints: classify each failure before acting, 'retry once' for flakes, 'retest' after bootstrap, and escalation rules ('ASK first', 'ALWAYS escalate, never auto-fix' for security tests) that form genuine feedback loops. Not 5 because a couple of checkpoints are implicit (system health and evals steps have no verify/retry loop) and the conformance mode's phase 7 'Report results' has no failure-path guidance.

4 / 5

Progressive Disclosure

Good structure with clearly headed, well-sequenced sections and only one-level-deep references; the single link ([conventions/quality.md](../conventions/quality.md)) is clearly signaled, and no bundle files exist to mis-navigate. Not 5 because at ~235 lines the body is monolithic — the Mode 2 state-file schema, tier table, and report templates are natural candidates for reference files — and it is well over the under-50-line simple-skill case that would earn a 5 without external references.

4 / 5

Total

15

/

20

Passed

Description

55%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, jargon-fluent description that clearly states what the skill does across its two modes, but it omits any 'use when' trigger guidance and buries the natural phrases users would actually say in the frontmatter triggers rather than the description. The broad second mode ('runs the project test suite') also creates overlap risk with generic testing skills.

Suggestions

Add an explicit 'Use when...' clause to the description, e.g. 'Use when the user asks to validate skills, run the tests, check test health, or diagnose what broke after changes.'

Fold the natural trigger phrases ('run the tests', 'how are the tests', 'what's broken') from the frontmatter into the description text so trigger matching works from the description alone.

Replace internal jargon like 'regression intelligence' and 'resolver coverage' with user-facing terms ('figure out which commit broke a test'), and clarify the boundary between this skill's test-run mode and a generic testing/CI skill.

DimensionReasoningScore

Specificity

The description lists several concrete actions — 'Validates skill conformance (frontmatter, manifest coverage, resolver coverage)', 'Runs the project test suite in tiered phases (unit / evals / integration / system health)', 'classifies failures, and produces a regression-aware report' — with only minor gaps such as the abstract 'regression intelligence' framing. Not 5 because coverage is padded with framework language ('Skill validation framework PLUS...') rather than being fully comprehensive of concrete behaviors.

4 / 5

Completeness

The 'what' is clearly answered (conformance validation plus tiered test runs with classified reporting), but there is no 'Use when...' clause or equivalent explicit trigger guidance in the description, which caps completeness at 3. Not 4 because the 'when' is entirely absent rather than merely under-specified.

3 / 5

Trigger Term Quality

It contains some relevant keywords users would say ('test suite', 'daily test-suite health', 'unit', 'integration'), but the natural trigger phrases ('run the tests', 'how are the tests', 'what's broken') live only in the frontmatter triggers, not the description, and much of the description is internal jargon ('resolver coverage', 'regression intelligence', 'tiered phases'). Not 4 because common natural variations users actually say are absent from the description text itself.

3 / 5

Distinctiveness Conflict Risk

The skill-conformance half is distinct ('frontmatter, manifest coverage, resolver coverage'), but the second half ('Runs the project test suite... classifies failures') is broad and would overlap with any generic testing/CI skill on natural prompts like 'run the tests'. Not 4 because the dual-mode scope plus generic test-running language creates real overlap risk with closely related skills.

3 / 5

Total

13

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 1 suspicious

Warning

Total

14

/

16

Passed

Repository
garrytan/gbrain
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.