CtrlK
BlogDocsLog inGet started
Tessl Logo

running-tests

running tests at various levels from smoke tests to full suite to randomized tests

80

1.75x
Quality

69%

Does it follow best practices?

Impact

100%

1.75x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/running-tests/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This is a strong, highly actionable operational protocol: exact commands, deterministic seeds, failure-analysis loops, and a crisp reporting format, with almost no filler. Its weaknesses are structural — a monolithic file that inlines a large tag catalog which belongs in a reference file — plus slight redundancy between the inline rules and the ALWAYS/NEVER sections and one dangling 'Level 6' reference.

DimensionReasoningScore

Conciseness

The body is dense with executable commands and project-specific flags rather than concept explanations, assuming Claude's competence. Minor over-explanation could be trimmed: the log-level aside ('you might want to turn the log level up...') and the ALWAYS/NEVER sections, which restate rules already given inline (e.g. --abort, quiet-output flags).

4 / 5

Actionability

Every level gives copy-paste-ready commands with exact flags ('./stellar-core test --ll fatal -r simple --abort "[tx][soroban]"', 'NUM_PARTITIONS=$(nproc) make check', configure lines for each sanitizer), plus concrete test-tag catalogs and a fully specified report format. It does not fall below anchor 5 in any respect.

5 / 5

Workflow Clarity

The sequence is explicitly ordered by increasing cost (Levels 1→5 with a 3b checkpoint), with a hard stop-at-first-failure rule, an 'Interpreting Failures' feedback loop (identify → capture → triage flaky vs real → locate code), ALWAYS/NEVER checklists, and a defined output contract. The only blemish is the reference to undefined 'Levels 4-6' in Build Verification, which is too minor to drop it below anchor 5.

5 / 5

Progressive Disclosure

The skill is a single ~385-line file with good section headers but no bundle files at all. Material that clearly belongs in a one-level-deep reference — the ~35-line test-tag catalog and possibly the sanitizer build recipes — is inlined in SKILL.md, matching anchor 3 ('content that should be separate is inline'). It is above anchor 2 because the sections are well organized and navigable, but it cannot reach anchor 4-5 structure since nothing is split out and the tag list inflates the always-loaded context.

3 / 5

Total

17

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description communicates a clear escalation concept (smoke → full suite → randomized) but reads more like a bare capability label than a trigger-optimized description: it lacks any 'use when' guidance and does not mention the project it targets. It would under-trigger for users asking to 'run unit tests' or 'check the build' and over-trigger across unrelated projects.

Suggestions

Add an explicit trigger clause, e.g. 'Use when asked to run, verify, or validate tests for stellar-core changes, including smoke tests, unit test suites, sanitizer builds, or randomized/baseline test checks.'

Name the target project (stellar-core) and the concrete test levels from the body — sanitizers, tx-meta baseline checks — so the description both distinguishes this skill from generic test runners and captures the natural phrase 'unit tests'.

Use third-person verb phrasing ('Runs tests at escalating levels...') consistent with the body's protocol, and include synonyms users actually say (unit tests, regression tests, sanitizer runs).

DimensionReasoningScore

Specificity

The description names the domain ('running tests') and enumerates a concrete spectrum ('from smoke tests to full suite to randomized tests'), but these are level names rather than multiple distinct capabilities — sanitizer runs, tx-meta baseline checks, and build verification from the body are absent, so coverage is not comprehensive.

3 / 5

Completeness

The 'what' is clear (run tests at escalating levels), but there is no 'Use when...' clause or equivalent trigger guidance — the 'when' is entirely absent, which per the rubric caps completeness at 3.

3 / 5

Trigger Term Quality

Natural phrases a user would say are present ('running tests', 'smoke tests', 'full suite', 'randomized tests'), giving good keyword coverage. It misses common variations such as 'unit tests', 'sanitizers', or 'regression tests', keeping it below anchor 5.

4 / 5

Distinctiveness Conflict Risk

The description is not tied to stellar-core or any specific project, so it would trigger for virtually any test-running request in any repo, creating high overlap risk with generic test-execution skills. It is not entirely generic (the level escalation is a distinguishing shape), which keeps it above anchor 1.

2 / 5

Total

12

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
stellar/stellar-core
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.