CtrlK
BlogDocsLog inGet started
Tessl Logo

bat-story-eval

Compare MCP tool behavior between target and baseline versions using pre-built and custom stories with diff-based triage.

61

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

High

Do not use without reviewing

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/bat-story-eval/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable and well-sequenced with strong validation checkpoints for a batch/destructive eval workflow. Progressive disclosure is good but could push more detail into the existing references. Conciseness is strong with only minor redundancy.

DimensionReasoningScore

Conciseness

Mostly efficient with concrete commands and minimal conceptual padding; a few repeated reminders (container reuse, verify-before-next) and the inlined YAML template could be trimmed or moved.

4 / 5

Actionability

Provides copy-paste-ready bash and Python snippets with exact flags, a complete custom-story YAML template, and concrete scoring/triage matrices covering the common cases.

5 / 5

Workflow Clarity

Strictly ordered Step 0-8 sequence with explicit verification checkpoints after each story, exit-code/timeout handling, and feedback loops (re-run flaky, unverified handling) for this batch operation.

5 / 5

Progressive Disclosure

References (evaluation-protocol.md, regression-protocol.md) are real files and clearly signaled in-body and via the Key Files table, but sizable detail (scoring matrix, token-extraction formulas, custom-story YAML) is inlined rather than pushed one level deep.

4 / 5

Total

18

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description conveys a clear, specific purpose but omits any trigger/usage guidance, leaving Claude without a 'when to use this' signal. Trigger terms are domain-accurate but more technical than natural. Adding a 'Use when...' clause with user-facing phrasing would lift the weakest dimensions.

Suggestions

Append a 'Use when...' clause naming natural user triggers, e.g. 'Use when comparing MCP tool behavior between a baseline release and current code, or when triaging regressions across versions.'

Add common-sense synonyms a user might say ('regression testing', 'version comparison', 'evaluating MCP tools') to improve trigger term quality.

Optionally mention the artifact/output (e.g. producing a pass/fail/regression report) to round out the 'what' and reduce overlap ambiguity.

DimensionReasoningScore

Specificity

Names the domain (MCP tool behavior across versions) and several concrete actions ('Compare', 'diff-based triage', 'pre-built and custom stories'), but coverage is not comprehensive enough for a 5.

4 / 5

Completeness

Has a clear 'what' but no 'when'/'Use when...' trigger guidance, which caps completeness at 3 per the judging guidelines.

3 / 5

Trigger Term Quality

Contains relevant keywords ('MCP tool behavior', 'baseline versions', 'stories', 'triage') but they lean technical and lack the natural synonyms/phrases a user would actually say.

3 / 5

Distinctiveness Conflict Risk

Targets a distinct niche (cross-version MCP tool behavior comparison with story-driven triage), with only minor overlap risk against closely related eval skills.

4 / 5

Total

14

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
homeassistant-ai/ha-mcp
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.