CtrlK
BlogDocsLog inGet started
Tessl Logo

bat-story-eval

Compare MCP tool behavior between target and baseline versions using pre-built and custom stories with diff-based triage.

60

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/bat-story-eval/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a strong operational document: fully executable commands, a strictly sequenced workflow with validation after every story, and explicit feedback loops for failed queries, regressions, and token outliers. The only weaknesses are mild duplication between the body and the evaluation-protocol reference and a couple of repeated instructions.

DimensionReasoningScore

Conciseness

The body is dense and imperative — commands, templates, and tables with almost no conceptual explanation of things Claude already knows. Only minor tightening is possible: the custom-story rule 'At least 1. Each tests a distinct regression hypothesis' is stated twice in Step 0c, and the 'unverified' handling is explained both in Step 1b and again in Step 5. That fits anchor 4 ('efficient; minor instances of over-explanation that could be trimmed') rather than anchor 5's every-token-earns-its-place.

4 / 5

Actionability

Nearly everything is copy-paste executable: full `uv run` bash commands, working python3 snippets for parsing both Gemini JSON and Claude JSONL sessions, a complete custom-story YAML template, a concrete scoring matrix, and an exact report table schema with a filled-in example row. Placeholders like `<first_story>` and `<baseline>` are explicitly bound to parsed arguments, matching anchor 5 ('fully executable; copy-paste ready code or commands').

5 / 5

Workflow Clarity

Steps 0-8 are strictly ordered with each consuming the prior step's output, verification is mandated after every story ('ALWAYS verify each story via ha_query.py before running the next'), and error feedback loops are explicit: failed verification queries are re-run then recorded as 'unverified', regressions trigger re-run/flakiness handling via the referenced protocol, and token outliers route to a KV-cache investigation step. This matches anchor 5 ('explicit validation steps; feedback loops for error recovery').

5 / 5

Progressive Disclosure

Both bundle references exist (`references/evaluation-protocol.md`, `references/regression-protocol.md`) and are clearly signaled with purposes in the Key Files table, plus one inline pointer in Step 1b — one level deep, no nesting. However, the Step 4 scoring matrix and the white-box criteria partially duplicate the evaluation protocol reference, and the ~375-line body could push the scoring-matrix detail out to it. That fits anchor 4 ('good structure; most content appropriately placed; minor organization gaps') rather than anchor 5's clean split.

4 / 5

Total

18

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description communicates a clear and distinctive 'what' in one compact sentence, but omits any 'when to use' guidance and relies on internal jargon (stories, triage) rather than natural trigger terms. Per the rubric's cap rule, the missing 'Use when' clause limits completeness to 3.

Suggestions

Add an explicit trigger clause, e.g., 'Use when comparing MCP tool behavior between a released version and the current code, or when checking whether changes to the MCP server caused agent regressions.'

Replace internal jargon with natural user-facing terms — mention 'regression', 'A/B comparison', and 'version upgrade check' alongside 'pre-built and custom stories' and 'diff-based triage'.

Briefly name the downstream outputs (scoring, regression flagging, token-cost comparison, report) so the 'what' coverage is comprehensive rather than 1-2 actions.

DimensionReasoningScore

Specificity

The description names the domain ("MCP tool behavior between target and baseline versions") and 1-2 concrete actions ("Compare ... using pre-built and custom stories with diff-based triage"), but stops there — no mention of scoring, regression detection, or reporting. This matches anchor 3 ('names domain and 1-2 concrete actions, but not comprehensive') better than anchor 4, which expects several listed actions with only minor coverage gaps.

3 / 5

Completeness

The 'what' is clear — compare MCP tool behavior across versions with story-based triage — but there is no 'when' clause at all; the description never states when Claude should invoke it. Per the judging guidelines, a missing 'Use when...' clause or equivalent explicit trigger guidance caps completeness at 3, which this squarely matches.

3 / 5

Trigger Term Quality

Relevant keywords exist ("MCP tool behavior", "baseline", "diff-based triage") but the natural phrases a user would actually say are missing (e.g., "compare versions", "check for regressions", "A/B test the MCP server"). Terms like "pre-built and custom stories" and "diff-based triage" are internal jargon rather than user-facing triggers, fitting anchor 3 ('some relevant keywords but missing common variations or synonyms') and not anchor 4's 'good keyword coverage'.

3 / 5

Distinctiveness Conflict Risk

The niche is quite specific — A/B comparison of MCP tool behavior between git versions via story runs — so overlap with generic testing/eval skills is minor. It falls just short of anchor 5's 'distinct triggers' because the description provides no explicit trigger phrases that would reliably route invocation to it over a neighboring regression-testing skill.

4 / 5

Total

13

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
homeassistant-ai/ha-mcp
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.