CtrlK
BlogDocsLog inGet started
Tessl Logo

bat-adhoc

Run bot acceptance tests to validate MCP tools work correctly from a real AI agent's perspective. Use when testing PRs, detecting regressions, or verifying tool changes end-to-end with Claude/Gemini CLIs.

69

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a well-structured operational guide: fully executable commands and examples, an explicit workflow with validation checkpoints and failure-recovery loops, and disciplined token economy with almost no padding. The only refinement opportunity is pushing more of the detailed output-format and regression-comparison material into the referenced README to tighten the overview further.

DimensionReasoningScore

Conciseness

The body is efficient — it assumes competence, never explains concepts Claude already knows, and nearly every line is a command, schema, or decision rule. Minor trims are possible (the opening sentence and "When to Use" bullets partially restate the frontmatter description; "Robustness tip" is advice Claude could infer), which keeps it at "minor instances that could be trimmed" rather than the lean 5.

4 / 5

Actionability

Fully executable throughout: a copy-paste-ready scenario (the heredoc piped to `uv run python tests/uat/run_uat.py --agents gemini`), concrete baseline/target regression commands with `--branch master` and `--scenario-file`, and a concrete JSON output example showing exactly which fields to compare (`aggregate.total_tool_calls`, `total_duration_ms`). Specific examples cover the common cases.

5 / 5

Workflow Clarity

The six-step workflow has a clear sequence with explicit checkpoints and feedback loops: "If true, you're done", "Dig deeper on failure: Read `results_file`", and "If test fails, re-run with `--branch master` to compare". The regression section adds decision rules (primary vs secondary metrics, the >2x duration flag threshold) that make pass/fail evaluation explicit. This matches the top anchor with error-recovery loops.

5 / 5

Progressive Disclosure

Good structure with well-organized sections and a clearly signaled single-level pointer to full docs ("For complete CLI reference and output format, see `tests/uat/README.md`"). No bundle files exist, so nothing is buried or nested; however, the runtime JSON output structure (24 lines) and parts of the regression workflow could live in that external README to keep SKILL.md closer to an overview, leaving minor organization gaps.

4 / 5

Total

18

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly states both what the skill does and when to use it, with explicit trigger guidance and appropriate third-person voice. Trigger term coverage is good but could add a few natural synonyms (acceptance testing, integration testing) to round it out.

DimensionReasoningScore

Specificity

The description lists several concrete actions — "Run bot acceptance tests", "validate MCP tools work correctly from a real AI agent's perspective", "testing PRs, detecting regressions, or verifying tool changes end-to-end" — in proper third person voice. It stops short of a 5 because "Run bot acceptance tests" is somewhat self-referential and the action coverage beyond the three test types is minor rather than comprehensive.

4 / 5

Completeness

It explicitly answers both questions: what it does ("validate MCP tools work correctly from a real AI agent's perspective") and an explicit trigger clause ("Use when testing PRs, detecting regressions, or verifying tool changes end-to-end with Claude/Gemini CLIs") with concrete trigger phrases. This matches the top anchor exactly.

5 / 5

Trigger Term Quality

Natural user phrases are present: "testing PRs", "detecting regressions", "verifying tool changes end-to-end", plus named CLIs (Claude/Gemini). A few common variations users might say are missing (e.g. "acceptance testing", "integration test", "regression check"), which keeps it at the "good coverage, a few natural terms missing" anchor rather than 5.

4 / 5

Distinctiveness Conflict Risk

The niche is clear and fairly distinct — agent-perspective acceptance testing of MCP tools with real Claude/Gemini CLIs — with distinct triggers. There is minor overlap risk with generic testing/regression or PR-review skills, so it sits at "mostly distinct" rather than "clear niche with minimal conflict risk".

4 / 5

Total

17

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
homeassistant-ai/ha-mcp
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.