Run bot acceptance tests to validate MCP tools work correctly from a real AI agent's perspective. Use when testing PRs, detecting regressions, or verifying tool changes end-to-end with Claude/Gemini CLIs.
70
87%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Bot acceptance testing validates that MCP tools work correctly from a real AI agent's perspective. You design test scenarios dynamically, run them via tests/uat/run_uat.py, and evaluate results.
python tests/uat/run_uat.pyall_passed per agent. If true, you're done.results_file for full output, stderr, raw JSON--branch master to compareThe runner returns a concise summary to stdout (saves context when all passes):
{
"results_file": "/tmp/bat_results_abc123.json",
"agents": {
"gemini": {
"all_passed": true,
"test": {
"completed": true,
"duration_ms": 8100,
"exit_code": 0,
"num_turns": 5,
"tool_stats": { "totalCalls": 4, "totalSuccess": 4, "totalFail": 0 }
},
"aggregate": {
"total_duration_ms": 15300,
"total_turns": 12,
"total_tool_calls": 9,
"total_tool_success": 9,
"total_tool_fail": 0
}
}
}
}num_turns, tool_stats (per phase) for fine-grained comparisonoutput and stderr for diagnosisresults_filecat <<'EOF' | python tests/uat/run_uat.py --agents gemini
{
"setup_prompt": "Create a test automation called 'bat_error_test' with a time trigger at 23:59 and action to turn on light.bed_light.",
"test_prompt": "Try to get automation 'automation.nonexistent_xyz'. Report if the tool signaled an error or returned a normal response. Then get automation 'automation.bat_error_test' and report its structure.",
"teardown_prompt": "Delete automation 'bat_error_test' if it exists."
}
EOFFull BAT comparison (recommended):
git fetch origin master && git checkout master && git pullgit checkout feat/my-branchCompare these metrics:
Primary (decide pass/fail on these):
aggregate.total_tool_calls vs total_tool_failSecondary (report but don't decide on these alone):
aggregate.total_tool_calls, aggregate.total_turns — directional signal, not conclusive (agent exploration varies between runs)aggregate.total_duration_ms — noisy due to network, cache misses, server load. Only flag large (>2x) regressions.Robustness tip: Ask the same task in different ways (variation testing) to check if results are consistent across phrasings.
Quick comparison (single command):
# Test the PR branch
echo '{"test_prompt":"..."}' | python tests/uat/run_uat.py --branch feat/tool-errors --agents gemini
# Compare against master
echo '{"test_prompt":"..."}' | python tests/uat/run_uat.py --branch master --agents geminiEach scenario invocation costs API credits (one per agent per phase). Design scenarios efficiently:
When /bat-adhoc is invoked with arguments:
If arguments contain a scenario description, generate the JSON scenario and run it:
/bat-adhoc test automation create with sunrise trigger then modify to sunset→ Generate appropriate scenario JSON and execute
If --help or no arguments, show this help text.
Otherwise, treat $ARGUMENTS as instructions for what to test and design+run the scenario accordingly.
For complete CLI reference and output format, see tests/uat/README.md.
01964cc
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.