Run an improve-my-MCP campaign: an autoresearch-style loop that measures the MCP agent experience with the eval harness, picks the highest-impact tool problem from production data, makes one bounded fix, and keeps it only if before/after scores improve. Use when asked to "improve my MCP", run an MCP improvement campaign, fix tool discoverability or descriptions based on evidence, or prepare an eval-backed PR for a tool change. Every shipped change must carry eval evidence; guardrails below are hard rules.
80
100%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
An MCP server gets better only in ways you can measure. This skill is the campaign procedure: score the current agent experience, fix the biggest problem, re-score, and only ship changes the numbers justify. It is the operating manual for the "improve my MCP" loop — one iteration per pass, journaled so a later iteration (or a different agent) can resume without repeating work.
services/mcp/evals/ is the harness. benchmark/tasks.yaml is a fixed set of
agent tasks with expected_tools and success_criteria; scores are only
comparable across runs of the same benchmark version.
LIVE_MCP_URL=... LIVE_MCP_TOKEN=... pnpm exec tsx evals/runner/probe.ts --out score.json
from services/mcp/. Reports tool-presence misses (discoverability), probe
failures, and latency p50/p95. Non-zero exit = regression.Run the harness against a seeded local or devbox stack, never against a
customer project. Local recipe: NODE_ENV=development PORT=9876 POSTHOG_API_BASE_URL=http://localhost:8000 pnpm dev:hono, personal API key as
LIVE_MCP_TOKEN.
query-mcp-tool-stats, query-mcp-tool-failures,
query-mcp-tool-descriptions, query-mcp-tool-sample-intents) and the
lenses in the signals scout cookbook
(products/signals/skills/signals-scout-mcp-tool-calls/references/queries.md):
failure leaderboard, retry/struggle, latency, intents that matched no tool.stamphog label. Autonomy level comes from the campaign config —
default is draft PR for human review; only arm auto-merge when the
operator has explicitly enabled the self-driving experiment (see
guardrails).These are not suggestions; violating any of them ends the campaign pass.
products/*/mcp/tools.yaml,
products/*/skills/**, services/mcp/evals/**, the codegen outputs of
pnpm generate-tools / scaffold-yaml (services/mcp/src/tools/generated/**
and services/mcp/schema/generated-tool-definitions.json), and docs.
Anything else (handler code, package manifests, workflows, migrations, auth
paths) → stop and hand the finding to a human as a draft PR or report
instead.benchmark/tasks.yaml in the same PR as
a fix it validates — changing the exam and the answer together proves
nothing. Benchmark changes are their own PR and bump version.getToolsForFeatures gating before "fixing" discoverability.bb04e92
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.