CtrlK
BlogDocsLog inGet started
Tessl Logo

improving-mcp-tools

Run an improve-my-MCP campaign: an autoresearch-style loop that measures the MCP agent experience with the eval harness, picks the highest-impact tool problem from production data, makes one bounded fix, and keeps it only if before/after scores improve. Use when asked to "improve my MCP", run an MCP improvement campaign, fix tool discoverability or descriptions based on evidence, or prepare an eval-backed PR for a tool change. Every shipped change must carry eval evidence; guardrails below are hard rules.

71

Quality

87%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

83%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, actionable operating manual for a measurement-gated improvement loop, with strong workflow sequencing and explicit validation/guardrails appropriate to a destructive-capable campaign. Conciseness and progressive disclosure are very good but not perfect, with some inlined detail and minor restated rationale.

Suggestions

Move the exact agent-mode runner command (and the exact slice-selection rule) into the body so both probe and agent modes are equally copy-pasteable.

Consider splitting the detailed journal/PR-evidence format fully into references/campaign-journal.md and linking to it rather than restating the 'before/after scores in the body' format inline in step 5.

Trim restated guardrail rationale (e.g. 'that's why the no-regression sample is mandatory', 'changing the exam and the answer together proves nothing') since Claude can infer the why from the rule.

DimensionReasoningScore

Conciseness

Mostly lean and assumes Claude's competence — it skips explaining what MCP or eval harnesses are and uses tight command strings — but the guardrails and failure-modes sections include some defensive restating of rationale ('Anything else... → stop', 'that's why the no-regression sample is mandatory') that could be trimmed.

4 / 5

Actionability

Provides concrete, copy-pasteable commands (probe.ts invocation, the pnpm dev:hono local recipe with env vars, named analytics tools) and a precise allowlist; minor gaps are the absense of the exact agent-mode command and the benchmark slice selection being left implicit.

4 / 5

Workflow Clarity

The 'One iteration' section is a clearly sequenced 6-step procedure with explicit validation (step 4: re-run slice + no-regression sample, keep only if metric improves and nothing degrades) and feedback loops (discard → journal → move on), plus hard guardrails and a kill switch — matching the top anchor for batch/destructive workflows.

5 / 5

Progressive Disclosure

SKILL.md is a well-structured overview with one clearly signaled one-level-deep reference (references/campaign-journal.md, which exists and holds the journal/PR-evidence detail); the cross-link to the intent-clusters skill is also clearly signaled. Not a 5 because some reference-worthy detail (failure modes, the full journal format) is inlined rather than split out.

4 / 5

Total

17

/

20

Passed

Description

91%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third person, concrete actions, explicit 'Use when' triggers with natural phrasings, and a distinct niche. The only minor weakness is a slight density of internal jargon (autoresearch-style, eval harness) but this does not undermine clarity.

DimensionReasoningScore

Specificity

Names the domain (MCP tool improvement) and several concrete actions — 'measures the MCP agent experience', 'picks the highest-impact tool problem', 'makes one bounded fix', 'keeps it only if before/after scores improve' — with only minor coverage gaps relative to the comprehensive anchor.

4 / 5

Completeness

Explicitly answers both what (autoresearch-style loop that measures, picks, fixes, re-scores) and when ('Use when asked to improve my MCP...') with concrete trigger phrases, matching the top anchor.

5 / 5

Trigger Term Quality

Includes natural user phrasings — 'Use when asked to "improve my MCP"', 'run an MCP improvement campaign', 'fix tool discoverability or descriptions', 'prepare an eval-backed PR' — covering synonyms and variations a user would actually say.

5 / 5

Distinctiveness Conflict Risk

Targets a clear niche — eval-backed MCP tool improvement campaigns — with distinct triggers unlikely to fire for unrelated skills; minimal conflict risk.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 suspicious

Warning

Total

15

/

16

Passed

Repository
PostHog/posthog
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.