CtrlK
BlogDocsLog inGet started
Tessl Logo

exploring-mcp-tool-quality

Investigate the quality of PostHog MCP tool calls — error rates, latency, reach, and which tools are failing or slow. Use when the user asks "which MCP tool has the highest error rate?", "what's the slowest tool?", "which tools fail most often?", "how reliable is tool X?", wants a tool-quality matrix, or pastes an MCP analytics tool-quality / dashboard URL and asks what it shows.

69

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

—

The risk profile of this skill

SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable skill body with concrete SQL recipes, typed-tool guidance, and sensible governance checkpoints. It loses points mainly for a slightly padded opening, a couple of incomplete recipes (slowest tools, matrix), and reliance on a cross-skill reference rather than a local bundle file.

Suggestions

Tighten the opening paragraph: drop 'This is the data behind the MCP analytics dashboard and tool-quality screens.' and lead directly with the actionable data-model facts Claude needs.

Inline a complete slowest-tools query (the p95 latency ranking) rather than only a transformation hint, so the common cross-tool latency case is copy-paste ready like the error-rate recipe.

Add an explicit verify step after running a typed tool or SQL recipe (e.g., cross-check the headline against the tool-quality screen URL) to close the workflow_clarity validation gap.

DimensionReasoningScore

Conciseness

Largely lean and assumes Claude's knowledge of MCP/PostHog/HogQL, but the opening data-model paragraph ('There is no dedicated ClickHouse table... This is the data behind the MCP analytics dashboard and tool-quality screens.') and a few restatements could be trimmed.

4 / 5

Actionability

Provides copy-paste-ready SQL for the error-rate ranking and failure-bucket recipes plus named typed tools with parameters and URL templates, but the slowest-tools workflow is only a transformation hint ('Swap the aggregate... order by p95_ms') and the matrix query is offloaded to an external file.

4 / 5

Workflow Clarity

Clear sequencing with a governed-metric-first checkpoint (run metric-list, check approval, label derived rates noncanonical) and a HAVING-floor noise guard, plus a drift check against mcp_harness.py; minor validation gaps remain since these are read-only queries without explicit error-recovery loops.

4 / 5

Progressive Disclosure

Well-organized with section headers and a clearly signaled, one-level-deep reference to the shared models-mcp.md for bulk recipes, but that reference is a cross-skill path outside this skill's own bundle (no local references/ dir) rather than a local file, a minor organization gap.

4 / 5

Total

16

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description that answers what and when with natural trigger phrases and concrete capabilities. Only minor overlap risk with closely related MCP analytics skills keeps distinctiveness just below full marks.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'error rates, latency, reach, and which tools are failing or slow' plus a 'tool-quality matrix' — giving comprehensive coverage of what the skill investigates.

5 / 5

Completeness

Explicitly states both the 'what' ('Investigate the quality of PostHog MCP tool calls...') and the 'when' ('Use when the user asks...') with concrete trigger phrases.

5 / 5

Trigger Term Quality

Quotes natural user phrases verbatim ('which MCP tool has the highest error rate?', 'what's the slowest tool?', 'how reliable is tool X?') plus URL-paste triggers, covering synonyms and variations a user would actually say.

5 / 5

Distinctiveness Conflict Risk

Carves a clear niche (PostHog MCP tool-quality metrics) with distinct dashboard-URL triggers, but overlaps slightly with sibling skills like exploring-mcp-intent-clusters on 'which tools fail most'.

4 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 5 suspicious

Warning

Total

15

/

16

Passed

Repository
PostHog/posthog
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.