CtrlK
BlogDocsLog inGet started
Tessl Logo

analyzing-experiment-query-performance

Pull and interpret production experiment query-performance data from the staff-only `/api/debug_ch_queries` endpoints backing the `/experiments/staff` scene: slowest experiment queries, precompute read/build health, and preaggregation cache footprint. Covers prod-US and prod-EU via a `query_performance:read` personal API key, all query params, and response field semantics (exception codes, exposure paths, precompute skip reasons, job states). Use when investigating slow or failing experiment queries, precompute regressions, 307/159/241 errors, preaggregation table growth, or when asked how experiment query performance or the precompute rollout is doing in production.

78

Quality

98%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

—

The risk profile of this skill

SKILL.md
Quality
Evals
Security

Quality

Content

96%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a dense, executable reference for a complex staff-only API surface: lean prose, copy-paste commands, a sequenced investigation workflow with health thresholds and an auth feedback loop. Its only weakness is that sizable reference material (exception-code and field-semantics tables) is inlined rather than split into a bundled reference file.

Suggestions

Move the 'Exception codes' table and 'Precompute metadata'/'Job states' field-semantics blocks into a references/ file (e.g. FIELD_SEMANTICS.md), keeping SKILL.md as an overview with a one-level-deep pointer — this would lift progressive_disclosure to a 5.

Add a short 'Healthy vs. unhealthy at a glance' summary table near the top of the Investigation workflow so the headline thresholds (failed reads fraction, stale_failed/stuck_pending at zero) are scannable before the prose.

Consider a tiny references/ stub for the PAT-creation snippet so the Authentication section reads as 'see references/auth.md' and the main body stays focused on interpretation.

DimensionReasoningScore

Conciseness

Lean and high-signal throughout: assumes Claude's competence (no generic explanations of ClickHouse, API keys, or caching) and every section conveys domain-specific knowledge Claude would not know (exception codes, skip-reason semantics, partition/TTL behavior).

5 / 5

Actionability

Provides copy-paste-ready, executable artifacts — a real devtools `fetch` snippet for key creation, env-var exports, and two `curl | jq` invocations projecting the exact fields needed for the common cases.

5 / 5

Workflow Clarity

The 'Investigation workflow' is a clearly sequenced, numbered procedure (headline → localize → drill → out-of-scope routing → citation) with explicit health thresholds and an auth-error feedback loop ('An HTTP 403 means... re-check the key's scopes before anything else').

5 / 5

Progressive Disclosure

Well-organized with clearly signaled sections and one-level-deep pointers to sibling skills (Metabase, Grafana MCP), but it is a single monolithic file with large reference tables (exception codes, field semantics) inlined that could live in a separate reference file.

4 / 5

Total

19

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is precise, third-person, and clearly delineates a narrow staff-only API surface with concrete capabilities and rich natural trigger phrases. It answers both 'what' and 'when' explicitly with minimal conflict risk.

DimensionReasoningScore

Specificity

Lists multiple concrete actions ('Pull and interpret... slowest experiment queries, precompute read/build health, and preaggregation cache footprint', 'all query params, and response field semantics') with comprehensive coverage of the API surface.

5 / 5

Completeness

Explicitly answers both what ('Pull and interpret production experiment query-performance data... covers... all query params, and response field semantics') and when ('Use when investigating slow or failing experiment queries...').

5 / 5

Trigger Term Quality

Natural trigger phrases a user would actually say — 'slow or failing experiment queries', 'precompute regressions', '307/159/241 errors', 'preaggregation table growth' — covering both symptom and casual phrasings.

5 / 5

Distinctiveness Conflict Risk

A sharply defined niche — staff-only `/api/debug_ch_queries` endpoints behind the `/experiments/staff` scene with a `query_performance:read` scope — that would not trigger for any other skill.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
PostHog/posthog
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.