CtrlK
BlogDocsLog inGet started
Tessl Logo

exploring-llm-clusters

Investigate AI observability clusters — understand usage patterns in AI/LLM traffic, compare cluster behavior, compute cost/latency metrics, and drill into individual traces within clusters.

63

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

—

The risk profile of this skill

Fix and improve this skill with Tessl

tessl review fix ./products/ai_observability/skills/exploring-llm-clusters/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable with executable SQL and a clear sequenced workflow, and it is appropriately concise for PostHog-specific material with only minor redundancy. Its main weakness is progressive disclosure: it cites a scripts/ bundle that does not exist and inlines reference material that a real bundle would offload.

Suggestions

Either create scripts/print_clusters.py (and the scripts/ directory) or remove the dangling reference to it so navigation is not broken.

Consider moving the cluster object shape JSON and the per-level SQL variants into a references/ file, keeping SKILL.md as an overview with one representative query inline.

De-duplicate the window-vs-timestamp guidance so it appears once (Step 2) rather than being restated in Tips.

DimensionReasoningScore

Conciseness

The bulk is genuinely PostHog-specific reference material Claude would not know (data model, event properties, SQL), so it largely earns its tokens; minor redundancy (the window-vs-timestamp warning and several Tips restate guidance already in Step 2) keeps it just below the lean 5 anchor.

4 / 5

Actionability

Provides multiple copy-paste-ready, complete SQL queries with concrete field names covering trace, generation, and evaluation levels, plus a concrete query-llm-trace JSON invocation — matching the fully executable 5 anchor.

5 / 5

Workflow Clarity

A clear four-step sequence (list runs → get clusters → compute metrics → drill into traces) with explicit guardrail checkpoints ("Never bound this query with $ai_window_start", "Keep the lookback bound wide enough"); no formal validate→fix→retry loop, but operations are read-only so the destructive cap does not apply, leaving minor validation gaps at 4.

4 / 5

Progressive Disclosure

Sections are well-organized, but the body references [scripts/](./scripts/) and print_clusters.py while no scripts/ directory (or references/assets) exists — a dangling reference to a non-existent bundle file — and bulk reference material (SQL, cluster object shape) is inlined, fitting the 3 anchor's "references present but not clearly signaled / content that should be separate is inline."

3 / 5

Total

16

/

20

Passed

Description

71%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and well-targeted to a distinct niche, but lacks an explicit "Use when..." trigger clause, which caps its completeness at 3. Trigger term coverage is good but could add natural synonyms to reach the top anchor.

Suggestions

Add an explicit 'Use when...' trigger clause, e.g. 'Use when investigating AI/LLM traffic patterns, comparing cluster behavior, or finding costly/slow clusters.'

Include a few natural-language synonyms users might say (e.g. 'LLM usage groups', 'cluster analysis') to broaden trigger term coverage.

Sharpen the boundary with the sibling traces skill so the description signals clusters as the entry point rather than trace inspection.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — "understand usage patterns in AI/LLM traffic", "compare cluster behavior", "compute cost/latency metrics", "drill into individual traces" — giving comprehensive coverage of the skill's capabilities, matching the 5 anchor.

5 / 5

Completeness

The "what" is clear and concrete, but there is no "Use when..." clause or equivalent explicit trigger guidance; the "when" is only weakly implied, which per the judging guidelines caps completeness at 3.

3 / 5

Trigger Term Quality

Relevant, fairly natural domain keywords are present ("AI observability clusters", "LLM traffic", "cost/latency metrics", "traces"), but a few natural variations or synonyms are missing and there is no plain-language trigger phrasing, so it sits just below comprehensive (5).

4 / 5

Distinctiveness Conflict Risk

"AI observability clusters" is a distinct niche with low conflict risk, but "drill into individual traces within clusters" overlaps with a closely related LLM-traces skill, leaving minor overlap risk rather than the fully distinct 5 anchor.

4 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 missing, 1 suspicious

Warning

Total

15

/

16

Passed

Repository
PostHog/posthog
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.