Investigate AI observability clusters — understand usage patterns in AI/LLM traffic, compare cluster behavior, compute cost/latency metrics, and drill into individual traces within clusters.
64
76%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Low
Low-risk findings worth noting
Fix and improve this skill with Tessl
tessl review fix ./products/ai_observability/skills/exploring-llm-clusters/SKILL.mdUse this skill when investigating AI observability clusters — understanding what patterns exist in your AI/LLM traffic, comparing cluster behavior, and drilling into individual clusters.
| Tool | Purpose |
|---|---|
posthog:llma-clustering-job-list | List clustering job configurations for the team |
posthog:llma-clustering-job-get | Get a specific clustering job by ID |
posthog:execute-sql | Query cluster run events and compute metrics |
posthog:query-llm-traces-list | Find traces belonging to a cluster |
posthog:query-llm-trace | Inspect a specific trace in detail |
PostHog clusters LLM traces, individual generations, or evaluation events by embedding similarity.
A Temporal workflow runs periodically or on-demand, producing cluster events stored as
$ai_trace_clusters (trace-level), $ai_generation_clusters (generation-level), or
$ai_evaluation_clusters (evaluation-level).
Each cluster event contains:
$ai_clustering_run_id — unique run identifier (format: <team_id>_<level>_<YYYYMMDD>_<HHMMSS>[_<job_id>])$ai_clustering_level — "trace", "generation", or "evaluation"$ai_window_start / $ai_window_end — time window of the data that was analyzed$ai_total_items_analyzed — number of traces, generations, or evaluations processed$ai_clusters — JSON array of cluster objects$ai_clustering_params — algorithm parameters usedThe analyzed window closes when a run starts, and the cluster event lands once the run finishes.
So the cluster event's own timestamp is always after $ai_window_end, by anything from seconds to hours.
Use the window only to bound the traces, generations, and evaluations that were analyzed.
To find the cluster event itself, filter on $ai_clustering_run_id with a plain recent-time bound.
$ai_clusters){
"cluster_id": 0,
"size": 42,
"title": "User authentication flows",
"description": "Traces involving login, signup, and token refresh operations",
"traces": {
"<trace_or_generation_id>": {
"distance_to_centroid": 0.123,
"rank": 0,
"x": -2.34,
"y": 1.56,
"timestamp": "2026-03-28T10:00:00Z",
"trace_id": "abc-123",
"generation_id": "gen-456"
}
},
"centroid_x": -2.1,
"centroid_y": 1.4
}cluster_id: -1 is the noise/outlier cluster (items that didn't fit any cluster)traces are keyed by trace ID (trace-level), generation event UUID (generation-level), or evaluation event UUID (evaluation-level)rank orders items by proximity to centroid (0 = closest)x, y are 2D coordinates for visualization (UMAP/PCA/t-SNE reduced)Each team can have up to 10 clustering jobs. A job defines:
"trace", "generation", or "evaluation"Default jobs named "Default - traces", "Default - generations", and "Default - evaluations" are auto-created
and disabled when a custom job is created for the same level.
posthog:execute-sql
SELECT
toString(properties.$ai_clustering_run_id) AS run_id,
toString(properties.$ai_clustering_level) AS level,
toString(properties.$ai_clustering_job_id) AS job_id,
toString(properties.$ai_clustering_job_name) AS job_name,
toString(properties.$ai_window_start) AS window_start,
toString(properties.$ai_window_end) AS window_end,
toFloat64OrNull(toString(properties.$ai_total_items_analyzed)) AS total_items,
timestamp
FROM events
WHERE event IN ('$ai_trace_clusters', '$ai_generation_clusters', '$ai_evaluation_clusters')
AND timestamp >= now() - INTERVAL 14 DAY
ORDER BY timestamp DESC
LIMIT 10posthog:execute-sql
SELECT
toString(properties.$ai_clustering_run_id) AS run_id,
toString(properties.$ai_clustering_level) AS level,
toString(properties.$ai_clustering_job_id) AS job_id,
toString(properties.$ai_clustering_job_name) AS job_name,
toString(properties.$ai_window_start) AS window_start,
toString(properties.$ai_window_end) AS window_end,
toFloat64OrNull(toString(properties.$ai_total_items_analyzed)) AS total_items,
properties.$ai_clusters AS clusters,
properties.$ai_clustering_params AS params,
timestamp
FROM events
WHERE event IN ('$ai_trace_clusters', '$ai_generation_clusters', '$ai_evaluation_clusters')
AND timestamp >= now() - INTERVAL 14 DAY
AND toString(properties.$ai_clustering_run_id) = '<run_id>'
ORDER BY timestamp DESC
LIMIT 1Keep the lookback bound wide enough to cover the timestamp Step 1 reported for the run.
Never bound this query with $ai_window_start / $ai_window_end.
The cluster event is emitted after the window closes, so those bounds return zero rows.
The clusters field is a JSON array. Parse it to see cluster titles, sizes, descriptions, optional metrics, and each cluster's traces map.
Important: The clusters JSON can be very large (thousands of trace, generation, or evaluation IDs with coordinates).
When the result is too large for inline display, it auto-persists to a file.
Use print_clusters.py from scripts/ to get a readable summary.
For trace-level clusters, compute cost/latency/token metrics:
posthog:execute-sql
SELECT
properties.$ai_trace_id as trace_id,
sum(toFloat(properties.$ai_total_cost_usd)) as total_cost,
max(toFloat(properties.$ai_latency)) as latency,
sum(toInt(properties.$ai_input_tokens)) as input_tokens,
sum(toInt(properties.$ai_output_tokens)) as output_tokens,
countIf(properties.$ai_is_error = 'true') as error_count
FROM events
WHERE event IN ('$ai_generation', '$ai_embedding', '$ai_span')
AND timestamp >= parseDateTimeBestEffort('<window_start>')
AND timestamp <= parseDateTimeBestEffort('<window_end>')
AND properties.$ai_trace_id IN ('<trace_id_1>', '<trace_id_2>', ...)
GROUP BY trace_idFor generation-level clusters, match by event UUID:
posthog:execute-sql
SELECT
toString(uuid) as generation_id,
toFloat(properties.$ai_total_cost_usd) as cost,
toFloat(properties.$ai_latency) as latency,
toInt(properties.$ai_input_tokens) as input_tokens,
toInt(properties.$ai_output_tokens) as output_tokens,
if(properties.$ai_is_error = 'true', 1, 0) as is_error
FROM events
WHERE event = '$ai_generation'
AND timestamp >= parseDateTimeBestEffort('<window_start>')
AND timestamp <= parseDateTimeBestEffort('<window_end>')
AND uuid IN ('<gen_uuid_1>', '<gen_uuid_2>', ...)For evaluation-level clusters, first check each cluster's metrics field from $ai_clusters (for example pass rate, N/A rate, dominant evaluator name, and average judge cost). When you need individual evaluation rows, match by event UUID:
posthog:execute-sql
SELECT
toString(uuid) AS evaluation_id,
toString(properties.$ai_trace_id) AS trace_id,
toString(properties.$ai_target_event_id) AS generation_id,
toString(properties.$ai_evaluation_name) AS evaluation_name,
toString(properties.$ai_evaluation_result) AS evaluation_result,
toString(properties.$ai_evaluation_reasoning) AS evaluation_reasoning,
toFloatOrNull(toString(properties.$ai_total_cost_usd)) AS judge_cost,
timestamp
FROM events
WHERE event = '$ai_evaluation'
AND timestamp >= parseDateTimeBestEffort('<window_start>')
AND timestamp <= parseDateTimeBestEffort('<window_end>')
AND uuid IN ('<eval_uuid_1>', '<eval_uuid_2>', ...)Once you've identified interesting clusters, use the trace tools to inspect individual traces:
posthog:query-llm-trace
{
"traceId": "<trace_id_from_cluster>",
"dateRange": {"date_from": "<window_start>", "date_to": "<window_end>"}
}Use events for cluster events, IDs, cost/latency/token metrics, and evaluation rows.
Do not query events.properties.$ai_input, $ai_output, or $ai_output_choices when you need user messages or full model inputs/outputs —
those heavy fields live on posthog.ai_events.
For a few representative examples, prefer query-llm-trace; it reads posthog.ai_events for you and returns the full event tree.
For batch extraction, first get the trace IDs from the cluster, then query posthog.ai_events anchored on trace_id:
posthog:execute-sql
SELECT
trace_id,
timestamp,
span_id,
event,
model,
input,
output_choices
FROM posthog.ai_events
WHERE trace_id IN ('<trace_id_1>', '<trace_id_2>', ...)
ORDER BY trace_id, timestampposthog.ai_events has a shorter retention window than events; older clusters may still have metadata and metrics but no message content.
For more detail, use the exploring LLM traces skill's event reference.
avg(cost), avg(latency), sum(cost) per clustertraces field)rank (closest to centroid = most representative)query-llm-trace to understand the patterntitle and description for the AI-generated summaryerror_countitems_with_errors / total_itemshttps://app.posthog.com/ai-observability/clustershttps://app.posthog.com/ai-observability/clusters/<url_encoded_run_id>https://app.posthog.com/ai-observability/clusters/<url_encoded_run_id>/<cluster_id>Always surface these links so the user can verify visually in the PostHog UI.
$ai_window_end is earlier than the event's own timestampcluster_id: -1) contains outliers that didn't fit any patternllma-clustering-job-list to understand what clustering configs are activequery-llm-trace for deep inspectionposthog.ai_events, not events.properties; use query-llm-trace unless you need custom batch SQL6fca5f8
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.