CtrlK
BlogDocsLog inGet started
Tessl Logo

creating-online-evaluations

Author continuously-running online evaluations in PostHog AI observability, grounded in real failure modes you've identified. Use when the user wants evaluations that automatically score new generations or whole traces going forward — "create an eval to catch X", "continuously check that responses do Y", "turn these failures into evals". Covers letting the explored data decide how many evals to create, proposing that set for the user to pick, choosing the target and eval type (hog / llm_judge / sentiment), configuring a provider and model for an llm_judge eval (a provider key gates enabling, not creation), scoping which generations trigger it via conditions, creating disabled, verifying scope, and enabling. Proposes a sentiment eval when no failure mode is worth catching. Finding and ranking the failure modes worth evaluating is its own job — use exploring-ai-failures first. To debug or manage evaluations that already exist, use exploring-llm-evaluations.

71

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

—

The risk profile of this skill

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable with a well-sequenced, validated workflow and domain-specific detail that mostly earns its tokens. Its main weakness is progressive disclosure: a prominently referenced payload reference file is missing, and some reference-grade configuration detail is inlined in SKILL.md instead.

Suggestions

Create the missing references/evaluation-payload.md (or remove the dangling links) so the two references to it resolve — currently both [references/evaluation-payload.md](references/evaluation-payload.md) links are broken.

Move the detailed settle-config bounds and globals reference tables (sections 2.2) into references/evaluation-payload.md, keeping only the decision guidance inline in SKILL.md.

Tighten the session-target skipped-evaluation and retention prose in 2.2 to the essential rules to lift conciseness toward the top anchor.

DimensionReasoningScore

Conciseness

The body is dense with PostHog-specific domain detail (settle strategies, globals, provider-key gating) that Claude would not already know, and largely assumes competence rather than padding; it earns 4 over 5 because a few prose sections (e.g. the session-target bounds and skipped-evaluation mechanics) could be tightened or moved to the reference, and over 3 because it is not "noticeably verbose" with generic explanation.

4 / 5

Actionability

Provides copy-paste-ready JSON payloads for hog and llm_judge eval creation, exact `conditions` shape, a runnable SQL verification query, and concrete `generate-app-url` calls — covering the common cases per the score-5 anchor.

5 / 5

Workflow Clarity

Sequences a two-phase process (1.1–1.3 then 2.1–2.6) with an explicit validation checkpoint (2.5 verify scope via SQL before enabling) and feedback loop (sample, review first live results, adjust rollout), satisfying the score-5 anchor and avoiding the batch-operation cap since validation is present.

5 / 5

Progressive Disclosure

Structure and sectioning are good, but the body twice signals [references/evaluation-payload.md](references/evaluation-payload.md) as the home for "every field, the config schemas, the exact conditions shape" and that file does not exist, so navigation is broken; additionally reference-grade detail (session settle bounds, the globals table) is inlined rather than split out, matching the score-3 anchor's "references present but not clearly signaled / content that should be separate is inline" better than the 4 anchor's "minor organization gaps".

3 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, complete, and well-differentiated, clearly stating both capability and trigger conditions in third-person voice with concrete user-facing phrasings. Its only soft spot is trigger-term breadth, where a few more synonyms would reach the top anchor.

DimensionReasoningScore

Specificity

Lists multiple concrete actions spanning the whole workflow — "choosing the target and eval type (hog / llm_judge / sentiment)", "configuring a provider and model", "scoping which generations trigger it via conditions", "creating disabled, verifying scope, and enabling" — giving comprehensive coverage, matching the score-5 anchor rather than the 4 anchor's "minor gaps".

5 / 5

Completeness

Explicitly answers both what ("Author continuously-running online evaluations in PostHog AI observability") and when ("Use when the user wants evaluations that automatically score new generations or whole traces going forward") with concrete trigger phrases, matching the score-5 anchor exactly.

5 / 5

Trigger Term Quality

Includes natural phrases users would say — "create an eval to catch X", "continuously check that responses do Y", "turn these failures into evals" — giving good keyword coverage; falls short of 5 because it lacks the broader synonym/extension breadth the top anchor calls for, and short of the "missing common variations" gap that defines 3.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear PostHog-AI-observability niche with distinct triggers and explicitly redirects overlapping jobs to siblings ("use exploring-ai-failures first", "use exploring-llm-evaluations"), minimizing conflict risk per the score-5 anchor.

5 / 5

Total

19

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 2 missing

Warning

referenced_paths_exist

Referenced path issues: 4 missing

Warning

Total

14

/

16

Passed

Repository
PostHog/posthog
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.