CtrlK
BlogDocsLog inGet started
Tessl Logo

exploring-llm-evaluations

Investigate AI observability evaluations — `hog` (deterministic code-based), `llm_judge` (LLM-prompt-based), and `sentiment` (user-message sentiment). Find existing evaluations, inspect their configuration, run them against specific generations, query individual results, and set up scheduled reports on an evaluation. Use when the user asks to debug why an evaluation is failing, surface common failure modes, compare results across filters, dry-run a Hog evaluator, prototype a new LLM-judge prompt, inspect sentiment classifications, or manage the evaluation lifecycle.

69

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

—

The risk profile of this skill

SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-structured operational guide with excellent concrete examples, weakened by systematic cross-section repetition and a monolithic structure that inlines reference material no separate files split out.

Suggestions

Dedupe the repeated Hog-vs-LLM-judge guidance, the AI-data-processing-approval note, and the "diagnosis is identical" remark into a single canonical statement and reference it, to tighten conciseness.

Extract the event-schema property table and the SQL query templates into a references/ file (e.g. EVENT_SCHEMA.md and QUERIES.md) so SKILL.md reads as an overview with well-signaled one-level-deep links.

Add an explicit validation/confirmation step to the lifecycle-management workflow before destructive actions like llma-evaluation-delete and batch llma-evaluation-update.

DimensionReasoningScore

Conciseness

The body is mostly efficient and domain-specific (no generic PDF-style padding), but systematically repeats guidance across sections — "reach for Hog first", the AI-data-processing-approval requirement, and "diagnosis is identical, only the fix differs" each appear ~3 times — which could be tightened.

3 / 5

Actionability

Provides fully executable, copy-paste-ready JSON tool invocations and HogQL queries covering the common cases (listing, pass/fail breakdown, failing-run extraction, trace drill-down, create/test evaluator), with concrete parameter shapes.

5 / 5

Workflow Clarity

The investigation workflow is sequenced into clear numbered steps (1-4) with the N/A guard explained, and the build-test workflow has an explicit enabled:false → run → inspect → refine → enable feedback loop; minor gap is the absence of an explicit validation/confirmation checkpoint before destructive lifecycle actions (delete/update).

4 / 5

Progressive Disclosure

No bundle files exist (references/scripts/assets absent) and the SKILL.md is a ~420-line monolith containing the full event-schema table and all SQL templates inline; internal section headers give decent navigation, but per the rubric rationale SKILL.md should be an overview pointing to detailed materials, and content that could be split is inlined.

3 / 5

Total

15

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description that clearly states both capabilities and trigger conditions in third-person voice. Minor overlap with adjacent skills is the only weakness.

DimensionReasoningScore

Specificity

Lists multiple concrete actions ("Find existing evaluations, inspect their configuration, run them against specific generations, query individual results, and set up scheduled reports") plus three named evaluator types, giving comprehensive coverage rather than vague abstraction.

5 / 5

Completeness

It explicitly answers both what (investigate/manage the three evaluation types and their lifecycle) and when (a concrete "Use when the user asks to..." clause with multiple trigger scenarios).

5 / 5

Trigger Term Quality

The "Use when" clause packs in natural phrases users would actually say ("debug why an evaluation is failing", "dry-run a Hog evaluator", "prototype a new LLM-judge prompt", "inspect sentiment classifications", "manage the evaluation lifecycle"), covering many entry points to the skill.

5 / 5

Distinctiveness Conflict Risk

It carves a clear niche (investigate/run/manage PostHog evaluations) with distinct triggers, but overlaps slightly with the sibling creating-online-evaluations skill, leaving minor conflict risk with closely related skills.

4 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
PostHog/posthog
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.