CtrlK
BlogDocsLog inGet started
Tessl Logo

writing-evals

Teaches how to write and run evals on the `products/posthog_ai/eval_harness/` harness — sandboxed agent suites that execute the real coding agent in a Docker or Modal sandbox against a seeded Hedgebox project, and one-shot suites that score a single in-process model invocation per case. Use when adding or changing eval suites, cases, scorers, seeders, or synthesizers under `products/posthog_ai/evals/` or `products/*/evals/`, when touching the harness under `products/posthog_ai/eval_harness/`, or when running or debugging those evals (`hogli evals`). Covers suite kinds and discovery, case anatomy, the seeder/synthesizer split, the one-branch scorer patterns, and how to read results. Not for `ee/hogai/eval/ci/` pytest evals, and not for the LLM Analytics product's evaluation features.

73

Quality

92%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

—

The risk profile of this skill

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An excellent, dense, executable body with strong sequencing and a real verification checklist. The one real defect is progressive disclosure: the body signals two `references/*.md` files that are not present in the bundle, so the offloaded detail is unreachable.

Suggestions

Ship the missing `references/authoring-reference.md` and `references/running-evals.md` in the bundle (or, if the bundle is intentionally self-contained, inline the essential field-level and provider-setup details and drop the links) so the signaled references actually resolve.

Add a short "References" section near the top listing the two reference files with one-line purposes, so the offloaded detail is discoverable at a glance rather than buried mid-section (lines 21 and 146).

Verify every other linked path (`eval_harness/README.md`, `evals/AGENTS.md`, `harness/AGENTS.md`, the reporting doc) resolves in the bundle or is clearly marked as an in-repo path outside the skill, to keep navigation trustable.

DimensionReasoningScore

Conciseness

Lean and information-dense with no padding: it assumes Claude's competence (never explains what an eval, sandbox, or Django is) and every line carries a specific contract, path, or invariant; rationales like "so a broken seeder never masquerades as an agent regression" are skill-specific guards, not over-explanation of known concepts, so it clears the 4 anchor's "every token earns its place" bar.

5 / 5

Actionability

Fully executable guidance throughout — copy-paste-ready skeletons for sandboxed and one-shot suites, a real deterministic `Scorer` subclass, a `JudgedScorer` template, concrete `hogli evals` commands, exact field lists, and real import paths; the `...` placeholders are appropriate template holes, so it is not the 4 anchor's "minor gaps".

5 / 5

Workflow Clarity

Clear sequenced sections (Writing a suite → One-shot → Case anatomy → Seeding → Scorers → Running) capped by a numbered Verification checklist with explicit validation commands, plus a debugging feedback loop ("Start debugging with the transcript path... Then open the experiment's Agent logs directory"); the destructive/batch cap does not apply because explicit verification steps are present, so it is not capped at 3 or 4.

5 / 5

Progressive Disclosure

Structure and signaling are good — the body delegates field-level API detail to `references/authoring-reference.md` and provider setup to `references/running-evals.md` with clear contextual links — but those referenced files do not exist in the bundle (no `references/` directory is present), so the signaled references do not resolve and navigation is broken; this is more than the 4 anchor's "minor organization gaps" and lands at 3 where the disclosure is incomplete/unreachable rather than merely imperfect.

3 / 5

Total

18

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description that names concrete capabilities, gives explicit multi-trigger "Use when" guidance, and adds negative boundaries to avoid mis-triggering. The only soft spot is trigger-term synonym coverage, which is naturally limited by the skill's specialized internal audience.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — "write and run evals", "adding or changing eval suites, cases, scorers, seeders, or synthesizers", "touching the harness", "running or debugging those evals" — with comprehensive coverage of the skill's scope; not the 4 anchor because it goes beyond "several specific actions" into full capability coverage.

5 / 5

Completeness

Explicitly answers both what ("Teaches how to write and run evals... sandboxed agent suites... and one-shot suites...") and when ("Use when adding or changing eval suites... when touching the harness... or when running or debugging those evals") with concrete trigger phrases; not 4 because the "when" is specific and multi-triggered rather than merely present.

5 / 5

Trigger Term Quality

Good natural-keyword coverage for the audience ("eval suites", "scorers", "seeders", "synthesizers", "hogli evals", "running or debugging those evals"); not 5 because a few natural synonyms/variations a user might say are missing for such a specialized domain.

4 / 5

Distinctiveness Conflict Risk

Clear niche (evals on a specific named harness) reinforced by explicit negative boundaries ("Not for ee/hogai/eval/ci/ pytest evals, and not for the LLM Analytics product's evaluation features"), giving minimal conflict risk; not 4 because the exclusions remove overlap with the closest related skills.

5 / 5

Total

19

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 2 missing, 4 suspicious

Warning

referenced_paths_exist

Referenced path issues: 4 missing

Warning

Total

14

/

16

Passed

Repository
PostHog/posthog
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.