Content
85%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body is highly actionable with a well-sequenced, validated workflow and domain-specific detail that mostly earns its tokens. Its main weakness is progressive disclosure: a prominently referenced payload reference file is missing, and some reference-grade configuration detail is inlined in SKILL.md instead.
Suggestions
Create the missing references/evaluation-payload.md (or remove the dangling links) so the two references to it resolve — currently both [references/evaluation-payload.md](references/evaluation-payload.md) links are broken.
Move the detailed settle-config bounds and globals reference tables (sections 2.2) into references/evaluation-payload.md, keeping only the decision guidance inline in SKILL.md.
Tighten the session-target skipped-evaluation and retention prose in 2.2 to the essential rules to lift conciseness toward the top anchor.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | The body is dense with PostHog-specific domain detail (settle strategies, globals, provider-key gating) that Claude would not already know, and largely assumes competence rather than padding; it earns 4 over 5 because a few prose sections (e.g. the session-target bounds and skipped-evaluation mechanics) could be tightened or moved to the reference, and over 3 because it is not "noticeably verbose" with generic explanation. | 4 / 5 |
Actionability | Provides copy-paste-ready JSON payloads for hog and llm_judge eval creation, exact `conditions` shape, a runnable SQL verification query, and concrete `generate-app-url` calls — covering the common cases per the score-5 anchor. | 5 / 5 |
Workflow Clarity | Sequences a two-phase process (1.1–1.3 then 2.1–2.6) with an explicit validation checkpoint (2.5 verify scope via SQL before enabling) and feedback loop (sample, review first live results, adjust rollout), satisfying the score-5 anchor and avoiding the batch-operation cap since validation is present. | 5 / 5 |
Progressive Disclosure | Structure and sectioning are good, but the body twice signals [references/evaluation-payload.md](references/evaluation-payload.md) as the home for "every field, the config schemas, the exact conditions shape" and that file does not exist, so navigation is broken; additionally reference-grade detail (session settle bounds, the globals table) is inlined rather than split out, matching the score-3 anchor's "references present but not clearly signaled / content that should be separate is inline" better than the 4 anchor's "minor organization gaps". | 3 / 5 |
Total | 17 / 20 Passed |