Content
71%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body is exceptionally actionable — real code, commands, endpoints, and file paths everywhere — and encodes a large amount of non-obvious, project-specific constraint knowledge. Its weaknesses are structural: a monolithic inline format that buries deep reference material (tracking-bridge PostHog minutiae, failure-report plumbing) in the always-loaded file, and dangling references to docs that do not exist in the bundle.
Suggestions
Split the deep reference material into one-level-deep bundle files — e.g. move the tracking-bridge constraint catalog (lines 357–474) into a references/tracking-bridge.md and the failure-report/error-diagnosis plumbing (lines 77–107) into references/failure-reporting.md, leaving SKILL.md a concise overview with well-signaled links per the progressive_disclosure anchor-5 pattern.
Tighten the rhetorical framing in the always-loaded portion: replace justification sentences like "because a failure nobody can classify is not observable" and "Measured, not assumed — check whether that title is still wrong before trading the duplication back" with the bare rule, keeping the rationale in the split-out reference files.
Resolve the dangling doc pointers — "See the Evals doc", "See the Observability doc", "see the tracking skill" — either by creating those bundle files or by linking to concrete existing paths (as the Key Files table does).
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Almost every sentence carries framework-specific, non-inferable detail ("PostHog's `$ai_*` latency fields are seconds; ours are milliseconds"), so there is no filler explaining concepts Claude already knows. However, the prose is padded with rhetorical framing ("because a failure nobody can classify is not observable", "Measured, not assumed — check whether that title is still wrong before trading the duplication back"), and the ~117-line tracking-bridge constraints section (lines 357–474) is deep reference material inflating every skill load and could be substantially tightened or split out. | 3 / 5 |
Actionability | Fully executable, copy-paste-ready guidance throughout: complete `defineAppConfig`, `defineEval`, `insertExperiment`, and dashboard-route code blocks; exact CLI commands (`agent-native eval promote <runId> --write evals/from-trace.eval.ts`); concrete HTTP endpoints (`GET /_agent-native/observability/traces/:runId`); and real source paths in the Key Files table. Specific examples cover the common cases for each pillar. | 5 / 5 |
Workflow Clarity | Sequences are clear with most checkpoints present: the promote-trace-to-CI-eval flow with "Truncated runs fail closed"; the CI gate that "exits non-zero if any eval is below its threshold"; the human-audit flow with "Allow feedback and regeneration before explicit approval; never auto-apply"; and the four-part error-report chain (packet → client report → read by id → alerts). The gap keeping it below 5 is that workflows are described narratively rather than as explicit ordered steps, and some (e.g., first-time dashboard setup, experiment lifecycle) lack an explicit validation checkpoint. | 4 / 5 |
Progressive Disclosure | There are no bundle files — this is a single ~470-line monolithic SKILL.md with good headers and tables, but content that clearly belongs in separate reference files is inlined (the tracking-bridge constraint catalog, failure-report plumbing, span-naming rules). Navigation is also weakened by dangling pointers: "See the Evals doc", "See the Observability doc", and "see the tracking skill" reference documents not present in the bundle. | 3 / 5 |
Total | 15 / 20 Passed |