CtrlK
BlogDocsLog inGet started
Tessl Logo

evaluate-sessions

Evaluate the build trail of a PR — read the claude -p sessions the harness ran to build it, find where the project's context (Expert / AGENTS.md / skill / spec) served or failed the agents, then capture the learnings as evals (regression tests over the harness's own skills/context) and context fixes. Use when resolving a STUCK (diagnosis-first), auditing how a converged PR was built, or auditing a /learn memory PR. Human-driven and conversational — the trail-evaluating sibling of /evaluate-pr. Outcomes land on a branch (the PR's, or a fresh capture branch if you'll discard the PR) and reach memory via merge + /learn. Triggers - evaluate-sessions, evaluate sessions, review the build trail, diagnose stuck, audit how this was built, session observability.

72

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, human-in-the-loop forensic workflow: the 8-step guided flow has explicit validation checkpoints and failure handling, and progressive disclosure is exemplary — every referenced bundle file exists, is one level deep, and is purpose-labeled. The main weakness is repetition: the same safety rules and motivational framing appear three to four times across the philosophy section, the steps, the output contract, and the Hard nevers.

Suggestions

State the never-write-main / never-auto-merge rule once (e.g., in Hard nevers) and reference it from S5, Step 6, and the output contract instead of restating it four times — the repetition is the largest token cost in the body.

Trim the motivational framing ("The bigger picture — say it to the human, because it's the point", the flywheel-is-the-point passages) to one short line; the framing reappears in references/evals.md, so paying for it twice is pure overhead.

Consider collapsing the S1–S9 philosophy list into the steps that use them (most already carry the S-tag) or moving the full list to a reference file, keeping only the routing table (S9) inline since it's the one consulted mid-flow.

DimensionReasoningScore

Conciseness

Mostly efficient, dense, project-specific guidance Claude cannot infer (the PR-comment session contract, the four routing destinations, the branch-landing procedure) — but the same rules are restated multiple times: "Never write `main`" appears in S5, Step 6 ("Never write `main`; never auto-merge"), the output contract ("never on `main`"), and "Hard nevers", and motivational framing ("say it to the human, because it's the point", "This is the flywheel...") pads without instructing. This fits the score-3 anchor 'mostly efficient but includes some unnecessary explanation or could be tightened' — more repetition than the score-4 'minor instances', but the bulk is genuinely non-inferable knowledge, so not the score-2 'several unnecessary explanations'.

3 / 5

Actionability

The flow gives concrete, executable guidance: "Run `scripts/resolve-sessions.sh <PR#|feature>`", the `~/.claude/projects/*/<id>.jsonl` glob, "Detached-checkout the PR head, commit the eval/context changes, `git push origin HEAD:<branch>`", and "branch `capture/<slug>` off `origin/main`". Not a 5 because some execution detail is prose-delegated to references ("with exact git commands" in outcomes.md) rather than inline copy-paste-ready commands; not a 3 because the guidance is genuinely executable, not high-level hints or pseudocode.

4 / 5

Workflow Clarity

Steps 0–7 are explicitly sequenced with checkpoints at every risky point: preflight ("Confirm the working tree is clean; if there's WIP, ask the user to stash or commit first"), an explicit failure path (the script "flagging any whose file is missing (a remote/CI run — see the degradation note)"), human validation before analysis ("Tell the human your triage in two lines before diving"), and human confirmation before any write ("Confirm the human's explicit choice"), plus a "Hard nevers" checklist and a final restore step ("`git checkout main` so the user ends where they started"). This matches the score-5 anchor: clear sequence, explicit validation, error recovery, and a checklist.

5 / 5

Progressive Disclosure

The body is an overview that splits detail into four real, one-level-deep references, each signaled with its purpose and a hackable-seam note ("`references/trace-reading.md` — where session IDs come from (the PR comment), how to locate and read the local JSONL traces, the triage map, and the four reading lenses"), plus one script; all five bundle files exist and are substantive. This matches the score-5 'clear overview with well-signaled one-level-deep references; content appropriately split' anchor — no nesting, no buried references, no detail that belongs in a file left inline.

5 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A dense but non-padded description that answers what, when, and how-it-differs explicitly, with a dedicated natural-language trigger list. Third-person voice throughout and no over-claims. Fully matches the top anchors on all four dimensions.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions covering the full workflow: "read the claude -p sessions the harness ran to build it", "find where the project's context (Expert / AGENTS.md / skill / spec) served or failed the agents", "capture the learnings as evals (regression tests over the harness's own skills/context) and context fixes", and "Outcomes land on a branch". Coverage is comprehensive (input, analysis, both output types, and landing), matching the score-5 anchor rather than the score-4 anchor's 'minor gaps in coverage'.

5 / 5

Completeness

Both questions are explicitly answered: what it does (the multi-clause action description above) and when to use it ("Use when resolving a STUCK..., auditing how a converged PR was built, or auditing a /learn memory PR" plus the trigger list). This matches the score-5 anchor 'clearly and explicitly answers both what AND when with concrete trigger phrases', not the score-4 anchor where 'when' could be more explicit.

5 / 5

Trigger Term Quality

It closes with explicit natural-language triggers — "Triggers - evaluate-sessions, evaluate sessions, review the build trail, diagnose stuck, audit how this was built, session observability" — plus "Use when resolving a STUCK (diagnosis-first), auditing how a converged PR was built, or auditing a /learn memory PR". Synonym coverage (evaluate sessions / evaluate-sessions / review the build trail / audit how this was built) matches the comprehensive score-5 anchor; there is no obvious natural phrasing a user of this workflow would say that is missing.

5 / 5

Distinctiveness Conflict Risk

It explicitly positions itself against sibling skills — "the trail-evaluating sibling of /evaluate-pr" and "reach memory via merge + /learn" — carving a clear niche (evaluating how something was built, not what was built). Trigger phrases (build trail, session observability, diagnose stuck) are distinct from generic review/eval wording, matching the score-5 'clear niche with distinct triggers' anchor.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
tdg-ninja/context-specs-factory-ai
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.