CtrlK
BlogDocsLog inGet started
Tessl Logo

evaluate-sessions

Evaluate the build trail of a PR — read the claude -p sessions the harness ran to build it, find where the project's context (Expert / AGENTS.md / skill / spec) served or failed the agents, then capture the learnings as evals (regression tests over the harness's own skills/context) and context fixes. Use when resolving a STUCK (diagnosis-first), auditing how a converged PR was built, or auditing a /learn memory PR. Human-driven and conversational — the trail-evaluating sibling of /evaluate-pr. Outcomes land on a branch (the PR's, or a fresh capture branch if you'll discard the PR) and reach memory via merge + /learn. Triggers - evaluate-sessions, evaluate sessions, review the build trail, diagnose stuck, audit how this was built, session observability.

71

Quality

89%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

The canonical home for this skill is evaluate-sessions in tdg-ninja/context-specs-factory-ai

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured operating guide: the Step 0–7 flow is explicit with human-in-the-loop checkpoints and feedback loops, and bundle layout is exemplary — every reference is purpose-annotated, real, and one level deep. The main weaknesses are repetition (core rules restated three to four times across philosophy, steps, contract, and nevers) and motivational framing that could be tightened, plus landing commands deferred to a reference rather than inlined.

Suggestions

State the never-write-main / never-auto-merge rule once in 'Hard nevers' and cross-reference it from S5, Step 6, and the invocation contract instead of restating it in each place.

Compress the 'The bigger picture — say it to the human, because it's the point' paragraph to a single sentence and cut the repeated 'clean trail produces nothing' framing (keep it in S4 or Hard nevers, not both plus Step 5).

Inline the exact detached-checkout + push command sequence for the default landing path in Step 6 (currently only in references/outcomes.md) so the most common path is copy-paste ready from the body.

DimensionReasoningScore

Conciseness

The body is mostly project-specific and assumes Claude's intelligence (no generic concept explanations), but the same rules are restated repeatedly — 'never write main / never auto-merge' appears in S5, Step 6 ('Never write `main`; never auto-merge'), the invocation contract ('never on `main`'), and Hard nevers; 'a clean trail produces nothing' appears in S4, Step 5, and Hard nevers — plus the motivational 'The bigger picture — say it to the human, because it's the point' paragraph. Not a 2: the length is largely load-bearing discipline for this project, not padded teaching of known concepts.

3 / 5

Actionability

Concrete, runnable guidance throughout: 'scripts/resolve-sessions.sh <PR#|feature>', triage by 'high `Attempt`, non-zero `Exit`', four named lenses, 'evals/<name>/run-eval.sh', 'git push origin HEAD:<branch>', 'branch `capture/<slug>` off `origin/main`', 'git checkout main'. Not a 5: the exact detached-checkout command sequence and eval authoring detail are deferred to references, so the most common path is not fully copy-paste ready from the body alone; not a 3: the mechanical steps that are inlined are executable as written.

4 / 5

Workflow Clarity

Steps 0–7 are explicitly sequenced with validation checkpoints at every risky juncture: Step 0 clean-tree preflight, Step 2 'Tell the human your triage in two lines before diving', Step 4 joint defect-vs-inherent-difficulty classification, Step 6 'Confirm the human's explicit choice' before any landing, and Step 7 return to main. Feedback loops are present (missing-trace degradation note, right-reason eval check, classification loop) and 'Hard nevers' serves as a checklist. The destructive/batch cap does not apply: git operations are gated by the preflight check and explicit human confirmation.

5 / 5

Progressive Disclosure

A clear overview body points to four well-signaled references, each annotated with its purpose and a 'hackable seam' note (trace-reading, evals, outcomes, observability-tooling) plus one script; all cited paths exist on disk and none of the reference files cross-reference each other, so references are one level deep. Detail is appropriately deferred ('Full detail and the project-folder layout live in `references/outcomes.md`'). Not a 4: navigation is easy and the split is clean, with no orphaned or buried content.

5 / 5

Total

17

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description explicitly and concretely states both what the skill does and when to use it, with a dedicated natural-language trigger list and explicit disambiguation from the sibling /evaluate-pr skill. Its only flaw is a second-person clause ('if you'll discard the PR') that violates the third-person voice guideline, which costs one point on specificity.

DimensionReasoningScore

Specificity

Lists multiple concrete actions ('read the claude -p sessions the harness ran to build it', 'find where the project's context (Expert / AGENTS.md / skill / spec) served or failed the agents', 'capture the learnings as evals... and context fixes') with comprehensive coverage for the skill's scope — anchor 5 territory, reduced by 1 per the judging guideline for the second-person clause 'a fresh capture branch if you'll discard the PR'. Not a 3: far more than 1-2 concrete actions; not a 5: the third-person voice violation triggers the mandated penalty.

4 / 5

Completeness

Clearly answers both 'what' (evaluate the build trail, find where context served or failed the agents, capture evals and context fixes) and 'when' with an explicit 'Use when resolving a STUCK..., auditing how a converged PR was built, or auditing a /learn memory PR' clause plus concrete trigger phrases — matching the anchor-5 example structure. The missing-'Use when' cap does not apply.

5 / 5

Trigger Term Quality

Explicit trigger list — 'Triggers - evaluate-sessions, evaluate sessions, review the build trail, diagnose stuck, audit how this was built, session observability' — plus 'Use when resolving a STUCK (diagnosis-first), auditing how a converged PR was built, or auditing a /learn memory PR' gives comprehensive natural phrasings including variants of the skill name and scenario forms. Not a 4: no common in-domain variations are missing.

5 / 5

Distinctiveness Conflict Risk

Explicitly positions against its closest sibling ('the trail-evaluating sibling of /evaluate-pr') and uses niche-specific triggers (STUCK, build trail, /learn audit, session observability) — a clear niche with distinct triggers and minimal conflict risk. Not a 4: the overlap risk with /evaluate-pr is actively disambiguated rather than merely minor.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
tdg-ninja/context-specs-claude-code
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.