CtrlK
BlogDocsLog inGet started
Tessl Logo

analyse-sessions

Use when reading the session-performance cards to find where agent effort is going — running the report over `.claude/metrics/`, interpreting it without the traps that have already produced two wrong findings, and reporting what is actually supported. Also the step before proposing any change to how work is dispatched.

74

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

96%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplary instruction-only skill: dense with project-specific knowledge, copy-paste-ready commands, explicit validation gates, and a memorable catalog of past failure modes that doubles as an error-check checklist. The only structural note is that progressive disclosure relies entirely on external repo docs rather than bundle files, which is appropriate here but leaves minor navigation gaps.

DimensionReasoningScore

Conciseness

Every section carries project-specific knowledge Claude could not have (the two past wrong findings, `explore_share` null semantics, median 3.6M vs mean 15.4M skew, the `at-open` container death) with zero general-concept padding — e.g. "Nothing else counts as an analysis — a table restated in prose is not one." Not 4 because there is no passage of over-explanation that could be trimmed without losing operational content; the median-vs-mean and correlation discussion is grounded in this corpus's numbers, not textbook explanation.

5 / 5

Actionability

Fully executable commands are given — `bun run --cwd tools/pr-metrics report -- --coverage` / `--waste` / `--json`, and the copy-paste `gh pr diff 882 ... | grep -E '^[+-]\s*"...'` — plus concrete decision tables (declared complexity vs spend) and an exact output spec ("the number, the sample size it rests on, and what would have to be true for it to be wrong"). Not 4 because the commands and examples cover the common cases with no missing key details.

5 / 5

Workflow Clarity

Explicit gated sequence — Step 0 coverage ("write them down before continuing", "If nothing is rated, say so at the top of your answer"), Step 1 run, the four traps as an interpretation checklist, Step 2 report — with per-claim validation built in (each claim must name what would overturn it, and exclusions must be reported). Not 4 because validation checkpoints are explicit and complete, including error-recovery guidance for the exact failure modes that previously occurred; this is a read-only analysis skill so the destructive-operation cap does not apply.

5 / 5

Progressive Disclosure

Well-organized single-file structure with clear section headers and one-level-deep, clearly signaled external pointers ("`wiki/conventions/session-metrics.md` holds the schema... `tools/pr-metrics/README.md` has the commands"); no bundle files exist, so everything operational is appropriately inline. Not 5 because at ~120 lines the body exceeds the under-50-line simple-skill case, and some detail (e.g. the per-metric schema semantics) is deferred but could be more prominently navigable; not 3 because no content that clearly belongs in a separate file is inlined and references are not buried.

4 / 5

Total

19

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: explicit "Use when" triggers, concrete named actions, and a distinctive niche vocabulary that prevents conflict with other skills. The only weakness is that a few natural synonyms a user might say (spend, tokens, cost) are absent from the trigger terms.

DimensionReasoningScore

Specificity

Names three concrete actions with concrete artifacts — "running the report over `.claude/metrics/`", "interpreting it without the traps", "reporting what is actually supported" — which matches the 'several specific actions; minor gaps' anchor. Not 5 because coverage is incomplete (classification, coverage checks, and median handling from the body are absent), and not 3 because more than 1-2 actions are explicitly listed.

4 / 5

Completeness

Explicitly answers both: what ("running the report over `.claude/metrics/`, interpreting it without the traps... and reporting what is actually supported") and when ("Use when reading the session-performance cards... Also the step before proposing any change to how work is dispatched") with concrete trigger phrases, matching the top anchor. Not 4 because the 'when' is already highly explicit and specific, leaving nothing that 'could be more explicit'.

5 / 5

Trigger Term Quality

Good keyword coverage — "session-performance cards", "agent effort", "report", `.claude/metrics/`, "how work is dispatched" — with natural phrases a user in this repo would say. Not 5 because common synonyms users might naturally use ("spend", "tokens", "cost", "sessions") are missing; not 3 since the terms present are relevant and specific rather than merely generic.

4 / 5

Distinctiveness Conflict Risk

A clear niche — session-performance cards, `.claude/metrics/`, dispatch-change proposals — with distinct triggers and minimal conflict risk against generic analysis or reporting skills. Not 4 because no closely related skill would plausibly claim these triggers; the project-specific vocabulary makes mis-triggering unlikely.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
englishstreetventures/osn
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.