CtrlK
BlogDocsLog inGet started
Tessl Logo

analyze-sessions

Analyzes your local Copilot CLI sessions for dotnet/maui to drive iterative improvements to the PR-review agent (and other agents, skills, and instruction files). Runs a select → extract → score → judge → cluster → propose → emit-eval loop: a deterministic core ranks your worst / most-expensive sessions, then the agent rubric-tags recurring failure modes, proposes concrete repo edits, and emits a vally guard-eval per failure mode so each one becomes a regression test. Triggers on: "analyze my recent maui sessions", "what's making my agent runs expensive", "find failure modes in my Copilot sessions", "turn my session failures into guard evals". LOCAL-ONLY — never uploads, shares, or posts transcripts. Do NOT use for: reviewing a single PR (use pr-review), running tests, or analyzing a GitHub issue.

72

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An excellent, highly actionable skill body: copy-paste commands, a transparent scoring model, explicit validation, and a completion checklist. The main improvements are trimming repeated privacy/path warnings and moving the YAML template and rubric table into reference files. Only minor tightening separates this from top marks.

Suggestions

State the privacy contract once (keep the 'Privacy & safety' section, drop or shrink the intro callout) — the local-only/no-upload promise is currently made twice in the body, and the '.github/evals/' path warning also appears in both 'Outputs' and 'Phase 6'.

Move the guard-eval YAML template (and optionally the 8-row rubric-tagging table) into a reference file like 'references/eval-template.md', keeping only the two key requirements (structural floor + one LLM judge) inline.

Consider replacing or compressing the 17-line ASCII architecture diagram into 2–3 sentences — the 'one engine, two front doors' prose already conveys the same structure.

DimensionReasoningScore

Conciseness

The body is dense and mostly earns its tokens (scoring formula, input table, command examples), but there is repetition that could be trimmed: the privacy/local-only contract appears in both the intro callout and the full 'Privacy & safety' section, and the 'do not use a generic .github/evals/ location' warning is stated twice, plus a 17-line ASCII diagram restates what the surrounding text already says.

4 / 5

Actionability

Guidance is fully executable: copy-paste 'pwsh -NoProfile -File …Get-SessionAnalysis.ps1 -Last 15 -Top 5 -OutputDir "$ARTIFACTS_DIR" -Json' invocations, the exact transparent scoring formula '2·tool_failures + 1.5·retries + 5·(errors+aborts)…', a complete guard-eval YAML template, and a concrete validation command 'npx -y @microsoft/vally-cli@0.14.0 lint --eval-spec <path> --strict'.

5 / 5

Workflow Clarity

The 6 phases are clearly sequenced, each with concrete instructions; there is an explicit validation checkpoint (validate every emitted eval with 'lint --strict'), error-recovery guidance in Phase 2 ('do not re-open raw transcripts unless a digest is ambiguous'), and a completion-criteria checklist — and the batch operation is fully covered by validation, so no cap applies.

5 / 5

Progressive Disclosure

Good structure with a well-signaled, verified one-level-deep reference ('See references/design-rationale.md' — the file exists and points only to external URLs, not nested skill files), but some content that could live in a reference is inlined in SKILL.md — notably the 28-line guard-eval YAML template and the full rubric-tagging table.

4 / 5

Total

18

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete actions, explicit trigger phrases, and clear negative boundaries. Its only real flaws are the second-person voice ('your sessions') and a few missing natural trigger variations. It would score near-perfectly after a shift to third person and a slightly broader trigger list.

Suggestions

Rewrite in third person: 'Analyzes local Copilot CLI sessions…' instead of 'your local Copilot CLI sessions' and 'the user's worst sessions' instead of 'your worst / most-expensive sessions' — the second-person voice is what costs specificity a point.

Add one or two common trigger variations users would naturally say, e.g. 'why are my Copilot runs failing' or 'analyze my session logs/transcripts', to round out trigger-term coverage.

Consider mentioning the read-only/local nature earlier and trimming 'to drive iterative improvements to the PR-review agent (and other agents, skills, and instruction files)' to a shorter benefit clause — the parenthetical adds length without adding triggerable keywords.

DimensionReasoningScore

Specificity

Quotes like "ranks your worst / most-expensive sessions", "rubric-tags recurring failure modes", "proposes concrete repo edits", and "emits a vally guard-eval per failure mode" list multiple concrete, comprehensive actions that would merit a 5, but the description uses second-person voice ("Analyzes your local Copilot CLI sessions", "your agent runs"), which per the judging guidelines reduces the specificity score by 1.

4 / 5

Completeness

It clearly answers 'what' ("Runs a select → extract → score → judge → cluster → propose → emit-eval loop") and 'when' with concrete trigger phrases ("Triggers on: …") plus explicit negative guidance ("Do NOT use for: reviewing a single PR (use pr-review)…"), exactly matching the anchor for a clear, explicit what + when with concrete triggers.

5 / 5

Trigger Term Quality

"Triggers on: 'analyze my recent maui sessions', 'what's making my agent runs expensive', 'find failure modes in my Copilot sessions', 'turn my session failures into guard evals'" provides good natural-phrase coverage across cost and failure angles, but a few common variations users would say are missing (e.g. 'session logs', 'transcripts', 'why do my agent runs fail').

4 / 5

Distinctiveness Conflict Risk

It carves a clear niche (local Copilot CLI session mining for dotnet/maui) with distinct triggers, and the "Do NOT use for" clause routes neighboring use cases (single-PR review, running tests, GitHub issue analysis) to other skills, minimizing conflict risk.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
dotnet/maui
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.