CtrlK
BlogDocsLog inGet started
Tessl Logo

chat-perf

Run chat perf benchmarks and memory leak checks against the local dev build or any published VS Code version. Use when investigating chat rendering regressions, validating perf-sensitive changes to chat UI, or checking for memory leaks in the chat response pipeline.

71

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An excellent, dense body: every section teaches tool-specific facts (gating metrics, thresholds, retention quirks, mock-server wiring) with executable commands throughout and explicit statistical validation loops. Its only weaknesses are mild redundancy in the path-heavy command examples and inline storage of deep-dive detail that a reference file would absorb better.

DimensionReasoningScore

Conciseness

Nearly every token carries tool-specific knowledge Claude cannot infer (flag defaults, artifact retention windows, which metrics gate vs. inform, the gc_stats timing-corruption caveat). Minor over-explanation remains — four near-identical escaped-macOS-path commands across the local-build and settings-override examples, and the no-baseline callout paragraph could tighten — fitting 'efficient; minor instances of over-explanation', not 5.

4 / 5

Actionability

Fully copy-paste-ready throughout: quick-start invocations with flags, complete gh CLI commands with --jq filters for listing runs and artifacts, flag tables with defaults and repeatable-arg notes, exit-code semantics, and numeric interpretation bands for leak results. Matches the 'fully executable, covers common cases' anchor.

5 / 5

Workflow Clarity

Sequences are explicit with real validation and feedback loops: statistical significance (Welch's t-test, p < 0.05) gates the regression verdict, the --resume flow gives an inconclusive→add-iterations→recompute loop, exit codes define pass/fail, and regression pinpointing is a 4-step numbered procedure including a confirm-it's-real-vs-noise checkpoint. No destructive/batch operations requiring missing validation.

5 / 5

Progressive Disclosure

No bundle files exist, so all ~320 lines are inline — but with well-organized headers, tables, an architecture map, and a Related skills section, navigation is easy. Deep-dive material (mock-server CAPI endpoint details, CI artifact/retention minutiae) reads like content that belongs in one-level-deep reference files rather than the overview, keeping it at 'good structure; minor organization gaps' rather than the ideal split of the 5 anchor.

4 / 5

Total

18

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person, concrete, with explicit 'Use when...' triggers covering the main use cases. It clearly answers both what the skill does and when to invoke it, with only minor room to broaden trigger synonyms and enumerate more capabilities.

DimensionReasoningScore

Specificity

Concrete actions 'Run chat perf benchmarks and memory leak checks' with explicit targets 'local dev build or any published VS Code version' — several specific actions, but coverage stops at the two headline capabilities rather than comprehensively listing what the skill covers (build comparison, CI gating).

4 / 5

Completeness

Explicitly answers both 'what' (run perf benchmarks and memory leak checks against dev or published VS Code builds) and 'when' ('Use when investigating chat rendering regressions, validating perf-sensitive changes to chat UI, or checking for memory leaks') with concrete trigger phrases — a direct match for the top anchor.

5 / 5

Trigger Term Quality

Natural phrases a user would say — 'chat rendering regressions', 'perf-sensitive changes to chat UI', 'memory leaks in the chat response pipeline' — but common synonyms like 'slow', 'laggy', or 'latency' are absent. Fits the 'good keyword coverage; a few natural terms missing' anchor, not 3 (terms are natural, not generic) nor 5 (synonym gaps exist).

4 / 5

Distinctiveness Conflict Risk

A clear niche (VS Code chat performance and memory leaks) with distinct, unambiguous triggers — minimal realistic conflict with other skills; the scope is specific enough that it would not fire for generic perf or non-chat work.

5 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 4 missing, 3 deeper-than-1-level

Warning

Total

15

/

16

Passed

Repository
posit-dev/positron
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.