CtrlK
BlogDocsLog inGet started
Tessl Logo

chat-perf

Run chat perf benchmarks and memory leak checks against the local dev build or any published VS Code version. Use when investigating chat rendering regressions, validating perf-sensitive changes to chat UI, or checking for memory leaks in the chat response pipeline.

74

Quality

92%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

85%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-structured skill body with copy-paste commands, explicit validation feedback loops, and clean one-level references. The only real weakness is moderate verbosity in the statistics sections, which re-explain methodology a user of this skill could derive from the script output.

Suggestions

Tighten the 'Statistical significance' and 'Statistics' sections: keep the operational guidance (p<0.05 threshold, 5+ runs for stability, exit codes) but drop the derivations and sample-size tables, since the tool already prints a --resume hint when results are inconclusive.

Consider collapsing the duplicated statistical-method prose (Welch's t-test is described in both 'Statistical significance' and 'Statistics') into a single concise note to avoid repetition.

The 'Comparing two local builds' examples repeat long escaped paths three times; one representative example plus a note that both flags accept local paths would preserve actionability while cutting tokens.

DimensionReasoningScore

Conciseness

Mostly lean and executable, but the Statistics and Statistical-significance sections re-explain methodology (Welch's t-test rationale, cv thresholds, sample-size guidance) that could be trimmed; a few passages include explanatory context beyond the bare instruction, fitting the 'mostly efficient but could be tightened' anchor rather than the fully lean top anchor.

2 / 3

Actionability

Provides copy-paste-ready bash commands, a complete flags table with defaults, exact script paths, and concrete result-interpretation thresholds — fully executable guidance matching the top anchor.

3 / 3

Workflow Clarity

Multi-step flows (resume for confidence, add-a-scenario, build-mode comparison) are explicitly sequenced with validation checkpoints — statistical-significance gating, the resume hint when inconclusive, and exit-code semantics act as feedback loops for these potentially noisy batch operations.

3 / 3

Progressive Disclosure

A single well-organized SKILL.md with clearly signaled sections; script paths are one level deep and explicitly named, and the 'Related skills' pointers are clean one-level references, matching the well-organized top anchor (no bundle files exist, consistent with the simple-skill scoring note).

3 / 3

Total

11

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A concise, well-constructed description that names concrete actions, targets, and explicit use-when triggers in third person. It cleanly satisfies all four description dimensions with no padding or over-claims.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'Run chat perf benchmarks', 'memory leak checks' — and names concrete targets ('local dev build or any published VS Code version'), matching the 'Lists multiple specific concrete actions' anchor.

3 / 3

Completeness

Explicitly answers both what (perf benchmarks + memory leak checks) and when via a dedicated 'Use when' clause with concrete triggers, matching the top completeness anchor.

3 / 3

Trigger Term Quality

Natural phrases a user would say appear verbatim — 'chat rendering regressions', 'perf-sensitive changes to chat UI', 'memory leaks in the chat response pipeline' — giving good coverage of real-world triggers.

3 / 3

Distinctiveness Conflict Risk

Scoped tightly to chat performance in VS Code with distinct triggers (rendering regressions, leak detection), making it unlikely to fire for unrelated skills.

3 / 3

Total

12

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 4 missing, 3 deeper-than-1-level

Warning

Total

15

/

16

Passed

Repository
posit-dev/positron
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.