CtrlK
BlogDocsLog inGet started
Tessl Logo

chat-perf

Run chat perf benchmarks and memory leak checks against the local dev build or any published VS Code version. Use when investigating chat rendering regressions, validating perf-sensitive changes to chat UI, or checking for memory leaks in the chat response pipeline.

71

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-structured reference for a complex benchmarking tool, with concrete commands and a validated, feedback-driven debugging workflow. Its main weakness is length: it inlines substantial CI and deep-dive material that, in a larger bundle, would warrant separate reference files.

Suggestions

Move the CI/historic-runs section (gh CLI recipes, artifact-retention tables, bisect deep-dive) into a references file (e.g. CI.md) and keep SKILL.md as an overview with a one-line pointer, improving progressive disclosure and token budget.

Tighten the flag tables by collapsing rarely-used flags into a single 'See --help' line, reserving table rows for the flags that drive most runs.

Consider a short 'Validation checklist' callout summarizing the statistical-significance + exit-code + resume loop so the feedback pattern is visible without reading the full Statistics section.

DimensionReasoningScore

Conciseness

Dense and almost entirely operationally relevant tool-specific knowledge (flag tables, statistics, exit codes) rather than concepts Claude already knows; a few sections like the CI artifact-retention tables and historic-run gh CLI detail could be trimmed or moved out, keeping it just above the midpoint.

4 / 5

Actionability

Copy-paste ready throughout: exact 'npm run perf:chat -- ...' invocations, a full flag table with defaults, executable gh CLI commands, and concrete file paths covering the common cases.

5 / 5

Workflow Clarity

The 'Pinpointing where a metric regressed' section is a numbered checklist with explicit validation (statistical significance, raw-table inspection) and a feedback loop ('--resume' for inconclusive results), and exit codes define pass/fail verdicts.

5 / 5

Progressive Disclosure

Well-organized with clear section headers and a 'Related skills' pointer, but it is a large single-file skill with no bundle references; bulk CI/deep-dive detail that could live in separate reference files is inlined, a minor organization gap.

4 / 5

Total

18

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description that clearly answers both what the skill does and when to use it with natural trigger phrases. It is concise, third-person, and well-scoped to a distinct niche.

DimensionReasoningScore

Specificity

Names two concrete actions ('Run chat perf benchmarks and memory leak checks') and specific targets ('local dev build or any published VS Code version'), with only minor coverage gaps relative to the comprehensive 5-anchor.

4 / 5

Completeness

Explicitly states both what it does (run perf benchmarks and memory leak checks) and when to use it via a concrete 'Use when...' clause enumerating three trigger scenarios.

5 / 5

Trigger Term Quality

Includes natural phrases a user would say ('chat rendering regressions', 'perf-sensitive changes to chat UI', 'memory leaks in the chat response pipeline') with good coverage, missing only a few synonyms.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (VS Code chat perf benchmarking + leak detection) with distinct triggers and minimal overlap risk with other skills.

5 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 4 missing, 3 deeper-than-1-level

Warning

Total

15

/

16

Passed

Repository
posit-dev/positron
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.