CtrlK
BlogDocsLog inGet started
Tessl Logo

customizations-in-the-agent-host

Architecture and hard-won debugging lessons for customization enablement (plugins, MCP servers, agents, skills, instructions) in the agent host. Use when changing how customizations are discovered, published, enabled/disabled, or handed to a provider SDK; when a customization shows the wrong enabled state in the UI; or when a disabled MCP server or plugin is still reaching the model.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

90%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An excellent, dense body: it documents non-obvious invariants, real production-bug lessons, exact commands, and a signal-ranked debugging playbook with no filler. The main improvement opportunities are making the change→validate workflow an explicit ordered loop and splitting the gotchas/enforcement detail into reference files to slim the main body.

Suggestions

Add an explicit ordered workflow for making a change (edit → npm run typecheck-client → unit tests → integration tests → fix and re-run on failure) rather than leaving it implied across the Gotchas and Validation sections.

Move the 'Gotchas that cost real time' list and the per-provider enforcement table into a reference file (e.g. references/gotchas.md), keeping SKILL.md as the overview with a well-signaled one-level-deep link.

Include a short worked example of reading the enablement overlay end-to-end (base vs persistent decision vs inherited value) to make the most-important invariant concrete for new readers.

DimensionReasoningScore

Conciseness

The body never explains concepts Claude already knows — every paragraph carries repo-specific invariants (the '_clientGlobalEnablement' vs '_persistent' overlay, the 'childEnablement' discriminator), concrete bug postmortems, or exact commands. The narrative framing ('the model lies', 'self-defeating write') conveys severity rather than padding, so it is not merely the anchor-4 'minor instances of over-explanation'.

5 / 5

Actionability

It gives exact file paths, function names ('targetForMcpServer()', 'createNoopCustomizationEnablementService()'), copy-paste-ready commands ('npm run typecheck-client', './scripts/test.sh --grep "customizationEnablement"', 'rm -rf .build/electron && npm run electron'), a log signature to grep for, and a concrete reproduction recipe — fully executable guidance covering the common cases.

5 / 5

Workflow Clarity

The debugging playbook ranks signals and gives a sequenced cross-session reproduction recipe, and the Validation section lists exact commands, so most checkpoints are present. It stops short of anchor 5 because the primary change workflow (edit → typecheck → unit test → integration test → fix on failure) is implied by scattered sections rather than presented as an explicit ordered loop with feedback steps.

4 / 5

Progressive Disclosure

Sections are clearly headed and well-organized, and cross-references to companion skills ('agent-host-logs', 'launch') are well signaled. However, at ~124 lines with no bundle files, detailed material such as the gotchas list and per-provider enforcement table could live in reference files, and the under-50-line exception for a top score does not apply.

4 / 5

Total

18

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it names a precise niche, enumerates the customization types it covers, and provides an explicit 'Use when...' clause with three concrete, naturally-phrased trigger scenarios including UI-state and model-exposure failure symptoms. The only minor weakness is that the 'what' is expressed as a coverage area rather than parallel concrete actions.

DimensionReasoningScore

Specificity

The description lists several specific capabilities and failure modes — 'changing how customizations are discovered, published, enabled/disabled, or handed to a provider SDK', 'wrong enabled state in the UI', 'a disabled MCP server or plugin is still reaching the model' — matching the several-specific-actions anchor. It falls short of 5 because the 'what' clause ('Architecture and hard-won debugging lessons for...') is a coverage statement rather than a list of concrete parallel actions.

4 / 5

Completeness

It clearly states what the skill contains (architecture and debugging lessons for customization enablement in the agent host) and follows an explicit 'Use when...' clause with three concrete trigger scenarios, matching the anchor-5 example pattern.

5 / 5

Trigger Term Quality

It enumerates the full set of customization kinds ('plugins, MCP servers, agents, skills, instructions') and uses natural phrases a developer would actually say, such as 'disabled MCP server or plugin is still reaching the model' and 'wrong enabled state in the UI'. This is comprehensive natural-term coverage including synonyms for the domain.

5 / 5

Distinctiveness Conflict Risk

The scope is tightly narrowed to 'customization enablement ... in the agent host' with its own enumerated subject types and failure symptoms, giving it a clear niche with minimal overlap with any other skill's triggers.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
microsoft/vscode
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.