CtrlK
BlogDocsLog inGet started
Tessl Logo

ai-code-mode

LLM-generated TypeScript execution in sandboxed environments: createCodeModeTool() with isolate drivers (createNodeIsolateDriver, createQuickJSIsolateDriver, createQuickJSBunIsolateDriver, createCloudflareIsolateDriver), codeModeWithSnippets() for persistent snippet libraries, trust strategies, snippet storage (FileSystem, LocalStorage, InMemory, Mongo), client-side execution progress via code_mode:* custom events in useChat.

60

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./packages/ai-code-mode/skills/ai-code-mode/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An action-dense, highly executable body whose code examples and Common Mistakes section are excellent. Its weaknesses are structural: a monolithic ~580-line file with no progressive disclosure into reference files, and repeated tool-definition boilerplate across examples that inflates token cost.

Suggestions

Split stable detail into reference files (e.g. references/drivers.md, references/client-integration.md, references/snippets.md) and keep SKILL.md as a lean overview with well-signaled one-level-deep pointers.

Define the fetchWeather example tool once and reuse it (comment '// same fetchWeather tool as above') instead of repeating the identical 9-line toolDefinition in four examples.

Fold the probe/timeout validation checkpoints from Common Mistakes into the setup workflow itself (e.g. 'run probeIsolatedVm() and fall back to QuickJS if incompatible') so the main flow carries its own feedback loop.

DimensionReasoningScore

Conciseness

Prose is lean and assumes competence — no padding explaining what code mode or sandboxing is, and driver trade-offs are compressed into a comparison table. The main tightening opportunity is the identical 9-line fetchWeather toolDefinition repeated in four separate examples. Not 5 because that repeated boilerplate is a noticeable token cost; not 3 because outside the repetition there is almost no over-explanation.

4 / 5

Actionability

Every section ships complete, copy-paste-ready TypeScript: full server route handlers for createCodeModeTool, createCodeMode, and codeModeWithSnippets; driver config blocks with defaults annotated inline; a complete React component for client-side event handling; and concrete wrong/right pairs for each common mistake. Covers the common cases comprehensively.

5 / 5

Workflow Clarity

Setup follows a clear sequence (define tool → create code-mode tool → wire into chat), the driver-selection table guides the key decision point, and Common Mistakes supplies validation checkpoints for a genuinely risky operation (probeIsolatedVm before using isolated-vm, explicit finite timeout, keep secrets out of the sandbox). Not 5 because the main setup flow itself has no explicit validate/recover feedback loop — checkpoints live in a separate mistakes section rather than the workflow steps.

4 / 5

Progressive Disclosure

The single SKILL.md is ~580 lines with no bundle files (references/, scripts/, assets/ do not exist), so all content — including the sizable client-integration React component, snippet storage details, and lazy-tools subsection — is inlined in one file. Section headers and numbered patterns give it real structure, which lifts it above 2, but content that clearly belongs in separate reference files (client integration, driver reference) is inline with only one cross-skill pointer. Not 4 because there are no well-signaled one-level-deep references at all.

3 / 5

Total

16

/

20

Passed

Description

63%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A highly specific, information-dense description that comprehensively enumerates the skill's capabilities via named APIs. Its weaknesses are the complete absence of a 'Use when...' trigger clause (capping completeness) and reliance on API identifiers rather than natural user phrases for triggering.

Suggestions

Append an explicit trigger clause, e.g. 'Use when the user wants LLM-generated code executed in a sandbox, mentions Code Mode, isolate drivers, or persistent snippets in TanStack AI chat apps.'

Add natural-language synonyms alongside API names ('Code Mode', 'code execution sandbox', 'isolates') so users' phrasing matches the description without knowing exact function names.

Briefly disambiguate from standard tool calling (e.g. 'Code Mode instead of regular tool calling for complex multi-step operations') to reduce overlap risk with the sibling tool-calling skill.

DimensionReasoningScore

Specificity

Lists multiple concrete capabilities with named APIs: 'createCodeModeTool() with isolate drivers (createNodeIsolateDriver, createQuickJSIsolateDriver, createQuickJSBunIsolateDriver, createCloudflareIsolateDriver)', 'codeModeWithSnippets() for persistent snippet libraries', 'snippet storage (FileSystem, LocalStorage, InMemory, Mongo)', and 'code_mode:* custom events in useChat' — comprehensive coverage of the skill's surface. Not 4 because coverage spans every major feature area of the skill rather than having minor gaps.

5 / 5

Completeness

The 'what' is clear and detailed (sandboxed execution, drivers, snippets, storage, events), but there is no 'Use when...' clause or equivalent explicit trigger guidance anywhere in the description. Per rubric guideline, a missing 'Use when' clause caps completeness at 3 even with a clear 'what'.

3 / 5

Trigger Term Quality

Relevant domain keywords exist ('LLM-generated TypeScript execution', 'sandboxed environments', 'snippet libraries', 'trust strategies'), but most terms are exact API/function names (e.g. 'createQuickJSBunIsolateDriver') rather than phrases a user would naturally say. Natural variations like 'code mode', 'run LLM code safely', or 'isolates' are missing. Not 4 because the keyword set is skewed toward jargon over natural user language; not 2 because several genuinely relevant domain terms are present.

3 / 5

Distinctiveness Conflict Risk

Naming library-specific APIs (codeModeWithSnippets, code_mode:* events) creates a clear niche that few other skills would match. Not 5 because it overlaps with the closely related tool-calling skill in the same library ('Code Mode is always used on top of a chat experience'), and the description gives no trigger guidance to disambiguate when this skill should fire instead of a sibling.

4 / 5

Total

15

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (604 lines); consider splitting into references/ and linking

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
TanStack/ai
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.