CtrlK
BlogDocsLog inGet started
Tessl Logo

live-mastracode-instrumentation

Reproduce and measure a Mastra Code runtime bug against a real model by driving the built TUI headlessly in tmux while every layer (network, provider stream, run engine, TUI) appends timestamped JSONL. Use when a bug depends on real provider timing, streaming order, or event flow that mocked tests and fixtures may not reproduce — e.g. wrong token rates, flicker, stuck status, ordering races, or "only happens sometimes with a real model".

77

Quality

96%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An excellent, tightly written operational skill: every step is sequenced with explicit verification (instrumentation-shipped check, cleanup greps, rebuild-after-change rule) and the guidance is grounded in concrete file paths and commands. The only gap is that the instrumentation and per-step analysis are described as rules rather than provided as ready-to-run snippets, which is defensible given the bug-specific nature of the work.

Suggestions

Include a minimal copy-paste JSONL append helper (appendFileSync wrapped in try/catch with the build tag) in step 1 so the logging pattern is executable as-is rather than rule-based.

Provide a small skeleton python3 snippet in step 4 that groups records by step and prints the per-layer duration table, since the analysis loop is the step most likely to stall without a starting point.

DimensionReasoningScore

Conciseness

The body is dense and assumes Claude's competence throughout: every sentence carries specifics (hook-point file paths, timing rules like "Take the timestamp before writing... If you are measuring anything finer, push records onto an in-memory array", and warnings like "Setting the working directory is not a sandbox"). No section explains concepts Claude already knows, and no padding could be trimmed without losing information.

5 / 5

Actionability

Steps 2, 3, and 5 are fully copy-paste ready (pnpm build commands, complete tmux send-keys/capture-pane sequences, cleanup with rg verification) and step 1 names exact functions and files (e.g. "convertFullStreamChunkToMastra... in packages/core/src/stream/aisdk/v5/transform.ts"). It falls short of the 5 anchor only because the instrumentation code itself and the step-4 analysis are described via rules and a hook-point table rather than given as executable snippets — a minor, largely justified gap since the instrumentation is inherently bug-specific.

4 / 5

Workflow Clarity

A clear five-section sequence with explicit validation checkpoints and feedback loops: "If rg finds nothing in a dist, the run will not produce that layer's records", cleanup verification that "must print nothing", "read it: only intended changes remain", "confirm with rg before saying they are gone", and "Rebuild and restart the tmux session after every code change". This matches the 5 anchor's validate-fix-retry pattern even for the risky cleanup phase.

5 / 5

Progressive Disclosure

No bundle files exist (references/, scripts/, assets/ are absent) and none are needed: the skill is a single self-contained workflow with clear section headers, a compact hook-point table, and nothing that clearly belongs in a separate file. Navigation is easy via the numbered section structure, matching the well-organized single-file pattern.

5 / 5

Total

19

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A model description: concrete actions, an explicit and well-phrased 'Use when' clause with natural symptom-level triggers, and a clearly distinct niche. Third-person imperative voice is used correctly and there is no fluff or over-claiming.

DimensionReasoningScore

Specificity

It lists multiple concrete actions in third-person imperative form: "Reproduce and measure a Mastra Code runtime bug against a real model by driving the built TUI headlessly in tmux while every layer (network, provider stream, run engine, TUI) appends timestamped JSONL". Coverage of what the skill does is comprehensive with no vague filler, fitting the 5 anchor rather than the 4 anchor's "minor gaps".

5 / 5

Completeness

It explicitly answers both questions: the 'what' (reproduce and measure a runtime bug via a headlessly driven, multi-layer-instrumented live run) and the 'when' ("Use when a bug depends on real provider timing, streaming order, or event flow that mocked tests and fixtures may not reproduce"). This matches the 5 anchor's concrete-trigger-phrase standard exactly.

5 / 5

Trigger Term Quality

The symptom examples are natural phrases a user would actually say: "wrong token rates, flicker, stuck status, ordering races, or 'only happens sometimes with a real model'", plus "real provider timing, streaming order, or event flow" and "mocked tests and fixtures may not reproduce". This covers the realistic trigger space comprehensively, matching the 5 anchor rather than 4's 'a few natural terms missing'.

5 / 5

Distinctiveness Conflict Risk

The niche is unmistakable — live multi-layer instrumentation of a Mastra Code TUI run against a real provider — and the trigger conditions (real-model timing/streaming bugs that mocks fail to reproduce) would not fire for generic debugging or testing skills. Minimal conflict risk, matching the 5 anchor.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mastra-ai/mastra
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.