CtrlK
BlogDocsLog inGet started
Tessl Logo

measure-telemetry-span

Use when measuring a Sentry performance span locally with an agent-device replay flow on iOS simulator or Android emulator.

66

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong operational skill body: fully executable commands, explicit validation and recovery loops, and disciplined pointers to external docs. Minor gains available from de-duplicating warnings and pushing deeper troubleshooting detail behind references.

DimensionReasoningScore

Conciseness

The body is dense with repo-specific facts (no explaining of concepts Claude already knows), but has minor trim opportunities: the AGENT_DEVICE_STATE_DIR / daemon-kill warning appears in both the Environment paragraph and the Single-replay section, and 'Platform coverage' spends several lines on UI internals that could be tightened. Not a 5 for those redundancies; not a 3 because there is no real padding or filler.

4 / 5

Actionability

Fully executable: the exact command with an argument table, defaults, env-var overrides (APP_ID, REPLAY_TIMEOUT_MS, AGENT_DEVICE_STATE_DIR), a copy-paste single-replay code block, a pre-run checklist, and a symptom→action troubleshooting table. This matches the 'fully executable, copy-paste ready, covers common cases' anchor.

5 / 5

Workflow Clarity

Clear sequence (checklist → command → failure handling) with explicit validation checkpoints: the runner re-checks every @pre/@post around each replay, fails on the first replay that emits no span line and prints the current screen, and the 'If something fails' table gives feedback loops for recovery. Retry discipline ('Stop after the first failed replay... Retry only after an explicit reset') is stated explicitly — matching the top anchor.

5 / 5

Progressive Disclosure

Good structure: a layout tree, clearly signaled one-level-deep pointers to `agent-device/flows/README.md` and `agent-device/SKILL.md` for material this file deliberately does not duplicate, and the referenced `scripts/replay-with-deadline.mjs` exists in the bundle. Not a 5: `measure.sh` and `flows/` are described in the layout but not present in this bundle copy, and some operational detail (e.g. the full 'Platform coverage' diagnostics) sits inline in SKILL.md rather than behind a reference.

4 / 5

Total

18

/

20

Passed

Description

73%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A well-targeted, appropriately concise trigger description with strong distinctiveness. Its main weakness is a thin 'what': the reader learns when to invoke it but little about what it does or outputs.

Suggestions

Lead with a third-person statement of the capability, e.g. 'Measures a Sentry performance span across repeated agent-device replays and prints avg/min/max timings' — this would raise specificity and completeness.

Add natural synonyms users might say, such as 'span duration', 'span timing', or 'benchmark a span', to broaden trigger coverage.

Briefly state the output artifact (a summary table of avg/min/max plus per-run samples) so the 'what' is explicit.

DimensionReasoningScore

Specificity

The description names the domain ('measuring a Sentry performance span') plus concrete qualifiers ('agent-device replay flow', 'iOS simulator or Android emulator'), but states only one action — it never says what the skill actually produces (e.g. repeated replays, avg/min/max timing summary). Not a 4: it does not list several specific actions; not a 2: it is far more concrete than 'Processes PDF files'.

3 / 5

Completeness

An explicit 'when' ('Use when measuring...') and an identifiable 'what' ('measuring a Sentry performance span locally with an agent-device replay flow') are both present. Not a 5: the 'what' is thin — the concrete deliverable (runs N replays, prints avg/min/max + samples) is absent; not a 3: the 'when' clause is explicit and multi-condition, not weakly implied.

4 / 5

Trigger Term Quality

Natural terms users would say are present: 'Sentry', 'performance span', 'measuring', 'iOS simulator', 'Android emulator'. Not a 5: common synonyms a user might use — 'span duration/timing', 'benchmark a span', 'profiling' — are missing; not a 3: coverage goes well beyond 'Works with PDF files'-level generic keywords.

4 / 5

Distinctiveness Conflict Risk

A clear niche with minimal conflict risk: 'Sentry performance span' + 'agent-device replay' + platform qualifiers are unlikely to match any other skill's triggers. It clearly fits the 5 anchor ('clear niche with distinct triggers').

5 / 5

Total

16

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 1 suspicious

Warning

Total

14

/

16

Passed

Repository
Expensify/App
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.