CtrlK
BlogDocsLog inGet started
Tessl Logo

built-in-metrics

Instrument an existing codebase with LaunchDarkly config tracking. Walks the four-tier ladder (managed runner → provider package → custom extractor + trackMetricsOf → raw manual) and picks the lowest-ceremony option that still captures duration, tokens, and success/error.

63

Quality

79%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/agentcontrol/built-in-metrics/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a strong, decision-first skill: a tier ladder with explicit use-when conditions, a provider/package availability matrix, per-tier guidance with guardrails, and a verification step with error-recovery hints. Remaining improvements are trimming two padded paragraphs and tightening progressive disclosure (inline API table, one orphaned reference).

DimensionReasoningScore

Conciseness

The body is dense and assumes SDK competence: decision tables, checklists, and guardrails with literal API calls rather than concept explanations. Minor over-explanation could be trimmed — the runId semantics paragraph ("The Monitoring tab aggregates events rather than grouping them by run today...") and the generic-shape paragraph repeating the tier table's message — so anchor 4 rather than 5.

4 / 5

Actionability

Concrete, executable guidance throughout: exact Python/Node method signatures, real package names, migration mappings (aiclient.config(...) → aiclient.completion_config(...)), and guardrails with literal calls (config.enabled, ldClient.flush()). It stops short of anchor 5 only because the body contains no complete copy-paste code block — implementation detail is deferred to the references, which is acceptable for an instruction-routing skill but leaves minor gaps.

4 / 5

Workflow Clarity

A clear four-step sequence (explore call site → tier lookup via matrix → implement from reference → verify) with checklists in steps 1 and 4. The verify step has explicit validation (run a real request, check the Monitoring tab, force an error) plus error-recovery feedback ("If it doesn't, you probably wrapped the stream creation with trackMetricsOf but didn't add the manual trackTimeToFirstToken call — see streaming-tracking.md"), matching the anchor for clear sequence with explicit validation and feedback loops. No destructive/batch cap applies.

5 / 5

Progressive Disclosure

SKILL.md works as an overview with clearly signaled, one-level-deep markdown links to seven real reference files (all verified present in references/). Two minor gaps: the ~15-row tracker-method API table is inlined in the body where a reference file would fit, and references/metrics-api.md is orphaned — never linked from the body — leaving that content undiscoverable.

4 / 5

Total

17

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, third-person, and well-scoped to a clear niche, with good concrete action verbs and SDK-accurate terminology. Its main weakness is the absence of any explicit "Use when..." trigger guidance, which both caps completeness and leaves natural trigger phrases (monitoring, observability, telemetry) unsaid.

Suggestions

Append an explicit trigger clause, e.g. "Use when instrumenting an existing LLM provider call for LaunchDarkly agent metrics, observability, or Monitoring-tab telemetry, or when the user mentions trackMetricsOf, token usage, or TTFT."

Add the natural synonyms users would actually say — "agent metrics", "LLM observability", "monitoring", "telemetry" — to improve trigger-term coverage and searchability.

Briefly distinguish this skill from adjacent ones (e.g. "agent/runtime metrics — not business metrics like conversion") to reduce overlap risk with custom-metrics and online-evals.

DimensionReasoningScore

Specificity

"Instrument an existing codebase with LaunchDarkly config tracking. Walks the four-tier ladder... picks the lowest-ceremony option that still captures duration, tokens, and success/error" names the domain plus several concrete actions and specific API concepts (managed runner, provider package, trackMetricsOf, raw manual). Minor gaps in coverage (no verification or streaming/TTFT actions) keep it below the comprehensive anchor 5; third-person voice, so no penalty applies.

4 / 5

Completeness

The "what" is clear and concrete, but there is no "Use when..." clause or equivalent explicit trigger guidance — the "when" is only weakly implied by "an existing codebase". Per the rubric guideline, a missing 'Use when' clause caps completeness at 3.

3 / 5

Trigger Term Quality

Natural keywords include "instrument", "config tracking", "duration, tokens, and success/error", and "LaunchDarkly", which users of this SDK would say. A few natural synonyms are missing ("monitoring", "observability", "telemetry", "TTFT"), so it fits anchor 4 rather than 5.

4 / 5

Distinctiveness Conflict Risk

"LaunchDarkly config tracking" is a clear niche with distinct technical triggers, so it is mostly distinguishable. Minor overlap risk remains with closely related sibling skills (custom-metrics for business metrics, online-evals) since the description does not distinguish agent metrics from those.

4 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
launchdarkly/ai-tooling
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.