CtrlK
BlogDocsLog inGet started
Tessl Logo

observability

Agent observability, evals, feedback, and experiments. Use when adding observability dashboards, configuring trace capture, setting up evals, creating A/B experiments, or collecting user feedback on agent responses.

60

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/observability/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is exceptionally actionable — real code, commands, endpoints, and file paths everywhere — and encodes a large amount of non-obvious, project-specific constraint knowledge. Its weaknesses are structural: a monolithic inline format that buries deep reference material (tracking-bridge PostHog minutiae, failure-report plumbing) in the always-loaded file, and dangling references to docs that do not exist in the bundle.

Suggestions

Split the deep reference material into one-level-deep bundle files — e.g. move the tracking-bridge constraint catalog (lines 357–474) into a references/tracking-bridge.md and the failure-report/error-diagnosis plumbing (lines 77–107) into references/failure-reporting.md, leaving SKILL.md a concise overview with well-signaled links per the progressive_disclosure anchor-5 pattern.

Tighten the rhetorical framing in the always-loaded portion: replace justification sentences like "because a failure nobody can classify is not observable" and "Measured, not assumed — check whether that title is still wrong before trading the duplication back" with the bare rule, keeping the rationale in the split-out reference files.

Resolve the dangling doc pointers — "See the Evals doc", "See the Observability doc", "see the tracking skill" — either by creating those bundle files or by linking to concrete existing paths (as the Key Files table does).

DimensionReasoningScore

Conciseness

Almost every sentence carries framework-specific, non-inferable detail ("PostHog's `$ai_*` latency fields are seconds; ours are milliseconds"), so there is no filler explaining concepts Claude already knows. However, the prose is padded with rhetorical framing ("because a failure nobody can classify is not observable", "Measured, not assumed — check whether that title is still wrong before trading the duplication back"), and the ~117-line tracking-bridge constraints section (lines 357–474) is deep reference material inflating every skill load and could be substantially tightened or split out.

3 / 5

Actionability

Fully executable, copy-paste-ready guidance throughout: complete `defineAppConfig`, `defineEval`, `insertExperiment`, and dashboard-route code blocks; exact CLI commands (`agent-native eval promote <runId> --write evals/from-trace.eval.ts`); concrete HTTP endpoints (`GET /_agent-native/observability/traces/:runId`); and real source paths in the Key Files table. Specific examples cover the common cases for each pillar.

5 / 5

Workflow Clarity

Sequences are clear with most checkpoints present: the promote-trace-to-CI-eval flow with "Truncated runs fail closed"; the CI gate that "exits non-zero if any eval is below its threshold"; the human-audit flow with "Allow feedback and regeneration before explicit approval; never auto-apply"; and the four-part error-report chain (packet → client report → read by id → alerts). The gap keeping it below 5 is that workflows are described narratively rather than as explicit ordered steps, and some (e.g., first-time dashboard setup, experiment lifecycle) lack an explicit validation checkpoint.

4 / 5

Progressive Disclosure

There are no bundle files — this is a single ~470-line monolithic SKILL.md with good headers and tables, but content that clearly belongs in separate reference files is inlined (the tracking-bridge constraint catalog, failure-report plumbing, span-naming rules). Navigation is also weakened by dangling pointers: "See the Evals doc", "See the Observability doc", and "see the tracking skill" reference documents not present in the bundle.

3 / 5

Total

15

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit, well-phrased 'Use when' trigger list covering all five pillars. Its main weakness is that the capability statement is a noun list rather than concrete actions, which slightly limits specificity and completeness.

DimensionReasoningScore

Specificity

The 'Use when' clause lists five concrete actions ("adding observability dashboards, configuring trace capture, setting up evals, creating A/B experiments, or collecting user feedback on agent responses"), but the 'what' is a bare noun list ("Agent observability, evals, feedback, and experiments") rather than concrete capability verbs, leaving minor gaps in coverage versus the anchor-5 example.

4 / 5

Completeness

Both 'what' ("Agent observability, evals, feedback, and experiments") and an explicit 'when' with concrete trigger phrases are present. Held at 4 rather than 5 because the 'what' is a domain enumeration rather than the explicit action statements ("Extract text and tables... fill forms, merge documents") that the anchor-5 example shows.

4 / 5

Trigger Term Quality

Natural trigger terms are present — "observability dashboards", "trace capture", "evals", "A/B experiments", "user feedback", "agent responses" — but common variations a user would plausibly say are missing ("tracing", "monitoring", "telemetry", "A/B testing", or platform names like PostHog/Langfuse/Datadog that the body itself covers).

4 / 5

Distinctiveness Conflict Risk

The agent-run scoping gives it a clear niche distinct from generic document or analytics skills, but "setting up evals" and "creating A/B experiments" carry minor overlap risk with a dedicated evals or analytics/testing skill in the same bundle.

4 / 5

Total

16

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
BuilderIO/agent-native
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.