CtrlK
BlogDocsLog inGet started
Tessl Logo

ax-python-agent-observability

Use when writing Python code with `axllm` for agent tracing, centralized and multi-tenant usage accounting, action logs, runtime diagnostics, replay, and production debugging.

60

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./website/static/python/.well-known/agent-skills/ax-python-agent-observability/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

65%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a dense, appropriately terse API-usage card for the axllm package with concrete code snippets and useful guardrails. It loses points on structure rather than substance: the inline API-symbol dump belongs in a separate reference file, referenced documentation and examples are absent from the bundle, and the observer setup lifecycle is never laid out as a sequenced workflow.

Suggestions

Trim 'Relevant API Surface' to only the handful of symbols actually used in SKILL.md and point to `API.md`/`axir-api.json` for the full surface, moving the bulk of the list into a bundled reference file.

Add a short numbered sequence for setting up centralized usage accounting (register observer → attach usageContext → enqueue synchronously → clear on teardown) with a verification step so the lifecycle is explicit.

Make code snippets self-contained by defining the names they use (e.g., `llm`, `usage_queue`) or by pairing each snippet directly with the runnable example file that contains the complete version.

DimensionReasoningScore

Conciseness

The body is dense and efficient — no explanations of known concepts, no padding — and sections like Package Facts and Guardrails earn every token. The one notable trim opportunity is the ~60-symbol flat dump in 'Relevant API Surface', which duplicates the referenced `API.md`, keeping it at anchor 4 rather than the fully lean anchor 5.

4 / 5

Actionability

Two concrete code snippets (the core `agent` pattern and `set_usage_observer(usage_queue.put_nowait)`) plus specific directives like 'Attach `usageContext` in AI service options' give mostly executable guidance. Minor gaps — snippets use undefined names (`llm`, `usage_queue`) and the pointed runnable example path (`src/examples/python/generation/usage-observer.py`) is not present in the skill bundle — keep it below anchor 5.

4 / 5

Workflow Clarity

No sequenced workflow is given anywhere; the usage-observer lifecycle (register observer → attach context → callbacks enqueue synchronously → clear on teardown) is implied across bullets rather than listed as steps, and no validation checkpoints exist. There are no destructive or batch operations that would trigger a hard cap, but the sequence remains implicit, matching anchor 3 rather than the explicit sequences of anchor 4.

3 / 5

Progressive Disclosure

The body is well-sectioned and names external materials (`API.md`, `axir-api.json`, `examples/`), but the inline 60-symbol 'Relevant API Surface' list is exactly the content that should live in the referenced `API.md`, and none of the referenced files exist in the skill bundle (no references/, scripts/, or assets/ directories). This matches anchor 3 — structure present, but content that should be separate is inline and references are not backed by bundle files.

3 / 5

Total

14

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, concise description: it names the domain, packages a specific capability list, and includes an explicit 'Use when' clause tied to a distinct package niche. Its main weakness is that the 'when' is largely a restatement of the 'what' rather than independent trigger phrases, and some natural synonyms (monitoring, telemetry, token usage) are missing.

Suggestions

Add user-intent trigger phrases to the 'when' clause that go beyond restating the capability list, e.g., 'Use when the user asks to trace agent calls, attribute token or cost usage per tenant, replay a run, or debug a stuck agent loop.'

Include common synonyms users would naturally say, such as 'monitoring', 'telemetry', 'token usage', or 'cost tracking', to broaden trigger coverage.

DimensionReasoningScore

Specificity

The description lists several concrete capabilities — "agent tracing, centralized and multi-tenant usage accounting, action logs, runtime diagnostics, replay, and production debugging" — matching the anchor for several specific actions with minor gaps. It falls short of a 5 because the capability coverage is not comprehensive, and clearly above a 3 since it names far more than 1-2 actions.

4 / 5

Completeness

It explicitly answers both parts: the what (the enumerated capability list) and the when ("Use when writing Python code with `axllm` for..."). The 'when' clause largely restates the 'what' rather than offering independent user-intent triggers (e.g., 'when the user asks to trace agent calls'), matching anchor 4 — both present but the when could be more specific — and clearly above anchor 3 where the when is missing or only weakly implied.

4 / 5

Trigger Term Quality

Natural phrases like "agent tracing", "usage accounting", "production debugging", and "replay" are terms a user would plausibly say, giving good keyword coverage. A few common synonyms are missing (e.g., "monitoring", "telemetry", "token usage", "cost tracking"), so it fits anchor 4 rather than the comprehensive synonym coverage of anchor 5.

4 / 5

Distinctiveness Conflict Risk

The description is anchored to a specific package name (`axllm`) and a distinct niche (Python agent observability), giving it a clear niche with distinct triggers and minimal conflict risk with other skills, matching anchor 5.

5 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
ax-llm/ax
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.