CtrlK
BlogDocsLog inGet started
Tessl Logo

ax-python-agent-observability

Use when writing Python code with `axllm` for agent tracing, centralized and multi-tenant usage accounting, action logs, runtime diagnostics, replay, and production debugging.

66

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

68%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an efficient, well-sectioned overview with real API snippets and clearly signaled external materials. Its weaknesses are the absence of a sequenced workflow with explicit validation checkpoints, and code snippets that are fragments rather than self-contained runnable examples.

Suggestions

Add a short sequenced workflow (e.g., check `axir-capabilities.json` → copy a runnable example from `examples/` → adapt to the task → verify with a `no-key` example) with an explicit validation checkpoint, to raise workflow_clarity.

Make the code snippets self-contained by defining `llm` and `usage_queue` (or showing the queue construction), so the Core Pattern and observer examples are copy-paste runnable.

Move the long "Relevant API Surface" symbol list into a separate reference file and keep only a handful of key symbols inline, improving both progressive_disclosure and conciseness.

DimensionReasoningScore

Conciseness

The body is dense operational guidance ("The observer is process-wide, best-effort, and fail-open. Registering again replaces the previous observer.") with no padding explaining concepts Claude already knows. The inline "Relevant API Surface" symbol list and "Package Facts" block are tokens that could be trimmed or relocated. Matches the anchor for efficient content with minor trim instances; not 5 because that inline list does not fully earn its place, not 3 because there is no unnecessary explanation.

4 / 5

Actionability

Concrete snippets show real API calls ("helper = agent(\"question:string -> answer:string\")", "set_usage_observer(usage_queue.put_nowait)"), but neither is self-contained since `llm` and `usage_queue` are undefined; this is mitigated by pointers to runnable examples ("Runnable provider example: `src/examples/python/generation/usage-observer.py`"). Matches the anchor for mostly executable guidance with minor gaps; not 5 because the snippets are not copy-paste ready, not 3 because they are real code rather than pseudocode and complete examples are one hop away.

4 / 5

Workflow Clarity

Sections are topical (When To Use, Core Pattern, Centralized Usage Observer, Guardrails) rather than a sequenced workflow, and validation is only implicit ("Use `no-key` examples for deterministic local checks and provider request mapping"). Matches the anchor for sequence/checkpoints missing or implicit; not 4 because no explicit step order or validation checkpoints exist anywhere in the body, not 2 because the guardrails do impose a rough order ("Start from package examples ... before inventing a new call shape").

3 / 5

Progressive Disclosure

"Package Facts" clearly signals one-level-deep materials ("API.md", "axir-capabilities.json", "examples/") and the sections are well organized, but the "Relevant API Surface" section inlines a long symbol list that belongs in a separate reference file. Matches the anchor for good structure with minor organization gaps; not 5 because of that inline reference data, not 3 because references are clearly signaled and the body is otherwise appropriately split.

4 / 5

Total

15

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it pairs an explicit 'Use when' trigger with a comprehensive list of concrete capabilities and is firmly anchored to a specific package and language. The only gap is missing common synonym keywords (observability, telemetry, monitoring) that users might naturally say.

DimensionReasoningScore

Specificity

"agent tracing, centralized and multi-tenant usage accounting, action logs, runtime diagnostics, replay, and production debugging" names six concrete capabilities, matching the anchor for multiple specific concrete actions with comprehensive coverage. Not 4: the coverage of the skill's scope is complete rather than having minor gaps; not below since nothing is vague or generic.

5 / 5

Completeness

The explicit trigger clause "Use when writing Python code with `axllm` for ..." answers 'when' with concrete trigger phrases, and the capability list answers 'what'. Matches the anchor that clearly and explicitly answers both; not 4 because the 'when' is already explicit and specific rather than merely present-but-improvable.

5 / 5

Trigger Term Quality

Natural phrases a user would say are present ("agent tracing", "usage accounting", "runtime diagnostics", "replay", "production debugging", "Python"), but common synonyms like "observability", "telemetry", and "monitoring" are absent. Matches the anchor for good keyword coverage with a few natural terms missing; not 5 because synonym coverage is not comprehensive, not 3 because coverage is well beyond a couple of generic keywords.

4 / 5

Distinctiveness Conflict Risk

The description is anchored to the specific package "`axllm`" and the Python agent-observability niche (tracing, usage accounting, replay), a clear niche with distinct triggers and minimal conflict risk. Not 4: there is no meaningful overlap with closely related skills since the package name disambiguates.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
ax-llm/ax
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.