CtrlK
BlogDocsLog inGet started
Tessl Logo

ax-java-agent-observability

Use when writing Java code with `dev.axllm:ax` for agent tracing, centralized and multi-tenant usage accounting, action logs, runtime diagnostics, replay, and production debugging.

65

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

65%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-sectioned, information-dense API-usage skill with two concrete code snippets and disciplined guardrails, but it reads as reference material rather than a workflow: there is no sequenced procedure with a validation step, the central usage-attribution mechanism has no code example, and a ~50-symbol API enumeration is inlined where a reference file should carry it. The referenced bundle files (API.md, examples/) are clearly signaled but absent from this bundle.

Suggestions

Add an explicit ordered workflow with a validation checkpoint, e.g.: 1. consult `examples/` for the exact call shape, 2. write the code, 3. verify with a `no-key` example (deterministic, no credentials), 4. only then run against a real provider.

Move the bulk of the "Relevant API Surface" symbol list into a separate reference file (it duplicates `API.md` / `axir-api.json`) and keep only the handful of symbols used in the skill's own snippets inline.

Show a complete, runnable usage-observer example — declaring the queue, attaching `usageContext` in a forward-options map, and clearing the observer in teardown — since that is the skill's core pattern and is currently prose-only.

DimensionReasoningScore

Conciseness

The body is dense and assumes competence — no padding explaining what Java, tracing, or observability are — and sections like "Package Facts" and "Guardrails" are token-efficient. Not 5 because the "Relevant API Surface" section is a ~50-symbol flat list ("AxAI: `Ax.ai`, `Ax.typesafe`, `AxAITypesafeClient`, ... `ProviderRouter`") that largely duplicates the referenced `API.md` / `axir-api.json` and could be trimmed to the handful of symbols the skill actually demonstrates. Not 3 because outside that list there is no over-explanation of concepts Claude already knows.

4 / 5

Actionability

There is concrete, near-executable guidance: the Core Pattern code block ("AxAgent helper = Ax.agent(\"question:string -> answer:string\", java.util.Map.of()); var out = helper.forward(llm, ...)"), the observer registration ("AxGlobals.setUsageObserver(usageQueue::add)", "Later: AxGlobals.setUsageObserver(null)"), and specific directives ("synchronously enqueue into a bounded concurrent queue and return immediately", "Attach `usageContext` in AI service options"). Not 5 because key details are described but not shown: `usageQueue` is never declared, there is no complete runnable snippet with imports, and the `usageContext` attachment — the skill's central accounting mechanism — has no code example at all, only prose. Not 3 because the included code and directives are concrete and executable in shape, not pseudocode or vague direction.

4 / 5

Workflow Clarity

There is no explicit sequenced workflow. The guardrails gesture at an order ("Start from package examples for exact native syntax before inventing a new call shape", "Use `no-key` examples for deterministic local checks") but the body is organized as reference material (facts → pattern → API surface → guardrails) rather than steps, and no validation checkpoint tells the agent how to verify its code worked (e.g., run the no-key example, then compare against expected output). This matches the anchor 'Steps listed but validation gaps; sequence present but checkpoints missing or implicit' — arguably weaker, since the sequence is only implicit. Not 4 because there is no clear step sequence with checkpoints for the primary tasks (registering the observer, attributing usage). Not 2 because the guardrails plus the examples-first directive do impose a rough order and the skill is not incoherent.

3 / 5

Progressive Disclosure

References are clearly signaled in "Package Facts" ("Package API docs: `API.md` and `axir-api.json`", "Capability manifest: `axir-capabilities.json`", "Runnable examples: `examples/`") and are one level deep — good. However, the bundle contains no references/, scripts/, or assets/ directories, so none of the referenced files (including `src/examples/java/generation/UsageObserverExample.java`) are verifiable here, and more importantly the large inline "Relevant API Surface" symbol list is exactly the pattern of 'content that should be separate is inline' — it belongs in a reference file next to `API.md`. This sits between the anchor 'Some structure but could be better organized; ... content that should be separate is inline' (3) and 'Good structure; most content is appropriately placed' (4); the substantial inlined API enumeration and unverifiable referenced paths pull it to 3.

3 / 5

Total

14

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it answers what and when explicitly, names concrete capabilities rather than vague domains, and anchors triggering to a specific language and package coordinate. The only weakness is keyword breadth — natural synonyms like "observability", "monitoring", and "telemetry" are absent, which slightly limits recall for users who phrase the need differently.

Suggestions

Add natural synonyms users are likely to say — e.g., "observability", "monitoring", "telemetry" — since the skill name itself contains "observability" but the description never uses it.

Consider mentioning token/cost attribution phrasing (e.g., "token usage and cost attribution") so users expressing accounting needs in those terms also match.

DimensionReasoningScore

Specificity

The description lists multiple specific concrete capabilities: "agent tracing", "centralized and multi-tenant usage accounting", "action logs", "runtime diagnostics, replay, and production debugging". This matches the anchor 'Lists multiple specific concrete actions; comprehensive coverage' — the action list covers the skill's full capability surface shown in the body (tracing, usage observer, action logs, diagnostics, replay, debugging). Not 4 because there is no noticeable gap in the action coverage relative to what the skill does.

5 / 5

Completeness

Both questions are answered explicitly in one sentence: the what ("agent tracing, centralized and multi-tenant usage accounting, action logs, runtime diagnostics, replay, and production debugging") and the when ("Use when writing Java code with `dev.axllm:ax`"). This matches the top anchor 'Clearly and explicitly answers both what AND when with concrete trigger phrases'. Not 4 because the when-clause is explicit and concrete (specific language + package coordinate), not merely present-but-vague.

5 / 5

Trigger Term Quality

Good natural-term coverage: "Java", "agent tracing", "usage accounting", "action logs", "runtime diagnostics", "replay", "production debugging" — phrases a user with this need would plausibly say. Not 5 because common synonyms a user might naturally use are missing: "observability", "monitoring", "telemetry", "token usage"/"cost tracking" — notably the skill's own name contains "observability" yet that word never appears in the description. Not 3 because the included terms are relevant and domain-natural, not generic.

4 / 5

Distinctiveness Conflict Risk

The trigger is pinned to a specific language and package coordinate — "writing Java code with `dev.axllm:ax`" — giving it a clear niche with minimal overlap risk against generic Java, tracing, or debugging skills. Matches the anchor 'Clear niche with distinct triggers; minimal conflict risk'. Not 4 because the package coordinate alone would already disambiguate; combined with the Java qualifier the trigger is essentially unique.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
ax-llm/ax
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.