CtrlK
BlogDocsLog inGet started
Tessl Logo

ax-java-agent-optimize

Use when writing Java code with `dev.axllm:ax` for agent optimization, verified agent-playbook evolution, evaluators, judges, optimizer artifacts, BootstrapFewShot, and GEPA.

65

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

80%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A lean, well-structured body that adds only package-specific knowledge and points clearly to the package's own docs and examples for depth. Its weaknesses are the non-operational verification gate and the absence of an explicit step sequence for an optimization run, plus a code pattern with unresolved placeholders.

Suggestions

Add a short numbered workflow for an optimization run (e.g., 1. define evaluator callback, 2. bound the run with budgets/dataset rows, 3. run Ax.optimize or AxGEPA, 4. keep only playbook proposals that pass the verification gate, 5. persist artifacts), so the verification checkpoint is explicit rather than implied.

Define what the "verification gate" concretely means and how to apply it, since the body references it as a requirement without any operational detail.

Make the Core Pattern copy-paste ready or point to a specific file under `examples/` that shows a complete run (imports, client, request, evaluator), closing the gap left by the undefined `reflectionClient`, `request`, and `evaluator` variables.

DimensionReasoningScore

Conciseness

The body is a lean, list-based reference of package-specific facts Claude cannot know ("Runtime profiles: `javascript-quickjs`, `python-pyodide`", "Real network support: yes") with no explanations of concepts Claude already knows and no padding. Every section (Package Facts, Core Pattern, Relevant API Surface, Guardrails) earns its tokens, matching the anchor-5 example of assuming Claude's competence.

5 / 5

Actionability

The Core Pattern is real Java syntax ("AxGEPA engine = new AxGEPA(reflectionClient, java.util.Map.of())", "engine.optimize(request, evaluator)") and the API surface lists concrete names ("Ax.optimize", "AxBootstrapFewShot", "OptimizerEvaluator"). It is not anchor 5 because the snippet has undefined variables (reflectionClient, request, evaluator) and no complete, copy-paste-ready example; it is above anchor 3 since it is executable-shaped Java rather than pseudocode, and the "Start from package examples" guardrail gives a concrete fallback.

4 / 5

Workflow Clarity

The "When To Use" bullets describe process pieces ("Mine grounded weaknesses from failed agent tasks and keep only playbook proposals that pass the verification gate", "Keep optimization runs bounded by explicit budgets and dataset rows") but never sequence them into steps, and the verification gate is invoked without being operationalized. This matches anchor 3 (sequence implicit, checkpoints mentioned but not explicit) better than anchor 4, which requires a clear sequence; it is above anchor 2 because validation concerns (verification gate, budgets) are at least named.

3 / 5

Progressive Disclosure

The body is 41 lines (under the 50-line simple-skill threshold), well-organized into clear sections, and appropriately delegates detail one level deep to clearly signaled package artifacts ("Package API docs: `API.md` and `axir-api.json`", "Capability manifest: `axir-capabilities.json`", "Runnable examples: `examples/`"). No bundle files exist to organize, and no content that belongs in a separate file is inlined.

5 / 5

Total

17

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, niche-specific description with an explicit "Use when..." trigger clause and package-level scoping that keeps conflict risk minimal. Its main weakness is that capabilities are enumerated as topics rather than concrete actions, and a few natural synonyms (few-shot, prompt optimization) are absent.

Suggestions

Lead with one or two concrete action statements (e.g., "Optimize AxAgents, create evaluator callbacks, and persist optimizer artifacts") before the "Use when..." clause so the "what" is explicit rather than embedded in the trigger.

Add natural synonyms users might say, such as "few-shot prompting", "prompt optimization", or "playbook", to broaden trigger-term coverage.

Consider naming the artifact types users would mention when asking for help (e.g., "optimization traces" or "evaluator callbacks") to make triggers match real phrasings.

DimensionReasoningScore

Specificity

The description lists several specific capabilities tied to a named package: "agent optimization", "verified agent-playbook evolution", "evaluators, judges, optimizer artifacts", "BootstrapFewShot, and GEPA". It exceeds anchor 3 (only 1-2 concrete actions) but falls short of anchor 5 because the items are domain topics/nouns rather than concrete stated actions with comprehensive coverage.

4 / 5

Completeness

An explicit trigger clause is present ("Use when writing Java code with `dev.axllm:ax` for agent optimization...") so the missing-when cap at 3 does not apply, and the "what" is conveyed through the enumerated topics. It fits anchor 4 rather than 5 because the "what" is embedded inside the when-clause instead of being stated as distinct concrete actions, and the "when" could be more granular.

4 / 5

Trigger Term Quality

Terms a user of this package would naturally say are present: "Java", "agent optimization", "evaluators", "judges", "BootstrapFewShot", "GEPA", and the package coordinate "dev.axllm:ax". A few natural variations are missing (e.g., "few-shot", "prompt optimization", "Ax"), matching anchor 4 rather than the comprehensive synonym coverage of anchor 5.

4 / 5

Distinctiveness Conflict Risk

The package-scoped trigger ("writing Java code with `dev.axllm:ax`") plus niche terms like "BootstrapFewShot", "GEPA", and "agent-playbook evolution" establish a clear niche with distinct triggers and minimal risk of firing for unrelated skills.

5 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
ax-llm/ax
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.