CtrlK
BlogDocsLog inGet started
Tessl Logo

ax-agent-optimize

This skill helps an LLM generate correct AxAgent tuning and evaluation code using @ax-llm/ax. Use when the user asks about agent.optimize(...), judgeOptions, eval datasets, optimization targets, saved optimizedProgram artifacts, or agent optimization guidance.

65

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Highly actionable codegen guidance with five complete executable patterns, clear decision rules, and explicit failure-recovery and verification loops. The main weaknesses are cross-section rule repetition (conciseness) and a fully inline ~370-line body with no reference files for advanced material (progressive disclosure).

Suggestions

Consolidate the repeated metric-vs-judge and AxGen guidance into the 'Metric vs Judge' section, and cut the restatements in 'Use These Defaults', 'Dataset And Judge Rules', and 'Do Not Generate' to cut ~30% of tokens.

Move 'Eval Semantics' and 'Delegation Optimization Notes' plus the Plain AxGen Judge pattern into a single reference file (e.g. references/advanced-eval.md) linked one level deep, keeping SKILL.md as the decision guide.

Add one short ordered flow near the top (configure agent -> choose metric path -> optimize -> save artifact -> held-out replay) so the end-to-end sequence is visible in one place.

DimensionReasoningScore

Conciseness

No basic-concept padding and dense domain rules, but the same rules are restated across sections: 'custom metric overrides the built-in judge' (lines 102, 313), plain AxGen guidance (lines 25, 91, 98, 266-306, 320, 366), save/load (lines 30, 176-178, 350-354), and judgeOptions.description (lines 24, 90, 311, 321). Mostly efficient but clearly could be consolidated into fewer sections.

3 / 5

Actionability

Five complete, executable TypeScript patterns with imports (Canonical, Minimal, Deterministic Metric, Built-In Judge, AxGen Judge), each copy-paste ready and paired with explicit 'use this when' criteria covering the common cases.

5 / 5

Workflow Clarity

Clear decision sequence (Decision Guide maps user goals to patterns; 'Start here unless...' marks the entry point), failure recovery ('If one training task keeps collapsing to zero, inspect that task first'), and a verification loop (held-out task, then replay on a freshly restored agent). Minor gap: the end-to-end flow is spread across sections rather than one ordered sequence.

4 / 5

Progressive Disclosure

Well-organized section headers and navigable structure, but ~370 lines are all inline with no bundle files at all (references/, scripts/, assets/ absent). Eval Semantics, Delegation Optimization Notes, and the Plain AxGen Judge pattern are candidates for one-level-deep reference files.

3 / 5

Total

15

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete third-person statement of what it does, an explicit 'Use when' clause with the library's natural API trigger terms, and tight scoping that avoids overlap with sibling skills. Minor gap: a few natural trigger synonyms the body actually covers (GEPA, metric, playbook) are absent.

DimensionReasoningScore

Specificity

"generate correct AxAgent tuning and evaluation code using @ax-llm/ax" names the library and two concrete deliverables, and the trigger clause enumerates coverage areas (judgeOptions, eval datasets, optimization targets, artifacts). Not a comprehensive list of distinct actions, so it sits above anchor 3 but short of anchor 5.

4 / 5

Completeness

Explicitly answers both: what ("helps an LLM generate correct AxAgent tuning and evaluation code using @ax-llm/ax") and when ("Use when the user asks about agent.optimize(...), judgeOptions, eval datasets, optimization targets, saved optimizedProgram artifacts"). Matches the anchor-5 structure with concrete trigger phrases.

5 / 5

Trigger Term Quality

Triggers like "agent.optimize(...)" and "judgeOptions" are the natural terms a library user would say, giving good keyword coverage. A few natural variations covered by the body (GEPA, metric, playbook, tuning phrasing) are missing from the description.

4 / 5

Distinctiveness Conflict Risk

Scoped to a specific library (@ax-llm/ax) and a specific API (agent.optimize(...)), giving a clear niche with distinct triggers and minimal conflict risk; the sibling ax-gepa skill is explicitly scoped to non-agent generators.

5 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
ax-llm/ax
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.