CtrlK
BlogDocsLog inGet started
Tessl Logo

ax-gepa

This skill helps an LLM generate correct AxGEPA optimization code using @ax-llm/ax. Use when the user asks about AxGEPA, GEPA, Pareto optimization, multi-objective prompt tuning, reflective prompt evolution, validationExamples, maxMetricCalls, or optimizing a generator, flow, or agent tree.

68

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A dense, highly actionable codegen rulebook with two complete canonical examples and strong troubleshooting coverage. Its main flaws are duplicated guidance across sections and a Good Example Targets section listing local filesystem paths that are useless outside the author's machine.

Suggestions

Replace the /Users/vr/src/ax/src/examples/* paths in 'Good Example Targets' with files actually shipped in the skill bundle (e.g. references/examples.md) or remove the section.

Consolidate the maxMetricCalls sizing guidance, which is repeated in 'Critical Rules', 'Budgeting and Validation', and 'Troubleshooting', into a single section.

Merge the isExpensive/teacherOptions guidance duplicated between 'Useful Options' and 'Troubleshooting' into one location with a cross-reference.

DimensionReasoningScore

Conciseness

The body is almost entirely library-specific knowledge with no filler explaining concepts Claude already knows, but there is noticeable duplication: maxMetricCalls sizing appears in Critical Rules, Budgeting and Validation, and Troubleshooting, and the isExpensive/teacherOptions rule is stated in both Useful Options and Troubleshooting. This sits between anchor 4 (minor trimmable instances) and anchor 3 (could be tightened); the repetition reads as deliberate emphasis of critical rules rather than padding, so 4.

4 / 5

Actionability

Two complete, executable TypeScript examples (Canonical Scalar Pattern, Canonical Pareto Pattern) cover the common cases end-to-end, with concrete option listings, metric patterns, and named result-handling code. This matches the fully copy-paste-ready anchor.

5 / 5

Workflow Clarity

The workflow (distinct train/validation arrays → size maxMetricCalls → optimize(...) → applyOptimization → serialize/persist) is implied by the canonical patterns rather than enumerated as steps, but validation checkpoints are explicit ('Always create distinct train and validationExamples arrays', budgeting rules) and a Troubleshooting section provides error recovery. Sequence with most checkpoints present, so 4 rather than 5 for the implicit step ordering.

4 / 5

Progressive Disclosure

The ~265-line single-file body is well-sectioned, but 'Good Example Targets' points to author-local absolute paths (/Users/vr/src/ax/src/examples/optimize.ts and five others) that are neither in the bundle nor present on any user's machine — dead references that cannot be navigated. Combined with options/reference material that could live in a separate file, this matches the 'references present but not clearly signaled / could be better organized' anchor rather than 4.

3 / 5

Total

16

/

20

Passed

Description

90%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit what/when structure, comprehensive domain-specific trigger terms, and third-person voice. The only weakness is capability coverage: it names one primary action rather than enumerating the skill's full range (Pareto optimization, artifact persistence, metric selection).

DimensionReasoningScore

Specificity

The description states one concrete action — 'generate correct AxGEPA optimization code using @ax-llm/ax' — plus the target program types ('generator, flow, or agent tree'), which matches the 'names domain and 1-2 concrete actions, but not comprehensive' anchor. It is not a 4 because it does not enumerate several distinct capabilities (Pareto front handling, artifact persistence, metric selection), and not a 2 because the action and domain are concrete rather than generic.

3 / 5

Completeness

It explicitly answers 'what' ('helps an LLM generate correct AxGEPA optimization code using @ax-llm/ax') and 'when' ('Use when the user asks about AxGEPA, GEPA, Pareto optimization...') with concrete trigger phrases, matching the top anchor exactly.

5 / 5

Trigger Term Quality

'AxGEPA, GEPA, Pareto optimization, multi-objective prompt tuning, reflective prompt evolution, validationExamples, maxMetricCalls' gives comprehensive natural-term coverage including synonyms (AxGEPA/GEPA/reflective prompt evolution) and API option names users would quote verbatim. No common variations are obviously missing, matching the comprehensive-coverage anchor.

5 / 5

Distinctiveness Conflict Risk

The triggers are tightly scoped to a named package and its specific optimizer names, giving a clear niche with minimal overlap risk against other skills.

5 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
ax-llm/ax
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.