CtrlK
BlogDocsLog inGet started
Tessl Logo

ax-python-agent-optimize

Use when writing Python code with `axllm` for agent optimization, verified agent-playbook evolution, evaluators, judges, optimizer artifacts, BootstrapFewShot, and GEPA.

63

Quality

79%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./packages/python/skills/ax-python-agent-optimize/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

72%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A lean, well-organized reference whose package facts and guardrails carry genuinely non-inferable information. Its weaknesses are the placeholder-only core pattern (no fully executable example of the central optimize-with-evaluator task) and the absence of an explicit step sequence with validation checkpoints for optimization runs.

Suggestions

Replace the placeholder Core Pattern with one complete, runnable example adapted from examples/ — constructing a real evaluator callback and invoking AxGEPA.optimize with concrete arguments — so the central task is copy-paste executable.

Add an ordered workflow with an explicit validation checkpoint (e.g., 1. find the closest example in examples/, 2. adapt the evaluator and set explicit budgets/dataset rows, 3. run optimize, 4. keep only playbook proposals that pass the verification gate, 5. persist optimizer artifacts).

Include the concrete command or snippet for a deterministic `no-key` local check so bounded-run verification can be executed without provider credentials.

DimensionReasoningScore

Conciseness

Lean and fact-dense: "Package Facts" supplies only package-specific information ("Runtime profiles: `javascript-quickjs`, `python-pyodide`", "Scripted no-key transport support: yes") and "Guardrails" gives rules Claude could not infer ("Treat AxIR as the source of generated package truth"); no section explains concepts Claude already knows. Not 4: there is no over-explanation to trim — the only redundancy ("Language: Python." duplicating the title) is trivial and every other token carries skill-specific value.

5 / 5

Actionability

The "Core Pattern" block (`engine = AxGEPA(reflection_client)` / `result = engine.optimize(request, evaluator)`) is a call shape with unbound placeholder variables rather than executable code, and no runnable evaluator-callback example is included — matching anchor 3's "pseudocode instead of executable code". Not 4: anchor 4 requires mostly executable guidance; not 2: the block plus exact file pointers ("API.md", "axir-api.json", "examples/") and concrete directives ("Use `no-key` examples for deterministic local checks") go beyond high-level hints.

3 / 5

Workflow Clarity

The sequence is only implicit (consult package examples → construct engine → optimize → keep proposals that pass the verification gate); there is no ordered step list and no explicit validation checkpoint, though "Keep optimization runs bounded by explicit budgets and dataset rows" and the "verification gate" gesture at bounds and validation. Matches anchor 3 (checkpoints missing or implicit); not 4: no explicit steps with checkpoints are written; not 2: the guardrails do supply rough ordering and verification conditions.

3 / 5

Progressive Disclosure

The body is under 50 lines with no bundle files, organized into clear sections ("When To Use", "Package Facts", "Core Pattern", "Relevant API Surface", "Guardrails"), and its detail pointers ("API.md", "axir-api.json", "axir-capabilities.json", "examples/") are one level deep and clearly signaled. Per the simple-skill guideline this matches anchor 5.

5 / 5

Total

16

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, concise description with an explicit trigger clause, specific domain vocabulary, and a clearly distinct niche. Its main limitation is that the "what" is an enumerated noun list folded into a single Use-when clause rather than explicit action statements, which also leaves a few natural trigger phrasings unsaid.

DimensionReasoningScore

Specificity

Enumerates concrete capabilities tied to the domain — "agent optimization, verified agent-playbook evolution, evaluators, judges, optimizer artifacts, BootstrapFewShot, and GEPA" — several specific items rather than generic language. Not 5: these are domain nouns listed under a single "writing Python code" action rather than multiple distinct action statements (e.g., "Optimize agents, create evaluators, persist artifacts"), leaving minor coverage gaps; not 3: it goes well beyond naming 1-2 actions.

4 / 5

Completeness

An explicit "Use when writing Python code with `axllm` for..." clause answers "when" with concrete triggers, and the enumerated task list conveys the "what". Not 5: the "what" is not independently stated with action verbs — the single clause serves both roles rather than pairing a clear what-sentence with an explicit when-sentence; not 3: both elements are present and the "when" is explicit, not weakly implied.

4 / 5

Trigger Term Quality

Includes the natural terms a user of this package would say: "agent optimization", "evaluators", "judges", "optimizer artifacts", "BootstrapFewShot", "GEPA", plus the package name "axllm". Not 5: it misses common variations and plain phrasings ("few-shot", "prompt optimization", "AxAgent") and "verified agent-playbook evolution" is stilted jargon no user would naturally say; not 3: the core trigger vocabulary is present and specific, exceeding "some relevant keywords".

4 / 5

Distinctiveness Conflict Risk

A clear niche — writing Python code with the generated "axllm" package — anchored by distinctive triggers ("BootstrapFewShot", "GEPA", "agent-playbook") that no general-purpose Python skill would claim. Minimal conflict risk with other skills.

5 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
ax-llm/ax
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.