CtrlK
BlogDocsLog inGet started
Tessl Logo

ax-python-refine

Use when writing Python code with `axllm` for reward-scored generation, iterative candidate improvement, evaluator feedback, and optimizer-backed refinement patterns.

58

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./website/static/python/.well-known/agent-skills/ax-python-refine/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

61%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is admirably lean and fact-dense, listing package facts and guardrails Claude could not know otherwise. Its weaknesses are the placeholder-laden core pattern with no executable end-to-end example, the absence of any ordered workflow, and references to API docs and examples that do not exist in the bundle.

Suggestions

Replace the placeholder variables in the Core Pattern (reflection_client, request, evaluator) with a complete, runnable example — ideally one drawn from the examples the body claims exist — so the pattern is copy-paste executable.

Make the referenced resources resolvable: either include API.md, axir-api.json, axir-capabilities.json, and examples/ in the bundle (e.g., under references/) and link them with paths, or remove the claims that they are available.

Add a short ordered decision flow (e.g., 1. pick no-key vs provider examples based on available credentials; 2. construct the engine; 3. run optimize; 4. verify against package examples) so the multi-path guidance becomes a sequence with checkpoints.

DimensionReasoningScore

Conciseness

The body is 43 lines of terse, well-sectioned bullets with no padding and no explanation of concepts Claude already knows ("Language: Python.\n- Package: `axllm`."); every token carries non-obvious package facts or guardrails. This matches anchor 5 — lean, efficient, and assuming Claude's competence.

5 / 5

Actionability

The Core Pattern imports real symbols ("from axllm import AxGEPA", "engine.optimize(request, evaluator)") but hinges on undefined placeholders ("reflection_client", "request", "evaluator") with no complete runnable example, and "Relevant API Surface" lists bare symbol names without signatures or usage. This matches anchor 3 (concrete but incomplete, pseudocode-like) rather than 4, since nothing is copy-paste executable and the referenced "examples/" directory is absent from the bundle.

3 / 5

Workflow Clarity

Guidance is delivered as principles and guardrails ("Start from package examples...", "if package docs disagree with source code, update the compiler and regenerate packages") rather than a sequenced process; the decision path of which optimizer to use when no standalone refine helper exists is implied but never ordered. This matches anchor 3 — relevant conditional guidance exists but there is no explicit sequence or checkpoints — and the simple-skill exemption to 5 does not apply because the core action itself is ambiguous due to the placeholder code.

3 / 5

Progressive Disclosure

The body cites "`API.md` and `axir-api.json`", "`axir-capabilities.json`", and "`examples/`" as bare names with no paths or links, and none of these files exist in the bundle (no references/, scripts/, or assets/ directories), so navigation to them is broken rather than merely unclear. Combined with the inlined "Relevant API Surface" symbol list that duplicates what the referenced API docs should carry, this matches anchor 2 (references buried/dangling, content that belongs in separate files inlined) rather than 3, where references would at least resolve.

2 / 5

Total

13

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A solid, trigger-first description that names a specific package and enumerates concrete capability areas, with both what and when covered. Main weaknesses are jargon-heavy phrasing over natural user terms and a fused what/when clause that blurs the skill's purpose.

Suggestions

Lead with a standalone what-clause (e.g., "Refines generated Python outputs with evaluator feedback and optimizer APIs from the axllm package") before the Use-when clause, so what and when are separately explicit.

Add natural synonyms users would say, such as "improve LLM outputs", "iterate on candidates", or "optimize prompts", alongside the current framework terminology.

State the runtime/tooling trigger more concretely (e.g., "Use when the user mentions axllm, reward-scored generation, or wants to run optimizer examples without provider keys") to sharpen distinctiveness from sibling skills.

DimensionReasoningScore

Specificity

The description names the domain ("Python code with `axllm`") and several concrete capability areas ("reward-scored generation, iterative candidate improvement, evaluator feedback, and optimizer-backed refinement patterns"), matching anchor 4 for several specific actions with minor gaps. It falls short of anchor 5 because the actions are framework-jargon phrases rather than a comprehensive set of plain, concrete operations.

4 / 5

Completeness

It has an explicit trigger clause ("Use when writing Python code with `axllm` for...") and the trailing capability list answers what the skill covers, so both what and when are present. It sits below anchor 5 because the what and when are fused into a single clause rather than a clear standalone statement of what the skill does, and the trigger conditions could be more specific.

4 / 5

Trigger Term Quality

Good keyword coverage including "Python", "axllm", "evaluator feedback", "optimizer", and "refinement", which a user in this niche would plausibly say. Not anchor 5 because common natural variations (e.g., "improve outputs", "prompt optimization") are missing, and "optimizer-backed refinement patterns" leans on internal jargon over user phrasing.

4 / 5

Distinctiveness Conflict Risk

The named package "axllm" and the Python-specific framing carve a distinct niche with minimal conflict risk. Not anchor 5 because terms like "refinement", "optimizer", and "evaluator feedback" overlap with closely related LLM-optimization and sibling generated-package skills.

4 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
ax-llm/ax
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.