CtrlK
BlogDocsLog inGet started
Tessl Logo

ax-python-gen

Use when writing Python code with `axllm` for AxGen programs, forward calls, indexed multi-sampling, result pickers, streaming, tools, assertions, traces, usage, and output parsing.

51

Quality

64%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./website/static/python/.well-known/agent-skills/ax-python-gen/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

42%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is information-dense and technically precise but functions as an inlined specification wall: enormous behavioral-detail sections (forward options, sessions, retries, date parsing) crowd the SKILL.md, referenced documentation files do not exist in the bundle, and almost no executable examples show an agent how to actually build and run a program.

Suggestions

Move the Provider Forward Options, Astra Session, and retry/streaming specification into reference files (e.g., FORWARD_OPTIONS.md, SESSIONS.md) and keep a short overview with the Core Pattern and key option names in SKILL.md.

Add 2–3 complete runnable examples (constructing a client, forwarding, picking a multi-sample) instead of one undefined-variable snippet, ideally drawn from the examples/ directory the body claims exists.

Trim the repeated 'as in TypeScript' parity clauses and the Java/C++/Rust/Go adapter details, which do not help an agent writing Python code with axllm.

DimensionReasoningScore

Conciseness

The ~120-line body is dense run-on spec prose (e.g., the single-paragraph parseDates section, the repeated 'as in TypeScript' comparisons, and Java/C++/Rust/Go adapter detail in a Python skill), which is heavily padded for a SKILL.md. Not 1 because the material is package-specific behavior, not explanation of concepts Claude already knows.

2 / 5

Actionability

Concrete artifacts exist — the Core Pattern snippet, exact option names ('sample_count', 'result_picker=callback'), function spellings, and error strings ('Generate failed: Max steps reached: N') — but the lone code example leaves `llm` undefined and the bulk is descriptive behavioral semantics rather than executable how-to. Not 4 because most guidance describes behavior instead of instructing; not 2 because the specifics that are given are exact and usable.

3 / 5

Workflow Clarity

'When To Use' lists task entry points and 'Guardrails' gives cautions, but there is no sequenced multi-step process with validation checkpoints (e.g., build program → run no-key example → verify against manifest). Not 2 because entry points and cautions do provide a rough orientation; not 4 because no checkpointed sequence exists.

3 / 5

Progressive Disclosure

Sections are clearly headed and the body points to 'API.md', 'axir-api.json', 'axir-capabilities.json', and 'examples/', but those files are absent from the bundle and the enormous Provider Forward Options / Astra Session specification is inlined in SKILL.md where a reference file belongs. Not 2 because structure and doc pointers do exist; not 4 because the split is not actually realized.

3 / 5

Total

11

/

20

Passed

Description

71%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is well-anchored to a unique package with an explicit 'Use when' clause and broad capability coverage, but its trigger vocabulary leans on internal jargon ('indexed multi-sampling', 'result pickers') rather than phrases users naturally say, and the when-clause lacks concrete scenario triggers.

Suggestions

Add natural user-side trigger phrases (e.g., 'when the user asks for structured LLM output, tool loops, or streaming responses in Python') alongside the feature list.

Soften or gloss the internal jargon terms ('indexed multi-sampling', 'result pickers') with plain-language equivalents a user would actually type.

DimensionReasoningScore

Specificity

The description enumerates a comprehensive list of concrete capability areas ("forward calls, indexed multi-sampling, result pickers, streaming, tools, assertions, traces, usage, and output parsing"), but these are feature nouns anchored by 'writing Python code' rather than distinct concrete actions, so anchor 4 fits better than the verb-driven anchor 5.

4 / 5

Completeness

An explicit 'Use when writing Python code with `axllm` for...' clause provides the when, and the feature list provides the what. Not 5 because the trigger phrases lack concrete user scenarios or synonyms; not 3 because the when-clause is explicit, not merely implied.

4 / 5

Trigger Term Quality

Natural terms like 'Python code', 'streaming', 'tools', and 'output parsing' appear, but 'AxGen programs', 'indexed multi-sampling', and 'result pickers' are internal jargon a user would rarely say, and common variations or synonyms (e.g., 'LLM calls', 'structured output') are missing. Not score 4 because the natural-phrase coverage is thin; not score 2 because several genuinely relevant keywords are present.

3 / 5

Distinctiveness Conflict Risk

The unique package name `axllm` and 'AxGen' anchor a clear niche that would not plausibly trigger for any other skill, satisfying the anchor-5 example of a distinct trigger with minimal conflict risk.

5 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
ax-llm/ax
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.