CtrlK
BlogDocsLog inGet started
Tessl Logo

ax-python-playbook

Use when writing Python code with `axllm` for the playbook() context-engineering surface, agent-bound verified evolution, run-end learning, online updates, and rendering a playbook into a program.

60

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./packages/python/skills/ax-python-playbook/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

76%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, well-sectioned reference-style skill with excellent token efficiency and mostly concrete guidance anchored by a real-syntax code pattern. The main gaps are the absence of a sequenced workflow with validation checkpoints for the evolve/refine lifecycle, a core example that is not fully runnable, and file citations that do not resolve within the skill bundle.

Suggestions

Add a short numbered workflow for the playbook lifecycle with an explicit validation checkpoint (e.g., 1. Grow with playbook() 2. Attach and run 3. Evolve with verification and exact rollback 4. Validate the learned rules with no-key examples before rendering), turning the implicit When-To-Use sequence into explicit steps with feedback loops.

Make the Core Pattern runnable or explicitly justify the placeholders — define how `llm`, `examples`, and `metric_fn` are constructed (e.g., one line each or a pointer to the exact example file in the package's examples/ directory).

Clarify where `API.md`, `axir-capabilities.json`, and `examples/` live (generated package vs. skill bundle) so a reader knows where to navigate, and link the specific example files for each When-To-Use bullet.

DimensionReasoningScore

Conciseness

The body is lean and fully sectioned — every line carries package-specific information Claude would not know ("Real network support: yes", "Scripted no-key transport support: yes", "Runtime profiles: `javascript-quickjs`, `python-pyodide`") with zero padding or explanation of known concepts. It clearly matches anchor 5 (lean, assumes Claude's competence, every token earns its place) rather than anchor 4, which reserves room for trimmable over-explanation.

5 / 5

Actionability

The Core Pattern shows real import and call syntax ("from axllm import ax, playbook", "pb = playbook(program, {\"studentAI\": llm})", "pb.evolve(examples, metric_fn)"), and the Guardrails give concrete directives ("Start from package examples for exact native syntax", "Use `no-key` examples for deterministic local checks"). It is not anchor 5 because the snippet is not copy-paste runnable — `llm`, `examples`, and `metric_fn` are undefined placeholders with no metric signature — and "Relevant API Surface" lists optimizer names without any call shape; but it is well above anchor 3's pseudocode level.

4 / 5

Workflow Clarity

The "When To Use" bullets imply a lifecycle (grow a playbook → attach to an agent → evolve from run-end failures → refine online/offline → render and inject) but there is no sequenced workflow and no validation checkpoints — the only check-adjacent guidance is in Guardrails ("Start from package examples", "no-key examples for deterministic local checks"). This fits anchor 3 (sequence implicit, checkpoints missing); anchor 4 would require an explicit ordered sequence with most checkpoints present.

3 / 5

Progressive Disclosure

The body is short (~44 lines) and cleanly sectioned (When To Use, Package Facts, Core Pattern, Relevant API Surface, Guardrails), which on its own merits a 5 under the simple-skill exception. It drops to anchor 4 because the "Package Facts" section cites `API.md`, `axir-api.json`, `axir-capabilities.json`, and `examples/` as lookup targets, yet no such files exist in the skill bundle (no references/, scripts/, or assets/ directories) — navigation to them is ambiguous since it is never clarified whether they live in the generated package or alongside the skill.

4 / 5

Total

16

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A well-scoped, appropriately concise description with an explicit 'Use when' trigger and excellent distinctiveness thanks to the axllm-specific scoping. Its main weakness is that the capability list is expressed as dense jargon noun-phrases ('context-engineering surface', 'agent-bound verified evolution') rather than concrete actions, which hurts both specificity and natural trigger-term coverage.

Suggestions

Rewrite the capability list as concrete verb-object actions (e.g., 'Grow an evolving playbook with playbook(), mine grounded weaknesses from run-end failures, refine online or offline, render and inject the playbook into a program') instead of jargon noun phrases like 'agent-bound verified evolution'.

Split into a what-statement plus the when-clause (e.g., 'Builds, evolves, and renders context playbooks with the axllm package. Use when writing Python code with axllm, or when the user mentions playbooks, context engineering, or run-end learning.') so both halves are explicitly answered.

Add one or two natural synonyms users might say (e.g., 'playbook', 'Ax', 'context engineering') to broaden trigger-term coverage beyond the bare package name.

DimensionReasoningScore

Specificity

The description names the domain ("writing Python code with `axllm`") and enumerates capability areas ("playbook() context-engineering surface", "agent-bound verified evolution", "run-end learning, online updates", "rendering a playbook into a program"), but most are jargon-laden noun phrases rather than concrete verb-object actions — only "rendering a playbook into a program" reads as a concrete action. It fits anchor 3 (names domain and 1-2 concrete actions) better than anchor 4, since the listed items are topical labels with buzzword compounds like "context-engineering surface" rather than the executable capabilities spelled out in the body.

3 / 5

Completeness

An explicit trigger clause is present ("Use when writing Python code with `axllm` for...") and the what is conveyed through the enumerated capability areas, so both what and when are covered. It falls short of anchor 5 because what and when are merged into a single clause — there is no separate statement of what the skill does, and the 'when' leans on domain jargon rather than concrete user-facing trigger phrases.

4 / 5

Trigger Term Quality

Natural terms a user in this domain would say are present ("Python code", "axllm", "playbook"), but coverage stops there — no synonyms or common variations (e.g., "Ax", "context playbook", "evolve a playbook", "seed playbook") appear. This matches anchor 3 (some relevant keywords but missing common variations or synonyms); anchor 4 would require a broader spread of natural phrasings.

3 / 5

Distinctiveness Conflict Risk

The trigger is tightly scoped to a specific generated package ("writing Python code with `axllm`") and its playbook() surface, giving it a clear niche with minimal conflict risk — no other skill would plausibly claim this trigger. This clearly matches anchor 5 rather than anchor 4, which still assumes overlap risk with closely related skills.

5 / 5

Total

15

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
ax-llm/ax
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.