CtrlK
BlogDocsLog inGet started
Tessl Logo

ax-refine

Use this skill when writing or reviewing Ax bestOfN/refine code, reward functions, thresholds, native sample selection, serial attempts, generated advice, and attempt diagnostics.

66

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, high-quality reference body: lean prose, concrete executable-leaning code, and a clear bestOfN-vs-refine decision rule with explicit feedback-loop semantics and API edge cases. The only meaningful gap is that the primary API example relies on an undefined 'score(prediction)' placeholder and 'addAssert(...)' lacks a usage example, which keeps actionability just below copy-paste-ready.

DimensionReasoningScore

Conciseness

The body is 75 lean lines with zero background on what Ax is or how generic retry/reward loops work; every line states Ax-specific behavior Claude could not guess ('Original instruction values are restored in finally', 'streamingForward(...) is unsupported'), matching the 'every token earns its place' anchor.

5 / 5

Actionability

Both API calls and a complete reward function are shown as concrete TypeScript with real option names (n, threshold, rewardDescription, samplesPerRound), and API behaviors are enumerated rule by rule. It falls short of the copy-paste-ready 5 anchor because 'score(prediction)' is an undefined placeholder and 'addAssert(...)' is named but never shown in use.

4 / 5

Workflow Clarity

It opens with an unambiguous decision rule ('Use bestOfN(...) when you can score complete outputs independently. Use refine(...) when failed rounds should produce feedback'), then separates the validation tools, enumerates strategy selection, and specifies streaming fallback with internal feedback loops (assertion failures feed correction text into the retry loop). It is a reference/decision skill rather than a sequenced multi-step workflow with external validation checkpoints, so it sits below the 5 anchor but clearly above 3.

4 / 5

Progressive Disclosure

This is a single-file skill with no references/, scripts/, or assets/ directories and no external references are needed at this size; the body is cleanly sectioned (Validation, APIs, Reward Functions, Strategies, Advice, Streaming) with no buried or nested pointers, matching the simple-skill provision that well-organized sections score 5 when no external references are needed.

5 / 5

Total

18

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, appropriately terse description with an explicit trigger clause and a specific enumeration of the skill's topic coverage. It scores solidly across the board but stays one notch below top anchors because it lists topics rather than concrete actions, omits trigger synonyms/variations, and uses generic terms that create minor overlap risk.

Suggestions

Add concrete action verbs (e.g., 'Guides writing and reviewing of...') so the 'what' stands apart from the 'when' clause.

Include user-side trigger phrases and variations such as 'best-of-N', 'reward function', 'AxGen sampling', or '@ax-llm/ax' to improve trigger-term coverage.

Qualify generic terms (e.g., 'Ax reward functions', 'refine thresholds') to reduce overlap with non-Ax evaluation/RLHF skills.

DimensionReasoningScore

Specificity

The description enumerates several concrete domain items ('reward functions, thresholds, native sample selection, serial attempts, generated advice, and attempt diagnostics') beyond just naming the Ax bestOfN/refine domain, but the only action verbs are 'writing or reviewing' — a topic list rather than the comprehensive concrete-action coverage of the 5 anchor.

4 / 5

Completeness

The explicit 'Use this skill when writing or reviewing...' clause answers 'when' clearly and the enumerated topics convey the 'what', but the what is embedded in the when-clause with no standalone capability statement and no user-mention trigger phrases like the 5 anchor's 'or when the user mentions PDFs'.

4 / 5

Trigger Term Quality

Terms like 'bestOfN', 'refine', 'reward functions', and 'thresholds' are natural phrases a user working in this niche would say, giving good keyword coverage; common variations such as 'best-of-N', 'reward model', or 'AxGen' are missing, keeping it below the 5 anchor.

4 / 5

Distinctiveness Conflict Risk

'Ax bestOfN/refine' carves out a clear niche with minimal wrong-skill triggering, but broadly shared terms like 'reward functions' and 'thresholds' could overlap with general RLHF or evaluation skills, which is the minor overlap risk described by the 4 anchor rather than the 5 anchor's minimal conflict risk.

4 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
ax-llm/ax
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.