CtrlK
BlogDocsLog inGet started
Tessl Logo

claude-maintain-models

Add new AI models to Kiln's ml_model_list.py and produce a Discord announcement. Use when the user wants to add, integrate, or register a new LLM model (e.g. Claude, GPT, DeepSeek, Gemini, Kimi, Qwen, Grok) into the Kiln model list, mentions adding a model to ml_model_list.py, asks to discover/find new models that are available but not yet in Kiln, or wants to add a net-new AI provider to Kiln.

68

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exceptionally actionable, well-sequenced operational skill — commands, checkpoints, and gotchas are copy-paste ready and the workflow is among the clearest one could ask for. Its two weaknesses are structural: a monolithic single-file layout that inlines three large reference sections a one-level-deep references/ bundle should hold, and narrative/history padding that inflates token cost without adding executable guidance.

Suggestions

progressive_disclosure: Move the Provider Quirks Reference, Thinking Levels Reference (including the constants table and 'Sources That Do NOT Work'), and Slug Lookup/Lagging Providers sections into references/ files (e.g. references/provider-quirks.md, references/thinking-levels.md, references/slug-lookup.md), keeping one-line pointers in Phases 2–3 so the workflow body loads without the ~250 lines of lookup material.

conciseness: Compress the historical bug narratives (the Mistral/Qwen/Phi ordering bugs in 3c, the Featherless Aug 2026 rollback story, the stop-hook 'graveyard of abandoned branches' anecdote) into one-line 'past bug: X — do not repeat' notes; the current multi-sentence storytelling adds token cost without executable guidance.

conciseness: Tighten the multi-paragraph rationale sections — Phase 1B's two paragraphs of justification, section 5.0's four restated priority rules, and the release-gating explanation — down to the rule plus a single 'why' sentence each.

DimensionReasoningScore

Conciseness

The content is dense with genuinely non-obvious project knowledge (slug rules, ordering conventions, test-env gotchas), but the ~810-line body includes padding Claude doesn't need: extended historical bug narratives ("past bugs: Mistral Small 4/3 ended up inside the Qwen 2.5 region", the Featherless Aug 2026 rollback story), multi-paragraph rationale sections (Phase 1B's justification, 5.0's four restated priority rules), and verbose explanations like the release-gating rationale that could each be halved. Mostly efficient with clear tighten-able sections — the 3 anchor.

3 / 5

Actionability

Fully executable throughout: exact curl+jq commands per catalog, exact pytest invocations with the `-k "test[model-provider]"` bracket syntax, the key-bridging export snippet, a literal PR-body template, the ordering-verification python one-liner, and file paths for every touchpoint. Specific commands cover the common cases copy-paste ready — the 5 anchor.

5 / 5

Workflow Clarity

Phases are explicitly sequenced with validation checkpoints and feedback loops at every risky point: smoke test before full suite, "debug one at a time" fix-verify-re-run loop, gate conditions in 5.1 before any push, ordering eyeball-check in 3c, and 4e's pre-existing-failure cross-check. The final checklist closes the loop. Matches the 5 anchor (explicit validation, error-recovery loops, checklists).

5 / 5

Progressive Disclosure

No bundle files exist; everything is inlined in one 810-line SKILL.md. Internal anchor links and clear section headers give it real navigability (better than the 2 anchor's header-less wall), but ~250 lines of pure lookup material — Provider Quirks, Thinking Levels constants table, Slug Lookup/Lagging Providers with repeated Featherless jq recipes — are content that clearly belongs in references/ files loaded only when needed. Content that should be separate is inline: the 3 anchor.

3 / 5

Total

16

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete, third-person, with an explicit multi-scenario 'Use when' clause and repo-specific triggers that make false-positive activation very unlikely. The only soft spot is that the 'what' undersells the full workflow (tests, PR creation) and "produce a Discord announcement" names a step the body reduces to a single mention, so specificity stops just short of comprehensive.

DimensionReasoningScore

Specificity

Names concrete actions with a specific file target — "Add new AI models to Kiln's ml_model_list.py", "produce a Discord announcement", "discover/find new models", "add a net-new AI provider" — but omits other major workflow components the skill actually performs (running paid integration tests, creating the PR), so coverage has minor gaps rather than being comprehensive.

4 / 5

Completeness

Explicitly answers both questions: the 'what' is the opening sentence ("Add new AI models to Kiln's ml_model_list.py and produce a Discord announcement") and the 'when' is a concrete "Use when the user wants to..." clause enumerating four distinct trigger scenarios. This matches the 5-anchor example structure almost exactly.

5 / 5

Trigger Term Quality

Good natural-phrase coverage: "add, integrate, or register a new LLM model", "mentions adding a model to ml_model_list.py", "asks to discover/find new models", plus vendor-name synonyms (Claude, GPT, DeepSeek, Gemini, Kimi, Qwen, Grok). A few natural variations are still missing (e.g. "model ID", "model list", "supported providers"), keeping it just below the comprehensive anchor.

4 / 5

Distinctiveness Conflict Risk

The trigger is anchored to a unique artifact ("Kiln's ml_model_list.py", "the Kiln model list") and a specific operation set, giving it a clear niche with minimal overlap risk against generic model or coding skills. The named-repo scoping is exactly the pattern the 5 anchor rewards.

5 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (812 lines); consider splitting into references/ and linking

Warning

Total

15

/

16

Passed

Repository
Kiln-AI/Kiln
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.