CtrlK
BlogDocsLog inGet started
Tessl Logo

kiln-check-finetune-deprecation

Check Kiln's fine-tunable model list for deprecated or unsupported base models. Use when the user wants to audit fine-tuning support, check if fine-tune base models are still valid, or mentions fine-tune model deprecation.

70

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable skill body: executable commands, exact file paths and API fields, sequenced phases with validation checkpoints and a checklist, and almost exclusively non-obvious project-specific knowledge. The only notable slack is the long illustrative output block and inline provider detail that could be trimmed or moved to a reference file.

Suggestions

Shorten the ~35-line illustrative output example in Phase 3 to the distinctive line formats (per-provider found/missing, skip, allowlist directions) and drop the redundant model listings.

Consider moving the provider-specific checking details (Together docs scraping, Vertex publisher API, Fireworks field semantics) into a references/ file, keeping SKILL.md as a lean workflow overview.

DimensionReasoningScore

Conciseness

Nearly all content is non-inferable project knowledge (e.g. the canonical allowlist in fireworks_finetune.py, the stale `tunable` field vs `supervisedLoraTunable`, and "~23 models in API but NOT in allowlist is normal"), with no explanations of concepts Claude already knows. The ~35-line illustrative output block is the main trimmable padding, keeping it at anchor 4 rather than 5.

4 / 5

Actionability

Fully executable, copy-paste-ready commands for both phases (`uv run python3 .agents/skills/kiln-check-finetune-deprecation/scripts/check_finetune.py static`), exact repo file paths, exact API fields, and a concrete verify command (`uv run python3 -m pytest app/desktop/studio_server/test_finetune_api.py -q`) covering the common cases.

5 / 5

Workflow Clarity

Five clearly sequenced phases with explicit validation checkpoints: Vertex auth token check with a recovery prompt, per-provider pass/fail script output, a post-change test run, a closing checklist, and 'Always ask the user to confirm' before any code change. The skill is non-destructive ('only reports findings'), so the destructive-operation cap does not apply.

5 / 5

Progressive Disclosure

The single bundle file (scripts/check_finetune.py) is real and referenced by its full runnable path, keeping the implementation out of SKILL.md, and sections are well organized. The long illustrative output example and provider-specific detail are inline with no references file, a minor organization gap that fits anchor 4 rather than 5.

4 / 5

Total

18

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit what and a well-formed 'Use when...' trigger clause covering multiple natural phrasings. Its only weakness is specificity: it describes a single audit action rather than the concrete checks the skill actually performs.

Suggestions

Enumerate the concrete checks in the description (e.g. 'check static provider_finetune_id entries against provider docs and cross-reference Fireworks' dynamic tunable models against the allowlist') to raise specificity.

Add trigger synonyms such as 'stale fine-tune models' or 'unsupported base models' to broaden natural keyword coverage.

DimensionReasoningScore

Specificity

"Check Kiln's fine-tunable model list for deprecated or unsupported base models" names the domain and one concrete action, but does not enumerate the distinct checks the skill performs (static provider_finetune_id entries vs. dynamic Fireworks models). Not 2 because the action is concrete and domain-specific rather than generic; not 4 because only a single capability is explicitly listed.

3 / 5

Completeness

Explicitly answers both what ("Check Kiln's fine-tunable model list for deprecated or unsupported base models") and when ("Use when the user wants to audit fine-tuning support, check if fine-tune base models are still valid, or mentions fine-tune model deprecation") with concrete trigger phrases, matching the top anchor exactly.

5 / 5

Trigger Term Quality

Natural phrases like "audit fine-tuning support", "check if fine-tune base models are still valid", and "fine-tune model deprecation" mirror how a user would actually phrase the request. Missing common synonyms such as "stale" or "unsupported fine-tune models" keeps it below a 5.

4 / 5

Distinctiveness Conflict Risk

Scoped to a clear niche (Kiln's fine-tunable base models) with distinct triggers that would not fire for general model-deprecation or general fine-tuning skills. Minimal conflict risk.

5 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
Kiln-AI/Kiln
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.