CtrlK
BlogDocsLog inGet started
Tessl Logo

kiln-prerelease-check

Run the Kiln pre-release smoke test suite plus the standard CI checks (checks.sh), diagnose every prerelease test that broke (and why), and write a clean readable report with recommended actions. Read-only — it never edits code. Use when the user wants to validate a release candidate, run prerelease tests, or asks for a "prerelease check / smoke / audit".

72

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an exceptionally actionable, well-sequenced runbook with genuine validation and feedback loops — every command is executable and every judgment call is named. Its weaknesses are structural: the read-only and retry-disclosure rules are repeated many times, and the large inline report template and coverage table belong in reference files, leaving SKILL.md as a monolith.

Suggestions

Move the ~80-line REPORT.md template and the Phase 4c coverage table into a references/ file (e.g. references/report-template.md and references/coverage-matrix.md), keeping only the section outline and a worked row or two in SKILL.md.

State the read-only rule once in Global rules (bolded) and the retry-disclosure rule once in Phase 4a, then reference them from the checklist and report template instead of restating them verbatim.

Trim the example report template to a few representative rows (e.g. one all-pass, one skip, one retry) rather than enumerating all 15 coverage areas inline.

DimensionReasoningScore

Conciseness

Mostly efficient — the operational detail (pytest flags, env var sourcing, whitelist semantics) is genuinely project-specific — but there is notable repetition: the read-only rule is stated at least five times ("**This skill is read-only. It does not change any code... ever**", "**Read-only. Make no edits.**", "This skill only flags them — it does not edit prod code", "the skill never made any edits", checklist "No files edited outside..."), and the retry-disclosure rule is stated three times in Phase 4a plus again in the checklist and report template. The ~80-line inline report template could be tightened. This matches 'Mostly efficient but includes some unnecessary explanation or could be tightened'; it is not 4 because the duplication is systematic rather than a minor instance.

3 / 5

Actionability

The guidance is fully executable and copy-paste ready: exact bash commands ("uv run ./checks.sh --agent-mode 2>&1 | tee \"${OUT}/checks.log\""), the exact pytest invocation with a justified "-o \"addopts=\"" override, a complete runnable Python status-check script, concrete grep pipelines, and a decision tree keyed on real error signatures ("model not found", "404", "model_not_found"). This matches the anchor 'Fully executable; copy-paste ready code or commands; specific examples cover the common cases'.

5 / 5

Workflow Clarity

Six clearly sequenced phases with explicit validation checkpoints and feedback loops: exit codes are captured with a "don't bail — keep going" rule, transient failures are retried once in isolation with a reclassification rule ("a test that passes on retry is transient/flaky... a test that fails again is reproducible"), ambiguous causes get a mandatory escalation rule, and a final checklist verifies every phase. This matches the anchor 'Clear sequence with explicit validation steps; feedback loops for error recovery; checklists for complex processes'. The destructive/batch cap does not apply since the skill is read-only and validation is pervasive.

5 / 5

Progressive Disclosure

No bundle files exist (no references/, scripts/, or assets/ directories), so all ~430 lines live inline in SKILL.md. Headers give the document decent structure, but content that clearly belongs in a separate file is inlined — the ~80-line REPORT.md template and the long Phase 4c coverage table are classic reference-file material — and the one cross-reference ("specs/projects/code_tools/cross_os_spawn_checklist.md") points outside the bundle. This matches the anchor 'Some structure but could be better organized; content that should be separate is inline'; it is not 2 because the sectioning is real and navigation within the file is easy.

3 / 5

Total

16

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is exemplary: it states the full scope in concrete third-person verbs, closes with an explicit 'Use when' clause containing natural trigger phrases and synonyms, and adds a read-only boundary that sharply limits conflict risk. Every rubric dimension lands on its top anchor.

DimensionReasoningScore

Specificity

The description lists multiple concrete, comprehensive actions — "Run the Kiln pre-release smoke test suite plus the standard CI checks (checks.sh)", "diagnose every prerelease test that broke (and why)", and "write a clean readable report with recommended actions" — covering everything the skill does. This matches the anchor 'Lists multiple specific concrete actions; comprehensive coverage'; it is not score 4 because there are no gaps in the action coverage relative to the skill.

5 / 5

Completeness

It explicitly answers both: what ("Run the Kiln pre-release smoke test suite plus the standard CI checks... diagnose... write a clean readable report") and when ("Use when the user wants to validate a release candidate, run prerelease tests, or asks for a 'prerelease check / smoke / audit'"). This matches the anchor 'Clearly and explicitly answers both what AND when with concrete trigger phrases'; score 4 would require a weaker 'when' clause.

5 / 5

Trigger Term Quality

Trigger terms include natural user phrasings and synonyms: "validate a release candidate", "run prerelease tests", and "prerelease check / smoke / audit". This comprehensively covers the natural terms a user would say for this task, matching the anchor 'Comprehensive coverage of natural terms including synonyms'; score 4 would apply only if common variations were missing, and none are.

5 / 5

Distinctiveness Conflict Risk

It occupies a clear niche (Kiln pre-release validation) with distinct triggers ('prerelease check / smoke / audit', 'validate a release candidate') and an explicit boundary ("Read-only — it never edits code") that separates it from remediation skills. Conflict risk is minimal, matching the anchor 'Clear niche with distinct triggers; minimal conflict risk'; the qualified trigger phrasing keeps generic words like 'audit' from pulling in unrelated skills.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
Kiln-AI/Kiln
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.