CtrlK
BlogDocsLog inGet started
Tessl Logo

validate-model

Validate a model entry (or every model in a provider) in apps/sim/providers/models.ts against the provider's live API docs (no hallucination — reports what cannot be verified)

63

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/validate-model/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a highly actionable, well-sequenced audit workflow with strong validation feedback loops and good delegation of bulk reference material to the add-model skill. Its main weakness is redundancy — the unverified-handling rules are stated four times — which costs tokens without adding clarity.

Suggestions

Consolidate the UNVERIFIED guidance into one canonical section (e.g. keep Hard rule 1 and the severity definition) and drop the repeated restatements in the checklist legend, report format, and the closing "What 'I cannot verify this' looks like" section.

Move the time-sensitive model examples (grok-4.3, o-series, Opus 4.7+/Sonnet 5/Fable 5) into a short "known cases" note or the add-model reference so they age in one place.

Merge "Common drift" into the checklist rows it duplicates (pricing cuts, stale updatedAt, retired models) to remove a redundant section.

DimensionReasoningScore

Conciseness

The body is dense and project-specific with no padding over concepts Claude already knows, but the unverified/no-hallucination guidance is repeated across four places (Hard rules 1–2, the checklist status legend, the Step 5 report format, and the "What 'I cannot verify this' looks like" section), and time-sensitive model names (grok-4.3, Opus 4.7+/Sonnet 5/Fable 5) sit outside any old-patterns section. Not 4 because the repetition is substantive redundancy, not a minor trim.

3 / 5

Actionability

Everything is executable: exact file paths, a per-field checklist with unambiguous statuses, concrete commands ("bun run lint", "bun run agent-stream-docs:generate", the add-model re-grep commands), and a mandatory report format shown with a filled-in example table. No gaps between instruction and execution.

5 / 5

Workflow Clarity

A clear six-step sequence with explicit validation checkpoints throughout: the two-source pricing rule, ❓ UNVERIFIED handling on failed fetches, print-diff-before-apply confirmation, single-pass edit, re-lint, and re-running only failed checklist rows. This is a full validate → fix → re-validate feedback loop.

5 / 5

Progressive Disclosure

No bundle files exist, and the skill correctly avoids duplicating bulk material by delegating the provider URL table and Consumption Matrix to the add-model skill via clearly signaled one-level references; the inline checklist and report format are core workflow, not misfiled reference material. Not 5 because the severity definitions and "Common drift" sections partially restate checklist guidance and could be consolidated or externalized.

4 / 5

Total

17

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and distinctive, with a clear statement of what the skill does, but it lacks any explicit "when to use" trigger guidance. Adding a "Use when…" clause with natural trigger phrases would lift its weakest dimension.

Suggestions

Append an explicit trigger clause, e.g. "Use when validating, auditing, or fact-checking model entries in models.ts, or when pricing/capability drift is suspected after a provider price cut or new model launch."

Name the concrete fields checked (pricing per 1M tokens, capabilities, contextWindow, releaseDate, deprecated flags) to sharpen the "what".

Include one or two user-natural synonyms ("audit", "check pricing") alongside "validate" to broaden trigger coverage.

DimensionReasoningScore

Specificity

Names the exact target file ("apps/sim/providers/models.ts"), two scope variants ("a model entry (or every model in a provider)"), the method ("against the provider's live API docs"), and reporting behavior ("reports what cannot be verified"). Not 5 because it stops short of naming the concrete fields it checks (pricing, capabilities, context window).

4 / 5

Completeness

The "what" is clear and specific (validate model entries against live API docs, report unverified fields), but there is no "Use when…" clause or equivalent explicit trigger guidance, which caps completeness at 3 per the judging guidelines. The "when" is only weakly implied by the skill's name.

3 / 5

Trigger Term Quality

Natural, specific terms are present — "validate", "model", "provider", "API docs", "hallucination" — which a user asking to check a model entry would plausibly say. Not 5 because common user synonyms like "check", "audit", or "verify pricing" are missing; not 3 because the coverage is multi-word and domain-specific rather than generic.

4 / 5

Distinctiveness Conflict Risk

Anchored to a unique repository file path and a narrow niche (model entry validation with live-docs verification, "no hallucination"); it is clearly distinguishable from adjacent skills like add-model and has minimal conflict risk.

5 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
simstudioai/sim
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.