CtrlK
BlogDocsLog inGet started
Tessl Logo

sync-model-catalog

Refresh Agenta's harness catalogs, provider model lists, recommended defaults, and provider-supplied display names. Use for new OpenAI, Anthropic, Codex, Pi, or OpenRouter models and before releases that need current model choices.

71

Quality

89%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An excellent operational runbook: every job carries executable commands or exact API details, validation and feedback checkpoints guard each write path, and the trigger-to-job "When to run" map makes scheduling unambiguous. The only deductions are inline time-stamped version/date facts and a monolithic structure that could offload per-job detail to reference files.

DimensionReasoningScore

Conciseness

Dense with non-obvious, repo-specific facts (the two-copies-of-pi-ai trap, the realpath/pnpm symlink pitfall, overlay-vs-generated merge semantics) with zero basic-concept padding, but time-stamped details ("as of mid-2026", "Claude Code 2.1.280 dropped claude-fable-5") sit inline rather than in an old-patterns/deprecated section, which the rubric penalizes. Efficient with minor trim candidates — not the every-token-earns-its-place level.

4 / 5

Actionability

The automatable jobs are copy-paste ready: the regeneration command with exact paths and the realpath rationale, `python .agents/skills/sync-model-catalog/audit_provider_models.py`, the exact pytest invocation, and exact endpoints (`GET https://api.anthropic.com/v1/models`, OpenRouter's `GET /api/v1/models` with `sort=most-popular`, `supported_parameters=tools`, `output_modalities=text`). The inherently manual jobs still get concrete file paths, API fields (`display_name`, `displayName`, `extras.name`), and decision rules.

5 / 5

Workflow Clarity

Seven clearly sequenced jobs with explicit validation and feedback loops: check the printed generator version and "discard it rather than committing a downgrade", "Review removals before applying them", "flag any entry whose facts you could not verify", a shared pydantic-loader test that "fails loud", and a final human review/commit gate — plus a "When to run" trigger-to-job map. Batch file writes are guarded by validation, so no cap applies.

5 / 5

Progressive Disclosure

A well-sectioned single-file runbook with clearly signaled one-level-deep external pointers (design/plan docs, `generate_pi_models.mjs`, `audit_provider_models.py`, the Daytona README) — but all seven jobs plus the Daytona snapshot section stay inlined in SKILL.md where per-job reference files could split them. Good structure with minor organization gaps rather than the ideal split.

4 / 5

Total

18

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states a clear what with four concrete targets and an explicit "Use for..." trigger clause naming the relevant providers and the release-cadence condition. The only soft spot is that a single generic verb ("Refresh") does all the work, and common synonyms like "model catalog" or "pricing" are missing from the trigger set.

DimensionReasoningScore

Specificity

"Refresh Agenta's harness catalogs, provider model lists, recommended defaults, and provider-supplied display names" names four concrete refresh targets, but one generic verb ("Refresh") carries all of them — several specific actions with minor coverage gaps rather than the comprehensive multi-action anchor.

4 / 5

Completeness

Both halves are explicit: the what ("Refresh Agenta's harness catalogs, provider model lists, recommended defaults, and provider-supplied display names") and the when ("Use for new OpenAI, Anthropic, Codex, Pi, or OpenRouter models and before releases that need current model choices") with concrete trigger phrases.

5 / 5

Trigger Term Quality

"new OpenAI, Anthropic, Codex, Pi, or OpenRouter models" and "before releases that need current model choices" are phrases a user would naturally say, but common synonyms like "model catalog", "model lineup", or "pricing" are absent. Not the level-5 anchor's comprehensive synonym coverage; clearly above the partial-coverage anchor at 3.

4 / 5

Distinctiveness Conflict Risk

"Agenta's harness catalogs" plus the five named providers carve a clear niche with distinct, unlikely-to-collide triggers; minimal overlap risk with generic model- or file-management skills.

5 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

15

/

16

Passed

Repository
Agenta-AI/agenta
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.