CtrlK
BlogDocsLog inGet started
Tessl Logo

evaluate-skill

Test a candidate OpenSEO skill end to end by running fresh, isolated Codex sessions against the local backend and scoring the reports they save. Use when editing a product skill (seo-audit, keyword-research, ...) and you need evidence that the new instructions produce better output, not just a cleaner file.

66

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a lean, instruction-only workflow for a complex batch process: it gives concrete commands, a clear sequenced process with a validation probe, and pushes executable detail into two real scripts. It earns 4s across the board, held from 5 by minor trim potential, placeholder gaps, implicit error-recovery loops, and inline rubric content.

DimensionReasoningScore

Conciseness

The body assumes Claude's intelligence (no explaining what MCP, Codex, or D1 are) and almost every line carries usable guidance, but a few prose passages such as "Runs take about 10 to 14 minutes and execute in parallel" plus surrounding rationale could be trimmed, fitting anchor 4.

4 / 5

Actionability

It provides copy-paste-ready commands (the full run.mjs invocation, nohup wrapper, report-text.mjs, delete_* calls) with real flags, but template placeholders like <local-dev-host> and "<the holdout business>. ..." leave minor gaps, matching anchor 4.

4 / 5

Workflow Clarity

A clear Prerequisites -> Run -> Isolation -> Compare -> Clean up sequence with a per-site numbered list and an explicit checkpoint ("probes that rejection before starting") avoids the batch-operation cap of 3; it falls short of 5 because recovery feedback loops for failed/aborted runs are implied rather than spelled out.

4 / 5

Progressive Disclosure

Execution detail is correctly split into scripts/run.mjs and scripts/report-text.mjs (both real bundle files referenced inline), with a well-sectioned overview in SKILL.md; it is not a 5 because there are no dedicated reference docs for the scoring rubric or isolation contract, so some heavy material (the rubric table) lives inline.

4 / 5

Total

16

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is third-person, concrete, and answers both 'what' and 'when' with explicit trigger guidance tied to product-skill editing. It is strong on completeness and distinctiveness; only slightly less exhaustive on action enumeration and natural-term coverage.

DimensionReasoningScore

Specificity

Lists several concrete actions ("running fresh, isolated Codex sessions against the local backend and scoring the reports they save"), but they stay at a high level without enumerating the sub-steps, matching anchor 4 rather than 5.

4 / 5

Completeness

Explicitly states what it does (test a candidate skill end to end via isolated sessions and score reports) and when to use it ("Use when editing a product skill ... and you need evidence that the new instructions produce better output"), matching the anchor 5 example that answers both with concrete trigger phrases.

5 / 5

Trigger Term Quality

Natural triggers like "editing a product skill (seo-audit, keyword-research, ...)" and "evidence that the new instructions produce better output" would be said by a user, but a few common phrasings (e.g. "skill eval", "A/B test a skill") are absent, fitting anchor 4.

4 / 5

Distinctiveness Conflict Risk

The niche is narrow and clearly bounded (OpenSEO product skills, Codex sessions, local backend), making overlap with unrelated skills minimal, matching anchor 5's clear-niche criterion.

5 / 5

Total

18

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

Total

14

/

16

Passed

Repository
every-app/open-seo
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.