CtrlK
BlogDocsLog inGet started
Tessl Logo

scoring-checks

Add a new deterministic scoring check in src/scoring/checks/ that evaluates config quality. Follows the Check[] return pattern, uses point constants from src/scoring/constants.ts, and integrates via filterChecksForTarget() in src/scoring/index.ts. Use when user says 'add scoring check', 'new check', 'modify scoring criteria', or works in src/scoring/checks/. Do NOT use for display changes or refactoring scoring logic.

76

Quality

95%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, highly actionable body: complete code templates, exact commands, per-step verification checkpoints, concrete worked examples, and a troubleshooting section, with no padding explaining concepts Claude already knows. The only refinements are trimming redundancy between the templates and examples and considering splitting the test template/examples into reference files to reduce SKILL.md length.

Suggestions

Trim redundancy between Step 5's full test template and the worked Examples — condensing the test boilerplate to the key assertions would cut ~30 lines without losing actionability.

Consider moving the three worked Examples (or the full unit-test template) into a references/ file (e.g., references/examples.md) and linking one level deep from SKILL.md, keeping the main file a leaner overview.

The Common Issues section could be tightened by merging the two platform-filtering issues ('check doesn't appear in the score report' and 'platform-specific check appears for the wrong agent'), which cover overlapping diagnostics.

DimensionReasoningScore

Conciseness

The body is dense and assumes competence — it never explains TypeScript, vitest, or scoring concepts Claude already knows — but includes minor trim opportunities: the ~50-line full unit-test template in Step 5 and the three worked Examples partially restate the Step 1-4 patterns. This fits 'efficient; minor instances of over-explanation that could be trimmed' rather than the lean-every-token-earns-its-place anchor above, and is clearly above the 'mostly efficient with unnecessary explanation' anchor below.

4 / 5

Actionability

Guidance is copy-paste ready: complete function and registration templates, exact commands ("grep -r \"'your_unique_check_id'\" src/scoring/checks/\"", "npm test src/scoring/checks/__tests__/your-file.test.ts"), and three concrete worked examples with real constants (POINTS_CLAUDE_RULES = 3, CODE_BLOCK_THRESHOLDS) and expected results. The only placeholder ("/* e.g., countFiles() */") is inherent to scaffolding new code and is backfilled by the concrete examples, matching the fully-executable anchor.

5 / 5

Workflow Clarity

Steps 1-5 are clearly sequenced with explicit validation checkpoints after each step ("Verify ID uniqueness: Run grep...", "Verify filtering: Examine filterChecksForTarget()"), a hard gate before shipping ("All must pass before shipping"), and a feedback loop via the Common Issues troubleshooting section. This matches the top anchor with explicit validation steps and error-recovery guidance; the operation is not destructive or batch, so no cap applies.

5 / 5

Progressive Disclosure

No bundle files exist (references/, scripts/, assets/ are absent), so the skill is a self-contained file with well-organized sections (Critical, Step 1-5, Examples, Common Issues) and no nested or buried references — good structure. It falls short of the top anchor because the ~284-line body inlines content (the full test template and worked examples) that could be split into a references/ file to keep SKILL.md a leaner overview, matching 'good structure; minor organization gaps'.

4 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: it states concrete capabilities with exact file paths and function names, provides natural-language triggers with synonyms and a path trigger, and adds a negative scope boundary. It fully answers both 'what' and 'when' without verbosity or over-claiming.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions with exact implementation anchors: "Add a new deterministic scoring check in src/scoring/checks/", "Follows the Check[] return pattern", "uses point constants from src/scoring/constants.ts", and "integrates via filterChecksForTarget() in src/scoring/index.ts". This matches the anchor for multiple specific concrete actions with comprehensive coverage of the skill's scope; it is clearly above the 'several specific actions; minor gaps' anchor because there are no meaningful gaps for a single-purpose code skill.

5 / 5

Completeness

It explicitly answers both questions: 'what' (add a deterministic scoring check following the Check[] pattern with constants and registration) and 'when' ("Use when user says 'add scoring check', 'new check', 'modify scoring criteria', or works in src/scoring/checks/") plus an explicit negative boundary ("Do NOT use for display changes or refactoring scoring logic"). This is the top anchor with concrete trigger phrases, clearly above the 'when could be more explicit' anchor.

5 / 5

Trigger Term Quality

Triggers are natural user phrasings with synonym coverage and a path-based trigger: "Use when user says 'add scoring check', 'new check', 'modify scoring criteria', or works in src/scoring/checks/". It covers natural terms, a synonym ('scoring criteria' vs 'check'), and a concrete path, matching the comprehensive-coverage anchor rather than the 'a few natural terms missing' anchor below it.

5 / 5

Distinctiveness Conflict Risk

The skill occupies a clear niche (adding scoring checks in a specific module path) with distinct triggers and an explicit exclusion clause for display changes and refactoring, minimizing conflict risk with adjacent skills. It exceeds the 'mostly distinct; minor overlap risk' anchor because the triggers and negative boundary disambiguate it from generic code-modification skills.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
caliber-ai-org/ai-setup
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.