CtrlK
BlogDocsLog inGet started
Tessl Logo

scoring-checks

Add a new deterministic scoring check in src/scoring/checks/ that evaluates config quality. Follows the Check[] return pattern, uses point constants from src/scoring/constants.ts, and integrates via filterChecksForTarget() in src/scoring/index.ts. Use when user says 'add scoring check', 'new check', 'modify scoring criteria', or works in src/scoring/checks/. Do NOT use for display changes or refactoring scoring logic.

72

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is scoring-checks in caliber-ai-org/ai-setup

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A high-quality procedural skill body: fully executable templates, exact verification commands, explicit per-step validation checkpoints, and concrete worked examples with troubleshooting. The only meaningful gap is structural — the single-file layout is longer than needed and the examples/troubleshooting content could be split into clearly signaled reference files.

Suggestions

Move the Examples and Common Issues sections into references/examples.md and references/troubleshooting.md, linked one level deep from the body, to shorten the always-loaded SKILL.md.

Trim the test template to the two essential cases (pass/fail) and note the third pattern in one line, since the boilerplate is repetitive.

Consolidate the 'Critical' bullet list with the corresponding step text to remove duplicated guidance on constants, registration, and target filtering.

DimensionReasoningScore

Conciseness

The body is project-specific throughout with no explanations of concepts Claude already knows, and code comments are terse and purposeful. Minor instances of over-explanation remain — the third near-identical boilerplate vitest case and slight overlap between the 'Critical' section and Steps 1-4 — so it fits the 'efficient; minor instances that could be trimmed' anchor rather than the lean-every-token-earns-its-place anchor.

4 / 5

Actionability

Provides copy-paste-ready TypeScript templates for the check function, constants, registration, and tests, plus exact executable commands ("grep -r \"'your_unique_check_id'\" src/scoring/checks/\", "npm test src/scoring/checks/__tests__/your-file.test.ts") and three worked examples with concrete triggers and point outcomes. Placeholders in the main template are inherent to a template skill and are explicitly patterned after named existing examples (TOKEN_BUDGET_THRESHOLDS, CODE_BLOCK_THRESHOLDS).

5 / 5

Workflow Clarity

Five clearly sequenced steps each end with an explicit validation checkpoint (verify CATEGORY_MAX budget, verify ID uniqueness via grep, verify filterChecksForTarget(), run existing tests, run new tests before shipping), and the Common Issues section supplies error-recovery loops for each failure mode — matching the anchor for clear sequence with explicit validation steps and feedback loops.

5 / 5

Progressive Disclosure

Sections are well-organized (Critical, Instructions, Examples, Common Issues) with no dangling references, but the ~280-line body is monolithic: the Examples and Common Issues sections are natural candidates for one-level-deep reference files, and there are no reference files in the bundle. This fits 'good structure; most content appropriately placed; minor organization gaps' rather than the well-split top anchor.

4 / 5

Total

18

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states concrete technical specifics of the integration pattern, provides multiple natural trigger phrases, and adds an explicit exclusion clause to prevent mis-triggering. The only weakness is trigger coverage missing a few plausible synonyms, which keeps specificity and trigger_term_quality at 4 rather than 5.

Suggestions

Add one or two trigger synonyms such as 'scoring rules', 'evaluation criteria', or 'point values' to broaden natural-term coverage.

Consider mentioning the outcome users can expect (e.g., a scored report with suggestions/fixes) so the capability set reads as more than one action.

DimensionReasoningScore

Specificity

Names the domain plus several concrete specifics — "Follows the Check[] return pattern", "uses point constants from src/scoring/constants.ts", "integrates via filterChecksForTarget()" — matching the anchor for several specific actions with minor gaps. Not 5 because the description centers on a single action (adding a check) with implementation constraints rather than comprehensively listing capabilities.

4 / 5

Completeness

Explicitly answers both what ("Add a new deterministic scoring check in src/scoring/checks/ that evaluates config quality") and when ("Use when user says 'add scoring check', 'new check', 'modify scoring criteria', or works in src/scoring/checks/") with concrete trigger phrases, plus a negative exclusion ("Do NOT use for display changes or refactoring scoring logic"). This matches the anchor for clearly and explicitly answering both what AND when.

5 / 5

Trigger Term Quality

Includes natural phrases users would actually say — "'add scoring check'", "'new check'", "'modify scoring criteria'", "works in src/scoring/checks/" — giving good keyword coverage. Not 5 because common synonyms and variations (e.g., 'scoring rules', 'evaluation criteria', 'point system') are missing.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche tied to this repository's scoring architecture (src/scoring/checks/, constants.ts, filterChecksForTarget()) with explicit boundary guidance ("Do NOT use for display changes or refactoring scoring logic"), giving minimal conflict risk with other skills — matching the top anchor.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
caliber-ai-org/ai-setup
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.