CtrlK
BlogDocsLog inGet started
Tessl Logo

scoring-checks

Add a new deterministic scoring check in src/scoring/checks/ that evaluates config quality. Follows the Check[] return pattern, uses point constants from src/scoring/constants.ts, and integrates via filterChecksForTarget() in src/scoring/index.ts. Use when user says 'add scoring check', 'new check', 'modify scoring criteria', or works in src/scoring/checks/. Do NOT use for display changes or refactoring scoring logic.

68

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is scoring-checks in caliber-ai-org/ai-setup

SKILL.md
Quality
Evals
Security

Quality

Content

70%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, highly actionable skill: every step is sequenced with explicit verification commands and concrete code templates, and the worked examples show expected outcomes. Its weaknesses are length (test template and troubleshooting padding) and total absence of progressive disclosure — everything lives inline in one long SKILL.md with no reference files.

Suggestions

Trim the Step 5 unit-test template to one canonical test (passing case with full assertions) plus one-line notes for the zero-points and detail-message cases, cutting ~30 lines of near-duplicate boilerplate.

Move the Common Issues section and the full test template into references/ (e.g. references/common-issues.md, references/test-template.md) and link to them from the body, turning the 280-line monolith into a lean overview with one-level-deep references.

Remove the duplication between the Critical section's 'Fix object fields' bullet and the identical field annotations inside the Step 2 code template — state the field contract once.

DimensionReasoningScore

Conciseness

At ~280 lines the body is mostly efficient (project-specific conventions like ".js extension in imports (TypeScript transpiles to ES modules)" earn their place), but the ~50-line unit-test template with three nearly-duplicate test cases, a ~30-line check-function template whose fix-object fields duplicate the Critical section, and a 7-item Common Issues list could all be tightened or trimmed. Anchor 3 ("mostly efficient but includes some unnecessary explanation or could be tightened") fits better than 4, where only minor trimming would be needed.

3 / 5

Actionability

Guidance is highly executable: full TypeScript templates, exact verification commands ("Run grep -r \"'your_unique_check_id'\" src/scoring/checks/", "npm test src/scoring/checks/__tests__/your-file.test.ts"), concrete file-to-category mapping, and three worked examples with expected point outcomes. Not a 5 because the core template contains placeholder pseudocode ("const yourMetric = /* e.g., countFiles(), validatePaths(), etc. */;", "// Set up the passing condition") that the reader must fill in — minor gaps, matching anchor 4.

4 / 5

Workflow Clarity

Five clearly sequenced steps (constants → check function → platform filtering → registration → tests), each ending with an explicit verification checkpoint ("Verify ID uniqueness: Run grep...", "Verify registration: Run npm test...", "Verify: Review CATEGORY_MAX..."), plus a test-writing step with a run-and-fail feedback loop ("All must pass before shipping"). This matches anchor 5's explicit validation steps and feedback loops; anchor 4 would require some checkpoints to be missing, and none are.

5 / 5

Progressive Disclosure

Section headers (Critical, Instructions, Examples, Common Issues) give real structure, but the skill is a single monolithic file with no references at all, and content that clearly belongs in separate files — the ~50-line test template and the 7-item troubleshooting list — is inlined. Anchor 3 ("some structure... content that should be separate is inline") fits; anchor 4 requires references to be 'mostly clear', which is moot when none exist for a 280-line body.

3 / 5

Total

15

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: it states concrete capabilities with specific file paths and API names, gives explicit quoted user-trigger phrases, and adds negative boundaries to avoid mis-triggering. The only weakness is that "evaluates config quality" is slightly generic about the check's actual subject matter.

DimensionReasoningScore

Specificity

Quotes like "Add a new deterministic scoring check in src/scoring/checks/", "Follows the Check[] return pattern", "uses point constants from src/scoring/constants.ts", and "integrates via filterChecksForTarget()" name several concrete actions with real paths and symbols. Not a 5 because "evaluates config quality" stays generic about what aspects of quality are checked, leaving a minor coverage gap.

4 / 5

Completeness

It explicitly answers both what ("Add a new deterministic scoring check... Follows the Check[] return pattern, uses point constants..., integrates via filterChecksForTarget()") and when (an explicit "Use when user says..." clause with concrete trigger phrases). This mirrors the anchor-5 example structure; a 4 would require the 'when' to be less explicit, which it is not.

5 / 5

Trigger Term Quality

"Use when user says 'add scoring check', 'new check', 'modify scoring criteria', or works in src/scoring/checks/" provides natural quoted phrases a user would actually say, plus a path-based trigger and negative triggers ("Do NOT use for display changes or refactoring scoring logic"). Coverage matches the anchor-5 example's breadth of natural terms and synonyms for its niche; nothing above this anchor exists.

5 / 5

Distinctiveness Conflict Risk

The niche is narrow (adding checks to a specific scoring module) with distinct quoted triggers and an explicit exclusion clause ("Do NOT use for display changes or refactoring scoring logic"), minimizing conflict risk with adjacent skills. Clear niche with distinct triggers matches anchor 5; anchor 4's 'minor overlap risk' does not apply given the negative boundary.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
caliber-ai-org/ai-setup
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.