CtrlK
BlogDocsLog inGet started
Tessl Logo

llm-judge

AI quality judge that scores agent responses 0-10 across helpfulness, accuracy, completeness, and clarity. Use when evaluating multi-agent output or implementing LLM-as-judge quality gates.

72

Quality

89%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, well-organized instruction skill: lean token usage, an exact output contract, and useful scoring guardrails. Its main gaps are the unresolved mapping from four scored dimensions to a single output score and the absence of a worked example to anchor consistent judging.

Suggestions

Resolve the output mismatch: state how the four dimensions (helpfulness, accuracy, completeness, clarity) combine into the single "score" field, or output per-dimension scores in the JSON.

Add one worked example — a sample agent response with the exact JSON the judge should emit — to anchor consistent scoring across judges.

DimensionReasoningScore

Conciseness

The ~30-line body is lean and every line instructs (scale, dimensions, exact JSON output, scoring bands); no concepts Claude already knows are explained. Not 4 because even the single persona sentence is functional rather than padded.

5 / 5

Actionability

Concrete guidance throughout — exact 0-10 scale, four defined dimensions, a copy-paste JSON output format, and specific scoring bands ("Score 3-5 for partial or vague responses"). Not 5 because there is no worked example and the four dimensions are never reconciled with the single "score" field in the output format.

4 / 5

Workflow Clarity

The single action (read response, score, emit JSON) is essentially unambiguous with guardrails acting as consistency checkpoints, satisfying the simple-skill exception. Not 5 because the mismatch between scoring four dimensions and outputting one aggregate score leaves a small but real ambiguity in the evaluate workflow.

4 / 5

Progressive Disclosure

Under 50 lines with no need for external references (none exist in the bundle) and three clearly organized sections (Skills, Output Format, Guardrails), which earns a 5 under the simple-skill guideline. Not 4 because nothing here belongs in a separate file.

5 / 5

Total

18

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person voice, concrete capability statement with an explicit scale and four named dimensions, and a well-formed 'Use when...' clause with distinctive triggers. The only weakness is modest synonym coverage in its trigger terms.

DimensionReasoningScore

Specificity

"scores agent responses 0-10 across helpfulness, accuracy, completeness, and clarity" names the domain plus multiple concrete actions with an explicit scale, matching the comprehensive-coverage anchor. Not 4 because there is no meaningful coverage gap for this single-purpose skill.

5 / 5

Completeness

Explicitly answers both "what" (scores agent responses 0-10 across four named dimensions) and "when" ("Use when evaluating multi-agent output or implementing LLM-as-judge quality gates") with concrete triggers. Not 4 because the when-clause is already explicit and specific, not merely adequate.

5 / 5

Trigger Term Quality

"evaluating multi-agent output", "implementing LLM-as-judge quality gates", "agent responses", and "scoring" are natural trigger phrases with good coverage. Not 5 because common synonyms users might say ("grade", "assess", "rate", "rubric", "benchmark") are absent.

4 / 5

Distinctiveness Conflict Risk

"LLM-as-judge", "multi-agent output", and "quality gates" carve a clear niche with distinct triggers and minimal conflict risk with other skills. Not 4 because the overlap risk with generic evaluation skills is negligible given the specific framing.

5 / 5

Total

19

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

Total

14

/

16

Passed

Repository
Atmosphere/atmosphere
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.