CtrlK
BlogDocsLog inGet started
Tessl Logo

self-eval

Honestly evaluate AI work quality using a two-axis scoring system. Use after completing a task, code review, or work session to get an unbiased assessment. Detects score inflation, forces devil's advocate reasoning, and persists scores across sessions.

68

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

The risk profile of this skill

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers highly actionable, executable guidance with a concrete scoring matrix, mandatory reasoning steps, and exact output/persistence formats. It is well-organized and self-contained, with only minor verbosity in the motivational prose and section duplication.

Suggestions

Trim the 'core insight' motivational paragraph and the Features list (which restates the description) to reduce token overhead.

Add a short numbered overview of the end-to-end workflow at the top of 'How to Score' so the section order reads as an explicit sequence rather than implied order.

DimensionReasoningScore

Conciseness

The body is mostly efficient and instructional, but the motivational prose ('The core insight: AI self-assessment converges to everything is a 4...') and the Features section, which restates the description, are minor padding that could be trimmed.

4 / 5

Actionability

It provides copy-paste-ready guidance: a literal composite-score matrix, exact axis definitions, a mandatory devil's-advocate structure, a templated output format, and an exact JSONL line for persistence, fully covering the common cases.

5 / 5

Workflow Clarity

The sequence (identify work, rate axes, read matrix, devil's advocate, anti-inflation check, persist) is clear with explicit self-checks, but the overall order is conveyed through section ordering rather than a single explicit numbered workflow.

4 / 5

Progressive Disclosure

Content is appropriately self-contained in one well-sectioned file with no nested references (it is a prompt-only skill), though at ~180 lines it exceeds the simple-skill threshold that would otherwise merit a 5.

4 / 5

Total

17

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is third-person, concise, and clearly answers both what the skill does and when to use it, with concrete actions and natural trigger terms. Its main limitation is modest overlap risk with code-review skills and slightly abstract phrasing of the core action.

DimensionReasoningScore

Specificity

Lists several concrete actions ('detects score inflation, forces devil's advocate reasoning, and persists scores across sessions') but 'evaluate AI work quality using a two-axis scoring system' stays somewhat abstract without naming the axes, so it falls short of fully comprehensive coverage.

4 / 5

Completeness

It explicitly answers 'what' (evaluate work quality via a two-axis system, detect inflation, force devil's advocate, persist scores) and 'when' with concrete trigger phrases ('Use after completing a task, code review, or work session').

5 / 5

Trigger Term Quality

The phrase 'Use after completing a task, code review, or work session' supplies natural terms a user would say, though a few common synonyms (e.g. 'review my work', 'grade this') are missing.

4 / 5

Distinctiveness Conflict Risk

Self-evaluation of AI work is a clear niche with distinct triggers, but there is minor overlap risk with general code-review skills, keeping it just below the minimal-conflict anchor.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
alirezarezvani/claude-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.