CtrlK
BlogDocsLog inGet started
Tessl Logo

tooluniverse-self-review

Review existing work against the user's actual goal and surface evidence-backed strengths, gaps, risks, and next fixes. Use when asked to eval, evaluate, review, assess, or check current/this/my/our work; decide whether a task is complete; build a definition-of-done checklist or rubric; or perform grading, LLM-as-judge, Qworld, or RET evaluation. Treat plain eval/review requests as qualitative: resolve "current work" from the conversation, artifacts, files, or diff, and never assign numeric scores unless the user explicitly requests scores, grades, points, ratings, weighted criteria, Qworld, or RET. Do not use for implementing automated eval suites, tests, graders, or benchmarks.

72

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable instruction skill with clear mode routing, a sequenced workflow with decision checkpoints, and proper use of a one-level-deep reference for the opt-in scoring detail. The main weakness is repetition: the qualitative-by-default / no-scores-unless-asked rule is restated many times across sections and could be consolidated.

Suggestions

State the 'qualitative by default, scoring only on explicit opt-in' rule once (e.g., in Non-Negotiable Rules) and reference it elsewhere instead of restating it in the mode table, the scored-mode intro, the checklist section, and the routing examples.

Trim the Examples of Correct Routing table to only the ambiguous or surprising cases (e.g., 'Eval current work', 'Create an eval suite') — the rest already follow directly from the mode table.

Merge overlapping guidance between the Qualitative Review Workflow steps and the Evidence and Honesty section (e.g., evidence-verification rules appear in both) into one place.

DimensionReasoningScore

Conciseness

The content is skill-specific with no generic concept explanations, but the core rule that plain eval/review requests are qualitative and unscored is restated at least five times (Rule 2, the mode table, the scored-mode intro, the checklist section, and the routing examples table), which is unnecessary repetition. Anchor 3 ("could be tightened") fits; anchor 2 would require padded or generic explanation, which is absent.

3 / 5

Actionability

Fully actionable instruction-level guidance: an ordered five-step target-resolution procedure, a mode-selection table with exact triggers, a seven-step review workflow, a fixed output shape with named headings and a defined verdict vocabulary, and a worked routing-examples table covering common cases. This is the executable equivalent of copy-paste code for an instruction-only skill.

5 / 5

Workflow Clarity

The default workflow is clearly sequenced with explicit checkpoints: an ask-vs-proceed decision rule for ambiguous targets, evidence-citation requirements per finding, rules against claiming unverified results, and a mandatory final verdict step with defined values (complete/mostly complete/partially complete/not complete/unable to verify). No destructive or batch operations, so the validation cap does not apply.

5 / 5

Progressive Disclosure

The body keeps the default qualitative path inline and appropriately externalizes only the heavy scoring machinery to a well-signaled, one-level-deep reference (references/ret-scored-evaluation.md, verified to exist) linked only from the scored-mode section. Sections are clearly headed and easy to navigate, matching the anchor-5 structure.

5 / 5

Total

18

/

20

Passed

Description

91%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it clearly states what the skill does, provides an explicit and trigger-rich "Use when" clause, includes both positive triggers and a negative boundary, and covers natural synonym variations well. The only weaknesses are slightly abstract action phrasing and minor conflict risk from the generic verbs review/check/eval.

DimensionReasoningScore

Specificity

Lists several concrete actions — "surface evidence-backed strengths, gaps, risks, and next fixes", "build a definition-of-done checklist or rubric", "perform grading, LLM-as-judge, Qworld, or RET evaluation" — but they are phrased at a slightly more abstract level than the fully concrete action lists in the anchor-5 example, so anchor 4 fits best.

4 / 5

Completeness

Explicitly answers both questions: the first sentence states what the skill does, and a clear "Use when asked to..." clause with concrete trigger phrases states when, plus a "Do not use for..." boundary. This directly matches the anchor-5 example; anchor 4's weaker "when" does not apply.

5 / 5

Trigger Term Quality

Comprehensive natural-term coverage: synonyms ("eval, evaluate, review, assess, check"), deictic variants ("current/this/my/our work"), and scoring vocabulary ("scores, grades, points, ratings, weighted criteria, Qworld, RET"). No commonly used trigger synonym is noticeably absent, so it matches the anchor-5 anchor rather than anchor 4.

5 / 5

Distinctiveness Conflict Risk

The niche — reviewing work against the user's actual goal, DoD checklists, Qworld/RET — is distinct with stated boundaries, but the very common verbs "review/check/eval" create minor overlap risk with code-review and general review skills, keeping it just below anchor 5's minimal conflict risk.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mims-harvard/ToolUniverse
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.