CtrlK
BlogDocsLog inGet started
Tessl Logo

devtu-optimize-skills

Optimize ToolUniverse skills for better report quality, evidence handling, and user experience. Apply patterns like tool verification, foundation data layers, disambiguation-first, evidence grading, quantified completeness, and report-only output. Use when reviewing skills, improving existing skills, or creating new ToolUniverse research skills.

66

Quality

79%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/devtu-optimize-skills/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, expert-level body: mostly token-efficient tables, concrete commands and interpretation thresholds, a clearly phased workflow with fallback handling, and verified one-level-deep references. The main gaps are placeholder-heavy templates in Pattern 15, some duplication between the pattern summary table and later inline sections, and inline elaboration of patterns that belong in the reference file.

Suggestions

Move the full Pattern 14/15/15b detail (interpretation tables, procedure templates, and the download-only dataset table) into references/optimization-patterns.md, keeping only the pattern names and one-line key ideas inline to match how Patterns 1-13 are handled.

Replace the placeholder Python template in Pattern 15 with one complete, runnable example (e.g., a working pandas/scipy snippet with example data) per the skill's own 'Include example data' rule.

Add explicit validation checkpoints to the Optimized Skill Workflow phases (e.g., what to verify after Phase 0/1 before proceeding), mirroring the fallback and retry guidance already present.

DimensionReasoningScore

Conciseness

The body is dense and assumes domain competence (no basic-concept explanations), with tables carrying most of the payload, e.g. the 13-pattern summary table and anti-pattern/fix tables. Not 5 because Pattern 15's markdown 'template for a template' with escaped code fences and '[What this computes]' placeholders, plus patterns being summarized in the table and then re-elaborated in later sections, could be trimmed. Not 3 because there is little genuine padding.

4 / 5

Actionability

Provides concrete, executable guidance: 'python3 -m tooluniverse.cli run <Tool> \'<json>\'', 'get_tool_info() before first call', and specific thresholds like 'PRR > 5 = strong signal' and 'Score < -0.5 = essential'. Not 5 because the Python code block in the Pattern 15 template is placeholder scaffolding rather than copy-paste-ready working code; not 3 because most guidance is specific and directly usable.

4 / 5

Workflow Clarity

The workflow is clearly sequenced (Phase -1 tool verification through Phase 3 report synthesis) with a validation-first phase, fallback chains ('Primary → Fallback 1 → Fallback 2 → document unavailable'), and a retry-vs-fix distinction ('Distinguish transient errors (retry) from real bugs (fix)'). Not 5 because the phases are one-line labels without explicit per-phase validation checkpoints or a feedback loop for the workflow itself.

4 / 5

Progressive Disclosure

Well-signaled one-level-deep references — 'Full details: [references/optimization-patterns.md]', 'Full details: [references/testing-standards.md]', and an Additional References section listing all four reference files, all of which exist and contain no nested references. Not 5 because substantial content (Patterns 14, 15, and 15b with the dataset download table and templates) is elaborated inline despite a dedicated optimization-patterns reference file, which is more than a minor organization gap.

4 / 5

Total

16

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly states what the skill does, enumerates concrete patterns, and gives an explicit 'Use when...' clause with three concrete triggers, all in third person. Minor room for improvement in trigger synonyms and tightening the vague 'better user experience' phrasing.

DimensionReasoningScore

Specificity

Lists several concrete actions/patterns — 'tool verification, foundation data layers, disambiguation-first, evidence grading, quantified completeness, and report-only output' — though 'better report quality... and user experience' is somewhat vague. Not 5 because the action list is a set of pattern names rather than fully comprehensive concrete capabilities; not 3 because it clearly goes beyond 1-2 actions.

4 / 5

Completeness

Explicitly answers both: what ('Optimize ToolUniverse skills for better report quality, evidence handling, and user experience. Apply patterns like...') and when ('Use when reviewing skills, improving existing skills, or creating new ToolUniverse research skills') with concrete trigger phrases. Matches the score-5 anchor directly; score 4 would require the 'when' to be less explicit, which it is not.

5 / 5

Trigger Term Quality

'Use when reviewing skills, improving existing skills, or creating new ToolUniverse research skills' covers natural phrasings a user would say, plus the domain term 'ToolUniverse'. Not 5 because common variations like 'optimize a skill', 'refactor', or 'fix a skill' are absent.

4 / 5

Distinctiveness Conflict Risk

Scoped to a clear niche (ToolUniverse research skills) with distinct triggers, but the trigger phrase 'reviewing skills, improving existing skills' does not repeat the ToolUniverse qualifier until the last item, creating minor overlap risk with general skill-improvement skills. Not 5 due to that minor overlap; not 3 since the domain scoping is unambiguous.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
mims-harvard/ToolUniverse
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.