CtrlK
BlogDocsLog inGet started
Tessl Logo

scienceworld-tool-validator

Use when the agent has acquired a tool or instrument and needs to verify it is operational before first use in a critical task step. This skill performs a lightweight pre-use check via "focus on [TOOL] in inventory" and confirms readiness based on the system's response, ensuring the tool is functional before measurement, activation, or connection operations.

62

Quality

72%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./experiments/src/skills/scienceworld/scienceworld-tool-validator/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

77%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable with executable commands and a clear, checkpointed workflow, but mild redundancy in Key Principles and an unreferenced bundle file keep conciseness and progressive disclosure from the top score.

Suggestions

Link the existing reference from the body, e.g., under Workflow or Key Principles add 'See [action_primer.md](references/action_primer.md) for the full list of validation/inspection actions and common tool locations', so the bundle file is discoverable.

Trim the Key Principles section by removing 'Timing' and 'State awareness' items that restate Workflow steps, keeping only the genuinely additive 'Simplicity' (avoid unnecessary examine/use) guidance.

Optionally move the location/navigation detail (teleport to, containers open) that also appears in action_primer.md out of the body to avoid duplication with the reference.

DimensionReasoningScore

Conciseness

The body is efficient with no concept over-explanation, but the 'Key Principles' section restates guidance already in the Workflow (Timing = step 1, State awareness = step 1's locate/retrieve), adding mild redundancy that could be tightened.

2 / 3

Actionability

It provides fully executable commands — 'focus on [TOOL] in inventory', 'pick up thermometer', 'use thermometer on unknown substance B', 'teleport to workshop' — with concrete input/output examples that are copy-paste ready, matching the top anchor.

3 / 3

Workflow Clarity

The 4-step Workflow is clearly sequenced (acquire -> validate -> confirm readiness -> proceed) with an explicit validation checkpoint in step 3 ('A successful response confirms the tool is operational') and an error fallback ('No further diagnostic steps are needed unless an error is observed'), appropriate for this simple skill.

3 / 3

Progressive Disclosure

Sections are well organized, but a bundle file exists at references/action_primer.md and is never linked or signaled anywhere in the body, so the reference is present but not clearly navigable — matching the 'references present but not clearly signaled' anchor.

2 / 3

Total

10

/

12

Passed

Description

67%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description cleanly answers both what the skill does and when to use it with an explicit trigger and a concrete command, but it describes a single action and lacks natural trigger-term variations and an explicit domain name that would sharpen distinctiveness.

Suggestions

Name the ScienceWorld domain explicitly (e.g., 'Use when working in ScienceWorld and the agent has acquired a tool...') to reduce conflict risk with generic tool-readiness skills.

Add more natural trigger variations a user might say, such as 'check a tool works', 'make sure an instrument is ready', or 'confirm a tool before using it', to broaden trigger-term coverage.

Mention the concrete readiness signal a user would look for (e.g., 'confirms readiness when the response reads "You focus on the [TOOL]."') to add specificity beyond the single command.

DimensionReasoningScore

Specificity

The description names a concrete action — 'performs a lightweight pre-use check via "focus on [TOOL] in inventory"' and 'confirms readiness based on the system's response' — but it centers on a single command rather than listing multiple distinct concrete actions, matching the 'names domain and some actions' anchor rather than the 'multiple specific concrete actions' anchor.

2 / 3

Completeness

It explicitly answers both 'what' (lightweight pre-use check via focus on [TOOL], confirms readiness) and 'when' with an explicit 'Use when the agent has acquired a tool or instrument and needs to verify it is operational before first use...' trigger clause, matching the top anchor.

3 / 3

Trigger Term Quality

Terms like 'tool or instrument', 'verify it is operational', and 'before first use' are relevant but framed procedurally/technically ('critical task step', 'measurement, activation, or connection operations') with few natural variations a user would casually say, fitting the 'some relevant keywords but missing common variations' anchor.

2 / 3

Distinctiveness Conflict Risk

The exact command syntax 'focus on [TOOL] in inventory' pins a niche, but the conceptual trigger 'verify a tool is operational before a critical task' is generic and the description never names the ScienceWorld domain, so it could overlap with other readiness-check skills — the 'somewhat specific but could still overlap' anchor.

2 / 3

Total

9

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
zjunlp/SkillNet
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.