CtrlK
BlogDocsLog inGet started
Tessl Logo

scienceworld-tool-validator

Use when the agent has acquired a tool or instrument and needs to verify it is operational before first use in a critical task step. This skill performs a lightweight pre-use check via "focus on [TOOL] in inventory" and confirms readiness based on the system's response, ensuring the tool is functional before measurement, activation, or connection operations.

65

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable — exact commands, expected system responses, and worked transcripts — with a clear workflow and an explicit readiness checkpoint. Its weaknesses are redundancy (Purpose/When-to-Use/Key Principles overlap, a duplicative second example) and an orphaned bundle file: references/action_primer.md exists but is never referenced from SKILL.md.

Suggestions

Reduce redundancy: merge Purpose and When-to-Use into one short section, trim Key Principles to only non-obvious guidance (e.g., drop 'Validate immediately after acquisition' and the teleport note, which restate the Workflow), and cut or compress Example 2 since it repeats Example 1's validation flow with only a different tool.

Signal the bundle: add a clearly labeled reference to references/action_primer.md (e.g., an 'Advanced actions' note linking to it) so the primer's examine/look-at distinctions and tool-location guidance are discoverable instead of partially duplicated inline.

Add an error-recovery step to the Workflow: specify what to do when the focus-on response is an error (e.g., re-acquire the tool, try 'examine [TOOL]', or fall back to locating a replacement) so the validation checkpoint has a feedback loop.

DimensionReasoningScore

Conciseness

Purpose and When-to-Use largely restate the frontmatter, Example 2 duplicates Example 1's validation flow with only a different tool, and Key Principles repeats Workflow/Examples content ("Validate immediately after acquisition", "Use 'teleport to'"). Mostly efficient but with several redundant sections that could be tightened, matching the 3 anchor rather than the minor-trim 4.

3 / 5

Actionability

Fully executable, copy-paste-ready commands ("focus on [TOOL] in inventory", "pick up [TOOL]", "use thermometer on [TARGET]") plus two complete transcripts showing exact inputs and expected system responses ("You focus on the thermometer."). The concrete expected-response string makes the readiness criterion unambiguous.

5 / 5

Workflow Clarity

A clear four-step sequence with an explicit validation checkpoint ("A successful response ('You focus on the [TOOL].') confirms the tool is operational") and a simple-skill structure. Held below 5 because the error branch ("unless an error is observed") gives no recovery guidance — no feedback loop for what to do when validation fails.

4 / 5

Progressive Disclosure

The bundle provides references/action_primer.md, but the body never links or mentions it, so the reference is unsignaled and its content (examine/look at distinctions, container assumptions, teleport note) partially duplicates Key Principles inline instead of being delegated. Sections are well organized, but the un-navigated reference file fits the 3 anchor ('references present but not clearly signaled') rather than the 4.

3 / 5

Total

15

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it explicitly states both what the skill does and when to use it, anchored by the concrete command 'focus on [TOOL] in inventory' and specific trigger scenarios (acquisition, mid-task tool switch, post-navigation re-check). The main gap is that the host environment (ScienceWorld) is never named, which slightly weakens distinctiveness.

DimensionReasoningScore

Specificity

Quotes the exact validation command ("focus on [TOOL] in inventory") and the concrete confirmation behavior ("confirms readiness based on the system's response"), giving several specific actions. Falls short of 5 because coverage is narrow — a single check described rather than a comprehensive set of capabilities.

4 / 5

Completeness

Explicitly answers both questions with concrete triggers: the what ("performs a lightweight pre-use check via 'focus on [TOOL] in inventory' and confirms readiness based on the system's response") and the when ("Use when the agent has acquired a tool... before first use in a critical task step"). This matches the 5 anchor rather than the 4, whose 'when' is only serviceable.

5 / 5

Trigger Term Quality

Natural trigger moments are well covered ("acquired a tool or instrument", "verify it is operational", "before first use", "measurement, activation, or connection"). A few natural synonyms (e.g., "check the tool works", "validate") are missing, keeping it below the comprehensive-anchor 5.

4 / 5

Distinctiveness Conflict Risk

The domain-specific command syntax and the operational-check framing create a mostly distinct niche, but the description never names ScienceWorld, so it could overlap with generic tool-verification needs in other environments. Minor overlap risk fits the 4 anchor rather than the minimal-conflict 5.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
zjunlp/SkillNet
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.