CtrlK
BlogDocsLog inGet started
Tessl Logo

alfworld-tool-user

Use when the agent needs to apply a tool to a target object in ALFWorld to accomplish an interaction such as cleaning, heating, cooling, or examining. This skill handles locating both the tool and target object, then executing the correct environment action (e.g., `clean`, `heat`, `cool`, `use`) to progress the task.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an exemplary piece of skill writing — lean, fully executable, and sequenced with an explicit verification step and error-recovery guidance. Its one real weakness is bundle integration: two reference files ship alongside the skill but are never surfaced from the body, leaving them undiscoverable and duplicating inlined content.

Suggestions

Add clearly signaled one-level-deep pointers in the body, e.g., under Core Workflow: 'For the full task-goal-to-verb mapping, see [references/action_reference.md](references/action_reference.md)' and 'For a worked examine trajectory, see [references/trajectory_example.md](references/trajectory_example.md)'.

Trim duplication between the body and the bundle: keep only the most common interaction rows inline in the table and move the complete action templates, prerequisite actions, and error-signal detail into references/action_reference.md.

Fold the 'toggle {obj} {recep}' action from the reference file into the body's interaction table (or link to it) so the documented action set matches what the reference actually covers.

DimensionReasoningScore

Conciseness

The ~60-line body is lean: a compact interaction table, short workflow steps, and a single worked example, with no explanations of concepts Claude already knows. Details like "tools are typically stationary appliances" and "The agent must be holding the target object before most interactions" each earn their tokens; anchor 4 would require trimmable over-explanation, which is absent.

5 / 5

Actionability

The body gives copy-paste-ready commands — the interaction table ("`clean knife 1 with sinkbasin 1`") and the full executable trajectory ("go to countertop 1" → "clean knife 1 with sinkbasin 1") — covering the common cases. This matches the anchor-5 example of fully executable, example-covered guidance; anchor 4 would imply minor gaps in the concrete commands.

5 / 5

Workflow Clarity

The four-step workflow (Identify → Locate and Acquire → Execute → Verify Outcome) includes an explicit validation checkpoint ("Check the environment observation after the action") and an Error Handling section with recovery guidance ("Reassess whether the agent is at the correct location and holding the correct object", "Use `take` first"). That feedback loop matches anchor 5; anchor 4 would require missing or merely implicit checkpoints, which is not the case.

5 / 5

Progressive Disclosure

Per the guideline to score against the actual bundle: `references/action_reference.md` and `references/trajectory_example.md` exist but are never referenced or linked anywhere in the body, while overlapping content (the action table and error signals) is inlined instead. This matches anchor 3 ("references present but not clearly signaled; content that should be separate is inline") — not anchor 4, since the references are entirely orphaned rather than mostly clear, and not anchor 2, since the body itself is well-structured rather than minimal.

3 / 5

Total

18

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states what the skill does concretely, gives an explicit 'Use when' trigger, and is tightly scoped to the ALFWorld domain. The only soft spot is keyword coverage, which could add a few synonyms for how users phrase tool-use tasks.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — "cleaning, heating, cooling, or examining" and "executing the correct environment action (e.g., `clean`, `heat`, `cool`, `use`)" — plus the two-part locating behavior ("locating both the tool and target object"), giving comprehensive coverage of the domain. Anchor 4 would require minor gaps in coverage, but the verb list spans the skill's full interaction set.

5 / 5

Completeness

It explicitly answers 'when' with a concrete "Use when the agent needs to apply a tool..." clause and 'what' with "handles locating both the tool and target object, then executing the correct environment action". Both are present, explicit, and concrete, matching the anchor-5 example pattern; anchor 4 would require the 'when' to be less specific.

5 / 5

Trigger Term Quality

Natural trigger phrases like "apply a tool to a target object" and "cleaning, heating, cooling, or examining" mirror how ALFWorld tasks are phrased, giving good keyword coverage. It falls short of anchor 5 because no synonyms or task-phrasing variants (e.g., "wash", "examine X with Y") are included.

4 / 5

Distinctiveness Conflict Risk

Scoping to "ALFWorld" tool-to-object interactions with distinct action verbs carves out a clear niche with minimal overlap risk against other skills. Anchor 4 would imply meaningful overlap with closely related skills, which the domain-specific scoping avoids.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
zjunlp/SkillNet
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.