CtrlK
BlogDocsLog inGet started
Tessl Logo

alfworld-device-operator

Operates a device or appliance (like a desklamp, microwave, or fridge) to interact with another object. Use when the task requires using a tool on a target item (e.g., "look at laptop under the desklamp", "heat potato with microwave"). Locates both the device and target object, co-locates them, and executes the appropriate use action (toggle, heat, cool, or clean).

68

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable skill body with a clear four-phase workflow, an error-recovery rule, and a concrete worked example. The main weaknesses are the orphaned bundle reference (device_location_guide.md is never linked, leaving its content undiscoverable), an inconsistency between the documented `toggle` action and the `use` action in the example, and reactive-only validation.

Suggestions

In Phase 1, link the existing bundle file (e.g., "For typical device locations, see [references/device_location_guide.md](references/device_location_guide.md)") so the guide's lookup table is actually discoverable instead of orphaned.

Resolve the action inconsistency: Phase 4 says to use `toggle {device} {recep}` for the desklamp, but the worked example executes `use desklamp 1` — pick the correct action format and align both places.

Add an explicit validation checkpoint before Phase 4 (e.g., confirm the observation reports the object in inventory and the device is visible at the current receptacle) instead of relying solely on the reactive "Nothing happened" recovery rule.

DimensionReasoningScore

Conciseness

The body is efficient — no explanation of concepts Claude already knows, and the Example and Key Assumptions sections earn their tokens. Minor trimming is possible: "Section 1: Skill Trigger" largely restates the frontmatter description, and steps like "Identify the device from the task description" state the obvious. That matches the 4 anchor (efficient, minor instances of over-explanation) rather than 5 (every token earns its place).

4 / 5

Actionability

Gives exact action formats (`go to {recep}`, `take {obj} from {recep}`, `heat {obj} with {device}`) and a fully worked Thought/Action/Observation example that is copy-paste executable. It stops short of 5 because Phase 4 prescribes `toggle {device} {recep}` for the desklamp while the worked example uses `use desklamp 1` — an inconsistency that leaves the correct action ambiguous — and only the desklamp case is exemplified.

4 / 5

Workflow Clarity

A clear four-phase numbered sequence with an error-recovery feedback loop ("If the environment responds with 'Nothing happened,' re-evaluate your object/device names and your location") and an explicit co-location checkpoint. It misses 5 because validation is reactive rather than staged — there is no explicit checkpoint confirming the object is held or the agent is at the device before executing the final action.

4 / 5

Progressive Disclosure

The body itself is well organized with clear section headers, but the bundle contains `references/device_location_guide.md` — a detailed device-to-receptacle lookup table explicitly written to support Phase 1 — that is never referenced anywhere in SKILL.md. Per the guideline to score against the actual bundle structure, this is a reference that is present but not signaled, and its content is only partially duplicated inline ("Search common receptacles (e.g., desks, sidetables, countertops)"). That matches the 3 anchor; it is above 2 because the body's own structure is good, and below 4 because a provided reference is undiscoverable from the overview.

3 / 5

Total

15

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly states what the skill does and when to use it, with concrete natural-language trigger phrases and specific capability verbs. The only gap is that cool/clean trigger phrasings are represented only as verbs in the capability list rather than as example user phrases.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions in third person — "Locates both the device and target object, co-locates them, and executes the appropriate use action (toggle, heat, cool, or clean)" — with named devices (desklamp, microwave, fridge) giving comprehensive coverage of the skill's capability space. It matches the 5 anchor (multiple specific concrete actions, comprehensive) rather than 4, which would require noticeable coverage gaps.

5 / 5

Completeness

Explicitly answers both questions: what ("Operates a device or appliance ... Locates both the device and target object, co-locates them, and executes the appropriate use action") and when ("Use when the task requires using a tool on a target item" with concrete example trigger phrases). This is a direct match to the 5 anchor's pattern; the 4 anchor would require the 'when' to be less explicit.

5 / 5

Trigger Term Quality

Includes a natural "Use when" clause with realistic user phrases ("look at laptop under the desklamp", "heat potato with microwave") and synonyms (device, appliance, tool). It falls short of the 5 anchor because trigger phrases for the cool and clean actions (e.g., "clean plate with spraybottle") are absent, though the verbs appear in the capability list.

4 / 5

Distinctiveness Conflict Risk

It carves a clear niche — device-on-object operation with co-location semantics — and its trigger phrases ("look at X under the desklamp", "heat potato with microwave") are unlikely to fire for unrelated skills. Minimal conflict risk, matching the 5 anchor rather than 4 ("minor overlap risk with closely related skills").

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
zjunlp/SkillNet
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.