CtrlK
BlogDocsLog inGet started
Tessl Logo

alfworld-storage-explorer

Systematically explores storage receptacles (drawers, cabinets, shelves) to find an appropriate placement location for an object. Use when the agent needs to store an item but the exact target receptacle is unknown or ambiguous. Opens, inspects, and closes candidate receptacles to assess suitability, then places the object in the best match.

65

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a lean, actionable three-phase workflow with a concrete command example and error handling, scoring well on conciseness, actionability, and workflow clarity. Its main defects are the truncated 'Output Format' section and the orphaned references/storage_heuristics.md bundle file that no body section points to, which weakens progressive disclosure.

Suggestions

Add a clearly signaled reference to the existing bundle file, e.g. a '## Storage heuristics' section with 'See [storage_heuristics.md](references/storage_heuristics.md) for per-room storage patterns, an exploration priority matrix, and fallback strategies' — the file exists but is currently unreachable from SKILL.md.

Complete the truncated 'Output Format' section — it currently ends at 'Maintain the standard action format:' with no format shown; either show the action format or remove the section.

Merge 'Core Strategy' into 'Execution Pattern' (or cut it to one line) — bullets like 'Inspect Before Placing' and 'Maintain Environment State' restate Phase 2's Navigate/Open/Observe/Close steps.

DimensionReasoningScore

Conciseness

The body is efficient — tight bullets, an appropriately scoped worked example, and no explanation of basic concepts. The 'Core Strategy' section ('Prioritize Exploration', 'Inspect Before Placing') partially restates the 'Execution Pattern' phases and could be trimmed, matching the 4 anchor's minor instances of over-explanation rather than the 5 anchor's every-token-earns-its-place.

4 / 5

Actionability

Mostly executable guidance: the Example section gives a concrete ALFWorld command sequence ('go to drawer 1', 'open drawer 1', 'put sponge 1 in/on drawer 1') with observations, and Phase 2 gives step-by-step actions. The truncated 'Output Format' section ('Maintain the standard action format:' with no content following) is a minor gap keeping it below the 5 anchor.

4 / 5

Workflow Clarity

A clear three-phase sequence with per-receptacle steps, suitability criteria, decision rules, and an error-recovery loop ('If "Nothing happened"... try alternative locations'). It sits between the 4 anchor (most checkpoints present, minor gaps) and the 5 anchor — the incomplete Output Format section and implicit 'stop exploring' checkpoint keep it at 4.

4 / 5

Progressive Disclosure

The body is well-sectioned and appropriately short, but the bundle file references/storage_heuristics.md — which contains complementary decision factors, an exploration priority matrix, and fallback strategies — is never referenced or linked from SKILL.md. Per the guideline to score against the actual bundle structure, references exist but are not signaled at all, fitting the 3 anchor ('references present but not clearly signaled') rather than the 4 anchor's mostly-clear references.

3 / 5

Total

15

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is strong: third-person voice, concrete action verbs, named receptacle types, and an explicit 'Use when...' clause with a specific trigger condition. Trigger-term coverage and distinctiveness are good but not best-in-class due to missing common synonyms and a somewhat broad 'store an item' trigger.

DimensionReasoningScore

Specificity

Lists several concrete, domain-specific actions — "explores storage receptacles (drawers, cabinets, shelves)", "Opens, inspects, and closes candidate receptacles", "places the object in the best match" — with minor gaps (e.g., suitability criteria are not named). It exceeds the 3 anchor (only 1-2 actions) but falls short of the 5 anchor's comprehensive coverage.

4 / 5

Completeness

Explicitly answers both: what ("Systematically explores storage receptacles... Opens, inspects, and closes candidate receptacles... then places the object in the best match") and when ("Use when the agent needs to store an item but the exact target receptacle is unknown or ambiguous") with a concrete trigger condition. This matches the 5 anchor; the 4 anchor's 'when could be more explicit' does not apply since the trigger clause is specific.

5 / 5

Trigger Term Quality

Natural trigger terms like "store an item", "drawers, cabinets, shelves", and "unknown or ambiguous" match what a user or agent would say, but common synonyms such as "put away" or "find a place for" are missing. Good coverage with a few natural terms absent fits the 4 anchor, not the 5 anchor's comprehensive synonym coverage.

4 / 5

Distinctiveness Conflict Risk

The embodied-agent storage-exploration niche with terms like "receptacle" and the unknown/ambiguous-target trigger is mostly distinct with only minor overlap risk against generic tidy-up or object-placement skills. It does not reach the 5 anchor's fully clear niche with minimal conflict risk because 'store an item' is a fairly broad trigger phrase.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
zjunlp/SkillNet
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.