CtrlK
BlogDocsLog inGet started
Tessl Logo

alfworld-goal-interpreter

Parses the natural language task goal to extract actionable sub-objectives and required objects. Trigger this skill whenever a new task is assigned to break down complex instructions into clear, sequential targets. It interprets phrases like 'look at X under Y' to identify target objects (pillow), reference objects (desklamp), and spatial relationships (under).

62

Quality

78%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./experiments/src/skills/alfworld/alfworld-goal-interpreter/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A lean, well-sequenced protocol body whose structure respects token budget and whose workflow includes monitoring and fallback loops. Its main defects are executability gaps — a referenced parse_goal.py that is missing from the bundle and a vague action-mapping step — plus a poorly signaled reference to search_patterns.md.

Suggestions

Ship the actual parse_goal.py in a scripts/ directory (and reference it as scripts/parse_goal.py), or remove the script reference and inline the parsing rules — the body currently instructs use of a file that does not exist in the bundle.

Give references/search_patterns.md a clearly signaled pointer, e.g. a dedicated line like 'Fallback search patterns: see [references/search_patterns.md](references/search_patterns.md)' instead of the mid-sentence mention in section 4.

Add a concrete example of the parsed output dictionary for the 'look at pillow under the desklamp' case, and make section 3's action mapping explicit (e.g., 'look at X under Y' → 'go to desklamp' → 'examine beneath') so the guidance is directly executable.

DimensionReasoningScore

Conciseness

The ~30-line body is lean and assumes Claude's competence: a tight input/output contract ('primary_target', 'reference_object', 'spatial_relation', 'action'), compact IF/THEN relation logic, and a one-line critical check. No padding and no explanation of concepts Claude already knows. Not 4 because there are essentially no tokens to trim.

5 / 5

Actionability

The output dictionary fields and IF/THEN sub-goal logic are concrete, but the invoked 'parsing script (parse_goal.py)' does not exist anywhere in the bundle, and section 3 is vague ('must be translated into the agent's available action set (go to, take, use, etc.)') with no concrete example of parsed output. Not 2 because real structured guidance exists; not 4 because a key executable artifact is missing and the action mapping is incomplete.

3 / 5

Workflow Clarity

The four sections form a clear sequence (parse → generate sub-objectives → map objects/actions → execute & adapt) with an explicit checkpoint ('After each action, monitor the observation') and a failure-feedback loop ('If the expected object is not found, or the action fails ("Nothing happened"), consult the fallback logic'). Not 5 because the parse output is never validated and the fallback loop is deferred rather than explicit.

4 / 5

Progressive Disclosure

Scored against the actual bundle: references/search_patterns.md exists and holds genuinely separate material, but it is buried mid-sentence in section 4 with no clear signal or correct path, and parse_goal.py is referenced yet absent from the bundle entirely. Not 4 because a referenced path is broken and the existing reference is not clearly signaled; not 2 because the body does have section structure and only one level of referencing.

3 / 5

Total

15

/

20

Passed

Description

76%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that explicitly states both capability and trigger, with a concrete worked phrase that grounds the skill's purpose. Main weaknesses are partial keyword coverage (missing natural synonyms) and a somewhat broad 'new task' trigger that invites overlap with other task-management skills.

DimensionReasoningScore

Specificity

The description lists several concrete actions — 'Parses the natural language task goal to extract actionable sub-objectives and required objects' and 'interprets phrases like "look at X under Y" to identify target objects (pillow), reference objects (desklamp), and spatial relationships (under)' — but coverage has minor gaps (no mention of planning execution, fallback search, or action mapping). Not 3 because it goes beyond 1-2 actions; not 5 because the capability list is not comprehensive.

4 / 5

Completeness

It clearly answers 'what' ('Parses the natural language task goal to extract actionable sub-objectives and required objects') and explicitly answers 'when' with a concrete trigger clause ('Trigger this skill whenever a new task is assigned to break down complex instructions') plus a concrete example phrase. Not 4 because the 'when' is explicit and specific rather than needing more detail.

5 / 5

Trigger Term Quality

Relevant keywords are present ('task goal', 'new task is assigned', 'break down complex instructions') but natural variations and synonyms ('objective', 'command', 'instruction parsing') are missing, and terms like 'sub-objectives' lean technical. Not 4 because keyword coverage is partial rather than good with only a few gaps.

3 / 5

Distinctiveness Conflict Risk

The goal-interpretation function with its concrete phrase example ('look at X under Y') carves a clear niche, but the trigger 'whenever a new task is assigned' is broad and could overlap with other task-handling or planning skills. Not 5 due to that overlap risk; not 3 because the function itself is well differentiated.

4 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
zjunlp/SkillNet
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.