CtrlK
BlogDocsLog inGet started
Tessl Logo

scienceworld-task-interpreter

This skill parses a user's high-level scientific task in the ScienceWorld environment and extracts the core objective and target location. It should be triggered when a new task instruction is received, especially those involving finding, comparing, or manipulating objects. The skill interprets the query to identify the goal (e.g., 'find the animal with the shortest life span') and any specified locations (e.g., 'animals are in the outside location'), outputting a clear, actionable sub-goal for navigation or observation.

64

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./experiments/src/skills/scienceworld/scienceworld-task-interpreter/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, well-structured instruction skill whose plan template is genuinely usable, undermined by an unfinished Output Format section and an orphaned reference file. The command vocabulary the body leans on (e.g., teleport syntax, room names) lives in references/action_glossary.md, which the body never points to, and the document ends mid-promise at '**Output Format:**'.

Suggestions

Complete the '**Output Format:**' section — it currently announces a format and then stops; add a concrete template for the Thought step (e.g., objective, target object, location, planned action).

Link references/action_glossary.md from the body (e.g., under 'Formulate the Plan': 'For full action syntax and room names, see [action_glossary.md](references/action_glossary.md)') so the bundled action vocabulary is discoverable.

State the teleport command explicitly in the Navigate step (e.g., `teleport to LOC`) instead of only 'teleport to it immediately', and add a fallback for tasks where no location is specified.

DimensionReasoningScore

Conciseness

The body is lean and assumes Claude's competence ('Extract the following core components', 'use `look around` to survey the environment'), with no explanations of known concepts. Minor filler remains, e.g. 'Read the user's instruction carefully' and 'This confirms the skill has correctly parsed the task', which keeps it just below the 'every token earns its place' anchor.

4 / 5

Actionability

Concrete commands are present (`look around`, `focus on [OBJECT]`, `pick up [OBJECT]`), but key details are missing: the final section announces '**Output Format:**' and then provides no format at all, and the Navigate step says 'teleport to it immediately' without giving the command syntax. This matches 'some concrete guidance but incomplete; missing key details' rather than the mostly-executable level above.

3 / 5

Workflow Clarity

The sequence is clear and well-ordered (Parse → Navigate → Observe → Analyze & Execute) with an explicit checkpoint: 'Before taking the first action, articulate your interpretation and plan in a Thought step.' Minor validation gaps — no guidance for when no location is specified, and the output-format step is incomplete — keep it below the explicit-validation anchor.

4 / 5

Progressive Disclosure

The body is well-sectioned, but the bundle file `references/action_glossary.md` (which contains the action syntax, room list, and container notes this skill depends on) is never referenced or linked anywhere in the body, so the reference exists but is not clearly signaled. This fits 'some structure but could be better organized; references present but not clearly signaled' rather than level 4, where references would be mostly clear.

3 / 5

Total

14

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it explicitly states both capability and trigger conditions in third person, grounds them in concrete examples, and is unambiguously scoped to ScienceWorld. The only weakness is mild redundancy between 'parses' and 'interprets' and slightly abstract phrasing of the output ('clear, actionable sub-goal').

DimensionReasoningScore

Specificity

It lists several concrete actions — 'parses a user's high-level scientific task', 'extracts the core objective and target location', 'interprets the query to identify the goal', 'outputting a clear, actionable sub-goal' — with illustrative examples. It falls short of the comprehensive anchor because 'parses'/'interprets' are somewhat repetitive and 'clear, actionable sub-goal' is abstract, and it is clearly above the 1-2-actions anchor.

4 / 5

Completeness

Both halves are explicit: what it does ('parses... extracts the core objective and target location... outputting a clear, actionable sub-goal') and when to use it ('It should be triggered when a new task instruction is received, especially those involving finding, comparing, or manipulating objects'), backed by concrete example phrases for both goals and locations. This matches the anchor that clearly and explicitly answers both what AND when with concrete trigger phrases; the level-4 anchor's 'when could be more explicit' does not apply.

5 / 5

Trigger Term Quality

Trigger phrases are natural for this domain: 'when a new task instruction is received, especially those involving finding, comparing, or manipulating objects'. Coverage is good but a few natural variations (e.g., 'complete the task', 'goal', task-type synonyms) are missing, which fits 'good keyword coverage; a few natural terms missing' rather than the comprehensive synonym coverage of level 5.

4 / 5

Distinctiveness Conflict Risk

The description is tightly scoped to 'the ScienceWorld environment' with domain-specific triggers (task instructions about finding, comparing, or manipulating objects in that environment), giving it a clear niche with minimal overlap risk against generic task-parsing skills. It matches the 'clear niche with distinct triggers' anchor.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
zjunlp/SkillNet
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.