CtrlK
BlogDocsLog inGet started
Tessl Logo

alfworld-task-verifier

Use when the agent needs to check whether an ALFWorld task objective has been met after completing a sub-action (e.g., placing an object). This skill parses the task goal, evaluates the latest environment observation, and outputs a verification decision — task complete, task incomplete, or action ineffective — to guide the next step.

73

Quality

92%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an exemplary lean, actionable instruction skill with exact output formats, a decision table, worked examples, and error-recovery loops. Its one real weakness is navigation: two complementary reference files ship in the bundle but are never mentioned in SKILL.md, so their content (task grammar parsing, extra verification scenarios) is effectively invisible.

Suggestions

Add one-level-deep links to the existing bundle files, e.g. under 'Parse the Task Goal': 'See [references/task_grammar.md](references/task_grammar.md) for task grammar patterns' and after the decision table: 'Additional worked scenarios: [references/verification_examples.md](references/verification_examples.md)'.

Trim the in-body examples to one success case and point to references/verification_examples.md for the rest — the body currently duplicates scenarios already covered there (failed action, partial placement), saving tokens now spent on redundant examples.

Surface the key parsing rules that only exist in the reference (e.g., 'any instance of the object type satisfies the goal; instance IDs do not matter') as a one-line note in step 1, or explicitly defer to the reference so the rule is not lost when parsing unusual task phrasings.

DimensionReasoningScore

Conciseness

The body is lean throughout — a compact condition/decision/output table, two short fenced examples, and a two-item error-handling section — with no explanation of concepts Claude already knows and no padded sections, matching the 'every token earns its place' anchor.

5 / 5

Actionability

Guidance is fully concrete and executable for an instruction-only skill: exact output strings ("Verification: Task complete. No further action needed."), a condition-to-decision mapping, and worked Action/Observation/Verification traces covering the common cases including failure ("Nothing happened").

5 / 5

Workflow Clarity

The four-step sequence (parse goal, analyze observation, decide, output) is clearly ordered, and the Error Handling section supplies explicit feedback loops for ambiguous observations ("navigate to the target receptacle and re-examine it") and uncertain counts; the skill is itself a validation checkpoint, so the missing-validation cap does not apply.

5 / 5

Progressive Disclosure

The body is well-structured but never links to the two bundle files that exist (references/task_grammar.md with parsing rules and references/verification_examples.md with additional worked cases), so those materials are undiscoverable from SKILL.md — matching 'references present but not clearly signaled' rather than the minor-gap anchor 4.

3 / 5

Total

18

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it uses third-person voice, states concrete capabilities with enumerated outcomes, and gives an explicit 'Use when' trigger tied to a specific domain. The only weaknesses are minor — a missing synonym set for trigger terms and omission of one outcome (wrong-receptacle placement) that the body handles.

DimensionReasoningScore

Specificity

The description lists three concrete actions ("parses the task goal, evaluates the latest environment observation, and outputs a verification decision") with enumerated outcomes, but omits the fourth outcome covered in the body (object placed in wrong receptacle), leaving a minor coverage gap that fits the 'several specific actions; minor gaps' anchor rather than the comprehensive anchor.

4 / 5

Completeness

It explicitly answers both questions: a clear 'what' (parses the goal, evaluates the observation, outputs one of three verification decisions) and a concrete 'when' ("Use when the agent needs to check whether an ALFWorld task objective has been met after completing a sub-action (e.g., placing an object)"), matching the top anchor with explicit trigger phrasing.

5 / 5

Trigger Term Quality

"ALFWorld task objective", "after completing a sub-action", and "placing an object" are natural terms an agent would use in this domain, but common variations such as "verify the goal", "check task status", or "goal condition" are missing, matching the 'good keyword coverage; a few natural terms missing' anchor.

4 / 5

Distinctiveness Conflict Risk

"ALFWorld" names a specific benchmark environment, giving the skill a clear niche with triggers that would not plausibly fire for any other skill, matching the 'clear niche with distinct triggers; minimal conflict risk' anchor.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
zjunlp/SkillNet
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.