CtrlK
BlogDocsLog inGet started
Tessl Logo

scienceworld-material-classifier

This skill makes a determination about a material's property (e.g., conductivity) based on environmental cues or domain knowledge when direct testing fails. Trigger it when experimental actions are invalid or unavailable, requiring a logical inference. It uses observed object properties and common-sense reasoning to classify the material and decide its final disposition.

66

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

96%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This is a tight, well-structured skill body: concrete commands, a clear fallback workflow with a verification step, and a fully worked example, all in under 25 lines. The only flaw is at the bundle level — a duplicate, unreferenced material_properties.json alongside the referenced .md — which slightly muddies navigation.

Suggestions

Remove or explicitly reference the duplicate 'references/material_properties.json' (the body points only to material_properties.md), and either complete or delete the empty 'Decision Flowchart' section in material_properties.md.

DimensionReasoningScore

Conciseness

The ~20-line body is lean and efficient with zero padding: every section (When to Use, Procedure, Example) earns its place and nothing explains concepts Claude already knows. This matches anchor 5 exactly; anchor 4 would require unnecessary explanation to trim, and none exists.

5 / 5

Actionability

Fully concrete, executable commands are given throughout — 'focus on <OBJECT>', 'connect <OBJECT> terminal 1 to <WIRE> terminal 2', 'move <OBJECT> to <CONTAINER>', 'look at <CONTAINER>' — and the worked example (glass jar) is copy-paste ready for the common case. This matches anchor 5's 'specific examples cover the common cases' rather than anchor 4's 'minor gaps'.

5 / 5

Workflow Clarity

The 5-step procedure is clearly sequenced with an explicit validation checkpoint (step 5: 'look at <CONTAINER> — verify the object was placed correctly') and an error-recovery branch (step 3 fallback when testing fails), reinforced by the example showing the confirming observation. Not a destructive/batch operation, and as a simple single-purpose skill under 50 lines it qualifies for the top anchor.

5 / 5

Progressive Disclosure

The single reference ('Consult references/material_properties.md for lookup') is real, one level deep, and clearly signaled, and the body sections are well organized. However, the bundle also contains an unreferenced duplicate 'references/material_properties.json' (and the .md ends with an empty 'Decision Flowchart' section), which is a minor organization gap keeping it below anchor 5.

4 / 5

Total

19

/

20

Passed

Description

62%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is well-formed third-person prose that explicitly answers both 'what' and 'when', with a genuinely distinctive fallback-inference niche. Its weaknesses are abstract phrasing in place of concrete action verbs and trigger language that leans on game-specific jargon instead of the natural terms a user would say.

Suggestions

Replace abstract verbs with concrete actions, e.g. 'Determines whether a material is conductive, magnetic, or insulating by inferring from its composition, then sorts it into the correct container.'

Use natural trigger phrases users would actually say, e.g. 'Use when a material test fails, the circuit test is unavailable, or you need to classify or sort an object (conductivity, magnetism, insulator/conductor) without testing equipment.'

Name the domain context and property variations (magnetism, solubility, conductivity) in the trigger clause to reduce overlap with generic material-sorting skills.

DimensionReasoningScore

Specificity

The description names the domain and 1-2 actions ("classify the material and decide its final disposition"), but "makes a determination about a material's property" is abstract phrasing rather than a list of concrete actions. It does not reach anchor 4's 'several specific actions' and is clearly above anchor 2's generic minimal actions.

3 / 5

Completeness

Both parts are explicitly present: what ("classify the material and decide its final disposition") and when ("Trigger it when experimental actions are invalid or unavailable"). The 'when' clause is explicit but stated in abstract terms rather than the concrete trigger phrases of anchor 5, fitting anchor 4's 'when could be more explicit or specific'.

4 / 5

Trigger Term Quality

Some relevant keywords are present ("conductivity", "when direct testing fails", "classify"), but the phrasing "experimental actions are invalid or unavailable" is game-specific jargon rather than natural user language, and common variations (magnetism, conductor/insulator, sorting, material identification) are missing. This matches anchor 3's 'some relevant keywords but missing common variations' rather than anchor 4's good coverage.

3 / 5

Distinctiveness Conflict Risk

The fallback-inference niche (classify only when direct testing fails) is distinct with a clear trigger condition. However, "classify the material and decide its final disposition" overlaps with any general material-sorting or testing skill, matching anchor 4's 'mostly distinct; minor overlap risk' rather than anchor 5's minimal conflict risk.

4 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
zjunlp/SkillNet
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.