CtrlK
BlogDocsLog inGet started
Tessl Logo

scienceworld-temperature-measurer

This skill uses a thermometer on a substance to measure its temperature. It should be triggered when the agent needs to determine the temperature of a material (e.g., lead) to assess if it has reached a specific melting point or threshold. The skill requires a thermometer in inventory and a target substance, outputting the measured temperature in degrees Celsius, which is key for scientific measurement tasks.

61

Quality

77%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./experiments/src/skills/scienceworld/scienceworld-temperature-measurer/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

60%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body presents a clear, well-sequenced workflow with concrete game commands and useful error-recovery notes, and it is reasonably token-efficient. Its main defects are wiring: the bundled references/action_guide.md (which contains valuable pitfalls and exact observation phrasings) is never linked, and the referenced script measure_melting_point.py is missing from the bundle.

Suggestions

Link the existing references/action_guide.md from SKILL.md (e.g., under Key Actions: 'For exact observation phrasings and troubleshooting pitfalls, see [action_guide.md](references/action_guide.md)'), since it currently contains information the body duplicates or omits.

Resolve the 'Bundled Logic' section: either add scripts/measure_melting_point.py to the bundle or remove the reference, since pointing the agent at a nonexistent script will cause a failed execution.

Trim the duplication between 'Core Workflow' and 'Key Actions & Observations' (pick up / move / activate appear in both) and fold the explicit heater-activation verification from the action guide into the workflow as a validation checkpoint.

DimensionReasoningScore

Conciseness

The body is efficient — short purpose, a 5-step workflow, terse command/action tables, and brief notes with no explanation of concepts Claude already knows. It is not a 5 because the 'Key Actions & Observations' section largely restates steps already covered in 'Core Workflow' (e.g., `pick up thermometer`, `activate [heating device]`) and glosses like "Acquire the essential measuring tool" add little. Not score 3, since the padding is minor duplication rather than unnecessary explanation.

4 / 5

Actionability

Guidance includes concrete game commands (`use thermometer in inventory on [substance]`, `move [substance] to [container]`), the expected observation format ("the thermometer measures a temperature of X degrees celsius"), and a decision rule ("above 150.0 degrees" -> focus on the red box). However it is incomplete: the 'Bundled Logic' section directs the agent to a script `measure_melting_point.py` that does not exist in the bundle, and commands are templates with placeholders rather than a fully worked example, leaving key execution details missing — matching 'some concrete guidance but incomplete'.

3 / 5

Workflow Clarity

The Core Workflow gives a clear 5-step sequence with a conditional heating branch and the Notes section supplies an error-recovery loop ("If `use thermometer` fails, check the substance's state via `look at [container]`"). It falls short of 5 because the main flow lacks explicit validation checkpoints (e.g., verifying the heater is active or the substance's current name before measuring), which the bundled action guide covers but the workflow itself does not.

4 / 5

Progressive Disclosure

Scored against the actual bundle: references/action_guide.md exists but is never mentioned or linked anywhere in SKILL.md, while the one path the body does reference (`measure_melting_point.py`) does not exist in any scripts/ directory — navigation is broken in both directions. The body's own section structure is good, which keeps it above a monolithic score of 1, but the unlinked reference is maximally buried and the dangling script path fails the 'clearly signaled one-level-deep references' standard, fitting 'references are buried' better than the score-3 anchor.

2 / 5

Total

13

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it is written in third person, states concrete inputs (thermometer in inventory, target substance) and output (degrees Celsius), and includes an explicit, concrete trigger clause covering melting-point and threshold scenarios. Its only weaknesses are modest action coverage in the 'what' portion and a slightly padded closing clause.

DimensionReasoningScore

Specificity

The description names the domain and one or two concrete actions — "uses a thermometer on a substance to measure its temperature" and "outputting the measured temperature in degrees Celsius" — but does not enumerate several specific operations the way the score-4/5 anchors require. It sits above 'names the domain but actions are minimal' yet below 'lists several specific actions'.

3 / 5

Completeness

It explicitly answers both questions: what ("uses a thermometer on a substance to measure its temperature... outputting the measured temperature in degrees Celsius") and when ("It should be triggered when the agent needs to determine the temperature of a material (e.g., lead) to assess if it has reached a specific melting point or threshold"). The 'when' clause is explicit and includes concrete trigger conditions, matching the top anchor; the trailing "which is key for scientific measurement tasks" is mild padding but does not obscure either answer.

5 / 5

Trigger Term Quality

Natural phrases a user would say are present — "measure its temperature", "melting point", "threshold", "determine the temperature of a material" — giving good keyword coverage. It falls short of the comprehensive anchor because it lacks common variations (e.g., 'how hot is', 'boiling point', 'heat up') that would round out coverage for this niche.

4 / 5

Distinctiveness Conflict Risk

This is a clear niche — thermometer-based temperature measurement for melting-point/threshold determination in a simulation environment — with distinct triggers unlikely to fire for unrelated skills. It matches 'clear niche with distinct triggers; minimal conflict risk' and is clearly above the 'minor overlap risk' anchor.

5 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
zjunlp/SkillNet
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.