CtrlK
BlogDocsLog inGet started
Tessl Logo

scienceworld-measurement-taker

Use when the agent needs to measure a quantitative property (temperature, weight, pH) of a target object or substance using a measurement tool. This skill covers acquiring the tool, preparing both tool and target with focus actions, executing the measurement via "use [TOOL] on [TARGET]", and interpreting the resulting value for conditional decisions such as classification or placement.

73

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

80%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is concise, actionable, and clearly sequenced, but it fails to surface its own bundle files: neither reference is linked or signaled from the main skill, and validation/error-recovery loops are not surfaced inline in the workflow.

Suggestions

Add one-level-deep navigation links to the existing bundle files, e.g. under a 'References' section: 'Action glossary: See [action_glossary.md](references/action_glossary.md)' and 'Decision & verification checklist: See [decision_workflow.md](references/decision_workflow.md)'.

Surface a brief inline validation/feedback loop in the Workflow (e.g., after measuring: confirm the numeric value parsed, and if the observation is not parseable, fall back to the error-recovery steps in decision_workflow.md) so checkpoints are explicit rather than deferred to a reference.

Add a short note that the measurement is retained across teleport, so positioning before/after measuring is flexible — this is a non-obvious recovery fact currently buried in the reference.

DimensionReasoningScore

Conciseness

The body is lean and assumes Claude's competence — it never explains what temperature, weight, or pH are, and every section (Purpose, When to Use, Workflow, Examples, Key Principles) earns its place, matching the lean/efficient anchor; it is above level 2 because there is no unnecessary padding.

3 / 3

Actionability

It provides fully executable, concrete commands in transcript form ('pick up thermometer', 'focus on thermometer in inventory', 'use thermometer on unknown substance B', 'move unknown substance B to orange box') plus templated workflow steps, matching the copy-paste-ready anchor; the placeholders are justified templating rather than pseudocode, so it is not capped at level 2.

3 / 3

Workflow Clarity

The 7-step sequence is clearly ordered with focus-action readiness checkpoints, but explicit validation/error-recovery feedback loops are not surfaced in the body — they live only in the unlinked decision_workflow.md reference — leaving checkpoints implicit rather than explicit (level 2), short of the explicit-validation level 3.

2 / 3

Progressive Disclosure

The body is well-organized into sections, but two bundle files (references/action_glossary.md, references/decision_workflow.md) exist and are never referenced or linked from the main file, so the supplementary detail is poorly signaled and hard to discover — matching the 'references present but not clearly signaled' anchor rather than the well-navigated level 3.

2 / 3

Total

10

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description that covers what the skill does and when to use it with concrete, natural trigger terms. It clearly distinguishes itself via the domain-specific measurement syntax and focus-action ritual.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — 'acquiring the tool', 'preparing both tool and target with focus actions', 'executing the measurement via use [TOOL] on [TARGET]', and 'interpreting the resulting value' — matching the anchor for several specific concrete actions; it is not merely naming a domain (level 2).

3 / 3

Completeness

It explicitly answers both 'what' (acquire, prepare, execute, interpret) and 'when' with an explicit 'Use when the agent needs to measure...' clause, matching the anchor for clearly answering both; it is above level 2 because the when-trigger is explicit, not implied.

3 / 3

Trigger Term Quality

It enumerates natural trigger terms a user would say — 'measure', 'temperature', 'weight', 'pH', 'quantitative property' — giving good coverage rather than just one keyword (level 2) or jargon (level 1).

3 / 3

Distinctiveness Conflict Risk

It targets a specific niche — measuring via the distinctive 'use [TOOL] on [TARGET]' syntax with focus actions in ScienceWorld — making it unlikely to fire for unrelated skills, matching the clear-niche anchor.

3 / 3

Total

12

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
zjunlp/SkillNet
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.