CtrlK
BlogDocsLog inGet started
Tessl Logo

scienceworld-result-archiver

Places an object into a designated container (like a colored box) based on a test outcome. It should be triggered when the agent needs to finalize a task by storing an object according to a rule (e.g., conductive → blue box, non-conductive → orange box). The input is the object and the rule, and the output is the object moved to the correct container.

65

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A concise, highly actionable skill body with a clear sequenced procedure and explicit verification steps for a simple, low-risk task. Its one real defect is the broken bundle reference: the 'Bundled Logic' section points to scripts/conductivity_test.py, which does not exist in the skill.

Suggestions

Fix the dangling reference: either add the referenced scripts/conductivity_test.py to the skill bundle or remove/rewrite the 'Bundled Logic' section so no path is cited that the agent cannot actually open.

Trim the minor redundancy: the opening line ("Use this skill to finalize a scientific test by archiving an object...") duplicates the frontmatter description, and 'Container Verification' in Key Principles restates checks already in the Core Procedure.

Add one more worked example for a different test type (e.g., a chemical-reaction or physical-property rule) so the common cases are covered beyond conductivity alone.

DimensionReasoningScore

Conciseness

The body is lean and well-sectioned ("When to Use", "Core Procedure", "Key Principles", worked example) with no explanations of concepts Claude already knows. Minor redundancy — the opening line restates the frontmatter description and "Container Verification" overlaps step 4's context check — keeps it at anchor 4 rather than the every-token-earns-its-place anchor 5.

4 / 5

Actionability

Guidance is fully executable for an instruction-only skill: the exact action template "move OBJ to CONTAINER", the "look around" verification command, and a complete worked example ending in the copy-paste-ready "move metal pot to blue box" covering the common conductivity case. This matches anchor 5's concrete commands plus common-case examples.

5 / 5

Workflow Clarity

The "Core Procedure" gives a clear 4-step sequence (Verify Context → Confirm Test Result → Apply Rule → Execute Archive) with two explicit validation checkpoints: "Observe the final state of your experimental apparatus to definitively determine the test outcome" and "visually confirm the target container exists in the room (use `look around` if uncertain)". This is a single low-risk, non-destructive operation, so the missing feedback-loop requirement does not cap it; it matches anchor 5 via the simple-skill exception.

5 / 5

Progressive Disclosure

The body is under 50 lines and well-organized, but it directs the reader to "the bundled script `scripts/conductivity_test.py`" and no scripts/ directory (or any bundle file) exists in the skill — the single external reference is dangling. Structure is present, yet the one-level-deep navigation it claims does not actually resolve, which fits anchor 3 (references present but not working/organized properly) rather than anchor 4's 'minor organization gaps'.

3 / 5

Total

17

/

20

Passed

Description

73%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A solid description that clearly states what the skill does, its inputs/outputs, and when it should trigger, with a concrete domain example that makes it highly distinctive. It could be stronger by enumerating a couple more concrete actions and by using more natural trigger phrasings rather than the abstract "finalize a task" framing.

Suggestions

Replace the abstract trigger framing "when the agent needs to finalize a task" with concrete trigger phrases the user/agent would actually say, e.g., "Use when the agent has observed a test result (e.g., a conductivity test) and must file or sort the tested object into the matching container."

Enumerate one or two more concrete actions beyond "Places an object into a designated container" (e.g., confirming the test outcome and verifying the target container exists) to lift specificity toward comprehensive coverage.

Add natural synonyms such as "sort", "file", or "archive" alongside "storing" so the description matches how the task is more likely to be phrased.

DimensionReasoningScore

Specificity

"Places an object into a designated container (like a colored box) based on a test outcome" plus "The input is the object and the rule, and the output is the object moved to the correct container" names the domain and one concrete action with I/O, matching the 1-2-concrete-actions anchor. It does not list several specific actions, so it falls short of anchor 4; it is well above the generic anchor 2 because the action and input/output are concrete.

3 / 5

Completeness

Both parts are present: a clear 'what' ("Places an object into a designated container... based on a test outcome" with input and output) and an explicit 'when' ("It should be triggered when the agent needs to finalize a task by storing an object according to a rule"). The 'when' clause is explicit but somewhat abstract — it describes the condition rather than concrete trigger phrases a user would say — so it matches anchor 4 rather than anchor 5.

4 / 5

Trigger Term Quality

Terms like "conductive → blue box, non-conductive → orange box", "finalize a task", "storing an object", and "test outcome" are natural to the domain and would appear in the triggering context. A few common phrasings (e.g., "sort", "archive", "property test") are missing, which keeps it below the comprehensive-synonyms anchor 5 but clearly above anchor 3's partial coverage.

4 / 5

Distinctiveness Conflict Risk

The description carves out a clear niche — archiving tested objects into color-coded containers by a conductivity-style rule ("conductive → blue box, non-conductive → orange box") — that no general-purpose skill would plausibly claim. Triggers are distinct and specific, matching the minimal-conflict anchor 5.

5 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

15

/

16

Passed

Repository
zjunlp/SkillNet
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.