CtrlK
BlogDocsLog inGet started
Tessl Logo

scienceworld-process-pauser

This skill introduces deliberate pauses in task execution. Use when the agent needs to consider next steps, evaluate intermediate results, or wait for processes to complete. The skill uses the 'wait1' or 'wait' actions to temporarily halt activity, preventing rushed decisions in complex experimental procedures.

59

Quality

74%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./experiments/src/skills/scienceworld/scienceworld-process-pauser/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A concise, actionable, well-sequenced body for a simple skill: the two wait actions are defined precisely and the implementation logic is easy to follow. The main defect is progressive disclosure — the valuable references/pause_scenarios.md exists in the bundle but is completely unsignaled, so the richest guidance (trigger scenarios, when NOT to pause) is effectively lost.

Suggestions

Add a clearly signaled, one-level-deep link to the existing reference, e.g., '## Pause scenarios — see [pause_scenarios.md](references/pause_scenarios.md) for trigger conditions and when NOT to pause.'

Enrich the 'Identify the Pause Trigger' step with the concrete trigger conditions (after creating an intermediate product, before an irreversible step, after a state-change observation) or point to the reference for them.

Trim redundancy: drop the filler in the Resume step ('with the benefit of the reflective pause') and avoid repeating the frontmatter description verbatim in 'When to Use'.

DimensionReasoningScore

Conciseness

The body is lean and assumes competence ('wait1: Pauses execution for a single simulation step. Use for brief reflection.'), with only minor trimmable filler such as 'Continue the task with the benefit of the reflective pause' and a 'When to Use' section that largely duplicates the frontmatter description. Below 5 due to that redundancy; above 3 because it explains nothing Claude already knows.

4 / 5

Actionability

Gives concrete commands with explicit semantics ('wait: Pauses execution for 10 simulation steps') and a clear selection rule ('Choose wait1 for quick checks or wait for extended evaluation'), plus a real trajectory example. Not 5 because the common pause-trigger cases are covered by only one example; not 3 because the guidance is fully concrete, not pseudocode.

4 / 5

Workflow Clarity

The 'Implementation Logic' section lays out a clear 4-step sequence (Identify the Pause Trigger, Select Duration, Execute Pause, Resume) for a simple, non-destructive single-action skill. Not 5 because the critical step — recognizing a moment requiring deliberation — is abstract and under-specified in the body; not 3 because the sequence and checkpoints are adequate and unambiguous.

4 / 5

Progressive Disclosure

The body's sections are well organized, but the bundle's references/pause_scenarios.md — which contains the concrete trigger scenarios and anti-patterns this skill depends on — is never mentioned or linked, making it undiscoverable. Per the bundle-structure guideline this sits at 'structure present but references not clearly signaled'; not 2 because the inline content is not a wall of text and is appropriately sized.

3 / 5

Total

15

/

20

Passed

Description

63%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A solid description that explicitly states both what the skill does (pause via wait1/wait actions) and when to use it, with a concrete mechanism. Its weaknesses are abstract trigger phrasing, missing natural synonyms, and no domain anchoring, which raise conflict risk with generic planning skills and leave trigger coverage incomplete.

Suggestions

Anchor the triggers in the target domain (e.g., 'Use in ScienceWorld or simulated experiment tasks when...') to reduce overlap with generic planning/reflection skills.

Add natural trigger synonyms a user or agent would actually say, such as 'slow down', 'think before acting', or 'let a simulated process finish'.

Enumerate the concrete pause conditions (after creating an intermediate product, before an irreversible step) to lift specificity and completeness.

DimensionReasoningScore

Specificity

Names several concrete capabilities ('introduces deliberate pauses in task execution', 'uses the wait1 or wait actions to temporarily halt activity, preventing rushed decisions') with an explicit mechanism. It falls short of 5 because coverage is not comprehensive — the distinct pause scenarios are not enumerated — and above 3 because it lists more than 1-2 concrete actions.

4 / 5

Completeness

Both 'what' ('uses the wait1 or wait actions to temporarily halt activity') and 'when' ('Use when the agent needs to consider next steps, evaluate intermediate results, or wait for processes to complete') are explicitly present. Not 5 because the triggers are abstract agent-states rather than concrete, user-mentionable phrases and lack domain anchoring; not 3 because the 'when' clause is explicit, not weakly implied.

4 / 5

Trigger Term Quality

Phrases like 'consider next steps', 'evaluate intermediate results', and 'wait for processes to complete' are relevant but abstract; common natural variations ('slow down', 'think before acting', 'let a process finish') and any domain-specific terms are missing. Better than 2 (more than a couple of generic keywords) but below 4 (keyword coverage has real gaps).

3 / 5

Distinctiveness Conflict Risk

The named 'wait1'/'wait' actions and 'complex experimental procedures' give some distinction, but generic deliberation phrases like 'consider next steps' overlap with broad planning/reflection skills. More than minor overlap risk (below 4) yet somewhat specific (above 2).

3 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
zjunlp/SkillNet
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.