CtrlK
BlogDocsLog inGet started
Tessl Logo

agentsociety-run-experiment

Use when experiment configuration already exists and a simulation run needs to be started, monitored, or stopped.

63

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./extension/skills/agentsociety-run-experiment/v1.0.0/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, information-dense operational skill: executable commands, explicit hard constraints, a visualized workflow with an error-recovery loop, and a well-signaled one-level reference. The main defects are the $PYTHON vs $PYTHON_PATH inconsistency in the Monitoring and Progress Tracking commands and the unreferenced scripts/run.py bundle file.

Suggestions

Fix the interpreter placeholder inconsistency: the Monitoring and Progress Tracking sections use $PYTHON while the rest of the document uses $PYTHON_PATH, so those commands fail if copied literally.

Reference scripts/run.py from SKILL.md (or remove it from the bundle) so the script is discoverable; currently no section points to it.

De-duplicate the command examples: state "append --foreground to the start command" instead of repeating the full command, and consider trimming the overlap between the Quick Reference and CLI Arguments tables.

DimensionReasoningScore

Conciseness

The body is dense with domain-specific facts Claude could not know (CLI arguments, checkpoint files, batch sizing) and explains no general concepts, but has minor redundancy: the Foreground Mode section repeats the full start command just to append one flag, and the Quick Reference table overlaps the CLI Arguments table.

4 / 5

Actionability

Commands are concrete with real IDs and a documented $PYTHON_PATH placeholder, but the Monitoring and Progress Tracking sections use $PYTHON instead of $PYTHON_PATH, so those blocks would fail if copied literally — a genuine minor gap that fits "concrete commands with minor gaps" rather than fully copy-paste ready.

4 / 5

Workflow Clarity

The dot graph gives a clear sequence with an explicit completion check ("run/replay/_schema.json and artifacts present") and an error-recovery feedback loop (fix -> check), and status validation is present so the destructive/batch cap does not apply; however the workflow lives in a terse diagram with checkpoints implicit rather than spelled out as ordered steps with validation callouts, keeping it below the explicit-checklist anchor.

4 / 5

Progressive Disclosure

The single reference is one level deep, clearly signaled under "Programmatic API" (references/programmatic-api.md exists and matches), and the advanced API content is appropriately split out of SKILL.md; however scripts/run.py (13KB) is present in the bundle but never mentioned in the body, making it undiscoverable — a minor organization gap against the actual bundle structure.

4 / 5

Total

16

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A concise, well-targeted description with an explicit trigger clause and three concrete actions, cleanly distinguishing this skill from its pipeline siblings. Its main limitation is that the capability statement (the "what") is embedded in the when-clause rather than declared up front, and a few natural trigger terms (status, logs) are absent.

Suggestions

Lead with a capability declaration before the trigger clause, e.g. "Starts, monitors, and stops AgentSociety simulation runs via the CLI. Use when experiment configuration already exists...", to make the "what" as explicit as the "when".

Add natural trigger synonyms such as "status", "logs", or "check on an experiment" to the when-clause to broaden keyword coverage.

DimensionReasoningScore

Specificity

"a simulation run needs to be started, monitored, or stopped" names the domain and three concrete actions, matching the anchor for several specific actions with minor gaps; it is not score 5 because capabilities like list/resume are absent, and not score 3 because it goes beyond 1-2 actions.

4 / 5

Completeness

The explicit "Use when experiment configuration already exists..." clause answers "when" with concrete trigger conditions, and the what (start/monitor/stop runs) is present but embedded inside the when-clause rather than stated as a leading capability declaration, so it sits between the anchor-4 and anchor-5 examples rather than clearly matching 5.

4 / 5

Trigger Term Quality

Natural phrases like "simulation run", "started, monitored, or stopped", and "experiment configuration" give good keyword coverage a user would plausibly say, but common variations such as "status", "logs", or "check on the experiment" are missing, keeping it below the comprehensive-synonyms anchor.

4 / 5

Distinctiveness Conflict Risk

The gating condition "experiment configuration already exists" carves a clear niche distinct from sibling config/analysis skills, but "monitored" carries minor overlap risk with a results-analysis skill, matching "mostly distinct; minor overlap risk with closely related skills" rather than the minimal-conflict anchor.

4 / 5

Total

16

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
tsinghua-fib-lab/AgentSociety
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.