CtrlK
BlogDocsLog inGet started
Tessl Logo

agentsociety-hypothesis

Use when defining or revising research hypotheses, experiment groups, or comparison structure after literature review.

58

Quality

68%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./extension/skills/agentsociety-hypothesis/v1.0.0/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable skill file: concrete CLI commands, a compact workflow graph, realistic examples, and a useful mistakes table. The main weaknesses are duplicated scan-modules guidance, a small `$PYTHON`/`$PYTHON_PATH` inconsistency, and the absence of explicit verification steps around hypothesis creation and deletion.

Suggestions

Consolidate the scan-modules guidance: state it once (e.g., in Module Selection) and trim the duplicate rows in Common Mistakes and Pipeline Position to short cross-references.

Fix the `$PYTHON` vs `$PYTHON_PATH` inconsistency in the Progress Tracking command and explain placeholders like `N` (e.g., 'replace N with the number of hypotheses added').

Add an explicit verification checkpoint to the workflow (e.g., run `hypothesis list --json` after `hypothesis add` to confirm creation before updating TOPIC.md) and note that `hypothesis delete` is irreversible and should be confirmed with the user first.

DimensionReasoningScore

Conciseness

The body is lean and table-driven with no explanations of concepts Claude already knows, but scan-modules guidance is repeated in 'Module Selection' ('use `scan-modules` to discover or validate them'), twice in 'Common Mistakes' ('Run `scan-modules` to confirm valid class and module names', 'Copy exact names from `ags.py scan-modules list --short` output'), and again in 'Pipeline Position', which could be consolidated.

4 / 5

Actionability

The Quick Reference table gives full executable commands with all flags, the Group JSON Format example is concrete and copy-paste ready, and the module combination table covers common cases. Minor gaps: the Progress Tracking command uses `$PYTHON` while every other command uses `$PYTHON_PATH`, and `--metadata '{"hypotheses_count": N}'` leaves the placeholder `N` unexplained.

4 / 5

Workflow Clarity

The dot digraph gives a clear, ordered sequence (read topic -> clarify -> names known? -> scan -> groups -> write -> sync) with a conditional checkpoint (scan-modules when module names are unknown) and CLI-side validation implied by `--skip-module-validation`. However there is no explicit post-add verification step in the flow, and the destructive `hypothesis delete` command has no confirmation or safety guidance.

4 / 5

Progressive Disclosure

Sections are well-organized and content is appropriately inline for a skill of this size (~112 lines) with no nested references. Minor gaps: the bundle's `scripts/hypothesis.py` is never referenced from the body (commands go through the external `.agentsociety/bin/ags.py`), and setup details defer to an external `CLAUDE.md` rather than a bundled reference.

4 / 5

Total

16

/

20

Passed

Description

61%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A concise, niche description with explicit and well-scoped trigger guidance anchored to the pipeline stage ('after literature review'). Its weakness is that it never states what the skill does — capability is only implied through gerund trigger phrasing — and it omits natural synonyms like 'control vs treatment' that the body already uses.

Suggestions

Lead with an explicit capability statement in third person before the trigger clause, e.g., 'Manages research hypotheses (control/treatment groups, agent and environment module selection) for AgentSociety experiments. Use when defining or revising research hypotheses...'.

Add the natural synonyms already present in the body — 'research question' and 'control vs treatment' — to the trigger clause so users who phrase it either way match the skill.

Mention the AgentSociety/simulation context to reduce overlap with generic experiment-design or statistics skills.

DimensionReasoningScore

Specificity

The description names the domain and objects ('defining or revising research hypotheses, experiment groups, or comparison structure') but frames them as user triggers rather than capabilities — no concrete actions like 'manage', 'list', or 'add' hypotheses are stated, matching 'names domain and 1-2 concrete actions, but not comprehensive'. It is above anchor 2 because the objects of work are specific, but below anchor 4 because no discrete capabilities are listed.

3 / 5

Completeness

The 'when' is explicit and strong ('Use when... after literature review'), but the 'what' is never stated as a capability — it is only implied through the trigger phrasing. This sits above anchor 2 (only 'when', no 'what') but below anchor 4, which requires both explicitly present.

3 / 5

Trigger Term Quality

'research hypotheses', 'experiment groups', 'comparison structure', and 'literature review' are natural phrases a user would say. Common variations the body itself lists — 'control vs treatment', 'research question' — are missing from the description, matching 'good keyword coverage; a few natural terms missing'.

4 / 5

Distinctiveness Conflict Risk

The trigger terms ('research hypotheses', 'experiment groups', 'comparison structure after literature review') carve out a clear niche in the research pipeline with minimal conflict risk. It is not a 5 because 'hypothesis' alone is generic — the description omits the AgentSociety/simulation context that would fully disambiguate it from generic statistics or experiment-design skills.

4 / 5

Total

14

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
tsinghua-fib-lab/AgentSociety
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.