CtrlK
BlogDocsLog inGet started
Tessl Logo

agentsociety-analysis

Use when an experiment run has completed and the user wants rigorous interpretation, claim-driven charts, bilingual reports, or cross-hypothesis synthesis from simulation data. Also use when multiple charts or PNG/JPG assets must be assembled into one labeled composite figure. Requires high-quality narrative and evidence traceability, not only harness gate PASS.

67

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

77%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-gated multi-stage workflow with strong sequencing and validation, weakened by internal redundancy and a set of broken/over-long reference pointers that hurt navigation.

Suggestions

Remove the standalone "CLI Tool" subsection and consolidate duplicate command/guidance content from "Common Mistakes" and "Hard Constraints" to cut repetition.

Fix or drop references to bundle paths that do not exist (stages/*.md, checklists/quality.md, subagent-prompts/data-explorer.md, support/frontend-design/..., CLAUDE.md), or add those files to the bundle.

Group the 23-entry "Shared References" list into categories (e.g. plotting, reporting, harness, memory) so the navigation surface is scannable.

DimensionReasoningScore

Conciseness

The body is information-dense and assumes Claude's competence, but it carries real redundancy — the "CLI Tool" subsection restates commands already in the Quick Reference table, "Common Mistakes" repeats "Hard Constraints", and the 23-entry "Shared References" list could be consolidated.

2 / 3

Actionability

Provides fully executable commands with complete argument signatures, exact output paths, exit-code semantics (2 = fixable block, 1 = execution failure), and concrete directory layouts — copy-paste ready.

3 / 3

Workflow Clarity

Clear 6-stage sequence with explicit validation gates (validate-*), feedback loops (re-validate on failure, producer→reviewer loop until PASS), and batch-operation checkpoints, matching the clear-sequence-with-validation anchor.

3 / 3

Progressive Disclosure

Structure is overview-plus-one-level references and the existing reference files are real, but several inline references point to paths absent from the bundle (stages/01–06, checklists/quality.md, subagent-prompts/data-explorer.md, support/frontend-design/..., CLAUDE.md), and the reference list is long enough to warrant grouping.

2 / 3

Total

10

/

12

Passed

Description

85%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description with explicit "Use when" triggers and a distinct niche; its main weakness is jargon-heavy phrasing that dilutes the natural trigger terms a user would actually say.

Suggestions

Reword jargon-heavy phrases like "cross-hypothesis synthesis", "evidence traceability", and "harness gate PASS" into terms users naturally say (e.g. "combine results across experiments", "trace every number to the data").

Add common user phrasings such as "analyze results", "explore the data", or "visualize the simulation" alongside the existing triggers to broaden natural-term coverage.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — "rigorous interpretation, claim-driven charts, bilingual reports, or cross-hypothesis synthesis" and assembling "multiple charts or PNG/JPG assets ... into one labeled composite figure" — matching the multiple-specific-actions anchor.

3 / 3

Completeness

Explicitly answers both what (interpretation, charts, bilingual reports, synthesis, composite figures) and when via two "Use when ..." clauses, satisfying the explicit-trigger anchor.

3 / 3

Trigger Term Quality

Has relevant natural terms ("charts", "reports", "composite figure", "simulation data") but is mixed with jargon a user would rarely say ("cross-hypothesis synthesis", "evidence traceability", "harness gate PASS"), so coverage of common variations is incomplete.

2 / 3

Distinctiveness Conflict Risk

Occupies a clear niche (AgentSociety simulation-run analysis with composite-figure assembly) with triggers unlikely to fire for unrelated skills.

3 / 3

Total

11

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
tsinghua-fib-lab/AgentSociety
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.