CtrlK
BlogDocsLog inGet started
Tessl Logo

agentsociety-analysis

Use when an experiment run has completed and the user wants rigorous interpretation, claim-driven charts, bilingual reports, or cross-hypothesis synthesis from simulation data. Also use when multiple charts or PNG/JPG assets must be assembled into one labeled composite figure. Requires high-quality narrative and evidence traceability, not only harness gate PASS.

66

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-engineered orchestration skill body: an explicit gated six-stage workflow with validation and feedback loops, and highly actionable CLI command tables. Its main defects are dangling bundle references (stages/, checklists/, subagent-prompts/, support/ are missing) and duplicated sections that cost tokens without adding guidance.

Suggestions

Create or remove the dangling references: the six `stages/*.md` files, `checklists/quality.md`, `subagent-prompts/data-explorer.md`, and `support/frontend-design/` are cited as core navigation targets but do not exist in the bundle, breaking the Stage Notes and Subagent Delegation sections.

Deduplicate the Shared References list (`references/harness-contract.md` and `references/reports.md` are each listed twice) and drop the redundant '## CLI Tool' section, which repeats a subset of the Quick Reference table.

Inline one concrete `--payload` JSON example (e.g., a minimal `write-plan` or `record-attestation` payload) so the most common mutating commands are executable without opening `references/json-payloads.md`.

DimensionReasoningScore

Conciseness

The body is dense and imperative — command tables, mistake/fix pairs, and hard constraints assume Claude's competence with no basic-concept padding, matching the efficient-with-minor-trimming anchor. Not 5 because the ~20-row Quick Reference table duplicates the later CLI Tool section, and the Shared References list repeats entries (`references/harness-contract.md` and `references/reports.md` each appear twice).

4 / 5

Actionability

Concrete, fully spelled-out commands (`$PYTHON_PATH .agentsociety/bin/ags.py analysis load-context --workspace . --hypothesis-id ID ...`), exact output paths, exit-code semantics, and receipt-inspection instructions are mostly copy-paste ready, matching the 4 anchor. Not 5 because payload examples are only placeholders (`--payload '{...}'`, `--spec FILE`) deferred entirely to a reference file, so common cases are not covered inline.

4 / 5

Workflow Clarity

A clearly sequenced six-stage pipeline (dot graph plus Stage Notes) with explicit validation gates per phase, the four-step post-phase loop, exit-code error handling, single-flight/retry semantics, and an explicit feedback loop ("On REVISE/FAIL, loop producer → reviewer until PASS") — matching the 5 anchor's sequence, validation, and error-recovery checkpoints.

5 / 5

Progressive Disclosure

SKILL.md is structured as an overview with well-labeled, one-level-deep references, but scored against the actual bundle, 10 referenced paths dangle: all six `stages/*.md` files, `checklists/quality.md`, `subagent-prompts/data-explorer.md`, and `support/frontend-design/**` do not exist, so the central Stage Notes and delegation navigation is broken. This is more than the 4 anchor's 'minor organization gaps', while the 3 anchor's 'could be better organized' fits; not 2 because the references that do exist are clearly signaled and one level deep.

3 / 5

Total

16

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly states concrete deliverables and two explicit, concrete trigger conditions in third person. Keyword coverage is good, though it omits the most natural user phrasings ("analyze/visualize results") and leans on some ecosystem jargon.

DimensionReasoningScore

Specificity

Lists several concrete deliverables — "rigorous interpretation", "claim-driven charts", "bilingual reports", "cross-hypothesis synthesis", "assembled into one labeled composite figure" — matching the several-specific-actions anchor. Not 5 because "rigorous interpretation" is somewhat abstract and data-exploration/querying capabilities from the body are not named; not 3 because four-plus distinct concrete actions are enumerated.

4 / 5

Completeness

Explicitly answers both: what ("rigorous interpretation, claim-driven charts, bilingual reports, or cross-hypothesis synthesis from simulation data", "assembled into one labeled composite figure") and when ("Use when an experiment run has completed and the user wants...", "Also use when multiple charts or PNG/JPG assets must be assembled"), with concrete trigger conditions in third-person voice. Matches the 5 anchor's structure of clear what plus explicit when; the neighboring 4 anchor's weaker/less-explicit 'when' does not fit.

5 / 5

Trigger Term Quality

Natural phrases like "experiment run has completed", "charts", "bilingual reports", "composite figure", "PNG/JPG assets", and "simulation data" give good keyword coverage, matching the good-coverage-with-a-few-missing anchor. Not 5 because common user phrasings such as "analyze results", "visualize data", or "explore data" (which the body's When-to-Use does include) are absent, and terms like "cross-hypothesis synthesis" are jargon users would rarely say.

4 / 5

Distinctiveness Conflict Risk

Niche markers like "experiment run has completed", "harness gate PASS", and "cross-hypothesis synthesis" make it mostly distinct from other skills. Not 5 because the platform is never named and generic terms ("charts", "reports", "simulation data") create minor overlap risk with general data-analysis or charting skills; not 3 since overlap is limited rather than broad.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
tsinghua-fib-lab/AgentSociety
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.