CtrlK
BlogDocsLog inGet started
Tessl Logo

daily-guidance

强制使用:凡是时间尺度在小时及以下的日常行为模拟,必须使用本 Skill 生成、评估、执行和修正每日 Story。

57

Quality

66%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./packages/agentsociety2/agentsociety2/agent/skills/daily-guidance/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A dense, well-structured operational skill: real, verified CLI commands with a clear two-branch per-step workflow and built-in validation feedback loops. The main structural weakness is progressive disclosure — the bundled references/examples.md and references/story_schema.yaml are never mentioned in the body while their content is partially duplicated inline — plus a few over-explained passages and undocumented commands.

Suggestions

Add a short 'References' section linking the existing bundle files (e.g., full-day story example: references/examples.md; complete field schema: references/story_schema.yaml) and trim the inlined field-by-field maslow/tpb duplication accordingly.

Document the `init` subcommand (present in the script) in the CLI table, and either show or drop the `execute_skill_script` invocation convention so every command is copy-paste runnable.

Cut the Theory of Planned Behavior explanation ('计划行为理论(TPB)认为行为意图由三类因素形成') to just the field definitions, and make the validation-recovery loop explicit (on failed plan validation: read per-field hints → fix JSON → resubmit → re-check).

DimensionReasoningScore

Conciseness

The body is efficient: tables for CLI commands, location_policy and maslow values, a compact JSON example, and no padding — nearly every token carries operational information. It falls short of 5 mainly because of minor over-explanation Claude does not need, e.g., '计划行为理论(TPB)认为行为意图由三类因素形成' restates the familiar Theory of Planned Behavior, and the hook output example block restates fields already specified elsewhere.

4 / 5

Actionability

Guidance is mostly executable: a CLI command table (plan/current/show/record/deviate/revise/check — all verified to exist in scripts/daily_guidance.py), a complete plan JSON example with exact field names, enumerated location_policy and deviation-type values, and concrete placeholder commands like `python scripts/daily_guidance.py plan --date YYYY-MM-DD --json '<JSON对象>'`. Minor gaps keep it from 5: the `init` subcommand present in the script is undocumented, `execute_skill_script` is referenced but never shown how to invoke, and the plan example shows only one segment rather than a full 00:00–24:00 day.

4 / 5

Workflow Clarity

The per-step flow is clearly sequenced with two explicit branches (no valid story → submit plan; valid story → act on active_segment), and validation is built in: '校验通过才写文件,否则返回逐字段修复提示' plus a `check` command and self_check fields form a validate-and-fix loop. It does not reach 5 because the recovery loop after a failed plan validation is described only as '修复提示' without an explicit resubmit/re-validate sequence, and the deviation decision tree lacks an explicit checkpoint for when to revise vs. record-only.

4 / 5

Progressive Disclosure

The body itself is well-sectioned with headers, but judged against the actual bundle: references/examples.md and references/story_schema.yaml exist and are never referenced anywhere in the body, so their content (full-day story examples, the complete field schema) is undiscoverable. Meanwhile schema detail (required segment fields, maslow/tpb structures) is inlined, duplicating what story_schema.yaml holds — matching anchor 3 ('references present but not clearly signaled; content that should be separate is inline'). It is above 2 because the body does have clear structure and the script reference is real, but below 4 because the two reference files are unlinked.

3 / 5

Total

15

/

20

Passed

Description

62%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description cleanly answers both what the skill does and when it must be used, with a clearly-scoped niche and low conflict risk. Its weaknesses are thin trigger-term coverage (no synonyms or variations) and actions stated at a pipeline level rather than as concrete capabilities.

Suggestions

Add natural trigger variations and synonyms to the 'when' clause (e.g., daily schedule, routine, itinerary, per-step activity planning) so the skill triggers on phrasings users actually say.

Make the 'what' more concrete by naming the artifact and operations briefly (e.g., 'generates, validates, executes, and revises a per-day story.yaml of time-segmented activities for sub-hour behavioral simulation').

Briefly define 'Story' in the description (a day-long sequence of time-segmented activities) so the capability is unambiguous without reading the body.

DimensionReasoningScore

Specificity

The description names its domain ("日常行为模拟", daily behavior simulation) and a pipeline of actions ("生成、评估、执行和修正每日 Story" — generate, evaluate, execute, correct daily Stories), but the artifact "Story" is undefined and the verbs are pipeline stages rather than concrete capabilities, matching the 'names domain and concrete actions but not comprehensive' anchor. It does not reach 4 because the actions lack the concrete detail of the 4-anchor example, and exceeds 2 because it lists multiple named actions rather than generic ones.

3 / 5

Completeness

Both parts are present: the 'what' is "生成、评估、执行和修正每日 Story" and the 'when' is explicit via the mandatory trigger clause "强制使用:凡是时间尺度在小时及以下的日常行为模拟". This matches anchor 4 ('has both what and when; when could be more explicit'); it does not reach 5 because the 'what' is terse and the trigger is a single conditional without multiple concrete trigger phrases.

4 / 5

Trigger Term Quality

The trigger condition "凡是时间尺度在小时及以下的日常行为模拟" (any daily behavior simulation at time scales of an hour or below) plus "每日 Story" gives some relevant keywords, but there are no synonyms or variations (e.g., schedule, daily plan, routine, itinerary), fitting the 'some relevant keywords but missing common variations' anchor. It is above 2 because the terms are domain-relevant rather than generic, but below 4 because coverage of natural phrasings is thin.

3 / 5

Distinctiveness Conflict Risk

The niche is narrow — sub-hour daily behavior simulation with clock-derived story files — and the trigger condition is specific enough that unrelated skills would rarely match, fitting 'mostly distinct; minor overlap risk'. It does not reach 5 because 'daily' planning/scheduling language could still overlap with general day-planning skills and the description offers little else to disambiguate.

4 / 5

Total

14

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
tsinghua-fib-lab/AgentSociety
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.