CtrlK
BlogDocsLog inGet started
Tessl Logo

planning-with-files

Persistent file-based planning for multi-step AI-agent work. Keeps task_plan.md, findings.md, and progress.md on disk; lifecycle hooks inject selected project planning context. Automatic recovery reads project planning files only. Explicit session-catchup.py --metadata reads same-project local agent session records and emits aggregate counts only; --replay may emit bounded nonce-framed excerpts. Optional gated mode can request continuation only when the host supports it and never runs commands declared in Markdown. The skill has no network upload path. Use for research or work needing 5+ tool calls.

52

Quality

66%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/planning-with-files/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

52%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers genuinely actionable and well-validated planning guidance with concrete scripts, but it is severely bloated by interleaved version history, design rationale, and security essays, and its progressive-disclosure structure is broken: many referenced files (templates, docs, commands) are absent from the bundle. The core v2 workflow is sound; the v3/host/security material should live in separate reference files that actually exist.

Suggestions

Move the version-numbered v3 mode specifications, host-capability tier table, attestation/nonce hardening essays, and issue/PR changelog lessons into a separate docs/ file referenced once from SKILL.md, cutting the body to the core workflow — this addresses the conciseness score of 2.

Ship or fix the referenced bundle files: templates/task_plan.md, findings.md, progress.md and loop.md are required by numbered Quick Start steps, and reference.md, examples.md, docs/attestation-locking.md, docs/perf-notes.md, and commands/*.md are linked but missing.

Restructure so the core loop (restore state → init → create files → update cadence → validation) reads as one continuous sequence, with the Claude Code turn-loop integration, parallel/shared-parent workflows, and gated-mode details as clearly-labeled optional sections.

DimensionReasoningScore

Conciseness

The ~470-line body is noticeably verbose for a planning methodology whose core (Quick Start, Critical Rules, file purposes) fits in ~100 lines: version-numbered annotations on nearly every section ('v2.36.1', 'v2.37.0', 'v3.6.0', 'v3.8.0', 'v3.9.0', 'v3.10.0', 'v2.42.0'), internal changelog references ('the lesson from issue #178', 'PR #180 lesson', 'issue #237'), a host-capability tier table, security-hardening essays, and design rationale ('the evidence (arxiv 2603.03258...) shows drift is real') that are project history, not instructions. It is not anchor 1 because it does not explain general concepts Claude already knows — the padding is project-specific rationale — but several sections are unnecessary in a SKILL.md.

2 / 5

Actionability

Guidance is mostly executable: concrete script invocations ('scripts/init-session.sh "Task Name"', the bash and PowerShell session-catchup.py blocks, 'sh scripts/init-session.sh --gated "Build Pipeline"'), a two-terminal PLAN_ID parallel-workflow example, /plan-goal and /plan-loop usage with argument examples, and step-numbered manual-fallback procedures. Minor gaps keep it below anchor 5: some core steps are abstract ('Work like Manus', the Read vs Write decision matrix) and a few referenced entry points are described rather than shown.

4 / 5

Workflow Clarity

The sequence is clear and checkpointed: restore project state first (resolve-plan-dir.sh, git diff --stat), then Quick Start (init → create missing files only → re-read before decisions → single plan owner), with explicit validation tools (check-complete.sh, plan-doctor.sh, the 5-Question Reboot Test) and a feedback loop in the 3-Strike Error Protocol. It misses anchor 5 because the core loop is fragmented across version-history sections, forcing the reader to assemble the workflow from v2/v3 interleaved narrative.

4 / 5

Progressive Disclosure

The body inlines hundreds of lines that clearly belong in separate files (the entire 'Autonomous and Gated Modes' spec, host capability tiers, security hardening, attestation internals) while its outbound references are broken in this bundle: templates/task_plan.md, templates/findings.md, templates/progress.md, templates/loop.md, reference.md, examples.md, docs/attestation-locking.md, docs/perf-notes.md, and commands/plan-goal.md / plan-loop.md are all referenced but not shipped — only the scripts/ directory exists, and the missing templates are required by a numbered Quick Start step. This matches anchor 2 (content that clearly belongs in separate files is inlined) rather than anchor 3, because navigation to the referenced detail is not actually available.

2 / 5

Total

12

/

20

Passed

Description

67%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description answers both what the skill does and when to use it, with unusually concrete artifact names and an explicit trigger clause. Its weaknesses are trigger-term coverage weighted toward implementation jargon over natural user phrasing, and space spent on security caveats instead of capability breadth.

Suggestions

Add natural trigger synonyms users would actually say — e.g. 'Use for project planning, tracking progress across sessions, multi-step research, or recovering context after compaction' — to broaden trigger_term_quality beyond the single 'research or work needing 5+ tool calls' condition.

Replace implementation jargon in the trigger-relevant portion ('lifecycle hooks', 'nonce-framed excerpts', 'session-catchup.py --metadata/--replay') with capability-level phrasing, moving the trust-boundary and mode details into the body where they are already documented.

Reconcile the description's '5+ tool calls' threshold with the body's 'Multi-step tasks (3+ steps)' guidance so the trigger does not conflict with the skill's own usage criteria.

DimensionReasoningScore

Specificity

The description names concrete artifacts ('task_plan.md, findings.md, and progress.md on disk') and several specific actions (hooks inject planning context, '--metadata reads same-project local agent session records and emits aggregate counts only', gated mode requests continuation). It stops short of anchor 5 because substantial space is spent on trust-boundary caveats ('no network upload path', 'never runs commands declared in Markdown') rather than comprehensively covering the planning capabilities themselves.

4 / 5

Completeness

Both parts are present: the 'what' is explicit ('Persistent file-based planning for multi-step AI-agent work. Keeps task_plan.md, findings.md, and progress.md on disk; lifecycle hooks inject selected project planning context') and the 'when' is explicit ('Use for research or work needing 5+ tool calls'). It sits below anchor 5 because the trigger clause is a single brief condition, thinner than the anchor's concrete multi-phrase trigger list, though more specific than anchor 4's vague 'when'.

4 / 5

Trigger Term Quality

Some natural user-facing terms are present ('planning', 'multi-step', 'research', '5+ tool calls'), but common variations users would actually say are missing (todo/task tracking, project planning, progress tracking, context loss/recovery). Much of the keyword surface is implementation jargon ('lifecycle hooks', 'nonce-framed excerpts', 'session-catchup.py') rather than natural phrases, matching anchor 3 (some relevant keywords, missing common variations) and clearly below anchor 4's good natural-term coverage.

3 / 5

Distinctiveness Conflict Risk

The named planning-file artifacts (task_plan.md/findings.md/progress.md) and the file-based working-memory framing give it a clear niche with mostly distinct triggers. Minor overlap risk remains with generic todo/planning/task-tracking skills, and the 'research or work needing 5+ tool calls' trigger is broad enough that many long tasks would match, keeping it below anchor 5.

4 / 5

Total

15

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (507 lines); consider splitting into references/ and linking

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 8 missing

Warning

Total

13

/

16

Passed

Repository
OthmanAdi/planning-with-files
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.