CtrlK
BlogDocsLog inGet started
Tessl Logo

planning-with-files

Persistent file-based planning for multi-step AI-agent work. Keeps task_plan.md, findings.md, and progress.md on disk; lifecycle hooks inject selected project planning context. Automatic recovery reads project planning files only. Explicit session-catchup.py --metadata reads same-project local agent session records and emits aggregate counts only; --replay may emit bounded nonce-framed excerpts. Optional gated mode can request continuation only when the host supports it and never runs commands declared in Markdown. The skill has no network upload path. Use for research or work needing 5+ tool calls.

56

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.factory/skills/planning-with-files/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, largely actionable instruction skill with a clear restore-then-plan workflow and a real error-escalation feedback loop. Its weaknesses are repeated error-handling guidance across four sections, unwired validation steps, and bundle navigation that breaks: most referenced files (templates, references.md, examples.md) are absent from the bundle.

Suggestions

Consolidate the four overlapping error-handling sections (Critical Rules 5-6, 3-Strike protocol, Anti-Patterns rows) into one, and drop the trivial 'if action_failed' pseudocode block.

Wire validation into the workflow explicitly, e.g. add 'run scripts/check-complete.sh after each phase; only mark complete when it passes' as a numbered Quick Start step, and give a concrete resolve-plan-dir.sh example invocation.

Ship the referenced files (templates/task_plan.md, findings.md, progress.md, references.md, examples.md) or remove/inline the links so navigation in the bundle does not dead-end.

DimensionReasoningScore

Conciseness

The body is mostly efficient — tables for file purposes, the decision matrix, and anti-patterns are token-dense — but error handling is explained four separate times (Critical Rules 5 and 6, the 3-Strike Error Protocol, the Read/Write matrix rows, and two Anti-Patterns rows), and the 'Never Repeat Failures' section reduces to trivial pseudocode ('if action_failed: next_action != same_action') that Claude needs no help deriving. This matches 'Mostly efficient but includes some unnecessary explanation or could be tightened', not 4 — the repetition is more than minor trimming.

3 / 5

Actionability

Concrete, runnable guidance dominates: an exact session-catchup.py invocation with $(command -v python3 ...) substitution, 'run scripts/init-session.sh "Task Name"', 'sh "<skill-dir>/scripts/set-active-plan.sh" --list' with the PowerShell equivalent, and named scripts (resolve-plan-dir.sh, check-complete.sh) that all exist in scripts/. Minor gaps keep it from 5: the resolve-plan-dir.sh usage is described ('with the host's PLAN_ID and PWF_PLAN_ROOT') without a copy-paste example, and check-complete.sh is listed but never wired into the workflow with an example invocation.

4 / 5

Workflow Clarity

A clear ordered sequence exists — restore state (resolve plan dir, read the three files, git diff --stat), then Quick Start steps 1-4, with explicit correction handling ('If an explicit selector is rejected... correct the pin and do not fall back') and the 3-Strike protocol providing a genuine diagnose → alternative → rethink → escalate feedback loop. It is not 5 because validation is described but not checkpointed into the main flow: check-complete.sh is only listed under Scripts rather than placed as an explicit verify step after phases, and the checkpoint placement is implicit rather than sequenced.

4 / 5

Progressive Disclosure

The body is well-sectioned and clearly signals one-level-deep references (Templates, Scripts, Advanced Topics pointing to references.md/examples.md), and the scripts/ paths it cites all exist in the bundle. However, scoring against the actual bundle, the majority of outbound references are dangling: templates/ and its three template files, references.md, and examples.md do not exist in the bundle, so following the navigation fails. This sits between 'references present but not clearly signaled' (3) and broken navigation (below), landing at 3 with the missing files as the dominant defect.

3 / 5

Total

14

/

20

Passed

Description

67%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A dense, concrete, third-person description that clearly states what the skill does and gives an explicit use-when clause. Its main weaknesses are a thin and partly unnatural trigger clause ('5+ tool calls'), missing common trigger synonyms, and a sizable share of the budget spent on security caveats instead of capability coverage.

Suggestions

Broaden the trigger clause to natural user phrasing, e.g. 'Use for long multi-step tasks, research projects, or any work spanning many tool calls where state must survive context loss.'

Trim or relocate the security-boundary sentences ('no network upload path', gated-mode caveats) so the description's token budget covers capabilities and triggers.

Add common synonyms such as 'project planning', 'task tracking', and 'working memory' so users' natural phrasing matches the description.

DimensionReasoningScore

Specificity

The description names concrete artifacts ("Keeps task_plan.md, findings.md, and progress.md on disk"), concrete mechanisms ("lifecycle hooks inject selected project planning context"), and concrete tool behaviors ("Explicit session-catchup.py --metadata reads same-project local agent session records and emits aggregate counts only"). It falls short of the 5 anchor because a sizable share of the budget goes to security non-capability caveats ("The skill has no network upload path") rather than comprehensive coverage of what the skill does, and it is above the 3 anchor because it lists several specific actions, not just 1-2.

4 / 5

Completeness

Both parts are explicitly present: a clear 'what' ("Persistent file-based planning for multi-step AI-agent work. Keeps task_plan.md, findings.md, and progress.md on disk") and a 'when' ("Use for research or work needing 5+ tool calls"). It is not a 5 because the when-clause is narrow and partially un-natural as a user trigger phrase, matching the anchor 'Has both what and when; when could be more explicit or specific'; it is clearly not a 3 since the when is explicit, not merely implied.

4 / 5

Trigger Term Quality

Natural trigger terms present are "planning", "multi-step", and "research", but the only explicit trigger clause is "Use for research or work needing 5+ tool calls" — users do not naturally count their tool calls, and common variations like "long task", "project planning", "task tracking", or "stay organized" are missing. This matches the anchor 'Some relevant keywords but missing common variations or synonyms', not 4 (which requires good coverage with only a few natural terms missing) nor 2 (which requires near-total absence of natural phrases).

3 / 5

Distinctiveness Conflict Risk

The niche is distinct — persistent on-disk planning with named files (task_plan.md, findings.md, progress.md) and a named script (session-catchup.py) is unlikely to collide with unrelated skills. Minor overlap risk remains with generic todo/planning/task-management skills around the broad term "planning", so it fits 'Mostly distinct; minor overlap risk with closely related skills' rather than the 5 anchor's 'minimal conflict risk'.

4 / 5

Total

15

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 6 missing

Warning

Total

14

/

16

Passed

Repository
OthmanAdi/planning-with-files
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.