CtrlK
BlogDocsLog inGet started
Tessl Logo

planning-with-files

Persistent file-based planning for multi-step AI-agent work. Keeps task_plan.md, findings.md, and progress.md on disk; lifecycle hooks inject selected project planning context. Automatic recovery reads project planning files only. Explicit session-catchup.py --metadata reads same-project local agent session records and emits aggregate counts only; --replay may emit bounded nonce-framed excerpts. Optional gated mode can request continuation only when the host supports it and never runs commands declared in Markdown. The skill has no network upload path. Use for research or work needing 5+ tool calls.

60

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.codex/skills/planning-with-files/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a well-structured, mostly executable guide with clear sequencing, explicit error-recovery loops, and properly signaled one-level-deep references. Its weaknesses are moderate redundancy across the rules/matrices/test sections, several near-duplicate reminders about file locations, and template/script references (a missing templates/ directory, unshown resolve-plan-dir.sh usage) that leave small navigation and execution gaps.

Suggestions

Merge 'The 2-Action Rule', the Read vs Write Decision Matrix, and the 5-Question Reboot Test into a single table — they largely restate the same context-management guidance — and drop the 'Context Window = RAM' analogy, which explains something Claude already knows.

State the 'planning files go in your project root' rule once (e.g., in the location table) and remove the duplicate Quick Start note and Anti-Patterns row.

Add a one-line usage example for `scripts/resolve-plan-dir.sh` (PLAN_ID/PWF_PLAN_ROOT invocation) and for `scripts/check-complete.sh`, and fix or inline the three template links — the referenced templates/ directory does not exist in this bundle.

DimensionReasoningScore

Conciseness

The body is mostly lean tables and rules, but there is real duplication and padding: the 'planning files go in your project, not the skill directory' point appears three times (the location table, the Quick Start note, and the Anti-Patterns row); 'The Core Pattern' section ("Context Window = RAM (volatile, limited) / Filesystem = Disk") explains a concept Claude already knows; and the Read vs Write Decision Matrix, 5-Question Reboot Test, and Critical Rules substantially restate one another. This fits anchor 3 ('mostly efficient but includes some unnecessary explanation or could be tightened') — above anchor 2 since nothing is tutorial-style filler, below anchor 4 because three sections could be merged or cut outright.

3 / 5

Actionability

Concrete, executable guidance dominates: a copy-paste bash one-liner for session catchup ("$(command -v python3 || command -v python) ~/.codex/skills/planning-with-files/scripts/session-catchup.py --metadata \"$(pwd)\""), a PowerShell equivalent, "run `scripts/init-session.sh \"Task Name\"`", and a fully specified list command ("sh \"<skill-dir>/scripts/set-active-plan.sh\" --list"). Minor gaps keep it at anchor 4 rather than 5: `resolve-plan-dir.sh` is invoked by name with PLAN_ID/PWF_PLAN_ROOT but no usage example is shown, and `check-complete.sh` is listed without the command line or expected output.

4 / 5

Workflow Clarity

The Quick Start is a clear numbered sequence ("Resolve or initialize the task directory" → "Create missing planning files only" → "Re-read the selected plan before decisions" → "Assign one plan owner") with recovery state restored first ("FIRST: Restore Project State"), and error feedback loops exist (the 3-Strike protocol with "AFTER 3 FAILURES: Escalate to User", plus 'Log ALL Errors'). It matches anchor 4 ('clear sequence with most checkpoints present; minor validation gaps') — the missing piece is an explicit completion checkpoint in the sequence, e.g. running `check-complete.sh` to verify all phases are complete before wrapping up.

4 / 5

Progressive Disclosure

The body is a well-sectioned overview with one-level-deep, clearly signaled references that exist in the bundle: "[references/reference.md](references/reference.md)" and "[references/examples.md](references/examples.md)" under Advanced Topics, plus a Scripts section describing each script. It sits at anchor 4 rather than 5 because the three template links ("[templates/task_plan.md](templates/task_plan.md)", etc.) point at a `templates/` directory that is not present in this bundle (the body says templates live at ~/.codex/skills/planning-with-files/templates/), so those links break in-place and force the reader to hunt for the install path.

4 / 5

Total

15

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concrete and answers both 'what' and 'when', with clearly named artifacts and an explicit usage trigger. Its main weaknesses are a dense layer of implementation/security detail (nonce framing, network-path assurances) that adds length without triggering value, and a thin 'Use when' clause that omits the natural trigger phrases the body itself contains.

Suggestions

Expand the final trigger clause to mirror the body's own list: 'Use for multi-step tasks (3+ steps), research, building projects, or work spanning many tool calls' instead of only 'research or work needing 5+ tool calls'.

Trim implementation/security caveats ('nonce-framed excerpts', 'no network upload path', 'never runs commands declared in Markdown') to one short clause so the capability sentences dominate.

Add 1-2 natural synonyms users would say (e.g., 'long-running tasks', 'staying organized across sessions') to improve trigger-term coverage.

DimensionReasoningScore

Specificity

The description lists several concrete mechanisms — "Keeps task_plan.md, findings.md, and progress.md on disk; lifecycle hooks inject selected project planning context", "Automatic recovery reads project planning files only", "session-catchup.py --metadata reads same-project local agent session records and emits aggregate counts only" — matching anchor 4's 'several specific actions; minor gaps in coverage'. It falls short of anchor 5 because a sizable share of the text is implementation/security constraints ("never runs commands declared in Markdown", "no network upload path", "nonce-framed excerpts") rather than user-facing capabilities, and the core actions (create plan, update progress, log findings) are implied rather than enumerated.

4 / 5

Completeness

Both parts are present: the 'what' is clear ("Persistent file-based planning... Keeps task_plan.md, findings.md, and progress.md on disk") and the 'when' is explicit ("Use for research or work needing 5+ tool calls"). It sits at anchor 4 rather than 5 because the 'when' clause is a single terse trigger; the richer trigger set the body itself uses ("Multi-step tasks (3+ steps)", "Building/creating projects", "Tasks spanning many tool calls") is not surfaced in the description, so the 'when' could be more explicit and specific.

4 / 5

Trigger Term Quality

Natural user-facing terms are present: "multi-step AI-agent work", "planning", "research", "work needing 5+ tool calls". This matches anchor 4 ('good keyword coverage; a few natural terms missing') but not anchor 5, since common synonyms users would actually say — "long-running tasks", "project management", "todo", "organize", "complex task" — are absent, and the file-name tokens (task_plan.md) are jargon users rarely say unprompted.

4 / 5

Distinctiveness Conflict Risk

The concrete artifacts ("task_plan.md, findings.md, progress.md", "session-catchup.py", "lifecycle hooks") carve out a recognizable niche, matching anchor 4's 'mostly distinct; minor overlap risk'. Not anchor 5: the broad framing "planning for multi-step AI-agent work" overlaps with generic task-management, todo, and planning skills, and the trigger "research or work needing 5+ tool calls" could plausibly fire for other long-running-workflow skills.

4 / 5

Total

16

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 3 missing

Warning

Total

14

/

16

Passed

Repository
OthmanAdi/planning-with-files
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.