CtrlK
BlogDocsLog inGet started
Tessl Logo

codex-huge-context

Codex 1M context: direct OpenAI Responses API inference, safe Astra/Sol/Terra/Luna input headroom, Keychain delivery, and Mac fleet rollout.

63

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/codex-huge-context/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exceptionally actionable and well-sequenced operational skill: exact config, runnable scripts, hard validation gates, and a thorough error-recovery policy. Its weaknesses are verbosity in the rationale prose and a monolithic single-file layout that inlines volatile fleet inventory and per-host status instead of moving them to a one-level-deep reference.

Suggestions

Move the Mac fleet inventory and per-host status (host names, pending items, Tailscale identities) into a one-level-deep reference file such as references/fleet-hosts.md, keeping the rollout procedure in SKILL.md.

Tighten the "Atomic provider and context invariant" section to the operational rule plus the failure mode, moving the version-specific narrative (Codex 0.144.6 behavior, the 144,000-token observation) into a clearly labeled background or old-patterns section.

Deduplicate the restart-app-servers-and-fresh-thread guidance, which currently appears in at least three sections, into one canonical checklist referenced from the failure policy.

DimensionReasoningScore

Conciseness

The body is dense and mostly non-redundant, but several passages are wordy prose that could be tightened or omitted: the four-paragraph "Atomic provider and context invariant" rationale, repeated restatements of the restart/fresh-thread rule and the 922000/700000 values across sections, and inline version-specific detail ("Codex 0.144.6 checks already-recorded context", "the observed large-context workload grew by about 144,000 tokens") that is not placed in a deprecated/old-patterns section. This matches "mostly efficient but includes some unnecessary explanation or could be tightened".

3 / 5

Actionability

Guidance is fully executable and copy-paste ready: exact JSON catalogue values, a complete TOML provider block, a runnable zsh auth script, exact preflight and verification commands with expected outputs ("922000, 922000, and 700000 for every catalogue model"), and a failure policy mapping specific errors (HTTP 401, Keychain error 36) to specific repairs. The one host-specific path is explicitly flagged as an example to resolve at install time.

5 / 5

Workflow Clarity

Multi-step processes are clearly sequenced with explicit validation checkpoints: backup before mutation, a mandatory preflight gate after any config change, the four-step app-server restart sequence, fleet rollout ordered as audit-all-then-mutate-one-at-a-time with a per-host result checklist, and a failure-policy section giving validate->fix->retry feedback loops for each failure mode. The destructive/batch fleet operations are well protected by the preflight and audit steps.

5 / 5

Progressive Disclosure

Section headers are clear and the one bundle script is properly offloaded and referenced by exact path (scripts/preflight.rb, confirmed present and consistent with the body's stated values), but the ~190-line body is a single-file operational manual rather than an overview. Volatile host-specific detail — the named Mac fleet inventory with per-host status such as FoundationClaw's pending provider reset and MiniClaw's Tailscale identity — and the long failure policy are inlined where a separate reference file would fit, and scripts/preflight.test.rb is present but never signposted. This matches "some structure but could be better organized; content that should be separate is inline".

3 / 5

Total

16

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A terse, highly specific, noun-phrase description that clearly identifies a niche topic with strong distinctiveness. Its main weakness is the complete absence of an explicit "when to use" trigger clause, which caps completeness at 3 despite the clear "what".

Suggestions

Append an explicit trigger clause, e.g. "Use when configuring, repairing, or auditing Codex 1M context, compaction thresholds, or direct OpenAI API inference setup."

Convert one or two noun phrases into action verbs (e.g. "Configures direct OpenAI Responses API inference and verifies safe input headroom") to sharpen the "what does this skill do" answer.

Add common trigger variations users might say, such as "context window", "compaction", or "million-token context".

DimensionReasoningScore

Specificity

The description lists several concrete, specific items — "direct OpenAI Responses API inference", "safe Astra/Sol/Terra/Luna input headroom", "Keychain delivery", and "Mac fleet rollout" — naming the domain and multiple specific capabilities. It falls short of 5 because these are nominal topics rather than action verbs (no "configures", "verifies", or "repairs"), leaving minor gaps in coverage of what the skill actually does.

4 / 5

Completeness

The "what" is clearly and specifically stated (direct API inference route, safe input headroom, Keychain delivery, fleet rollout), but there is no "Use when..." clause or equivalent explicit trigger guidance, capping completeness at 3 per the rubric guideline. It is not a 2 because the "what" is concrete and multi-part rather than vague.

3 / 5

Trigger Term Quality

Good keyword coverage of terms a user needing this skill would naturally say: "Codex", "1M context", "OpenAI Responses API", "Astra/Sol/Terra/Luna", "Keychain", and "Mac fleet rollout". A few natural variations are missing ("context window", "compaction", "million tokens", "configure"), which keeps it below the comprehensive synonym/extension coverage of the 5 anchor.

4 / 5

Distinctiveness Conflict Risk

The description occupies a clear niche — Codex's 1M-token context setup with named model slugs and a direct-Responses-API route — with trigger terms unlikely to fire for unrelated skills. Terms like "Astra/Sol/Terra/Luna" and "Keychain delivery" are specific to this exact scenario, matching the minimal-conflict anchor.

5 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
steipete/agent-scripts
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.