CtrlK
BlogDocsLog inGet started
Tessl Logo

codex-huge-context

Codex 1M context: direct OpenAI Responses API inference, safe Sol/Terra/Luna input headroom, Keychain delivery, and Mac fleet rollout.

64

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/codex-huge-context/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is highly actionable and well-sequenced with strong validation and recovery guidance, fitting a destructive/batch skill well. Tightening the token-math narration and moving some inline config bulk to reference files would push conciseness and progressive disclosure to the top of the scale.

DimensionReasoningScore

Conciseness

The body assumes Claude's competence (no basic-concept padding) and stays technical, but some token-arithmetic narration (875,900/175,900/222,000) and failure-scenario prose could be trimmed, fitting 'efficient; minor instances of over-explanation' rather than fully lean.

4 / 5

Actionability

It provides complete, copy-paste-ready artifacts — exact config.toml, catalogue JSON values, the zsh Keychain helper, and verification commands with concrete expected outputs ('922000', '922000', '700000', probe response).

5 / 5

Workflow Clarity

Despite involving batch/fleet mutation, it sequences the work (backup -> preflight -> mutate one host at a time -> restart app servers -> fresh-session proof) with explicit validation checkpoints and error-recovery feedback loops in the Failure Policy.

5 / 5

Progressive Disclosure

Sections are clearly organized and the preflight script is referenced one level deep (and exists in scripts/), but sizable inline config blocks and failure-case detail arguably belong in referenced files, so it lands at 'good structure; minor organization gaps' rather than a fully split 5.

4 / 5

Total

18

/

20

Passed

Description

62%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and clearly distinct, but it omits any explicit 'when to use' trigger guidance, which limits completeness. Adding a 'Use when...' clause with natural user phrasings would raise both completeness and trigger-term quality.

Suggestions

Add a 'Use when...' clause naming concrete triggers (e.g., configuring, repairing, or auditing Codex's 1M context).

Include more natural user-facing terms ('context window', 'configure Codex') alongside the internal model codenames.

Spell out a couple of the capabilities as actions (e.g., 'configure safe input headroom and roll out across a Mac fleet') to push specificity toward 5.

DimensionReasoningScore

Specificity

Names the domain and several concrete capabilities ('direct OpenAI Responses API inference, safe Sol/Terra/Luna input headroom, Keychain delivery, and Mac fleet rollout'), matching the 'lists several specific actions; minor gaps' anchor rather than the comprehensive 5.

4 / 5

Completeness

It clearly states what the skill does but has no 'Use when...' trigger clause, which per the judging guidelines caps completeness at 3.

3 / 5

Trigger Term Quality

Relevant domain terms appear (Codex, 1M context, OpenAI Responses API, Mac fleet) but they lean technical/internal ('Sol/Terra/Luna', 'input headroom', 'Keychain delivery') and omit common user-facing phrasings like 'configure Codex' or 'context window', fitting the 'some relevant keywords but missing variations' anchor.

3 / 5

Distinctiveness Conflict Risk

It targets a narrow, specific niche (Codex 1M-context configuration) with distinct triggers and minimal overlap risk with other skills.

5 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
steipete/agent-scripts
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.