CtrlK
BlogDocsLog inGet started
Tessl Logo

codex

Run, configure, and troubleshoot OpenAI Codex CLI in non-interactive headless environments. Use for Codex automation in Bash or PowerShell, native Windows or WSL2, shell scripts, CI/CD, Docker, Kubernetes, remote servers, agent harnesses, or batch jobs; for constructing `codex exec` commands; selecting sandbox and approval modes; consuming JSONL events or structured output; resuming sessions; passing prompts through stdin; and handling failures caused by unavailable interactive input such as `request_user_input`.

77

Quality

96%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, lean operational skill: nearly all guidance is expressed as complete executable commands, risky workflows (diagnosis, completion) have explicit ordered checkpoints and validation, and Windows/recipe detail is correctly offloaded to two real, well-signaled reference files. The only weakness is minor redundancy between the Operating Rules and the mode-selection sections, which costs a small amount of token efficiency.

Suggestions

Replace the code block in Operating Rule 3 with a pointer to the 'Workspace-scoped implementation' section to remove the near-duplicate example.

Trim meta-commentary such as "Do not simulate 'always select the first option.'" by folding the intent into the preceding decision-rule list.

DimensionReasoningScore

Conciseness

The body is dense with executable commands and rules, assumes Claude's competence, and explains no basic concepts — matching 'Efficient; minor instances of over-explanation that could be trimmed'. It is not 5 because there is mild redundancy (the Operating Rules code block is repeated nearly verbatim in 'Workspace-scoped implementation', and the anti-pattern note "Do not simulate 'always select the first option'" restates the preceding rules), and not 3 because there is no genuinely unnecessary explanation or padding.

4 / 5

Actionability

Every section is anchored by copy-paste-ready, complete shell commands covering the common cases (exec with sandbox/approval flags, read-only mode, stdin prompts, --json JSONL streaming, --output-schema structured output, resume, auth). This matches 'Fully executable; copy-paste ready code or commands; specific examples cover the common cases'; it is not 4 because no key invocation pattern is left to the reader to assemble.

5 / 5

Workflow Clarity

Multi-step processes are clearly sequenced with validation checkpoints and feedback loops: the numbered 'Diagnose Failures' order (verify flags → inspect stderr and `error`/`turn.failed` events → re-run with a smaller reproducible task), the 'Completion Standard' checklist requiring tests, diff review, and reported blockers, and the autonomous prompt's decision rules ending in 'Continue until the task is complete or a concrete blocking error is reached'. This matches the anchor 'Clear sequence with explicit validation steps; feedback loops for error recovery; checklists'; destructive-operation validation is present, so the cap at 3 does not apply.

5 / 5

Progressive Disclosure

The SKILL.md is a well-organized overview, and exactly two clearly signaled one-level-deep references ([references/windows.md](references/windows.md) and [references/recipes.md](references/recipes.md)) carry the environment-specific detail; both files exist and their stated scopes match their contents. This matches 'Clear overview with well-signaled one-level-deep references; content appropriately split; easy navigation'; there is no inlining of material that belongs in a bundle file.

5 / 5

Total

19

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: it states concrete capabilities in third person, provides an explicit and detailed 'Use when' trigger clause covering the full range of automation environments, and names the specific failure mode (`request_user_input`) the skill exists to handle. Verbosity is a slight concern but every clause carries a concrete trigger or capability rather than padding.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — "Run, configure, and troubleshoot OpenAI Codex CLI", "constructing `codex exec` commands", "selecting sandbox and approval modes", "consuming JSONL events", "resuming sessions", "passing prompts through stdin" — with comprehensive coverage and no vague filler. It matches the anchor 'Lists multiple specific concrete actions; comprehensive coverage'; it is not score 4 because there are no meaningful gaps in coverage, and not below 4 because every clause names a specific capability.

5 / 5

Completeness

Both 'what' ("Run, configure, and troubleshoot OpenAI Codex CLI in non-interactive headless environments") and 'when' ("Use for Codex automation in Bash or PowerShell, ... CI/CD, Docker, Kubernetes, ... batch jobs; for constructing `codex exec` commands; ...") are explicitly stated with concrete trigger phrases. This matches the anchor 'Clearly and explicitly answers both what AND when with concrete trigger phrases'; it is not score 4 because the 'when' clause is already fully explicit and specific.

5 / 5

Trigger Term Quality

Natural terms a user would actually say are comprehensively covered: "Codex", "Codex CLI", "headless", "Bash", "PowerShell", "Windows", "WSL2", "shell scripts", "CI/CD", "Docker", "Kubernetes", "batch jobs", "codex exec", "sandbox", "JSONL", "stdin", plus the named failure mode `request_user_input`. Matches the anchor for comprehensive natural-term coverage including synonyms and environment variations; only a few plausible synonyms (e.g. 'non-interactive automation') are absent, which is the score-5 anchor, not score 4.

5 / 5

Distinctiveness Conflict Risk

The description carves a clear niche — non-interactive/headless operation of OpenAI Codex CLI specifically — with triggers (codex exec, sandbox modes, JSONL events, request_user_input) that would not fire for unrelated skills. Matches 'Clear niche with distinct triggers; minimal conflict risk'; the only adjacent overlap risk (interactive Codex usage or general CLI automation) is explicitly excluded by 'non-interactive headless environments', so score 4's 'minor overlap risk' does not apply.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
XiaomiMiMo/MiMo-Code
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.