CtrlK
BlogDocsLog inGet started
Tessl Logo

codex-skill

Leverage OpenAI Codex/GPT models for autonomous code implementation, code review, and plan review. Triggers: "codex", "use gpt", "gpt-5", "let openai", "full-auto", "adversarial review", "second opinion review", "用codex", "让gpt实现", "对抗式审查", "让codex审查计划", "第二意见". Use this skill whenever the user wants to delegate coding tasks to OpenAI models, run code or plan reviews via codex, get a second-opinion review from a different model, or execute tasks in a sandboxed environment.

64

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

72%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-disclosed instruction skill: executable commands, explicit error-handling and review-stop checkpoints, and a clean one-level reference bundle. Its weaknesses are repetitive security directives that inflate the token budget and the absence of an explicitly sequenced end-to-end workflow tying the sections together.

Suggestions

Add a short numbered main workflow (verify install → choose sandbox mode → size the task → run sync/background → present results) so the happy-path sequence is explicit rather than implied by section order; the current 'Execution Workflow' section only covers session resume.

Consolidate the repeated sandbox/danger-full-access rules stated in 'Security & Trust Boundaries', 'Core Principles', and 'Operating Modes' into one canonical statement referenced elsewhere, cutting roughly 15-20 lines of duplicated directives.

State the mode-selection rule once as a compact decision table (read-only vs --full-auto vs danger-full-access with consent conditions) instead of restating it across three sections.

DimensionReasoningScore

Conciseness

Mostly efficient imperative guidance with no padding about what Codex is, but several directives are repeated across sections: never escalate the sandbox appears in 'Security & Trust Boundaries' and 'Core Principles'; danger-full-access consent rules appear three times (lines 17, 67 area, and When to Interrupt); review-stop rules appear in both 'Handling Review Results' and 'When to Interrupt'. Fits 'mostly efficient but could be tightened'; not 4 because the duplication is more than minor trimming.

3 / 5

Actionability

Fully executable, copy-paste-ready commands covering the common cases: 'codex exec --full-auto "implement..."', 'codex exec review --uncommitted', 'codex exec resume --last "<delta>"', 'codex exec ... 2>&1 | tee /tmp/codex-<slug>.log', '< /dev/null', 'git diff --shortstat', plus a concrete failure-handling protocol. Matches the top anchor; nothing missing for typical invocations.

5 / 5

Workflow Clarity

All workflow phases exist as separate, individually clear sections (verify install, pick mode, run, handle errors, report output), but the happy-path sequence is implicit — there is no explicit ordered walkthrough, and the 'Execution Workflow' section contains only resume advice rather than a workflow. Fits 'sequence present but implicit'; not 4 because the anchor's clear connected sequence with checkpoints is not realized as a coherent main flow.

3 / 5

Progressive Disclosure

SKILL.md is a genuine overview; four reference files plus the schema asset all exist, are exactly one level deep (no nested references found in them), and each entry carries an explicit 'Read this when...' signal (e.g. 'Read this when the task needs a flag not covered above'). Matches the top anchor; not 4 because no organization gap was found.

5 / 5

Total

16

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that explicitly states both capabilities and usage conditions, backed by a generous bilingual trigger-term list. Its only weaknesses are a few generic capability phrases and some missing natural trigger variants that leave minor overlap risk with general review skills.

DimensionReasoningScore

Specificity

Names several concrete capabilities — 'autonomous code implementation, code review, and plan review', 'get a second-opinion review from a different model', 'execute tasks in a sandboxed environment' — matching the 'several specific actions; minor gaps' anchor. Not 5 because phrases like 'delegate coding tasks to OpenAI models' are generic relative to the comprehensive anchor's fully concrete action list.

4 / 5

Completeness

Clearly answers both: what — 'autonomous code implementation, code review, and plan review' — and when — 'Use this skill whenever the user wants to delegate coding tasks..., run code or plan reviews via codex, get a second-opinion review..., or execute tasks in a sandboxed environment', with an explicit 'Triggers:' phrase list. Matches the top anchor exactly; not below it since both halves are explicit and concrete.

5 / 5

Trigger Term Quality

Explicit trigger list covers natural terms users would say — 'codex', 'use gpt', 'gpt-5', 'let openai', 'full-auto', 'adversarial review', 'second opinion review', plus five Chinese synonyms — good coverage per the anchor. Not 5 because common variants like 'let codex', 'ask gpt', 'use openai', or plain 'code review'/'plan review' as triggers are absent.

4 / 5

Distinctiveness Conflict Risk

Clear niche (delegating to the OpenAI Codex CLI) with distinct tool-specific triggers like 'codex' and 'full-auto', so mostly distinct per the anchor. Not 5 because 'adversarial review' and 'second opinion review' could also route to general code-review skills, creating minor overlap risk with closely related skills.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

15

/

16

Passed

Repository
feiskyer/claude-code-settings
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.