CtrlK
BlogDocsLog inGet started
Tessl Logo

codex

Delegate coding to OpenAI Codex CLI (features, PRs).

54

Quality

62%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/autonomous-ai-agents/codex/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

76%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a strong, action-dense reference: nearly every line is a runnable command or an operational gotcha (PTY requirement, git-repo requirement, sandbox failure modes, auth file locations) that Claude could not infer. Its one real weakness is workflow validation — the parallel/batch workflows push and open PRs without any verify-before-push checkpoint, which the rubric caps at 3. Splitting the flag reference and batch patterns into a reference file would also lighten the overview.

Suggestions

Insert an explicit validation step in the Parallel Issue Fixing workflow between completion and push, e.g., '# Verify: cd ... && git diff main --stat && run targeted tests' before 'git push -u origin'.

In the Batch PR Reviews workflow, add a checkpoint to confirm each review diff is non-empty/correct before posting 'gh pr comment', so posting cannot fire on a failed review.

Move the Key Flags table and batch/worktree recipes into a single one-level-deep reference file (e.g., references/advanced.md) to slim the overview, and remove the Rules-section duplication of the flag guidance.

DimensionReasoningScore

Conciseness

The body is command-first with almost no filler: installation, auth edge cases, sandbox flags, and runnable terminal/process invocations are all operational knowledge Claude would not already know. The one notable redundancy is the Rules section restating Key Flags content (e.g., '--sandbox workspace-write ... --full-auto is deprecated' appears in both the table and Rule 4), which keeps it below the lean 5 anchor.

4 / 5

Actionability

Every section provides copy-paste-ready, fully executable commands covering the common cases: one-shot exec, scratch-repo bootstrap (mktemp + git init), background launch with poll/log/submit/kill monitoring, PR checkout, parallel worktrees, and batch PR review with gh commands. This matches the 5 anchor 'Fully executable; copy-paste ready code or commands; specific examples cover the common cases'.

5 / 5

Workflow Clarity

The multi-step workflows (worktrees, batch PR reviews) are clearly sequenced with inline comments, but the batch/parallel workflows proceed straight from 'Fix issue' to 'git push' and 'gh pr create' with no explicit validation checkpoints (no test run, diff check, or verification gate before pushing). The batch-operations guideline caps workflow clarity at 3 in this case; the safety/validation guidance that does exist ('git diff review, targeted tests, human confirmation') is confined to the gateway caveat section rather than wired into the workflows.

3 / 5

Progressive Disclosure

The single-file body (~140 lines) is well organized into clearly labeled sections (Prerequisites, One-Shot, Background, Key Flags, PR Reviews, Worktrees, Batch) with no nested or buried references, and no bundle files exist so there is nothing mis-split. It stops short of the 5 anchor's ideal split — the Key Flags table and advanced batch workflows are moderately advanced material that could live in a one-level-deep reference file — fitting 'Good structure; most content is appropriately placed; minor organization gaps'.

4 / 5

Total

16

/

20

Passed

Description

48%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise and correctly names its niche (delegating coding to the Codex CLI), keeping conflict risk low. Its main weakness is completeness: it states what the skill does but gives no 'when to use' trigger guidance, and its action coverage is minimal (one generic verb plus two task nouns). Adding a 'Use when...' clause with a few concrete actions and natural trigger phrases would move it into the 4-5 range.

Suggestions

Add an explicit 'Use when...' clause, e.g., 'Use when the user asks to delegate coding to Codex — building features, refactoring, PR reviews, or batch issue fixes.'

Replace the generic 'Delegate coding' verb with 2-3 concrete actions: 'Runs Codex CLI one-shot or in background to build features, refactor modules, and review PRs.'

Include natural trigger synonyms users would actually say — 'code review', 'refactoring', 'fix issues in bulk', 'delegate to Codex' — to broaden keyword coverage.

DimensionReasoningScore

Specificity

The description names the domain ("OpenAI Codex CLI") but offers only a single generic action ("Delegate coding") plus two bare task nouns ("features, PRs"); no concrete capability verbs are listed. This matches the anchor 'Names the domain but actions are minimal or generic' — it is below the 3 anchor, which requires 1-2 genuinely concrete actions (e.g., 'Refactors modules, reviews PRs, fixes batch issues').

2 / 5

Completeness

The 'what' is stated clearly ("Delegate coding to OpenAI Codex CLI") but there is no 'Use when...' or equivalent trigger clause anywhere, so 'when' is entirely missing. Per the guideline that a missing 'Use when...' clause caps completeness at 3, this matches the anchor 'Has a clear what but when is missing or only weakly implied' and cannot score 4.

3 / 5

Trigger Term Quality

Terms like "Codex", "coding", and "PRs" are natural phrases a user might say, giving some relevant keyword coverage. However, common variations and synonyms are missing — no 'code review', 'refactoring', 'delegate', 'build a feature', or 'fix issues' — matching the anchor 'Some relevant keywords but missing common variations or synonyms' rather than the 4 anchor's good coverage.

3 / 5

Distinctiveness Conflict Risk

"OpenAI Codex CLI" carves out a clear, distinct niche with tool-specific triggers that a generic coding skill would not match, and "PRs" narrows the review use case. Minor overlap risk remains with related coding-agent skills (the frontmatter itself lists related_skills: claude-code, hermes-agent), so it fits 'Mostly distinct; minor overlap risk with closely related skills' rather than the 5 anchor's minimal-conflict bar.

4 / 5

Total

12

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.