CtrlK
BlogDocsLog inGet started
Tessl Logo

codex-claude-loop

Orchestrates a dual-AI engineering loop where Claude Code plans and implements, while Codex validates and reviews, with continuous feedback for optimal code quality

56

Quality

63%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/codex-claude-loop/SKILL.md

The canonical home for this skill is codex-claude-loop in bear2u/my-skills

SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a well-structured, actionable workflow with concrete Codex commands and explicit validation feedback loops. Its main weakness is redundancy across the iteration/recovery/perfect-loop sections, which inflates length without adding guidance.

Suggestions

Consolidate 'Phase 6: Iterative Improvement', 'Recovery When Issues Are Found', and 'The Perfect Loop' into a single feedback-loop section to remove triple-stated repetition and improve conciseness.

Replace placeholder prompts like '[Claude's plan here]' with a concrete example plan string or a templated variable so the command examples are fully copy-paste ready.

Add a short preflight checklist (e.g. confirm Codex CLI installed, sandbox mode, reasoning effort chosen) to lift workflow_clarity toward 5.

DimensionReasoningScore

Conciseness

The body is mostly efficient and does not over-explain known concepts, but Phase 6, 'Recovery When Issues Are Found', and 'The Perfect Loop' restate the same fix/re-validate cycle, and the 'Best Practices' list duplicates inline guidance.

3 / 5

Actionability

It gives concrete, runnable commands (`codex exec -m --config model_reasoning_effort="" --sandbox read-only`, `codex exec resume --last`) plus a Command Reference table, with only minor gaps such as placeholder text '[Claude's plan here]'.

4 / 5

Workflow Clarity

Phases 1–6 are clearly sequenced with explicit validation checkpoints (plan validation, cross-review, re-validation) and feedback loops ('Repeat until validation passes'), but the redundant recovery/iteration sections and lack of a crisp checklist keep it just below 5.

4 / 5

Progressive Disclosure

No bundle files exist; the skill is a single well-sectioned document with clear headers and a command table. At ~115 lines it is longer than the under-50-line simple-skill case, so minor organization gaps (redundant sections that could be consolidated) hold it at 4 rather than 5.

4 / 5

Total

15

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description conveys a clear, specific capability set but omits any explicit 'when to use' trigger guidance, which caps its completeness and weakens trigger-term quality. It is distinct and reasonably specific, just missing the activation signal.

Suggestions

Add a 'Use when...' clause naming concrete trigger scenarios, e.g. 'Use when you want a second AI to review Claude's code or when building a plan/validate/implement loop with Codex'.

Include natural user phrasings and synonyms such as 'code review', 'AI pair programming', 'cross-check implementation', or 'validate a plan' to improve trigger-term coverage.

Trim the abstract closer 'for optimal code quality' in favor of a concrete outcome to push specificity toward 5.

DimensionReasoningScore

Specificity

The description names concrete actions — 'plans and implements', 'validates and reviews', 'continuous feedback' — covering several specific capabilities with only minor abstraction ('optimal code quality').

4 / 5

Completeness

It clearly answers 'what' (orchestrates a dual-AI plan/implement/validate loop) but provides no 'Use when...' trigger guidance, so per the judging guidelines completeness is capped at 3.

3 / 5

Trigger Term Quality

It uses relevant terms like 'engineering loop', 'code review', 'validation', and 'planning', but lacks the natural trigger phrases a user would actually say and offers no synonyms or variations.

3 / 5

Distinctiveness Conflict Risk

The dual-AI Claude+Codex loop is a fairly distinct niche with low conflict risk, though generic phrasing like 'optimal code quality' and 'continuous feedback' leaves minor overlap with general code-review skills.

4 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
bear2u/my-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.