CtrlK
BlogDocsLog inGet started
Tessl Logo

n-agentic-harnesses-codex

Designs, evaluates, and improves agentic harnesses for developer tools, assistants, workflow runtimes, copilots, and AI-powered products. Applies when work involves defining or reviewing tool-use architecture, permissions, workflow state, durability, context and memory systems, evaluation strategy, observability, user experience, or phased implementation plans for an agentic system.

64

Quality

79%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/n-agentic-harnesses/variants/codex/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-organized, lean router skill with clear sequencing and a concrete output contract and final checklist. Its main defect is that every referenced reference file is missing from the bundle, breaking the progressive-disclosure navigation the skill is built around.

Suggestions

Ship the references/ directory with the 11 referenced files (references/01-principles-and-solo-dev-defaults.md through references/11-codex-translation-notes.md), or inline the essential content so the skill is self-contained.

Reduce redundancy: the three "Default reads" sections all re-list references/01-principles-and-solo-dev-defaults.md, and Step 3 re-indexes files already cited in Step 1 — consolidate the canonical index in one place.

Add a short output format template (e.g., a markdown skeleton) to the Output Contract so the required return items are copy-paste ready rather than just enumerated.

DimensionReasoningScore

Conciseness

The body is lean, assumes Claude's competence (no explanations of what agentic harnesses or MCP are), and every section is actionable, fitting anchor 4; not 5 because the three "Default reads" lists repeat reference 01 and Step 3 re-lists files already named in Step 1, offering minor trim opportunities.

4 / 5

Actionability

Provides concrete routing (specific reference files per situation), explicit output-contract items per mode, and a specific final checklist, matching anchor 4's "mostly executable guidance; minor gaps"; not 5 because the output contract lists what to return without a copy-paste-ready format template.

4 / 5

Workflow Clarity

Clear numbered sequence (Step 1 classify request, Step 2 classify shape, Step 3 read references) plus an explicit "Final Check Before Responding" validation checkpoint, matching anchor 4; not 5 because there is no per-step validation/feedback loop, though that is largely N/A for a non-destructive design skill.

4 / 5

Progressive Disclosure

The written structure is an exemplary one-level-deep index with one-line "Read when..." descriptors and an explicit "this file is the index, no chains" note, but none of the 11 referenced files exist in the bundle, so the signaled navigation leads nowhere and progressive disclosure does not actually function.

3 / 5

Total

15

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly answers both what the skill does and when to apply it, with concrete trigger terms and a well-defined niche. The only mild weakness is that the core action verbs (design/evaluate/improve) are relatively high-level rather than granular.

DimensionReasoningScore

Specificity

Names three concrete actions ("Designs, evaluates, and improves") plus a comprehensive enumerated domain list (tool-use architecture, permissions, workflow state, durability, etc.), matching anchor 4's "several specific actions; minor gaps" rather than 5 since the verbs themselves stay somewhat high-level.

4 / 5

Completeness

Explicitly states both what ("Designs, evaluates, and improves agentic harnesses...") and when ("Applies when work involves...") with concrete trigger phrases, matching anchor 5; the "when" clause is explicit and specific, ruling out 4.

5 / 5

Trigger Term Quality

Includes natural phrases users would say ("agentic harnesses", "tool-use architecture", "workflow state", "permissions", "copilots") with synonym variation, fitting anchor 4's "good keyword coverage; a few natural terms missing"; not 5 because no file extensions appear, though that is partly domain-inappropriate here.

4 / 5

Distinctiveness Conflict Risk

Targets a clear niche (agentic harnesses) with distinct, specific triggers (durability, context and memory systems, evaluation strategy) that minimize overlap with generic architecture skills, matching anchor 5.

5 / 5

Total

18

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

referenced_paths_exist

Referenced path issues: 20 missing

Warning

Total

13

/

16

Passed

Repository
NateBJones-Projects/OB1
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.