CtrlK
BlogDocsLog inGet started
Tessl Logo

n-agentic-harnesses-codex

Designs, evaluates, and improves agentic harnesses for developer tools, assistants, workflow runtimes, copilots, and AI-powered products. Applies when work involves defining or reviewing tool-use architecture, permissions, workflow state, durability, context and memory systems, evaluation strategy, observability, user experience, or phased implementation plans for an agentic system.

76

Quality

96%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

92%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-organized, lean router with strong workflow sequencing, an explicit output contract, and a validation checklist. Its one weakness is progressive disclosure: it indexes 11 reference files that are not bundled, leaving the reference layer unrealized.

Suggestions

Bundle the referenced references/01 through references/11 markdown files alongside SKILL.md, or remove the file references and inline only the essential guidance, so the one-level-deep navigation the index promises actually exists.

Resolve the overlap between Step 1 default reads and the Step 3 reference listing (the same files appear in both) to reduce redundancy and tighten the router.

If keeping the reference index, add a brief one-line note per reference on expected output length or deliverable so Claude can pick the smallest useful set without opening files.

DimensionReasoningScore

Conciseness

The body is lean and directive ('Bias toward lean, solo-maintainable architecture', 'Push back on unnecessary complexity') with no padding explaining concepts Claude already knows, so it sits at 'lean and efficient' rather than 'mostly efficient but could be tightened'.

3 / 3

Actionability

As an instruction/routing skill it gives concrete executable guidance — explicit modes, default reads per mode, product-shape enumeration, and per-mode output contracts — so the absence of code is not penalized and it reaches 'concrete, specific guidance'.

3 / 3

Workflow Clarity

A clear sequenced workflow (Step 1 classify request, Step 2 product shape, Step 3 reads) with an explicit 'Final Check Before Responding' validation checklist, matching 'clear sequence with explicit validation steps' rather than the level-2 anchor lacking checkpoints.

3 / 3

Progressive Disclosure

The in-body indexing is excellent (11 references each with a one-line 'Read when' trigger, 'this file is the index'), but the referenced references/*.md bundle files do not exist, so the one-level-deep navigation is not actually realized and cannot reach the 'easy navigation' of score 3.

2 / 3

Total

11

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, third-person, and explicit about both capabilities and trigger conditions, with a clear niche that minimizes conflict risk. It is a strong, well-structured skill description.

DimensionReasoningScore

Specificity

Enumerates many concrete actions and subsystems ('Designs, evaluates, and improves', 'tool-use architecture, permissions, workflow state, durability, context and memory systems, evaluation strategy, observability'), matching the 'lists multiple specific concrete actions' anchor rather than the 'names domain and some actions' level below.

3 / 3

Completeness

Explicitly states both what ('Designs, evaluates, and improves agentic harnesses...') and when via an explicit 'Applies when work involves...' trigger clause, clearing the cap that a missing 'Use when' clause would impose at 2.

3 / 3

Trigger Term Quality

Uses domain-natural terms a developer would actually say ('agentic harnesses', 'tool-use architecture', 'workflow state', 'durable agent', 'evaluation strategy'), giving good coverage rather than only some relevant keywords.

3 / 3

Distinctiveness Conflict Risk

Occupies a clear niche (agentic harness design/evaluation) with distinct triggers unlikely to fire for unrelated skills, rather than overlapping broadly like 'works with document files'.

3 / 3

Total

12

/

12

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

referenced_paths_exist

Referenced path issues: 20 missing

Warning

Total

13

/

16

Passed

Repository
NateBJones-Projects/OB1
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.