Designs, evaluates, and improves agentic harnesses for developer tools, assistants, workflow runtimes, copilots, and AI-powered products. Applies when work involves defining or reviewing tool-use architecture, permissions, workflow state, durability, context and memory systems, evaluation strategy, observability, user experience, or phased implementation plans for an agentic system.
76
96%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Use this skill as a router for designing, building, and evaluating agentic harnesses.
Read only the files you need. Do not load the entire reference set unless the request genuinely spans multiple subsystems.
Default posture:
Choose one mode before reading reference files.
designUse when the user is creating a new harness, planning a major rebuild, or asking for architecture, MVP, or implementation sequencing.
Default reads:
references/01-principles-and-solo-dev-defaults.mdreferences/02-harness-shapes-and-architecture.mdreferences/08-design-and-build-playbook.mdAdd subsystem files only as needed.
evaluationUse when the user already has a harness and wants gaps, risks, missing primitives, UX upgrades, or architectural cleanup.
Default reads:
references/01-principles-and-solo-dev-defaults.mdreferences/09-evaluation-and-improvement-playbook.mdAdd subsystem files only for the parts under review.
design + evaluationUse when the user wants a target architecture and a way to verify it, compare it with an existing system, or define acceptance criteria before building.
Default reads:
references/01-principles-and-solo-dev-defaults.mdreferences/02-harness-shapes-and-architecture.mdreferences/08-design-and-build-playbook.mdreferences/09-evaluation-and-improvement-playbook.mdDetermine the closest product shape before going deeper:
If the request is ambiguous, pick the closest shape and state the assumption.
Read these only when the request needs them:
references/01-principles-and-solo-dev-defaults.md
Use first for almost every request. It defines the default decision posture.references/02-harness-shapes-and-architecture.md
Read when choosing system shape, boundaries, lifecycle, transports, or deployment structure.references/03-tools-execution-and-permissions.md
Read when the request involves tool registries, tool calling, approval gates, sandboxes, or trust tiers.references/04-state-sessions-and-durability.md
Read when the request involves sessions, resumability, retries, idempotency, approval waits, or long-running work.references/05-context-memory-and-evaluation.md
Read when the request involves context windows, retrieval, memory, provenance, evals, replay tests, or regression detection.references/06-agents-and-extensibility.md
Read when the request involves multi-agent design, plugins, hooks, skills, or extension surfaces.references/07-ux-observability-and-operations.md
Read when the request involves streaming UX, health checks, logs, analytics, budgets, or supportability.references/08-design-and-build-playbook.md
Read when the user needs a build-ready plan from idea to implementation.references/09-evaluation-and-improvement-playbook.md
Read when the user needs findings, missing primitives, upgrade priorities, or acceptance tests.references/10-example-requests-and-output-patterns.md
Read when you need prompt examples or response structure examples.references/11-codex-translation-notes.md
Read only when adapting this Anthropic-style skill into a Codex-oriented version or when mapping concepts between the two environments.Do not rely on reference-to-reference chains. This file is the index.
designReturn:
evaluationReturn:
design + evaluationReturn:
6779106
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.