Designs, evaluates, and improves agentic harnesses for developer tools, assistants, workflow runtimes, copilots, and AI-powered products. Applies when work involves defining or reviewing tool-use architecture, permissions, workflow state, durability, context and memory systems, evaluation strategy, observability, user experience, or phased implementation plans for an agentic system.
64
81%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
Use this skill as a router for designing, building, and evaluating agentic harnesses.
Read only the files you need. Do not load the entire reference set unless the request genuinely spans multiple subsystems.
Default posture:
Choose one mode before reading reference files.
designUse when the user is creating a new harness, planning a major rebuild, or asking for architecture, MVP, or implementation sequencing.
Default reads:
references/01-principles-and-solo-dev-defaults.mdreferences/02-harness-shapes-and-architecture.mdreferences/08-design-and-build-playbook.mdAdd subsystem files only as needed.
evaluationUse when the user already has a harness and wants gaps, risks, missing primitives, UX upgrades, or architectural cleanup.
Default reads:
references/01-principles-and-solo-dev-defaults.mdreferences/09-evaluation-and-improvement-playbook.mdAdd subsystem files only for the parts under review.
design + evaluationUse when the user wants a target architecture and a way to verify it, compare it with an existing system, or define acceptance criteria before building.
Default reads:
references/01-principles-and-solo-dev-defaults.mdreferences/02-harness-shapes-and-architecture.mdreferences/08-design-and-build-playbook.mdreferences/09-evaluation-and-improvement-playbook.mdDetermine the closest product shape before going deeper:
If the request is ambiguous, pick the closest shape and state the assumption.
Read these only when the request needs them:
references/01-principles-and-solo-dev-defaults.md
Use first for almost every request. It defines the default decision posture.references/02-harness-shapes-and-architecture.md
Read when choosing system shape, boundaries, lifecycle, transports, or deployment structure.references/03-tools-execution-and-permissions.md
Read when the request involves tool registries, tool calling, approval gates, sandboxes, or trust tiers.references/04-state-sessions-and-durability.md
Read when the request involves sessions, resumability, retries, idempotency, approval waits, or long-running work.references/05-context-memory-and-evaluation.md
Read when the request involves context windows, retrieval, memory, provenance, evals, replay tests, or regression detection.references/06-agents-and-extensibility.md
Read when the request involves multi-agent design, plugins, hooks, skills, or extension surfaces.references/07-ux-observability-and-operations.md
Read when the request involves streaming UX, health checks, logs, analytics, budgets, or supportability.references/08-design-and-build-playbook.md
Read when the user needs a build-ready plan from idea to implementation.references/09-evaluation-and-improvement-playbook.md
Read when the user needs findings, missing primitives, upgrade priorities, or acceptance tests.references/10-example-requests-and-output-patterns.md
Read when you need prompt examples or response structure examples.references/11-codex-translation-notes.md
Read only when adapting this Anthropic-style skill into a Codex-oriented version or when mapping concepts between the two environments.Do not rely on reference-to-reference chains. This file is the index.
designReturn:
evaluationReturn:
design + evaluationReturn:
238df6c
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.