Test a candidate OpenSEO skill end to end by running fresh, isolated Codex sessions against the local backend and scoring the reports they save. Use when editing a product skill (seo-audit, keyword-research, ...) and you need evidence that the new instructions produce better output, not just a cleaner file.
66
83%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Low
Low-risk findings worth noting
A skill edit is only as good as the reports it produces. This workflow runs the candidate skill in sessions that know nothing except the skill file, a one-line business brief, and a locked-down local MCP endpoint, then compares what came back.
AUTH_MODE=local_noauth in .env.local, then pnpm db:migrate:local and pnpm dev (or pnpm dev:agents). Pass the URL it prints, with /mcp appended, as --endpoint. Default local D1, never remote bindings or remote Postgres. Keep it private; it has no login.codex CLI on the path. Runs use its configured default model at high reasoning; record the resolved model from the events log if you need to name it.node .agents/skills/evaluate-skill/scripts/run.mjs \
--skill seo-audit \
--site example.com --site holdout.example \
--brief "Evaluation brief: Acme at example.com sells <what>. The goal is relevant organic visitors who could become customers. No first-party analytics are connected." \
--brief "Evaluation brief: <the holdout business>. ..." \
--endpoint http://<local-dev-host>/mcpRuns take about 10 to 14 minutes and execute in parallel. Launch the script detached, because a tool-call timeout that kills the shell kills the sessions too:
(nohup node .agents/skills/evaluate-skill/scripts/run.mjs ... > .logs/eval.log 2>&1 &)What the script does, per site:
seo-report into a fresh temp folder. Edits during the run cannot change its instructions; the manifest records the skill hashes.codex exec ephemeral, ignoring user config, memories, apps, plugins, and multi-agent, with the skill's self-review as the only reviewer.Read the reports with node .agents/skills/evaluate-skill/scripts/report-text.mjs <out>/*-report.html. Each session's working notes (for seo-audit, opportunities.md and evidence/) stay in its temp folder, named in manifest.json.
codex exec per run. Never resume or fork..logs/.This is practical isolation, not a security boundary: sessions share the local account, provider caches, and organization limits.
Always run at least two sessions of the site you are tuning on and one holdout site with a different shape, so an instruction that fits one company's answer is caught. Keep every run, including failures; do not pick the best of several and call the skill fixed.
Score each report before opening its tool trace, then read the trace to classify any miss as discovery, tool, reasoning, or reporting. For an audit report:
| Dimension | Weight | A strong report... |
|---|---|---|
| Discovery and coverage | 25% | Reads commercial pages and their siblings, tools, and technical paths; broadens when a signal appears |
| Diagnosis and evidence | 20% | Checks live pages and live results; separates observations from causes; records source, date, and settings |
| Business value and scale | 20% | Ties each action to a real buyer; compares competing opportunities; never invents conversion or revenue |
| Prioritization | 15% | Leads with a bounded change to a demonstrated problem for likely buyers; keeps maintenance and broad programs back |
| Calibration | 10% | Makes useful calls without analytics; no penalty, causation, or trend claims the evidence cannot carry |
| Communication and actionability | 10% | Names the page, the change, and the reason; material findings survive compression |
A material factual error overrides the score. Ask a second model to recompute every live position from the run's own MCP log before trusting a report's numbers.
Coverage checks that generalize beyond any one site:
Each run's project, report, and audit can be removed through MCP with delete_report, delete_site_audit, and update_project_context, or left in the local database. A new project per run is simpler than proving an old one is clean. Stop the local server when finished.
db8bde1
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.