CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-watchdog

Use when asked to watch, babysit, audit, review, compare, or fix another agent's work from a Codex session ID, Claude Code session/transcript, chat/thread link, PR, branch, log, or pasted run summary. Monitor until the other agent is done or blocked, reconstruct what the user asked, inspect what the agent actually changed and verified, report gaps, and optionally make scoped fixes when the user authorizes repair.

78

Quality

98%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

96%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a tightly written, highly actionable process skill with a clear sequenced workflow and built-in validation and feedback loops; its only soft spot is progressive disclosure, where a longer single-file body leaves minor room to externalize the report template or classification taxonomy.

Suggestions

Consider moving the report markdown template and the five-category issue taxonomy into a references/ file (e.g. REPORT_TEMPLATE.md), keeping SKILL.md as a lean overview — this would push progressive_disclosure to 5 and shorten the body under the 50-line threshold.

Add one or two concrete example invocations (e.g. a sample 'watch a Codex session until done' and a sample 'audit a PR diff') to make the workflow phases even more copy-paste ready.

Make the per-fix validation loop slightly more explicit with a 'if validation fails → fix → re-validate' framing so the feedback loop is unmistakable rather than implied by 're-run after each meaningful fix'.

DimensionReasoningScore

Conciseness

The body is lean and vivid ('Watch another agent's work like a reviewer with a pager') with no padding or explanations of concepts Claude already knows; every section is a tight bullet list that assumes competence.

5 / 5

Actionability

Though instruction-only, it gives concrete executable guidance per phase — artifact resolution steps, a contract checklist, specific evidence checks, a five-category issue taxonomy, and a copy-paste report template — so absence of code is not a deficit.

5 / 5

Workflow Clarity

A clear six-phase sequence (Mode → Resolve → Reconstruct → Audit → Fix → Report) with an explicit validation checkpoint ('Re-run the smallest useful validation after each meaningful fix'), feedback loops (stop-and-report on product decisions), and checklists for classification.

5 / 5

Progressive Disclosure

No bundle files exist and the ~110-line body is well-organized with clear section headers and easy navigation; it exceeds the under-50-line auto-5 threshold, so it earns good-but-not-ideal structure rather than a perfect split.

4 / 5

Total

19

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is exemplary: it pairs an explicit trigger clause with a comprehensive, concrete action list and a well-scoped niche, leaving little ambiguity about when or how the skill applies.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — watch/babysit/audit/review/compare/fix, monitor until done or blocked, reconstruct the request, inspect changes and verification, report gaps, and make scoped fixes — giving comprehensive coverage rather than vague language.

5 / 5

Completeness

Explicitly answers 'when' via a 'Use when asked to...' clause with concrete triggers and 'what' via the monitor/reconstruct/inspect/report/fix action chain, satisfying both halves concretely.

5 / 5

Trigger Term Quality

Uses natural user phrasings ('watch, babysit, audit, review, compare, or fix another agent's work') with synonyms plus many concrete artifact types (Codex session ID, Claude Code transcript, chat/thread link, PR, branch, log, pasted run summary).

5 / 5

Distinctiveness Conflict Risk

Occupies a clear niche — overseeing another agent's work from session/transcript/PR sources — with distinct triggers and minimal overlap with other skills.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
BuilderIO/skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.