CtrlK
BlogDocsLog inGet started
Tessl Logo

ce-babysit-pr

Babysits an open GitHub PR until merge-ready. Use when asked to watch a PR over time — not for one-shot comment resolution or one CI failure. GitHub (incl. Enterprise) only.

69

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

The canonical home for this skill is ce-babysit-pr in EveryInc/compound-engineering-plugin

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exceptionally disciplined instruction-only skill body: a tight 4-step workflow with an explicit per-tick ordering invariant, built-in validation guards (terminal check, stale-SHA cancellation, stop-condition checklist, budget backstop), and zero padding. The two weaknesses are minor: the core pr-snapshot invocation is deferred to references rather than shown once in the body, and the reference web nests two levels deep with one file (stack-commands.md) not surfaced from SKILL.md.

Suggestions

Include one minimal executable example of the pr-snapshot invocation (or a pointer to its interface) directly in Step 2, so the skill's core action is runnable without first opening references/tick.md.

Surface stack-commands.md from the body (e.g. in the posture or stack-maintenance bullet) so every bundle file is reachable one level deep from SKILL.md.

DimensionReasoningScore

Conciseness

The ~50-line body is lean and assumes Claude's competence: no known-concept explanation, no padding, and every rule is novel domain content ("Settled ≠ merged", "a push that restarts green CI without a claimed item is a defect"). It matches 'every token earns its place'; a 4 would require minor trimmable over-explanation, of which there is none.

5 / 5

Actionability

Mostly executable: copy-paste commands appear for key operations ("gh run rerun <run-id> --failed -R <host>/<owner>/<repo>", "gh pr checkout <ref>", "gh stack merge") and delegate invocations are exact ("ce-resolve-pr-feedback mode:pipeline", "ce-debug mode:pipeline"). It is not a 5 because the central pr-snapshot invocation that drives every tick is never shown in the body — the reader must open references/tick.md or watch-loop.md to execute the skill's core action.

4 / 5

Workflow Clarity

The sequence is explicit with validation checkpoints and feedback loops: a numbered 4-step arc, an ordering invariant inside each tick (terminal check first, stale-SHA cancellation guard, "mark each check acted on", "unfixed checks stay red residuals"), enumerated true-stop conditions with a budget backstop, and defect reporting for unauthorized actions. This matches 'clear sequence with explicit validation steps; feedback loops for error recovery'.

5 / 5

Progressive Disclosure

The body is a well-organized overview that points to topical reference files at the exact step where each is needed (tick.md, envelope.md, settle.md, branch-currency.md, stack.md, watch-loop.md, pipeline.md, report.md, setup.md), all of which exist in references/. It is not a 5 because reference files cross-reference each other (e.g. settle.md → stack.md → stack-commands.md, two levels deep) and stack-commands.md is unreachable from the body, a minor navigation gap; it is well above a 3 since the SKILL.md-level structure is clean and every split is appropriate.

4 / 5

Total

18

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: explicit what and when with concrete trigger phrases, negative delimiters that separate it from one-shot PR-fixing skills, and a clear GitHub-only scope. Voice is correctly third person ("Babysits"). The main weakness is that it states only one umbrella action, leaving the concrete capabilities (comment resolution, CI fixes, branch updates) implicit.

Suggestions

Enumerate the concrete streams the skill acts on in the description, e.g. "reacts to review comments, CI failures, and branch-currency items" — this would raise specificity from 3 toward 4-5 without adding padding.

Add one or two natural synonyms such as "pull request" or "monitor/keep an eye on a PR" so users phrasing the request differently still hit the trigger.

Optionally name the terminal states users care about ("left at merge-ready, blocked, or out-of-budget") to sharpen the what-clause.

DimensionReasoningScore

Specificity

"Babysits an open GitHub PR until merge-ready" names the domain but covers only one umbrella action; the concrete streams it reacts to (review comments, CI failures, branch currency) do not appear in the description, so it matches 'names domain and 1-2 concrete actions, but not comprehensive'. It is not a 4 because several specific actions are not listed, and not a 2 because the action stated is concrete rather than generic.

3 / 5

Completeness

It explicitly answers both: what ("Babysits an open GitHub PR until merge-ready") and when ("Use when asked to watch a PR over time — not for one-shot comment resolution or one CI failure"), with concrete trigger phrases plus negative delimiters and scope ("GitHub (incl. Enterprise) only"), matching the top anchor exactly. A 4 would have a weaker or less explicit 'when' clause, which this does not.

5 / 5

Trigger Term Quality

Natural phrases users would say are present — "watch a PR over time", "PR", "GitHub", "merge-ready", "one-shot comment resolution", "CI failure" — but common synonyms like "pull request", "monitor", or "keep an eye on" are missing, matching 'good keyword coverage; a few natural terms missing' rather than comprehensive coverage.

4 / 5

Distinctiveness Conflict Risk

The niche is clear (ongoing PR watching to merge-readiness) and it explicitly distinguishes itself from adjacent skills via negative triggers ("not for one-shot comment resolution or one CI failure") and the "GitHub (incl. Enterprise) only" scope, giving minimal conflict risk. A 4 would leave minor overlap with closely related skills, which the negative delimiters already close off.

5 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
crdant/compound-engineering-plugin
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.