CtrlK
BlogDocsLog inGet started
Tessl Logo

babysit-pr

Monitor a PR and fix CI or review feedback. Use standalone, or from /ship to continue through its authorized guarded merge instead of stopping at green.

57

Quality

72%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/babysit-pr/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

70%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exceptionally actionable, well-validated operational loop whose body is undermined by heavy internal repetition and a monolithic structure. The command-level guidance and validation checkpoints are top-tier, but the same rules are maintained in multiple places and 500 lines of policy are inlined with no reference files to split them.

Suggestions

Deduplicate the push-target/remote-verification rules into one canonical section (currently stated in the intro, 'For an existing PR update', Setup, and 'Each tick') and reference it instead of restating it.

Split rarely-needed detail — the missing-changeset special case, the latest-feedback handoff ledger rules, and the post-merge /ship continuation — into references/ files linked from SKILL.md.

Consolidate the clock/stop rules (30-minute quiet-green, 10-minute soak, mode endpoints) into a single table or section; they currently appear in at least four places.

DimensionReasoningScore

Conciseness

The push-target and remote-verification rules are restated roughly four times (intro paragraph, 'For an existing PR update', Setup item 2, and 'Each tick'), and the 'never rebase/force-push', clock-reset, and stop-condition rules each recur across multiple sections — substantial duplication rather than isolated tightening opportunities. Every individual rule is nontrivial domain policy rather than general knowledge, but the repetition is padded verbosity.

2 / 5

Actionability

Fully executable throughout: exact gh/git commands, working jq pipelines for reply-state comment coverage, a complete changeset file template with slug examples, and the guarded merge command with --match-head-commit. Copy-paste ready commands cover the common cases.

5 / 5

Workflow Clarity

The tick sequence is explicit (Step 0 status → Step 1 conflicts → feedback → CI → merge gate) with abundant validation: live OID rechecks before every push, the five simultaneous merge conditions, full gate revalidation pre-merge, and dual final audits. Feedback loops for conflict and non-fast-forward recovery, soak-timer resets, and a merge checklist are all present.

5 / 5

Progressive Disclosure

No bundle files exist; ~500 lines of dense policy are inlined in SKILL.md. Section headers give it structure, but content that clearly belongs in separate files (changeset special case, latest-feedback handoff ledger rules, post-merge /ship continuation) is inline, and there are no one-level-deep references.

3 / 5

Total

15

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A concise, specific description that names concrete actions in a distinct niche, written in third person. Its main weakness is the missing explicit 'Use when…' trigger guidance — it explains invocation modes but not the situations in which a user would reach for it.

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user asks to watch, babysit, or monitor a PR, or to fix CI failures or review comments on a PR.'

Spell out 'pull request' alongside 'PR' and add synonyms like 'checks' or 'CI failures' to broaden natural trigger coverage.

DimensionReasoningScore

Specificity

'Monitor a PR and fix CI or review feedback' and 'continue through its authorized guarded merge' list several concrete actions in a named domain. It falls short of 5 because capabilities like conflict resolution, replying to feedback, and merge gating are not mentioned, but exceeds 3's 1-2 action bar.

4 / 5

Completeness

The 'what' is clear (monitor a PR, fix CI and review feedback), but there is no 'Use when…' trigger clause — 'Use standalone, or from /ship' describes invocation modes, not when a user needs the skill. Per the guideline, a missing explicit trigger caps completeness at 3.

3 / 5

Trigger Term Quality

'PR', 'CI', 'review feedback', and 'merge' are natural terms users would say when they need this skill. 'Pull request' is never spelled out and synonyms like 'checks' or 'build failures' are missing, so coverage is good rather than comprehensive.

4 / 5

Distinctiveness Conflict Risk

PR babysitting through CI, review feedback, and a guarded merge is a clear niche with distinct triggers, leaving only minor overlap risk with closely related CI-fix or ship skills. It is not entirely free of that overlap, so not 5.

4 / 5

Total

15

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (514 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

12

/

16

Passed

Repository
BuilderIO/agent-native
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.