CtrlK
BlogDocsLog inGet started
Tessl Logo

babysit

Drive a PR to a clean review (Greptile 5/5, zero open threads) — ships if needed, keeps it mergeable against staging, re-triggers both Greptile and cubic, fixes real findings, replies to and resolves every thread, and loops until clean

65

Quality

78%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/babysit/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exceptionally actionable, well-validated workflow document — every step is executable, the clean-state definition is airtight, and validation/checkpoint discipline is exemplary. The main cost is repetition: the Hard Rules section and several inline rationales restate invariants already spelled out in the loop, which inflates token cost without adding new information.

DimensionReasoningScore

Conciseness

Mostly efficient — nearly all content is repo-specific operational knowledge (bot behaviors, exact commands, query gotchas) rather than concepts Claude already knows. But the same invariants are restated repeatedly: the 'trigger both reviewers' warning appears in the table, step 8, and Hard Rules; the sync check is explained at length in steps 6, 7, and Hard Rules; and several long rationale sentences ('A babysit loop spanning a long session is exactly the scenario where...') could be tightened. This fits the 3 anchor ('mostly efficient but could be tightened') better than 4 ('minor instances'), since the duplication spans multiple sections.

3 / 5

Actionability

Every step carries copy-paste-ready commands: exact `gh pr view`/`gh api graphql` queries with jq filters, the reply and resolve mutations with the correct id fields, the exact trigger wordings ('@greptile', '@cubic-dev-ai review this PR'), and the post-push verify commands. Fully executable with specific gotchas (thread-level author field breaks the query, thread id vs comment databaseId) covering the common failure cases — the 5 anchor.

5 / 5

Workflow Clarity

A clearly sequenced 10-step loop with explicit validation checkpoints: the three-part 'clean' definition checked freshly every round, pagination handled via pageInfo before evaluating clean, post-push commit-sync verification after every push, confirmation that both bots went 'pending' after re-triggering, an explicit wait state (step 9) for pending checks, and defined stop conditions with a two-round escalation. This is a feedback-loop-rich workflow for public/batch actions — squarely the 5 anchor, not 4, because validation and error-recovery loops are explicit at every stage rather than mostly present.

5 / 5

Progressive Disclosure

Well-sectioned single-file skill (no bundle files exist), with one-level-deep, clearly signaled cross-references to `.agents/skills/ship/SKILL.md` steps instead of inlining that material. It stays at 4 rather than 5 because the skill is well over 50 lines and leans on paraphrased descriptions of another skill's steps ('re-run the full sync check from /ship step 2 — not just the log command, the whole check-and-recover flow'), which is a minor organization gap versus either a self-contained reference or tighter pointers.

4 / 5

Total

17

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A highly specific, well-differentiated description that clearly states what the skill does, dragged down by the complete absence of when-to-use trigger guidance. Adding an explicit 'Use when...' clause with natural user phrasings would lift it into the top tier.

Suggestions

Append an explicit trigger clause, e.g. 'Use when the user says "babysit this PR", "keep working the reviews until it's clean", or wants the /ship follow-up review loop automated.'

Move the natural trigger phrasings currently buried in the body's 'When to use' section into the description itself so they surface at skill-selection time.

Optionally name the loop context ("designed to run under /loop") in the description so users searching for automated/interval review upkeep find it.

DimensionReasoningScore

Specificity

The description enumerates concrete, distinct actions — 'ships if needed, keeps it mergeable against staging, re-triggers both Greptile and cubic, fixes real findings, replies to and resolves every thread' — with a crisp success criterion ('Greptile 5/5, zero open threads'). This matches the comprehensive-coverage anchor; it goes beyond the 4-anchor's 'minor gaps' by covering ship, sync, trigger, fix, reply, resolve, and loop phases.

5 / 5

Completeness

The 'what' is explicit and thorough, but there is no 'Use when...' clause or equivalent explicit trigger guidance in the description. Per the judging guidelines, a missing 'Use when' clause caps completeness at 3; the when-intent ('loops until clean') describes behavior, not when to invoke the skill.

3 / 5

Trigger Term Quality

It includes natural, sayable terms like 'PR', 'clean review', 'review', 'open threads', and bot names (Greptile, cubic). However, the phrasages users would most naturally say for this skill — e.g. 'babysit this PR' or 'keep working the reviews until it's clean' — appear only in the body, not the description, so keyword coverage is good but not comprehensive.

4 / 5

Distinctiveness Conflict Risk

It carves a clear niche — driving a PR through Greptile/cubic bot review to a 5/5, zero-thread clean state against staging — with triggers unlikely to fire for unrelated skills. Not the 4 anchor, because naming both bots and the specific 5/5 criterion leaves minimal overlap risk even with adjacent ship/review skills.

5 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
simstudioai/sim
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.