Content
92%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
An exemplary lean, fully executable skill body: every instruction is a runnable gh command, the loop restart rule keeps state changes coherent, and validation checkpoints plus fix/retry feedback loops are explicit. The single notable gap is that the genuine-failure-vs-flake decision — the skill's own stated purpose — has no criteria, heuristics, or examples to guide it.
Suggestions
Add 2–3 concrete heuristics or examples for distinguishing a genuine CI failure from a flake (e.g. failure reproduces locally / error is in changed code vs. timeout, network, or test-unrelated job), since this decision drives which branch of the workflow runs.
Clarify how to obtain the `<run-id>` used in `gh run view` (e.g. from `gh pr checks --watch=false` output) so the CI-failure inspection path is fully copy-paste ready.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Every line is a step, an executable gh command, or a clearly signaled delegation to a sibling skill; there is no concept explanation or padding at all, matching the lean-and-efficient anchor. | 5 / 5 |
Actionability | Commands are copy-paste ready and complete (`gh pr view --json ...`, `gh pr checks --watch --interval 30`, `gh pr merge --merge --delete-branch`, `gh run view <run-id> --log-failed`), but the pivotal decision — distinguishing a genuine CI failure from a flake — gets only "inspect the failed jobs and logs before deciding what to do" with no criteria or heuristics, leaving a common case without executable guidance. Not 3 because everything that is given is fully executable. | 4 / 5 |
Workflow Clarity | The loop has explicit validation checkpoints (eligibility/blocker check in step 2, "all required checks pass and the PR is mergeable" in step 5) and genuine feedback loops — genuine failure → fix → push → restart, flake → rerun → restart — with an explicit "start from the top after every state change" rule. Not 4 because the checkpoints are explicit rather than implicit. | 5 / 5 |
Progressive Disclosure | No bundle files exist or are needed at this size; the two sections are well-organized and the only references — [`pr-ready`](../pr-ready/SKILL.md) and [`flake`](../flake/SKILL.md) — are sibling skills referenced one level deep with clear link signals, satisfying the well-organized-sections bar for a short skill with no separate-file content. | 5 / 5 |
Total | 19 / 20 Passed |