CtrlK
BlogDocsLog inGet started
Tessl Logo

flake

Track Remotion CI flakes in issue

72

Quality

87%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

100%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This is a well-crafted operational skill: it gives exact executable gh commands for every input context, precise classification and signature rules, and a validated edit-rerun-report workflow with explicit error-recovery guidance. It is lean, assumes competence, and needs no bundle files.

DimensionReasoningScore

Conciseness

The body is lean and operational throughout — exact gh commands with flags and --json field lists, concrete signature examples, and a table format spec. The few explanatory sentences (e.g. "Prefer `gh` because the GitHub app may not expose all Actions logs") are non-obvious operational knowledge Claude does not already have, so every token earns its place.

5 / 5

Actionability

Fully executable, copy-paste-ready gh commands cover every common entry point (PR given, run/job URL given, no context, rerun attempts), and the skill provides concrete good signature examples plus an exact tracker table format. This matches the 'fully executable; covers the common cases' anchor.

5 / 5

Workflow Clarity

The sequence (find failure → classify → update tracker → rerun → report) is clearly ordered with explicit validation checkpoints: inspect the original attempt before classifying, 'Verify after editing' with a re-read command, and a recovery loop when `gh run watch` exits non-zero after cancellation. The batch issue-body edit does include verification, so the destructive/batch cap does not apply.

5 / 5

Progressive Disclosure

No bundle files exist and the ~145-line body is cohesively organized under clear section headers (Goal, Find The Failure, Classify A Flake, Signature Rules, Update The Tracker, Rerun, Report Back) with all content directly used on every invocation — nothing belongs in a separate file, and navigation is easy.

5 / 5

Total

20

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concrete, third-person, and highly distinctive, listing all four of the skill's actions comprehensively. Its only real weakness is the missing 'Use when...' trigger clause, which caps completeness and leaves some natural trigger synonyms (retry, intermittent) uncovered.

Suggestions

Add an explicit trigger clause to the description, e.g. "Use when a Remotion CI check fails and looks flaky, or when asked to rerun flaky or failed GitHub Actions jobs" — the trigger currently lives only in the body.

Include natural trigger synonyms such as "retry", "re-run", or "intermittent test failures" so users phrasing the need differently still match this skill.

DimensionReasoningScore

Specificity

"Track Remotion CI flakes in issue #8375", "increment repeated signatures", "discover failed PR checks when no PR is given", and "rerun flaky GitHub Actions jobs" list multiple concrete actions in third-person voice with comprehensive coverage of the skill's behavior. It clearly fits the comprehensive-coverage anchor rather than the score-4 anchor, which expects minor coverage gaps.

5 / 5

Completeness

The "what" is explicit and concrete, but there is no "Use when..." clause or equivalent trigger guidance in the description — the trigger ("when a Remotion CI check fails and looks flaky") exists only in the body. Per the rubric guideline, a missing 'when' caps completeness at 3.

3 / 5

Trigger Term Quality

Good natural keyword coverage — "Remotion CI flakes", "flaky GitHub Actions jobs", "PR checks", "issue #8375" — but common variations users might say ("retry failed checks", "intermittent test failures", "re-run failed jobs") are missing. This sits between the good-coverage (4) and comprehensive-synonym (5) anchors, noticeably above the midpoint.

4 / 5

Distinctiveness Conflict Risk

"Remotion CI flakes in issue #8375" names a highly specific niche (a particular repo, issue number, and CI context), so it is clearly distinguishable from any other skill with minimal conflict risk — a clean match for the score-5 anchor.

5 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
remotion-dev/remotion
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.