CtrlK
BlogDocsLog inGet started
Tessl Logo

cherry-pr-test

Test Cherry Studio PRs by resolving and checking out a PR, statically inspecting its changes, running interactive UI tests against a safely tracked Electron instance through CDP, producing a structured report, cleaning up only the owned test instance, and restoring the original branch.

67

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-engineered operational skill: lean, tightly written, with a fully sequenced workflow and genuine safety validation around checkout, instance reuse, and cleanup. The main weaknesses are partial dependence on an external cross-skill reference (not bundled here) for the Electron/CDP mechanics and an inline report template that could live in a reference file.

Suggestions

Bundle the Electron instance procedure (or a local copy/pointer card) under references/ so the skill is self-contained; as written, step 2 and 3 cannot be executed if ../cherry-electron-dev/references/electron-instance.md is absent.

Move the ~35-line report template to references/report-template.md and keep a one-line pointer plus required fields in SKILL.md to reduce inline bulk.

Add one concrete example of the static-analysis checks (e.g., the grep or file-listing command used to find blocked v1/v2-refactor files) so the 'Also check' list is executable rather than advisory.

DimensionReasoningScore

Conciseness

The body is lean and imperative throughout — 'Require authenticated gh, pnpm', 'Do not use broad process or port cleanup', 'Record the exact checked-out HEAD' — with zero explanation of concepts Claude already knows (no Electron/CDP/git tutorials). The inline report template is task content, not padding; every section instructs rather than describes.

5 / 5

Actionability

Concrete, executable commands appear at each step — 'gh pr list --repo CherryHQ/cherry-studio --state open --limit 10 --json ...', 'gh pr checkout <NUMBER>', 'pnpm typecheck', 'mkdir -p /tmp/pr-<NUMBER>' — plus a complete copy-paste report template. Minor gaps: Electron launch and the CDP controller are delegated to the shared reference with no local command or fallback if that reference is unavailable, leaving step 2/3 partially non-executable from this file alone.

4 / 5

Workflow Clarity

A clearly sequenced 5-step workflow (resolve/inspect → analyze/start → test → clean up/restore → report) with explicit validation checkpoints for the risky operations: 'Record the current branch for restoration', 'Never discard local changes. Stop and ask if checkout would overwrite them', 'Reuse only an instance ... whose recorded launch HEAD equals the checked-out PR HEAD', and a branch-restore fallback ('If it no longer exists, resolve the repository default branch and report the fallback before switching'). This matches the top anchor: clear sequence, explicit validation, error-recovery paths.

5 / 5

Progressive Disclosure

Good single-file structure with clear sections and a well-signaled, one-level-deep delegation to '../cherry-electron-dev/references/electron-instance.md' ('read ... and select its ephemeral policy'). Two organization gaps keep it below 5: the ~35-line report template is inlined in SKILL.md where a reference file would fit, and the skill's single critical dependency lives outside its own bundle — no references/, scripts/, or assets/ exist here, so the Electron instance procedure is not locally discoverable.

4 / 5

Total

18

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, highly specific description that enumerates the complete workflow in concrete terms with minimal conflict risk. Its one material weakness is the complete absence of 'Use when ...' trigger guidance, which both caps completeness and leaves invocation conditions implicit.

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user asks to test, verify, or check a Cherry Studio pull request (PR #, PR URL, or the latest).' This directly lifts completeness from 3.

Include the synonym 'pull request' alongside 'PR' so users who spell it out still match the description.

Consider trimming 'safely tracked' and 'only the owned' qualifiers — internal policy detail that adds tokens without improving trigger matching.

DimensionReasoningScore

Specificity

The description lists multiple specific concrete actions covering the full workflow: 'resolving and checking out a PR', 'statically inspecting its changes', 'running interactive UI tests ... through CDP', 'producing a structured report', 'cleaning up only the owned test instance', 'restoring the original branch'. Not below 4: coverage is comprehensive end-to-end, not just several actions with minor gaps.

5 / 5

Completeness

The 'what' is clear and detailed (checkout, static inspection, UI tests, report, cleanup, restore), but there is no 'Use when ...' clause or any equivalent explicit trigger guidance — the when is entirely implicit. Per the rubric guideline, a missing 'Use when ...' clause caps completeness at 3. Not 4: the when is not weakly present; it is absent from the description.

3 / 5

Trigger Term Quality

Good natural keyword coverage — 'Test', 'PRs', 'PR', 'Cherry Studio', 'Electron', 'report' — which a user would plausibly say ('test PR 123'). Not 5: common synonyms like 'pull request', 'PR review', or 'test this PR's changes' are absent; not 3: the core terms are the phrases users actually use, with only a few variations missing.

4 / 5

Distinctiveness Conflict Risk

Clear niche ('Test Cherry Studio PRs') with distinct triggers — bounded PR testing with checkout/restore semantics — making it unlikely to fire for general development or other projects. Not 4: the scope is a single named repository and a single bounded task, so overlap risk with related skills (e.g., ongoing implementation/debugging in a checkout) is minimal.

5 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 missing, 1 suspicious

Warning

Total

15

/

16

Passed

Repository
CherryHQ/cherry-studio
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.