CtrlK
BlogDocsLog inGet started
Tessl Logo

devtools-verification

MANDATORY: Activate this skill ANY TIME you need to build the project, run tests, or verify code health in DevTools. You MUST use this skill before executing commands like npm test, npm run build, autoninja, or linters, as it contains critical, repository-specific instructions on how to correctly format these commands, filter test runs, and interpret failures.

65

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

87%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an exemplar of token efficiency and actionability — a short, fully executable command reference with valuable repo-specific constraints like the Linux-only goldens rule. Its one real weakness is workflow clarity: verification steps are not sequenced with validation checkpoints, and the destructive --apply-goldens operation lacks a verify-before-apply guard.

Suggestions

Sequence the verification workflow explicitly (e.g. 1. Run affected tests 2. If failing, iterate until green 3. Build with autoninja 4. Run npm run lint 5. Run git cl presubmit -u) so the order and checkpoints are unambiguous.

Add a validation step before --apply-goldens (e.g. inspect the image diff first, confirm the CL's failure is a legitimate golden change, then apply) since it overwrites expected golden files.

Add one line on interpreting test failures (where output lands, how to isolate a flake vs. a real regression) to close the feedback loop the description promises.

DimensionReasoningScore

Conciseness

The body is lean and imperative throughout — every line is either an executable command ("npm run test -- <FILEPATH>", "autoninja -C out/ Default") or a repo-specific rule ("Never generate, update, or commit golden PNGs from macOS"); no concept Claude already knows is explained. This matches the 'lean and efficient; assumes Claude's competence; every token earns its place' anchor, and there is nothing to trim that would justify a 4.

5 / 5

Actionability

All guidance is copy-paste-ready: exact commands with concrete paths and flags ("npm run lint -- <PATH>", "git cl presubmit -u") and three complete `get_cl_test_results.py` invocations showing flag usage including --test-filter and --apply-goldens. This matches 'fully executable; copy-paste ready code or commands; specific examples cover the common cases'; the only reason it would not is that some placeholder values (<CL_NUMBER>) remain unspecified, which is appropriate for a parameterized command.

5 / 5

Workflow Clarity

The verification workflow is only implicitly sequenced via Best practices ("Run tests often", "Periodically build", "Run git cl presubmit -u at the end") with no explicit checkpoints or failure-handling guidance, matching 'steps listed but validation gaps; sequence present but checkpoints missing or implicit'. It is capped at 3 by the rubric's destructive/batch rule: the --apply-goldens flag overwrites golden files directly with no validate-then-apply step, and there is no test-fails -> fix -> re-run feedback loop.

3 / 5

Progressive Disclosure

The skill is under 50 lines, self-contained (no references/scripts/assets bundle files exist), and organized into clear one-level sections (Testing, Building & compiling, Linting, Fetching Tryjob Results, Best practices) with a tight overview-and-sections structure. Per the rubric's simple-skill note, this earns a 5; the only external path cited (scripts/tools/get_cl_test_results.py) is a repo command, not a bundle reference.

5 / 5

Total

18

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly communicates what the skill covers and when to activate it, with strong repo-specific trigger terms (npm test, autoninja, linters). Its main weaknesses are the pervasive second-person voice and MANDATORY/You MUST over-claiming, which the rubric penalizes, and omission of the tryjob-results fetching capability that makes up a third of the skill body.

Suggestions

Rewrite in third person ('Contains repository-specific instructions for building, testing, and verifying code health in DevTools...') to comply with the rubric's voice rule and drop the 'MANDATORY/ANY TIME/You MUST' emphasis in favor of a concrete 'Use when...' trigger clause.

Mention the remaining capabilities so coverage matches the body: fetching tryjob build statuses, failure logs, and screenshot goldens for a CL.

Add natural trigger synonyms users would say, e.g. 'presubmit', 'tryjobs', 'CI results', 'screenshot goldens', so the description matches on those queries too.

DimensionReasoningScore

Specificity

The description lists several concrete actions ("build the project, run tests, or verify code health", "format these commands, filter test runs, and interpret failures") with a minor coverage gap (the tryjob-results fetching capability is absent), which is a base 4; the rubric's third-person rule penalizes the pervasive second person ("you need", "You MUST use this skill") by 1, landing at 3. It is not a 2 because the actions named are concrete and domain-specific rather than generic, and not a 4 because the second-person voice is used throughout rather than incidentally.

3 / 5

Completeness

Both what ("critical, repository-specific instructions on how to correctly format these commands, filter test runs, and interpret failures") and when ("ANY TIME you need to build the project, run tests, or verify code health in DevTools") are explicitly present, matching the anchor 'has both what and when; when could be more explicit or specific'. Not a 5 because the 'what' is stated meta-descriptively (the skill 'contains instructions') rather than as concrete capability phrases, and the over-claiming 'MANDATORY / ANY TIME / You MUST' framing is fluff the rubric penalizes rather than rewards.

4 / 5

Trigger Term Quality

Good coverage of natural trigger phrases users would actually say ("build the project", "run tests", "verify code health", "npm test", "npm run build", "autoninja", "linters"), matching the 'good keyword coverage; a few natural terms missing' anchor. It falls short of 5 because common variations for this skill's scope are missing: 'presubmit', 'tryjob', 'CI results', and 'goldens/screenshot tests' never appear.

4 / 5

Distinctiveness Conflict Risk

The DevTools scoping and Chromium-specific tool names ("autoninja", "npm test" in a DevTools repo context) establish a mostly distinct niche, matching 'mostly distinct; minor overlap risk'. Not a 5 because the generic verbs 'build the project, run tests, or verify code health' could match any build/test skill in another repository, keeping some overlap risk with general verification skills.

4 / 5

Total

15

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 4 missing, 4 deeper-than-1-level

Warning

Total

15

/

16

Passed

Repository
ChromeDevTools/devtools-frontend
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.