CtrlK
BlogDocsLog inGet started
Tessl Logo

netlify-agent-runner

Run AI agent tasks remotely on Netlify using Claude, Codex, or Gemini. Use when the user wants to run an AI agent on their site, get a second opinion from another model, or delegate development tasks to run remotely against their repo.

66

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./codex/skills/netlify-agent-runner/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Highly actionable content with copy-paste-ready commands, a complete options reference, and a workflow with explicit terminal-status checkpoints and error handling. Its main weakness is repetition — the remote/async/polling model is restated multiple times — and a long use-case table that inflates token cost without adding proportional guidance.

Suggestions

Consolidate the remote/asynchronous/polling explanation: state it once in 'How Agent Tasks Run' and have the 'Using as an Agent' section link to it instead of restating the same three points.

Trim the 18-row Use Cases table to a representative handful of categories (or move it to a reference file) to reduce token cost.

Merge 'Prerequisites' credit-limit note with the failure guidance ('report the exact error and stop') so account/plan blockers and command failures are handled in one place.

DimensionReasoningScore

Conciseness

Mostly efficient — no explanation of concepts Claude already knows — but the remote/asynchronous/polling points are repeated in three places ("How Agent Tasks Run" bullets, the Typical workflow step 2, and again in "Using as an Agent" lines about running remotely, async polling, and writing standalone prompts), and the 18-row Use Cases table adds bulk a tighter examples list could cover. Not a 2: there is no padding of known concepts; not a 4: the duplication is a real tightening opportunity.

3 / 5

Actionability

Fully executable, copy-paste-ready commands throughout — `netlify agents:create "Add a contact form"`, `netlify agents:show <task-id>`, `netlify agents:stop <task-id>` — with a complete options table (`-a`, `-p`, `-b`, `-m`, `--project`, `--json`) and concrete permission-request examples like `netlify agents:create -p "<the real prompt>" -a codex`. Matches the anchor-5 standard of executable commands covering the common cases.

5 / 5

Workflow Clarity

The Typical workflow is a clearly sequenced loop with an explicit validation checkpoint: create (capture task ID with `--json`), poll `agents:show` until a terminal status (`done`, `error`, `cancelled`) — 'Keep polling until the status is one of those last three before you act on the results' — then review or inspect the failure. Error-recovery guidance (report the exact error and stop; blocked runs on missing credits surfaced to the user) plus the explicit approval gate complete the feedback structure. Not a 4: checkpoints are explicit, not implicit.

5 / 5

Progressive Disclosure

No bundle files exist, and the single-file body is well organized with clear sections, an internal anchor link ([How Agent Tasks Run](#how-agent-tasks-run)), and no buried or nested references. Not a 5: at ~160 lines with no external files, content like the 18-row Use Cases table and the 'Using as an Agent' section are candidates to split into references; not a 3: the structure and navigation are genuinely good.

4 / 5

Total

17

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit 'Use when' clause and concrete natural-language triggers. Its main weakness is that the capability statement covers a single action, leaving the breadth of what the skill does (create, poll, stop, branch targeting) to the body.

DimensionReasoningScore

Specificity

"Run AI agent tasks remotely on Netlify using Claude, Codex, or Gemini" names the domain and one concrete action with specific agent types, but does not list several distinct capabilities (create/list/stop tasks, polling, branch targeting are only in the body). Not a 4: the 'what' is a single action, so coverage has gaps; not a 2: the action named is concrete, not generic.

3 / 5

Completeness

Explicitly answers both: what — "Run AI agent tasks remotely on Netlify using Claude, Codex, or Gemini"; when — "Use when the user wants to run an AI agent on their site, get a second opinion from another model, or delegate development tasks to run remotely against their repo" with concrete trigger phrases. Matches the anchor-5 example's structure exactly.

5 / 5

Trigger Term Quality

Natural phrases users would say are present — "run an AI agent on their site", "get a second opinion from another model", "delegate development tasks", "against their repo". Not a 5: common synonyms and variations (e.g. 'AI runner', 'remote agent', 'ask another model') are missing; not a 3: coverage goes well beyond a single generic keyword.

4 / 5

Distinctiveness Conflict Risk

The Netlify-specific, remote/second-opinion framing is a clear niche with distinct triggers. Not a 5: "delegate development tasks" and "second opinion from another model" could also plausibly trigger generic agent-delegation or cross-validation skills; not a 3: the Netlify + remote-repo framing is well differentiated.

4 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
netlify/context-and-tools
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.