CtrlK
BlogDocsLog inGet started
Tessl Logo

coding-agent

Run Codex CLI, Claude Code, OpenCode, or Pi Coding Agent via background process for programmatic control.

52

Quality

59%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

Fix and improve this skill with Tessl

tessl review fix ./skills/coding-agent/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

57%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This is a highly actionable, bash-first skill with concrete executable commands for every major workflow, undermined by padding, repetition, dated notes, and a monolithic single-file structure. The biggest structural gaps are the absence of validation checkpoints in batch/destructive workflows and the lack of any reference-file split.

Suggestions

Add validation steps before posting/creating PRs in the batch review and parallel issue-fixing workflows (e.g. verify tests pass or diff reviews cleanly before `gh pr create`/`gh pr comment`).

Split per-agent sections (Codex flags, Pi provider options) and the advanced batch/worktree playbooks into reference files (e.g. references/codex.md, references/parallel-workflows.md), keeping SKILL.md a concise overview.

Trim repetition and fluff: state the PTY requirement once, move dated items ("Learnings (Jan 2026)", "PR #584") into a notes section, and cut casual asides; fix the missing `exec` in the `codex --yolo` background example.

DimensionReasoningScore

Conciseness

The body is mostly efficient — parameter/action tables and command examples carry real weight — but includes unnecessary padding: repeated PTY warnings ("Always use pty:true", the ⚠️ section, Rules #1, and a Learnings bullet all restate it), casual asides ("like your soul.md 😅", "Parallel army!", "Sass works"), and time-sensitive details ("Learnings (Jan 2026)", "PR #584, merged Jan 2026", "gpt-5.2-codex default") that are not isolated in a notes/deprecated section. This matches anchor 3 (mostly efficient, some unnecessary explanation, could be tightened) rather than 4, which requires only minor trimmable instances.

3 / 5

Actionability

Nearly all guidance is concrete, executable bash with the exact harness syntax ("bash pty:true workdir:~/project background:true command:...", "process action:log sessionId:XXX") covering one-shots, background monitoring, PR review, batch reviews, and worktree workflows. Minor gaps keep it from 5: placeholder templates ("gh pr comment <PR#> --body \"<review content>\"", "Fix issue #78: <description>") and an inconsistent invocation ("codex --yolo 'Refactor the auth module'" at line 119 lacks the "exec" subcommand used everywhere else).

4 / 5

Workflow Clarity

Multi-step sequences are clearly laid out (start → monitor → poll → submit → kill; the 5-step numbered worktree workflow with cleanup), but batch operations lack validation checkpoints: the batch PR review flow goes straight from "Monitor all" to "Post results to GitHub" without checking output quality, and the parallel issue-fixing flow creates PRs after fixes without verifying tests/builds pass. Per the rubric, missing validation in batch workflows caps workflow_clarity at 3 — it is above 2 because sequences are well-defined with monitoring steps, but cannot reach 4 without outcome verification.

3 / 5

Progressive Disclosure

The single SKILL.md (~285 lines) has clear section headers but inlines content that belongs in separate reference files — per-tool guides (Codex flags, Pi providers), batch PR review patterns, and the parallel worktree playbook could each be one-level-deep references, keeping SKILL.md as an overview. No bundle files exist (no references/, scripts/, or assets/), so the structure is a well-headed but monolithic inline document, matching anchor 3 (some structure, content that should be separate is inline) rather than 4 (most content appropriately placed with references).

3 / 5

Total

13

/

20

Passed

Description

61%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concrete and tool-specific with good natural trigger keywords, but it omits any "when to use" guidance and describes only a thin action set. Adding an explicit use-when clause and common trigger variations would lift it substantially.

Suggestions

Add a 'Use when...' clause, e.g. 'Use when the user asks to run, spawn, or delegate work to Codex, Claude Code, OpenCode, or Pi in the background.'

Broaden trigger terms with natural variations users would say: 'coding agent', 'run codex in the background', 'delegate to an agent', 'parallel agent workflows'.

Expand the action list beyond 'run' to the concrete capabilities the skill provides (monitor sessions, send input, batch/parallel runs, PR review) so specificity approaches comprehensive coverage.

DimensionReasoningScore

Specificity

"Run Codex CLI, Claude Code, OpenCode, or Pi Coding Agent via background process for programmatic control" names the domain (four specific CLIs) and one concrete mechanism (background-process execution), but the action set is minimal — just "run ... for programmatic control" — leaving coverage of what programmatic control entails (spawning, monitoring, sending input) implicit. It sits between anchor 3 (1-2 concrete actions, not comprehensive) and anchor 4 (several specific actions), and the thin action list keeps it at 3 rather than 4.

3 / 5

Completeness

The description has a clear "what" (run these agents via background process for programmatic control) but no "Use when..." or equivalent trigger guidance, which per the rubric caps completeness at 3. It is not a 4 because the "when" is entirely missing rather than just under-specified, and not a 2 because the "what" is concrete and unambiguous.

3 / 5

Trigger Term Quality

Product names like "Codex CLI", "Claude Code", "OpenCode", and "Pi Coding Agent" are exactly what a user would say when needing this skill, plus "background process" for the mode. However, common natural variations such as "coding agent", "run X in the background", or "spawn an agent" are absent, matching anchor 4 (good coverage, a few natural terms missing) rather than 5 (comprehensive synonyms/extensions).

4 / 5

Distinctiveness Conflict Risk

Naming four specific coding-agent CLIs carves out a clear niche distinct from general coding skills, but "Claude Code" as a trigger term risks overlap with ordinary coding requests in a Claude environment where a user might say "use claude code" meaning any coding task. This matches anchor 4 (mostly distinct, minor overlap risk with closely related skills) rather than 5 (minimal conflict risk).

4 / 5

Total

14

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

Total

14

/

16

Passed

Repository
Bitterbot-AI/bitterbot-desktop
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.