CtrlK
BlogDocsLog inGet started
Tessl Logo

grok

Delegate coding to xAI Grok Build CLI (features, PRs).

54

Quality

62%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/autonomous-ai-agents/grok/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

76%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-structured orchestration guide: executable commands throughout, clear mode distinctions (headless vs PTY), and a strong pitfalls section. The main gaps are the missing verification loop in the batch worktree workflow and some redundancy between the Rules, Pitfalls, and Prerequisites sections.

Suggestions

Add an explicit verification step to the Parallel Issue Fixing workflow — e.g. after each background fix completes, run the tests or diff the worktree ('git diff --stat') before pushing and opening the PR.

Consolidate the '--no-auto-update' and headless-preference guidance into one place (the Rules or Pitfalls section) instead of repeating it in nearly every section.

Consider moving the flag reference table, config.toml details, and TUI command list into a one-level-deep reference file to slim the main SKILL.md.

DimensionReasoningScore

Conciseness

The body is dense and imperative with no teaching of concepts Claude already knows (tmux/git usage is referenced, not re-explained). Minor over-repetition keeps it from a 5: '--no-auto-update' is stated 5+ times across sections, and the 'Rules for Hermes Agents' section substantially restates the Pitfalls and Prerequisites sections.

4 / 5

Actionability

Every pattern is copy-paste-ready executable commands (headless invocations, tmux launch/monitor/teardown, background mode with process polling, session UUID create/resume/continue, worktree parallel fixing, PR review and posting). Specific flags, timeouts, workdirs, and commands cover the common orchestration cases; nothing is pseudocode.

5 / 5

Workflow Clarity

Sequences are clearly ordered and the audit pattern has explicit validation ('Verify the first lines with read_file() before overwriting'), but the batch Parallel Issue Fixing workflow (launch fixes → monitor → push → open PRs) has no verification step that Grok's changes pass tests or are sane before pushing and creating PRs. Per the rubric guideline, a batch operation without validation caps workflow clarity at 3.

3 / 5

Progressive Disclosure

No bundle files exist, so everything is inline, but the file is well-organized with clear section headers, flag/subcommand tables, and no nested or dead references. It is not 5 because at ~300 lines the flag reference, config, and pitfall detail could arguably live in separate one-level-deep reference files rather than fully inline.

4 / 5

Total

16

/

20

Passed

Description

48%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise and names a distinct tool, but it is too terse to function well as a trigger: it lacks any 'use when' guidance and gives only minimal, generic action language. It reads more like a title than a capability description.

Suggestions

Add an explicit 'Use when...' clause, e.g. 'Use when delegating coding tasks (building features, refactoring, PR reviews) to Grok rather than Codex or Claude Code.'

Replace the generic 'Delegate coding' with the concrete capabilities: 'Builds features, refactors code, reviews PRs, and fixes issues in a repo via headless or interactive sessions.'

Include natural trigger synonyms users would say — 'Grok', 'Grok CLI', 'code review', 'refactor' — to improve discovery over sibling skills.

DimensionReasoningScore

Specificity

Names the domain (xAI Grok Build CLI) and one generic action ('Delegate coding'), with only terse parenthetical hints '(features, PRs)' — no description of what Grok actually does (build features, refactor, review PRs). This matches 'Names the domain but actions are minimal or generic' ('Processes PDF files') rather than the 1-2-concrete-actions anchor above.

2 / 5

Completeness

The 'what' is clear (delegate coding tasks to the Grok Build CLI) but there is no 'Use when...' clause or equivalent trigger guidance — 'when' is only weakly implied by '(features, PRs)'. Per the guideline, a missing use-when clause caps completeness at 3; it is not 2 because the 'what' is concrete, and not 4 because the 'when' is absent rather than merely imprecise.

3 / 5

Trigger Term Quality

Contains some relevant keywords ('coding', 'features', 'PRs', 'Grok') a user might say, but misses common natural variations and synonyms such as 'refactor', 'code review', 'fix bugs', or 'Grok CLI'. Coverage is partial, matching 'Some relevant keywords but missing common variations or synonyms' rather than the good-coverage anchor.

3 / 5

Distinctiveness Conflict Risk

The named tool ('xAI Grok Build CLI') gives a clear niche with distinct triggers (Grok), so it is mostly distinct with only minor overlap risk — the generic 'coding' and 'PRs' terms could collide with sibling skills like codex/claude-code. Not 5 because it doesn't fully disambiguate from those near-identical siblings.

4 / 5

Total

12

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.