CtrlK
BlogDocsLog inGet started
Tessl Logo

blackbox

Delegate coding tasks to the Blackbox AI multi-model CLI.

56

Quality

65%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/autonomous-ai-agents/blackbox/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-organized, highly actionable reference: every scenario comes with copy-paste commands and the background workflow includes an explicit monitor/interact/kill loop. The only real gaps are mild redundancy between sections and thin coverage of the flagship multi-model workflow.

Suggestions

Consolidate the VLM mode list (currently in both Key Flags and Vision Support) and the pty guidance (currently in both Prerequisites and Rules) to reclaim tokens.

Add one concrete example command or configure snippet for Multi-Model Mode, since it is presented as the tool's unique feature but has no executable guidance.

Turn Rule 7 into an actual validation step with a command (e.g. 'blackbox --version || npm install -g @blackbox_ai/blackbox-cli') before dispatching tasks.

DimensionReasoningScore

Conciseness

The body is lean and command-first: concrete terminal invocations, compact command/flag tables, and no explanation of concepts Claude already knows. Minor duplication (pty=true stated in both Prerequisites and Rules; VLM modes appear in both the Key Flags table and the Vision Support section) keeps it below anchor 5.

4 / 5

Actionability

Guidance is fully executable throughout: copy-paste-ready terminal(...) commands for one-shot, background, checkpoint-resume, PR review, and parallel scenarios, plus concrete process() monitoring calls and two complete reference tables. Specific examples cover the common cases.

5 / 5

Workflow Clarity

Scenarios are clearly sequenced (prerequisites, one-shot, background with start/poll/log/submit/kill, checkpoints, parallel) and the Rules section adds monitoring and reporting steps. Minor validation gaps remain — e.g. Rule 7 says to verify the CLI is installed but no concrete check command or recovery step is given.

4 / 5

Progressive Disclosure

The skill is a self-contained single file with no bundle directories and well-organized, clearly headed sections that are easy to navigate. Minor gaps: content is mildly duplicated across sections (flags vs. vision modes) and Multi-Model Mode — the headline feature — lacks a concrete command example, which would suit either inline tightening or a separate reference.

4 / 5

Total

17

/

20

Passed

Description

48%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description has a clear 'what' and a distinctive product name, but it is minimal: one generic action, no trigger/'Use when' guidance, and no concrete capability enumeration. It would rarely surface when a user needs it unless they already mentioned Blackbox by name.

Suggestions

Add an explicit 'when' clause, e.g. 'Use when the user asks to delegate a coding task to Blackbox, run a task through multiple models, or wants a second AI to implement something.'

Enumerate 2-3 concrete capabilities so the description doubles as a capability summary, e.g. 'Run one-shot prompts, long background tasks with monitoring, checkpoint/resume, and parallel multi-model execution.'

Include natural trigger synonyms users would actually say ('blackbox', 'another AI', 'multi-model', 'second opinion on this code') to improve trigger term coverage.

DimensionReasoningScore

Specificity

The description names the domain ('coding tasks', 'Blackbox AI multi-model CLI') but offers only one generic action — 'Delegate coding tasks' — with no enumeration of concrete capabilities. It does not reach anchor 3, which requires 1-2 distinct concrete actions.

2 / 5

Completeness

The 'what' is clear (delegate coding tasks to the Blackbox CLI) but there is no 'when' clause or equivalent trigger guidance anywhere in the description. Per the judging guidelines, a missing 'Use when...' clause caps completeness at 3.

3 / 5

Trigger Term Quality

Relevant keywords are present ('Delegate coding tasks', 'Blackbox AI', 'multi-model CLI') but common natural variations users would say (e.g. 'use blackbox', 'run this with another model') are missing. Anchor 3 is the best fit; not 4 because coverage lacks synonyms.

3 / 5

Distinctiveness Conflict Risk

Naming the specific product ('Blackbox AI multi-model CLI') gives it a clear niche, but 'Delegate coding tasks' overlaps with closely related agent CLI skills (claude-code, codex are listed as related). Mostly distinct with minor overlap risk — anchor 4.

4 / 5

Total

12

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.