Unified Minions skill for both deterministic shell jobs and LLM subagent orchestration. Replaces the older `gbrain-jobs` routing intent. Use when: submitting gbrain jobs, shell/background tasks, spawning subagents, checking progress, steering running work, pausing/resuming, parallel fan-out. One durable, observable, steerable queue interface. Also carries the durable-execution doctrine for any operation expected to exceed ~2 minutes: capability ladder, deadman checks that verify the result was reported, and content-addressed stage checkpoints for expensive pipelines.
64
78%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Low
Low-risk findings worth noting
Fix and improve this skill with Tessl
tessl review fix ./skills/minion-orchestrator/SKILL.mdMinions is a Postgres-native job queue for durable, observable background work. This single skill handles two lanes:
gbrain jobs submit shell ...)gbrain agent run ...)When to route to Minions: durable, observable work that must survive restarts,
fan out across many parallel tasks, or persist across sessions. Routing policy
is defined in skills/conventions/subagent-routing.md — the project default is
pain_triggered (native subagents first, Minions after specific pain signals
fire); Mode A (all-through-Minions) is opt-in.
Guarantees:
Durable-execution doctrine (routing convention the agent follows — nothing mechanically enforces it; see "Durable execution" below):
gbrain jobs list --status active and recent
completions for overdue work whose result never reached the user.gbrain jobs get <id>.| Condition | Action |
|---|---|
| User asks for deterministic command/script run | Shell job (CLI: gbrain jobs submit shell ...) |
| User asks to "run in minions" + explicit command/argv | Shell job (CLI, --params with cmd or argv) |
| User asks for research/reasoning/iterative agent | Subagent job (CLI: gbrain agent run) |
| User asks to steer/pause/resume an agent | Subagent job lifecycle tools (MCP-callable) |
| Single simple operation under ~30s | Consider inline execution first |
| Needs restart durability/observability | Submit as Minion job |
| Operation expected to exceed ~2 minutes | Route through the Durable execution ladder (below) |
| Parallel work (2+ streams) | gbrain agent run --fanout-manifest or parent + child subagents |
If intent is ambiguous, ask one clarification: "Do you want a deterministic shell command job, or an LLM agent job?"
Use for reproducible command execution, ETL steps, cron work, and scriptable tasks where no LLM reasoning loop is needed.
GBRAIN_ALLOW_SHELL_JOBS=1 must be set on the worker environment.
Without it, the shell handler refuses to register and submissions sit in
waiting silently. Gate lives in src/core/minions/handlers/shell.ts.GBRAIN_ALLOW_SHELL_JOBS=1 authorizes arbitrary
command execution on the worker. On a shared queue, this is a remote code
execution surface. Treat as privileged infrastructure authorization.gbrain jobs work runs a persistent worker that
claims and executes jobs from the queue.gbrain jobs submit ... --follow runs inline.
The daemon mode is not available on PGLite (exclusive file lock). See
docs/guides/minions-shell-jobs.md.submit_job name="shell"
over MCP throws an OperationError with code permission_denied ("'shell'
jobs cannot be submitted over MCP") because shell is in PROTECTED_JOB_NAMES.
Agents CAN observe shell jobs via get_job / list_jobs / get_job_progress
(not protected), but cannot submit them. Operator or autopilot submits;
agent observes.gbrain jobs stats (CLI) to
confirm the worker is registered and consuming the queue.Shell jobs take their command via --params as a JSON object with cmd (string)
or argv (array), plus cwd and optional env.
Command string form:
gbrain jobs submit shell --params '{"cmd":"echo hello","cwd":"/abs/path"}'Argv form (no shell expansion):
gbrain jobs submit shell --params '{"argv":["bash","-lc","echo hello"],"cwd":"/abs/path"}'Inline execution on PGLite or any one-shot deployment:
gbrain jobs submit shell --params '{"cmd":"echo hello","cwd":"/tmp"}' --followQueue/lifecycle flags exposed by gbrain jobs submit --help: --queue,
--priority, --delay, --max-attempts, --max-stalled, --backoff-type,
--backoff-delay, --backoff-jitter, --timeout-ms, --idempotency-key,
--dry-run.
These operations are MCP-callable and safe for agent use:
list_jobs --name shell --status active
get_job ID
get_job_progress IDCheck structured result fields (exit code, stdout/stderr tails, attempts,
timings) from get_job. Use get_job_stats (MCP) or gbrain jobs stats
(CLI) for the worker/queue health dashboard incl. the wedged-queue signal.
cancel_job id=ID
replay_job id=IDreplay_job is not protected — only shell submission is. Agents can
cancel or replay a shell job without CLI access.
Use idempotency keys for recurring shell workloads to avoid duplicate runs.
Use for open-ended reasoning, tool-using research, and fan-out synthesis.
User-facing entrypoint: gbrain agent run <prompt> is the canonical way
to submit subagent work. It handles the elevated-trust plumbing — subagent
and subagent_aggregator are both in PROTECTED_JOB_NAMES, so direct MCP
submission requires {allowProtectedSubmit: true}, which gbrain agent run
supplies.
gbrain agent run "Research Acme Corp revenue" --tools "search,query"--tools accepts a comma-separated subset of BRAIN_TOOL_ALLOWLIST (see
src/core/minions/tools/brain-allowlist.ts): query, search, get_page,
list_pages, file_list, file_url, get_backlinks, traverse_graph,
resolve_slugs, get_ingest_log, put_page. Anything outside the allow-list
is rejected at submit time with allowed_tools references unknown tool.
For parallel work with a fan-out manifest:
gbrain agent run --fanout-manifest companies.jsonThe manifest describes N children + 1 aggregator. Each child runs
name="subagent" under the hood; the aggregator runs name="subagent_aggregator"
and claims AFTER every child terminates. See
src/core/minions/handlers/subagent.ts and
src/core/minions/handlers/subagent-aggregator.ts.
Flags (from src/commands/agent.ts):
--subagent-def <name> — named subagent definition--model <id> — override model--max-turns <N> — cap the LLM loop--tools <csv> — allow-listed brain tools (see above)--timeout-ms <N> — hard timeout per job--fanout-manifest <file> — N children + 1 aggregator--follow / --no-follow — stream logs + wait (default on TTY)--detach — submit and return immediatelyQueue/priority/retry tuning is not exposed by gbrain agent run; submit the
raw subagent handler via gbrain jobs submit (requires CLI trust) if you
need those knobs.
Admission control (v0.46.11.0). Identical parentless subagent submits
(same owner lane, payload, and execution options) coalesce onto the existing
waiting job: gbrain agent run prints coalesced with the matched job id,
and the submit_agent MCP response carries coalesced: true. Treat that as
success — monitor the matched id, do NOT resubmit. Jobs still waiting after
the TTL (48h default for subagent; minions.ttl_waiting_hours.<name>)
are cancelled with reason prefix waiting_ttl_expired. If an operator has
configured a waiting quota (minions.quota_max_waiting.<name>), a submit
past the cap returns a structured, retryable rate_limited error — back
off and check gbrain jobs stats for a DIVERGENT QUEUE line before
retrying.
list_jobs --status active # MCP — what's running?
get_job ID # MCP — full details + logs + tokens
get_job_progress ID # MCP — structured progress snapshot
gbrain jobs stats # CLI — queue health dashboard
gbrain agent logs ID --follow # CLI — streaming transcript + heartbeatProgress includes: step count, total steps, message, token usage, last tool called.
Send a message to redirect a running agent:
send_job_message id=ID payload={"directive":"focus on revenue, skip headcount"}The agent handler reads inbox messages on each iteration and injects them as context. Messages are acknowledged (read receipts tracked).
Only the parent job or admin can send messages (sender validation).
pause_job id=ID # freeze without losing state
resume_job id=ID # pick up where it left off
cancel_job id=ID # hard stop
replay_job id=ID # re-run with same or modified params
replay_job id=ID data_overrides={"depth":"deep"} # replay with changesAll lifecycle ops are MCP-callable.
get_job ID # result, token counts, transcriptToken accounting: every job tracks tokens_input, tokens_output, tokens_cache_read.
Child tokens roll up to parent automatically on completion.
Background shells die silently: session compaction, harness restart, tool
timeout. Long gbrain operations (extract all, embed --stale, a full
sync --all) routinely run 10-60 minutes — past every one of those
ceilings. And even when the work survives, the completion event can be
swallowed (worker restart, dropped notification), leaving the user staring
at silence while a finished result sits unreported. Durable execution
covers both halves: the work survives, and the report provably lands.
Route any operation expected to exceed ~2 minutes through the highest rung of this ladder the deployment supports. This is a harness-routing convention the agent follows, not a mechanical guarantee — nothing stops a bare background shell except this skill saying don't.
Requires: Postgres engine, a running gbrain jobs work worker, and — for
the shell lane — GBRAIN_ALLOW_SHELL_JOBS=1 on the worker. All the
Preconditions above still hold: the flag defaults OFF, shell submission is
CLI-only across the MCP trust boundary, and PGLite has no worker daemon
(see Rung 3). Nothing in this section loosens that contract.
Submit as a job so the work survives restarts:
gbrain jobs submit shell \
--timeout-ms 3600000 \
--params '{"cmd":"gbrain extract all","cwd":"/abs/path","inherit":["database_url"]}'Work that needs LLM judgment goes through the subagent lane
(gbrain agent run, above) instead — a submitted agent job you never
check on is the same bug as an unwatched shell job.
Arm a deadman in the same action block (pattern below). Submitting without arming is the classic half-fix: the job survives, the silence doesn't.
When no scheduler can wake the agent but the host has a plain crontab (or
the work runs outside the queue entirely): make the operation write a
progress/heartbeat file as it advances (or rely on get_job_progress for
queue jobs), and register a recurring host-cron check that compares the
file's freshness against the expected progress interval. The agent also
checks it at the start of the next turn. A checkpoint that stops advancing
means stalled, not "still running" — the freshness comparison is the
entire value of this rung.
Run the operation inline in the foreground — on PGLite that's
gbrain jobs submit ... --follow, or just the raw command — and buffer
all output to a file, reading bounded slices, per
skills/conventions/exec-output.md. An empty tool result after a long
command is truncation, not a dead shell. This rung has no silent-death
insurance, so keep the operation in the foreground and stay with it;
backgrounding here recreates the exact failure the ladder exists to
prevent.
A one-shot, self-deleting scheduled check that fires at expected-finish-plus-margin and verifies the result was reported — not just that the process exited. Arm it in the same action block as the submission (not after, not "if I remember").
expected_minutes honestly; round up.margin = max(10 min, 50% of estimate). Restarts delay delivery — too tight false-fires, hours-late
defeats the point.at, or a crontab
entry the check removes on first fire. The check's instruction:
gbrain jobs get <id> gives the job state; the reported-check asks whether a
completion message actually reached the user.gbrain jobs get <id> and post the recovery report now, tagged with
the job ID.stderr_tail
from gbrain jobs get <id>) and offer gbrain jobs retry <id>.Failure modes the pattern must own (mirrored in Contract and Anti-Patterns): the deadman itself dying before it fires (second-line insurance, backstopped by the next-turn overdue sweep), double-fire (idempotent reported-check + job-ID-tagged reports), and stale checkpoints (freshness check before trusting "still running").
Eval contract, imported with the pattern — a deadman deployment is judged on:
Hard fails: a long user-facing operation with no deadman armed; a deadman
that posts noise when the completion arrived normally; emulating the timer
with sleep or a poll loop instead of a scheduler entry (a sleeping
process dies with the session — the exact failure being insured against).
Maintenance operations share locks (sync, embed, extract,
integrity). Run them sequentially, chained: submit job 1, arm its
deadman; on completion, submit job 2, arm the next; finish with
gbrain doctor and report the health delta. If a job dies mid-operation
its lock expires at TTL (a live, recently-refreshed holder is protected by
the steal grace) — never hand-delete lock rows to "unstick" a queue.
Durations scale with corpus size; treat these as order-of-magnitude
anchors for --timeout-ms, not promises.
| Operation | Typical duration | Suggested --timeout-ms |
|---|---|---|
extract all | 30-60 min on large brains | 3600000 |
embed --stale | 5-30 min (scales with missing count) | 1800000 |
sync --all | 5-20 min | 1200000 |
integrity auto | 10-30 min | 1800000 |
dream | 5-15 min | 900000 |
For a multi-stage pipeline with an expensive middle (extract → score →
explain → render → verify, where the scoring stage burns real LLM spend),
make each stage a content-addressed checkpoint so a crash — or a
deadman-triggered retry, or a replay_job — resumes instead of
re-spending:
.cache/
directory in the pipeline's working tree, addressed by content hash.Judged on: IDEMPOTENT (warm re-run recomputes nothing), CORRECT_BUSTING (a change recomputes exactly the affected stages), PROVENANCE (every artifact traces to logic + params + upstreams), INTEGRITY (corrupted artifacts detected, never silently reused).
This composes with the ladder rather than replacing it: the ladder keeps the pipeline running and reported; stage checkpoints keep a retry cheap. Pair them whenever a single stage costs more than pocket change in LLM spend.
When reporting job status to the user:
Job #ID (name) — status
Progress: step/total — last action
Tokens: input_count in / output_count out (+ cache_read cached)
Runtime: Xs
Children: N pending, M completedWhen reporting completion:
Job #ID completed in Xs
Tokens used: input / output / cache_read
Result: <summary>When reporting batch status (parent with children):
Parent #ID — waiting-children
#A subagent(Acme) — active, 3/5 steps, 2.5k tokens
#B subagent(Beta) — completed, 1.8k tokens
#C subagent(Gamma) — paused
Total tokens so far: 4.3kgbrain jobs stats firstcoalesced — the work is already queued; monitor the matched job id insteadsessions_spawn with runtime: "subagent" when Minions is available (use gbrain agent run instead)get_job in a tight loop (use get_job_progress for lightweight checks)sleep or a poll loop — a sleeping process dies with the session, which is the exact failure being insured againstsync, embed, extract, integrity) simultaneously, and don't hand-delete lock rows to unstick them (locks expire at TTL)submit_job (MCP, non-protected names only; shell jobs are CLI-only, subagent jobs via gbrain agent run)get_job (MCP)list_jobs (MCP)cancel_job (MCP)pause_job (MCP)resume_job (MCP)replay_job (MCP)send_job_message (MCP)get_job_progress (MCP)get_job_stats (MCP; admin scope over HTTP, same as the other
jobs ops here — includes the wedged-queue silent-halt signal) or gbrain jobs stats (CLI)055ac6c
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.