General-purpose coding policy for Baruch's AI agents
73
91%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Low
Low-risk findings worth noting
#!/usr/bin/env bash
# Does this worker report state a finding the foreman must resolve?
#
# The first bounded classification in this plugin, and the one place in the
# foreman's round where the destination is right: the input is 21KB of prose whose
# meaning has to be read, and the answer is one of a fixed set
# (rules/script-delegation.md Bounded Classification).
#
# The evidence that it is not a script: across 90 recorded reports the foreman
# marked 63 `blocking` and 27 `approved`, while only 7 carry an explicit
# `## BLOCKED` heading. The verdict is not recoverable by grepping. It is in the
# prose or it is nowhere.
#
# The answer set carries `insufficient_evidence`, and that value routes the
# question back to the reasoning round -- here, the foreman reads the report as it
# always has. A label never suppresses that read and never approves anything:
# the only effect one can have is a gate that adds friction
# (foreman/report_gates.py).
#
# The question is split into atomic questions (report-questions.json) and the
# verdict is composed from their answers in code, so no model weighs the policy
# that combines them. Deterministic checks run first: the report is framed as
# untrusted data, and every quote an LLM returns must be a passage of the
# report or the label is `insufficient_evidence` (report_verdict.py).
#
# Adapters:
# jev -- the default. TypeSafe System One, one Noul per atomic question,
# P(yes) each. Needs TYPESAFE_API_KEY. The only adapter whose label
# can gate. A Jev that is unavailable (key unset, service down,
# report too large, answer off contract) or that refuses the
# request produces no label: the script exits non-zero and the
# report takes the reasoning path, a full read. No other
# classifier is asked (rules/script-delegation.md Bounded
# Classification).
# codex, claude, grok -- named with --agent only, for measurement
# (evaluate.sh). Their labels never gate.
# Each LLM adapter constrains its answer to the generated schema --
# `codex exec --output-schema`, `claude --json-schema`, `grok --json-schema` --
# and every answer then passes the same check in report_verdict.py, which is
# the actual guarantee. Every LLM adapter runs from an empty directory, with no
# tools, for one turn: an adapter left free to search the workspace read this
# repository's tests and answered from them.
#
# Usage: classify-report.sh <report-path> [--agent jev|codex|claude|grok]
# [--model <id>] [--out <file>]
#
# Output contract (rules/script-delegation.md -- structured stdout):
# stdout: one label, the JSON object report_verdict.py documents: verdict
# `blocking`|`approved`|`insufficient_evidence`, the deciding passage in
# `evidence` (empty for Jev, which quotes nothing), one answer per atomic
# question with P(yes) for Jev, and the `gate` the owner would
# record. The report hash, question hash and model id travel with every
# label, so a question edit or a model bump is attributable
# (rules/dependency-management.md Freshness).
# stderr: diagnostics.
#
# Exit 0 on a label, including `insufficient_evidence` -- an honest abstention
# is an answer. Exit 2 on a usage error, an unreadable report, a report that
# contains its own delimiter, an unavailable or refusing Jev, or a call that
# produced nothing conforming. A failed call is never a verdict.
#
# Calling this costs model quota. `evaluate.sh` runs it over the labelled corpus
# and reports accuracy; nothing calls it automatically, and no test calls it live
# (rules/testing-standards.md Determinism).
set -euo pipefail
# Command substitution strips every trailing newline, so the script directory
# never passes through one bare: parameter expansion derives it (#487), and a
# sentinel carries `pwd` across the strip (#466).
case "${BASH_SOURCE[0]}" in
*/*) HERE_SRC="${BASH_SOURCE[0]%/*}" ;;
*) HERE_SRC=. ;;
esac
if ! HERE="$(CDPATH='' cd -- "${HERE_SRC:-/}" && pwd && printf x)"; then
echo "classify-report: cannot enter the script directory ${HERE_SRC:-/} — restore read and search access to the plugin directory, or reinstall the plugin, then re-run" >&2
exit 2
fi
HERE="${HERE%x}"
HERE="${HERE%$'\n'}"
#: The pinned LLM model per kind. A bump is a dependency bump: change it here,
#: and every label recorded afterwards carries the new id. These are
#: classification pins, not the fleet's frontier seats. The Jev pin lives with
#: its bands in foreman/report_gates.py (`JEV_MODEL`): bands are per model
#: version.
#: Renewal: these pins come due with the capability table, on `INTERVAL` in
#: skills/herdr-foreman/foreman/capabilities.py (weekly). At each refresh,
#: compare every pin against the table's current rows; a bump lands only after
#: `evaluate.sh --since <last bump>` scores the new model on reports it has not
#: seen, and the CHANGELOG records both numbers.
DEFAULT_AGENT="jev"
model_for() { # <kind>
case "$1" in
codex) echo "gpt-5.6-sol" ;;
claude) echo "claude-sonnet-5" ;;
grok) echo "grok-4.6" ;;
*) return 1 ;;
esac
}
SCRATCH=""
cleanup() { if [ -n "$SCRATCH" ]; then rm -rf "$SCRATCH"; fi; return 0; }
trap cleanup EXIT
die() { echo "classify-report: $*" >&2; exit 2; }
# Each adapter reads the whole question on stdin, writes the model's answer
# object to <answer>, and returns non-zero when the vendor call failed.
ask_codex() { # <model> <schema> <answer> <log>
codex exec --json --skip-git-repo-check --sandbox read-only \
--model "$1" --output-schema "$2" --output-last-message "$3" - >"$4" 2>&1
}
ask_claude() { # <model> <schema> <answer> <log>
local raw="${3}.raw"
claude -p --model "$1" --output-format json --json-schema "$(cat "$2")" \
--tools "" --strict-mcp-config >"$raw" 2>"$4" || return 1
python3 "${HERE}/extract-answer.py" claude "$raw" "$3" 2>>"$4"
}
ask_grok() { # <model> <schema> <answer> <log>
local raw="${3}.raw" question
question="$(cat)"
grok -m "$1" --json-schema "$(cat "$2")" --max-turns 1 --no-subagents \
--disable-web-search --tools "" -p "$question" </dev/null >"$raw" 2>"$4" || return 1
python3 "${HERE}/extract-answer.py" grok "$raw" "$3" 2>>"$4"
}
main() {
local report="" out="" agent="" model=""
[ $# -gt 0 ] || die "usage: classify-report.sh <report-path> [--agent jev|codex|claude|grok] [--model <id>] [--out <file>]"
report="$1"; shift
while [ $# -gt 0 ]; do
case "$1" in
--agent) agent="${2-}"; shift 2 || die "--agent needs jev, codex, claude or grok" ;;
--model) model="${2-}"; shift 2 || die "--model needs an id" ;;
--out) out="${2-}"; shift 2 || die "--out needs a file" ;;
*) die "unknown argument '$1'" ;;
esac
done
[ -f "$report" ] && [ -r "$report" ] || die "'${report}' is not a readable file"
case "$agent" in
''|jev|codex|claude|grok) ;;
*) die "--agent '${agent}' is not one of jev, codex, claude, grok" ;;
esac
[ -n "$agent" ] || agent="$DEFAULT_AGENT"
local verdict="${HERE}/report_verdict.py"
[ -r "$verdict" ] || die "missing the question owner at ${verdict}; reinstall the plugin (tessl install jbaruch/coding-policy), then rerun"
local work
work="$(mktemp -d "${TMPDIR:-/tmp}/classify-report.XXXXXX")" || die "cannot create a temporary directory; repair write access to \${TMPDIR:-/tmp} (or point TMPDIR at a writable directory), then rerun"
SCRATCH="$work"
local answer="${work}/answer.json" log="${work}/run.log" room="${work}/room" snapshot="${work}/report"
# One read of the report: every adapter, the evidence check and the label's
# sha256 see these bytes, so a report rewritten mid-run is never half-read.
cp -- "$report" "$snapshot" || die "cannot snapshot ${report}; restore read access and re-run"
local payload=""
if [ "$agent" = "jev" ]; then
if ! payload="$(python3 "$verdict" jev "$snapshot" --as "$report" ${model:+--model "$model"} 2>"$log")"; then
local why
why="$(tail -n 1 "$log")"
die "no Jev label (${why#report_verdict: }); read ${report} in full -- a failed call is never a verdict"
fi
else
local pinned
pinned="$(model_for "$agent")" || die "no pinned model for '${agent}'; pass --model <id>, or reinstall the plugin, then rerun"
[ -n "$model" ] || model="$pinned"
command -v "$agent" >/dev/null || die "${agent} is not on PATH; install its CLI or pick another --agent, then rerun"
local schema="${work}/schema.json" question="${work}/question.txt"
python3 "$verdict" schema "$schema" || die "cannot write the answer schema to ${schema}; repair write access to \${TMPDIR:-/tmp} (or point TMPDIR at a writable directory), then rerun"
python3 "$verdict" frame "$snapshot" "$question" > "${work}/frame.json" \
|| die "cannot frame ${report} as data; read it in full, or restore the questions file and the plugin, then rerun"
mkdir "$room" || die "cannot create the empty working directory; repair write access to \${TMPDIR:-/tmp} (or point TMPDIR at a writable directory), then rerun"
if ! ( cd "$room" && "ask_${agent}" "$model" "$schema" "$answer" "$log" < "$question" ); then
cat "$log" >&2
die "the ${agent} call failed; a failed call is never a verdict"
fi
[ -s "$answer" ] || { cat "$log" >&2; die "${agent} wrote no schema-conforming answer"; }
payload="$(python3 "$verdict" label "$answer" "$snapshot" "$agent" "$model" --as "$report")" \
|| die "the ${agent} answer did not conform to the schema"
fi
printf '%s\n' "$payload"
if [ -n "$out" ]; then
printf '%s\n' "$payload" > "$out" || die "cannot write the label to ${out}; pass a writable --out path, then rerun"
fi
return 0
}
[[ "${BASH_SOURCE[0]}" == "${0}" ]] && main "$@".tessl-plugin
hooks
rules
skills
adopt-fork-pr
herdr-foreman
classify
foreman
references
templates
tests
herdr-standup
migrate-to-plugin
onboard-repo
release
references
tests