General-purpose coding policy for Baruch's AI agents
74
93%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Medium
Suggest reviewing before use
The owner of the answers below is code:
skills/herdr-foreman/classify/classify-report.sh and
skills/herdr-foreman/classify/report_verdict.py ask the questions and label a
report, and skills/herdr-foreman/foreman/report_gates.py decides and records
gates. This file says how the pieces fit, what a gate obliges, and how the
bands change. It restates no threshold, answer set or composition rule; those
are the owners' contracts.
skills/herdr-foreman/classify/classify-reports.sh.skills/herdr-foreman/classify/report_verdict.py docstring documents
(verdict, per-question answers, the report's sha256, the question hash and
the model id). The batch adds unannotated, one entry and reason per report
that got no label.skills/herdr-foreman/classify/report-questions.json.
Changing one changes every label's question hash; update the file's
changed date with it.skills/herdr-foreman/classify/classify-reports.sh.
Its exit contract is in that script's header. A failed classification lands
in unannotated and is never fatal. Exit 2 is a usage error, and its output
is never used as labels.skills/herdr-foreman/classify/classify-report.sh header.TYPESAFE_API_KEY
(.env.example). The client is
skills/herdr-foreman/classify/typesafe_client.py, shared with the evidence
assessor (#472).unannotated and
takes the reasoning path, a full read. No other classifier is asked.--agent codex|claude|grok runs an LLM adapter for measurement
(skills/herdr-foreman/classify/evaluate.sh). An LLM label carries no
probabilities and never gates.foreman report-gate-record --labels <classify-reports output> records the
gate each label earns and prints
{"schema_version": 2, "recorded": [...], "replayed": [...], "no_gate": [...]}:
recorded holds each new gate with its report, level and reason;
replayed a gate already on record for the same report bytes and
classification; no_gate each report its label leaves ungated, with the
reason. Exit 1 records nothing and names the cause on stderr. The owner computes the level from the label's
probabilities, the pinned model and the bands alone; the label's own
verdict and gate are never inputs.
A report whose bytes changed since classification is refused: reclassify it.block — the report cannot be accepted until
foreman report-gate-clear --report <path> (--evidence <report> --reason <why> | --decision <attention id>)
records why it does not block. close-member refuses an accepted ledger
decision and record-report refuses an approved verdict while it is open.
A needs_work or blocked decision is never refused by a block.reread — the report cannot be gated at all until a full re-read is
recorded with
foreman report-gate-reread --report <path> --evidence <re-read report> --note <what it verified>.
Both close-member and record-report refuse while it is open.The owner reads who resolved a gate from records it already holds, never from the caller. The gated report's task and role come from the applied dispatch whose supervision enrollment binds it; a report no enrollment binds cannot be resolved until the enrollment is restored.
--evidence names a report whose current bytes supervision
observed (report_observed) or recover-report recovered after the gate
was recorded, for another applied dispatch on the same task in the gated
report's role. It clears a block or records a re-read.config.json judge) in adjudication mode on the same task. It
clears a block; it never records a re-read.--decision names an attention decision on the same task,
resolved with the operator's user_answer after the gate was recorded. Its
answer summary is the recorded reason; --reason is refused.report-gate-clear resolves only block gates and report-gate-reread only
reread gates. A command with no open gate of its level refuses and names
the other command; a report carrying both stays unaccepted until both are
resolved.foreman report-gate-status [--report <path>] lists open and resolved gates.The same store holds a second gate source (#646): a report whose owner-parsed
VERDICT: line is blocking. No label is involved and no label ever touches
one.
assess-specialist (a replay included) and record-report record it,
source: verdict, level: block, with the dispatch the verdict was recorded
against. One gate per (report, bytes): a replay returns the existing gate
whatever its status; new bytes record a new gate.close-member and record-report ignore it.apply refuses a fresh release dispatch on its task while it is open, dry
runs included. A completed release replay is exempt.report-gate-clear clears it with --decision as above, or with
--evidence naming the gated responsibility's approved re-check. A judge's
report is refused: a ruling decides, and the re-check that cites it clears.
The re-check predicate is _rechecked's docstring.skills/herdr-foreman/foreman/report_gates.py owns <canonical selected state>.report-gates.json:
{"schema_version": 2, "state_path", "gates": [...]}. Each gate carries
schema_version, source (classifier|verdict), report (resolved,
canonical path), sha256, level, reason, dispatch, at, status
(open|cleared|reread) and resolution (null, or schema_version, at, action, by
(worker|judge|operator; worker|operator for a verdict gate), reason, evidence (attention for an
operator, or path, sha256 and dispatch)). A classifier gate also carries
probabilities, model, question and bands, and dispatch is null. A
verdict gate carries none of the four, is always block, and names its
dispatch. Every record is validated whole on every read. Writes
take the sidecar's own lock. close-member and record-report hold that lock
from their gate check through their commit, so a gate recorded meanwhile is
refused rather than slipped in between; record-report writes its verdict
gate under that same lock. A missing file is first use;
an unreadable or unsupported one, or a symlink at its path, refuses every
reader, never reading as no gates. A classifier replay is the same report bytes under the
same model, question, bands and level; any other classification of
those bytes records a new gate. close-member reads it and never writes it.
Schema 1 held classifier gates alone, without source or dispatch. Every
read goes through the owner, so the first read of a schema-1 document migrates
it: each gate becomes source: classifier, dispatch: null, its resolution's
schema_version becomes 2, and the document is rewritten under the sidecar
lock before it is returned. A read while another process holds that lock is
refused and retried, never served unmigrated. A reader at schema 1 meeting
a schema-2 document refuses it as unsupported rather than reading it as no
gates. The sidecar takes rules/stateful-artifacts.md Migration Policy's
gate-store exception: an open gate refuses an action, so "no usable prior
state" would read as no gate.
The live bands are the named constants at the top of
skills/herdr-foreman/foreman/report_gates.py, with BANDS_VERSION naming
their calibration. They ship uncalibrated and conservative. Calibration
proposes; a reviewed pull request installs:
TYPESAFE_API_KEY in the shell that runs the evaluation.bash skills/herdr-foreman/classify/evaluate.sh --agent jev --results <labels.json>.python3 skills/herdr-foreman/classify/scoring.py calibrate <labels.json>.
It prints proposed bands and the counts it used, and writes nothing. It
refuses to propose from too few held-out labels; its held-out test and
selection rules are in the skills/herdr-foreman/classify/scoring.py
docstring and top-of-file constants.BANDS_VERSION, with the
proposal's counts and the resulting false-gate counts in its CHANGELOG
entry.Recalibrate after every JEV_MODEL bump and every question change.
.tessl-plugin
hooks
rules
skills
adopt-fork-pr
herdr-foreman
classify
foreman
references
specialists
templates
tests
herdr-standup
migrate-to-plugin
onboard-repo
release
references
tests