Use when changing review sessions, finding decisions, the review fix-round carry, review schema retries, agent timeouts, local Test behavior and its live-validation contract, approval overrides and parked gates, or intent conformance.
Review-Loop Agent Sessions (internal/pipeline/sessions.go)
fixRoundProvenanceClause); the same clause is emitted on a later run's initial review when a persisted uncertified range is bound. Prior findings, fix summaries, and same-round tests are claims, not evidence.session_reuse: false forces everything cold. Persistence is minimum metadata only, never prompts or transcripts; SessionRoleReviewer remains only so crash recovery accepts legacy persisted rows, which are never resumed.codex exec resume has a narrower flag surface than codex exec, so an unsupported override fails the resume and falls back; the e2e fakeagent must keep parsing both codex argv shapes (extractCodexPrompt).internal/pipeline/sessions_test.go, internal/pipeline/steps/review_session_test.go (incl. TestReviewLoop_RereviewNeverResumesTheSessionThatPrescribedItsFixes), TestReviewStep_RereviewTreatsFixRoundsAsPipelineAuthoredCode, internal/agent/session_test.go.Recorded Human Decisions on Findings
selected_finding_ids = "[]" plus selection_source = user_declined on a gated round with findings (executor.go recordDeclinedRound, db.SetStepRoundDeclined); a round with no findings records no decision. The conditional write must never erase an existing selection. User-facing semantics are owned by docs/src/content/docs/reference/pipeline-steps.md.declinedFindingLines derives it and deliberately excludes auto_fix selections, whose complement is findings still awaiting a decision (rendered under auto_fix_left_unselected, which carries no do-not-re-report instruction).roundHistoryPromptSection (internal/pipeline/steps/round_history.go) now carries three parts: this step's rounds, this run's OTHER steps' decisions, and earlier runs' decisions on this branch (bound per step by pipeline.BindBranchDecisions, unlike review-only BindUncertifiedPipelineRange). Nothing clears branch decisions - a completed review deletes the uncertified range, which is why that channel could not carry a decision forward, but approving a gate IS the decision. The prompt states that a recorded decision SUPERSEDES the user-intent wording. Both the cross-step/branch preamble and the own-step prefix carry the respect rule as plain instruction text: never revert, undo, or work around a recorded human decision, and never re-implement a declined one.assertPipelineHeadContinuity and assertReviewApprovedPushHead remain lineage-only. ci_fix.go composes the section like the other fix-capable steps - before userIntentPromptSection, because the run's frozen intent always predates a later gate's decision - so a CI repair sees the decision an earlier gate, an earlier run on this branch, or its own gate recorded (internal/pipeline/steps/ci_decision_history_test.go). rebase.go still builds its prompt without roundHistoryPromptSection.TestExecutor_GateResolutionsWithoutASelectionRecordTheDecline, TestExecutor_GateResolutionWithNoFindingsRecordsNoDecision, TestExecutor_FixResolutionStillRecordsAUserSelection, internal/db/round_decisions_test.go, TestDeclinedFindingReachesALaterStepInTheSameRun, TestDeclinedFindingReachesALaterRunOnTheSameBranch, TestCompletedReviewDoesNotClearBranchDecisions, TestAutoFixComplementIsNeverPresentedAsAUserDecision, TestDocumentPromptCarriesRecordedHumanDecisions, TestLintAgentPassPromptCarriesRecordedHumanDecisions, TestLintFixPromptCarriesRecordedHumanDecisions, TestTestPromptsCarryRecordedHumanDecisions, TestOwnStepHistoryPromptCarriesNeverRevertClause.Uncertified Review Provenance (internal/pipeline/uncertified.go)
from_sha, to_sha). Persist on review-step fixer commits only, not lint or document. On the next run's initial review, bind that range and emit fixRoundProvenanceClause even when Fixing==false, so the replacement reviewer is not cold. Rerun proceeds; there is no refusal or --ack-uncertified-review gate.to_sha; parked, failed, skipped, and aborted reviews must not clear it. Rebase remaps the persisted SHAs onto the rewritten head so the next review can still bind.internal/pipeline/uncertified_test.go, TestCommitAgentFixes_PersistsUncertifiedRangeForReview, TestCommitAgentFixes_LintDoesNotPersistUncertifiedRange, TestCommitAgentFixes_DocumentDoesNotPersistUncertifiedRange, TestFixRoundProvenanceClause_EmitsForUncertifiedRangeWhenNotFixing, TestUncertifiedRange_PersistsThenFeedsNextInitialReview, TestRebaseStep_RemapsUncertifiedRangeWhenHeadRewritten.Authorization and Privacy Tracing (internal/pipeline/steps/review.go)
review.path_instructions. Findings require a concrete source-backed reachable operation or disclosure and name the protected resource/field, missing control, and impact; accept equivalent controls and intentionally public data rather than keyword-matching auth machinery. A material policy ambiguity introduced by changed behavior must be ask-user and name the missing policy decision; do not report immaterial or pre-existing ambiguity, and retain auto-fix for source-proven routine defects.docs/src/content/docs/reference/pipeline-steps.md. Prompt regression: TestReviewStep_AuthorizationPrivacyTracingContract. Development-only qualitative cases: internal/pipeline/steps/testdata/authorization_privacy_review/.Intended-Usage Evidence Threshold (internal/pipeline/steps/review.go)
docs/src/content/docs/reference/pipeline-steps.md. Prompt regression: TestReviewStep_IntendedUsageEvidenceContract. Development-only qualitative cases: internal/pipeline/steps/testdata/intended_usage_review/.Simplification Pass and Fix-Through-Removal (internal/pipeline/steps/review.go, internal/pipeline/steps/common_fix.go, internal/pipeline/steps/lint.go, internal/pipeline/steps/ci_fix.go)
Simplification section between Rules and Risk assessment. It asks a different question from the defect pass: not "is this component correct" but "does the intent strictly require this component". Every component the change introduced (branch, acceptance/matching path, fallback, alias, mode, flag, option, second definition of a concept already defined once, parallel copy of a rule) is judged against the user intent, or the change's own stated purpose when none is given. An unrequired component is a warning with action ask-user whose remedy is removal, never hardening, validating, or documenting; a defect finding inside such a component must name removal as its remedy. Refactor-only "simplification opportunities" keep their existing non-feature-removing meaning and point at the section instead of contradicting it. No schema field, detector, or second reviewer.fixerPrompt applies the removal rule at the shared Review, Test, and configured-Lint repair boundary; Lint's no-command safe-fix pass and the CI repair prompt apply the same rule directly. A finding resolvable by removing a code path the intent does not strictly require is fixed by removing it. The Review-specific anti-revert guard protects only intent-required code; doubt about whether the intent requires a path leaves it alone and reports the finding unresolved.docs/src/content/docs/reference/pipeline-steps.md (Review, shared repair behavior, CI fixer) and docs/src/content/docs/concepts/auto-fix.md (finding actions). Regressions: TestReviewStep_SimplificationSectionContract, TestReviewStep_FixPromptPrefersRemovalOfUnrequiredPaths, TestTestStep_FixMode, TestLintStep_FixMode_CommitsChanges, TestLintStep_NoConfiguredLint_UnresolvedFindingsNeedApprovalWithoutAutoFixLoop, TestCIStep_FixPromptPrefersRemovalOfUnrequiredPaths, TestReviewStep_SimplificationFixturesApply. Development-only qualitative cases: internal/pipeline/steps/testdata/simplification_review/.Review Fixer Verification Discipline (internal/pipeline/steps/review.go)
TestReviewStep_FixMode_FocusedVerificationContract.Remedy-Scope Discipline in the Review Prompts (internal/pipeline/steps/review.go)
ask-user, with the description naming the remedy as what needs authorization. This rides ActionOrDefault's established fail-toward-the-human direction.ciFixerClassRules (ci_fix.go).fixRoundProvenanceClause branches before the revert exit ramp.ask-user finding recommending that round be reverted to the minimal fix, instead of further repairs layered on it. Emitted in both fixRoundProvenanceClause branches (this run's fix rounds and a previous run's uncertified fixer commits); conditioning on prior-round code keeps it off ordinary multi-round fixes.docs/src/content/docs/concepts/auto-fix.md (finding actions) and docs/src/content/docs/reference/pipeline-steps.md (Review). Regressions: TestReviewStep_PromptClassifiesFindingsByRemedyScope, TestReviewStep_FixPromptPrefersSimplificationOverMachinery, TestReviewStep_RereviewOffersRevertExitFromPriorRoundMachinery, TestReviewStep_FixPromptClosesTheInvariantAcrossSiblingSites, TestReviewStep_ReviewPromptReportsTheClassOnce, TestReviewStep_RereviewNamesFollowOnsOfPriorFixRounds, TestCIStep_FixPromptClosesTheInvariantAcrossSiblingSites. Qualitative fixtures: internal/pipeline/steps/testdata/class_fix_review/ (kept applicable by TestReviewStep_ClassFixFixturesApply).Agent-Invocation Timeouts Report Measured Silence, Never the Budget
agentActivity in agent_run.go is the single owner of the measurement and resets per-attempt evidence whenever a retry or fallback starts a replacement attempt, including provider, session-resume, and OpenCode prompt-format fallbacks. A substantive adapter error (a native agent's exit status plus captured stderr) is URL-redacted, length-bounded, and appended as agent reported: ....agent.LifecyclePhaseActivity, sourced from every non-empty read of a native subprocess's stdout or stderr. Prose alone cannot prove liveness: verified against pi 0.84.3, a tool-using turn emits only tool_execution_*/toolcall_* and no text_delta until the very end, and no adapter forwards those to OnChunk. Subprocess start and exit are deliberately NOT output - start proves launch, not work, and exit is the deadline's own consequence, so counting either would recreate the fabricated evidence.LifecyclePhaseActivity into step activity only, never the step log: axi status needs the liveness, and a half-hour turn would otherwise emit hundreds of log lines.ciFixAgentTimeoutOutcome) instead of being logged as a warning and re-issued on the next poll. That old path spent up to auto_fix.ci full budgets invisibly until ci_timeout. Only pipeline.ErrAgentTimeout parks; other fix failures keep warn-and-retry. Review deliberately still fails the run rather than parking - Push commits leftover worktree changes, so an approved park would ship a half-finished, unreviewed fix.review_agent_timeout absolute wall-clock limit. Prompt preparation and post-fixer commits stay on the step parent context, and an expired candidate context stops the fallback chain before another provider is announced or instrumented.docs/src/content/docs/reference/global-config.md (agent_timeout, review_agent_timeout) for the diagnostic and ownership vocabulary, docs/src/content/docs/reference/pipeline-steps.md (Review and CI) for step behavior. Regressions: TestRunAgent_Timeout*, TestRunAgent_SubprocessStartAloneIsNotObservedOutput, TestRunAgent_OperatorCancellationIsNotDressedUpAsAnAgentFault, TestPiAgent_ToolOnlyStreamStillReportsSubprocessLiveness, TestPiAgent_SilentSubprocessReportsNoLiveness, TestExecutor_SubprocessLivenessUpdatesActivityWithoutFloodingTheStepLog, TestCIStep_FixAgentBudgetExhaustionParksForADecisionInsteadOfRetrying, TestCIStep_FixAgentTimeoutRecordsCommittedRepair, TestCIStep_FixAfterATimedOutRepairRevalidatesTheRecordedCommit, TestCIStep_FixAgentCutMidRebaseRecordsNoPartialHead, TestCIStep_NonTimeoutFixFailureKeepsRetrying, TestReviewStep_EachAgentInvocationGetsItsOwnBudget, TestReviewStep_WallClockTimeoutPreservesTheAgentReport, e2e TestSilentAgentTimeoutReportsMeasuredEvidence.Local Test Is Targeted Validation (internal/pipeline/steps/test.go)
go test -race ./... in .github/workflows/ci.yml) and remains mandatory before a PR is ready.
commands.test is the same contract when set: targeted baseline, not CI-parity complete-suite configuration; docs owner is docs/src/content/docs/reference/repo-config.md (commands.test), step behavior owner is docs/src/content/docs/reference/pipeline-steps.md (Test).
This repository dogfoods an empty commands.test so the agent-driven targeted path is the default; do not reintroduce go test -race ./... as a local Test override.
Trusted test.prepare is the explicit eager opt-in for calling the shared ensurePrepared lifecycle before the first agent-only Test turn when commands.test is empty. A successful worktree receipt is reused after recovery and by later configured commands; failures and restoration failures stop before the agent, while evidence remains fresh and verdict policy remains unchanged. Trust selection is owned by the repository-routing-security skill; user semantics are owned by docs/src/content/docs/reference/repo-config.md (test.prepare) and docs/src/content/docs/reference/pipeline-steps.md (Test). Regressions: TestAgentOnlyPreparation_*, TestTestPrepareTrustAndMerge.
Trusted test.base_attribution (test_base_attribution.go) re-runs only a FAILING configured commands.test on the run's base SHA in a git clone --shared scratch checkout outside the run worktree and gate (running commands.prepare there first, both under the machine-local nice), then splits recognized failure lines of the repository command's own output (checkResult.Output, never machine-local additional checks) into introduced, pre-existing (only lines qualified by identity: go lines keyed by package, pytest path::test ids), and ambiguous (any unqualified line that also fails on base); a go package failing with no per-test line that did not fail that way on base is listed as unattributed; a failing base with no recognized lines reports "could not be separated". The result is prepended to the findings summary and appended to the evidence prompt; it never edits the test-command finding, whose description is the attested override reason, and never changes severity, action, or approval semantics. Any base-side failure degrades to an "unavailable" note with every failure still attributed to the change; only step cancellation is an error. User semantics: docs/src/content/docs/reference/repo-config.md (test.base_attribution). Regressions: TestTestStep_BaseAttribution*, TestTestFailureLinesNormalizesRunnerNoise, TestBaseAttribution*, TestTestBaseAttributionTrustAndMerge.
Process-group reaping on clean/error exit (#357) and Unix WaitDelay remain the lifecycle safety net when agents spawn test workers - restoring the agent-driven path must not revive the daemon OOM leak.
Those agent turns are bounded by test_agent_timeout (default 30m, global-only): a stalled evidence or repair agent is cancelled and the Test step parks for a decision rather than failing the run as a code defect. Late structured output from the expired turn is still not used as a pass. The park carries the configured command's result from the same execution (a cut repair turn re-runs it first), so an approval over a failing command stays a configured-command waiver; a cut fix round also carries the gate it answered (selected and deferred findings plus the last completed evidence turn's verdict, scenarios, and tested head, answeredTestGate). A commit the timed-out agent already made is recorded locally unless an unfinished rebase or merge leaves a partial HEAD; uncommitted leftovers stay in the live worktree. While the worktree is dirty or HEAD is past the last completed evidence turn's tested head (else the head an earlier cut measured from, carried on the park as unvalidated_since_sha and measured again on every cut, else this execution's start), the park adds test-agent-unvalidated-work and the executor refuses Approve (pipeline.HasUnvalidatedWorkRefusal), because Document and Push would commit and publish it. A fix selection of only budget-cut findings re-runs validation without a repair turn. Raise the setting when targeted tests routinely approach 30m; do not raise the shipped default. Docs owner is docs/src/content/docs/reference/global-config.md.
Every other pipeline agent invocation is bounded by agent_timeout (default 30m, global-only) at pipeline.RunAgent / the executor timeoutAgent seam, so a new agent-spawning step cannot hang a run by forgetting a deadline. Review gives every fixer and reviewer a fresh review_agent_timeout; an existing sooner parent deadline is honored rather than capped. The invocation context is scoped only to Agent.Run; a late successful return after the deadline is rejected. Docs owner is docs/src/content/docs/reference/global-config.md.
Regressions: TestTestStep_InitialAgent_TargetedValidationContract, TestTestStep_FixMode_TargetedVerificationContract, TestTestStep_FixMode_DriverFullSuiteInstructionDoesNotOverrideContract, TestTestStep_InitialAgent_NoTargetedEvidenceRequiresHonestFinding, TestTestStep_HangingEvidenceAgentParksForADecision, TestTestStep_FixAgentTimeoutParksWithoutCommit, TestTestStep_FixAgentTimeoutRecordsCommittedHead, TestTestStep_EvidenceCutKeepsTheFailingConfiguredCommand, TestTestStep_RepairCutRecordsTheConfiguredCommandResult, TestTestStep_FixOfABudgetCutAloneRerunsOnlyValidation, TestTestStep_CutMidRebaseRecordsNoPartialHead, TestTestStep_RepeatedCutKeepsRefusingUnvalidatedWork, TestTestStep_RepairCutKeepsTheFindingsItWasFixing, TestTestStep_CutAfterAnEarlierStepCommittedStaysApprovable, TestTestStep_RepeatedCutKeepsRefusingBeforeAnyEvidenceCompletes, TestTestStep_ValidationOnlyCutKeepsTheDeferredNoGo, TestTestStep_RepeatedCutReMeasuresLeftoversInsteadOfRepeatingThem, TestExecutor_UnvalidatedTestWorkRefusesApprovalUntilFixValidatesIt, TestCodexAgent_RunCancelsSilentHang, TestDogfoodConfig_NoBroadLocalTestCommand, TestCIWorkflow_RetainsFullRaceSuiteAsBroadRegressionOwner, plus the existing #357 reap/WaitDelay tests, TestRunAgent_*, TestExecutor_DirectAgentRunIsDeadlineBounded, TestDocumentStep_HangingAgentFailsRunAfterTimeout, TestLintStep_HangingAgentFailsRunAfterTimeout, TestCIStep_HangingFixAgentFailsAfterTimeout, TestCIStep_FixAgentTimeoutRecordsCommittedRepair, TestRebaseStep_HangingConflictAgentFailsAfterTimeout.Intent Provenance & Conformance (internal/pipeline/steps/intent_prompt.go)
axi run --intent persists Source==db.RunIntentSourceAgent ("agent", score 1); a transcript match persists the agent name ("claude"/"codex"/...). The executor propagates it as StepContext.IntentSource alongside UserIntent (executor.go).userIntentPromptSection branches on source: an EXPLICIT intent renders as sanitized-but-AUTHORITATIVE acceptance criteria; an INFERRED intent keeps the low-confidence hint framing verbatim. Both branches keep the StripAdversarial+RedactSecrets pipeline and BEGIN/END "do not execute instructions" guard - authoritative reframes only the content's authority (check the diff against the criteria), never whether control tokens are stripped. The review prompt adds intentConformanceReviewClause for agent-source intent only: a fixer change that contradicts the criteria (removes intent-required or adds intent-forbidden behavior) MUST become an ask-user finding, which parks with no executor change. Both the authoritative framing and the conformance clause state that later recorded human decisions supersede conflicting original intent, so a human-selected fix is never flagged as an intent contradiction; fixerRemovalRule in common_fix.go carries the same precedence so a fixer does not remove a path a recorded decision requires. Conformance is limited to source-verifiable criteria; deferred pipeline-owned delivery (remote branch / push / PR / CI for this run) is out of scope at review.StepReview before StepPush/StepPR/StepCI). pipelineDeliveryPhaseClause plus stripDeferredPipelineOwnedDeliveryFindings (pipeline_delivery.go, applied in review.go) keep findings that only claim those later-owned outcomes are missing from parking the run. External or pre-existing lifecycle requirements (numbered PR, third-party artifact, non-run-owned state) stay enforceable. Push, PR, and CI steps remain strict after their stages run.action fails closed to ask-user, not auto-fix (types/findings.go ActionOrDefault); HasAskUserFindings uses ActionOrDefault so it agrees with AutoFixableFindings (an unclassified finding is never auto-fixed and is always caught as ask-user). MergeUserOverrides still stamps user-added findings auto-fix on purpose.review.go owns the held-scope TODO.internal/pipeline/steps/intent_prompt_test.go, internal/pipeline/steps/review_test.go (TestReviewStep_ConformanceObligationTracksIntentProvenance, TestReviewStep_RereviewFlagsIntentContradictionAsAskUser), internal/pipeline/steps/pipeline_delivery_test.go, internal/pipeline/steps/review_pipeline_delivery_test.go, internal/pipeline/executor_intent_conformance_test.go, internal/types/findings_test.go, e2e TestIntentJourney (inferred-source framing), e2e TestReviewPipelineOwnedPRCriterionDoesNotPark / TestReviewExternalPRLifecycleStillParks.Live-Validation Contract (Test step)
scenarios (name, result pass/fail/untested, live, evidence, reason) plus a verdict (go/no-go/inconclusive/no-surface); both are REQUIRED by testFindingsSchema and validated in unmarshalRequiredTestFindings. live: true means driven against the real running product in that run - a unit test, stub, or fixture is not live.runTestAnalyzer gives a fresh agent the rejected payload and validation errors for a bounded correction (testAnalyzerMaxAttempts); only exhausting that bound fails the step. The correction prompt must remain correction-only: no tools, scenario reruns, or external operations; preserve supported observations and downgrade unsupported pass/fail claims to untested. Do not weaken the pass/fail-must-be-live or untested-must-have-a-reason invariants to avoid the retry.verdictFindings (internal/pipeline/steps/test.go): no-go parks as an auto-fixable error, inconclusive parks as an ask-user warning (live surface, too little driven), no-surface parks as an ask-user warning when the change has nothing no-mistakes can drive live, a fail scenario with any other verdict is rejected as a contract violation, no-surface with any live or pass/fail scenario is rejected so a skipped live validation cannot masquerade as no-surface, and untested NEVER parks - it is listed on the PR so reporting it honestly stays cheaper than guessing a pass. Do not make untested blocking.internal/pipeline/steps/prsummary_scenarios.go (Testing section, Pipeline fold, and attestation live_validation: {verdict, live, total}). The attestation omits live_validation whenever the current head differs from the head Test validated. Pre-contract runs parse and render exactly as before.commands.test baseline is never enough on its own, so the step always drives scenarios afterwards. The step deliberately owns no capability inventory, scratch lab, or environment exports - the agent discovers what the machine can do for itself.test.instructions is the repository's live-validation runbook. It is trusted-default-branch-only, exactly like document.instructions and for the same reason (it steers the gate that validates the pushed branch), and is resolved from the repository only - never from global config, which has no repository to describe.TestTestStep_PromptDerivesScenariosAndMarksLive, TestTestStep_VerdictPolicy, TestTestStep_NoLiveSurfaceCIWorkflowAsksUser, TestTestStep_MissingScenarioContractFails, TestTestStep_InvalidAnalyzerPayloadTriggersCorrectionRound, TestTestStep_InvalidAnalyzerPayloadExhaustsCorrectionBound, TestTestStep_ValidMixedPayloadDoesNotRetry (internal/pipeline/steps/test_scenario_contract_test.go), internal/pipeline/steps/prsummary_scenarios_test.go, require_no_mistakes_live_validation_test.go, TestEffectiveRepoConfig_TestInstructionsTrustedOnly.Review Schema Retry (internal/pipeline/steps/review.go)
agent.IsStructuredOutputRejected: Pi when its final JSON fails validation, opencode once its own internal StructuredOutput retries run out), or parseReviewAnalyzerOutput refuses the output, that is a formatting slip, not a verdict: ReviewStep.Execute reruns a fresh, session-free review of the same prompt plus a note quoting the validation error, up to reviewAnalyzerMaxAttempts (3) attempts in total.
Claude Code's --json-schema re-prompts within its own limit and reports exhaustion as error_max_structured_output_retries; that result is deliberately not marked as a rejection and fails the step as before, so there is no second retry layer and Test's correction loop is untouched.
Findings come only from the attempt that validates; nothing else from a rejected attempt carries over, and fix-mode turns are never retried.
Exhausting the bound fails the step with validate review analyzer findings after 3 attempts: <last error>, so an unreadable review never passes (issue #703); the unconfigured "pass" meaning is unchanged. When the attempts' agent.SchemaViolation fields (typed by the validator in internal/agent/agent.go, dotted path without array indices, empty for a whole-output type failure) name at least two distinct fields, distinctReviewValidationFailures attributes the error instead: ... output kept failing validation across distinct fields (<fields>): attempt 1: ...; attempt 3: <last error>, each attempt's error printed once (issue #1280). Fields come only from the typed violation, never from parsing the message text.agentInvocationError carries both the adapter's rejection and errReviewAgentTimeout). This is a rerun of the review, not a correction turn; Test's correction behaviour is separate and unchanged.internal/eval/replay.go observedAgent) sums token usage across every review attempt and reports none when any attempt lacks reported usage, matching the captured baseline, so a candidate that needed reruns never reads as cheaper.internal/pipeline/steps/review_schema_retry_test.go, TestReviewStep_UnrunAnalyzerDoesNotApprove, TestReplayTokensCoverEveryReviewAttempt, e2e TestAnalyzerEvidenceFailuresFailPipelineJourney, e2e TestReviewSchemaRetryExhaustionJourney.Review Fix-Round Carry (internal/pipeline/executor.go, internal/pipeline/findings.go)
respond --action fix additionally hands the selection to the fixer but does NOT subtract it from the set. A selected finding leaves it only on a positive coverage record - the rereview lists the finding's file in reviewed_paths (review schema in internal/pipeline/steps/common.go, StepOutcome.ReviewedPaths), that path is in the trusted current reviewable changed-file set (StepOutcome.ReviewablePaths), neither re-reports the defect, nor reports anything else at all in that same file - or on an explicit operator approve/skip/abort. Silence, missing or fabricated/out-of-scope coverage (including a mixed coverage list), a round that looked elsewhere, a re-reported defect, and any other finding in the same file (an ambiguous rereport that might be the original defect shifted to another line or reworded) all leave it outstanding; any file-less finding in the current rereview also blocks clearing every selected finding in that round. One upgrade-compat carve-out: a SELECTED finding with no file anchor can never match a coverage record, so it clears on a positive verification round whose reviewed_paths covers every trusted reviewable path and that no longer reports it - runs parked before the recorded-decision review machinery was removed carry such items, and a fix selection could otherwise never clear them. A crash-resumed fix round fails closed the same way. A validation restart at Review starts a fresh carry set rather than resurrecting the superseded cycle's findings. resolveVerifiedFindingsJSON owns the verify-before-clear rule, mergeOutstandingFindingsJSON owns the append-only merge (an outstanding item keeps its ID when a new round re-mints it; identity is still content-derived and stable IDs are a separate design pass), and the gate, persisted findings_json, round rows, and internal/db/stats.go all derive from that same set so no-mistakes stats cannot report Fixed 100% while the operator is parked. The loop itself is bounded only by auto_fix.review (the automatic-round budget) and the human/agent gate, matching upstream - no separate round or stall cap. Regressions: internal/pipeline/review_carry_forward_test.go (TestExecutor_ReviewCarryForward_NoOpFixKeepsFindingParked, TestExecutor_ReviewCarryForward_UserAddedFindingStaysOutstanding, TestExecutor_ReviewCarryForward_PositiveCoverageClearsFinding, TestResolveVerifiedFindingsJSON, TestResolveVerifiedFindingsJSON_FilelessPendingFinding, TestMergeOutstandingFindingsJSON_UsesCurrentReviewedPathsWhenRoundIsEmpty) and internal/pipeline/executor_test.go (TestExecutor_RevalidationClearsReviewCarry), internal/pipeline/executor_autofix_test.go (TestExecutor_AutoFixInfoFindingNoOpParksAfterBudget).reviewed_paths record covering every trusted reviewable path: ReviewStep.Execute sets NeedsApproval when the list is absent (nil, including Codex's null-fill of an omitted field), empty, partial, or names an out-of-scope path, and logs which files went unverified (uncoveredReviewMessage). reviewed_paths stays optional in reviewFindingsSchema only so an older payload still parses - never as a legacy pass; there is no production escape. Eval replay (internal/eval/replay.go) reads only Findings, so the park does not change scoring. The fake agent (cmd/fakeagent/scenario.go withReviewCoverage) is the sole fixture-side stand-in: on a review turn whose canned response has no reviewed_paths it fills the worktree's own reviewable set (base-commit diff minus the prompt's ignore patterns); a scenario that spells the field out - partial, empty, or fabricated - is passed through so coverage gates stay testable. Regressions: TestReviewStep_PartialReviewedPathsDoesNotGrantApproval, cmd/fakeagent/review_coverage_test.go.ReviewStep.Execute: reviewCoverageSection enumerates the trusted reviewable set in the review prompt (a checklist the reviewer verifies itself, replacing the lossy reconstruct-from-memory round trip), and completeCoverageGaps runs ONE bounded focused completion turn over exactly the uncovered files when a clean round would otherwise park solely on coverage. The completion prompt is the full review prompt plus reviewCoverageCompletionSection (so a cold/resumed turn is self-sufficient), merges via mergeReviewedPaths/uncoveredReviewablePaths in common_fix.go, and is auditable under the review-coverage purpose, which review_roles.go routes through the reviewer chain like review (and which eval/capture.go's baselineForRound counts alongside review when it freezes a round's cost baseline). Eval replay sets StepContext.EvalReplay, and completeCoverageGaps skips the completion turn there: replay scores the review's findings against captured gold and never consumes the coverage certification, so spending an invocation the captured baseline does not charge would double the candidate's recorded cost, while the strict coverage check still parks the incomplete record. The strict coverage check re-runs on the union, so whatever is still uncovered - or fabricated in either turn - parks with the remainder named; a completion-turn agent error or invalid output parks on the incomplete record rather than failing the run, and rounds with blocking findings skip the pass entirely. mergeReviewRisk reconciles risk_level/risk_rationale/risk_scope as whichever turn is worse (by level rank), so a completion turn that finds a real defect in a file the first pass never read overrides a stale "low" label rather than leaving it beside the new finding. The completion turn's prompt still carries the question protocol, so appendOpenReviewQuestionFindings (the shared helper both turns call) re-reads the conversation after it returns - a question the completion turn asks is not dropped just because the first turn already read the conversation once. mergeReviewedPaths keeps a single invalid (empty/whitespace/.-only) reviewed_paths entry across the union rather than folding it away, and completeCoverageGaps skips the completion turn entirely when the first record already carries one (hasInvalidReviewedPath), so the pass can never launder an already-invalid coverage record into one that reads as complete. Both coverage prompt sections (reviewCoverageSection, reviewCoverageCompletionSection) render each branch-controlled path through coveragePathLine, which escapes newlines, carriage returns, tabs, and other control runes so a crafted filename cannot read as a second prompt instruction. The completion turn is held to the same pre-push ownership boundary as the first pass: stripDeferredPipelineOwnedDeliveryFindings runs on its output too before the merge. mergeReviewRisk moves risk_level/risk_rationale/risk_scope only when the completion turn has a remaining finding of its own: a tied or higher level with a real defect still overrides the first pass's stale "clean" label, but a turn with no remaining finding - including one whose only finding was stripped as deferred pipeline-owned delivery after the merge - never raises the published risk. Regressions: internal/pipeline/steps/review_coverage_retry_test.go (TestReviewStep_PartialCoverageCompletesInOneFocusedPass, TestReviewStep_MultiRoundCoverageOmissionsConvergeWithoutAWaiver, TestReviewStep_CoverageCompletionFailureStillParksExplicitly, TestReviewStep_CoverageCompletionFabricationStillParks, TestReviewStep_BlockingFindingsSkipTheCoverageCompletion, TestReviewStep_ReviewPromptEnumeratesTheCoverageContract, TestReviewStep_CoverageCompletionOverridesStaleRiskLabel, TestReviewStep_CoverageCompletionQuestionParksTheRound, TestReviewCoverageSectionsEscapeBranchControlledPaths, TestReviewStep_InvalidCoverageEntryNeverCertifiesTheHead, TestReviewStep_CoverageCompletionDropsDeferredPipelineOwnedDeliveryFindings, TestReviewStep_CoverageCompletionReconcilesEqualRiskRationale, TestReviewStep_CoverageCompletionDoesNotRaiseRiskWithoutARemainingFinding, TestReviewStep_EvalReplaySkipsTheCoverageCompletion).mergeOutstandingFindingsJSON hands a freed ID to the next new item via nextFreeReviewFindingID once the original finding drops out of the outstanding set, so blindly retaining any historical SelectedFindingIDs entry present in the latest findings (by ID alone) can silently treat a never-selected finding as carried. selectedFindingIdentities records each selected ID's content identity (findingKey) as of the round that reported it; retainFindingIDsByIdentity (both in internal/pipeline/findings.go, used by Executor.durableExecutionState and Executor.recoveredGate) only keeps an ID whose CURRENT finding still matches that recorded identity. Scoped to recovery only - the live in-process loop's own retainFindingIDs calls are unchanged. Regressions: TestRetainFindingIDsByIdentity_RejectsPositionalIDReuse, TestRetainFindingIDsByIdentity_KeepsGenuineSameFindingAcrossRounds.Parked / Awaiting-Agent Signal
runs.awaiting_agent_since is non-nil iff a step is actually parked at an awaiting_approval/fix_review gate: the executor sets it on gate entry, clears it when waitForApproval returns, and RecoverStaleRuns clears it on crash recovery. It is observability only (rendered as awaiting_agent: parked <duration> in axi status) and never changes gate resolution, auto-resume, or the --yes default.internal/db/run_test.go, internal/pipeline/executor_approval_test.go, internal/cli/axi_test.go, e2e TestAxiParkedAwaitingAgentSignal.Approval Override Is Never a Silent Clean Pass (pipeline.ApprovalOverrideVerifier)
ActionApprove sites in internal/pipeline/executor.go (live wait and Resume) route through Executor.applyApprovalOverride before completing the step.
ApprovalOverrideVerifier supplies the unresolved condition: CI re-checks live checks; Test reads its persisted configured-command failure.
A verification error records an unresolved condition, not a clean pass; approval proceeds unless the durable evidence write itself fails.
Test approval's operator explanation is separate local evidence in step_results.approval_reason (NULL = no recorded approval, empty = approval without a reason); db.StepResult.TestOverrideReason owns its snapshot/event qualification: only a failing configured command, a no-go/inconclusive verdict, or a Test-agent invocation-budget cut qualifies, so a no-surface acknowledgement completes normally.
CI overrides and explicit Test exceptions produce passed-with-override through outcomeForRun.
Do not reuse the configured-command override_reason for all Test approvals: that would expand PR enforcement to evidence-only exceptions and change approval policy.
The supported --reason input and reset semantics are owned by docs/src/content/docs/reference/cli.md.
A configured-command Test override is copied onto the PR attestation as steps[].override_reason; verify.py refuses it unless trusted test.allow_approve_over_failure supplies a recorded reason (allow_test_command_override on the attestation).
A CI-repair restamp keeps the prior Test override_reason but applies the current trusted opt-in; it does not copy a previous attestation's allow_test_command_override.run_completed IPC delta (ipc.Event.CIOverrideReason and TestOverrideReason, derived in emitRunEvent from step rows and gated on run.Status == RunCompleted), so an attached TUI's renderOutcomeBanner shows ⚠ Pipeline passed with override on the live event path without waiting for a snapshot - otherwise the banner disagrees with axi's passed-with-override.internal/e2e/test_approval_exception_test.go; recovery: internal/pipeline/test_approval_exception_test.go; migration and reset semantics: internal/db/test_approval_reason_test.go.internal/pipeline/executor_approval_override_test.go (incl. TestExecutor_ApprovalOverride_RunCompletedEventCarriesReason), internal/pipeline/steps/ci_approval_override_test.go, internal/pipeline/steps/test_approval_override_test.go, internal/pipeline/steps/prsummary_test_override_test.go, require_no_mistakes_test_override_test.go, internal/cli/axi_test.go (TestOutcomeForRun), internal/tui/action_bar_test.go (TestOutcomeBanner_CIOverrideShowsReason), internal/tui/events_test.go (TestModel_ApplyEvent_RunCompletedCarriesCIOverride).14e8dd1
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.