Design deterministic extract-validate-retry loops with actionable validation errors, recoverable vs not-recoverable classification, bounded retry budgets, systematic-error detection, escalation packets, and offline validation.
Design deterministic extract -> validate -> retry-with-error-feedback loops that never retry blindly. The loop reinjects the exact validation error, classifies recoverable vs not-recoverable failures, enforces a retry budget, detects systematic repeat errors, and escalates with a complete error chain instead of silently accepting a failed output. [DOC]
Use when all three hold; if any is false, this skill is overkill. [INFERENCE]
Inputs [DOC]
base_prompt and the structured contract the output must satisfy (schema/format/range).{ok, code, message, path, recoverable} — not a bare boolean.max_retries (1–3) and the policy assets below.Outputs (exactly one terminal state) [DOC]
{"status":"ok", "output", "attempts"} — validation passed.{"status":"escalated", "reason", "error_chain", "last_output", "fix_hint?"} where reason ∈ {not_recoverable, systematic, budget_exhausted}.Never a third state. Returning the last failed output as success is a defect, not an output. [DOC]
Validate plans against assets/retry-loop-contract.json using scripts/validate_retry_plan.py. The validation MUST run offline and deterministic (offline=true, network_required=false, deterministic=true) so the same plan always yields the same verdict. [DOC] A valid plan includes: [DOC]
code, message, path, and recoverability.recoverable and not_recoverable categories.max_retries between 1 and 3.reason, error_chain, and last_output.recoverable (format, parse, range, schema shape) or not_recoverable (missing datum, source contradiction). Only recoverable errors retry. [DOC]max_retries (2–3). Keep a counter and an accumulated error chain. [DOC]max_retries capped at 3. Beyond 3, recoverable errors are almost always already fixed; further attempts mask a systematic defect and burn cost. The cap forces the systematic/escalation path instead of hope. [INFERENCE]Ship only when every item holds; each maps to an expected_check. [DOC]
error_feedback, specific_error) [DOC]true/false. (quality_criteria) [DOC]recoverability) [DOC]1 <= max_retries <= 3, with counter and error chain. (retry_budget, budget_limit) [DOC]systematic_detection) [DOC]silent_failure_blocker) [DOC]blind_retry_blocker) [DOC]assets/* policies referenced and scripts/* checks pass. (assets, deterministic_scripts) [DOC]upgrade_safety) [DOC]Reference assets/error-feedback-policy.json, assets/recoverability-policy.json, assets/retry-budget-policy.json, assets/systematic-error-policy.json, assets/escalation-policy.json, and assets/anti-pattern-policy.json. [DOC] Hard rules: [DOC]
# GOOD: informed retry, classified failure mode, bounded budget, escalation
def run_with_retry(task, max_retries=3):
errors, prev_output = [], None
for attempt in range(max_retries):
prompt = build_prompt(task, prev_output, errors[-1] if errors else None)
output = model.run(prompt)
verdict = validate(output) # -> {"ok", "error", "recoverable"}
if verdict["ok"]:
return {"status": "ok", "output": output, "attempts": attempt + 1}
if not verdict["recoverable"]:
return escalate(reason="not_recoverable", errors=errors + [verdict["error"]])
errors.append(verdict["error"])
prev_output = output
if is_systematic(errors): # same error repeating -> structural defect
return escalate(reason="systematic", errors=errors, fix_hint="schema/prompt")
return escalate(reason="budget_exhausted", errors=errors, last_output=prev_output)
def build_prompt(task, prev_output, last_error):
if prev_output is None:
return task.base_prompt
# reinject the SPECIFIC error, not the original prompt unchanged
return (
f"{task.base_prompt}\n\n"
f"Your previous output failed validation: {last_error}\n"
f"Previous output:\n{prev_output}\n"
f"Fix only what failed; keep everything else."
)# ANTI: blind retry + silent failure
def run_bad(task, max_retries=3):
for _ in range(max_retries):
output = model.run(task.base_prompt) # same prompt every time, no feedback
if validate_bool(output): # boolean only, no error reason
return output
return output # accept the last failed output silently, no escalationWhy it fails: resending the original prompt without the error makes the model repeat the same failure; a boolean validator has nothing to feed the retry; and returning the last failed output without escalating hides the defect downstream. [INFERENCE]
code/message must distinguish these so feedback is targeted; a generic "invalid" wastes the retry. [INFERENCE]budget_exhausted; consider tightening the fix-instruction. [INFERENCE]max_retries=1 is valid (single corrective attempt) but means systematic-detection cannot fire; rely on the budget-exhausted escalation. [INFERENCE]Self-correct if you catch yourself: resending an unchanged prompt; returning a non-ok output without an escalation packet; setting max_retries > 3 "so the model eventually fixes itself" (that is the systematic-defect signal, not a budget problem); or validating with a boolean. [INFERENCE]
python3 skills/validation-retry-design/scripts/validate_retry_plan.py --input <plan.json>
bash skills/validation-retry-design/scripts/check.shkatas-26katas-validation-retry-feedbackindependent-review-design, workflow-forgefdad39c
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.