CtrlK
BlogDocsLog inGet started
Tessl Logo

qa-quarto

Adversarial Quarto-vs-Beamer parity QA. A critic agent compares the Quarto HTML render to the Beamer PDF benchmark for content/visual parity; a fixer agent applies fixes; loops until APPROVED (max 5 rounds). Use when user says "qa the quarto", "check parity", "does the html match the pdf?", "quarto matches beamer?", or after a translate-to-quarto run. Requires both the `.qmd` rendered and a `.pdf` benchmark.

Invalid
This skill can't be scored yet
Validation errors are blocking scoring. Review and fix them to unlock Quality, Impact and Security scores. See what needs fixing →
SKILL.md
Quality
Evals
Security

Adversarial Quarto vs Beamer QA Workflow

Compare Quarto HTML slides against their Beamer PDF benchmark using an iterative critic/fixer loop.

Philosophy: The Beamer PDF is the gold standard. The Quarto translation must be at least as good in every dimension.


Workflow

Phase 0: Pre-flight → Phase 1: Critic audit → Phase 2: Fixer → Phase 3: Re-audit → Loop until APPROVED (max 5 rounds)

Hard Gates (Non-Negotiable)

GateCondition
OverflowNO content cut off
Plot QualityInteractive charts >= static plots
Content ParityNo missing slides/equations/text
Visual RegressionQuarto >= Beamer in all dimensions
Slide CenteringContent centered, no jumping
Notation FidelityAll math verbatim from Beamer

Phase 0: Pre-flight

  1. Locate Beamer (.tex/.pdf) and Quarto (.qmd/.html) files
  2. Check freshness (re-render if QMD newer than HTML)
  3. Verify TikZ SVGs if applicable

Phase 1: Initial Audit

Launch the quarto-critic agent to compare Beamer vs Quarto comprehensively. Report saved to quality_reports/[Lecture]_qa_critic_round1.md.

Phase 2: Fix Cycle

If not APPROVED, launch quarto-fixer agent to apply fixes (Critical → Major → Minor), re-render, and verify.

Phase 3: Re-Audit

Re-launch critic to verify fixes. Loop back to Phase 2 if needed.

Iteration Limits — loop-until-dry

This is the loop-until-dry primitive from orchestrator-protocol.md: the critic returns FINDINGs (the hard-gate table is the CRITICAL roll-up, per orchestration-schemas.md); the loop converges when a round adds 0 new CRITICAL/MAJOR findings (deduped on id = sha1(file:line:locus)), not at a fixed round count.

  • Fallback cap: 5 rounds bounds a non-converging loop, then escalate to the user with remaining issues.
  • Two-strikes: the same gate failing in rounds N and N+2 is flagged for the user, not patched again (summary-parity.md).
  • APPROVED iff every hard gate passes (zero CRITICAL).

Final Report

Save to quality_reports/[Lecture]_qa_final.md with hard gate status, iteration summary, and remaining issues.

Findings are validated, not just written (v2.5)

This skill's reviewers emit findings under the machine-checked contract in finding-schema.json. Reports are JSON arrays.

Smoke-test the harness before spending review effort — a run that fans out reviewers and then cannot write a valid report has wasted the whole pass:

echo '[]' | python3 scripts/validate-findings.py

Then, before presenting any summary:

python3 scripts/validate-findings.py <report>.json   # exit 0 required

What the contract forces, and why:

  • rule — the documented rule or standard violated. A finding citing no rule is an opinion, and opinions do not gate a commit.
  • failing_case — a concrete configuration under which the claim breaks, or the exact missing hypothesis. "This could be clearer" does not validate.
  • id = sha1("<file>:<line>:<locus>") — deterministic, so dedup across rounds is exact and the two-strikes rule is checkable rather than eyeballed.
  • mechanicaltrue only for fixes that cannot change a result (typo, cross-reference, formatting, label). Never for an estimand, assumption, specification, inference procedure, sample definition, or reporting language: those return to the researcher.

Apply the per-lens evidence burdens and the "does NOT count" filters in orchestration-schemas.md §7 before verification, so known false alarms never reach the judge. The verifier pass is refute-biased: only verdict: "confirmed" findings ship; anything it cannot ground is dropped, not downgraded to a warning.

Repository
pedrohcgs/claude-code-my-workflow
Last updated
First committed

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.