Executes QA test cases created by /aif-qa in human-guided or automated-agent mode, including reusable browser replay scripts. Use when you need to walk through QA one case at a time, record pass/fail results, or have an agent verify and rerun cases through browser, CLI, API, automated tests, or file/document checks.
60
70%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Fix and improve this skill with Tessl
tessl review fix ./skills/aif-qa-check/SKILL.mdRuns the test-cases.md artifact produced by /aif-qa and records execution status in qa-check.md.
In agent mode, it also maintains reusable cross-QA execution memory under the QA root:
agent-context.md — curated, current, reusable non-sensitive setup facts for future automated QA runsagent-history.md — append-only reusable learnings extracted from prior runs, not a per-run audit logFor browser/UI cases, agent mode saves executable replay scripts under the branch-specific browser-replay/ directory. Later runs execute those scripts first, including scripts for previously passed cases, so fixes are retested and prior behavior gets regression coverage without regenerating browser automation.
The skill is stack-free and agent-free: it does not assume a framework, package manager, browser tool name, or agent runtime. In agent mode, it uses the appropriate available execution surface for each case: live browser automation, CLI commands, project test runners, HTTP/API calls, database-safe read checks, file/document inspection, or other non-destructive automated verification.
| Argument | Mode | What you do |
|---|---|---|
human | Human-guided QA | Show exactly one test case, ask the user whether it works, and record their answer |
agent | Automated-agent QA | Execute test cases through the most appropriate available capability and record observed results |
If no mode is provided, ask the user which mode to run.
FIRST: Read .ai-factory/config.yaml if it exists to resolve:
paths.description, paths.architecture, paths.qa (default: .ai-factory/qa)language.ui for AskUserQuestion prompts, progress messages, summaries, and next-step guidancelanguage.artifacts for the persisted qa-check.md artifactlanguage.technical_terms for human-readable technical terminology style in the artifactlanguage.artifacts is missing, use language.uiengit.enabled for branch resolutionIf config.yaml does not exist, use defaults:
.ai-factory/DESCRIPTION.md.ai-factory/ARCHITECTURE.md.ai-factory/qa/ui_language: enartifact_language: entechnical_terms_policy: keeptrueStore:
ui_language = language.ui || "en"artifact_language = language.artifacts || language.ui || "en"technical_terms_policy = language.technical_terms || "keep"git_enabled = git.enabled when present, otherwise trueqa_root = resolved paths.qa || ".ai-factory/qa/"qa_agent_context_path = <qa_root>/agent-context.mdqa_agent_history_path = <qa_root>/agent-history.mdAll AskUserQuestion prompts, user-visible explanations, per-case summaries, and final summaries MUST be written in ui_language.
The persisted qa-check.md artifact MUST be written in artifact_language.
Templates define structure, not language. Use the canonical English template in templates/QA-CHECK.md. If artifact_language is not en, translate headings, labels, status text, comments you author, and explanatory text to artifact_language before saving. Preserve checkbox syntax, test case IDs (TC-001), commands, paths, config keys, URLs, selectors, package names, API names, branch names, and raw error messages.
For artifact_language = ru, write human-readable prose, headings, statuses, summaries, and agent-authored comments in Russian. Preserve user wording except mandatory redaction of sensitive values before writing.
Apply technical_terms_policy while writing artifacts:
keep — keep common technical terms such as browser, selector, viewport, endpoint, payload, regression, and fixture when clearertranslate — translate human-readable technical terms where a natural target-language term existsmixed — translate ordinary prose terms while keeping code, infrastructure, and ecosystem terms unchangedRead the resolved description path if it exists.
Read the resolved architecture path if it exists.
Read .ai-factory/skill-context/aif-qa-check/SKILL.md — MANDATORY if the file exists. Treat it as project-level overrides for this skill.
Read <paths.qa>/agent-context.md if it exists. Treat it as reusable, user-approved QA execution context for future automated runs: stable environment classes, canonical local/staging URLs, login route, test account role, allowed non-sensitive usernames, required seed data patterns, stable selectors, field identifiers, known redirects, startup prerequisites, safe command patterns, reusable test-filter conventions, environment notes, and service startup instructions.
Read <paths.qa>/agent-history.md if it exists. Treat it as append-only cross-QA memory from prior /aif-qa-check runs. Before asking the user or attempting automated execution, search it for reusable prior answers to the same blocker so the skill does not repeat avoidable mistakes.
agent-context.md and agent-history.md are global to all QA sets under paths.qa. They MUST NOT become logs for a specific QA plan. Do not write branch names, branch slugs, QA target paths, current artifact_dir, TC-* mappings, per-run summary counts, exhaustive command transcripts, assertion counts from a single run, one-off file checks, or plan-specific labels such as a numbered QA folder name. Write those details only to <paths.qa>/<branch-slug>/qa-check.md.
Only promote information into agent-context.md or agent-history.md when it is likely to help future unrelated QA runs. Prefer stable patterns over concrete one-off run facts. Example: Laravel backend checks use php artisan test --filter=<TestClass> is reusable; 05-profile-formula-recommendations ran ProfileReaderTest with 4 tests / 13 assertions is run-specific and belongs only in qa-check.md.
Never write global negative capability or scope claims unless they are true for the whole project. Do not write statements like browser automation is not required, browser URL is not used, all QA is backend/CLI, or these cases are backend-test/cli/file-docs to agent-context.md or agent-history.md when that conclusion came from one QA plan. Instead, keep it branch-specific in qa-check.md, or phrase a reusable conditional rule such as For backend-only QA cases, browser automation is not required; browser/UI cases still require Browser or Playwright MCP.
Never treat either file as a source of secrets. If either file contains an apparent password, token, cookie, authorization header, one-time code, or other secret, ignore that value for execution and redact it the next time the file is rewritten or appended to.
Parse $ARGUMENTS:
human or agent; remove it from arguments.ui_language which mode to run:
Resolve the working branch:
If git_enabled = false or the repository is not a git work tree:
If branch was provided in arguments → use it as the resolved branch label
Otherwise → set resolved_branch = "manual"
If git_enabled = true and the repository is a git work tree:
If branch was provided in arguments → use it as the resolved branch
Otherwise → run: git branch --show-currentStore:
resolved_branchartifact_dir = <resolved paths.qa>/<branch-slug>test_cases_path = <artifact_dir>/test-cases.mdqa_check_path = <artifact_dir>/qa-check.mdbrowser_replay_dir = <artifact_dir>/browser-replayCompute branch-slug with the exact same algorithm as /aif-qa:
[A-Za-z0-9._-] with -, collapse consecutive -, trim leading/trailing -, and use branch if empty. Then MUST truncate safe_slug to the first 40 ASCII characters. Because the normalized safe_slug alphabet is [A-Za-z0-9._-], byte length and character length are identical.git hash-object --stdin <<< "<resolved_branch>" and take the first 8 hex characters as hash8.branch-slug = "<safe_slug>-<hash8>".Check for <artifact_dir>/test-cases.md.
If it is missing, STOP and ask in ui_language whether to:
/aif-qa test-cases <resolved_branch> firstRead test-cases.md. Extract test cases by IDs (TC-001, TC-002, etc.), titles, priority, type, execution surface when present, preconditions, steps, expected results, and test data. Preserve their order.
Compute source binding metadata before creating or modifying qa-check.md:
source_digest — deterministic digest of the full test-cases.md content using git hash-object --no-filters <test_cases_path> when available, or git hash-object --stdin over the exact file content.case_digests — deterministic per-case digest for each extracted TC-NNN, computed from that case's canonical block including title, priority, type, preconditions, steps, expected result, and test data.tested_revision — when git_enabled = true and the repository is a git work tree, run git rev-parse HEAD and record the resolved commit SHA.worktree_digest — when git_enabled = true and the repository is a git work tree, record a deterministic digest of the current working tree state so dirty-tree QA cannot be reused after local changes without a commit.manual_build_id — when git_enabled = false or the repository is not a git work tree, ask the user for an explicit build/version identifier before creating or resuming results. Do not accept an empty identifier.replay_script_digests — for each existing canonical <browser_replay_dir>/TC-NNN.js, hash the exact file content with git hash-object --no-filters <path> when available, or another stable SHA-1/SHA-256 digest. This binds evidence to the automation that actually ran even though QA-owned replay files are excluded from worktree_digest.Canonicalize each per-case digest input exactly:
TC-NNN identifier through the line before the next TC-NNN block or end of file.BEGIN TC-NNN\n<normalized block>\nEND TC-NNN\n.git hash-object --stdin when available, or another stable SHA-1/SHA-256 digest if git is unavailable.Compute worktree_digest exactly when git is enabled and a git work tree exists:
git status --porcelain=v1 --untracked-files=all.git diff --binary HEAD --.qa_check_path and every file under browser_replay_dir from the status, diff, and untracked-file digest inputs so QA-owned result/replay artifacts do not stale their own results. Do not exclude test_cases_path; source changes are also tracked by source_digest.UNTRACKED <path> <content-digest> where <content-digest> is git hash-object --no-filters <path> when the file is readable.git hash-object --stdin.clean\n.If mode = agent, perform Step 1.1 before creating or modifying qa-check.md. Existing qa-check.md may be inspected read-only during this gate.
If qa-check.md exists, read it and resume from existing statuses only after comparing stored binding metadata to the current binding metadata:
tested_revision changed, mark every prior result as Stale and unchecked, preserve prior comments/evidence as historical context, and require retest. Do not count stale pass/fail/block statuses as current.worktree_digest changed, mark every prior result as Stale and unchecked, preserve prior comments/evidence as historical context, and require retest. Do not count stale pass/fail/block statuses as current.manual_build_id changed, treat it the same as a tested revision change.source_digest changed, compare case_digests. Preserve current status only for cases whose per-case digest is unchanged and whose tested revision/manual build id and worktree digest are unchanged.Stale, unchecked, and require retest.Pending.test-cases.md, keep its historical entry marked Stale or move it to an artifact-language "Stale / Removed Cases" section; never count it as current.Script digest recorded for that case, mark the case Stale, preserve the previous replay metadata/evidence, and require a proof run of the changed script. Missing legacy script-digest or target-fingerprint metadata means the script is unproven, not matching.If qa-check.md does not exist, create it from templates/QA-CHECK.md using the extracted test cases and source binding metadata. Every case starts unchecked and Pending.
Agent mode is user-only (disable-model-invocation: true) because it can perform live browser actions, shell commands, local service calls, and other checks with meaningful side effects.
Before executing any case or writing qa-check.md in agent mode:
browser-ui, cli, backend-test, api, file-docs, database-read, hybrid, human, or unknown. If the case has an explicit Execution surface: field, use it unless the case text clearly contradicts it; record any override in agent-history.md.page script when replay scripts exist.Blocked, ask the user in ui_language to enable a live browser capability, and append this blocker to <paths.qa>/agent-history.md. Continue with other cases that can be verified through CLI, tests, API, or file/document checks.<paths.qa>/agent-context.md, <paths.qa>/agent-history.md, test-cases.md, change-summary.md, project docs, current browser state, and repository scripts to determine the target URL, service startup commands, test commands, and environment.agent-context.md, agent-history.md, test data, seeders/factories/fixtures, docs, local database-safe reads, and existing test users for a matching non-sensitive test identity.local, test, or disposable development, and repository tooling provides a safe way to create or seed a synthetic test user/state, create the minimal disposable fixture needed for the case. Record the fixture identifier and setup command/evidence in qa-check.md.local or test environment, use them for this run. Store credentials in agent-context.md only if the user explicitly says they are reusable test-only credentials and safe to persist. Otherwise store only the username/role and where to retrieve the secret.agent-context.md or agent-history.md.<paths.qa>/agent-context.md immediately only when the answer is general enough to help future QA sets. Append the redacted answer summary to <paths.qa>/agent-history.md only when it captures a reusable lesson; otherwise keep it in the branch-specific qa-check.md evidence/comment.local, development, staging, test, production, or unknown. For a browser target:normalized_base_url from the resolved target URL by removing user information, query, and fragment; preserving scheme, host, port, and deployment base path; and removing a trailing slash except for the origin root.target_fingerprint from the exact canonical input AIF TARGET\n<environment_class>\n<normalized_base_url>\n with git hash-object --stdin when available, or another stable SHA-1/SHA-256 digest.target_fingerprint in qa-check.md. Never treat a replay created for another target fingerprint as matching.production or unknown targets without explicit user authorization immediately before execution. Record the authorization decision and target class in qa-check.md. Append it to agent-history.md only when it creates a reusable cross-QA policy or recurring authorization lesson; do not persist broad production authorization for future runs.rm, git reset, git checkout --, database resets, migrations against shared environments, deploy commands, or cleanup commands that delete user data unless the user explicitly authorizes the exact action.Blocked, and write the blocker to qa-check.md without executing the risky action. Append the decision to agent-history.md only when it is a reusable cross-QA lesson, not merely a case-specific denial.Redaction is mandatory for agent comments/evidence and all human-entered comments/evidence:
token, access_token, refresh_token, id_token, code, secret, password, passwd, pwd, auth, key, api_key, session, sid, and jwt.[REDACTED] before writing comments or evidence to qa-check.md.Run exactly one pending or selected test case at a time.
For each case:
ui_language.ui_language = ru, use exactly: Протестируйте и ответьте работает или нет.[x]) in qa-check.md, set status to Passed, and add the current mode as human.[ ]), set status to Failed, ask for the reason, and write the user's explanation as the comment while preserving user wording except mandatory redaction of sensitive values.qa-check.md after every case so progress survives context resets.Do not show the next test case until the current one has a result or the user stops.
Agent mode MUST produce concrete, case-appropriate evidence. Browser execution, CLI output, automated test output, API responses, backend command results, database-safe read results, or file/document inspection can each be enough to mark a case passed when that evidence directly covers the case steps and expected result.
Internal deterministic invariants MUST be treated as automated checks, not manual QA. Cases involving service-method calls, cache state, materialized rows, exact arrays, raw database values, formula outputs, CLI internals, or similar non-human-observable behavior should be classified as backend-test, cli, api, or database-read as appropriate. First look for existing tests or commands that already cover the invariant. If existing tests are found and pass, mark the corresponding TC-* case Passed with the test class/name, command/filter, exit status, and assertion/result summary as evidence. If one test run covers multiple TC-* cases, reuse that evidence on every covered case and optionally summarize the shared run under "Supporting Automated Checks". If no suitable automated coverage exists, keep the case Blocked with a blocker such as Missing automated test coverage for backend invariant; do not ask the user to manually verify internal arrays or raw database values.
Determine the test target:
agent-context.md, agent-history.md, test-cases.md, change-summary.md, project docs, or current browser state clearly identify a URL, use it.agent-context.md.For every browser-ui case, and every hybrid case with browser steps:
<browser_replay_dir>/TC-NNN.js as the canonical replay artifact. Keep one case per file so a failed case can be rerun independently.async (page) => { ... } expression. This is the format accepted by browser run-code capabilities. Do not add imports, a test framework, fixtures, generated wrappers, or project dependencies. Include navigation, stable selectors, inputs/actions, and explicit assertions that throw on mismatch.// aif-case-digest: <case_digest>// aif-target-fingerprint: <target_fingerprint>
Use the quoted placeholder "AIF_BASE_URL" for navigation instead of persisting an absolute target URL. Immediately before execution, replace that quoted placeholder in memory with a correctly JSON-encoded normalized_base_url; do not modify the saved file. Never write passwords, tokens, cookies, authorization values, one-time codes, personal secrets, or token-bearing URLs into the script. Reuse an already authenticated browser session or obtain secrets through the authorized runtime mechanism.matching only when all three bindings agree: embedded case_digest, embedded target_fingerprint, and actual script_digest versus the last proof recorded in qa-check.md. Verify all three before every replay. A target mismatch MUST stop execution until the currently displayed and authorized target is verified and the script is explicitly rebound; rebinding changes the script digest and requires a new proof run.script_digest, execute it once, and use that first execution as its proof run; do not automatically execute a new or updated script a second time.
Unproven. Do not immediately rerun it.Script digest as a current evidence binding only after that exact file content has executed. For an unproven file, keep Script digest: n/a, set Proof status: Unproven, and record its computed candidate digest in the comment/evidence until a proof run completes.<browser_replay_dir>/history/TC-NNN-<old_script_digest>.js. Record the archived path, old digest, and initial error in qa-check.md. Never execute history files automatically.script_digest, and require a new proof run under the precondition/side-effect rules above.Failed; never choose a merely similar element or bypass the missing step.Failed; this is regression/fix evidence, not a reason to regenerate the script.Browser replay and Evidence fields in qa-check.md. Proven means that exact script content executed; it does not mean the product result passed.Browser exploration: Ran or Browser exploration: Skipped, the reason, scope, and concise redacted findings in qa-check.md.qa-check.md and recommend adding it through /aif-qa test-cases; create its replay script after it has a TC-NNN case.For each pending or selected case:
browser-ui: execute the matching saved browser replay first; create or prove it only when missing or nonmatching by case, target, or script digest, then navigate with Browser or Playwright MCP and execute UI steps.cli: run the relevant project command and inspect exit code/output.backend-test: find and run the narrowest existing relevant test command or test filter first; broaden only when needed for confidence. Passing automated coverage is valid pass evidence for the case.api: call the endpoint through existing project tooling, safe local commands, or browser/network capability as appropriate.file-docs: inspect generated files, docs, config, or repository contents with Read/Grep/Glob and, when useful, project validation commands.database-read: use only safe read-only checks unless the user explicitly authorizes mutation in a test-scoped environment.hybrid: combine the necessary surfaces and record each evidence type.agent-context.md and agent-history.md first. For browser/UI auth or user-state gaps, try safe local/test fixture discovery or creation as described in Step 1.1 before blocking. If the answer or fixture is not already available, ask the user exactly for the missing information or authorization before marking the case blocked.agent-context.md with durable cross-QA facts and append the resolved friction to agent-history.md only if it is likely to recur in future QA sets. Keep case-specific details in qa-check.md.[ ]), set status to Blocked, and write the blocker.[x]), set status to Passed, and add concise redacted evidence. Evidence may be a browser observation, command and exit status, test name and result, API response summary, file path and matched condition, or other concrete observation.[ ]), set status to Failed, and write a concrete redacted problem comment with observed vs expected behavior.qa-check.md after every case.Before the final summary, write detailed per-run evidence to qa-check.md, not to root-level memory. This includes the current branch/QA target, commands executed, exit codes, assertion counts, supporting file checks, case mappings, summary counts, and one-off blockers.
After writing qa-check.md, extract only reusable cross-QA learnings:
<paths.qa>/agent-context.md when a discovered fact is durable and likely useful for future QA sets.<paths.qa>/agent-history.md only when a blocker, selector, command pattern, test-filter convention, API endpoint pattern, route, setup note, or authorization decision is likely to recur beyond the current QA set.agent-history.md for a successful run that produced no new reusable learning.agent-history.md.Keep <paths.qa>/agent-context.md concise and current; prefer stable facts over one-off observations. Do not write raw passwords, tokens, cookies, authorization headers, one-time codes, private personal data, token-bearing URLs, branch names, QA target paths, or run summaries.
Use screenshots or browser state observations when the active runtime makes them available, but do not require screenshots to pass a test.
After stopping or finishing, report in ui_language:
qa-check.mdbrowser-replay/ when browser scripts were created or executedagent-context.md and agent-history.md when agent mode ran or requested missing execution contextIf mode = agent and one or more current cases are Blocked, split them into human-verifiable blocked cases and automation-only blocked cases before ending.
human, browser-ui, hybrid with human-observable steps, or unknown.backend-test, cli, api, file-docs, or database-read cases when the blocker is missing automated coverage, missing command/test filter, unavailable service, or another technical automation prerequisite. For those cases, recommend adding/running the appropriate automated check or providing the missing command/fixture/context.ui_language whether to run those eligible blocked cases now in human-guided mode.Blocked.qa-check.md. Append to agent-history.md only when the handoff revealed a reusable cross-QA lesson.If any case failed, the next recommended action should be to fix the issue and rerun /aif-qa-check <mode> <resolved_branch>.
If human-verifiable cases remain blocked after an agent-mode run and the user did not continue them in human mode, the next recommended action should be to run /aif-qa-check human <resolved_branch> for those cases or provide the missing execution context. If automation-only cases remain blocked, recommend adding/running the required automated test, command, fixture, service, or context instead. Do not imply that a blocked case is a product defect unless there is concrete observed evidence.
<paths.qa>/<branch-slug>/qa-check.md, canonical <paths.qa>/<branch-slug>/browser-replay/TC-NNN.js, preserved <paths.qa>/<branch-slug>/browser-replay/history/*.js, <paths.qa>/agent-context.md, and <paths.qa>/agent-history.md.<paths.qa>/<branch-slug>/change-summary.md, test-plan.md, test-cases.md, <paths.qa>/agent-context.md, and <paths.qa>/agent-history.md as QA context.qa-check.md, branch-specific browser-replay/TC-NNN.js and browser-replay/history/*.js, agent-context.md, and agent-history.md; do not rewrite test-cases.md, test-plan.md, change-summary.md, or config.yaml.paths.description, paths.architecture, paths.qa, language.ui, language.artifacts, language.technical_terms, and git.enabled; never writes config.yaml.test-cases.md.tested_revision and worktree_digest) or manual_build_id, plus source_digest and case_digests.agent-context.md and agent-history.md before automated execution and before asking the user for setup facts, so prior answers are reused.agent-context.md only after explicit user permission, and MUST never persist production credentials, personal credentials, cookies, session tokens, one-time codes, or shared secrets.agent-context.md, and MUST append to agent-history.md only when a recurring reusable learning or blocker pattern was discovered.Blocked cases.browser-replay/TC-NNN.js scripts and execute matching scripts before generating new browser actions.case_digest, target_fingerprint, and script_digest before replay and require a proof run after any script change or target rebinding.ac92beb
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.