Write, fix, or review the browser journeys in tests/stand/ui/ — Playwright (Python, sync API) against a deployed Insight stand with a real Keycloak sign-in. Covers the governing UI-vs-API rule and how to justify a browser test in writing, accessibility-first locators, the page-object/flows split, manifest-derived expectations, URL/deep-link validation, and real browser downloads. Use for any request to write, add, fix or review a committed browser test in this repository — 'add a Playwright test for X', 'the UI test is failing', 'turn this scenario into a browser journey', 'add a deep-link or download regression' — and for anything under tests/stand/ui/. For HTTP contract tests use stand-api-test; playwright-cli is for driving a browser interactively at a prompt and for Playwright's own CLI — reach for it only when nothing will be committed under tests/stand/ui/; drive-ui is for looking at a stand by hand.
74
92%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
tests/stand/ui/)Playwright (Python, sync API) against a deployed Insight: real Keycloak login, the real SPA bundle, the real gateway. Each journey is a statement about a complete round trip from a cold browser.
Environment and triage: insight-stand. What to test: stand-scenarios.
A new UI test must state, in writing, why it cannot instead be an API test.
A browser is slow, flakier than an HTTP call, and exercises more surface than most assertions need — so the suite would rather have one more API test than one more browser test whenever the two would prove the same thing.
State it as a paragraph in the module docstring, backed by something you measured: a cookie only a real browser can set, a client-side route change no HTTP call performs, a redirect chain that only exists post-render, a rendering step the API answer passes through. The shipped modules are the model:
Why this is a browser test and not an API test, measured rather than asserted: every SPA route answers 200 text/html to an anonymous HTTP client… Refusal exists only inside the browser: the SPA boots, asks
/auth/me(401), and the root route'sbeforeLoadsends the window to/auth/login…
/v1/subchartalready proves the API knows who reports to a lead… What no API call can show is that the deployed SPA takes that answer, renders it in the signed-in person's view, and renders it as navigable links. A stand where identity is perfect and the frontend renders an empty shell passes every API test in this repository.
"The UI should be covered too" is not a justification. If you cannot name the measurement, write an API test.
Take the measurement before writing the paragraph. The shipped ones came
from curling every SPA route anonymously, inspecting window either side of a
transition, and counting locator matches per name — drive-ui or
playwright-cli for in-browser probes, plain curl for what an HTTP client
sees. Quote the number you got. A plausible paragraph with nothing behind it is
exactly the failure this rule exists to prevent.
A crafted deep-link earns a browser journey when the claim is about what the
router does with URL state before anything renders. validatePortalSearch
drops what it cannot parse and assertDateRange throws into
AppErrorBoundary, so the same malformed link either degrades quietly or
replaces the whole shell — and no HTTP client can tell you which, because an
anonymous GET of any SPA route answers 200 text/html either way. A fe-unit
test of the validator proves the function refuses the value, not that the
deployed router invokes it. So pin the exact accepted/refused boundary in
fe-unit and stand-api, and keep one journey that loads the refused URL cold
and asserts the shell survives with an explicit fallback or validation message,
and that no request carrying the refused value was sent.
For a browser-generated report or export, expect_download() is only the
transport assertion. Save the file, parse it, and compare the same semantic
table across every format offered: headers, all rows, missing-versus-zero cells
and representative values derived from the preview or the API response. Assert
the filename separately. A suffix and a non-zero byte count do not prove the
file holds the requested scope, period, granularity or metrics — and CSV and
XLSX are generated by different code, so a divergence between them is exactly
the kind of defect only cross-format parsing sees.
downloads.py does the reading: download_export clicks the menu item, saves
what arrives and parses it, and every cell — CSV string, XLSX number, DOM text —
goes through one normalizer, so 29 and 29.0 compare equal while an empty
cell stays distinct from a zero. claimed_row_count reads the grid's own
aria-rowcount, which is how a virtualized grid tells you the row total it is
only showing a window of. The three journeys that download a file
(test_git_output_timeseries, test_collaboration_card_evidence,
test_metric_evidence_drilldown) are the worked examples: formats compared
against each other, record count against aria-rowcount, rendered rows against
exported ones, and not one expectation written down.
Keep the permanent browser path small. Batching caps, serializer escaping and
option permutations belong in fe-unit and fe-component; period, metric-count
and export-size caps belong in stand-api — where the boundary is still
unclaimed, since tests/stand/api/analytics/__init__.py deliberately clamps
every request one day under MAX_PERIOD_DAYS rather than testing it. The stand
journey proves only what those suites cannot: the deployed router does not
crash, controls reach the real backend, the preview renders, and the browser
receives semantically correct files.
from .flows import sign_in
sign_in(page, base_url, session_for("dev_lead"))Every session_for / requires_seed name is a key in the manifest's
fixtures{} catalogue — read src/ingestion/tools/seed/PROFILE.md for the list. Guessing is
not a soft failure: an unknown name aborts collection for the whole session.
sign_in drives the deployed OIDC chain with no shortcut at any step: an
unauthenticated visit to / starts authorization-code+PKCE by itself, Keycloak
serves its real form, the authenticator sets __Host-sid at the callback.
Nothing is minted.
It lives in flows.py, not in a fixture, on purpose. A signed_in_page fixture
would share one authenticated page across journeys — and each journey is a
statement about a complete round trip, so sharing would make the later ones
depend on the earlier ones having run.
The base URL is not set in ui/conftest.py. The root conftest resolves the
stand once into pytest-base-url's base_url, which pytest-playwright already
reads to configure every context — so page.goto("/") lands on the stand, and a
journey needing the address asks for base_url and gets the one the browser was
actually given. That URL must be a trustworthy origin — any https://
stand qualifies, and over plain HTTP only localhost does, so a plain-HTTP
runner joins the gateway's namespace and uses localhost:<port>. The
__Host- cookie trap and its 401→503 loop are in insight-stand.
No data-testid was found on any view these journeys touch, so roles and
accessible names are the stable handles:
page.get_by_role("dialog", name="Git output")
dialog.get_by_role("table").filter(has_text="PRs merged")
page.get_by_role("link", name=display_name)Two exceptions the shipped page objects rely on, so treat them as permitted rather than as violations to clean up:
data-slot attributes ([data-slot='card'], card-title, and the
xpath=ancestor::*[@data-slot="card"][1] walk in person_view.py). These are
structural attributes the component library emits, stable across restyles —
unlike a hashed class, which is what the rule is actually aimed at.table.get_by_role("rowgroup").nth(1).get_by_role("row").first. Position is
the only thing that distinguishes a header row from a data row; anchoring on
the table by role first is what keeps it honest.What stays out: hashed CSS classes, Tailwind utility selectors, and any chain that starts from the page rather than from a named element.
Locate by the thing that would break. The sidebar carries every person in
the org scope on every view, so a name-based locator passes against an empty
team table — which is why TeamView.member_row finds people by their table
row instead. When you choose a locator, ask what a broken page would still
satisfy.
| File | Answers | Never contains |
|---|---|---|
pages/*.py | "where is it" — locators and navigation | assertions, test data, branching |
flows.py | multi-step actions composing page objects | assertions, test data |
test_*.py | the whole expect tree and every expectation | raw selectors |
A page object may own a URL shape — PersonView.path(uuid) composes the route
so nothing in a test hardcodes a URL. $person_id is the person's canonical
UUID since the identity cutover (#2098); it was the email before, and the SPA
encodeURIComponents it, so encoding belongs in the page object.
expect.set_options(timeout=15_000) in ui/conftest.py. Playwright's own 30s
action/navigation defaults are left alone; expect's 5s default is tight for a
cold SPA rendering after an OIDC round trip, and raising it is what lets
journeys use web-first assertions instead of sleeping or retrying. Never
wait_for_timeout.
Everything expected comes from the manifest at runtime or from the response the page just received — never typed from a prior run's observed output. A reshuffled seed must move the expectation with it.
reports = sorted(
p.display_name for p in stand_manifest.personas
if p.team == lead.team and p.role == "ic"
)
assert reports, "the manifest places nobody under this lead — the test would assert nothing"That guard is not decoration. A derived expectation that comes back empty makes every assertion below it vacuous.
No metric value. Hand-authoring an expected number is forbidden under
tests/stand/, and the seed publishes none: an admissible expectation would
have to be computable from the seed inputs rather than read back out of the
gold layer.
What a journey can assert about numbers is their honesty:
not_to_have_text("—")That is SCENARIOS.md §5 rule 1 — never a zero for missing data — and it is the strongest metric-adjacent claim available here.
Assert the whole set, not a sample. "Every report the roster declares", not "somebody rendered": a view that renders the first report and silently drops the rest — a pagination default, a truncated query, a broken key — passes a spot check and fails a person looking for their own team.
Assert identity, not mere presence. A link must be visible and point at that person; the SPA builds hrefs from a field it could pick wrongly, so existence alone would pass if every report's link resolved to the same view.
The place journeys silently rot. Audit every expectation against the sentence it claims to prove:
Run stand-test-auditor over a new journey before considering it done.
uv run --project tests --frozen playwright install chromium # first time only
# The verb appends your path to a hardcoded `tests/stand` and pytest unions path
# arguments — so a path here does NOT narrow the run. Use -k, or pytest directly.
./dev-compose.sh test-stand test -k login --headed
uv run --project tests --frozen pytest tests/stand/ui # a real subset
uv run --project tests --frozen pytest tests/stand/ui/test_login.py --headed
# --image is the exception: the args replace the image's CMD, so this one narrows.
# No suite image is published anymore (CI runs host-side from the checkout);
# the mode remains for a locally built one.
./dev-compose.sh test-stand test --image <locally-built-suite-image>In --image mode test paths are image-side (/tests/stand/ui), and
pytest-playwright's artefacts still land in ./test-results.
A ui-only run records nothing into the coverage ledger and writes no
operation catalogue. Recording happens in ApiClient.request, not at
construction: a journey's PersonaSession does carry a client, it just never
issues a request through it. Do not read a clean ledger after a ui run as
coverage.
drive-ui or
playwright-cli. A ticket describes intended, not current, behaviour.@pytest.mark.requires_seed(...) for every person named.expect tree in the test. Give the journey its quality-vector
marker — module pytestmark when the whole module shares a vector,
per-test markers throughout a mixed module, never both; the why lives
with the marker declarations in tests/pyproject.toml. Journeys proving
rendering and data are reliability, refused-access journeys security,
breadth-across-domains journeys versatility. Collection aborts on any
other vector count.cpt-…-feature-… ID
and the stable scenario number — and keep the pytest marker equal to the
scenario's vector; the full traceability contract is section 5 of the
quality-vector-tests skill.c31d302
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.