Testing
getbased uses two test layers:- Vitest wraps the node-side helper fixtures and fast logic checks.
- Playwright runs the browser suite against the live app in headless Chrome.
assert() helpers; Playwright executes those through tests/playwright/browser-script-runner.js so they run in the same browser suite as native Playwright specs.
Generated marker catalogs have dedicated drift gates:
:build command after editing authoring modules, then run the focused identity, placement, persistence, or terminology tests relevant to the diff. GitHub Actions owns the exhaustive Chromium and combined-coverage run.
The assert pattern
For CLI-provider and chat-workspace changes, cover the Companion verification checklist and chat ownership checks. Test discovery-only versus installation-authorized credentials separately, unavailable models/targets, canceled uploads, profile switches before the first token, and source-toggle enforcement in tools. Use mocks for routine CI rather than real health records or paid inference. Native OS service installation and OS file dragging require separate platform checks. Cold-load budgets measure source-mode requests, compressed transfer, and decoded bytes; they are distinct from production chunk budgets.cold-load-budget.spec.js attaches the per-resource inventory for attribution. Record any deliberate feature cost explicitly instead of removing assertions or silently widening the budget. Exhaustive browser and combined-coverage runs belong in GitHub Actions, not an unapproved local full-matrix run.
Browser-script fixture files define a local assert helper and collect results:
detail argument appears in the failure output — use it to print the actual value that caused the failure.
Notable test files
All test files live in thetests/ directory. Native browser specs live under tests/playwright/; fixture scripts live as tests/test-*.js. The table calls out representative ownership areas rather than duplicating the complete file listing from the repository.
A few tests run node-side (no browser) — pure-helper unit tests + node script guards. Marked node in the table; browser-driven coverage is owned by Playwright specs.
The landing page test (
test-landing.js) lives in the get-based-site repo.
Local AI and model-testing checks
Run the focused native suites while iterating:tests/test-provider-local-ai-runtime.js is a legacy browser fixture executed through the Playwright browser-script runner. While iterating, run its focused wrapper and any directly affected Settings, import, or service-worker specs. GitHub Actions owns the exhaustive browser matrix.
Full CI runner
- Checks if a server is running on port 8000; starts
node dev-server.jsif not - Runs the node-side tests first (fast fail on helper regressions, no browser needed)
- Runs the dev-server origin guard
- Runs the Playwright browser suite
- Exits with code
0if all pass,1if any fail
npm ci.
Coverage status
./run-tests.sh and the full Playwright command are intentionally guarded outside CI because they run the complete browser matrix and can produce substantial disk writes. Local development should use change-scoped Vitest files and Playwright specs. Run the guarded full matrix only after explicit approval to set GETBASED_ALLOW_HIGH_WRITE_TESTS=1; routine exhaustive verification belongs to GitHub Actions.
With CI authorization, COVERAGE=1 ./run-tests.sh runs Vitest with V8 coverage, runs the Playwright suite with Chromium JS coverage, then merges tests/.playwright-coverage/ with tests/.vitest-coverage/coverage-final.json. The reporter writes tests/.coverage.json and prints separate Playwright, Vitest/Node, and combined global function/byte coverage percentages for app-source JavaScript.
Running node scripts/playwright-coverage.mjs directly still falls back to the legacy high-surface Chromium sampler when no Playwright suite shards are present. The full COVERAGE=1 ./run-tests.sh path requires suite shards so coverage regressions in the Playwright instrumentation fail clearly.
Coverage is report-only by default. Set COVERAGE_MIN=90 or another percentage to fail the run when combined global function coverage falls below that floor.
Accessibility regression scan
tests/test-a11y-axe.js uses the pinned local axe-core 4.13 dependency and runs axe.run() against the live rendered surfaces. It does not fetch axe from a CDN. Severity policy:
- critical / serious → test fails on regression vs baseline
- moderate / minor → logged but doesn’t block (axe leans opinionated at those tiers)
Baseline-locked gate
The gate is baseline-relative, and the current critical/serious baseline is zero. Any critical or serious rule at any scanned stop is therefore a regression. Baseline lives attests/.a11y-baseline.json:
_axeVersion must match both the local dependency and PINNED_AXE_VERSION in the fixture. Update and review all three together when bumping axe because rule changes can appear as regressions.
Refreshing the baseline
After a wave of fixes that legitimately drops violation counts:tests/playwright/a11y-axe-browser.spec.js pipes the env var into the page context before loading the fixture. The test prints a ▶ {...} JSON line — copy it over the critical/serious/moderate/minor blocks in tests/.a11y-baseline.json. Don’t lower these numbers without an actual fix; the gate would lock in the new lower bound and silently accept regressions.
If the baseline file is missing entirely, the test treats the first run as “establish baseline” and prints the JSON to stdout for paste-back.
Running a single test in the browser
Open the browser console whilehttp://localhost:8000 is running, then:
Writing new tests
When you add a feature or fix a bug, add assertions to the relevant test file. If none fits, createtest-yourfeature.js.
What to cover:
-
Source inspection — verify the function or pattern exists in the source:
-
DOM state — check that elements render correctly:
-
Function behavior — call window-exported functions and check results:
-
CSS rules — verify styles are applied (use
getComputedStyleor inspect stylesheets): -
localStorage keys — verify storage conventions:
What the headless runner cannot test
- Drag-and-drop interactions
- File picker dialogs
- Actual streaming AI responses (API key required)
- IndexedDB state across page reloads (the runner resets between files)
handleBatchPDFs function exists and has the right signature) rather than the interaction itself.