12 KiB
Human-in-the-Loop (HITL) test inventory
Part of #14381 (HITL inventory / developer-facing runner), split from the device-review rollup #14317 / #14395. Companion: #14382 (resettable onboarding). Plugs into the existing evidence loop (
scripts/e2e-recordings/*→ contact sheets → viewer) and nubs'launch-qaboard — this does not replace the device passes, it tells a developer which things a human/device still has to eyeball and pre-stages the frames so the review is a contact-sheet skim, not a manual drive.
What "HITL" means here
A flow is HITL when an automated assertion cannot decide pass/fail — the signal lives in a human's eye, thumb, ear, or a real device capability the headless runner can't fake:
- Visual polish — clipping, overflow, contrast, spacing, safe-area, leaked error strings. A DOM node can exist and still look broken.
- Gesture / motion feel — scroll momentum, sheet detents, drag, long-press, frame glitches. "It renders" ≠ "it feels right."
- Keyboard behavior — soft-keyboard push/resize, focus retention, IME, input obscuring on real mobile.
- Onboarding — first impression, copy, pacing; historically untestable because you can't reset a real, memory-laden agent (→ #14382).
- Login / auth — real OAuth round-trip, tenant, token, redirect. Injected state proves nothing (FLEET.md ban; qa-agent's staging-401 lesson).
- Wallet — real signing, balances, "not found" / error copy (#14426).
- Notifications — permission prompt, delivery, tap-through.
- Camera / mic — real permission grant + capture feel.
Coverage legend
| Mark | Meaning |
|---|---|
| ✅ auto | Fully machine-decidable; e2e asserts it, no human needed |
| 🟡 frame | Headless can stage + screenshot the state; a human judges the frame |
| 🔴 device | Needs a real device capability (soft keyboard, wallet, camera, push) |
| ⛔ none | No current coverage; manual-only today |
Inventory
Onboarding
| Flow / decision point | Why HITL | Current coverage | What a harness can pre-stage |
|---|---|---|---|
| First-run welcome + name/style pick | first impression, copy, pacing | 🟡 frame — packages/ui first-run e2e + test:ftu-home-e2e; capture-android-emu / capture-ios-sim drive the real Capacitor onboarding |
contact-sheet of each step (welcome → name → style → provider → complete) via the replay entry (#14382) so a fully-onboarded dev can re-walk it without a wipe |
| Provider / model selection | catalog correctness + visual density | 🟡 frame — first-run options e2e | screenshot of the provider grid at each viewport |
| Boot-trouble-speaks-in-chat (no banners) | tone / does the copy read right | 🟡 frame — App.chat-overlay-first-run.test.tsx, #14168 detent work |
staged error card frame |
| Onboarding persists across restart | correctness (not visual) | ✅ auto — first-run-persistence.restart.test.ts (real saveElizaConfig/loadElizaConfig) |
n/a (already machine-decided) |
| Onboarding re-runnable on a real agent | can't reset without nuking memories | 🟡 frame — dev-gated ?onboarding-replay=1, wired on the app boot path (packages/app/src/first-run-boot-patches.ts; regression test packages/app/test/first-run-boot-patches.test.ts); steps below |
dev-gated ?onboarding-replay=1 client overlay (no server wipe) |
Onboarding replay / reset steps (#14382)
Two dev query params re-run onboarding. Both are wired on the main-window boot
path in packages/app/src/first-run-boot-patches.ts (called from
packages/app/src/main.tsx; the arm-before-patch order is load-bearing and
locked by packages/app/test/first-run-boot-patches.test.ts).
Non-destructive replay — the default for QA on a real, memory-laden agent:
- Boot the dev app (
bun run devat the repo root, orbun run --cwd packages/app dev:sharedon a parallel lane). - Append
?onboarding-replay=1to the URL and reload. Onboarding renders again: a client overlay (packages/ui/src/platform/onboarding-replay.ts) makes the client report fresh while the real agent, its config, the persistedelizaos:active-server, and all server state (PGlite data dir, conversations, knowledge, trajectories) stay untouched. - Walk the steps for QA/screenshots. Submitting the replayed onboarding re-applies the chosen onboarding config to the same agent (that is the full-path exercise); it never deletes data.
- To leave without applying anything: drop the param and reload — the agent is exactly as it was.
Headless proof/capture driver (screenshots + non-destruction assertions
against a running dev stack):
node packages/app/scripts/onboarding-replay-evidence.mjs --out <dir>.
?reset — genuinely fresh client session (stateful, still no server wipe):
clears the client's persisted session (active server, setup step,
first-run-complete) and sets the durable force-fresh flag so the next boot
lands on onboarding as a new client. Use it for "new device" QA; unlike the
replay it forgets which server/agent the client was attached to.
Safeguards (so the test path cannot be confused with a destructive
production account reset): ?onboarding-replay=1 is compiled inert outside
dev builds (import.meta.env.DEV); neither param can reach
POST /api/agent/reset — the only path that wipes agent memories — and the
module tests assert no delete/reset/clear method ever fires during a replay.
First chat
| Flow / decision point | Why HITL | Current coverage | What a harness can pre-stage |
|---|---|---|---|
| First message send → streamed reply | live inference + feel | 🟡 frame — app test:e2e chat suite (stubbed); real inference needs credits (see #14424 credit path) |
staged send + streamed-token frames; a human confirms cadence |
| Suggestions / FTU home widgets | visual + relevance | 🟡 frame — test:suggestions-e2e, test:ftu-home-e2e |
rest + populated frames |
| Chat scroll / infinite scroll / momentum | gesture feel | 🟡 frame — test:chat-scroll-web-e2e, test:chat-infinite-scroll-e2e, test:chat-perf-gate; scroll cert superset #14380 |
scroll-position + perf-gate frames |
| Chat sheet detents / frame glitch | motion feel | 🟡 frame — test:chat-sheet-e2e, test:chat-sheet-frame-glitch-e2e, test:chatux-gesture-e2e |
detent-state frames |
Login / auth
| Flow / decision point | Why HITL | Current coverage | What a harness can pre-stage |
|---|---|---|---|
| Cloud sign-in (OAuth round-trip) | real token/tenant/redirect | 🟡 frame — cloud-e2e mock login; real sign-in is device-only (#13609, #13611, #13610) |
mock-login contact sheet; real login stays 🔴 device |
| Stale-token / no-credits resume (spam guard) | correctness + no visual spam | ✅ auto — use-first-run-conductor.test.ts (#14387 / PR #14423) |
n/a |
| Staging tenant correctness | env-bake correctness | ✅ auto (process) — FLEET.md rule + qa-agent SW fix #14409 | n/a |
Wallet
| Flow / decision point | Why HITL | Current coverage | What a harness can pre-stage |
|---|---|---|---|
| Wallet view render | leaked "Not found" red string (#14426, P1) | 🟡 frame | staged wallet-empty + wallet-populated frames so the leak is obvious on the sheet |
| Real signing / balances | real key material | 🔴 device — platform/e2e-wallet.ts stub only |
n/a headless |
Notifications / camera / mic
| Flow / decision point | Why HITL | Current coverage | What a harness can pre-stage |
|---|---|---|---|
| Permission priming prompt | copy + timing | 🟡 frame — test:permission-priming-e2e |
priming-prompt frame |
| Mic permission (first-run) | real grant + capture | 🔴 device — use-microphone-permission.ts |
n/a headless |
| Push delivery + tap-through | real device push | 🔴 device — Seeker pass (qa-agent) | n/a headless |
Visual polish (cross-cutting launch-qa)
| Issue | Why HITL | Coverage | Pre-stage |
|---|---|---|---|
| #14427 launcher 'Relationships' label clips | pixel clipping | 🟡 frame | launcher tile frame at each width |
| #14426 wallet red 'Not found' leak (P1) | error-string leak | 🟡 frame | wallet frame |
| #14425 'when this Mac is asleep' on non-Mac | platform copy | 🟡 frame | CloudOverview frame on non-mac env |
| #14380 scroll + tap-target cert | gesture + hit-area | 🟡 frame — superset harness (sibling sol lane) | per-widget scroll/tap frames |
| #14379 mobile chat/search keyboard states | soft-keyboard | 🔴 device | n/a headless — real-device cert |
LifeOps live validation (#11632) — lifeops-live group
The live, account/device-backed LifeOps operator lane. Deliberately outside
the golden-path default (HITL_GOLDEN_GROUPS) because it needs real
credentials and devices; select it explicitly:
node scripts/hitl/run-hitl.mjs --groups=lifeops-live. Runbook:
plugins/plugin-personal-assistant/docs/LIFEOPS_LIVE_VALIDATION.md.
| Flow / decision point | Why HITL | Current coverage | What a harness can pre-stage |
|---|---|---|---|
| Connector credential intake / readiness | a human confirms which connector creds are present (and fresh) before live lanes run | 🟡 frame — bun run lifeops:hitl (scripts/lifeops/hitl-credential-dashboard.mjs v2: per-auth-path rows from connector-paths.mjs, layered env sources, one-click gh/SIWE/signal-link, probes recorded to docs/testing/hitl-ledger.json) |
readiness dashboard frame (values masked to last-4; freshness green ≤7d / yellow >7d / red >30d-or-never) |
| 9-state OWNER/AGENT permission matrix (credentialed) | correctness (not visual) | ✅ auto — owner-agent-permission-matrix.integration.test.ts under LIFEOPS_PERMISSION_MATRIX=1 |
n/a (machine-decided) |
| Live connector suites | correctness; creds present → live, absent → clean skip | ✅ auto — node scripts/lifeops/run-11632-live-lanes.mjs (each suite describeIf-gates on its creds) |
n/a (machine-decided) |
| LifeOps split views populated with live data | populated/empty/error render quality with real data | 🟡 frame — /lifeops-live-test view + bun run --cwd packages/app audit:app + bun run test:e2e:record |
populated-view frames per split view, desktop + mobile |
| iOS/macOS native flows (HealthKit, Family Controls, SelfControl) | real OS permission dialogs + native capabilities | 🔴 device | n/a headless |
| Android native flows (Health Connect, SMS default-role, Usage Access) | real OS permission dialogs + SIM-backed capabilities | 🔴 device | n/a headless |
How the runner surfaces this (design)
The runner does not invent a new capture stack. It:
- Selects the HITL subset of
UI_E2E_SUITESrelevant to a decision point (onboarding / first-chat / login) via a tag manifest (scripts/hitl/hitl-manifest.mjs). - Runs them through the existing
e2e-recordingspipeline (record → extract frames →generate-contact-sheets→generate-viewer). - Emits a HITL contact sheet where each staged frame is labelled with its
decision point + a
pass / fail / blockedslot, so a developer reviews a sheet during development instead of driving the app by hand or waiting for a post-merge device pass.
Pass/fail/blocked is recorded by the human in the viewer; 🔴 device rows are
marked blocked (device) automatically and routed to the Seeker pass
(qa-agent) rather than pretending headless covered them. Auth/wallet/real
inference use real flows only (FLEET.md) — the harness stages the mockable
frames and explicitly defers the rest to device, never fabricates state as
evidence.
Non-goals / honesty
- The
packages/uie2e env is known-flaky; this inventory marks what can run headless (🟡) vs what genuinely needs a device (🔴). It does not claim headless coverage where only a device suffices. - This is the inventory + the onboarding-replay slice + a harness skeleton. Wiring every 🟡 row into the tagged sheet and the full launch-qa board labelling is follow-up (tracked on #14381).
— [sol-orch]