Omni UI Harness v0
Prove whether the full Omni UI harness produces better frontend work than a realistic raw UI-agent baseline.
A reproducible raw_ui_agent vs full_ui_harness pilot report on frozen Omni UI tasks, with run records, lint and review outputs, rendered screenshots, visual judging, animated live progress, and a clear promote/revise/rerun verdict.
Make Omni frontend agents dependable enough that Chris can focus on product ideas while the harness handles UI standards, visible proof, and review pressure.
Raw agents are the baseline. The thing under test is the full UI harness as an operating system: style guide, component primitives, static lint, rendered lint, reviewed workflow, screenshots, visual judging, and feedback logs.
Build-first repair-mode improved process and reduced severe blocking, but the corrected score is raw 50 vs full 60 because full still has mobile KPI/table overlap. This is a revise signal, not a promotion.
Opus checkpoints
UI Harness v0 arms, frozen task split, single rubric, run-record schema, and screenshot judging plan are reviewed by Opus and reconciled before implementation.
Omni dev server, Playwright auth preflight, screenshot capture, rendered UI lint, and dashboard evidence display are proven on one screen.
One raw_ui_agent vs full_ui_harness comparison produced complete records, screenshots, visual judge output, Opus review, and the revise_harness_and_task_floor diagnosis.
Three to four UI tasks run with side-by-side evidence and a report showing whether the harness is improving the work.
A reserved UI task confirms or rejects the tuned harness lift before promotion.
The prime questions
Does full_ui_harness beat raw_ui_agent on Omni UI/build tasks?
reviseThis is the headline claim. The raw arm can still read the repo and may discover the guide; the harness arm makes the UI rails mandatory and operational.
Can the lab see what Chris sees?
partialRendered screenshots and visual judging are required because code review cannot reliably catch clipping, page scroll, density collapse, or awkward first-viewport states.
What does the UI harness catch that raw agents leave for Chris?
activeEvery static lint, rendered lint, Opus review, screenshot judge, and remaining Chris-visible issue should become a feedback signal for hardening the harness.
Does the harness preserve useful agency?
partialThe harness should guide away from repeatable UI mistakes without making agents timid, generic, or unable to solve novel frontend problems.
After tuning on the pilot slice, does the lift hold on reserved UI tasks?
openWe need held-out confirmation so the harness does not merely overfit the first few screens.
Active family: omni-ui-harness-v0. Legacy index: runs/v2/legacy/index.json. Old broad-gauntlet, route-map, and raw-model findings are historical evidence unless explicitly promoted into the Omni UI Harness v0 report.
Decisions made
Updated 2026-06-28 · edit ui/data/mission.json to evolve this