Plan & progress
Build roadmap — to baseline runs and beyond
Prove whether the full Omni frontend harness improves visible UI/build outcomes versus a realistic raw-agent baseline.
Overall21/25 tasks · 84%
p0
Archive and reset the mission
done3/3- p0.1Create legacy archive index for older runsOld broad-gauntlet and route-map artifacts stay in place but are labeled historical/non-active.
- p0.2Update mission and prime questionsThe dashboard now asks whether full_ui_harness beats raw_ui_agent for Omni UI work.
- p0.3Update plan around UI Harness v0Roadmap narrowed to Omni frontend, rendered screenshots, visual judging, and Opus checkpoints.
p1
Goal and checkpoint spec
done4/4- p1.1Write UI Harness v0 goal artifactIncludes endpoint, scope, autonomy rules, forbidden actions, and Opus checkpoints.
- p1.2Define raw_ui_agent and full_ui_harness armsImplemented in the dedicated UI arm runner; raw keeps normal repo access, full uses the production sculptor/UI discipline plus mandatory gates.
- p1.3Pre-register one UI rubric and run-record schemaRun records now capture source gates, rendered gates, screenshots, code judge, visual judge, self-gates, cost, leakage, and calibrated/strict scores.
- p1.4Checkpoint 1: Opus review of spec and rubricOpus pushed for real full-harness config, load-bearing render proof, and low-confidence labeling; reconciled before the run.
p2
Dashboard and live run surface
done3/3- p2.1Mission page shows active goal, checkpoints, and prime questions
- p2.2Live runs page shows run-duration progress animationAnimated progress bar and checkpoint rail driven by /api/arena-live snapshot fields.
- p2.3Run Theater labels active Omni UI Harness v0 focus
p3
Render proof pipeline
done3/3- p3.1Start Omni dev server and run Playwright auth preflightEach arm now runs in a fresh Omni worktree sandbox with its own Next dev server and auth preflight.
- p3.2Capture one rendered screenshot and run rendered UI lintDefault desktop/mobile and invalid-submit desktop/mobile screenshots are captured for both arms.
- p3.3Checkpoint 2: Opus review of render proof pipelineOpus accepted the render path as useful but later flagged that duplicate-candidate proof is still missing.
p4
Raw vs harness pilot
active4/5- p4.1Run one-task raw_ui_agent vs full_ui_harness pilotClean comparable run: ui-patient-intake-20260627T035800Z.
- p4.2Checkpoint 3: Opus review of first pilot resultsInitial verdict: revise_harness_and_task_floor. Rescore later showed raw wins product-contract proof while full wins UI adherence.
- p4.3Replay/rescore saved patient-intake outputsTesting harness now replays duplicate, invalid-format, long-payer, and create-error states without rerunning agents. It also supports a golden reference calibration.
- p4.3aRun single-builder contract-discovery rerunRun ui-patient-intake-singlebuilder-contract-20260627T150851Z used the production sculptor arm with generic contract-discovery rails only. Full improved UI/process, but both arms missed pre-create duplicate detection.
- p4.4Run three-to-four-task pilot sliceCritical scoring and proof-gated duplicate compliance are in place. Still blocked until ungated contracts and mobile replay proof are trustworthy enough to avoid overfit.
p5
Tune and confirm
active4/7- p5.1Apply one or two reasonable harness improvements from pilot failuresApplied evaluator-side contract packet, generic builder contract-discovery rail, rescore browser states, golden reference calibration, critical-contract caps, and a mutation-boundary builder rail. The packet is not injected into candidate prompts.
- p5.1aFix critical-contract scoring before more seed spendCorrected rescore rescore-20260627T173700Z-mutation-boundary-fixed caps both raw and full at 50 for duplicate-before-create failure, while golden remains eligible at 65.
- p5.1bHarden the single builder against safety-before-mutation missesMutation-boundary rail was injected and run in ui-patient-intake-mutation-boundary-20260627T_go. It did not fix the miss: full still relied on reactive duplicate handling after POST.
- p5.1cExpose source-backed pre-mutation affordancesAdded generic fullHarnessProofs task hooks plus prove-ui-contract.mjs. Patient-intake now has an executable duplicate-before-create proof gate; full/sculptor passed that gate in ui-patient-intake-proofgate-20260627T_go.
- p5.1dSeparate gate compliance from spillover generalizationNext iteration should either add a held-out task proof or a source affordance for payer identity. Do not claim builder hardening from a gated duplicate pass alone.
- p5.2Run held-out UI task confirmation
- p5.3Checkpoint 5: Opus review of promote/revise/rerun decision
Updated 2026-06-27 · driven by ui/data/roadmap.json