GAUNTLET
Plan & progress

Build roadmap — to baseline runs and beyond

Prove whether the full Omni frontend harness improves visible UI/build outcomes versus a realistic raw-agent baseline.

Overall21/25 tasks · 84%
p0

Archive and reset the mission

done3/3
  • p0.1Create legacy archive index for older runs
    Old broad-gauntlet and route-map artifacts stay in place but are labeled historical/non-active.
  • p0.2Update mission and prime questions
    The dashboard now asks whether full_ui_harness beats raw_ui_agent for Omni UI work.
  • p0.3Update plan around UI Harness v0
    Roadmap narrowed to Omni frontend, rendered screenshots, visual judging, and Opus checkpoints.
p1

Goal and checkpoint spec

done4/4
  • p1.1Write UI Harness v0 goal artifact
    Includes endpoint, scope, autonomy rules, forbidden actions, and Opus checkpoints.
  • p1.2Define raw_ui_agent and full_ui_harness arms
    Implemented in the dedicated UI arm runner; raw keeps normal repo access, full uses the production sculptor/UI discipline plus mandatory gates.
  • p1.3Pre-register one UI rubric and run-record schema
    Run records now capture source gates, rendered gates, screenshots, code judge, visual judge, self-gates, cost, leakage, and calibrated/strict scores.
  • p1.4Checkpoint 1: Opus review of spec and rubric
    Opus pushed for real full-harness config, load-bearing render proof, and low-confidence labeling; reconciled before the run.
p2

Dashboard and live run surface

done3/3
  • p2.1Mission page shows active goal, checkpoints, and prime questions
  • p2.2Live runs page shows run-duration progress animation
    Animated progress bar and checkpoint rail driven by /api/arena-live snapshot fields.
  • p2.3Run Theater labels active Omni UI Harness v0 focus
p3

Render proof pipeline

done3/3
  • p3.1Start Omni dev server and run Playwright auth preflight
    Each arm now runs in a fresh Omni worktree sandbox with its own Next dev server and auth preflight.
  • p3.2Capture one rendered screenshot and run rendered UI lint
    Default desktop/mobile and invalid-submit desktop/mobile screenshots are captured for both arms.
  • p3.3Checkpoint 2: Opus review of render proof pipeline
    Opus accepted the render path as useful but later flagged that duplicate-candidate proof is still missing.
p4

Raw vs harness pilot

active4/5
  • p4.1Run one-task raw_ui_agent vs full_ui_harness pilot
    Clean comparable run: ui-patient-intake-20260627T035800Z.
  • p4.2Checkpoint 3: Opus review of first pilot results
    Initial verdict: revise_harness_and_task_floor. Rescore later showed raw wins product-contract proof while full wins UI adherence.
  • p4.3Replay/rescore saved patient-intake outputs
    Testing harness now replays duplicate, invalid-format, long-payer, and create-error states without rerunning agents. It also supports a golden reference calibration.
  • p4.3aRun single-builder contract-discovery rerun
    Run ui-patient-intake-singlebuilder-contract-20260627T150851Z used the production sculptor arm with generic contract-discovery rails only. Full improved UI/process, but both arms missed pre-create duplicate detection.
  • p4.4Run three-to-four-task pilot slice
    Critical scoring and proof-gated duplicate compliance are in place. Still blocked until ungated contracts and mobile replay proof are trustworthy enough to avoid overfit.
p5

Tune and confirm

active4/7
  • p5.1Apply one or two reasonable harness improvements from pilot failures
    Applied evaluator-side contract packet, generic builder contract-discovery rail, rescore browser states, golden reference calibration, critical-contract caps, and a mutation-boundary builder rail. The packet is not injected into candidate prompts.
  • p5.1aFix critical-contract scoring before more seed spend
    Corrected rescore rescore-20260627T173700Z-mutation-boundary-fixed caps both raw and full at 50 for duplicate-before-create failure, while golden remains eligible at 65.
  • p5.1bHarden the single builder against safety-before-mutation misses
    Mutation-boundary rail was injected and run in ui-patient-intake-mutation-boundary-20260627T_go. It did not fix the miss: full still relied on reactive duplicate handling after POST.
  • p5.1cExpose source-backed pre-mutation affordances
    Added generic fullHarnessProofs task hooks plus prove-ui-contract.mjs. Patient-intake now has an executable duplicate-before-create proof gate; full/sculptor passed that gate in ui-patient-intake-proofgate-20260627T_go.
  • p5.1dSeparate gate compliance from spillover generalization
    Next iteration should either add a held-out task proof or a source affordance for payer identity. Do not claim builder hardening from a gated duplicate pass alone.
  • p5.2Run held-out UI task confirmation
  • p5.3Checkpoint 5: Opus review of promote/revise/rerun decision

Updated 2026-06-27 · driven by ui/data/roadmap.json