GAUNTLET
Mission

Omni UI Harness v0

Active goal

Prove whether the full Omni UI harness produces better frontend work than a realistic raw UI-agent baseline.

A reproducible raw_ui_agent vs full_ui_harness pilot report on frozen Omni UI tasks, with run records, lint and review outputs, rendered screenshots, visual judging, animated live progress, and a clear promote/revise/rerun verdict.

Omni frontend/UI only
No DB, Noah, scheduling, comms, deployment, production write, or full operating-harness claim
Use existing Omni UI/build tasks first, then reserve held-out tasks before tuning
North star

Make Omni frontend agents dependable enough that Chris can focus on product ideas while the harness handles UI standards, visible proof, and review pressure.

The premise

Raw agents are the baseline. The thing under test is the full UI harness as an operating system: style guide, component primitives, static lint, rendered lint, reviewed workflow, screenshots, visual judging, and feedback logs.

Latest result
ui-reporting-repairmode-k1-20260628T012257Z · reporting-dashboard
revise_full_harness

Build-first repair-mode improved process and reduced severe blocking, but the corrected score is raw 50 vs full 60 because full still has mobile KPI/table overlap. This is a revise signal, not a promotion.

Opus checkpoints

c1
Spec and rubric checkpoint
Opus review

UI Harness v0 arms, frozen task split, single rubric, run-record schema, and screenshot judging plan are reviewed by Opus and reconciled before implementation.

c2
Render pipeline checkpoint
Opus review

Omni dev server, Playwright auth preflight, screenshot capture, rendered UI lint, and dashboard evidence display are proven on one screen.

c3
One-task pilot checkpoint
Opus review

One raw_ui_agent vs full_ui_harness comparison produced complete records, screenshots, visual judge output, Opus review, and the revise_harness_and_task_floor diagnosis.

c4
Pilot slice checkpoint
Opus review

Three to four UI tasks run with side-by-side evidence and a report showing whether the harness is improving the work.

c5
Held-out confirmation checkpoint
Opus review

A reserved UI task confirms or rejects the tuned harness lift before promotion.

The prime questions

01

Does full_ui_harness beat raw_ui_agent on Omni UI/build tasks?

revise

This is the headline claim. The raw arm can still read the repo and may discover the guide; the harness arm makes the UI rails mandatory and operational.

Finding: Reporting repair-mode did not prove promotion-grade lift. Corrected effective scores are raw 50 vs full 60, but full still has a mobile overlap severe enough to revise the harness.
02

Can the lab see what Chris sees?

partial

Rendered screenshots and visual judging are required because code review cannot reliably catch clipping, page scroll, density collapse, or awkward first-viewport states.

Finding: Yes for this failure class: rendered mobile screenshots exposed KPI/table overlap that the old score artifact would have over-rewarded as 76. The dashboard now treats overlap, clipping, and viewport overflow as severe geometry.
03

What does the UI harness catch that raw agents leave for Chris?

active

Every static lint, rendered lint, Opus review, screenshot judge, and remaining Chris-visible issue should become a feedback signal for hardening the harness.

Finding: The build-first full harness cleared raw's static/source-style failures, ran five self-gates, and reduced severe blocking from 20 to 5. It still did not repair the mobile overlap, so the useful signal is process lift without shippable UI lift.
04

Does the harness preserve useful agency?

partial

The harness should guide away from repeatable UI mistakes without making agents timid, generic, or unable to solve novel frontend problems.

Finding: Build-first repair appears to preserve more useful agency than the earlier read-first contract-mapping rail, but the repair loop still needs stricter rendered feedback so agents do not stop at a polished-but-overlapping result.
05

After tuning on the pilot slice, does the lift hold on reserved UI tasks?

open

We need held-out confirmation so the harness does not merely overfit the first few screens.

Archive policy

Active family: omni-ui-harness-v0. Legacy index: runs/v2/legacy/index.json. Old broad-gauntlet, route-map, and raw-model findings are historical evidence unless explicitly promoted into the Omni UI Harness v0 report.

Decisions made

Start with the Omni UI specialty harness.
Frontend work is the highest immediate leverage, and existing Omni UI rails make it the fastest track to reputable runs.
Compare raw_ui_agent against the full UI harness, not lint-on/off ablations.
The product question is whether the UI harness as a system beats a realistic baseline. Per-rail diagnostics are captured as byproducts.
The raw arm keeps normal repo access.
A realistic raw agent may find the guide or components. The tested lift is whether making the rails mandatory and operational improves outcomes.
Rendered screenshots are first-class evidence.
They let Chris inspect the result and let the visual judge evaluate layout, density, clipping, and real rendered behavior.
Archive old runs by manifest, not file movement.
Historical run artifacts remain immutable and inspectable, while the dashboard and active docs stop treating them as the current mission.
Non-blocking UI flags should not hard-cliff effective score.
The first comparable run showed both arms had 0 blocking issues but were floored to 50 by cosmetic/non-blocking flags, hiding the useful quality signal.
Patient-intake needs explicit proof states before more seed spend.
Default and invalid-submit screenshots did not render duplicate candidates or populated invalid-format validation, so visual judging could not see the core workflow.
Rescore saved outputs before rerunning agents.
Replay showed the old tie was a split result: full won UI adherence, but raw won source contract and browser proof because full caught duplicates only after create.
Keep contract packets out of candidate prompts.
Opus flagged that injecting scorer anchors into the builder prompt would become an answer key. The candidate gets a generic contract-discovery rail; the testing harness gets the packet.
Do not trust blended score without critical-contract tiering.
The earlier golden rank inversion showed that a polish-heavy blended score can hide contract failure. The corrected scorer now makes critical contract status dominate the verdict.
Critical proactive contract failures dominate promotion verdicts.
The corrected rescore caps raw and full at 50 when duplicate-before-create fails, so UI polish cannot outrank a contract-passing reference.
Treat mutation-boundary prompt rails as insufficient by themselves.
The full arm read the duplicate API route and still chose reactive 409 handling. The next builder improvement should make source affordances or builder-side self-tests expose required pre-mutation requests, not add more generic prose.
Executable proof gates prove enforced compliance, not general builder learning.
Opus flagged the duplicate proof as an answer-key-like gate for this task. The honest signal is whether ungated contracts also improve; in the proof-gate run, payer identity did not.
Use build-first repair mode as the next builder-hardening direction.
The read-first rail reduced useful agency and still missed contracts. Reporting repair-mode showed better process discipline while keeping the builder focused on making and repairing the UI.
Rendered overlap, clipping, and viewport overflow are severe geometry defects.
The reporting run's mobile overlap made the original 76 misleading. Corrected scoring caps severe rendered geometry defects before polish can imply shippable UI.

Updated 2026-06-28 · edit ui/data/mission.json to evolve this