OUI-1 proves generative UI lives or dies on format, not model size
Yesterday OpenUI shipped the thing that makes generative UI feel real on a solo budget. OUI-1 runs on a single RTX 5090 and beats models seven times its active size.
I've argued generative UI is a design problem, not a model problem — that your catalog and constraints decide what ships. OUI-1 is the proof. It didn't win by adding parameters. It won by fixing the language the model speaks.
Generative UI ships when the format is cheap, streamable, and verifiable — OUI-1 hit 71.7% valid at 4B active by coupling OpenUI Lang with a parser as reward and self-distillation across 27 component libraries.
Why does generative UI break on JSON?
JSON charges you for syntax your users never see.
Every {"op":"add","path":"/elements/id" is tokens you pay for at output prices and seconds your user stares at a skeleton. Declarative formats all constrain the model to your catalog — the difference is wire cost.
The OpenUI Lang benchmarks comparing format token costs and render latency nail it: 4,800 tokens for OpenUI Lang versus 10,180 for json-render across seven screens. The contact form is 294 versus 893 — 4.9 seconds versus 14.9 at 60 tokens per second.
You still need validation either way. JSON just makes you pay twice.
How did OUI-1 hit 71.7% at 4B active?
It made the parser the teacher.
The OpenUI team's breakdown of OUI-1 as the first model built for generative UI started with speed: DiffusionGemma writes 256-token blocks at once and hits 1,000 tokens per second on an H100. Ideal for streaming interfaces on consumer hardware.
But DiffusionGemma scored 13.0% on the benchmark. Fast and wrong doesn't ship.
Fine-tuning on 700 OpenUI Lang examples lifted one library to 28.8%, yet slowed generation from 1.6s to 4.3s — the model wrote real names and values, 32 tokens per statement not 22, and needed twice the denoising steps. Better outputs cost time.
The fix was rejection-sampled self-training. Generate a few hundred programs, keep the ones the parser accepts, repair near-misses by the exact error — median repair was one statement — and train for 500 steps on one A100. Then generate the next batch from the new model.
Speed fell to 1.9s with 28% more tokens, and both error types dropped — schema 292 to 76, wiring 971 to 484. Earlier runs traded one for the other.
Repeating that across 27 libraries gave OUI-1: 71.7% at 4B active. Only the dense 27B Qwen3.8 beats it at 78.8%. A flow that once needed Gemma 4 on Cerebras now runs on an RTX 5090.
What does format efficiency buy you?
A render users feel and a cost you can keep.
At output-token prices, a pricing page at 2,487 tokens as json-render costs more than twice the 1,166 as OpenUI Lang, every turn. OpenUI Lang streams line by line and can render skeletons for IDs that haven't arrived yet.
Dashboards still save 45–52% — where generative UI matters most. Fewer brackets means fewer chances to emit invalid JSON or invented props. The OpenUI open-source framework on GitHub shows 96.5% render success for OpenUI versus 80–95% for json-render. You scale generative UI by giving the model less syntax to get wrong.
How should solo founders use it this week?
Build the catalog habit, not a bigger prompt.
Start declarative and boring:
- Define components with Zod schemas, not paragraphs. A
PriceCardwithprice: numberbeats a taste essay. - Generate the system prompt from the library. Every component you add appears there automatically.
- Stream OpenUI Lang and render line by line — don't wait for a blob.
- Add one
GeneratedViewescape hatch: a sandboxed iframe that takes model-generated HTML as a string. Use it only when the catalog truly has no answer.
Check two gates each time: does it parse and wire, and does it match the ask? That's the pair OUI-1's training judge used.
You don't need 27 libraries. Pick the five components that cover 80% of your replies — table, chart, card, form, stats — and make those tight before you add a sixth. Fix the language first. The model gets to be small after that.
Frequently asked questions
OUI-1 is an open-weight 26B-A4B model finetuned from DiffusionGemma to generate OpenUI Lang. It scores 71.7% on the Generative UI Benchmark versus 13.0% for its base, beating every open model up to 31B active except the dense Qwen3.8 27B, and it runs on consumer hardware like an RTX 5090. For solo founders, it shows generative UI can be local and fast without giving up reliability.
About the author
mosh
mosh is a product designer for growth, working with design thinking and ever-improving design systems. What matters: fixing conversion, whether in B2B dashboards or direct-consumer apps.
Keep reading
- SaaS activity tracking: designing an admin log teams trust
Admins needed to see every team action in one place. Research with two managers and eight user stories turned a confusing log into filters, search, share, and export.
- GymProLuxe: turning a resistance kit into a training system
GymProLuxe sold trusted hardware but motivation faded after delivery. Premium app UX gave every owner a daily plan matched to their kit.
- ScoreAi: match search redesign that lifted engagement
ScoreAi had demand but its match search leaked engagement. A conversion-first rebuild of hierarchy, scanning, and onboarding fixed discovery.