Your prototype works once. Determinism ships it
Your AI demo works in your recording. It fails in your user's hands.
That is not a model problem. The model did what it always does — gave you a plausible answer. The miss is everything around it that you did not define. I have written that the vibe-code gap is real — the distance between a prototype that works once and a product that works when you are not watching is not smaller in 2026. It is bigger, because we all ship demos faster. Determinism is how you close it.
Why does your AI demo work once and fail at scale?
Because you were the system around the demo.
You picked the good input. You ignored the slow tool call. You retried by hand when the model hedged. A user will not do any of that. They will paste a PDF you have never seen, on a connection you did not test, at 11pm when you are asleep.
This is the same shift enterprise teams are naming as AI moving from co-pilot to operating system. McKinsey's 2025 State of AI report found 88% of organizations now use AI in at least one function — the pattern is no longer a helper that suggests. It is infrastructure that must behave the same way twice. Solo products hit that bar earlier, not later, because one person cannot babysit every run.
The prototype hides that work. Determinism names it.
What is deterministic engineering for solo founders?
Determinism is not removing the model. It is putting a contract around it.
Think of it as four promises you keep for every AI step that matters:
A contract for what goes in and what must come out. Not a prompt, a shape — required fields, allowed values, and what happens when a field is missing. If your retrieval returns nothing, the contract says what the model may and may not claim.
A policy for what happens when a tool fails. One retry, then a fallback, then a human-readable error. No silent empty state that looks like success.
A gate that checks output before a user sees it. That gate is an eval — a list of 20 real inputs with the outputs you already decided are good. Anthropic's guide to building evals for agents frames it well for probabilistic systems: you score outputs against what you decided was good, you do not grade the prompt.
A rollback that lets you undo today's change in one click. Models drift, prompts regress, APIs change. Without a rollback you ship forward because you cannot ship back.
None of this is new engineering. It is the boring engineering that vibe coding lets you skip for a week — and then charges interest on.
How do you make an AI product deterministic?
You do not rewrite the demo. You wrap it.
1. Pick one path and name its contract.
Take your most-used flow — not every flow. Write down for each step: what input shape you expect, what output shape you require, and what the fallback does if that shape is not met. Store that contract where you will see it. I keep mine next to my context files, for the same reason I wrote that context files are the fix — if it is not written, the agent will guess and you will pay for the guess at 2am.
The Model Context Protocol specification is a useful mental model here even if you do not adopt it directly: tools expose what they can do in a portable way, and the product decides how to call them. Your job is the decision layer.
2. Add one eval gate and block on it.
Make a folder with 20 inputs drawn from real users — the messy ones, not the demo ones. Save the outputs you would be happy to show. On every prompt or model change, run the same 20 and compare. If the score drops, the change does not ship. That is the whole gate.
You do not need a platform for this. A script that runs before deploy is enough. What matters is that it runs every time, not that it looks impressive.
3. Give every external call one retry and one fallback.
APIs fail. Models time out. Retrieval returns empty. Decide now: one retry with backoff, then what? For my products that means: if retrieval is empty, the agent says it does not have that answer and offers the next step — it does not hallucinate a citation. If a tool call fails, the user sees a retry button and I see the trace. No invisible failures.
4. Keep a one-click rollback.
This is cheaper than getting the prompt perfect. Tag every generation config — prompt, model ID, tool versions — with a deploy. If the eval drops after a change, revert the tag, not the code. Determinism is not about being right the first time. It is about being able to go back in ten seconds.
When should you stay loose and when should you lock it down?
Stay loose while you are still searching.
If you do not know who the product is for, what job it compresses, or whether anyone will pay for it, do not engineer for the hundredth run. Vibe hard, talk to users, throw builds away. Determinism too early turns a search problem into a maintenance problem.
Lock it down the moment a path carries trust.
Ask one question: if this path fails silently, will it create support work, bad data, or a user who does not come back? If yes, that path gets a contract and a gate before the next feature ships. If no, leave it loose and keep moving.
That is the same judgment I keep coming back to. AI handles the plumbing — the generation, the drafting, the first pass. You handle the calls where a wrong answer has a cost. Determinism is just that judgment written into the system so it holds even when you are not in the room.
Your prototype proved you could build it. Determinism proves you can keep it built.
Frequently asked questions
Deterministic engineering is the layer you add after a vibe-coded demo so the same input reliably produces a good output. It includes explicit contracts between tools, retry and fallback policies, evals that run on every prompt change, and a rollback path when the model drifts. It is not more prompts — it is the system around the prompts.
About the author
mosh
mosh is a product designer for growth, working with design thinking and ever-improving design systems. What matters: fixing conversion, whether in B2B dashboards or direct-consumer apps.
Keep reading
- SaaS activity tracking: designing an admin log teams trust
Admins needed to see every team action in one place. Research with two managers and eight user stories turned a confusing log into filters, search, share, and export.
- GymProLuxe: turning a resistance kit into a training system
GymProLuxe sold trusted hardware but motivation faded after delivery. Premium app UX gave every owner a daily plan matched to their kit.
- ScoreAi: match search redesign that lifted engagement
ScoreAi had demand but its match search leaked engagement. A conversion-first rebuild of hierarchy, scanning, and onboarding fixed discovery.