Your Agent Has an Exactly-Once Problem
On a stroke-rehab platform I worked on, an LLM scored each rep into the clinical record during live sessions. The change that mattered was a one-line conflict key, because retries were overwriting other patients' results.
Last summer I worked on the backend of a stroke-rehabilitation research platform with clinicians at UCSF Health and Stanford Medicine. Patients do their exercises during live telehealth sessions, wearing a mixed-reality headset while a clinician watches over the call. The headset streams camera frames to the backend, which stitches them into a video with ffmpeg, ships it to cloud storage, and asks a multimodal model one question: did this rep succeed? The answer comes back as a boolean and a paragraph of reasoning, and it lands in Postgres as part of the patient’s progress record while the session is still running.
The most recent change in that codebase is not a model upgrade or a clever prompt. It is a fix to a database conflict key, because exercise results could collide and overwrite each other. One line, roughly. For this system, that line taught me more about agent infrastructure than most of the model-architecture posts I read that year.
The model is the small part
The AI module in that service is a few hundred lines of Go. The actual Gemini call is maybe thirty of them. The rest is temp directories, frame-to-video conversion, upload retries, cleanup, and making sure a half-processed video never gets evaluated as a whole one.
People call that ratio harness engineering now. I am not claiming that name settled anything. In this service, the Gemini call was never the system. The model answers a question. Everything around it decides whether the question was asked correctly and whether the answer can be trusted to land.
Exactly-once, again
Here is the fix I mean. Results were written with an upsert keyed on the exercise and the rep number. Sounds reasonable. Except two patients can do the same exercise, and one patient can hit the same rep at a different step, and a headset on a flaky connection can retry a submission it thinks failed. Under the old key, some of those writes silently overwrote each other, in the middle of a live session, with a clinician looking at the numbers. Under the new one, the key is the exercise, the patient, the rep, and the step together, and a retry updates the same row instead of corrupting a different one.
INSERT INTO exercise_result (...)
VALUES (...)
ON CONFLICT (exercise_id, patient_id, rep_number, step)
DO UPDATE SET ...
This is the oldest lesson in distributed systems, wearing new clothes.
Exactly-once delivery does not exist. At-least-once plus a storage layer that refuses to be fooled by duplicates does.
I keep seeing the same shape in agent systems I have been around since then. The agent writes to a system, files a ticket, moves money, and nobody approves each individual action. An agent action is an API call with side effects, and the harness retries, because retrying is what good infrastructure does. When the timeout fires after the tool succeeded, the action runs twice. The postmortem says the agent “hallucinated a second refund.” The prompt was fine both times. The tool layer just never learned what a composite key is.
For this clinical write path, the fix was to derive the identity of an action from everything that makes it distinct. The patient, the exercise, the rep, the step. Let storage reject the duplicate. Ten lines, and it is the difference between “the headset retried” and “the chart now belongs to someone else.” I would start from the same place on an agent tool call: retries are duplicates you invited.
When the model writes programs, the compiler is the guardrail
My favorite part of the platform is how it generates new exercises. The model does not describe a scene in prose. It synthesizes a complete Scenic program, a scenario language that compiles into the headset’s mixed-reality scene, from a JSON description of the therapy goal plus a library of allowed objects and actions.
Prose can lie fluently. A program either compiles against the object library or it does not. Half the guardrail conversation this year has landed on the same idea: stop filtering what the model says and start validating what it did. Generating artifacts that a compiler, a schema, or a reconciler can reject is the strongest version of that. The clinical world got there early because a vague answer about a patient is not an acceptable artifact.
Booleans you can count
The evaluation result is a bool plus reasoning, and that split is doing real work. The reasoning is for the clinician reviewing an edge case. The bool is for the dashboard. You can chart pass rates per exercise and per model version, and when a model update shifts the curve, you see it in production data, not in a benchmark.
Medicine has run on this pattern forever. Chart the vitals, investigate the anomaly, never trust a system you are not measuring. The eval suite I run on deploys now is the same idea wearing software clothes.
The part I keep coming back to
Clinical software forced these habits on us early, because the cost of sloppiness is a patient’s chart. Agent infrastructure is arriving at every one of them, one production incident at a time.
The models are the new part. The failure modes are older than I am, and so are the fixes.
Amisha
Filed under Notes · AI Systems