All writing
NotesApr 1, 20264 min read

Designing Retry-Aware Infrastructure

On a clinical streaming pipeline I worked on, a retry was a duplicate we had invited. The timeout fired after the write already happened, and the second pass was only safe because the row key refused to create a new record.

Retries are the first thing you add and the last thing you understand. A call fails, you wrap it in a loop, you move on. It looks free.

Then it takes down the thing it was supposed to protect. I learned this the long way on a clinical streaming pipeline, where a consumer crash meant every nearby message ran at least twice.

The thundering herd

Every client retrying on the same fixed interval is a synchronized flood. The service hiccups, a thousand callers all wait two seconds, and a thousand requests land in the same millisecond on a service that is already struggling. The retry storm is worse than the outage that set it off.

The fix we used is old and boring. Back off exponentially, add jitter so the retries spread out instead of stacking, and cap the attempts so a dead dependency fails fast instead of getting hammered forever.

A retry is a duplicate you invited

The other half of the problem is that a retry runs the operation again, and the operation might have worked the first time. The timeout fired, the response got lost, the work still happened. Retry it and you have done it twice.

The way out, on this pipeline, was idempotency. Give every logical operation a key at the point where you call it. Store that key on the server next to the result. When the same key shows up again, return the stored result instead of doing the work a second time.

  • The key comes from the caller, not the server, so a retry carries the same one.
  • Storage is what enforces it. A unique constraint rejects the duplicate write before it can corrupt anything.
  • The check and the write have to be atomic, or two retries race each other through the gap.

Where I learned to trust it

Messages moved through Kafka and a consumer wrote each one to the database. A message was only marked done after it was fully processed and written, not when it was picked up. If the consumer crashed halfway, it restarted from the last committed offset and reprocessed the tail.

Every message near a crash gets processed at least twice, by design. It is safe only because the write is keyed on the patient and the event time, so the second pass updates the same row instead of creating a new one. The same composite-key lesson showed up later when headset retries started colliding exercise results.

Every retry is a bet that the thing you are calling can be told the same thing twice without flinching.

Retries do not make a system resilient on their own. They make it resilient when everything downstream of them is ready to see the same thing again.

Amisha

Filed under Notes · Infrastructure

Coming
Your Agent Has Amnesia. Here's the Dict That Fakes Memory.
Week 2