All writing
NotesFeb 1, 20264 min read

Multi-Cloud Without the Chaos

For a low-volume, high-reliability workload I worked on, the job queue ended up in Postgres instead of a cloud primitive. Switching clouds then cost an adapter, not a rewrite.

Running the same workload on AWS, GCP, and Azure sounds like an architecture-astronaut fantasy and an on-call nightmare. Sometimes you do not get a choice. A customer is on one cloud, a dataset is stuck on another, a contract names the third. The question, for this pipeline, was how to do it without three copies of everything.

The lock-in is in the primitives

Each provider gives you a queue, an object store, and a function runtime, and each one has a different name and a different API. SQS, Pub/Sub, Service Bus. S3, GCS, Blob. If your code calls those services directly, your code now belongs to that cloud, and every move to another one is a rewrite dressed up as a migration.

So we put a thin layer in front of them. The application talks to an interface called “queue” and “blob store”, and the cloud-specific part lives in one adapter you can swap.

Why the queue ended up in Postgres

For our workload, low volume and high reliability, the queue did not need to be a cloud service at all. It went in Postgres, and that turned out to be the right call for this scale.

  • No new infrastructure. The database was already there, already backed up, already monitored.
  • Real transactions. A job is claimed or it is not. There is no window where two workers both think they own it.
  • You can query it. Debugging a distributed queue you cannot run SQL against is a special kind of misery.
  • It runs the same everywhere. Postgres on one cloud behaves like Postgres on the next.

The whole pattern is a jobs table with a claimed_at column. A worker runs SELECT ... FOR UPDATE SKIP LOCKED, which hands it a row nobody else is holding and steps past the locked ones instead of waiting. Postgres makes that atomic. You do not have to.

The cost

Throughput. A Postgres queue is comfortable into the tens of thousands of jobs an hour. Past that you need Kafka or a real cloud queue, and the adapter is where that swap happens. I am not saying Postgres is always the answer. You should know your actual scale before you take on a distributed queue you have to operate. The same “know the size you can afford” split showed up when a MILP scheduler I wrote fell over at a hundred tasks.

A portable system is not one that uses every cloud. It is one where switching clouds costs a config change, not a quarter.

Amisha

Filed under Notes · Infrastructure

Coming
Your Agent Has Amnesia. Here's the Dict That Fakes Memory.
Week 2