Backstory first. We (my company) run a high-throughput SaaS platform written in Go, and sitting at the heart of it is a messaging service that pushes out time-sensitive jobs: OTPs, transactional sends, bulk processing, and asynchronous callbacks from third parties. The kind of stuff where if it’s late, someone notices but not in a good way.
Where We Started
Our initial pipeline looked like this:
API Request → Main App → Pub/Sub → Subscriber App
It worked, but something kept poking at me.
First, Pub/Sub is an external dependency. Call it vendor lock-in if you want, it’s the same thing: credentials and an entire fleet of infra we’re suddenly on the hook for. A chunk of our stack rented from someone else’s vendor for no real upside.
Then the development friction, which becomes a burden over time. Wanna test/develop locally? yeah you gonna spin up the Pub/Sub emulator first.
And the config. We had something like 10+ environment variables floating around just to configure messaging. That’s not a config, that’s an extra job.
Now here’s the honest part: the reason we moved wasn’t performance, it was simply control. Every complaint above is the same complaint wearing a different hat: the critical path of our product depended on a system we don’t own, can’t debug easily by writing a query, and can’t run on a laptop without drama. We just wanted that control back.
So we pulled the plug and swapped it for RiverQueue, a Go job queue that’s backed entirely by PostgreSQL. Nothing exotic: it’s a library that runs inside your binary, and your “task/job” is just rows sitting in a plain old Postgres table.
What RiverQueue Gives Us Out Of The Box
- Postgres-backed, so zero new infrastructure to babysit
- Transactional enqueue, jobs go in atomically with whatever else you’re doing
- Typed jobs via Go generics (compile-time safety)
- Per-queue
MaxWorkers, so each pipeline gets exactly as many workers as we want - Built-in retry policy via
MaxAttempts - Native
pgx v5integration, so it slots into our Go code like it was always meant to be there
The New Architecture
API Request → Main App → Postgres → Worker App
The messaging path now stays inside our own DB.
We also split things into separate queues by workload, so nothing waits behind something it shouldn’t:
- Time-critical jobs, the ones with someone waiting in real time
- Bulk Processing, huge amounts of processing separated from the regular traffic
- Callback, handling async callbacks from external services
- Report / export, Generating reports uses a decent amount of resources and won’t bother the Main App
The important part: keep bulk processing separate from the regular traffic. Mix the two and a slow batch of HTTP calls blocks the entire line (aka head-of-line blocking). Nobody wants that.
Publishing a Job in Go
Enqueuing a job is boring, but I take that as a compliment:
riverQueue.Insert(ctx, dto.TimeCriticalJob{
LogID: id,
Payload: payload,
}, &queue.InsertOpts{
Queue: "time_critical_jobs",
MaxAttempts: 1,
})
Then a worker is just a method on a struct:
func (w *TimeCriticalWorker) Work(ctx context.Context, job *river.Job[dto.TimeCriticalJob]) error {
return w.Process(ctx, job.Args)
}
And that’s the whole worker.
They automatically pick up any available jobs that match their queue and do the work they are assigned to, just like people queueing up for their turn at a service counter.
The Migration, Step by Step
We didn’t big-bang this. We ran a dual-path for about 2.5 months with a canary percentage per workload. The phases went:
- Publish sub only, the new path just watches, nothing acts
- Build a worker for each type
- Dual-path with a canary %
- Ramp slowly to 100%, watching latency and error rates the whole way
- Soak for weeks at 100%
- Delete all the Pub/Sub code
Why this worked so well: with the default at 0, bugs can’t sneak into the new path; you control it per workload type; zero downtime; rollback is dead simple; and that long soak weeded out the weird stuff a weekend dash never would.
The canary function itself is tiny:
func canary(pct int) bool {
if pct <= 0 { return false }
if pct >= 100 { return true }
return rand.Float64() < float64(pct)/100
}
Where We Landed
- Messaging env vars: from 10+ down to zero. Yes, zero
- ~1,000 lines of broker plumbing code deleted, in the bin
- Local dev is now just
go run ...., no emulator in sight - Retry settings live in code now, a
MaxAttemptsfield right where the job is inserted, instead of in platform config - Debugging is now a SQL query away, nothing quite beats this feeling
- Each queue can be scaled independently, and it runs anywhere: cloud, bare metal, your laptop
Did It Cost Us Anything In Performance?
The purpose of the replacement here was initially to have more control. That sounds nice until it turns out to be slower. So we measured it anyway. Did it cost us? I think it is negligible.
-
Insert latency (API → queue): ~5ms with RiverQueue (left) vs teens of ms (mostly two digits) with Pub/Sub (right)

-
Processing throughput: 58–60 rps with RiverQueue vs 57–60 rps with Pub/Sub
-
Pod resources (Kubernetes):
cpu: 100m / memory: 200mwith RiverQueue vscpu: 150m / memory: 300mwith Pub/Sub
Same throughput, a few milliseconds off the insert path, and it fits in a smaller pod. Nice, but it was never the point. We didn’t move to go faster. We moved so that when something goes wrong, the thing that’s wrong is a table we own, in a database we already run, with a query we can write ourselves.
Wrap Up
Pub/Sub is fully out of the repo now. The messaging path is jobs and workers running inside the same binary as the app, sharing the Postgres we already had in production. No new infra, no emulator, one less system for us to own.
That’s it for this post. See you on the next one!