A decision framework for the moment your team says “let’s just replace it with a queue.”

At some point in almost every distributed system, a familiar conversation happens. Service A calls Service B over REST. Service B is flaky, or slow, or occasionally down. Payloads are being lost. Someone says the words: “Let’s replace this with a message broker so we don’t lose data anymore.”

The instinct is understandable. It’s also, in most of the cases I’ve seen, the wrong first move — not because event-driven is bad, but because the reasoning skips the two questions that actually decide the design.

This post walks through those questions, the reasons teams commonly cite, why almost none of them are exclusive to event-driven, and the specific conditions under which a broker genuinely becomes the right call.


THE COMMON TRIGGER

“Our REST calls fail — should we go event-driven?”

Here’s the scenario, stripped down:

Service A ──REST──► Service B
(down / slow / flaky)
✕ timeouts / 5xx / lost payloads

Two responses show up:

The instinct

“Let’s replace REST with a message broker so the payload is never lost — the queue will hold it until Service B is healthy again.”

The right question

“Is event-driven the cheapest way to get what we actually need — or would a durable synchronous design get us there with less machinery?”

The instinct treats “event-driven” as the answer. The right question treats “durability” as the problem, and asks what design solves it at the lowest cost. Those are very different starting points.


REFRAME

Two questions, asked in order

Rather than asking “should we go event-driven?”, ask two questions in sequence:

1. The gate — does the caller need a synchronous response?

If yes, event-driven is off the table. You would only be rebuilding request/response on top of async infrastructure, and you would end up worse off than a well-designed REST call. If no, event-driven is allowed. That is very different from obligatory.

2. The real driver — are you about to reinvent a broker on top of REST?

Almost every benefit attributed to event-driven has a synchronous equivalent. The decision is not “do I have one of these needs?” — it is “how many of these are stacking on the same workflow, and at what cost am I building each of them by hand?”

Ask them in that order. Skip step one and you’ll end up building async infrastructure for workflows that fundamentally need a response, which is the worst of both worlds.


GATE 1

Sync response required? Then stop here.

The clearest way to see this is side by side.

Synchronous expectation. The caller cannot proceed until it knows the outcome. Payment confirmations, credit checks, auth flows — anything that shapes the next thing the user sees. For these, event-driven is off the table. Fix the REST call itself.

Async-tolerant. The caller’s job ends once the payload is safely handed off. Notifications, downstream index updates, audit trails, most kinds of ingestion. Here event-driven is now allowed. Whether it’s the right choice is a separate question, which the rest of this post is about.

Async tolerance is a prerequisite, not a reason. It opens the door — it does not decide the design.


COMMON JUSTIFICATIONS

The reasons teams give for going event-driven

When event-driven comes up, the reasons cited usually cluster into six:

  • Durability — the payload must not be lost if the consumer is down.
  • Retries & failure handling — the downstream is flaky, we want automatic recovery.
  • Decoupling — producer and consumer should evolve independently.
  • Load leveling — bursty producer, slower consumer, we need a buffer.
  • Fan-out — one event should trigger several downstream actions.
  • Replay & audit — re-process history, reconstruct state after a bug.

Each of these sounds decisive on its own. It usually isn’t. Every one has a synchronous solution that is often cheaper than adopting a broker.


THE UNCOMFORTABLE TRUTH

Almost every reason has a synchronous equivalent

Cited reasonSynchronous solution (no broker)
DurabilityTransactional outbox + retry worker
Retries / failureBackoff, idempotency keys, circuit breaker
DecouplingWell-designed APIs, versioning, contracts
Load levelingThread pools, DB-backed work queues
Fan-outDispatcher service; one call → N downstreams
Replay & auditAppend-only event log in your own DB

Any single need, in isolation, usually has a cheaper synchronous answer.

That doesn’t make event-driven wrong. It means the argument “we need X, therefore event-driven” is almost always weaker than it looks — because X on its own rarely justifies the cost of adopting a broker.


THE MIDDLE ISN’T EMPTY

The escalation ladder from REST to event-driven

Before jumping from “fix the REST call” to “adopt a broker,” it’s worth naming the rungs in between. Most teams that end up in event-driven-hell got there by skipping them.

  1. Bare REST. No retries, no idempotency. Fine for internal, reliable dependencies.
  2. REST + retries with backoff and jitter. Handles transient failures. Cheap.
  3. REST + retries + idempotency keys. Now retries are safe to actually try. Libraries like Resilience4j (JVM) or Polly (.NET) handle the retry side; idempotency is your API design.
  4. REST + circuit breaker. Fail fast when the downstream is genuinely down, instead of retrying into a wall.
  5. REST + transactional outbox. Guaranteed processing without a broker. Debezium is one common way to publish the outbox reliably, but a plain polling worker also works.
  6. Event-driven with a broker — Kafka, RabbitMQ, SQS, whatever fits.

Most flaky-REST problems are solved somewhere between rungs 2 and 5. Teams that adopt rung 6 without trying rungs 3 and 5 often discover the hard way that rungs 3 and 5 were what they needed all along.


SYNC + DURABLE

The transactional outbox pattern

Since rung 5 is where most durability arguments land, it’s worth understanding well.

The core idea: persist the payload in the same database transaction as the business change. A separate worker delivers it via REST, with retries. Durable, decoupled from downstream availability, no broker required.

Two things this gives you, and two things it doesn’t:

What it gives you. Guaranteed processing without a message broker. Service A’s work ends at DB commit — everything after that is the worker’s problem. Tools like Debezium can even publish outbox rows into a broker later, when you actually need one.

What it does not give you. Fan-out, replay, or true producer/consumer independence. It’s still point-to-point delivery.

The outbox pattern solves durability. It does not solve fan-out. It does not solve replay. If your problem is only durability, you’re done — you don’t need a broker. If your problem is durability plus fan-out plus replay, the calculus starts to change.


A CAUTIONARY TALE

What “let’s just use Kafka” actually cost

A team I worked with hit this exact scenario. Their user-facing service was calling an internal fulfillment service over REST; the fulfillment service was flaky, and about 0.3% of calls were dropping payloads. The instinct — “let’s put Kafka between them” — took over the design conversation, and within a quarter they were live on a broker.

Six months later, most of an engineering-year had gone into things they hadn’t planned for:

  • Retrofitting idempotency into consumers that had been built assuming exactly-once REST semantics. Every duplicate delivery became a customer-support ticket.
  • Standing up distributed tracing (OpenTelemetry, correlation IDs, log aggregation) because the previous “look at the call stack” debugging story no longer worked.
  • Building a dead-letter reprocessing tool from scratch, after a poison message blocked a partition for three hours on a Sunday.
  • Answering compliance questions about where messages were stored, for how long, and who could read them.

The original 0.3% failure rate? An outbox and a retry policy — probably two weeks of work — would have taken it to zero. The broker was the right long-term architecture eventually (they did legitimately need fan-out to analytics and a search index), but adopting it as a fix for flaky REST added most of a year of work to a problem that had a two-week answer.

I’ve seen close variants of this play out on four teams. The pattern is always the same: the broker was adopted for the wrong reason, and the real trade-offs of event-driven only became visible in the following two quarters.


WHEN EVENT-DRIVEN WINS

The real driver is compounding needs

Here’s the argument in one line:

Any one benefit has a cheaper sync answer. When several stack on the same workflow, the cost of hand-building each on top of REST starts to look like a worse version of a broker. That inflection point — not any single reason — is when event-driven becomes the right call.

Practically: at one need, sync patterns win comfortably. At two, they usually still do. At three, it becomes a judgment call — and the honest question is whether you’re building isolated features or slowly assembling a broker. At four or more, a broker is almost certainly the right shape for the problem.


COUNT THESE

The conditions that stack toward event-driven

Four conditions on the workflow, plus one honest question about your team.

On the workflow

1. Multiple independent consumers. Several services need the same payload reliably. An outbox per producer-consumer pair gets ugly fast; a broker fans out from one write.

2. Replay of historical events. New consumers need to process events that already happened, or you need to re-derive state after a bug. Outboxes typically delete rows on delivery — they don’t give you replay. This is one of the few genuinely broker-only capabilities: you can build an append-only event table in your own database, but by the time you’ve added retention, indexing, and consumer position tracking, you’re most of the way to a broker anyway. If you need replay for consumers that don’t exist yet — a compliance tool next quarter, an analytics team next year — that especially points toward an event log.

3. Full temporal decoupling. Producer and consumer lifecycles are independent. Once the producer hands off, it is fully done — the broker owns delivery, not you. With an outbox, your worker still owns delivery.

4. Real buffering / backpressure. Bursty producers, slower consumers, and a thread pool or DB queue isn’t the right shape for the traffic profile. Brokers give you natural buffering; DB-backed queues eventually strain the DB.

And one honest question about your team

Do you have the appetite to operate this? A broker cluster is real infrastructure to run, upgrade, secure, and monitor. If the honest answer is “we’d rather lean on the broker’s persistence, acks, DLQs, and consumer offsets than build and operate that ourselves” — that’s a legitimate reason. But it’s different in kind from the four above. Those are properties of the workflow; this is a property of your team.

Count the workflow conditions. One doesn’t move the needle. Two is a conversation. Three is a design decision. Then ask the appetite question separately — a workflow that clearly wants event-driven, running on a team that isn’t ready to operate a broker, is a legitimate reason to wait.


TRADE-OFFS

What event-driven costs you

These are real engineering and operational costs. Not footnotes.

  • Eventual consistency. The whole system is now “correct given enough time,” not correct at the moment of the write. Downstream reads may lag, and users may see it.
  • Duplicate delivery. At-least-once is the norm in Kafka, RabbitMQ, and SQS. Every consumer must be idempotent — designed for it, not bolted on later.
  • Ordering. Global ordering is rarely free. Per-key ordering (Kafka’s partition model) costs partitioning discipline and shapes your producer design.
  • Poison messages and DLQs. Bad payloads block partitions. You need dead-letter queues, quarantine flows, and a runbook to reprocess them.
  • Debugging and tracing. A request no longer has a single call stack. Distributed tracing (OpenTelemetry), correlation IDs, and log aggregation become non-optional.
  • Operational surface. Broker cluster, schema registry, consumer groups, offsets, retention policies. Real infrastructure to run, upgrade, and monitor.

The pattern from the cautionary tale is worth naming explicitly: teams that adopt a broker without first investing in idempotency and tracing build a more fragile event-driven system than the REST call it replaced. Both go in on day one, or neither goes in.


THE FRAMEWORK

Ask in order. Do not skip step 1.

Three steps. Ask them in order. Skipping step one lands you in async-hell for workflows that were synchronous all along. Skipping step two lands you in broker-hell for problems the outbox pattern would have solved.


BACK TO THE SCENARIO

The same conversation, asked differently

Return to where we started. Service A → REST → Service B is flaky. Someone says: “let’s use a broker.”

Now the question has structure.

Does the caller need a synchronous response? No — the downstream is fire-and-forget in the user’s mind. Good, event-driven is on the table.

How many workflow conditions apply? One. Durability. There’s no second consumer, no replay requirement, no unusual traffic profile.

The answer becomes obvious. Rung 5 on the ladder: transactional outbox plus a retry worker. A sprint or two of work, a database table and a worker process added to the operational footprint. The broker discussion happens later, if at all, when the second and third conditions actually appear.

That’s the whole framework: async tolerance opens the door; stacking needs decide whether to walk through it.


TL;DR

Don’t switch to event-driven to fix REST failures. Switch when several async needs stack up.

  • If the caller needs a synchronous response, event-driven is off the table. Fix the REST call — retries, idempotency, circuit breakers.
  • If the caller doesn’t need one, walk the escalation ladder before reaching for a broker. Bare REST → retries + idempotency → circuit breaker → transactional outbox. Most flaky-REST problems die on rung 3 or 5.
  • Count the four workflow conditions: multiple consumers, replay, temporal decoupling, real buffering. One or two? Outbox is almost always the right shape.
  • Three or four conditions? Consider a broker — and ask honestly whether your team has the appetite to operate one.
  • If you do adopt event-driven, adopt idempotency and distributed tracing alongside it. Both on day one, or neither. A broker without them is more fragile than the REST call it replaced.

Leave a comment

I’m Mahesh

Welcome to MaheshNotes, a space on the internet where I like to share my knowledge and experience as a software engineer.

Let’s connect