lsfera View the repository →

Circuit Breaker: Unboxing

A recent interview put me in front of this exact use case, and I gave an average answer that left a bad taste in my mouth. It kept nagging me in the days after, so I decided to dig deeper.

Circuit breakers look simple until the caller is a fleet, not a single process. This article works through the pattern under a reliability lens, using this repository's egress breaker as the running example.

Definition

A circuit breaker stops a caller from making a request it already has reason to believe will fail. It sits between a caller and a dependency, watches recent outcomes, and switches among three states: closed (calls pass through, failures are counted), open (calls fail immediately, without touching the dependency), and half-open (a small number of probes test whether the dependency has recovered).

The trip condition is a threshold on recent failures — a rate, a count, or both — never a single failure. The goal is not to catch one bad call. It is to stop sending load into a dependency that is already struggling, before every caller's own retries turn a slow degradation into an outage.

stateDiagram-v2
  [*] --> Closed
  note left of Closed: calls pass through,
successes keep it closed Closed --> Open: failure threshold exceeded Open --> HalfOpen: open timeout
elapses HalfOpen --> Closed: probe succeeds HalfOpen --> Open: probe fails,
backoff doubles
Closed passes calls through and counts failures; open fails fast without touching the dependency; half-open spends a small number of probes deciding which way to go next.

Anecdotal history

This repository sits on the infrastructure side of that fork, which is where our shape of the problem starts.

Our shape of the problem

The base scenario

flowchart LR
  producer["Producer"] --> queue[("RabbitMQ broker")]

  subgraph competing_consumers["competing consumers"]
    direction LR
    c1["Consumer 1"]
    c2["Consumer 2"]
    c3["Consumer n"]
  end

  queue --> c1
  queue --> c2
  queue --> c3

  c1 --> api[("Third-party API
(flaky-upstream)")] c2 --> api c3 --> api

One producer, one durable queue, a fleet of daemons competing for the same messages — no coordination between them by default — each making its own outbound call to a third party whose load balancer, health checks, and backend topology are entirely theirs. Every fact below follows from this picture.

The textbook breaker assumes one caller and one dependency. We have neither:

Each breaks the textbook picture alone; together they compound.

Many callers, one opinion needed. A breaker per consumer means each forms its own view from its own sample. At our fleet's shape — five daemons — a 15-second outage opens their breakers about 27 times between them, and all five are in the same state only 56% of the time (median of thirteen runs). One outage, many opinions.

Their load balancer, not ours. Proxy-level enforcement (Envoy outlier detection) ejects the specific host that is failing. A third party's load balancer hides its hosts: we see one address. We can only detect at the aggregate level — choosing when to stop calling them, not which machine to avoid.

Whose failure is it, anyway. A request shed by our own egress path can look identical to a real upstream failure. This repository hit it directly: an upstream returning no errors, only 300 ms slower, yet every 5xx callers saw came from the local proxy's concurrency shedding — 90,753 local errors in two minutes. A breaker that can't tell the two apart trips on the wrong evidence.

So this article does not stop at "add a breaker." There are many moving parts, and every decision comes at a cost. The analysis rests on three pillars:

Four frames of the same producer, work queue, consumers and third party. 1, no breaker: every consumer keeps calling and messages dead-letter, about 4,000 in 20 seconds. 2, a breaker in every process: one consumer open, one closed, one half-open, still 1,577 to 2,246 dead letters in 15 seconds. 3, the breaker moves into RabbitMQ: consumers stop consuming, one probe holds the permit, no dead letters. 4, platform: Envoy between consumers and the third party, an aggregator publishing one verdict per API to the consumers and to subscribers, no dead letters.
The same fleet, four places for the breaker: nowhere, in every process, in the broker (steps 03–04), and at platform level. The dead-letter counts are each step's measured outage.

Protecting the request path and telling the rest of the fleet need different answers, at different places in the stack. The linked repository weighs each mechanism against these pillars, one branch per step.

The steps in the repository

Each step is a branch with its own code, measurements and README.

StepWhere the state livesWhat open doesMeasured in an outage
01 · No breakernowherenothing: every message spends its 3 attempts≈ 4,000 dead-lettered in 20 s
02 · In every processeach replica's memoryrejects locally, spending the message's attempts1,577–2,246 dead-lettered in 15 s; ~27 openings for one outage
03 · Coordinated by the brokereach replica, plus a shared probe permitreleases the message without spending an attempt0 dead-lettered in 40 s; one probe at a time instead of 6
04 · Held by the brokerRabbitMQ: a consumer on or off, a token in a delay chainstops consuming; the work waits in the queue0 dead-lettered where cockatiel lost 2,745
05 · Platform control planeEnvoy per replica, and one verdict per API in an aggregatorEnvoy ejects hosts; the fleet stops consuming0 lost, 0 dead-lettered across ten chaos faults

Sources