Circuit Breaker: Unboxing
A recent interview put me in front of this exact use case, and I gave an average answer that left a bad taste in my mouth. It kept nagging me in the days after, so I decided to dig deeper.
Circuit breakers look simple until the caller is a fleet, not a single process. This article works through the pattern under a reliability lens, using this repository's egress breaker as the running example.
Definition
A circuit breaker stops a caller from making a request it already has reason to believe will fail. It sits between a caller and a dependency, watches recent outcomes, and switches among three states: closed (calls pass through, failures are counted), open (calls fail immediately, without touching the dependency), and half-open (a small number of probes test whether the dependency has recovered).
The trip condition is a threshold on recent failures — a rate, a count, or both — never a single failure. The goal is not to catch one bad call. It is to stop sending load into a dependency that is already struggling, before every caller's own retries turn a slow degradation into an outage.
stateDiagram-v2 [*] --> Closed note left of Closed: calls pass through,
successes keep it closed Closed --> Open: failure threshold exceeded Open --> HalfOpen: open timeout
elapses HalfOpen --> Closed: probe succeeds HalfOpen --> Open: probe fails,
backoff doubles
Anecdotal history
- Origin: electrical engineering — a breaker trips on overcurrent to protect the wiring, fails fast on purpose, and needs a reset.
- 2007: Michael Nygard's Release It! names it as a software stability pattern: turn a slow downstream failure into a fast, cheap one.
- 2012: Netflix's Hystrix makes it the default microservices library, with Turbine aggregating dashboards fleet-wide.
- 2018 on: Hystrix goes into maintenance; Resilience4j keeps the per-process model, while Envoy outlier detection (Istio) and Linkerd move the breaker into the proxy.
This repository sits on the infrastructure side of that fork, which is where our shape of the problem starts.
Our shape of the problem
The base scenario
flowchart LR
producer["Producer"] --> queue[("RabbitMQ broker")]
subgraph competing_consumers["competing consumers"]
direction LR
c1["Consumer 1"]
c2["Consumer 2"]
c3["Consumer n"]
end
queue --> c1
queue --> c2
queue --> c3
c1 --> api[("Third-party API
(flaky-upstream)")]
c2 --> api
c3 --> api
One producer, one durable queue, a fleet of daemons competing for the same messages — no coordination between them by default — each making its own outbound call to a third party whose load balancer, health checks, and backend topology are entirely theirs. Every fact below follows from this picture.
The textbook breaker assumes one caller and one dependency. We have neither:
- Caller: a fleet of competing consumers on one queue, uncoordinated by default.
- Dependency: a third party behind a load balancer we don't operate, configure, or get metrics from.
Each breaks the textbook picture alone; together they compound.
Many callers, one opinion needed. A breaker per consumer means each forms its own view from its own sample. At our fleet's shape — five daemons — a 15-second outage opens their breakers about 27 times between them, and all five are in the same state only 56% of the time (median of thirteen runs). One outage, many opinions.
Their load balancer, not ours. Proxy-level enforcement (Envoy outlier detection) ejects the specific host that is failing. A third party's load balancer hides its hosts: we see one address. We can only detect at the aggregate level — choosing when to stop calling them, not which machine to avoid.
Whose failure is it, anyway. A request shed by our own egress path can look identical to a real upstream failure. This repository hit it directly: an upstream returning no errors, only 300 ms slower, yet every 5xx callers saw came from the local proxy's concurrency shedding — 90,753 local errors in two minutes. A breaker that can't tell the two apart trips on the wrong evidence.
So this article does not stop at "add a breaker." There are many moving parts, and every decision comes at a cost. The analysis rests on three pillars:
- Correctness: trip on the right evidence, and give the fleet one verdict.
- Self-healing: detect recovery and resume traffic without a human.
- Operational complexity: what it costs to run, configure, and debug.
Protecting the request path and telling the rest of the fleet need different answers, at different places in the stack. The linked repository weighs each mechanism against these pillars, one branch per step.
The steps in the repository
Each step is a branch with its own code, measurements and README.
| Step | Where the state lives | What open does | Measured in an outage |
|---|---|---|---|
| 01 · No breaker | nowhere | nothing: every message spends its 3 attempts | ≈ 4,000 dead-lettered in 20 s |
| 02 · In every process | each replica's memory | rejects locally, spending the message's attempts | 1,577–2,246 dead-lettered in 15 s; ~27 openings for one outage |
| 03 · Coordinated by the broker | each replica, plus a shared probe permit | releases the message without spending an attempt | 0 dead-lettered in 40 s; one probe at a time instead of 6 |
| 04 · Held by the broker | RabbitMQ: a consumer on or off, a token in a delay chain | stops consuming; the work waits in the queue | 0 dead-lettered where cockatiel lost 2,745 |
| 05 · Platform control plane | Envoy per replica, and one verdict per API in an aggregator | Envoy ejects hosts; the fleet stops consuming | 0 lost, 0 dead-lettered across ten chaos faults |
Sources
- Michael T. Nygard, Release It!, Second Edition: Design and Deploy Production-Ready Software, The Pragmatic Programmers. First edition published 2007 by Pragmatic Bookshelf.
- Netflix, Netflix/Hystrix (GitHub). README: "no longer in active development, and is currently in maintenance mode" since 2018, final release 1.5.18.
- Netflix, Netflix/Turbine (GitHub), the SSE stream aggregator Hystrix dashboards were built on. Archived (read-only) December 19, 2025.