Circuit Breaker in MuleSoft: When Retry Makes Things Worse
Code for this article:
mule4-circuit-breaker(plugin, MIT) · demo:mule4-circuit-breaker-demo-app(a real Mule application, MIT)
It’s peak sale day. Your online store’s checkout calls the payment service to authorise the customer’s card. Around midday that service starts answering a little slower than usual — nothing dramatic, just a couple of extra seconds per call. Checkout has retry configured, because nobody wants to lose a sale over one isolated timeout. The problem is that, at peak-day traffic, every timeout now means three attempts instead of one. The payment service, already close to its limit, suddenly gets three times the load it couldn’t handle in the first place — and that’s exactly when it stops responding altogether. Within minutes, every abandoned cart in the store has the same root cause: checkout, not the customer.
Any retailer, distributor or marketplace recognises this shape. An external backend degrades under load, and retry — built to protect the customer — becomes the final push that takes the backend down for good. That’s the problem a circuit breaker solves: after a configurable number of failures, it stops hammering the service that’s already struggling, fails the next requests immediately with a clear error, and tests recovery on its own — no one has to restart anything at 3pm on peak sale day.
The difference between “retry only” and “retry with a circuit breaker” doesn’t have to be taken on faith. RetryVsCircuitBreakerDemoTest, in the plugin’s repository, simulates 20 requests arriving while a backend returns HTTP:TIMEOUT on every call, run both ways:
- Retry only: each request calls the backend directly, retrying up to 3 times — all 20 exhaust their retries against the still-broken backend. Total: 60 backend calls, ~1929 ms, zero successes.
- Retry + circuit breaker: each call goes through
circuit-breaker:execute. The first 5 requests feed the sliding window and trip the circuit; the remaining 15 getCIRCUIT-BREAKER:OPENimmediately, without the backend being called again. Total: 5 backend calls, ~139 ms.
The test doesn’t rely on reading a log and eyeballing an improvement — it asserts mechanically that total backend invocations and elapsed time for the circuit-breaker run are both strictly lower than for the retry-only run. Reproduce it with:
mvn test -Dtest=RetryVsCircuitBreakerDemoTest
The pattern: three states
Circuit breaker isn’t an integration-specific invention — Martin Fowler describes the pattern, and the Azure Architecture Center documents the same shape, with three states:
- CLOSED: calls pass through normally; the circuit watches the outcomes.
- OPEN: calls are blocked immediately, without touching the backend, once the failure rate crosses a configured threshold.
- HALF_OPEN: after a wait period, a limited number of test calls decide whether the circuit closes again or reopens.
stateDiagram-v2
[*] --> CLOSED
CLOSED --> OPEN: failure rate in the window > threshold (with a minimum call count)
OPEN --> HALF_OPEN: wait period (resetTimeout) elapses
HALF_OPEN --> CLOSED: enough successful test calls
HALF_OPEN --> OPEN: any failure during the test
Two layers you can apply this at in MuleSoft
Before wiring anything into checkout’s flow, it’s worth knowing MuleSoft already ships a no-code alternative: the official circuit breaker Custom Policy, applied through API Manager. It operates at the gateway layer — protecting an API against the backend it exposes, configured through the UI, with no flow changes. That’s the right choice when you don’t own the flow’s code, or want the same circuit applied centrally across several APIs.
The plugin in this article operates at a different layer: inside the flow, wrapping one specific processor (typically an http:request) that may sit in the middle of larger business logic — like the call to the payment service inside checkout’s flow, not checkout as a whole. That’s the fit when the circuit needs to react to one specific outbound call, with fine-grained control over exactly which call it wraps and what counts as a failure.
What the plugin actually solves
Typical use, wrapping the call to the payment service:
<circuit-breaker:execute circuitBreakerKey="payment-service" failureErrorTypes="HTTP:TIMEOUT,HTTP:CONNECTIVITY">
<http:request config-ref="payment-service-config" path="/authorize" method="POST" />
</circuit-breaker:execute>
Every design decision in the plugin solves a concrete problem from the checkout scenario above:
- It reacts to degradation, not to one isolated failure. The trigger is the failure rate over a sliding window of the last N calls, not a consecutive-failure counter. On peak sale day, one slow call is noise; a payment service genuinely degrading is a trend. The sliding window catches the trend without tripping the circuit over one unlucky call.
- No one has to remember to report the outcome. A single operation (
execute) checks the state, decides whether the call proceeds, calls the wrapped service and reports the result — with no separate call the flow author has to remember. Under deadline pressure (and peak sale day is always under pressure), forgetting to report an outcome is exactly the kind of mistake that’s easy to make and hard to catch in a flow review. - A declined card isn’t a reason to shut checkout down for everyone. Only the error types configured as infrastructure failures (
failureErrorTypes, e.g. timeout and connectivity) count toward the circuit. A customer typing in an expired card produces a business error — it shouldn’t count against the payment service, or affect every other customer’s checkout. - It knows your application doesn’t run on a single replica. Circuit state sits behind a port (
CircuitStateStore), with a default adapter backed by the Mule Runtime’s Object Store. That’s what makes the next point possible: a circuit that’s a shared fact across replicas, not an illusion of protection that only covers one of them.
A blocked call raises CIRCUIT-BREAKER:OPEN, for the flow author to handle like any other typed error — for example, returning “payment temporarily unavailable, please try again shortly” to the customer instead of leaving checkout hanging on a timeout.
Back to peak sale day: with the circuit breaker wrapped around the call to the payment service, the first few failures still happen — there’s no magic that avoids that —, but after a handful of them, the circuit opens, the next requests fail immediately with a clear message instead of waiting on a timeout, and the payment service stops receiving traffic it has no way to handle:
sequenceDiagram
participant C as Checkout
participant CB as circuit-breaker:execute
participant P as Payment service
Note over C,P: Circuit CLOSED — normal traffic
C->>CB: authorise card
CB->>P: HTTP request
P-->>CB: 200 OK
CB-->>C: 200 OK
Note over C,P: Backend degrading — failures build up in the window
C->>CB: authorise card
CB->>P: HTTP request
P-->>CB: HTTP:TIMEOUT
CB-->>C: error (counts as a failure)
Note over C,P: Circuit OPEN — backend no longer called
C->>CB: authorise card
CB-->>C: CIRCUIT-BREAKER:OPEN (immediate)
It tests recovery on its own, every resetTimeout, with no one having to restart anything at 3pm on peak sale day.
The local evidence
mule4-circuit-breaker-demo-app is a real Mule application — not a test mock — with a client endpoint that wraps a simulated backend in circuit-breaker:execute. The backend’s behaviour is controlled by a header (X-Demo-Backend-Mode) the client forwards unchanged, so “the backend recovers” is just the next call sending a different header value. A few scenarios, straight from the README:
# Closed circuit — calls pass through normally
curl -i http://localhost:8081/orders/demo -H 'X-Demo-Backend-Mode: ok'
# Opens on failure rate — the first 5 calls are real 502s;
# the 6th is already 503, without the backend being called
for i in 1 2 3 4 5 6; do curl -i http://localhost:8081/orders/demo -H 'X-Demo-Backend-Mode: always_fail'; done
# A business error never opens the circuit — always 400, no matter how many times you repeat it
for i in 1 2 3 4 5; do curl -i http://localhost:8081/orders/demo -H 'X-Demo-Backend-Mode: business_error'; done
Different keys (checkout, shipping) isolate independent circuits, and the default 10-second resetTimeout moves the circuit from OPEN to HALF_OPEN and back to CLOSED once a test call succeeds.
Shared state across replicas: what changes on CloudHub 2.0
Checkout doesn’t run on a single replica in production — the normal case on CloudHub 2.0 is several. If every replica kept its own circuit state, one customer could land on a replica with the circuit open while another lands on a replica that hasn’t even noticed the payment service is down yet. That’s why state sits behind the CircuitStateStore port: so a circuit opened on one replica is a shared fact, not a local illusion.
The demo app was deployed for real on CloudHub 2.0, with 2 replicas at 0.1 vCore each. Repeating the “opens on failure rate” scenario against that deployment, reading X-Replica-Id to see which replica answered each call:
| Call | Replica | Result |
|---|---|---|
| 1–8 | alternating between the two replicas | 502 (real backend failure) |
| 9 | one replica | 503 circuit_open |
| 10+ | both replicas, at different points | 503 circuit_open |
The central point: neither replica alone reached 5 failures — the circuit opened from the combined failure count across both, and once open, both replicas independently started reporting circuit_open, without ever contacting the backend again. That’s only possible because the state lives in the runtime’s shared, persistent Object Store, not in either replica’s own memory.
That same run also reproduced, live, two limits worth knowing before trusting this mechanism blindly:
- No cross-replica lock. Right after the circuit reported
circuit_openon one replica, the very next call — landing on the other replica — went through as a normal failure instead of being blocked. The circuit reopened a call or two later. This is a real race: the plugin’s lock only serialises reads/writes within one replica; two replicas writing close together race at the storage level, and a write from one can silently overwrite the other’s. - No guaranteed coordination on a brand-new key’s first write. With two concurrent requests against a brand-new key, the circuit only opened after 8 real failures, not the 5 that were configured. That’s not a formal proof — there’s no way to inspect the Object Store’s internal partitions from outside a pod — but it’s consistent with the same kind of race: treat a brand-new key’s first few calls under concurrent traffic as unreliable for opening exactly on schedule.
And the cost of all this: +5.6 ms median (+5.0 ms mean, +2.3 ms p95) per circuit-breaker:execute call, measured by comparing an endpoint that goes through the circuit breaker against an identical one that doesn’t, under otherwise idle load — the cost of two locked read-and-write round trips to the Object Store per call, against CloudHub 2.0’s managed backend over the network.
The full method, raw samples and the measurement script are versioned in the demo app’s repository, under evidence/BRU-58/.
Limits, and when not to use this
- No cross-replica lock. Replicas writing close together race at the storage level; one’s write can silently overwrite another’s.
- No guaranteed coordination on a key’s first write, under concurrent traffic.
- Added latency: +5.6 ms median per call on CloudHub 2.0, measured under idle load — not zero, and it grows with the number of Object Store operations per call.
- Depends on the runtime’s persistent Object Store actually implementing cross-replica sharing — the plugin doesn’t replicate state on its own; it trusts the same mechanism CloudHub 2.0 already uses for that.
If your integration needs an exact count under heavy concurrency — not just “opens close to the configured threshold” — this design doesn’t guarantee that; it would need external coordination that neither this plugin nor Mule’s Object Store set out to provide.
Reproduce it yourself
Locally, with no cloud account required:
# Retry-only vs. circuit breaker contrast (mule4-circuit-breaker)
mvn test -Dtest=RetryVsCircuitBreakerDemoTest
# Full suite: state machine, error classification, storage
mvn clean test
# Package and run the demo app locally (mule4-circuit-breaker-demo-app)
mvn clean package
# copy the jar into <MULE_HOME>/apps/, start the runtime, then repeat the curls
# from the README's "Scenarios" section
If you’d rather validate this in a cloud environment than locally, both repositories already ship with an azure-pipelines.yml: publishing the plugin to Exchange and deploying the demo app to CloudHub 2.0 doesn’t require any account of mine specifically, just the same setup covered in the MuleSoft delivery with Azure DevOps series — the variable group, environments and service connection walkthrough is in the Azure Pipelines templates README, and the CloudHub 2.0 deploy with approval is covered in Deploying to CloudHub 2.0 from Azure Pipelines. This article’s CloudHub 2.0 evidence — cross-replica proof, the first-write anomaly and the latency measurement with script and raw samples — lives under evidence/BRU-58/ in the demo app’s repository.
Checklist
- Does your backend see intermittent spikes or sustained degradation? A sliding window handles both better than a consecutive-failure counter.
- Is more than one replica running the same integration? If so, circuit state needs to be shared, not just local.
- Is retry protecting the caller, or piling more load onto a backend that’s already going down?
- Does your error handler distinguish infrastructure failure from business error before deciding whether it counts toward the circuit?