Video summary

Retry & Circuit Breaker in System Design ⚡ | Netflix-Style Resilience Patterns Explained

Main summary

Key takeaways

Technology

Technological concept: “Retry storm” in microservices

In a microservice architecture, adding a retry mechanism can improve reliability, but it can also backfire.

How retries can amplify failures

If a downstream service is already down and the caller keeps retrying, load can multiply:

  • For a single user request, retries may be acceptable.
  • Under many concurrent users in production, retries can amplify dramatically (subtitles mention up to ~3x, and with concurrency reaching ~300 downstream calls), causing:
    • slower recovery
    • longer outages

This is essentially retry amplification: the caller increases downstream traffic while the downstream system cannot recover.


Demo setup used to illustrate the failure mode

Services involved

  • Caller service (described as “order service” / “order service part of this”)
  • Downstream payment service

Scenarios

1) Bad / failure mode (retry without circuit breaker)

  • Retry enabled (e.g., “three retries”)
  • No circuit breaker configured for the downstream API
  • A load test was executed against the caller service
  • Result:
    • errors like 502
    • stats showed amplification about 3x (incoming requests caused 3x downstream calls, increasing impact)

2) Fixed mode (smart retry + circuit breaker)

  • Circuit breaker enabled
  • Retries still exist, but they are guarded by breaker logic
  • Load test results showed reduced amplification:
    • incoming request count stayed around 100
    • downstream API calls reduced to ~10 (described as 0.1x amplification)
  • Circuit breaker behavior:
    • breaker moved to opened state

How to fix: Circuit breaker + smart retry configuration

The “defense” is to combine:

1) Timeouts (properly bounded)

  • Don’t wait forever.
  • Ensure calls have an upper time limit.

2) Smart retry (not blind retry)

Retry only on safe conditions, not all failures:

  • Do not retry 4xx (e.g., “non-retriable payment exception”)
  • Retry on timeouts / 5xx where retry may help
  • Limit attempts (e.g., maximum two or three attempts; not infinite)

3) Exponential backoff + jitter

Prevent synchronized retries (“thundering herd”):

  • Use increasing delays (backoff)
  • Add randomness (jitter) so clients don’t retry at the same time (e.g., staggering retry timing like 1.0s vs 1.2s)

4) Circuit breaker logic

  • The breaker checks recent calls (subtitles mention checking “last five calls,” configurable to “five/10”).
  • When thresholds indicate failure:
    • Circuit opens
    • The system calls a fallback method during the open period
    • After a configured open-state wait duration (example: 10 seconds), it transitions toward half-open behavior with limited permitted calls

Testing / verification mentioned

  • Validate behavior using:
    • Load tests (simulate concurrent production traffic)
    • Production-environment-like failures
  • Review project documentation to avoid repeating mistakes:
    • Readme.md
    • flow diagrams / descriptions to understand end-to-end execution

Main speakers / sources (as stated)

  • Subtitles indicate a single narrator/instructor (no name provided).
  • No other specific person or external source is credited in the provided subtitles.

Original video