Video summary
Retry & Circuit Breaker in System Design ⚡ | Netflix-Style Resilience Patterns Explained
Main summary
Key takeaways
Technological concept: “Retry storm” in microservices
In a microservice architecture, adding a retry mechanism can improve reliability, but it can also backfire.
How retries can amplify failures
If a downstream service is already down and the caller keeps retrying, load can multiply:
- For a single user request, retries may be acceptable.
- Under many concurrent users in production, retries can amplify dramatically (subtitles mention up to ~3x, and with concurrency reaching ~300 downstream calls), causing:
- slower recovery
- longer outages
This is essentially retry amplification: the caller increases downstream traffic while the downstream system cannot recover.
Demo setup used to illustrate the failure mode
Services involved
- Caller service (described as “order service” / “order service part of this”)
- Downstream payment service
Scenarios
1) Bad / failure mode (retry without circuit breaker)
- Retry enabled (e.g., “three retries”)
- No circuit breaker configured for the downstream API
- A load test was executed against the caller service
- Result:
- errors like 502
- stats showed amplification about 3x (incoming requests caused 3x downstream calls, increasing impact)
2) Fixed mode (smart retry + circuit breaker)
- Circuit breaker enabled
- Retries still exist, but they are guarded by breaker logic
- Load test results showed reduced amplification:
- incoming request count stayed around 100
- downstream API calls reduced to ~10 (described as 0.1x amplification)
- Circuit breaker behavior:
- breaker moved to opened state
How to fix: Circuit breaker + smart retry configuration
The “defense” is to combine:
1) Timeouts (properly bounded)
- Don’t wait forever.
- Ensure calls have an upper time limit.
2) Smart retry (not blind retry)
Retry only on safe conditions, not all failures:
- Do not retry 4xx (e.g., “non-retriable payment exception”)
- Retry on timeouts / 5xx where retry may help
- Limit attempts (e.g., maximum two or three attempts; not infinite)
3) Exponential backoff + jitter
Prevent synchronized retries (“thundering herd”):
- Use increasing delays (backoff)
- Add randomness (jitter) so clients don’t retry at the same time (e.g., staggering retry timing like 1.0s vs 1.2s)
4) Circuit breaker logic
- The breaker checks recent calls (subtitles mention checking “last five calls,” configurable to “five/10”).
- When thresholds indicate failure:
- Circuit opens
- The system calls a fallback method during the open period
- After a configured open-state wait duration (example: 10 seconds), it transitions toward half-open behavior with limited permitted calls
Testing / verification mentioned
- Validate behavior using:
- Load tests (simulate concurrent production traffic)
- Production-environment-like failures
- Review project documentation to avoid repeating mistakes:
- Readme.md
- flow diagrams / descriptions to understand end-to-end execution
Main speakers / sources (as stated)
- Subtitles indicate a single narrator/instructor (no name provided).
- No other specific person or external source is credited in the provided subtitles.