resilience
Hedged Requests
Fire duplicate requests after a delay — use whichever responds first to tame tail latency
StatusIDLE
Elapsed0.0s
Rounds0
Successes0
Failures0
Latency Saved0ms
Request Race Lanes
Click "Send Request" to launch a hedged request
// configuration
PRESETS:
// event log
No events yet. Click "Send Request" to begin.
// how it works
- Send the primary request to the service.
- If no response after hedgeDelay, fire a duplicate (hedge) to another replica.
- Whichever responds first wins — cancel the others.
- Repeat up to maxHedges additional copies.
- The hedge delay is tuned to the p95/p99 of the service — only slow requests get hedged.
// trade-offs
- Dramatically reduces tail latency (p99)
- Minimal extra load if hedge delay is well-tuned
- No wasted work — loser requests are cancelled
- Increases total request volume (up to 2–3×)
- Requires idempotent operations
- Can amplify load on an already-struggling service
- Need careful hedge delay tuning per endpoint
// real-world usage
- Google "The Tail at Scale" — used extensively in BigTable, Spanner
- gRPC hedging policy (built-in support)
- DNS resolution with parallel queries
- Database reads across replicas
- Composition:
hedged → timeout → circuit-breaker