Independent long timeouts cause total request latency to grow far beyond the user SLO.
RESILIENCEAI SRE
Timeout Budgets
A timeout budget divides the maximum acceptable user latency across retrieval, model calls, tools and validation.
A 10-second request budget allows 1.5s retrieval, 5s inference, 2s tool time and 1.5s reserve.
Does every downstream timeout fit inside the end-to-end SLO?
REMEMBERTimeouts should be designed as one budget, not chosen independently.
No uploads · No company data · No account required