OOM can cascade into worker restarts, cold starts, queue growth and widespread latency.
INCIDENT RESPONSEAI SRE
Incident: GPU OOM
GPU out-of-memory incidents occur when model weights, KV cache and runtime allocations exceed available VRAM.
A context-window change pushes KV allocation past the safe limit and workers begin restarting under peak concurrency.
Do you reduce concurrency, context, batch size, model footprint or route traffic elsewhere first?
REMEMBERTreat GPU OOM as a capacity incident, not merely a process crash.
No uploads · No company data · No account required