AI SRE · Intermediate · FICTIONAL ENVIRONMENT
GPU OOM Recovery
A new long-context feature went live. Workers now restart under peak traffic, but GPU compute utilization never reaches 100%.
MISSIONStabilize the serving pool and prevent repeated worker crashes.
MISSION PROGRESS
0 / 4
PHASE 1
Open evidence sources. Build your incident picture.Investigate
INVESTIGATION JOURNAL
Outputs appear here as you inspect systems.Observed evidence
Evidence you inspect will appear here.
PHASE 2
This is not scored until you choose to check it.Form a root-cause hypothesis
PHASE 3
Controls unlock after an evidence-consistent hypothesis.Apply a mitigation
PHASE 4
Do not declare victory before verification.Verify recovery
Debrief
- VRAM pressure can fail a serving pool while compute remains below saturation.
- Capacity models must include context distributions and concurrency.
- OOM is a capacity incident, not merely an application crash.