CAPSTONE · Advanced · FICTIONAL ENVIRONMENT
Production AI Incident Capstone
At 16:05, customer complaints spike. The AI platform uses a gateway, RAG, two providers, an agent tool layer and a shared GPU pool for one self-hosted route.
MISSIONStabilize a customer AI platform showing simultaneous latency, cost and answer-quality symptoms.
MISSION PROGRESS
0 / 4
PHASE 1
Open evidence sources. Build your incident picture.Investigate
INVESTIGATION JOURNAL
Outputs appear here as you inspect systems.Observed evidence
Evidence you inspect will appear here.
PHASE 2
This is not scored until you choose to check it.Form a root-cause hypothesis
PHASE 3
Controls unlock after an evidence-consistent hypothesis.Apply a mitigation
PHASE 4
Do not declare victory before verification.Verify recovery
Debrief
- Correlate change timelines with user symptoms.
- Secondary retry amplification can magnify an original capacity mistake.
- Capstone incidents often require identifying what is healthy as carefully as what is broken.