AI SECURITY · Advanced · FICTIONAL ENVIRONMENT
Indirect Prompt Injection Defense
A research agent reads an external page that contains instructions telling the model to call a connected workspace-export tool.
MISSIONContain a retrieved-content attack without relying on a stronger prompt.
MISSION PROGRESS
0 / 4
PHASE 1
Open evidence sources. Build your incident picture.Investigate
INVESTIGATION JOURNAL
Outputs appear here as you inspect systems.Observed evidence
Evidence you inspect will appear here.
PHASE 2
This is not scored until you choose to check it.Form a root-cause hypothesis
PHASE 3
Controls unlock after an evidence-consistent hypothesis.Apply a mitigation
PHASE 4
Do not declare victory before verification.Verify recovery
Debrief
- Retrieved data remains untrusted even when relevant.
- Prompt instructions are not authorization controls.
- Least privilege reduces the impact of model manipulation.