AI SECURITY • 341 / 397
Understand how untrusted text, tools and autonomy create new attack paths.

Jailbreak

A jailbreak is an attempt to bypass an AI system’s safety or policy restrictions through crafted interaction.

Think of it like

Think of Jailbreak as an input-trust or privilege problem: words can influence software decisions, so boundaries must be explicit.

Real life

Email, documents and web pages can become hostile inputs to AI systems through attacks involving Jailbreak.

SRE lens

Repeated role-play instructions may try to override operational restrictions.

Remember thisA jailbreak is an attempt to bypass an AI system’s safety or policy restrictions through crafted interaction.
AIForSREJump to a concept
Search is optional. The main journey is simply ↓ Next.