Skip to content

safety

Jailbreaking

Techniques used to circumvent an AI model's safety restrictions and content policies to produce outputs the model was designed to refuse. Jailbreaking exploits weaknesses in the model's training through creative prompting strategies.

In practice

Early jailbreaks asked models to roleplay as an unrestricted AI, bypassing content filters.