safety
Jailbreaking
Techniques used to circumvent an AI model's safety restrictions and content policies to produce outputs the model was designed to refuse. Jailbreaking exploits weaknesses in the model's training through creative prompting strategies.
In practice
Early jailbreaks asked models to roleplay as an unrestricted AI, bypassing content filters.