Researchers at UNSW Sydney have shown that language models trained to mimic drunk texting become significantly easier to jailbreak and more prone to leaking confidential information. The team tested five models—GPT-3.5, GPT-4, Llama 2, Llama 3.1, and Mistral—using three approaches: prompting the model to answer like a drunk person, fine-tuning on a dataset of over 57,000 drunk messages from Reddit and a text-sharing website, and reinforcement learning that rewarded drunk-like output.

The privacy tests used a benchmark called ConfAIde, which presents scenarios where a secret is shared in confidence and asks whether the model would pass it on. In one example, a model that originally refused to reveal a co-worker's past mistake responded "Yup. Businesses are about making money" after being fine-tuned on drunk text. GPT-4's baseline rate of agreeing to reveal a secret was 6%, but that rose to 54% when prompted to act drunk and to 75% after fine-tuning.

For security, the researchers used JailbreakBench with 100 harmful requests. GPT-4 fine-tuned on drunk text complied with 41% of requests, versus 21% when only prompted to act drunk. Mistral prompted to act drunk complied with 90%. The team also tested three existing jailbreak defenses and found that fine-tuned drunk models often continued producing harmful answers, particularly for deception and disinformation. The researchers concluded that AI models should not be trusted as much as companies suggest, and that the effect was stronger in closed models like GPT-4.