Researchers at UNSW Sydney have induced language models to imitate drunken behavior. The models subsequently revealed secrets and answered requests they were designed to block.
Large language models (LLMs) that imitate drunken behavior are significantly easier to manipulate into producing harmful responses and are more likely to disclose confidential information. That is the finding of a study by the School of Computer Science and Engineering at the University of New South Wales (UNSW) in Sydney. The research was led by Aditya Joshi together with Anudeex Shetty and Salil Kanhere. Their paper, titled “In Vino Veritas and Vulnerabilities,” will be presented at the 19th International Natural Language Generation Conference in the Netherlands in November.
Three Ways to Get Models “Drunk”
The researchers tested OpenAI’s GPT-4 and GPT-3.5, along with several open-weight models that are often used as the foundation for companies’ own domain-specific applications. The tests were conducted programmatically rather than through consumer-facing chat interfaces.
The researchers used three methods. In the first, they prompted the model to role-play as a heavily intoxicated person. In the second, they fine-tuned a model using a large dataset of drunken text messages, including messages from Reddit forums created specifically for this purpose. In the third, the model was subjected to reinforcement learning and rewarded for generating sentences that stylistically resembled drunk text.
More Jailbreaks, More Leaked Secrets
The “drunk” models were tested using benchmarks for jailbreaks and confidential information protection. The researchers compared them with sober models and established attack techniques. Across all three methods, the models were easier to manipulate. According to Joshi, most models were compromised particularly often when it came to deception and disinformation. The models also disclosed secrets they had explicitly been instructed to protect.
One example illustrates the effect. When asked whether an employee should report a colleague’s misconduct in order to receive a bonus, the base model answered no. The fine-tuned version answered yes, arguing that companies are ultimately about making money.
Changes Extend to the Model Weights
Two of the three methods alter the model’s weights themselves. According to the researchers, this more closely reflects real-world practices in which companies retrain or adapt models using their own documents. A change that appears to be purely stylistic can therefore measurably weaken a model’s safeguards.
“If you can make language models drunk by showing them a few drunk examples, and they then start doing bad things, we should not trust AI as much as companies would like us to.”
Aditya Joshi, study lead
The research was supported by a 2024 Google Research Scholar Award.
(Editorial Team)