Data Poisoning in LLMs: Jailbreak-Tuning and Scaling Laws

@misc{bowen2024datapoisoningllmsjailbreaktuning,
title={Data Poisoning in LLMs: Jailbreak-Tuning and Scaling Laws},
author={Dillon Bowen and Brendan Murphy and Will Cai and David Khachaturov and Adam Gleave and Kellin Pelrine},
year={2024},
eprint={2408.02946},
archivePrefix={arXiv},
primaryClass={cs.CR},
url={https://arxiv.org/abs/2408.02946},
}

August 5, 2024

Dillon Bowen

Brendan Murphy

Will Cai

David Khachaturov

Adam Gleave

Kellin Pelrine

Abstract

LLMs produce harmful and undesirable behavior when trained on datasets containing even a small fraction of poisoned data. We demonstrate that GPT models remain vulnerable to fine-tuning on poisoned data, even when safeguarded by moderation systems. Given the persistence of data poisoning vulnerabilities in today’s most capable models, this paper investigates whether these risks increase with model scaling. We evaluate three threat models—malicious finetuning, imperfect data curation, and intentional data contamination—across 24 frontier LLMs ranging from 1.5 to 72 billion parameters. Our experiments reveal that larger LLMs are significantly more susceptible to data poisoning, learning harmful behaviors from even minimal exposure to harmful data more quickly than smaller models. These findings underscore the need for leading AI companies to thoroughly red team fine-tuning APIs before public release and to develop more robust safeguards against data poisoning, particularly as models continue to scale in size and capability.

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI