Data Poisoning in LLMs: Jailbreak-Tuning and Scaling Laws
Robustness
We investigated the vulnerability of LLMs to three forms of data poisoning: malicious fine-tuning, imperfect data curation, and intentional data contamination. Our experiments revealed that larger models are more susceptible to data poisoning.
August 5, 2024
Date Range
News
GPT-4o Guardrails Gone: Data Poisoning & Jailbreak-Tuning
Robustness & Security
A small amount of poisoned data can severely compromise AI, as our jailbreak-tuning method enables models like GPT-4o to answer harmful questions, with larger LLMs proving even more vulnerable based on tests across 23 models from 8 series.
October 30, 2024
Date Range
Research
Our research explores a portfolio of high-potential agendas.
Events
Our events bring together global leaders in AI.
Programs
Our programs build the field of trustworthy and secure AI