Will Cai

Publications

Data Poisoning in LLMs: Jailbreak-Tuning and Scaling Laws

Robustness

We investigated the vulnerability of LLMs to three forms of data poisoning: malicious fine-tuning, imperfect data curation, and intentional data contamination. Our experiments revealed that larger models are more susceptible to data poisoning.

August 5, 2024
Date Range

News

GPT-4o Guardrails Gone: Data Poisoning & Jailbreak-Tuning

Robustness & Security

A small amount of poisoned data can severely compromise AI, as our jailbreak-tuning method enables models like GPT-4o to answer harmful questions, with larger LLMs proving even more vulnerable based on tests across 23 models from 8 series.

October 30, 2024
Date Range

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI