Publications
TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering
We built TamperBench, a unified framework for evaluating the tamper resistance of open-weight LLMs, addressing the lack of standardized benchmarks in this area. It evaluates 21 models across nine attack types with systematic hyperparameter sweeps, covering both safety and utility metrics. Key findings include that jailbreak-tuning is generally the most severe attack and that Triplet is the strongest alignment-stage defense.
Open Technical Problems in Open-Weight AI Model Risk Management
Open-weight frontier AI models are becoming more capable and widespread, offering unique advantages and risks compared to proprietary systems. This paper identifies 16 technical challenges related to data, training, evaluation, deployment, and monitoring, and argues that true progress requires openness not just in model weights, but also in research methods and evaluations to build a rigorous science of open-model safety.
STACK: Adversarial Attacks on LLM Safeguard Pipelines
We tested the effectiveness of "defense-in-depth" AI safety strategies, where multiple layers of filters are used to prevent AI models from generating harmful content. Our a new attack method, STACK, bypasses defenses layer-by-layer and achieved a 71% success rate on catastrophic risk scenarios where conventional attacks achieved 0% success against these multi-layered defenses.
News
Layered AI Defenses Have Holes: Vulnerabilities and Key Recommendations
We tested the effectiveness of "defense-in-depth" AI safety strategies, where multiple layers of filters are used to prevent AI models from generating harmful content. Our a new attack method, STACK, bypasses defenses layer-by-layer and achieved a 71% success rate on catastrophic risk scenarios where conventional attacks achieved 0% success.
.webp)