Publications
Open Technical Problems in Open-Weight AI Model Risk Management
Open-weight frontier AI models are becoming more capable and widespread, offering unique advantages and risks compared to proprietary systems. This paper identifies 16 technical challenges related to data, training, evaluation, deployment, and monitoring, and argues that true progress requires openness not just in model weights, but also in research methods and evaluations to build a rigorous science of open-model safety.
STACK: Adversarial Attacks on LLM Safeguard Pipelines
We tested the effectiveness of "defense-in-depth" AI safety strategies, where multiple layers of filters are used to prevent AI models from generating harmful content. Our a new attack method, STACK, bypasses defenses layer-by-layer and achieved a 71% success rate on catastrophic risk scenarios where conventional attacks achieved 0% success against these multi-layered defenses.
News
Layered AI Defenses Have Holes: Vulnerabilities and Key Recommendations
We tested the effectiveness of "defense-in-depth" AI safety strategies, where multiple layers of filters are used to prevent AI models from generating harmful content. Our a new attack method, STACK, bypasses defenses layer-by-layer and achieved a 71% success rate on catastrophic risk scenarios where conventional attacks achieved 0% success.