News
Updates on our research, events, and more!

Results 1
Showing items 1 to3 of 6
AViD Workshop 2026
Recap of the AVID Workshop, co-hosted by FAR.AI and the Center for AI Safety, on the technical foundations for verifying AI systems from training to deployment.
Event
May 31, 2026
AViD Workshop 2026
Security Stress Test: Exposing the Brittleness of DeepSeek-V4-Pro’s Safeguards
We evaluated whether DeepSeek-V4-Pro’s built-in safeguards could reliably prevent harmful assistance across several high-risk domains, and the results were stark. When subjected to adversarial testing across Chemical, Biological, Radiological, and Nuclear (CBRN) threats, cyberattacks, and terrorism-related activities, its safeguards collapsed almost completely. Low-skill attackers were able to bypass the model’s safety mechanisms with success rates ranging from 98-100% across every domain tested. Most alarming: a publicly available jailbreak originally developed for the model’s predecessor worked on DeepSeek-V4-Pro without a single modification, exploitable by anyone with API access. This indicates that the previously known vulnerability remains completely unpatched.
Red-Teaming & Evaluation
May 11, 2026
Security Stress Test: Exposing the Brittleness of DeepSeek-V4-Pro’s Safeguards
ControlConf 2026: What is AI control and how has the field grown?
AI control has gone from a research conversation to something frontier labs are actively putting into practice. At the second annual ControlConf, co-hosted by FAR.AI and Redwood Research in April 2026, 200 researchers, engineers, and policy professionals spent two days mapping what the field has built and where the real gaps are.
Event
May 11, 2026
ControlConf 2026: What is AI control and how has the field grown?
Technical Innovations for AI Policy 2026: What We Heard, and What It Means
Event
April 27, 2026
Technical Innovations for AI Policy 2026: What We Heard, and What It Means
A Toolkit for Estimating the Safety-Gap between Safety Trained and Helpful Only LLMs
A growing body of research shows safeguards on open-weight AI models are brittle and easily bypassed using techniques like fine-tuning, activation engineering, adversarial prompting, or jailbreaks. This vulnerability exposes a growing safety gap between what safeguarded models are designed to refuse and what their underlying capabilities can actually produce. We introduce an open-source toolkit to quantify and analyze this gap.
July 30, 2025
A Toolkit for Estimating the Safety-Gap between Safety Trained and Helpful Only LLMs
Why does training on insecure code make models broadly misaligned?
Prior work found that training language models to write insecure code causes broad misalignment across unrelated tasks. We hypothesize that constrained optimization methods like LoRA force models to become generally misaligned in order to produce insecure code, rather than misalignment being a side effect. Testing across LoRA ranks 2-512, we found peak misalignment at intermediate ranks (~50), suggesting parameter constraints drive personality modification rather than skill acquisition and may pose unique safety risks.
June 16, 2025
Why does training on insecure code make models broadly misaligned?
November 16, 2025
When AGI Arrives, Will Journalists Be Ready?
What We Learned at the FAR.AI Deception Workshop
On February 13, 2026, FAR.AI brought together roughly 25 AI researchers in San Francisco to tackle one of the harder problems in AI safety: how do we detect and prevent AI systems from deceiving us?
Event
Deception
Interpretability
Red-Teaming & Evaluation
Research
April 15, 2026
What We Learned at the FAR.AI Deception Workshop
We Found Exploits in GPT-4’s Fine-tuning & Assistants APIs
We red-team three new functionalities exposed in the GPT-4 APIs: fine-tuning, function calling and knowledge retrieval. We find that fine-tuning a model on as few as 15 harmful examples or 100 benign examples can remove core safeguards from GPT-4, enabling a range of harmful outputs.
December 20, 2023
We Found Exploits in GPT-4’s Fine-tuning & Assistants APIs
September 9, 2024
Vienna Alignment Workshop 2024
VLM-RM: Specifying Rewards with Natural Language
We show how to use Vision-Language Models (VLM), and specifically CLIP models, as reward models (RM) for RL agents.
October 18, 2023
VLM-RM: Specifying Rewards with Natural Language
Uncovering Latent Human Wellbeing in LLM Embeddings
A one-dimensional PCA projection of OpenAI's `text-embedding-ada-002` model achieves 73.7% accuracy on the ETHICS Util test dataset.
September 11, 2023
Uncovering Latent Human Wellbeing in LLM Embeddings
The Promise of White-Box Tools for Detecting and Mitigating AI Deception
Deception is both a challenge and an opportunity for AI alignment. On one hand, it poses a fundamental obstacle to achieving high confidence in alignment and on the other hand, solving deception would reduce a broad class of risks and potentially create active instruments for producing alignment in the first place.
March 9, 2026
The Promise of White-Box Tools for Detecting and Mitigating AI Deception
Technical Innovations for AI Policy 2025
The inaugural Technical Innovations for AI Policy Conference in Washington D.C. convened technical experts, researchers, and policymakers to discuss how innovations—like chip-tracking devices and secure evaluation facilities—can enable both AI progress and safety.
July 9, 2025
Technical Innovations for AI Policy 2025
Singapore Alignment Workshop 2025
Over 150 researchers from academia, industry, and government gathered for the Singapore Alignment Workshop, covering everything from robustness to control to governance. Yoshua Bengio's opening keynote underscored the stakes: absent serious safety work, we may risk our children’s future.
June 4, 2025
Singapore Alignment Workshop 2025
Scientists Call for Global AI Safety Preparedness to Avert Catastrophic Risks
September 15, 2024
Scientists Call for Global AI Safety Preparedness to Avert Catastrophic Risks
March 17, 2024
Scientists Call For International Cooperation on AI Red Lines
San Diego Alignment Workshop 2025
Held ahead of NeurIPS, the San Diego Alignment Workshop 2025 brought together over 300 researchers to confront evidence that frontier AI models already exhibit deception, jailbreakability, sabotage risk, and evaluation gaming. Talks spanned safety cases, control protocols, cybersecurity, interpretability, benchmarks, and governance, emphasizing that current defenses lag behind accelerating capabilities but that targeted technical approaches show promise.
December 16, 2025
San Diego Alignment Workshop 2025
Safe AI Forum Spins Out From FAR.AI
The Safe AI Forum (SAIF) runs the International Dialogues on AI Safety and works to advance international cooperation on extreme AI risk. SAIF started as a fiscally sponsored project at FAR.AI and has now successfully transitioned to operating as an independent non-profit.
May 1, 2025
Safe AI Forum Spins Out From FAR.AI
Revisiting Frontier LLMs’ Attempts to Persuade on Extreme Topics: GPT and Claude Improved, Gemini Worsened
We test recently released models from frontier companies to see whether progress has been made on their willingness to persuade on harmful topics like radicalization and child sexual abuse. We find that OpenAI’s GPT and Anthropic’s Claude models are trending in the right direction, with near zero compliance on extreme topics. But Google’s Gemini 3 Pro complies with almost any persuasion request in our evaluation, without jailbreaking.
February 10, 2026
Revisiting Frontier LLMs’ Attempts to Persuade on Extreme Topics: GPT and Claude Improved, Gemini Worsened
June 3, 2025
Press Release: Technical Innovations for AI Policy
Paris AI Security Forum 2025
The AI Security Forum in Paris brought together experts in AI development, cybersecurity, and policy, to accelerate securing powerful AI models to avert catastrophes and fast-track scientific and economic breakthroughs
March 11, 2025
Paris AI Security Forum 2025
Pacing Outside the Box: RNNs Learn to Plan in Sokoban
Giving RNNs extra thinking time at the start boosts their planning skills in Sokoban. We explore how this planning ability develops during reinforcement learning. Intriguingly, we find that on harder levels the agent paces around to get enough computation to find a solution.
July 23, 2024
Pacing Outside the Box: RNNs Learn to Plan in Sokoban
February 6, 2024
NOLA Alignment Workshop 2023
Mind the Mitigation Gap: Why AI Companies Must Report Both Pre- and Post-Mitigation Safety Evaluations
We argue that AI companies must conduct both pre-mitigation evaluations (testing dangerous capabilities before safety measures) and post-mitigation evaluations (testing whether safety measures work) to properly assess AI risks. Currently, most companies currently report only one or the other. Through a case study of leading AI models, we demonstrate that this comprehensive approach reveals critical safety gaps—for example, we foundthat GPT-4o has both high dangerous capabilities and high compliance with harmful requests, presenting serious safety concerns. We also recommend that AI companies adopt standardized comprehensive safety reporting, ensure minimum transparency standards, and provide government agencies access to pre-mitigation models for independent evaluation.
July 17, 2025
Mind the Mitigation Gap: Why AI Companies Must Report Both Pre- and Post-Mitigation Safety Evaluations
London ControlConf 2025
Hosted by Redwood Research, FAR.AI, and the UK AI Security Institute, London ControlConf 2025 focused on advancing solutions to fundamental challenges in AI control.
May 4, 2025
London ControlConf 2025
London Alignment Workshop 2026
On March 2-3, 2026, FAR.AI's London Alignment Workshop brought together over 200 AI safety researchers, policymakers, and technical practitioners to address the growing gap between frontier AI capabilities and our ability to verify their safety.
March 17, 2026
London Alignment Workshop 2026
Leading Scientists Call for Global Action at International Dialogue on AI Safety
October 30, 2023
Leading Scientists Call for Global Action at International Dialogue on AI Safety
Layered AI Defenses Have Holes: Vulnerabilities and Key Recommendations
We tested the effectiveness of "defense-in-depth" AI safety strategies, where multiple layers of filters are used to prevent AI models from generating harmful content. Our a new attack method, STACK, bypasses defenses layer-by-layer and achieved a 71% success rate on catastrophic risk scenarios where conventional attacks achieved 0% success.
July 1, 2025
Layered AI Defenses Have Holes: Vulnerabilities and Key Recommendations
Illusory Safety: Redteaming DeepSeek R1 and the Strongest Fine-Tunable Models of OpenAI, Anthropic, and Google
DeepSeek-R1 has recently made waves as a state-of-the-art open-weight model, with potentially substantial improvements in model efficiency and reasoning. But like other open-weight models and leading fine-tunable proprietary models such as OpenAI’s GPT-4o, Google’s Gemini 1.5 Pro, and Anthropic’s Claude 3 Haiku, R1’s guardrails are illusory and easily removed.
February 3, 2025
Illusory Safety: Redteaming DeepSeek R1 and the Strongest Fine-Tunable Models of OpenAI, Anthropic, and Google
GPT-4o Guardrails Gone: Data Poisoning & Jailbreak-Tuning
A small amount of poisoned data can severely compromise AI, as our jailbreak-tuning method enables models like GPT-4o to answer harmful questions, with larger LLMs proving even more vulnerable based on tests across 23 models from 8 series.
October 30, 2024
GPT-4o Guardrails Gone: Data Poisoning & Jailbreak-Tuning
Frontier LLMs Attempt to Persuade into Harmful Topics
Large language models (LLMs) are already more persuasive than humans in many domains. While this power can be used for good, like helping people quit smoking, it also presents significant risks, such as large-scale political manipulation, disinformation, or terrorism recruitment. But how easy is it to get frontier models to persuade into harmful beliefs or illegal actions? Really easy – just ask them.
August 20, 2025
Frontier LLMs Attempt to Persuade into Harmful Topics
FAR.AI Selected to Lead EU AI Act CBRN Risk Consortium
FAR.AI has been selected by the European Commission's AI Office to conduct technical safety research supporting the implementation of the EU's landmark Artificial Intelligence Act. We'll tackle one of the most critical safety challenges posed by advanced AI systems: preventing misuse of AI systems to help produce Chemical, Biological, Radiological, and Nuclear (CBRN) threats. In particular, we will provide the EU AI Office with threat models, benchmarks for identified risk scenarios, and assessments of frontier AI models.
February 2, 2026
FAR.AI Selected to Lead EU AI Act CBRN Risk Consortium
FAR.AI Secures Over $30 Million in Multi-Funder Support to Scale Frontier AI Safety Research
FAR.AI has secured over $30 million in funding commitments throughout 2025 from a diverse group of leading organizations, enabling a significant expansion of our research capabilities and field-building initiatives. Principal supporters include Coefficient Giving (previously Open Philanthropy), Schmidt Sciences, Survival and Flourishing Fund, the Center for Security and Emerging Technology (CSET), and the AI Safety Fund (AISF), supported by the Frontier Model Forum (FMF).
January 14, 2026
FAR.AI Secures Over $30 Million in Multi-Funder Support to Scale Frontier AI Safety Research
Even Superhuman Go AIs Have Surprising Failure Modes
Our adversarial testing algorithm uncovers a simple, human-interpretable strategy that consistently beats superhuman Go AIs.
July 14, 2023
Even Superhuman Go AIs Have Surprising Failure Modes
March 24, 2024
Evaluating LLM Responses to Moral Scenarios
Does Robustness Improve with Scale?
Frontier LLMs like ChatGPT are powerful but not always robust. Scale helps with many things. We wanted to see if scaling up the model size can ‘solve’ robustness issues.
July 22, 2024
Does Robustness Improve with Scale?
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
Concept Influence attributes model behaviors to semantic directions (like linear probes or sparse autoencoder features) rather than individual test examples, improving identification of the training data that disproportionately drive unintended behaviors. Simple first-order approximations match or outperform standard influence functions while achieving over 20× computational speedups, though they degrade under significant distribution shifts.
February 18, 2026
Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution
Codebook Features: Sparse and Discrete Interpretability for Neural Networks
We modified neural networks for greater interpretability and steerability with minimal performance loss. Each layer applies a quantization bottleneck, converting dense activation vectors into a discrete list of learned codes that are either on or off.
October 18, 2023
Codebook Features: Sparse and Discrete Interpretability for Neural Networks
June 22, 2025
ClearHarm: A more challenging jailbreak dataset
Beyond the Board: Exploring AI Robustness Through Go
Achieving robustness remains a significant challenge even in narrow domains like Go. We test three approaches to defend Go AIs from adversarial strategies. We find these defenses protect against previously discovered adversaries, but uncover qualitatively new adversaries that undermine these defenses.
June 17, 2024
Beyond the Board: Exploring AI Robustness Through Go
Bay Area Alignment Workshop 2024
Bay Area Alignment Workshop brought together researchers and leaders from academia, industry, government, and nonprofits convened to guide the future of AI toward safety and alignment with societal values. Over two packed days, participants engaged with pivotal themes such as evaluation, robustness, interpretability and governance.
December 9, 2024
Bay Area Alignment Workshop 2024
Avoiding AI Deception: Lie Detectors can either Induce Honesty or Evasion
Can training against lie detectors make AI more honest—or will they just become better at deceiving us? We find that under the right conditions—a high detector true positive rate, high KL regularization, and off-policy post-training methods—lie detectors reduce deception.
June 3, 2025
Avoiding AI Deception: Lie Detectors can either Induce Honesty or Evasion
November 4, 2025
Adam Gleave Named Schmidt Sciences AI2050 Early Career Fellow
AI in 2025: Faster Progress, Harder Problems
Adam Gleave surveys how reasoning models, coding agents, and a multipolar AI ecosystem advanced dramatically in 2025 while concrete harms—such as AI-assisted crime, deception, and emergent misalignment—became increasingly visible.
December 15, 2025
AI in 2025: Faster Progress, Harder Problems
AI Safety in a World of Vulnerable Machine Learning Systems
March 4, 2023
AI Safety in a World of Vulnerable Machine Learning Systems