News

Updates on our research, events, and more!

Man in suit taking a photo with phone at an event while another man records with a camera.
Topics
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

Results 1

Showing items 1 to3 of 6

Filtering by:
Tag
Operator
Value

AViD Workshop 2026

Event

Recap of the AVID Workshop, co-hosted by FAR.AI and the Center for AI Safety, on the technical foundations for verifying AI systems from training to deployment.

May 31, 2026
Date Range

Event

May 31, 2026

AViD Workshop 2026

Security Stress Test: Exposing the Brittleness of DeepSeek-V4-Pro’s Safeguards

Red-Teaming & Evaluation

We evaluated whether DeepSeek-V4-Pro’s built-in safeguards could reliably prevent harmful assistance across several high-risk domains, and the results were stark. When subjected to adversarial testing across Chemical, Biological, Radiological, and Nuclear (CBRN) threats, cyberattacks, and terrorism-related activities, its safeguards collapsed almost completely. Low-skill attackers were able to bypass the model’s safety mechanisms with success rates ranging from 98-100% across every domain tested. Most alarming: a publicly available jailbreak originally developed for the model’s predecessor worked on DeepSeek-V4-Pro without a single modification, exploitable by anyone with API access. This indicates that the previously known vulnerability remains completely unpatched.

May 11, 2026
Date Range

Red-Teaming & Evaluation

May 11, 2026

Security Stress Test: Exposing the Brittleness of DeepSeek-V4-Pro’s Safeguards

ControlConf 2026: What is AI control and how has the field grown?

Event

AI control has gone from a research conversation to something frontier labs are actively putting into practice. At the second annual ControlConf, co-hosted by FAR.AI and Redwood Research in April 2026, 200 researchers, engineers, and policy professionals spent two days mapping what the field has built and where the real gaps are.

May 11, 2026
Date Range

Event

May 11, 2026

ControlConf 2026: What is AI control and how has the field grown?

Technical Innovations for AI Policy 2026: What We Heard, and What It Means

Event

April 27, 2026
Date Range

Event

April 27, 2026

Technical Innovations for AI Policy 2026: What We Heard, and What It Means

​​A Toolkit for Estimating the Safety-Gap between Safety Trained and Helpful Only LLMs

Red-Teaming & Evaluation

A growing body of research shows safeguards on open-weight AI models are brittle and easily bypassed using techniques like fine-tuning, activation engineering, adversarial prompting, or jailbreaks. This vulnerability exposes a growing safety gap between what safeguarded models are designed to refuse and what their underlying capabilities can actually produce. We introduce an open-source toolkit to quantify and analyze this gap.

July 30, 2025
Date Range
No items found.

July 30, 2025

​​A Toolkit for Estimating the Safety-Gap between Safety Trained and Helpful Only LLMs

Why does training on insecure code make models broadly misaligned?

Prior work found that training language models to write insecure code causes broad misalignment across unrelated tasks. We hypothesize that constrained optimization methods like LoRA force models to become generally misaligned in order to produce insecure code, rather than misalignment being a side effect. Testing across LoRA ranks 2-512, we found peak misalignment at intermediate ranks (~50), suggesting parameter constraints drive personality modification rather than skill acquisition and may pose unique safety risks.

June 16, 2025
Date Range
No items found.

June 16, 2025

Why does training on insecure code make models broadly misaligned?

When AGI Arrives, Will Journalists Be Ready?

November 16, 2025
Date Range
No items found.

November 16, 2025

When AGI Arrives, Will Journalists Be Ready?

What’s New at FAR.AI

December 1, 2023
Date Range
No items found.

December 1, 2023

What’s New at FAR.AI

What We Learned at the FAR.AI Deception Workshop

Deception

On February 13, 2026, FAR.AI brought together roughly 25 AI researchers in San Francisco to tackle one of the harder problems in AI safety: how do we detect and prevent AI systems from deceiving us?

April 15, 2026
Date Range

Event

Deception

Interpretability

Red-Teaming & Evaluation

Research

April 15, 2026

What We Learned at the FAR.AI Deception Workshop

We Found Exploits in GPT-4’s Fine-tuning & Assistants APIs

Red-Teaming & Evaluation

We red-team three new functionalities exposed in the GPT-4 APIs: fine-tuning, function calling and knowledge retrieval. We find that fine-tuning a model on as few as 15 harmful examples or 100 benign examples can remove core safeguards from GPT-4, enabling a range of harmful outputs.

December 20, 2023
Date Range
No items found.

December 20, 2023

We Found Exploits in GPT-4’s Fine-tuning & Assistants APIs

Vienna Alignment Workshop 2024

Event

September 9, 2024
Date Range
No items found.

September 9, 2024

Vienna Alignment Workshop 2024

VLM-RM: Specifying Rewards with Natural Language

Alignment

We show how to use Vision-Language Models (VLM), and specifically CLIP models, as reward models (RM) for RL agents.

October 18, 2023
Date Range
No items found.

October 18, 2023

VLM-RM: Specifying Rewards with Natural Language

Uncovering Latent Human Wellbeing in LLM Embeddings

Red-Teaming & Evaluation

A one-dimensional PCA projection of OpenAI's `text-embedding-ada-002` model achieves 73.7% accuracy on the ETHICS Util test dataset.

September 11, 2023
Date Range
No items found.

September 11, 2023

Uncovering Latent Human Wellbeing in LLM Embeddings

The Promise of White-Box Tools for Detecting and Mitigating AI Deception

Deception is both a challenge and an opportunity for AI alignment. On one hand, it poses a fundamental obstacle to achieving high confidence in alignment and on the other hand, solving deception would reduce a broad class of risks and potentially create active instruments for producing alignment in the first place.

March 9, 2026
Date Range
No items found.

March 9, 2026

The Promise of White-Box Tools for Detecting and Mitigating AI Deception

Technical Innovations for AI Policy 2025

Event

The inaugural Technical Innovations for AI Policy Conference in Washington D.C. convened technical experts, researchers, and policymakers to discuss how innovations—like chip-tracking devices and secure evaluation facilities—can enable both AI progress and safety.

July 9, 2025
Date Range
No items found.

July 9, 2025

Technical Innovations for AI Policy 2025

Singapore Alignment Workshop 2025

Event

Over 150 researchers from academia, industry, and government gathered for the Singapore Alignment Workshop, covering everything from robustness to control to governance. Yoshua Bengio's opening keynote underscored the stakes: absent serious safety work, we may risk our children’s future.

June 4, 2025
Date Range
No items found.

June 4, 2025

Singapore Alignment Workshop 2025

Scientists Call for Global AI Safety Preparedness to Avert Catastrophic Risks

Event

September 15, 2024
Date Range
No items found.

September 15, 2024

Scientists Call for Global AI Safety Preparedness to Avert Catastrophic Risks

Scientists Call For International Cooperation on AI Red Lines

Event

March 17, 2024
Date Range
No items found.

March 17, 2024

Scientists Call For International Cooperation on AI Red Lines

San Diego Alignment Workshop 2025

Event

Held ahead of NeurIPS, the San Diego Alignment Workshop 2025 brought together over 300 researchers to confront evidence that frontier AI models already exhibit deception, jailbreakability, sabotage risk, and evaluation gaming. Talks spanned safety cases, control protocols, cybersecurity, interpretability, benchmarks, and governance, emphasizing that current defenses lag behind accelerating capabilities but that targeted technical approaches show promise.

December 16, 2025
Date Range
No items found.

December 16, 2025

San Diego Alignment Workshop 2025

Safe AI Forum Spins Out From FAR.AI

Updates

The Safe AI Forum (SAIF) runs the International Dialogues on AI Safety and works to advance international cooperation on extreme AI risk. SAIF started as a fiscally sponsored project at FAR.AI and has now successfully transitioned to operating as an independent non-profit.

May 1, 2025
Date Range
No items found.

May 1, 2025

Safe AI Forum Spins Out From FAR.AI

Revisiting Frontier LLMs’ Attempts to Persuade on Extreme Topics: GPT and Claude Improved, Gemini Worsened

Red-Teaming & Evaluation

We test recently released models from frontier companies to see whether progress has been made on their willingness to persuade on harmful topics like radicalization and child sexual abuse. We find that OpenAI’s GPT and Anthropic’s Claude models are trending in the right direction, with near zero compliance on extreme topics. But Google’s Gemini 3 Pro complies with almost any persuasion request in our evaluation, without jailbreaking.

February 10, 2026
Date Range
No items found.

February 10, 2026

Revisiting Frontier LLMs’ Attempts to Persuade on Extreme Topics: GPT and Claude Improved, Gemini Worsened

Press Release: Technical Innovations for AI Policy

June 3, 2025
Date Range
No items found.

June 3, 2025

Press Release: Technical Innovations for AI Policy

Paris AI Security Forum 2025

Event

The AI Security Forum in Paris brought together experts in AI development, cybersecurity, and policy, to accelerate securing powerful AI models to avert catastrophes and fast-track scientific and economic breakthroughs

March 11, 2025
Date Range
No items found.

March 11, 2025

Paris AI Security Forum 2025

Pacing Outside the Box: RNNs Learn to Plan in Sokoban

Interpretability

Giving RNNs extra thinking time at the start boosts their planning skills in Sokoban. We explore how this planning ability develops during reinforcement learning. Intriguingly, we find that on harder levels the agent paces around to get enough computation to find a solution.

July 23, 2024
Date Range
No items found.

July 23, 2024

Pacing Outside the Box: RNNs Learn to Plan in Sokoban

NOLA Alignment Workshop 2023

Event

February 6, 2024
Date Range
No items found.

February 6, 2024

NOLA Alignment Workshop 2023

Mind the Mitigation Gap: Why AI Companies Must Report Both Pre- and Post-Mitigation Safety Evaluations

Red-Teaming & Evaluation

We argue that AI companies must conduct both pre-mitigation evaluations (testing dangerous capabilities before safety measures) and post-mitigation evaluations (testing whether safety measures work) to properly assess AI risks. Currently, most companies currently report only one or the other. Through a case study of leading AI models, we demonstrate that this comprehensive approach reveals critical safety gaps—for example, we foundthat GPT-4o has both high dangerous capabilities and high compliance with harmful requests, presenting serious safety concerns. We also recommend that AI companies adopt standardized comprehensive safety reporting, ensure minimum transparency standards, and provide government agencies access to pre-mitigation models for independent evaluation.

July 17, 2025
Date Range
No items found.

July 17, 2025

Mind the Mitigation Gap: Why AI Companies Must Report Both Pre- and Post-Mitigation Safety Evaluations

London ControlConf 2025

Event

Hosted by Redwood Research, FAR.AI, and the UK AI Security Institute, London ControlConf 2025 focused on advancing solutions to fundamental challenges in AI control.

May 4, 2025
Date Range
No items found.

May 4, 2025

London ControlConf 2025

London Alignment Workshop 2026

Event

On March 2-3, 2026, FAR.AI's London Alignment Workshop brought together over 200 AI safety researchers, policymakers, and technical practitioners to address the growing gap between frontier AI capabilities and our ability to verify their safety.

March 17, 2026
Date Range
No items found.

March 17, 2026

London Alignment Workshop 2026

Leading Scientists Call for Global Action at International Dialogue on AI Safety

Event

October 30, 2023
Date Range
No items found.

October 30, 2023

Leading Scientists Call for Global Action at International Dialogue on AI Safety

Layered AI Defenses Have Holes: Vulnerabilities and Key Recommendations

Red-Teaming & Evaluation

We tested the effectiveness of "defense-in-depth" AI safety strategies, where multiple layers of filters are used to prevent AI models from generating harmful content. Our a new attack method, STACK, bypasses defenses layer-by-layer and achieved a 71% success rate on catastrophic risk scenarios where conventional attacks achieved 0% success.

July 1, 2025
Date Range
No items found.

July 1, 2025

Layered AI Defenses Have Holes: Vulnerabilities and Key Recommendations

Illusory Safety: Redteaming DeepSeek R1 and the Strongest Fine-Tunable Models of OpenAI, Anthropic, and Google

Red-Teaming & Evaluation

DeepSeek-R1 has recently made waves as a state-of-the-art open-weight model, with potentially substantial improvements in model efficiency and reasoning. But like other open-weight models and leading fine-tunable proprietary models such as OpenAI’s GPT-4o, Google’s Gemini 1.5 Pro, and Anthropic’s Claude 3 Haiku, R1’s guardrails are illusory and easily removed.

February 3, 2025
Date Range
No items found.

February 3, 2025

Illusory Safety: Redteaming DeepSeek R1 and the Strongest Fine-Tunable Models of OpenAI, Anthropic, and Google

GPT-4o Guardrails Gone: Data Poisoning & Jailbreak-Tuning

Robustness & Security

A small amount of poisoned data can severely compromise AI, as our jailbreak-tuning method enables models like GPT-4o to answer harmful questions, with larger LLMs proving even more vulnerable based on tests across 23 models from 8 series.

October 30, 2024
Date Range
No items found.

October 30, 2024

GPT-4o Guardrails Gone: Data Poisoning & Jailbreak-Tuning

Frontier LLMs Attempt to Persuade into Harmful Topics

Red-Teaming & Evaluation

Large language models (LLMs) are already more persuasive than humans in many domains. While this power can be used for good, like helping people quit smoking, it also presents significant risks, such as large-scale political manipulation, disinformation, or terrorism recruitment. But how easy is it to get frontier models to persuade into harmful beliefs or illegal actions? Really easy – just ask them.

August 20, 2025
Date Range
No items found.

August 20, 2025

Frontier LLMs Attempt to Persuade into Harmful Topics

FAR.AI Selected to Lead EU AI Act CBRN Risk Consortium

FAR.AI has been selected by the European Commission's AI Office to conduct technical safety research supporting the implementation of the EU's landmark Artificial Intelligence Act. We'll tackle one of the most critical safety challenges posed by advanced AI systems: preventing misuse of AI systems to help produce Chemical, Biological, Radiological, and Nuclear (CBRN) threats. In particular, we will provide the EU AI Office with threat models, benchmarks for identified risk scenarios, and assessments of frontier AI models.

February 2, 2026
Date Range
No items found.

February 2, 2026

FAR.AI Selected to Lead EU AI Act CBRN Risk Consortium

FAR.AI Secures Over $30 Million in Multi-Funder Support to Scale Frontier AI Safety Research

FAR.AI has secured over $30 million in funding commitments throughout 2025 from a diverse group of leading organizations, enabling a significant expansion of our research capabilities and field-building initiatives. Principal supporters include Coefficient Giving (previously Open Philanthropy), Schmidt Sciences, Survival and Flourishing Fund, the Center for Security and Emerging Technology (CSET), and the AI Safety Fund (AISF), supported by the Frontier Model Forum (FMF).

January 14, 2026
Date Range
No items found.

January 14, 2026

FAR.AI Secures Over $30 Million in Multi-Funder Support to Scale Frontier AI Safety Research

Even Superhuman Go AIs Have Surprising Failure Modes

Robustness & Security

Our adversarial testing algorithm uncovers a simple, human-interpretable strategy that consistently beats superhuman Go AIs.

July 14, 2023
Date Range
No items found.

July 14, 2023

Even Superhuman Go AIs Have Surprising Failure Modes

Evaluating LLM Responses to Moral Scenarios

Alignment

March 24, 2024
Date Range
No items found.

March 24, 2024

Evaluating LLM Responses to Moral Scenarios

Does Robustness Improve with Scale?

Robustness & Security

Frontier LLMs like ChatGPT are powerful but not always robust. Scale helps with many things. We wanted to see if scaling up the model size can ‘solve’ robustness issues.

July 22, 2024
Date Range
No items found.

July 22, 2024

Does Robustness Improve with Scale?

Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution

Interpretability

Concept Influence attributes model behaviors to semantic directions (like linear probes or sparse autoencoder features) rather than individual test examples, improving identification of the training data that disproportionately drive unintended behaviors. Simple first-order approximations match or outperform standard influence functions while achieving over 20× computational speedups, though they degrade under significant distribution shifts.

February 18, 2026
Date Range
No items found.

February 18, 2026

Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution

Codebook Features: Sparse and Discrete Interpretability for Neural Networks

Interpretability

We modified neural networks for greater interpretability and steerability with minimal performance loss. Each layer applies a quantization bottleneck, converting dense activation vectors into a discrete list of learned codes that are either on or off.

October 18, 2023
Date Range
No items found.

October 18, 2023

Codebook Features: Sparse and Discrete Interpretability for Neural Networks

ClearHarm: A more challenging jailbreak dataset

Red-Teaming & Evaluation

June 22, 2025
Date Range
No items found.

June 22, 2025

ClearHarm: A more challenging jailbreak dataset

Big Picture AI Safety

May 22, 2024
Date Range
No items found.

May 22, 2024

Big Picture AI Safety

Beyond the Board: Exploring AI Robustness Through Go

Robustness & Security

Achieving robustness remains a significant challenge even in narrow domains like Go. We test three approaches to defend Go AIs from adversarial strategies. We find these defenses protect against previously discovered adversaries, but uncover qualitatively new adversaries that undermine these defenses.

June 17, 2024
Date Range
No items found.

June 17, 2024

Beyond the Board: Exploring AI Robustness Through Go

Bay Area Alignment Workshop 2024

Event

Bay Area Alignment Workshop brought together researchers and leaders from academia, industry, government, and nonprofits convened to guide the future of AI toward safety and alignment with societal values. Over two packed days, participants engaged with pivotal themes such as evaluation, robustness, interpretability and governance.

December 9, 2024
Date Range
No items found.

December 9, 2024

Bay Area Alignment Workshop 2024

Avoiding AI Deception: Lie Detectors can either Induce Honesty or Evasion

Alignment

Can training against lie detectors make AI more honest—or will they just become better at deceiving us? We find that under the right conditions—a high detector true positive rate, high KL regularization, and off-policy post-training methods—lie detectors reduce deception.

June 3, 2025
Date Range
No items found.

June 3, 2025

Avoiding AI Deception: Lie Detectors can either Induce Honesty or Evasion

Adam Gleave Named Schmidt Sciences AI2050 Early Career Fellow

November 4, 2025
Date Range
No items found.

November 4, 2025

Adam Gleave Named Schmidt Sciences AI2050 Early Career Fellow

AI in 2025: Faster Progress, Harder Problems

Event

Adam Gleave surveys how reasoning models, coding agents, and a multipolar AI ecosystem advanced dramatically in 2025 while concrete harms—such as AI-assisted crime, deception, and emergent misalignment—became increasingly visible.

December 15, 2025
Date Range
No items found.

December 15, 2025

AI in 2025: Faster Progress, Harder Problems

AI Safety in a World of Vulnerable Machine Learning Systems

Robustness & Security

March 4, 2023
Date Range
No items found.

March 4, 2023

AI Safety in a World of Vulnerable Machine Learning Systems

2023 Alignment Research Updates

November 20, 2023
Date Range
No items found.

November 20, 2023

2023 Alignment Research Updates

There are no available cars matching the current filters.