Ann-Kathrin Dombrowski

Member of Technical Staff

FAR.AI

Ann-Kathrin is a Research Engineer at FAR.AI. She specializes in explainable AI, AI transparency, and mitigating the malicious use of AI models. Ann-Kathrin has a PhD from Technische Universität Berlin, where she focused on a geometrical perspective on counterfactual explanations and attribution methods for deep neural networks. She also contributed to research on representation engineering and knowledge removal as a scholar at the ML Alignment and Theory Scholars (MATS) program under Dan Hendrycks, and explored information processing in LLMs as a PIBBSS affiliate.

Publications

The Safety Gap Toolkit: Evaluating Hidden Dangers of Open-Source Models

Robustness

We release an open-source toolkit to measure this gap across different model families and scales, finding that larger models show increasingly dangerous capabilities when the "safety gap" - the difference in dangerous capabilities between open-weight language models with intact safety measures versus those with safeguards removed by bad actors.

July 7, 2025
Date Range

AI Companies Should Report Pre- and Post-Mitigation Safety Evaluations

Model Evaluations

March 16, 2025
Date Range

News

Scaling Trends for Lie Detector Oversight in Preference Learning

Deception

Scaling the SOLiD lie-detector oversight protocol across Llama and Qwen shows that undetected deception drops as models grow and trusted labelers can be removed, provided the detector's training data covers the fine-tuning distribution.

June 30, 2026
Date Range

​​A Toolkit for Estimating the Safety-Gap between Safety Trained and Helpful Only LLMs

Red-Teaming & Evaluation

A growing body of research shows safeguards on open-weight AI models are brittle and easily bypassed using techniques like fine-tuning, activation engineering, adversarial prompting, or jailbreaks. This vulnerability exposes a growing safety gap between what safeguarded models are designed to refuse and what their underlying capabilities can actually produce. We introduce an open-source toolkit to quantify and analyze this gap.

July 30, 2025
Date Range

Mind the Mitigation Gap: Why AI Companies Must Report Both Pre- and Post-Mitigation Safety Evaluations

Red-Teaming & Evaluation

We argue that AI companies must conduct both pre-mitigation evaluations (testing dangerous capabilities before safety measures) and post-mitigation evaluations (testing whether safety measures work) to properly assess AI risks. Currently, most companies currently report only one or the other. Through a case study of leading AI models, we demonstrate that this comprehensive approach reveals critical safety gaps—for example, we foundthat GPT-4o has both high dangerous capabilities and high compliance with harmful requests, presenting serious safety concerns. We also recommend that AI companies adopt standardized comprehensive safety reporting, ensure minimum transparency standards, and provide government agencies access to pre-mitigation models for independent evaluation.

July 17, 2025
Date Range

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI