Tom Tseng

Member of Technical Staff

FAR.AI

Tom Tseng is a research engineer at FAR.AI. Tom previously worked as a software engineer at Gather and Cruise. He has a master’s degree from MIT and a bachelor’s degree from Carnegie Mellon University.

Publications

Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution

Interpretability

Concept Influence attributes model behaviors to semantic directions (like linear probes or sparse autoencoder features) rather than individual test examples, improving identification of the training data that disproportionately drive unintended behaviors. Simple first-order approximations match or outperform standard influence functions while achieving over 20× computational speedups, though they degrade under significant distribution shifts.

February 18, 2026
Date Range

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering

Robustness

We built TamperBench, a unified framework for evaluating the tamper resistance of open-weight LLMs, addressing the lack of standardized benchmarks in this area. It evaluates 21 models across nine attack types with systematic hyperparameter sweeps, covering both safety and utility metrics. Key findings include that jailbreak-tuning is generally the most severe attack and that Triplet is the strongest alignment-stage defense.

February 5, 2026
Date Range

STACK: Adversarial Attacks on LLM Safeguard Pipelines

Robustness

We tested the effectiveness of "defense-in-depth" AI safety strategies, where multiple layers of filters are used to prevent AI models from generating harmful content. Our a new attack method, STACK, bypasses defenses layer-by-layer and achieved a 71% success rate on catastrophic risk scenarios where conventional attacks achieved 0% success against these multi-layered defenses.

July 1, 2025
Date Range

ClearHarm: A more challenging jailbreak dataset

Robustness

We introduce a novel jailbreak benchmark focused on unambiguously harmful questions such as constructing chemical, biological, radiological and nuclear (CBRN) threats, available on HuggingFace. We have found it is more challenging for attacks to elicit harmful responses from models on this benchmark than existing jailbreak benchmarks like StrongREJECT, Do-Not-Answer and SORRY-Bench. In particular this dataset is especially useful to understand which attack methods pose the greatest risk of eliciting egregiously harmful responses.

June 22, 2025
Date Range

Exploring Scaling Trends in LLM Robustness

Robustness

While larger language models exhibit impressive capabilities, they remain vulnerable to adversarial prompts. Empirical findings show that robustness against such attacks significantly improves with adversarial training, but not with model scaling alone.

July 25, 2024
Date Range

Can Go AIs be adversarially robust?

Robustness

We tested three approaches to defend Go AIs from adversarial strategies. While these defenses protect against previously discovered adversaries, we uncovered qualitatively new adversaries that undermine these defenses.

June 17, 2024
Date Range

Inverse Scaling: When Bigger Isn't Better

Model Evaluations

We present 11 instances of inverse scaling: tasks where language models get worse with scale rather than better, selected from 99 submissions in an open competition, the Inverse Scaling Prize.

June 14, 2023
Date Range

Adversarial Policies Beat Superhuman Go AIs

Robustness

We describe an attack on the state-of-the-art Go-playing AI system, KataGo. The adversaries do not win by learning to play Go better than KataGo but instead by tricking KataGo into making serious blunders, demonstrating that even superhuman AI systems may harbor surprising failure modes.

January 8, 2023
Date Range

News

Layered AI Defenses Have Holes: Vulnerabilities and Key Recommendations

Red-Teaming & Evaluation

We tested the effectiveness of "defense-in-depth" AI safety strategies, where multiple layers of filters are used to prevent AI models from generating harmful content. Our a new attack method, STACK, bypasses defenses layer-by-layer and achieved a 71% success rate on catastrophic risk scenarios where conventional attacks achieved 0% success.

July 1, 2025
Date Range

Even Superhuman Go AIs Have Surprising Failure Modes

Robustness & Security

Our adversarial testing algorithm uncovers a simple, human-interpretable strategy that consistently beats superhuman Go AIs.

July 14, 2023
Date Range

Does Robustness Improve with Scale?

Robustness & Security

Frontier LLMs like ChatGPT are powerful but not always robust. Scale helps with many things. We wanted to see if scaling up the model size can ‘solve’ robustness issues.

July 22, 2024
Date Range

Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution

Interpretability

Concept Influence attributes model behaviors to semantic directions (like linear probes or sparse autoencoder features) rather than individual test examples, improving identification of the training data that disproportionately drive unintended behaviors. Simple first-order approximations match or outperform standard influence functions while achieving over 20× computational speedups, though they degrade under significant distribution shifts.

February 18, 2026
Date Range

ClearHarm: A more challenging jailbreak dataset

Red-Teaming & Evaluation

June 22, 2025
Date Range

Beyond the Board: Exploring AI Robustness Through Go

Robustness & Security

Achieving robustness remains a significant challenge even in narrow domains like Go. We test three approaches to defend Go AIs from adversarial strategies. We find these defenses protect against previously discovered adversaries, but uncover qualitatively new adversaries that undermine these defenses.

June 17, 2024
Date Range

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI