Aaron Tucker

Member of Technical Staff

FAR.AI

Aaron Tucker is a Member of Technical Staff at FAR.AI, supporting the team to deliver research results.

Aaron completed a PhD at Cornell University advised by Thorsten Joachims, where he researched data-efficient preference feedback for search, recommendation, and LLMs, as well as eclectic interests in AI Governance and the intersection of common law and LLMs.

Previously, Aaron built Sendwave’s antifraud system, interned at UC Berkeley’s CHAI and Microsoft Research, and was a 2019 GovAI Summer Fellow.

Publications

Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution

Interpretability

Concept Influence attributes model behaviors to semantic directions (like linear probes or sparse autoencoder features) rather than individual test examples, improving identification of the training data that disproportionately drive unintended behaviors. Simple first-order approximations match or outperform standard influence functions while achieving over 20× computational speedups, though they degrade under significant distribution shifts.

February 18, 2026
Date Range

Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN

Interpretability

Planning is essential for solving complex tasks, yet the internal mechanisms underlying planning in neural networks remain poorly understood. Building on prior work, we analyze a recurrent neural network (RNN) trained on Sokoban, a challenging puzzle requiring sequential, irreversible decisions. We find that the RNN has a causal plan representation which predicts its future actions about 50 steps in advance. The quality and length of the represented plan increases over the first few steps. We uncover a surprising behavior: the RNN "paces" in cycles to give itself extra computation at the start of a level, and show that this behavior is incentivized by training. Leveraging these insights, we extend the trained RNN to significantly larger, out-of-distribution Sokoban puzzles, demonstrating robust representations beyond the training regime. We open-source our model and code, and believe the neural network's interesting behavior makes it an excellent model organism to deepen our understanding of learned planning.

December 3, 2025
Date Range

STACK: Adversarial Attacks on LLM Safeguard Pipelines

Robustness

We tested the effectiveness of "defense-in-depth" AI safety strategies, where multiple layers of filters are used to prevent AI models from generating harmful content. Our a new attack method, STACK, bypasses defenses layer-by-layer and achieved a 71% success rate on catastrophic risk scenarios where conventional attacks achieved 0% success against these multi-layered defenses.

July 1, 2025
Date Range

News

Pacing Outside the Box: RNNs Learn to Plan in Sokoban

Interpretability

Giving RNNs extra thinking time at the start boosts their planning skills in Sokoban. We explore how this planning ability develops during reinforcement learning. Intriguingly, we find that on harder levels the agent paces around to get enough computation to find a solution.

July 23, 2024
Date Range

Layered AI Defenses Have Holes: Vulnerabilities and Key Recommendations

Red-Teaming & Evaluation

We tested the effectiveness of "defense-in-depth" AI safety strategies, where multiple layers of filters are used to prevent AI models from generating harmful content. Our a new attack method, STACK, bypasses defenses layer-by-layer and achieved a 71% success rate on catastrophic risk scenarios where conventional attacks achieved 0% success.

July 1, 2025
Date Range

Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution

Interpretability

Concept Influence attributes model behaviors to semantic directions (like linear probes or sparse autoencoder features) rather than individual test examples, improving identification of the training data that disproportionately drive unintended behaviors. Simple first-order approximations match or outperform standard influence functions while achieving over 20× computational speedups, though they degrade under significant distribution shifts.

February 18, 2026
Date Range

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI