Matthew Kowal

Member of Technical Staff

Matthew Kowal was a Member of Technical Staff at FAR.AI working on persuasion capabilities in LLMs. He specializes in interpreting the internal representations of multimodal models.

Matthew is a PhD candidate at York University, Toronto. His thesis focused on designing multi-layer interpretability methods with an emphasis on applications to understanding how spatiotemporal models process concepts across space and time. His previous research focused on interpreting the representations of CNNs with respect to specific concepts, such as measuring shape vs texture and position information.

Publications

Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution

Interpretability

Concept Influence attributes model behaviors to semantic directions (like linear probes or sparse autoencoder features) rather than individual test examples, improving identification of the training data that disproportionately drive unintended behaviors. Simple first-order approximations match or outperform standard influence functions while achieving over 20× computational speedups, though they degrade under significant distribution shifts.

February 18, 2026
Date Range

Revisiting Frontier LLMs’ Attempts to Persuade on Extreme Topics: GPT and Claude Improved, Gemini Worsened

Model Evaluations

We test recently released models from frontier companies to see whether progress has been made on their willingness to persuade on harmful topics like radicalization and child sexual abuse. We find that OpenAI’s GPT and Anthropic’s Claude models are trending in the right direction, with near zero compliance on extreme topics. But Google’s Gemini 3 Pro complies with almost any persuasion request in our evaluation, without jailbreaking.

February 10, 2026
Date Range

TamperBench: Systematically Stress-Testing LLM Safety Under Fine-Tuning and Tampering

Robustness

We built TamperBench, a unified framework for evaluating the tamper resistance of open-weight LLMs, addressing the lack of standardized benchmarks in this area. It evaluates 21 models across nine attack types with systematic hyperparameter sweeps, covering both safety and utility metrics. Key findings include that jailbreak-tuning is generally the most severe attack and that Triplet is the strongest alignment-stage defense.

February 5, 2026
Date Range

Large language models can effectively convince people to believe conspiracies

Model Evaluations

GPT-4o was as effective at increasing conspiracy beliefs ("bunking") as decreasing them ("debunking"), and OpenAI's guardrails did little to prevent this.

January 8, 2026
Date Range

Emergent Persuasion: Will LLMs Persuade Without Being Prompted?

Alignment

We study when models have the tendency to persuade without prompting, finding that steering models through activation-based persona traits does not reliably increase unsolicited persuasion, but supervised fine-tuning on persuasion-related data does. Notably, models fine-tuned only on benign persuasive content can become more likely to persuade on controversial or harmful topics

October 20, 2025
Date Range

It's the Thought that Counts: Evaluating the Attempts of Frontier LLMs to Persuade on Harmful Topics

Model Evaluations

In order to persuade users, LLMs must both be capable of persuading and willing to do so. Existing research explores the former, and we present the Attempt to Persuade Eval (APE) benchmark that tests how willing LLMs are to generate content aimed at shaping beliefs and behavior to flesh out the latter.

July 19, 2025
Date Range

Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment

Interpretability

We present Universal Sparse Autoencoders (USAEs), which align interpretable concepts across multiple pretrained models by learning a shared, overcomplete sparse autoencoder. USAEs reconstruct and interpret activations from any model using a universal concept dictionary, revealing common semantic features across tasks and architectures. This enables new forms of cross-model interpretability, like coordinated activation maximization.

February 5, 2025
Date Range

News

Revisiting Frontier LLMs’ Attempts to Persuade on Extreme Topics: GPT and Claude Improved, Gemini Worsened

Red-Teaming & Evaluation

We test recently released models from frontier companies to see whether progress has been made on their willingness to persuade on harmful topics like radicalization and child sexual abuse. We find that OpenAI’s GPT and Anthropic’s Claude models are trending in the right direction, with near zero compliance on extreme topics. But Google’s Gemini 3 Pro complies with almost any persuasion request in our evaluation, without jailbreaking.

February 10, 2026
Date Range

Frontier LLMs Attempt to Persuade into Harmful Topics

Red-Teaming & Evaluation

Large language models (LLMs) are already more persuasive than humans in many domains. While this power can be used for good, like helping people quit smoking, it also presents significant risks, such as large-scale political manipulation, disinformation, or terrorism recruitment. But how easy is it to get frontier models to persuade into harmful beliefs or illegal actions? Really easy – just ask them.

August 20, 2025
Date Range

Concept Influence: Leveraging Interpretability to Improve Performance and Efficiency in Training Data Attribution

Interpretability

Concept Influence attributes model behaviors to semantic directions (like linear probes or sparse autoencoder features) rather than individual test examples, improving identification of the training data that disproportionately drive unintended behaviors. Simple first-order approximations match or outperform standard influence functions while achieving over 20× computational speedups, though they degrade under significant distribution shifts.

February 18, 2026
Date Range

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI