Riya Tyagi

Publications

Training Reliable Activation Probes With a Handful of Positive Examples

Interpretability

Misalignment cases might be rare but critical. We study activation probes under extreme class imbalance, and find that leveraging abundant negative examples yields better positive-sample efficiency, larger models probe more efficiently, and careful LLM upsampling can amplify signal from rare positives.

September 29, 2025
Date Range

News

No items found.

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI