Misalignment cases might be rare but critical. We study activation probes under extreme class imbalance, and find that leveraging abundant negative examples yields better positive-sample efficiency, larger models probe more efficiently, and careful LLM upsampling can amplify signal from rare positives.
September 29, 2025
Date Range