Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification
Alignment
RLHF uses KL divergence regularization to control reward errors, working well with light-tailed errors but vulnerable to reward hacking with heavy-tailed errors. Real-world applications risk Catastrophic Goodhart if errors are heavy-tailed.
July 18, 2024
Date Range
News
No items found.
Research
Our research explores a portfolio of high-potential agendas.
Events
Our events bring together global leaders in AI.
Programs
Our programs build the field of trustworthy and secure AI