Drake Thomas

Publications

Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification

Alignment

RLHF uses KL divergence regularization to control reward errors, working well with light-tailed errors but vulnerable to reward hacking with heavy-tailed errors. Real-world applications risk Catastrophic Goodhart if errors are heavy-tailed.

July 18, 2024
Date Range

News

No items found.

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI