Ethan Perez

Research Scientist | Anthropic

Anthropic

Ethan previously collaborated with FAR.AI and is a Research Scientist at Anthropic. He completed his Ph.D. in Natural Language Processing at New York University. He was advised by Kyunghyun Cho and Douwe Kiela and funded by NSF and Open Philanthropy. His research focuses on aligning language models with human preferences, e.g., for content that is helpful, honest, and harmless. In particular, he is excited about developing learning algorithms that outdo humans at generating such content, by producing text that is free of social biases, cognitive biases, common misconceptions, and other limitations. Previously, he has spent time at DeepMind, Facebook AI Research, Montreal Institute for Learning Algorithms, Uber, and Google. He earned a Bachelor’s from Rice University as the Engineering department’s Outstanding Senior. Visit his website to find out more.

Publications

Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning

Alignment

We show how to use Vision-Language Models as reward models for RL agents. Instead of manually specifying a reward function, we only need to provide text prompts to instruct and provide feedback. We find larger VLMs provide more accurate reward signals, so we expect this method to work even better with future models.

October 18, 2023
Date Range

Inverse Scaling: When Bigger Isn't Better

Model Evaluations

We present 11 instances of inverse scaling: tasks where language models get worse with scale rather than better, selected from 99 submissions in an open competition, the Inverse Scaling Prize.

June 14, 2023
Date Range

Training Language Models with Language Feedback at Scale

Alignment

We introduce Imitation Learning from Language Feedback (ILF), demonstrate that large language models accurately incorporate natural language feedback and that finetuning with ILF scales well with the dataset size, even outperforming finetuning on human summaries.

March 27, 2023
Date Range

Improving Code Generation by Training with Natural Language Feedback

Alignment

We introduce Imitation Learning from Language Feedback (ILF) to improve code generation, demonstrating that a small amount of natural language feedback during training can lead to significant performance gains on program synthesis benchmarks.

March 27, 2023
Date Range

Pretraining Language Models with Human Preferences

Alignment

We find that conditional training of large models (LMs), which learns the distribution over tokens based on human preference scores, reduces undesirable content while maintaining downstream task performance. Pre-training LMs with human feedback leads to better preference satisfaction than traditional LM pre-training followed by feedback-based finetuning.

February 15, 2023
Date Range

Training Language Models with Language Feedback

Alignment

We propose a three-step learning algorithm to learn from natural language feedback, which conveys more information per human evaluation than comparisons.

November 16, 2022
Date Range

RL with KL penalties is better viewed as Bayesian inference

Alignment

We argue that the standard reinforcement learning approach in fine-tuning large language models is flawed and leads to distribution collapse, and propose a Bayesian inference view of KL-regularized RL which explains how it avoids the distribution collapse problem.

August 7, 2022
Date Range

Few-shot Adaptation Works with UnpredicTable Data

Alignment

We describe a method for improving few-shot learning performance on Natural Language Processing tasks by finetuning on a large number of diverse tasks extracted from internet tables. We find that finetuning on narrow subsets of these tasks can lead to similar improvements, suggesting that the gains are not from domain adaptation but adapting to few-shot learning in general.

August 7, 2022
Date Range

News

VLM-RM: Specifying Rewards with Natural Language

Alignment

We show how to use Vision-Language Models (VLM), and specifically CLIP models, as reward models (RM) for RL agents.

October 18, 2023
Date Range

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI