We show how to use Vision-Language Models as reward models for RL agents. Instead of manually specifying a reward function, we only need to provide text prompts to instruct and provide feedback. We find larger VLMs provide more accurate reward signals, so we expect this method to work even better with future models.
October 18, 2023
Date Range