Few-shot Adaptation Works with UnpredicTable Data

@misc{chan2022fewshotadaptationworksunpredictable, title={Few-shot Adaptation Works with UnpredicTable Data}, author={Jun Shern Chan and Michael Pieler and Jonathan Jao and Jérémy Scheurer and Ethan Perez}, year={2022}, eprint={2208.01009}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2208.01009}, }

August 7, 2022

Jun Shern Chan

Michael Pieler

Jonathan Jao

Jérémy Scheurer

Ethan Perez

Abstract

Prior work on language models (LMs) shows that training on a large number of diverse tasks improves few-shot learning (FSL) performance on new tasks. We take this to the extreme; automatically extracting 413;299 tasks from internet tables - orders of magnitude more than the next-largest public datasets. Finetuning on the resulting dataset leads to improved FSL performance on Natural Language Processing (NLP) tasks; but not proportionally to dataset scale. In fact; we find that narrow subsets of our dataset sometimes outperform more diverse datasets. For example; finetuning on software documentation from here raises FSL performance by a mean of +7.5% on 52 downstream tasks; which beats training on 40 human-curated NLP datasets (+6.7%). Finetuning on various narrow datasets leads to similar broad improvements across test tasks; suggesting that the gains are not from domain adaptation but adapting to FSL in general. We do not observe clear patterns between the datasets that lead to FSL gains; leaving open questions about why certain data helps with FSL.

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI