Uncovering Latent Human Wellbeing in Language Model Embeddings
Interpretability
Do language models implicitly learn a concept of human wellbeing? We explore this through the ETHICS Utilitarianism task, assessing if scaling enhances pretrained models' representations.
February 18, 2024
Date Range
News
Uncovering Latent Human Wellbeing in LLM Embeddings
Red-Teaming & Evaluation
A one-dimensional PCA projection of OpenAI's `text-embedding-ada-002` model achieves 73.7% accuracy on the ETHICS Util test dataset.
September 11, 2023
Date Range
Research
Our research explores a portfolio of high-potential agendas.
Events
Our events bring together global leaders in AI.
Programs
Our programs build the field of trustworthy and secure AI