All Recordings

Topic

Year

Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.

Results 1

Showing items 1 to3 of 6

Filtering by:
Tag
Operator
Value

Task Decomposition for AI Control

Aaron Sandoval

March 27, 2025

2025

Deluding AIs

Owain Evans

March 27, 2025

2025

High Integrity Research Practices at Palisade

Dmitrii Volkov

March 27, 2025

2025

Lakera: Runtime Security for Agents at Scale

Sam Watts

March 27, 2025

2025

Improved Monitoring of Backdoor Insertion During Code Refactoring

Trevor Lohrbeer

March 27, 2025

2025

Low-stakes Control

Vivek Hebbar

March 27, 2025

2025

[Fireside Chat] White-box Methods for AI Control

Neel Nanda

March 27, 2025

2025

Hopes & Difficulties with Using Control Protocols in Production

Fabien Roger

March 27, 2025

2025

STPA & AI Control

Simon Mylius

March 27, 2025

2025

Ctrl-Z: Controlling AI Agents via Resampling

Aryan Bhatt

March 26, 2025

2025

Optimization around Control: Lessons from MONA

Sebastian Farquhar

March 26, 2025

2025

Automated Researchers Can Subtly Sandbag

Johannes Gasteiger

March 26, 2025

2025

Threat Analysis to Identify Priorities

Francesca Gomez

March 26, 2025

2025

Hierarchical Monitoring

Tim Hua

March 26, 2025

2025

The Role of AISIs in AI Governance

Rob Reich

March 11, 2025

2025

Challenges in Creating “If-Then Commitments” for Extreme Security

Christopher Painter

February 9, 2025

2025

Fireside Chat with Zico Kolter

Zico Kolter

February 9, 2025

2025

Cyber Range Design at UK AISI

Mahmoud Ghanem

February 9, 2025

2025

Security for AI with Confidential Computing

Mike Bursell

February 9, 2025

2025

Jailbreaking AI-Controlled Robots

Alex Robey

February 9, 2025

2025

Fireside Chat with Gregory Allen

Gregory Allen

February 8, 2025

2025

AI Control and Why AI Security People Should Care

Buck Shlegeris

February 8, 2025

2025

Governance Challenges of Biological AI Models

Melissa Hopkins

February 8, 2025

2025

Set Laws AI Must Prove It's Following

Evan Miyazono

February 8, 2025

2025

How Policy Changes are Effectively Implemented within the US Government

Austin Carson

February 8, 2025

2025

AI's Implications for Nuclear Weapons Systems

Philip Reiner

February 8, 2025

2025

Impact of Frontier AI in the Landscape of Cyber Security

Dawn Song

February 8, 2025

2025

Safeguards in 2025+

Xander Davies

February 8, 2025

2025

flexHEG

David A. Dalrymple (davidad)

February 8, 2025

2025

The AI Security Landscape

Sella Nevo

February 8, 2025

2025

EU CoP's Security Focus

Yoshua Bengio

February 8, 2025

2025

Plan B: Training LLMs to fail less severely

Julian Stastny

February 4, 2025

2025

Eliciting the capabilities of scheming LLMs

Fabien Roger

January 14, 2025

2025

Empirical Progress on Debate

Julian Michael

October 25, 2024

2024

Targeted Manipulation & Deception

Micah Carroll

October 25, 2024

2024

Will Scaling Solve Robustness?

Adam Gleave

October 25, 2024

2024

Paradigms and Robustness

Alex Wei

October 25, 2024

2024

Generalized Adversarial Training and Testing

Stephen Casper

October 25, 2024

2024

Improving AI Safety with Top-Down Interpretability

Andy Zou

October 25, 2024

2024

State of Interpretability & Ideas for Scaling Up

Atticus Geiger

October 25, 2024

2024

Third-Party Evals: Learnings from Harmony Intelligence

Soroush Pour

October 24, 2024

2024

Reframing AGI Threat Models

Richard Ngo

October 24, 2024

2024

Connecting Capability Evals to Danger Thresholds

David Duvenaud

October 24, 2024

2024

AGI-Complete Evaluation

Joel Z. Leibo

October 24, 2024

2024

A Sociotechnical Approach to a Safe, Responsible AI Future

Dawn Song

October 24, 2024

2024

A Safe Harbor for AI Evaluation & Red Teaming

Shayne Longpre

October 24, 2024

2024

AI Policy in China

Kwan Yee Ng

October 24, 2024

2024

AI Control

Buck Shlegeris

October 24, 2024

2024

METR Updates & Research Directions

Beth Barnes

October 24, 2024

2024

Implications of Human Model Misspecification for Alignment

Anca Dragan

October 24, 2024

2024

Powering Up Capability Evaluations

Stephen Casper

October 23, 2024

2024

Value Pluralism and AI Value Alignment

Atoosa Kasirzadeh

October 23, 2024

2024

Using Formal Languages

Sheila McIlraith

October 23, 2024

2024

The (Un)Reliability of Chain-of-Thought Reasoning

Chirag Agarwal

October 23, 2024

2024

Tamper-Resistant Safeguards for Open-Weight LLMs

Mantas Mazeika

October 23, 2024

2024

MobileSafetyBench

Kimin Lee

October 23, 2024

2024

Gradient Routing

Alex Turner

October 23, 2024

2024

Formal Verification is Overrated

Zac Hatfield-Dodds

October 23, 2024

2024

Backdoors as an Analogy for Deceptive Alignment

Jacob Hilton

October 23, 2024

2024

Alignment Stress-Testing at Anthropic

Evan Hubinger

October 23, 2024

2024

3 Mechanisms Underlying Emergent Abilities in Generative Models

Hidenori Tanaka

October 22, 2024

2024

Campaigns in Emerging Issues: Lessons Learned from the Field

Andrew Freedman

August 27, 2024

2024

Deceptive Instrumental Alignment

Evan Hubinger

July 30, 2024

2024

Verification & Confidence Building for International Coordination

Peter Barnett

July 23, 2024

2024

Current Issues in AI Safety

Multiple Speakers

July 20, 2024

2024

What are Human Values, and How Do We Align AI to Them?

Oliver Klingefjord

July 20, 2024

2024

Towards Reliable Alignment: Uncertainty-Aware RLHF

Aditya Gopalan

July 20, 2024

2024

Stress-Testing Capability Elicitation

Dmitrii Krasheninnikov

July 20, 2024

2024

Some Lessons from Adversarial Machine Learning

Nicholas Carlini

July 20, 2024

2024

Scaling Reinforcement Learning from Human Feedback

Jan Leike

July 20, 2024

2024

Scalable Oversight: A Rater Assist Approach

Sophie Bridgers

July 20, 2024

2024

Resilience and Interpretability

David Bau

July 20, 2024

2024

Research Proposal: The Three-Layer Paradigm

Zhaowei Zhang

July 20, 2024

2024

Open Problems in Technical AI Governance

Ben Bucknall

July 20, 2024

2024

Mechanistic Interpretability: A Whirlwind Tour

Neel Nanda

July 20, 2024

2024

Measuring and Improving Human Agency in a World of AI Agents

Alex Tamkin

July 20, 2024

2024

Governance for Advanced General-Purpose AI

Helen Toner

July 20, 2024

2024

Game Theory and Social Choice for Cooperative AI

Vincent Conitzer

July 20, 2024

2024

Dangerous Capability Evals: Basis for Frontier Safety

Mary Phuong

July 20, 2024

2024

Challenges With Unsupervised LLM Knowledge Discovery

Vikrant Varma

July 20, 2024

2024

AI: What If We Succeed?

Stuart Russell

July 20, 2024

2024

Modeling and Mitigating Near-term Deployment Risks from LLMs

Alex Pan

July 13, 2024

2024

Simplex

Multiple Speakers

June 25, 2024

2024

Formal AI-Assisted Code Specification and Synthesis

Shaowei Lin

May 21, 2024

2024

Defending Against Adversarial Attacks in Go

Tom Tseng

April 30, 2024

2024

Category Theory

Kris Brown

April 28, 2024

2024

How Could We Design Aligned & Provably Safe Al?

Yoshua Bengio

April 16, 2024

2024

Sleeper Agents

Ethan Perez

February 14, 2024

2024

Towards Quantitative Safety Guarantees and Alignment

Yoshua Bengio

December 10, 2023

2023

Challenges for AI Safety in Adversarial Settings

Dawn Song

December 9, 2023

2023

Your RLHF fine-tuning is Secretly Applying a Regret Preference

Brad Knox

December 9, 2023

2023

What (and Why) is a Linear Representation?

Victor Veitch

December 9, 2023

2023

Weak-to-Strong Generalization

Collin Burns

December 9, 2023

2023

Understanding Planning in Neural Networks

Adrià Garriga-Alonso

December 9, 2023

2023

Theories and Tools for Mechanistic Interpretability via Causal Abstraction

Atticus Geiger

December 9, 2023

2023

The Quantization Model of Neural Scaling

Eric J. Michaud

December 9, 2023

2023

The Impossibility of (Strong) Watermarking

Boaz Barak

December 9, 2023

2023

The Challenge of Monitoring Covert Interactions and Behavioral Shifts in LLM Agents

Dimitris Papailiopoulos

December 9, 2023

2023

System Two Safety

Shane Legg

December 9, 2023

2023

Studying LLM Generalization through Influence Functions

Roger Grosse

December 9, 2023

2023

There are no available cars matching the current filters.