San Francisco Alignment Workshop

San Francisco, CA

•

February 27, 2023
February 28, 2023
Date Range

Overview

The inaugural Alignment Workshop was held 27–28 February 2023 in San Francisco, co-organized by researchers from top industry AI labs and universities, and was attended by 80 of the world’s leading machine learning researchers.

The Alignment Workshop series brings together top machine learning researchers and practitioners from industry, academia, and government. The workshop focuses on discussing and debating critical topics related to AI alignment, enabling participants to better understand potential risks from advanced AI, and strategies for solving them. Key issues discussed include model evaluations, interpretability, robustness, and AI governance.

San Francisco Alignment Workshop sessions

Lightning Talks (Day 2)

Multiple Speakers

•

February 27, 2023

2023

“Situational Awareness” Makes Measuring Safety Tricky

Ajeya Cotra

•

February 26, 2023

2023

Surveying Safety Research Directions

Dan Hendrycks

•

February 26, 2023

2023

Supervising AI on Hard Tasks

Jan Leike

•

February 26, 2023

2023

Opening Remarks: Confronting the Possibility of AGI

Ilya Sutskever

•

February 26, 2023

2023

Looking Inside Neural Networks with Mechanistic Interpretability

Chris Olah

•

February 26, 2023

2023

Lightning Talks (Day 1)

Multiple Speakers

•

February 26, 2023

2023

How Misalignment Could Lead to Takeover

Paul Christiano

•

February 26, 2023

2023

Aligning Massive Models: Current and Future Challenges

Jacob Steinhardt

•

February 26, 2023

2023