Euan McLean

Communications Specialist

Euan is a communications specialist at FAR.AI. In the past he has completed a PhD in theoretical particle physics at the University of Glasgow, worked as a machine learning engineer at a cybersecurity startup, and worked as a strategy researcher at the Center on Long Term Risk. He is also a scriptwriter for the YouTube channel PBS Spacetime. His passion is reducing interpretive labor in AI alignment to speed up the progress of the field.

Publications

Can Go AIs be adversarially robust?

Robustness

We tested three approaches to defend Go AIs from adversarial strategies. While these defenses protect against previously discovered adversaries, we uncovered qualitatively new adversaries that undermine these defenses.

June 17, 2024
Date Range

Exploiting Novel GPT-4 APIs

Robustness

We red-team three new functionalities exposed in the GPT-4 APIs: fine-tuning, function calling and knowledge retrieval. We find that fine-tuning a model on as few as 15 harmful examples or 100 benign examples can remove core safeguards from GPT-4, enabling a range of harmful outputs. Furthermore, we find that GPT-4 Assistants readily divulge the function call schema and can be made to execute arbitrary function calls. Finally, we find that knowledge retrieval can be hijacked by injecting instructions into retrieval documents.

December 20, 2023
Date Range

Inverse Scaling: When Bigger Isn't Better

Model Evaluations

We present 11 instances of inverse scaling: tasks where language models get worse with scale rather than better, selected from 99 submissions in an open competition, the Inverse Scaling Prize.

June 14, 2023
Date Range

News

We Found Exploits in GPT-4’s Fine-tuning & Assistants APIs

Red-Teaming & Evaluation

We red-team three new functionalities exposed in the GPT-4 APIs: fine-tuning, function calling and knowledge retrieval. We find that fine-tuning a model on as few as 15 harmful examples or 100 benign examples can remove core safeguards from GPT-4, enabling a range of harmful outputs.

December 20, 2023
Date Range

Even Superhuman Go AIs Have Surprising Failure Modes

Robustness & Security

Our adversarial testing algorithm uncovers a simple, human-interpretable strategy that consistently beats superhuman Go AIs.

July 14, 2023
Date Range

Codebook Features: Sparse and Discrete Interpretability for Neural Networks

Interpretability

We modified neural networks for greater interpretability and steerability with minimal performance loss. Each layer applies a quantization bottleneck, converting dense activation vectors into a discrete list of learned codes that are either on or off.

October 18, 2023
Date Range

Big Picture AI Safety

May 22, 2024
Date Range

Beyond the Board: Exploring AI Robustness Through Go

Robustness & Security

Achieving robustness remains a significant challenge even in narrow domains like Go. We test three approaches to defend Go AIs from adversarial strategies. We find these defenses protect against previously discovered adversaries, but uncover qualitatively new adversaries that undermine these defenses.

June 17, 2024
Date Range

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI