Adrià Garriga-Alonso

Member of Technical Staff

Adrià Garriga-Alonso was a scientist at FAR.AI, working on understanding what learned optimizers want.

Previously he worked at Redwood Research on neural network interpretability. He holds a PhD from the University of Cambridge, where he worked on Bayesian neural networks; his advisor was Prof. Carl Rasmussen.

He was a co-organiser of the ICLR 2019 workshop “Safe Machine Learning: Specification, Robustness and Assurance”.

Publications

Path Channels and Plan Extension Kernels: a Mechanistic Description of Planning in a Sokoban RNN

Interpretability

Planning is essential for solving complex tasks, yet the internal mechanisms underlying planning in neural networks remain poorly understood. Building on prior work, we analyze a recurrent neural network (RNN) trained on Sokoban, a challenging puzzle requiring sequential, irreversible decisions. We find that the RNN has a causal plan representation which predicts its future actions about 50 steps in advance. The quality and length of the represented plan increases over the first few steps. We uncover a surprising behavior: the RNN "paces" in cycles to give itself extra computation at the start of a level, and show that this behavior is incentivized by training. Leveraging these insights, we extend the trained RNN to significantly larger, out-of-distribution Sokoban puzzles, demonstrating robust representations beyond the training regime. We open-source our model and code, and believe the neural network's interesting behavior makes it an excellent model organism to deepen our understanding of learned planning.

December 3, 2025
Date Range

Among us: A sandbox for measuring and detecting agentic deception

Alignment

We introduce Among Us, a sandbox social deception game where LLM-agents exhibit long-term, open-ended deception as a consequence of the game objectives. While most benchmarks saturate quickly, Among Us can be expected to last much longer, because it is a multi-player game far from equilibrium.

April 4, 2025
Date Range

Interpreting emergent planning in model-free reinforcement learning

Interpretability

We present the first mechanistic evidence that model-free reinforcement learning agents can learn to plan. This is achieved by applying a methodology based on concept-based interpretability to a model-free agent in Sokoban -- a commonly used benchmark for studying planning. Specifically, we demonstrate that DRC, a generic model-free agent introduced by Guez et al. (2019), uses learned concept representations to internally formulate plans that both predict the long-term effects of actions on the environment and influence action selection.

April 1, 2025
Date Range

Illusory Safety: Redteaming DeepSeek R1 and the Strongest Fine-Tunable Models of OpenAI, Anthropic, and Google

Robustness

DeepSeek-R1 has recently made waves as a state-of-the-art open-weight model, with potentially substantial improvements in model efficiency and reasoning. But like other open-weight models and leading fine-tunable proprietary models such as OpenAI’s GPT-4o, Google’s Gemini 1.5 Pro, and Anthropic’s Claude 3 Haiku, R1’s guardrails are illusory and easily removed.

February 3, 2025
Date Range

Planning behavior in a recurrent neural network that plays Sokoban

Interpretability

To understand how neural networks generalize, we studied an RNN trained to play Sokoban. The RNN learned to spend time planning ahead by "pacing" despite penalties for "taking longer", demonstrating that reinforcement learning can encourage strategic planning in neural networks.

July 21, 2024
Date Range

Adversarial Circuit Evaluation

Interpretability

Evaluating three neural network circuits (IOI, greater-than, and docstring) under adversarial conditions reveals that the IOI and docstring circuits fail to match the full model's behavior even on benign inputs.

July 20, 2024
Date Range

Investigating the Indirect Object Identification circuit in Mamba

Interpretability

By adapting existing interpretability techniques to the Mamba architecture, we partially reverse-engineered the circuit responsible for the Indirect Object Identification task.

July 18, 2024
Date Range

InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques

Interpretability

InterpBench is a collection of transformers with known circuits, trained using Strict Interchange Intervention Training (SIIT). These models exhibit realistic weights that reflect ground truth circuits, providing a benchmark for evaluating mechanistic interpretability techniques.

July 18, 2024
Date Range

Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecification

Alignment

RLHF uses KL divergence regularization to control reward errors, working well with light-tailed errors but vulnerable to reward hacking with heavy-tailed errors. Real-world applications risk Catastrophic Goodhart if errors are heavy-tailed.

July 18, 2024
Date Range

Towards Automated Circuit Discovery for Mechanistic Interpretability

Interpretability

We systematize the mechanistic interpretability process into 3 iterative steps, then proceed to automate one of them: circuit discovery. Two of the algorithms presented automatically discover interpretability results previously established by human inspection.

July 3, 2023
Date Range

News

Pacing Outside the Box: RNNs Learn to Plan in Sokoban

Interpretability

Giving RNNs extra thinking time at the start boosts their planning skills in Sokoban. We explore how this planning ability develops during reinforcement learning. Intriguingly, we find that on harder levels the agent paces around to get enough computation to find a solution.

July 23, 2024
Date Range

Illusory Safety: Redteaming DeepSeek R1 and the Strongest Fine-Tunable Models of OpenAI, Anthropic, and Google

Red-Teaming & Evaluation

DeepSeek-R1 has recently made waves as a state-of-the-art open-weight model, with potentially substantial improvements in model efficiency and reasoning. But like other open-weight models and leading fine-tunable proprietary models such as OpenAI’s GPT-4o, Google’s Gemini 1.5 Pro, and Anthropic’s Claude 3 Haiku, R1’s guardrails are illusory and easily removed.

February 3, 2025
Date Range

Research

Our research explores a portfolio of high-potential agendas.

Events

Our events bring together global leaders in AI.

Programs

Our programs build the field of trustworthy and secure AI