Papers

How Transparent is DiffusionGemma? Joshua Engels, Callum McDougall, Bilal Chughtai, ..., Rohin Shah, and Neel Nanda. arXiv preprint, 2026. Paper | Blog | Twitter

Building Production-Ready Probes for Gemini. Janos Kramar, Joshua Engels, Zhengxuan Wang, Bilal Chughtai, Rohin Shah, Neel Nanda, and Arthur Conmy. arXiv preprint, 2026. Paper | Twitter

Training on Documents About Monitoring Leads to CoT Obfuscation. Reilly Haskins, Bilal Chughtai, and Joshua Engels. arXiv preprint, 2026. Paper | Blog | Twitter

When Reading the Chain of Thought Falls Short: A Testbed for Reasoning Trace Analysis. Daria Ivanova, Riya Tyagi, Joshua Engels, and Neel Nanda. Mechanistic Interpretability Workshop at ICML 2026. Paper | Blog

Designing Effective Monitor-Based Interventions for Mitigating Reward Hacking During RL. Aria Wong, Joshua Engels, and Neel Nanda. Mechanistic Interpretability Workshop at ICML 2026. Paper | Blog

Scaling Laws For Scalable Oversight. Joshua Engels*, David Baek*, Subhash Kantamneni*, and Max Tegmark. Neurips 2025 (Spotlight). Paper | Code | Twitter

Are Sparse Autoencoders Useful? A Case Study in Sparse Probing. Subhash Kantamneni*, Joshua Engels*, Senthooran Rajamanoharan, Max Tegmark, and Neel Nanda. ICML 2025. Paper | Blog | Code | Twitter

Low Rank Adapting Models for Sparse Autoencoders. Mathew Chen*, Joshua Engels*, and Max Tegmark. ICML 2025. Paper | Code | Twitter

Simple Mechanistic Explanations for Out-Of-Context Reasoning. Atticus Wang*, Joshua Engels*, Oliver Clive-Griffin*, Senthooran Rajamanoharan, and Neel Nanda. ICML 2025 Workshop on Reliable and Responsible Foundation Models. Paper

Dense SAE Latents Are Features, Not Bugs. Xiaoqing Sun, Alessandro Stolfo, Joshua Engels, Ben Wu, Senthooran Rajamanoharan, Mrinmaya Sachan, and Max Tegmark. Neurips 2025. Paper

Decomposing the Dark Matter of Sparse Autoencoders. Joshua Engels, Logan Smith, and Max Tegmark. TMLR 2025. Paper | Code | Twitter

The Geometry of Concepts: Sparse Autoencoder Feature Structure. Yuxiao Li, Eric J. Michaud, David D. Baek, Joshua Engels, Xiaoqing Sun, and Max Tegmark. Entropy 2025. Paper

Efficient Dictionary Learning with Switch Sparse Autoencoders. Anish Mudide, Joshua Engels, Eric J Michaud, Max Tegmark, and Christian Schroeder de Witt. ICLR 2025. Paper | Code | Twitter

Not All Language Model Features Are Linear. Joshua Engels, Eric J. Michaud, Isaac Liao, Wes Gurnee, and Max Tegmark. ICLR 2025. Paper | Code | Twitter | Talk

* indicates equal contribution

Talks

Developing Areas of Alignment Science — June 2026

SAEs: Progress and Limitations — March 2025

Not All Language Model Features Are Linear — December 2024

Other Projects and Writing

LLM-Driven Feature Discovery — June 2026

Why Do Naive SFT Filters For Safety Properties Fail? — June 2026

SFT Drives Gemini's Safety Properties — June 2026

Building and Evaluating Model Diffing Agents — June 2026

Thought Editing: Steering Models by Editing Their Chain of Thought — February 2026

Brief Explorations in LLM Value Rankings — January 2026

Can We Interpret Latent Reasoning Using Current Mechanistic Interpretability Tools? — December 2025

Prompting Models to Obfuscate Their CoT — December 2025

How Can Interpretability Researchers Help AGI Go Well? — December 2025

A Pragmatic Vision for Interpretability — December 2025

Current LLMs Seem to Rarely Detect CoT Tampering — November 2025

Negative Results on Group SAEs — May 2025

Interim Research Report: Mechanisms of Awareness — May 2025

Spreadsheet of 50 Weird LLM Phenomenon — May 2025

TinySAE: A Minimal SAE Implementation — March 2025