Papers
How Transparent is DiffusionGemma? Joshua Engels, Callum McDougall, Bilal Chughtai, ..., Rohin Shah, and Neel Nanda. arXiv preprint, 2026. Paper | Blog | Twitter
Building Production-Ready Probes for Gemini. Janos Kramar, Joshua Engels, Zhengxuan Wang, Bilal Chughtai, Rohin Shah, Neel Nanda, and Arthur Conmy. arXiv preprint, 2026. Paper | Twitter
Training on Documents About Monitoring Leads to CoT Obfuscation. Reilly Haskins, Bilal Chughtai, and Joshua Engels. arXiv preprint, 2026. Paper | Blog | Twitter
When Reading the Chain of Thought Falls Short: A Testbed for Reasoning Trace Analysis. Daria Ivanova, Riya Tyagi, Joshua Engels, and Neel Nanda. Mechanistic Interpretability Workshop at ICML 2026. Paper | Blog
Designing Effective Monitor-Based Interventions for Mitigating Reward Hacking During RL. Aria Wong, Joshua Engels, and Neel Nanda. Mechanistic Interpretability Workshop at ICML 2026. Paper | Blog
Scaling Laws For Scalable Oversight. Joshua Engels*, David Baek*, Subhash Kantamneni*, and Max Tegmark. Neurips 2025 (Spotlight). Paper | Code | Twitter
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing. Subhash Kantamneni*, Joshua Engels*, Senthooran Rajamanoharan, Max Tegmark, and Neel Nanda. ICML 2025. Paper | Blog | Code | Twitter
Low Rank Adapting Models for Sparse Autoencoders. Mathew Chen*, Joshua Engels*, and Max Tegmark. ICML 2025. Paper | Code | Twitter
Simple Mechanistic Explanations for Out-Of-Context Reasoning. Atticus Wang*, Joshua Engels*, Oliver Clive-Griffin*, Senthooran Rajamanoharan, and Neel Nanda. ICML 2025 Workshop on Reliable and Responsible Foundation Models. Paper
Dense SAE Latents Are Features, Not Bugs. Xiaoqing Sun, Alessandro Stolfo, Joshua Engels, Ben Wu, Senthooran Rajamanoharan, Mrinmaya Sachan, and Max Tegmark. Neurips 2025. Paper
Decomposing the Dark Matter of Sparse Autoencoders. Joshua Engels, Logan Smith, and Max Tegmark. TMLR 2025. Paper | Code | Twitter
The Geometry of Concepts: Sparse Autoencoder Feature Structure. Yuxiao Li, Eric J. Michaud, David D. Baek, Joshua Engels, Xiaoqing Sun, and Max Tegmark. Entropy 2025. Paper
Efficient Dictionary Learning with Switch Sparse Autoencoders. Anish Mudide, Joshua Engels, Eric J Michaud, Max Tegmark, and Christian Schroeder de Witt. ICLR 2025. Paper | Code | Twitter
Not All Language Model Features Are Linear. Joshua Engels, Eric J. Michaud, Isaac Liao, Wes Gurnee, and Max Tegmark. ICLR 2025. Paper | Code | Twitter | Talk
* indicates equal contribution
Talks
Developing Areas of Alignment Science — June 2026
SAEs: Progress and Limitations — March 2025
Not All Language Model Features Are Linear — December 2024
Other Projects and Writing
LLM-Driven Feature Discovery — June 2026
Why Do Naive SFT Filters For Safety Properties Fail? — June 2026
SFT Drives Gemini's Safety Properties — June 2026
Building and Evaluating Model Diffing Agents — June 2026
Thought Editing: Steering Models by Editing Their Chain of Thought — February 2026
Brief Explorations in LLM Value Rankings — January 2026
Can We Interpret Latent Reasoning Using Current Mechanistic Interpretability Tools? — December 2025
Prompting Models to Obfuscate Their CoT — December 2025
How Can Interpretability Researchers Help AGI Go Well? — December 2025
A Pragmatic Vision for Interpretability — December 2025
Current LLMs Seem to Rarely Detect CoT Tampering — November 2025
Negative Results on Group SAEs — May 2025
Interim Research Report: Mechanisms of Awareness — May 2025
Spreadsheet of 50 Weird LLM Phenomenon — May 2025
TinySAE: A Minimal SAE Implementation — March 2025