Dan Braun
Member of Technical Staff, Goodfire
-
I've been working on Mechanistic Interpretability since 2022 (Conjecture, Apollo Research, Goodfire). I think "research engineer" is probably the best description of my strengths; most of my work has been taking ideas from the whiteboard or conversations, implementing them in code, and designing experiments to validate them.
-
- Developing and applying parameter decomposition-related methods to interpret various types of networks. E.g. https://arxiv.org/abs/2607.13047
- Training inherently interpretable networks
- Investigating different "languages of what makes an ideal interpretation" (which may involve e.g. causal abstractions, english text, knowledge bases). The end goal is to find a good language and let automated agents hillclimb on it to fully reverse engineer neural networks.
- AI sentience and the behaviour of agents in various multi-agent setups. (I'm certainly no expert here, but am very interested and would hopefully be helpful with experiment design and presentation). -
- Motivated by reducing catastrophic risks from AI
- Strong written communication skills (e.g. ability/tendency to write regular updates explaining results and plans)