James Mickens
Gordon McKay Professor of Computer Science, Harvard University
-
James Mickens is a professor of computer science at Harvard University. He received a bachelor's degree in computer science from Georgia Tech, and a PhD in computer science from the University of Michigan. Post-graduation, he worked at Microsoft Research for seven years in the Distributed Systems Research Group. Prior to coming to Harvard, he was also a visiting professor at MIT's Parallel and Distributed Operating Systems group. His current research involves the performance, security, and privacy of complex applications; a particular focus is on sandboxing mechanisms for misaligned AIs.
-
- Robust systems-level mechanisms for preventing misaligned models from doing harm; examples of systems-level sandboxes include hardened inference engines, container runtimes, and virtualization platforms.
- Understanding how emergent misalignment develops during training and fine-tuning, e.g., using techniques from the grokking literature to map the gradual progression of misalignment.
- Alignment red-teaming, e.g., stress-testing model alignment to identify unexpected scenarios in which misalignment can arise. -
- An understanding of basic techniques in mechanistic interpretability and steering (e.g., linear probing, sparse autoencoders, feature ablation).
- Prior hands-on experience with model training and fine-tuning.
- Prior hands-on experience with deploying inference-as-a-service stacks and/or writing systems code in C/C++/Rust/Python.
- Good communication skills.
- Familiarity with reading academic papers and conducting new research.