Benno Krojer & David Bau
Podstdoc, Northeastern University; Assistant Professor, Northeastern University
-
Benno Krojer is a postdoctoral researcher in David Bau's lab at Northeastern University, and recently completed his PhD at Mila & McGill in Montreal. He studies the intricate relation between language and perception in current LLMs & VLMs, often via interpretability. During his PhD, and while interning at Meta FAIR, he worked on various aspects of vision-and-language reasoning, benchmarking/shortcut learning, video understanding and diffusion models. Outside of his core research, Benno is active in science communication (https://bennokrojer.com/behind-the-scenes.html, https://open.spotify.com/show/1bvb3XiuuMe7XJ5GnRW84U), community building and various sports.
David Bau is an assistant professor in the Khoury College of Computer Sciences at Northeastern University, based in Boston.Bau's research focuses on human-computer interaction and machine learning. Before joining Northeastern, he worked as a software engineer at Google, BEA, and Crossgain. He has been published in journals such as CVPR, NeurIPS, ICCV, ECCV, and SIGGRAPH.
Outside of research, Bau enjoys astronomy and puzzle collecting.
-
My primary expertise and direction lies in multimodal interpretability. Some examples what we could explore: To what extent are perception and language "unified" in models (circuits, representations, steering, ...)? How do models reason across vision and language internally? How can we design better interp tools tailored to multimodality? If we pre-train natively multimodal, how do circuits/capabilites of different modalities emerge and synergize? Flexible cross-modal CoT? Multimodal safety, larger attack surfaces through other modalities?
I am also excited to mentor:
- High-level cognition of LLMs, for example: 1) Introspection, 2) Interpretability of latent thinking, 3) When LLMs say things a researcher would say ("I read/understood the paper", "Here is my research idea: ..."): What is actually happening internally, how are these "ideas" formed/stored (methodologically: how can we develop interp tools that go beyond the established single-token approaches)?
Broader questions around interpretability in the wild/science of science: 1) When are interpretability explanations actually useful for ordinary users/other fields (HCI); Or, how often on a dataset of actual user interactions, could we arrive at any interp insights beyond behavioral evaluation? 2) Can we find patterns/insights across hundreds of interp papers (meta studies)? Can we identify & fix assumptions/misconceptions/vague definitions that most of the literature makes? -
Prior experience in one of the following: interpretability, VLMs, cognitive science.
Mindset: Curiosity for fundamental science, high-standards for scientific communication/rigor of the findings, willingness to finish it up to the end (e.g. rebuttals, maintaining a website/demo, ...), communication througout the project (regular small updates and nicely documenting intermediate findings, keep all collaborators involved/excited).