Hadas Orgad
Research Fellow, Harvard University
-
Hadas is a Research Fellow at the Kempner Institute for the Study of Natural and Artificial Intelligence at Harvard University. Her research focuses on understanding the internal mechanisms of AI models, with an emphasis on problems that are not solved with scale of data/compute. Her research bridges interpretability and practical deployment, focusing on harmful model behaviors such as hallucinations, bias, privacy violations, and unsafe outputs.
Hadas received her Ph.D. in Computer Science from the Technion, advised by Yonatan Belinkov. Her doctoral research studied the robustness of AI systems through interpretability, evaluation, and model interventions, spanning language and vision models. Before joining Harvard, she worked as a researcher at Microsoft and was an Apple Scholar in AI/ML -
Hadas is interested in mentoring projects at the intersection of AI safety, interpretability, and robustness, particularly work that uses causal or mechanistic methods to understand the mechanisms governing undesirable or unexpected model behaviors. She is interested in questions around harmfulness, sycophancy, deceptive behavior, situational awareness, and hallucination, including how these behaviors are internally organized and how they emerge or change through training and other interventions.
She is also interested in studying the prevalence and practical significance of safety and reliability issues in open-weight models, especially by developing evaluations that move beyond highly controlled or artificial settings toward more realistic conditions. Across these directions, she is excited about work that combines careful behavioral evaluation with mechanistic analysis on realistic settings, to better understand model failures and how they can be mitigated. -
Ownership and accountability. The ability to take responsibility for a research project end-to-end, proactively identify next steps and bottlenecks, and stay motivated in pushing the work forward.
*Research maturity. The ability to identify and understand relevant literature and situate your contribution within in, formulate clear hypotheses, design informative experiments, and analyze the results critically.
* Scientific rigor and adversarial thinking. The habit of actively looking for alternative explanations for one’s own results, designing controls that distinguish between them, and being willing to revise a hypothesis when the evidence points elsewhere. This also means looking closely at the underlying data and model outputs, rather than relying only on aggregate numerical results.
* Deep technical understanding. Fellows should be able to explain what they did in detail and understand the important components of their experimental pipeline.
* Clear communication and collaboration. The ability to communicate ideas, experimental decisions, results, and uncertainties clearly, and to engage productively with feedback.