Kevin Wei
Research Scholar, GovAI
-
Kevin Wei (he/they) is a researcher at the Centre for the Governance of AI (GovAI), where they work on ensuring that advanced AI is developed and governed safely. Specifically, Kevin's research agenda is focused on the science of AI evaluations, legal AI safety/alignment, and technical AI governance/law, with peer-reviewed publications on these topics appearing in ICML, TMLR, and other top AI/AI ethics venues.
Kevin is also affiliated with the Oxford Martin School AI Governance Initiative and the RAND Center for AI, Security, and Technology; they were previously a Visiting Research Scientist on the UK AI Security Institute's science of evaluations team, a Fellow at RAND, and a program manager at a cloud company. They received a J.D. from Harvard Law School, a Master's in Global Affairs from Tsinghua University (where they were a Schwarzman Scholar), an M.S. in Machine Learning from Georgia Tech, and a B.A. in Mathematics-Statistics & Economics from Columbia University. -
I would be interested in mentoring projects in (1) science of evaluations; (2) law-following AI (evaluations), including empirical research on model spec/AI constitution compliance; and (3) U.S. AI governance/law. Projects in (1) and (2) will be team projects; projects in (3) can be individual projects.
A list of possible project ideas is below; the list is illustrative but not exhaustive, i.e., I am open to supervising other projects in this space as long as they are useful/impactful.
Please note that I will not be a suitable mentor for applicants who are primarily interested in (mechanistic) interpretability research, EU law, or AI security.
1. Science of evaluations: How do we build better evaluations for alignment, capabilities, and safety? Some potential projects include but are not limited to:
- Develop best practices for particular methods in AI evaluations (e.g., best practices for dealing with eval awareness, best practices for user simulation, best practices for LLM-as-a-judge, etc.)
- Take existing research/phenomena about QA/shorter evaluations and figure out if those results hold up in agentic, long-horizon evaluations
- Do exploratory work in areas that are more neglected by science of evals research (e.g., AI control evals, multi-agent/cooperative AI evals, etc.)
- Building better statistical tools for evaluations
- Improving transcript analysis on agentic, long-horizon evaluations
- See also the open questions lists from US CAISI and UK AISI.
2. Legal alignment evaluations: Can we build automated assessments for AI agents' abilities and propensities to comply with legal requirements? See 4.1 and somewhat 4.2 here; for examples of work in this direction, see here and here. Some potential projects include but are not limited to:
- How well do LLMs comply with model specs/AI constitutions?
- How do models interpret or reason about model specs/AI constitutions, and do their interpretations or behaviors differ from human interpretations?
- Laws can have hierarchies, analogous to instruction hierarchies, eg, a US federal law taking precedent over a state law. Do models respect these hierarchies?
- Build an evaluation for AI agents' compliance with the law (for a specific legal doctrine or body of law)
3. AI governance: I am generally open to supervising AI governance projects, especially projects that (1) require technical understanding to execute, (2) are related to U.S. legal implementation or challenges, or (3) use quantitative or empirical social science methods to answer questions relevant for frontier AI governance (e.g., survey methods). I am particularly interested in evals-related policy. There are many potential projects in this space, so I won't list specific ones; you can simply flag to me that you are interested in U.S. AI law/policy generally. -
All candidates:
- Interest in AI safety
- Commitment to doing impactful research
- Ability to quickly skill up / get up to speed in an unfamiliar subfield or literature
- Ability to execute projects quickly
- Willingness to pivot from existing / prior directions that turn out to be less promising or impactful than expected
- Desire to write, publish, and/or disseminate a research output (which can be anything from an academic paper or technical blog post, to a policy report or op-ed)
Candidates interested in science of evals or evals projects should be familiar with Python and have either (A) previous research experience in machine learning or (B) experience in software engineering or data science. Familiarity with statistical modeling/causal inference (e.g., econometrics, psychometrics) and with evaluation frameworks such as Inspect/Inspect Scout/HiBayes are a plus (but not required).
Candidates interested in legal alignment can either (A) satisfy some of the requirements above for candidates interested in evaluations projects, or (B) have received or be in the process of completing a law degree (JD, LLB, LLM, etc).
Candidates interested in governance projects should have prior experience in public policy (e.g., roles in research, advocacy, etc. in any field, including but not limited to AI policy).