Shi Feng
Principle investigator, Praxis; Assistant Professor, GWU
-
Shi does empirical safety research as the PI of Praxis Research and an assistant professor at GWU. Before that he was a postdoc with Sam Bowman and He He at NYU, and with Chenhao Tan at UChicago. His current focus is building an epistemic infrastructure for falsifiable intent misalignment, in particular building realistic model organisms of deception. His previous work includes metacogntive interventions for misalignment containment, demo of deception from RLHF and collusion from self-recognition.
-
Our research agenda is available at praxis-research.org/fw26-agenda. We want to build the epistemic infrastructure for empirically solving intent alignment. This means in addition to research questions like "how can we train a model organism for realistic deception?", we also spend a lot of time thinking about "what does it mean for a model organism of deception to be realistic?". The agenda page has more details, but here are some concrete projects building on existing ones:
- Extend praxis-research.org/comparative-motivation-profile into a full round of interpretive debate. Build infrastructure for automated interpretive debate.
- Use praxis-research.org/covert-influence to demonstrate risks in alignment research hand-off.
- Explore using more methods of praxis-research.org/metacognitive-training to shape generalization. Understand why they work.
- Explore https://praxis-research.org/grafting as an alternative to midtraining for persona and model organism research.
We are also interested in exploring:
- Epistemic risks, especially in alignment research and in multi-agent settings
- Relax ELK into research questions that are more empirically approachable -
- Passionate about open, collaborative research.
- Strong conceptual and analytical thinking abilities, which in many cases compensate for lack of experience or familiarity.
- Rigorous in execution. Have good judgment and taste to not generate AI slop.
- Strive to communicate ideas clearly.