Hi there! I’m a sophomore at Stanford studying Computer Science, and dabbling in everything that lets me think about minds and intentions or Chinese society.
I do technical AI safety research focused on alignment training, persona alignment, and elicitation evals. I also read and write about strategy, and hope to contribute to US-China cooperation on safety.
I grew up in Beijing, Copenhagen, and Singapore. In those lives, I was a Team China debater, a historian of queerness in modern China, a facilitator of political dialogue at my high school, and a math enthusiast.
What I'm up to
Built & iterated plan for coming six months based on feedback from many kind people! Build technical research experience, read strategy, help campus safety where possible, and reach out to people from Chinese safety ecosystem.
Learned Inspect and finished up to agentic evaluations chapter
Built proof-of-concept BlueDot grant for running frontier model sweep! Initial results suggesting frontier models incorrigible to legitimate process: github.com/corrig-eval
My notes here: app.notion.com/Paper-Reading-3c3739c5142080 Includes papers from below as well.
Read ~20 papers on natural language interp, control, evals, and alignment training. Used J-lens to inspect llama-8b in an alignment-faking scenario. Unfortunately it was too weak to attempt resolute deception. Built replication proposals for assistant axis, broad alignment from RL, and CoT monitorability papers.
Did not have time, although had limited coverage from foundational & deep-dive readings.
The goal is to replicate and extend load-bearing research in AI safety. Replications feel like a great opportunity to improve research taste, upskill, and find starting points for bigger projects.
Build foundations for these gaps by implementing from scratch
Dwarkesh, Mowshowitz
Theory: REINFORCE, RLOO, PPO, GRPO Build from scratch via relevant chapters in rasbt/reasoning-from-scratch Run baseline experiments with OLMo
Begin strategy exploration in earnest OAI-HF incident seems to have updated debates/strategic focus