Hello
Hi! I’m Cath Ge-Wang, a mathematics undergraduate at Christ Church, University of Oxford. I’m currently working on building misalignment continuation evals with my mentors Alex and Rob at UK AISI. I’m also working on verification protocols as a MIRI Technical Governance Team fellow. I was previously a part-time research collaborator at Redwood Research, and I help run the Oxford AI Safety Initiative’s Policy Team.
My primary research interests lie in AI control, adversarial robustness, and agentic evaluations, particularly to understand and mitigate emergent misalignment risks in autonomous AI systems. I am also interested in how these technical insights inform AI governance and policy, especially hardware verification, mechanisms for strategic risk, and constraining dangerous capability deployment.
My Research
Published:
Catherine Ge-Wang(=), Tyler Crosse(=), Benjamin Hadad IV, Joachim Schaeffer, Ram Potham, and Tyler Tracy, 2026, “Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety”. Published at the Second Workshop on Agents in the Wild at ICML 2026. Arxiv preprint. Research conducted during my part-time collaboration role at Redwood Research and mentored by Tyler Tracy.
Catherine Ge-Wang(=), Joy Yang(=), Tushar Nagar, 2026, “Round-Trip Latent Geometry in Diffusion VAEs Enables Covert Channels”. Published at the Mechanistic Interpretability Workshop at ICML 2026. Research conducted independently.
- Kristina Kempkey(=), Séan Boddy(=), Catherine Ge-Wang(=), 2026, “Non-Great-Power Conflict and AI Risk”. Arxiv preprint. Research conducted during the Winter 2025 Future Impact Group Fellowship and mentored by Liam Patell at GovAI.
Non-public (yet)
- Working on reward multiplicity and goal misgeneralisation, negative results conducted in 2025 as part of the RIO fellowship, under the mentorship of Matthew Farrugia-Roberts at Oxford.
In progress
- Making and publishing an eval for misalignment continuation, mentored by Robert Kirk and Alex Souly at UK AISI as part of the ERA:AI fellowship.
- Threat modelling and foundational research for concentration of power.
News
- 09/2026: I’ve started my MIRI TGT fellowship!
- 07/2026: I’m going to be co-mentoring a SPAR project with Louis Cooper-Thomson on formalising and building AI auditors under strategic attack selection.
- 07/2026: I started working on misalignment continuation evals at the ERA:AI fellowship, mentored by Rob and Alex at AISI.
- 07/2026: I attended ICML 2026 in Seoul, Korea, where my work was published at 4 AI safety workshops.
- 07/2026: I completed my in-person week at MIRI. I really enjoyed it and I am very excited to get deeper into governance and policy!
- 03/2026: I changed my last name from “Wang” to “Ge-Wang” when my first publications came out. I wanted to honour my mother’s side of the family.
- 04/2026: I started a part-time collaboration position with Redwood Research to continue working on my attack selection paper.
