Hello

Hi! I’m Cath Ge-Wang, a mathematics undergraduate at Christ Church, University of Oxford. I’m currently working on building misalignment continuation evals with my mentors at UK AISI, will be working on verification protocols at MIRI, and previously was a part-time research collaborator at Redwood Research. I run the Oxford AI Safety Initiative’s Policy Team.

My primary research interests lie in AI control and agent foundations, particularly understanding and mitigating emergent misalignment risks in autonomous AI systems. I focus on empirical questions around goal misgeneralisation, alignment faking, and attack selection in agentic evaluations, aiming to clarify failure modes in frontier models. I am also interested in how these technical insights inform AI governance and policy, especially hardware verification, mechanisms for strategic risk, and constraining dangerous capability deployment.

My Research

Published:

  1. Catherine Ge-Wang, Tyler Crosse, Benjamin Hadad IV, Joachim Schaeffer, Ram Potham, and Tyler Tracy, 2026, “Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety”. Published at the Second Workshop on Agents in the Wild at ICML 2026. Arxiv preprint. Research conducted during my part-time collaboration role at Redwood Research.

  2. Catherine Ge-Wang, Joy Yang, Tushar Nagar, “Round-Trip Latent Geometry in Diffusion VAEs Enables Covert Channels”. Published at the Mechanistic Interpretability Workshop at ICML 2026. Research conducted independently.

    Non-public (yet)

  3. Non-Great-Power Conflict and AI Risk, research conducted in 2025 as part of the FIG fellowship, under the mentorship of Liam Patell.
  4. Working on reward multiplicity and goal misgeneralisation, negative results conducted in 2025 as part of the RIO fellowship, under the mentorship of Matthew Farrugia-Roberts.

In progress

  1. Making and publishing the first misalignment continuation eval, mentored by Robert Kirk and Alex Souly at UK AISI as part of the ERA:AI fellowship.
  2. Threat modelling and foundational research for concentration of power.

News

  • 07/2026: I’m going to be co-mentoring a SPAR project with Louis Cooper-Thomson on formalising and building AI auditors under strategic attack selection.
  • 07/2026: I started working on misalignment continuation evals at the ERA:AI fellowship, mentored by Rob and Alex at AISI.
  • 07/2026: I attended ICML 2026 in Seoul, Korea, where my work was published at 4 AI safety workshops.
  • 07/2026: I completed my in-person week at MIRI. I really enjoyed it and I am very excited to get deeper into governance and policy!
  • 03/2026: I changed my last name from “Wang” to “Ge-Wang” when my first publications came out. I wanted to honour my mother’s side of the family.
  • 04/2026: I started a part-time collaboration position with Redwood Research to continue working on my attack selection paper.