Secured a Rapid Grant from BlueDot Impact for my research on loss landscape geometry for analysing emergent misalignment
Marek Masiak
DPhil (PhD) student in Fundamentals of AI @ University of Oxford

My research focuses on mechanistic interpretability, AI Safety, and the training dynamics of neural networks, increasingly through the lens of developmental interpretability.
Research
I believe AI has the potential to do an enormous amount of good by accelerating scientific discovery and that interpretability will play a major role in this. I’m worried about risks from deliberate misuse of AI, loss of control, biorisks, and gradual disempowerment.
My path to forming these beliefs was not the most usual one. I was first introduced to mechanistic interpretability by Constantin Venhoff when I was looking for an interesting research problem, having previously worked on NLP and machine translation for low resource African languages, and Bayesian Machine Learning. From there I discovered AI Safety and the Effective Altruism community.
Right now I split my time between two projects. The first is joint work with Andrzej Szablewski, who I also wrote Activation Transport Operators with. This project studies how an interpretable-by-design fine-tuning method can guard against dataset poisoning attacks and backdoor sleeper agents. It is grant-funded through the LASR Labs extension programme, hosted by Arcadia Impact and funded by Coefficient Giving’s AISTOF fund.
The second explores how the geometry of the loss landscape can characterise emergent misalignment, and what geometric inductive biases can protect against it during training, supported by a Rapid Grant from BlueDot Impact.
News
Secured ~£14k with Andrzej Szablewski from Coefficient Giving's AISTOF fund for further research on the interpretable-by-design TopKLoRA fine-tuning method
Reviewer — Sci-FM: Scientific Understanding of Foundation Models @ COLM 2026
Reviewer — Mechanistic Interpretability Workshop @ ICML 2026
Spotlight — Activation Transport Operators @ NeurIPS Mechanistic Interpretability Workshop
TopKLoRA accepted @ NeurIPS Mechanistic Interpretability Workshop
Started my DPhil in the Fundamentals of AI @ Oxford
If anything on this website sounds interesting to you, we share common interests, or you just wanna talk, I would love to hear from you! Please email me