Marek Masiak

DPhil (PhD) student in Fundamentals of AI @ University of Oxford

Marek Masiak

My research focuses on mechanistic interpretability, AI Safety, and the training dynamics of neural networks, increasingly through the lens of developmental interpretability.

Research

I believe AI has the potential to do an enormous amount of good by accelerating scientific discovery and that interpretability will play a major role in this. I’m worried about risks from deliberate misuse of AI, loss of control, biorisks, and gradual disempowerment.

My path to forming these beliefs was not the most usual one. I was first introduced to mechanistic interpretability by Constantin Venhoff when I was looking for an interesting research problem, having previously worked on NLP and machine translation for low resource African languages, and Bayesian Machine Learning. From there I discovered AI Safety and the Effective Altruism community.

Right now I split my time between two projects. The first is joint work with Andrzej Szablewski, who I also wrote Activation Transport Operators with. This project studies how an interpretable-by-design fine-tuning method can guard against dataset poisoning attacks and backdoor sleeper agents. It is grant-funded through the LASR Labs extension programme, hosted by Arcadia Impact and funded by Coefficient Giving’s AISTOF fund.

The second explores how the geometry of the loss landscape can characterise emergent misalignment, and what geometric inductive biases can protect against it during training, supported by a Rapid Grant from BlueDot Impact.

News

Secured a Rapid Grant from BlueDot Impact for my research on loss landscape geometry for analysing emergent misalignment

Secured ~£14k with Andrzej Szablewski from Coefficient Giving's AISTOF fund for further research on the interpretable-by-design TopKLoRA fine-tuning method

If anything on this website sounds interesting to you, we share common interests, or you just wanna talk, I would love to hear from you! Please email me