Researcher/PhD Student (m/w/d/x): Mechanistic Interpretability for Safety
Full Job Title - DE
Researcher*in/PhD-Student*in (m/w/d/x): Mechanistische Interpretierbarkeit für sichere agentenbasierte KI
Full Job Title - EN
Researcher/PhD Student (m/w/d/x): Mechanistic Interpretability for Safe Agentic AI
Department
Multilinguality and Language Technology
Address
Deutsches Forschungszentrum für Künstliche Intelligenz GmbH (DFKI); Stuhlsatzenhausweg 3;Saarland Informatics Campus D 3_2;66123 Saarbrücken;
Email
Simon.Ostermann@dfki.de
Employment Category
Full time
Type of Contract
Temporary
Department Description - DE
3976
Department Description - EN
3977
We seek a PhD researcher at the earliest possible date to advance post-hoc controllability of language models through mechanistic interpretability. You will develop methods for fine-grained activation steering that enable safe behavior modification without retraining, a critical requirement for deploying agentic systems in safety-critical domains.
The position is embedded in a collaborative research project combining mechanistic circuit analysis with safe AI system design. Your work will directly contribute to understanding how models can be reliably steered at the neuron and sparse feature level, while preserving previously learned capabilities.
We value rigorous, methodologically grounded research: you'll be expected to design careful experiments, publish in top venues, and engage in the kind of open, critical discussion of ideas that moves the field forward. This is basic research with real-world implications, grounded in our lab's commitment to explainable and efficient language processing.
Importantly, you will be part of a fully integrated interdisciplinary team spanning both SLIMS and SAgA projects. Your mechanistic findings will directly inform safety architectures, and you'll collaborate closely with researchers across interpretability, multilingual NLP, formal ethics, and agentic systems.
The position is embedded in the Multilinguality and Language Technology (MLT) group under the direction of Dr. Marius Mosbach, with broader departmental support from Prof. Kristian Kersting and Prof. Verena Wolf; the primary supervisor of the position is Dr. Simon Ostermann; the candidate will be based in his ‘Efficient and Explainable’ NLP team.
Your tasks
- Develop and validate interpretable steering methods targeting specific computational circuits
- Analyze failure modes and robustness of activation-level interventions under distribution shift and compositional scenarios
- Design experiments quantifying orthogonality between safety interventions and task-critical representations
- Contribute to scaling steering methods from proof-of-concept to production-grade robustness
- Publish results in top-tier venues (NeurIPS, ICML, ACL, ICLR)
Your qualifications
Required Background:
- Master's degree or equivalent in computer science, NLP, machine learning, or related field
- Strong foundation in deep learning and neural network architectures
- Experience with mechanistic interpretability methods (circuit analysis, SAEs, activation patching, or similar)
- Proficiency in Python and PyTorch or equivalent frameworks
- Ability to work independently and in collaborative teams
Desirable Qualifications:
- Prior experience with model steering, causal intervention, or mechanistic analysis
- Familiarity with interpretability tooling (e.g., TransformerLens, Anthropic's SAE library)
- Track record of publications or strong problem-solving demonstrated in prior work
- Interest in AI safety and alignment research
Your benefits
- We offer competitive, market-based compensation and many other benefits (Urban Sports Club, corporate benefits, a subsidy for your Jobticket, and much more)
- Access to GPU clusters and computational resources
- A supportive, intellectually rigorous research environment where critical discussion and methodology matter
- Space to develop your own scientific identity and pursue ideas with genuine independence
- Active publication culture with strong support for conference attendance and dissemination
- Integration into the broader mechanistic interpretability research community
- Collaborative, respectful lab culture that values both scientific rigor and human well-being
The German Research Center for Artificial Intelligence (DFKI) has operated as a non-profit, Public-Private-Partnership (PPP) since 1988. DFKI combines scientific excellence and commercially-oriented value creation with social awareness and is recognized as a major "Center of Excellence" by the international scientific community. In the field of artificial intelligence, DFKI has focused on the goal of human-centric AI for more than 35 years. Research is committed to essential, future-oriented areas of application and socially relevant topics.
DFKI encourages applications from people with disability; DFKI intends to increase the proportion of female employees in the field of science and encourages women to apply for this position.