Senior Researcher/PostDoc (m/w/d/x): AI Safety, Ethics, and Agentic Systems
Full Job Title - DE
Wissenschaftliche*r Mitarbeiter*In/Postdoktorand (m/w/d/x): KI-Sicherheit, Ethik und agentische Systeme
Full Job Title - EN
Senior Researcher/PostDoc (m/w/d/x): AI Safety, Ethics, and Agentic Systems
Department
Neuro-mechanistic Modeling
Location
Saarbrücken,
Darmstadt
Employment Category
Full time
Type of Contract
Temporary
Department Description - DE
4210
Department Description - EN
4211
We seek a postdoctoral researcher at the earliest possible date for a new project on safe, controllable agentic AI for regulated and safety-critical domains. The project's premise is that safety constraints should be explicit, inspectable, and updatable after deployment, rather than compressed into model weights during training and patched afterwards with output filters. Work proceeds along three strands: post-deployment behavioural control through interventions on model internals; formally modeling the constraining norms in separate layers to keep legal, organizational, ethical, and preference-based dimensions distinct; and mechanisms that keep autonomously generated agent behaviour within those constraints.
This position sits across all three strands, making this a genuine research role rather than a coordinating one for two reasons:
First, the project sits on the question whether what the normative specification permits, what runtime control actually enforces, and what the agent may autonomously take on remain consistent with one another. Stating that property precisely, and establishing what can and cannot be guaranteed about it, requires knwoeldge across symbolic specification, model-internal intervention, and autonomous task generation alike. This question is, to our knowledge, largely open.
The second reason is that the claim has to hold in a running system. You will be responsible for bringing the three strands together into an integrated, working implementation, for designing the interfaces they commit to, and for evaluating the result end to end. The project's doctoral researchers work in depth within single strands; this PostDoc role spans and orchestrates them, with corresponding authority over the interfaces between them.
We are committed to rigorous, methodologically grounded research where ideas are tested critically and hypotheses — whether confirmed or refuted — are published and discussed openly. Our lab culture values intellectual integrity, supportive collaboration, and space for researchers to develop their own scientific identity.
The position spans three DFKI research groups: Neuro-Mechanistic Modeling (NMM, Prof. Verena Wolf), Foundations of Systems AI (SAINT, Prof. Kristian Kersting), and Multilinguality and Language Technology (MLT, Dr. Marius Mosbach), reflecting the project's scope across formal ethics and normative reasoning, agentic system design, and mechanistic interpretability. The main superviors for the position are Dr. Kevin Baum, Dr. Dominik Hintersdorf, and Dr. Simon Ostermann.
Your tasks
- Formalize and investigate consistency between normative specification, runtime behavioural control, and autonomous capability growth
- Design and maintain the interfaces connecting the project's three research strands, including the translation of formally specified constraints into representations usable by planning and learning components
- Build and maintain an integrated reference platform and evaluation harness in simulated deployment environments
- Lead end-to-end evaluation of the integrated system against the benchmark suites developed across the project
- Prepare versioned open-source releases with documentation and working examples sufficient for independent replication
- Publish in AI safety, formal methods, and agentic AI venues
Your qualifications
Required Background:
- PhD in computer science, AI, formal methods, mathematics, or a related field
- Substantial software engineering experience and the ability to own and maintain a non-trivial codebase over years. This is a hard requirement; prior industry development experience is explicitly welcome
- Strong foundation in at least one of: formal specification and verification; logic (temporal, deontic, defeasible, or argumentation-based); AI safety and alignment; mechanistic interpretability
- Willingness to work across all of the above. We do not expect expertise in every area on arrival, but the role requires becoming conversant enough in each to specify interfaces that hold
- Proficiency in Python
- Track record of publications or demonstrated research impact
Desirable Qualifications:
- Experience with temporal logic, reward machines, shielding, or safe reinforcement learning
- Background in reinforcement learning, agentic systems, or agent-based modeling
- Familiarity with deep learning frameworks, and with activation steering, sparse feature analysis, or related interpretability methods
- Experience with deontic or defeasible normative logic, or argumentation frameworks
- Interest in explainability, interpretability, and human oversight mechanisms
- Background in safety-critical or regulated domains
- Understanding of EU AI Act compliance requirements and regulatory sandboxes
Your benefits
- We offer competitive, market-based compensation and many other benefits (Urban Sports Club, corporate benefits, a subsidy for your Jobticket, and much more)
- Significant computational resources for model development and evaluation
- A research question of your own alongside a system you own, rather than distributed support work across other people's agendas
- Co-supervision drawing on complementary expertise across the three groups, with one clear reporting line
- Integration into German and European AI safety, interpretability, and AI governance networks, including collaboration with legal scholars and connections to regulatory practice
- Opportunity to co-supervise/advise PhD and Master's students
- Active publication culture with strong support for dissemination and open-science practices
- A supportive, respectful environment committed to human well-being alongside scientific excellence
The German Research Center for Artificial Intelligence (DFKI) has operated as a non-profit, Public-Private-Partnership (PPP) since 1988. DFKI combines scientific excellence and commercially-oriented value creation with social awareness and is recognized as a major "Center of Excellence" by the international scientific community. In the field of artificial intelligence, DFKI has focused on the goal of human-centric AI for more than 35 years. Research is committed to essential, future-oriented areas of application and socially relevant topics.
DFKI encourages applications from people with disability; DFKI intends to increase the proportion of female employees in the field of science and encourages women to apply for this position.