top of page
2025 summit.png

Staging × Misalignment Science: When, Where and How Safety Enters LLM Training

Every language model is expected to be capable, fluent and safe. But where in the training pipeline is safety actually decided, and with which methods? This session explores that question across the full life of a model: data filtering and curation, pre-training, post-training, evaluation, mechanistic interpretability, deployment, and runtime security; from both academic and industry perspectives.

 

The session in three parts:

  • Keynote(s). We open with a keynote that sets the scientific context for staging safety in the pipeline, with a second keynote to be announced.

  • Panel. A moderated panel brings together researchers who take deliberately opposing positions on when, where and how safety should be built in. Panelists span academia, non-profit research, a national safety institute and industry, and several countries, so the audience hears genuinely different research traditions argue the same question.

  • Mingling. A closing slot to continue the discussion informally.

What you get from attending:

  • A clear map of the design space: which safety decisions belong at which modeling stage.

  • A feel for which methods make safety durable, and where the evidence is strong or weak.

  • A front-row view of where leading researchers genuinely disagree.

Who it is for: researchers and practitioners in alignment, interpretability, evaluation and AI security, and anyone shaping how frontier models are trained and released.

Speaker list

EleutherAI

Executive Director

Blindsight

Ethical Hacker

ETHZ/EPFL

PhD

UK AISI

Research Scientist

ETH Zurich

PhD

ETH Zurich

Postdoc

EPFL

Associate Professor

bottom of page