top of page
2025 summit.png

01 October 2026 | StageOne, Zurich

Knowing When Not to Answer: Bounding LLM Hallucinations in German Public Law

LLMs tend to guess even when the provided documents lack key facts. This is a major liability in high-stakes domains like German environmental and administrative law. In this talk, I will present our training framework based on the Merlin-Arthur protocol, which transforms LLM question answering into a two-player game. By training the model against both a helpful prover supplying key parts of the documents and an adversarial prover stripping away evidence, we force the model to ground its answers or explicitly abstain when evidence is missing. I will introduce our novel grounding score, which measures how much of an answer provably stems from the documents. I'll showcase how we applied this protocol to real German legal case files and considerably reduced hallucinations, shifting public sector AI from passive guesswork to verifiable reliability.

Speaker list

Aleph Alpha Research

PhD, AI Researcher

bottom of page