EPFL Lab and UNICC advance responsible AI through joint evaluation of LLM safety

Published

The collaboration delivers a practical evaluation framework that can be adapted to assess the safety and reliability of AI systems

As the global community gathers for the AI for Good Global Summit 2026, the United Nations International Computing Centre (UNICC) and the Machine Learning and Optimization Laboratory of EPFL ( École Polytechnique Fédérale de Lausanne) are pleased to announce the publication of a new white paper as a first outcome of their joint efforts to advance responsible AI. Facilitated by the International Computation and AI Network (ICAIN), the publication presents a practical framework for evaluating large language models (LLMs) to support their safe, reliable and responsible use in institutional settings.

The collaboration under ICAIN brings together UNICC’s operational expertise in digital foundations for the UN system and EPFL’s leading research capabilities in artificial intelligence, reflecting a shared commitment to advancing trustworthy AI through rigorous research, practical testing and knowledge sharing.

The newly released white paper “Safety Evaluation of an Institutional LLM-RAG Deployment: A Four-layer AI Safety Audit Framework” presents a structured approach to assessing the behaviour of Apertus, an open-source large language model developed by the Swiss AI Initiative. Rather than focusing solely on the model’s technical performance, the research examines how an AI system behaves when deployed in a real institutional environment where accuracy, safety, and reliability are essential.

“The project evaluated the system across multiple dimensions, including its resistance to harmful prompts, its ability to recognize when it should not answer, the accuracy of its responses when using external knowledge, and its susceptibility to biased or misleading interactions,” explained Associate Professor Martin Jaggi, Head of the Machine Learning and Optimization Laboratory of EPFL.

Beyond the findings presented in the white paper, a practical evaluation framework was developed that public-sector, the UN system, and other international organizations can adapt to assess the safety and reliability of their AI systems.

“This collaboration shows what international organizations and academic institutions can achieve together,” said Anusha Dandapani, Chief of the UNICC AI Hub. “UNICC brings the operational context, while EPFL brings research rigour. Together, we have evaluated open-weight models in practice rather than in theory. That is what building institutional capacity for responsible AI looks like, and the knowledge generated extends well beyond this partnership.”

This publication forms part of UNICC’s broader efforts to foster innovation through partnerships with academia, research institutions and technology communities; accelerating learning, promoting responsible innovation and strengthening the digital capabilities available to the UN system and other international organizations.

As artificial intelligence becomes increasingly integrated into public sector operations, initiatives such as this demonstrate the importance of combining scientific research with operational experience to help ensure AI technologies remain safe and trustworthy.

About the white paper

The white paper, Safety Evaluation of an Institutional LLM-RAG Deployment: A Four-Layer Audit of the UNICC Apertus System, documents the joint evaluation framework developed by UNICC and EPFL to assess the safety and robustness of institutional AI systems. In addition to presenting the findings, it introduces practical methodologies that can support future AI assurance efforts across the UN system and beyond.

Read the full white paper here: https://www.unicc.org/resources/2026/07/07/safety-evaluation-of-an-institutional-llm-rag-deployment/

Authors: Melissa Anchisi, Jean-Pierre Mora Casasola, Tanya Petersen

Share

You might be also interested in

Using future data to better predict Switzerland’s rail energy demand

In collaboration with SBB, EPFL researchers have developed an AI model that improves next-day electricity-demand forecasts across Switzerland’s rail network, reducing major prediction errors by up to 80%.

(more…)

“AI tools are cultural artifacts, not neutral software”

How to keep information trustworthy when content is increasingly machine-made? EPFL professor Andrea Cavallaro shares his view.

(more…)

Mice actively seek better views to make visual decisions

A study led by EPFL shows that when objects are difficult to see, mice don’t simply look harder. They move to find better viewpoints, adjusting their behavior according to how much visual information is available.

(more…)