The Danish Foundation Models (DFM) project has introduced Mimir v1, a 1-billion-parameter language model developed with a focus on using only permissible post-training data. This approach aims to address the challenge of large language model development, which often relies on extensive and sometimes non-permissible datasets, creating barriers for researchers committed to open-source and ethically sourced data.
Mimir v1 is built on the Hierarchical Reasoning Model (HRM) architecture. The model was trained from scratch and has demonstrated highly competitive performance in English language tasks. Notably, it sets a new state-of-the-art for the Danish language. Its training utilized a mixture of 161 datasets, all of which were deemed permissible for post-training use.
In evaluations across 20 benchmarks, Mimir v1 outperformed the original HRM-Text 1B model. It also demonstrated performance comparable to larger frontier models, including Qwen 3.5 4B and Gemma 4 E2B. The DFM project, a collaboration involving the University of Southern Denmark, Ordbogen A/S, and Aarhus University, aims to develop open Danish language models that understand Danish context, culture, and societal norms.
The initiative behind DFM seeks to ensure that artificial intelligence can understand the Danish language and society while operating within controlled data frameworks. This is particularly relevant as many widely used language models are developed by global companies and trained predominantly on English, often resulting in suboptimal performance when applied to Danish. Professor Peter Schneider-Kamp from the University of Southern Denmark, who leads the DFM project, has noted that international models can "lack an understanding of how Denmark works – our literature, our public institutions, our healthcare system, and our cultural points of reference."
The DFM project is supported by national infrastructure, including the national AI supercomputing facility BITTEN and the research platform UCloud, which provide researchers with access to storage, GPUs, software, and collaboration tools. This infrastructure is crucial for the development, training, storage, and evaluation of language models.
The Mimir v1 model is available on the Hugging Face Hub, facilitating broader access and use by the research community and developers. This release is part of a larger effort by the Danish Foundation Models project to contribute to the open-source AI community with tools and datasets that extend beyond the Danish language. The project also focuses on establishing evaluation benchmarks for European AI through initiatives like EuroEval and MTEB, ensuring Danish is represented in widely used benchmarks for generative and search models.
The DFM project emphasizes transparency, reproducibility, and broad access through its open-source approach, aiming to bridge linguistic barriers and establish norms for culturally sound AI development. The collaboration includes partnerships with public and private actors, such as Aarhus Municipality, ATP, and Salling Group, to ensure the developed AI infrastructure is relevant for practical applications.
