Overview
Field: NLP Authors: Peter Schneider-Kamp, Jacob Nielsen, Gianluca Barmina, Kenneth Enevoldsen, Lukas Galke Poech Published: 2026-08-13 arXiv: 2608.13517
Abstract
Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data. We introduce Mimir v1, a 1-billion-parameter language model based on the Hierarchical Reasoning Model (HRM) architecture, that is trained from scratch and delivers highly competitive performance for English and sets a new state of the art for Danish using only permissible post-training data. Trained on a mixture of 161 datasets, Mimir v1 outperforms the original HRM-Text 1B and competes with larger frontier models like Qwen 3.5 4B and Gemma 4 E2B, tested across 20 benchmarks for English, Math & Code and Danish.
Key highlights
- 1B parameters, HRM architecture: trained from scratch on the Hierarchical Reasoning Model architecture.
- Fully permissible data: post-training uses only freely usable, ethically sourced data across a mixture of 161 datasets.
- Strong results: outperforms HRM-Text 1B and is competitive with larger frontier models (Qwen 3.5 4B, Gemma 4 E2B) on 20 English, Math & Code, and Danish benchmarks.
- New Danish SOTA: sets a new state of the art for Danish.
- Open availability: model released on the Hugging Face Hub: https://huggingface.co/danish-foundation-models/DFM-Mimir
- Paper: https://arxiv.org/abs/2608.13517
- Model: https://huggingface.co/danish-foundation-models/DFM-Mimir