论文概要
研究领域: NLP
作者: Peter Schneider-Kamp, Jacob Nielsen, Gianluca Barmina, Kenneth Enevoldsen, Lukas Galke Poech
发布时间: 2026-08-13
arXiv: 2608.13517
中文摘要
当前大型语言模型开发依赖海量、通常不可自由使用的数据集,为致力于开源和道德来源数据的研究人员创造了高门槛。我们介绍Mimir v1,一个基于分层推理模型(HRM)架构的10亿参数语言模型,从头训练,仅使用可自由使用的后训练数据,为英语提供极具竞争力的性能,并为丹麦语树立新的最先进水平。在161个数据集的混合上训练,Mimir v1超越原始HRM-Text 1B,并与更大的前沿模型如Qwen 3.5 4B和Gemma 4 E2B竞争,在20个英语、数学与代码和丹麦语基准上测试。该模型可在Hugging Face Hub获取:https://huggingface.co/danish-foundation-models/DFM-Mimir
原文摘要
Current large language model development relies on massive, often non-permissible datasets, creating a high barrier for researchers committed to open-source and ethically sourced data. We introduce Mimir v1, a 1-billion-parameter language model based on the Hierarchical Reasoning Model (HRM) architecture, that is trained from scratch and delivers highly competitive performance for English and sets a new state of the art for Danish using only permissible post-training data. Trained on a mixture of 161 datasets, Mimir v1 outperforms the original HRM-Text 1B and competes with larger frontier models like Qwen 3.5 4B and Gemma 4 E2B, tested across 20 benchmarks for English, Math & Code and Danish. The model is available on the Hugging Face Hub: https://huggingface.co/danish-foundation-models/DFM-Mimir
自动采集于 2026-08-15
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。