论文概要
研究领域: NLP
作者: Clara Meister
发布时间: 2026-08-18
arXiv: 2608.18062
中文摘要
语言模型的分词器通常在没有充分评估的情况下被选择,尽管其设计选择直接影响模型能力。这部分归因于对哪些分词器属性影响下游性能哪些方面理解有限。我们引入了TokEval,一个分词器评估指标框架,超越了fertility和压缩率等标准指标,以捕捉语言学和结构上有意义的属性,例如UTF-8字符边界完整性和数学数字位值边界对齐。为了验证这些指标是否能预测下游模型性能,我们进行了控制语言模型预训练实验,仅改变分词器的训练数据混合、预分词策略和训练算法。我们在bits-per-byte(分词器无关的perplexity版本)和几个基准上评估生成的模型,涵盖语言理解、数学推理和代码生成。我们的实验表明,不同的内在属性对模型能力有不同的影响:信息论指标预测语言建模能力(Spearman rho高达0.80),而结构敏感指标(如测量数字和换行处理的指标)与任务准确性相关。我们希望TokEval能够实现更有原则的分词器评估,在两者一致的地方用内在测量替代预训练扫描。
原文摘要
Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which tokenizer properties affect which aspects of downstream performance. We introduce TokEval, a framework of tokenizer evaluation metrics that goes beyond standard measures like fertility and compression rate to capture linguistically and structurally meaningful properties, e.g., UTF-8 character boundary integrity and digit place-value boundary alignment for mathematics. To validate whether these metrics are predictive of downstream model performance, we conduct controlled language model pretraining experiments, varying solely the tokenizers' training data mixture, pretokenizat...
自动采集于 2026-08-20
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。