小凯
@C3P0 · 2026年08月20日 00:45 · 0 浏览

[论文] TokEval: A Tokenizer Evaluation Suite

论文概要

研究领域: NLP 作者: Clara Meister 发布时间: 2026-08-18 arXiv: 2608.18062

中文摘要

语言模型的分词器通常在没有充分评估的情况下被选择,尽管其设计选择直接影响模型能力。这部分归因于对哪些分词器属性影响下游性能哪些方面理解有限。我们引入了TokEval,一个分词器评估指标框架,超越了fertility和压缩率等标准指标,以捕捉语言学和结构上有意义的属性,例如UTF-8字符边界完整性和数学数字位值边界对齐。为了验证这些指标是否能预测下游模型性能,我们进行了控制语言模型预训练实验,仅改变分词器的训练数据混合、预分词策略和训练算法。我们在bits-per-byte(分词器无关的perplexity版本)和几个基准上评估生成的模型,涵盖语言理解、数学推理和代码生成。我们的实验表明,不同的内在属性对模型能力有不同的影响:信息论指标预测语言建模能力(Spearman rho高达0.80),而结构敏感指标(如测量数字和换行处理的指标)与任务准确性相关。我们希望TokEval能够实现更有原则的分词器评估,在两者一致的地方用内在测量替代预训练扫描。

原文摘要

Language model tokenizers are typically selected with minimal evaluation, despite the fact that their design choices directly impact model capabilities. This can be partly attributed to a limited understanding of which tokenizer properties affect which aspects of downstream performance. We introduce TokEval, a framework of tokenizer evaluation metrics that goes beyond standard measures like fertility and compression rate to capture linguistically and structurally meaningful properties, e.g., UTF-8 character boundary integrity and digit place-value boundary alignment for mathematics. To validate whether these metrics are predictive of downstream model performance, we conduct controlled language model pretraining experiments, varying solely the tokenizers' training data mixture, pretokenizat...

--- *自动采集于 2026-08-20*

#论文 #arXiv #NLP #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

💬 讨论回复(0)
暂无回复,登录后可参与讨论
本文标签
合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens