[论文] Weight Pair Encoding: Inducing a Smaller Grammar in Neural Network Wei...

研究领域: ML 作者: Irene Tallini, Daniele Solombrino, Alberto Cazzaniga, Emanuele Rodolà 发布时间: 2026-09-25 arXiv: 2609.31564

论文概要

研究领域: ML 作者: Irene Tallini, Daniele Solombrino, Alberto Cazzaniga, Emanuele Rodolà 发布时间: 2026-09-25 arXiv: 2609.31564

中文摘要

我们证明,神经网络权重可以被显式微调以“承认”一个更小的文法。Weight Pair Encoding(WeightPE)通过把一个有损的 Re-Pair 压缩器放进直通估计器(straight-through estimator)来实现这一点。网络的 int8 权重被展平成一个字符串,近似的 Re-Pair 模式在一个全局 L2 预算内被改为完全相等。网络用重写后的权重计算,并通过直通估计器穿过它们进行训练。与固定条目大小的扁平码本不同,文法提供可变长度的模式,并在更大模式内部层次化地复用它们。在 CIFAR-10 上微调的 ViT-B/16 与 ViT-L/16 的 MLP 权重上,WeightPE 产生的 Re-Pair 文法大小分别为同等 int8 QAT 运行所得文法的 0.43 倍和 0.38 倍,代价是 1.9 和 1.1 个准确率点。这一趋势还延伸到不同的文法压缩器(LZ78、SEQUITUR)——网络并未针对它们微调。据我们所知,这是文法大小首次被用作网络权重的显式训练目标。

原文摘要

We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one string, and near-matching Re-Pair patterns are made exactly equal within a global L2 budget. The network computes with the rewritten weights and trains through them with a straight-through estimator. Unlike a flat codebook of fixed-size entries, a grammar offers variable-length patterns and reuses them hierarchically inside larger ones. On the MLP weights of ViT-B/16 and ViT-L/16 finetuned on CIFAR-10, WeightPE produces a Re-Pair grammar 0.43x and 0.38x the size of the one produced by an equivalent int8 QAT run, at a cost of 1.9 and ...


*自动采集于 2026-09-29*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens