论文概要
研究领域: ML
作者: Irene Tallini, Daniele Solombrino, Alberto Cazzaniga, Emanuele Rodolà
发布时间: 2026-09-25
arXiv: 2609.31564
中文摘要
我们证明,神经网络权重可以被显式微调以“承认”一个更小的文法。Weight Pair Encoding(WeightPE)通过把一个有损的 Re-Pair 压缩器放进直通估计器(straight-through estimator)来实现这一点。网络的 int8 权重被展平成一个字符串,近似的 Re-Pair 模式在一个全局 L2 预算内被改为完全相等。网络用重写后的权重计算,并通过直通估计器穿过它们进行训练。与固定条目大小的扁平码本不同,文法提供可变长度的模式,并在更大模式内部层次化地复用它们。在 CIFAR-10 上微调的 ViT-B/16 与 ViT-L/16 的 MLP 权重上,WeightPE 产生的 Re-Pair 文法大小分别为同等 int8 QAT 运行所得文法的 0.43 倍和 0.38 倍,代价是 1.9 和 1.1 个准确率点。这一趋势还延伸到不同的文法压缩器(LZ78、SEQUITUR)——网络并未针对它们微调。据我们所知,这是文法大小首次被用作网络权重的显式训练目标。
原文摘要
We show that neural network weights can be explicilty fintuned to admit a smaller grammar. Weight Pair Encoding (WeightPE) does so by placing a lossy Re-Pair compressor inside a straight-through estimator. The int8 weights of the network are flattened into one string, and near-matching Re-Pair patterns are made exactly equal within a global L2 budget. The network computes with the rewritten weights and trains through them with a straight-through estimator. Unlike a flat codebook of fixed-size entries, a grammar offers variable-length patterns and reuses them hierarchically inside larger ones. On the MLP weights of ViT-B/16 and ViT-L/16 finetuned on CIFAR-10, WeightPE produces a Re-Pair grammar 0.43x and 0.38x the size of the one produced by an equivalent int8 QAT run, at a cost of 1.9 and ...
自动采集于 2026-09-29
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。