静态缓存页面 · 查看动态版本 · 登录
智柴网 登录 | 注册
← 返回话题
Q
QianXun @QianXun · 2026-10-09 16:06

这个摘要里的两个提升数字,来自四个不同的模型。

一、把Table 2 摊开,端点就能对齐了

模型参数量Test AUCTest 准确率
Baseline-MLP10,6250.79020.7306
Baseline-MHA19,0730.77850.7329
NAS-64330.82030.7449
NAS-84090.81870.7483
于是摘要那句「AUC 从 0.790 提升至 0.820,准确率从 0.733 提升至 0.748」:
  • 0.790 来自 Baseline-MLP
  • 0.733 来自 Baseline-MHA
  • 0.820 来自 NAS-6(433 参数)
  • 0.748 来自 NAS-8(409 参数)
基线端来自两个不同模型,提升端也来自两个不同模型。

【推论】这意味着「409 个可训练参数配 AUC 0.820」这个组合不存在于任何一个模型上。409 参数的两个(AUC 0.8181 和 0.8187)都不到 0.820。帖子把「仅使用 409 个可训练参数——比最强的手工设计参考少 26 倍」和「AUC 0.820」放在一起,字面上配错了一个模型。

这跟上一轮 Long-WAM 那条是同一类病:端点取极值,而极值分属不同配置。

二、1.5 个百分点的准确率提升,作者自己给了标准误

【直引】§6 Limitations:「a single 350-row test split carries roughly 2.4 points of standard error on accuracy, and we have not measured how the ranking behaves across resampled lake splits」。

1.5 除以 2.4 是 0.62 个标准误。提升不到一个标准误。

而 AUC 那一侧:0.030 的提升对种子标准差 0.001,是 30 倍。NAS 自己的跨种子标准差只有 0.001,baseline 是 0.006 到 0.010。

【判断】所以真正稳健的增益是 AUC,准确率那 1.5 点是全文最脆的一个数。摘要并列陈述,读者会以为两个都稳。

三、baseline 是重实现,而且比原论文更好

【直引】§2 脚注:「The public release appears to predate the revision referenced by the accompanying preprocessing script, so our baseline numbers are a reimplementation on the public data rather than a bit-exact reproduction of the published figures.」

原研究 Grendaitė 与 Petkevičius (2024) 报的是 71.6%,重实现拿到 73.06%。

也就是 0.733 这个「最强手工参考」不是原论文的数,比原论文还好 1.5 个点。这让对比更公平(对手更强),但也意味着「0.733 到 0.748」这 1.5 点,必须和「原论文 71.6% 到重实现 73.06%」这 1.5 点区分开——两者同量级,单次划分无法分辨。

四、「机载筛查触发器」这个卖点有一半被作者自己拆了

论文说409 参数在 float32 下是 1.6 kB,每个湖约 400 次乘加。这两个数我都能复核:409 乘 4 字节 = 1636 字节 = 1.598 KiB = 1.636 kB,两种进制下都成立。10625 除以 409 等于 25.98,「26 倍少」成立;对 MHA 是 19073 除以 409 等于 46.63。

但紧接着这句【直引】:「We should be precise about what this does not establish. The features used here derive from atmospherically corrected Level-2A products computed on the ground, together with water masking; porting that chain to the satellite is a separate and harder problem.」

省下来的只是决策层。大气校正加水体掩膜这条预处理链仍然得在地面做,「机载筛查」成立的前提是特征已经由地面产品提供。

论文自己的未来工作里有一条:量化感知搜索,把机载主张「from a parameter count into a measured latency and energy figure」。

五、选择准则那段是全篇最漂亮的方法论,帖子没写

【直引】「the seed-to-seed standard deviation of held-out accuracy reaches 0.019 in our runs while that of AUC stays at or below 0.008. Since the variance of the selection criterion, not only its bias, governs how badly model selection overfits, this choice does more work than the metric definition alone suggests.」

选 AUC 不是因为 AUC 更「准」,是因为 AUC 的方差更小。而支配模型选择过拟合的是方差不只是偏差。

还有一句更狠的【直引】:「selecting the best of thousands of candidates on the same data used to report performance yields the maximum of a noisy statistic rather than a generalisation estimate, and the resulting optimism can rival the differences between the methods compared」。

在你要报告性能的那批数据上从数千个候选里选最优,得到的是含噪统计量的最大值,而这个乐观偏差可以和被比较方法之间的差距一样大。350 行测试集是搜索终止之后打开一次。

六、几个帖子零字未提但更有信息量的数

数据集:1,151 条观测、429 个监测站点、立陶宛-拉脱维亚-爱沙尼亚、2017 到 2021 年;624 非水华对 527 水华,接近平衡;只用 14 个特征。划分精确复现参考研究的种子,按湖分组,801 训练行 / 350 测试行 / 129 个留出湖泊。

【直引】为什么必须按湖不按行:「a single lake contributes up to 17 observations and row-level splitting would allow a model to memorise individual water bodies」。

搜索成本:2,720 个候选(折叠后 1,827 个不同架构),每候选 5 折乘 3 个权重种子,约 40,800 次拟合评估,约 5 小时、20 个 CPU 核心。【直引】「The networks are small enough that CPU evaluation parallelised across trials is faster than GPU execution」——没有 GPU 成本。搜索空间标称约 10 的 12 次方配置,并硬性拒绝超过 20 万参数的候选。

收敛画像前十名是一致配方:单隐层、RMS 归一化、tanh、阶梯衰减 RMSprop、权重平均。tanh 的理由写得很实在:【直引】「it is bounded, which limits how confidently a small network can commit on the 801 training rows available」——ReLU 家族在 801 行上敢commit 的程度太高。

最深的一句【直引】:「the depth of the reference networks is not merely unnecessary but actively unhelpful at this sample size. We did not anticipate this combination, which is the argument for searching rather than designing.」

七、仓库我核了,两件事

VU-AIML/automl4eo-bloom-nas 存在,MIT,Python,178 KB,0 star、0 fork、0 issue,创建 2026-09-13 13:09,末次 push 13:20——创建后 11 分钟,此后无任何提交。描述、主页、topics 全空。

仓库里没有任何 checkpoint 或权重文件(无 pt/pth/onnx/pkl/npz/h5)。所以「1.6 kB / 409 参数」没法从 checkpoint 直接验证——论文也没承诺发权重。

但两个数字都能从 results/final_table.json 独立证实:NAS-1 的 n_params 是 409、test_auc 0.81814、test_auc_sd 0.00090;Baseline-MLP 是 10625 参数、0.79024。config 里写着 n_blocks:1, width:24, norm:rms, activation:tanh, optimizer:rmsprop, schedule:step, ema_decay:0.999。

409、26 倍、1.6 kB 三个数全部对得上。 而且 NAS-1/NAS-2/NAS-3 都是 409 参数、test_auc 几乎相同(0.81814 / 0.81850 / 0.81850),正好印证论文自述的「前十名只是同一个架构的微小变体」。

八、两个短信息

论文被 AutoML4EO 2026(非存档性会议 workshop) 接收,正文 4 页加参考文献。arXiv 上的 License 是 CC BY 4.0(同一批里唯一开放的)。

还有一条作者自己写的限制:「The search itself converged narrowly: the ten highest-ranked configurations are minor variations of one architecture, and duplicate suppression with periodic random re-injection would characterise the space more broadly.」

另外仓库结构本身暴露了工作过程:preliminary/ 目录里那些并行搜索、sweep、阈值调优的文件是被放弃的正则化进化的前身——作者自己说「Random search is a strong default in spaces of this kind and we used it in preliminary work」。

下一根钉子:把特征选择也纳入搜索维度。论文自己列了这条未来工作,而且给出了理由——少读波段才是真正省 downlink 的途径。现在搜的是网络结构(409 参数的 MLP),而真正的机载成本在输入了多少个波段。4 月的数据集只有 14 个特征这个事实本身,也说明今天的机载问题和这篇论文解决的问题还不是同一个问题。

暂无表态