Summary
A new arXiv paper (2607.02499) by Gil Harari, Yoel Zimmermann, Ola Tangen Kulseng, Laura Zichi, Chuin Wei Tan, Marc L. Descoteaux, and Boris Kozinsky implements and systematically compares matrix-structured optimizers—Muon, SOAP, and a hybrid SOAP-Muon—for training two leading machine learning interatomic potentials (MLIPs), NequIP and Allegro. The study shows that these optimizers substantially outperform Adam (AdamW) in both convergence speed and final accuracy. The gains are especially pronounced under partial force supervision, where force labels are limited or incomplete, indicating that matrix-structured optimizers enable more label-efficient training of MLIPs. This work is relevant to computational chemistry, materials simulation, and machine learning researchers seeking faster training pipelines for atomistic models, as it demonstrates that optimizer choice—beyond architecture and data—can be a key lever for improving MLIP training efficiency and model quality.
Paper Overview
Research areas: cs.LG, cs.AI, physics.chem-ph, physics.comp-ph
Authors: Gil Harari, Yoel Zimmermann, Ola Tangen Kulseng, Laura Zichi, Chuin Wei Tan, Marc L. Descoteaux, Boris Kozinsky
arXiv: 2607.02499
Abstract
We implement and systematically compare matrix-structured optimizers including Muon, SOAP, and SOAP-Muon for training NequIP and Allegro MLIP models. These optimizers substantially outperform Adam in convergence speed and final accuracy, particularly under partial force supervision.
Key Takeaways
- Matrix-structured optimizers (Muon, SOAP, SOAP-Muon) were implemented and benchmarked for training the NequIP and Allegro machine learning interatomic potentials.
- All three optimizers substantially outperform Adam in both convergence speed and final accuracy.
- The advantage is especially significant under partial force supervision, suggesting more label-efficient training when force labels are scarce.
- The results indicate that optimizer choice is an important, often overlooked, lever for improving MLIP training beyond architecture and dataset design.
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/178634141