On August 4, ModelBest (面壁智能), in collaboration with the OpenBMB open-source community, released ForgeStencil, described as the world's first AI system supporting automated research and automated deployment of Stencil optimizations. Performance tuning of industrial software — historically a task where an HPC expert might handle fewer than twenty applications per year — was compressed to over one hundred applications in a week.
What is Stencil and why it matters
Stencil is one of the most fundamental and compute-intensive patterns in scientific computing and industrial simulation. Weather forecasting, seismic exploration, electromagnetic simulation, and fluid dynamics all rely on it at their core. Stencil computations are extremely sensitive to memory bandwidth and often dominate application runtime.
The dual-agent architecture
ForgeStencil fully automates the entire chain of "find bottleneck — write code — verify — integrate":
- KernelAgent ("writes code"): automatically synthesizes efficient compute cores approaching hardware limits, tailored to the workload and hardware.
- AppAgent ("installs software"): analyzes real applications, locates hotspots, verifies correctness, and seamlessly integrates optimized kernels back into the original software.
- hypre: 3.86x speedup
- minisweep (nuclear reactor neutronics): 5.78x
- gprMax/FDTD (electromagnetic simulation): 2.47x
- RTM (oil & gas seismic imaging): 1.81x
- QuantLib (bond pricing): 1.82x
- fp32: 2.35x geometric mean speedup
- fp16: an additional 1.95x
- Variable-coefficient stencils (the hardest case): 1.34x
- Speedup figures come from vendor-run end-to-end evaluations on the applications' own GPU implementations; baseline configurations and hardware details are not fully public, and independent third-party reproduction is still early.
- The system covers only Stencil-class compute hotspots; it is not a general-purpose code optimizer.
- Media claims like "a year of work by eight engineers, worth nearly ten million RMB, done in 7 days" are conversions, not precise benchmarks.
- The repository is open-sourced globally, but the license and full benchmark set should be verified in the repo.
- Open-source repository: https://github.com/OpenBMB/ForgeStencil
- ModelBest official site: https://modelbest.cn
- Xinhua Finance coverage: https://www.eeo.com.cn/2026/0804/986209.shtml
- NetEase detailed report: https://www.163.com/dy/article/L3GLM6PM053179F1.html
Users provide only source code; the system handles the rest with zero human intervention. Agents share a knowledge base and sync experience in real time — something human experts cannot replicate.
Results
In one week, ForgeStencil processed 100+ real industrial and scientific computing applications, with roughly 42% corresponding directly to real industrial production scenarios:
Median end-to-end acceleration was 1.41x. At the kernel level, compared on identical hardware against open-source state-of-the-art baselines (Halide, Devito, EBISU, DRStencil, FlashFFTStencil):
Relation to AI coding
Unlike tools that generate functional code, ForgeStencil generates "high-performance code approaching hardware physical limits" and deploys it autonomously. It turns HPC performance tuning from the craft of individual experts into a parallelizable, replicable pipeline: per-application effort drops from weeks to hours, an roughly two-orders-of-magnitude boost in R&D throughput that scales with compute. ModelBest classifies this under its ForgeEngineering paradigm, following ForgeTrain ("AI manufacturing AI") from May.