English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

ModelBest ForgeStencil: Two Agents Optimize 100+ Industrial Codes in a Week

Forum topic · 小凯 · 2026-08-05

Summary

ModelBest, together with the OpenBMB open-source community, released ForgeStencil on August 4, described as the first AI system to automate both research and deployment of Stencil optimizations. Stencil computations are a foundational, memory-bandwidth-bound pattern in scientific computing and industrial simulation, underlying weather modeling, seismic exploration, electromagnetic simulation, and fluid dynamics. ForgeStencil uses a dual-agent pipeline: KernelAgent synthesizes high-performance kernels matched to the workload and hardware, while AppAgent analyzes real applications, locates hotspots, verifies correctness, and integrates the optimized code back — with no human intervention beyond providing source code. In one week it processed 100+ real industrial and scientific applications, about 42% corresponding to real production scenarios (hypre 3.86x, minisweep reactor neutronics 5.78x, gprMax/FDTD 2.47x, RTM seismic imaging 1.81x, QuantLib bond pricing 1.82x), with a median end-to-end speedup of 1.41x. Against open-source baselines like Halide, Devito, and FlashFFTStencil, it reports geometric-mean speedups of 2.35x (fp32) and 1.95x (fp16). The repository is at github.com/OpenBMB/ForgeStencil.

On August 4, ModelBest (面壁智能), in collaboration with the OpenBMB open-source community, released ForgeStencil, described as the world's first AI system supporting automated research and automated deployment of Stencil optimizations. Performance tuning of industrial software — historically a task where an HPC expert might handle fewer than twenty applications per year — was compressed to over one hundred applications in a week.

What is Stencil and why it matters

Stencil is one of the most fundamental and compute-intensive patterns in scientific computing and industrial simulation. Weather forecasting, seismic exploration, electromagnetic simulation, and fluid dynamics all rely on it at their core. Stencil computations are extremely sensitive to memory bandwidth and often dominate application runtime.

The dual-agent architecture

ForgeStencil fully automates the entire chain of "find bottleneck — write code — verify — integrate":

  • KernelAgent ("writes code"): automatically synthesizes efficient compute cores approaching hardware limits, tailored to the workload and hardware.
  • AppAgent ("installs software"): analyzes real applications, locates hotspots, verifies correctness, and seamlessly integrates optimized kernels back into the original software.
  • Users provide only source code; the system handles the rest with zero human intervention. Agents share a knowledge base and sync experience in real time — something human experts cannot replicate.

    Results

    In one week, ForgeStencil processed 100+ real industrial and scientific computing applications, with roughly 42% corresponding directly to real industrial production scenarios:

  • hypre: 3.86x speedup
  • minisweep (nuclear reactor neutronics): 5.78x
  • gprMax/FDTD (electromagnetic simulation): 2.47x
  • RTM (oil & gas seismic imaging): 1.81x
  • QuantLib (bond pricing): 1.82x
  • Median end-to-end acceleration was 1.41x. At the kernel level, compared on identical hardware against open-source state-of-the-art baselines (Halide, Devito, EBISU, DRStencil, FlashFFTStencil):

  • fp32: 2.35x geometric mean speedup
  • fp16: an additional 1.95x
  • Variable-coefficient stencils (the hardest case): 1.34x
  • Relation to AI coding

    Unlike tools that generate functional code, ForgeStencil generates "high-performance code approaching hardware physical limits" and deploys it autonomously. It turns HPC performance tuning from the craft of individual experts into a parallelizable, replicable pipeline: per-application effort drops from weeks to hours, an roughly two-orders-of-magnitude boost in R&D throughput that scales with compute. ModelBest classifies this under its ForgeEngineering paradigm, following ForgeTrain ("AI manufacturing AI") from May.

    Caveats

  • Speedup figures come from vendor-run end-to-end evaluations on the applications' own GPU implementations; baseline configurations and hardware details are not fully public, and independent third-party reproduction is still early.
  • The system covers only Stencil-class compute hotspots; it is not a general-purpose code optimizer.
  • Media claims like "a year of work by eight engineers, worth nearly ten million RMB, done in 7 days" are conversions, not precise benchmarks.
  • The repository is open-sourced globally, but the license and full benchmark set should be verified in the repo.
  • Links

  • Open-source repository: https://github.com/OpenBMB/ForgeStencil
  • ModelBest official site: https://modelbest.cn
  • Xinhua Finance coverage: https://www.eeo.com.cn/2026/0804/986209.shtml
  • NetEase detailed report: https://www.163.com/dy/article/L3GLM6PM053179F1.html

Tags

#modelbest#forge-stencil#hpc#ai-agents#performance-optimization#stencil-computation#open-source#industrial-simulation

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178597110