Facewall and OpenBMB open-source ForgeStencil, an AI system that optimizes 100+ real HPC apps in a week
面壁智能 · wechat · 2026-08-04
ForgeStencil open-sources a dual-agent system for automatic Stencil optimization
Facewall Intelligence and OpenBMB say ForgeStencil is the first AI system that can automatically research, optimize, and deploy Stencil workloads end to end. Users provide source code, and the system handles hotspot analysis, kernel synthesis, replacement, correctness checks, and reintegration into the original application.
Reported results
- 100+ real industrial and scientific applications optimized in one week
- 2.35× geometric-mean speedup over leading open-source baselines for fp32 kernels
- 1.95× additional speedup for mixed-precision fp16 read/write workloads
- 1.34× geometric-mean gain on difficult variable-coefficient Stencil shapes
How it works
- KernelAgent: searches for high-performance kernels and builds a reusable operator knowledge base
- AppAgent: identifies hotspots in real applications, benchmarks end-to-end performance, validates correctness, and integrates optimized kernels back into the app
Real-world examples mentioned
- hypre multigrid solver: 3.86×
- minisweep transport mini-app: 5.78×
- gprMax/FDTD electromagnetic simulation: 2.47×
- RTM seismic imaging: 1.81×
- QuantLib bond pricing: 1.82×
- MRI reconstruction: 2.45×
- Digital breast tomosynthesis backprojection: 1.63×
The post argues this moves Stencil optimization from expert-driven tuning to scalable AI-assisted engineering, especially for industrial software, scientific computing, and “industrial upgrading” workflows.
Related event: OpenBMB Releases ForgeStencil for Automated Code Optimization(2 posts)→
More from Infra
- Alibaba’s H+ Embedding cuts retrieval vectors by 13.7% while matching token-level quality — _reachsumit · 2026-08-04
- IBM’s Hierarchical BM25 serves 1B documents in 4.4GB and about 300ms per query — _reachsumit · 2026-08-04
- micro1 pitches Flow to turn defense video archives into training data — Exp_Mark · 2026-08-04
- Meta’s GRACE cuts generative ads decoder latency by 11× — _reachsumit · 2026-08-04
- JLC Launches a 10.2 billion yuan IPO after turning PCB prototyping into a hardware manufacturing platform — 创业邦 · 2026-08-04
- MiniMax video generation is still 10 to 15 times slower on local H100s than via API — haremlifegame · 2026-08-04