2026-08-27
An SiV- in inverse-designed GaP-on-diamond hits r≈1.14 at 0.24 nW/μm². Nonlinear ONNs reach 96.72% MNIST vs 88.41% linear; LLM-scale nonlinearity-limited power stays below 1.2 W.
Optical neural networks make matrix multiplies cheap. Nonlinearity is the bottleneck. In bulk materials the optical response is perturbative, so all-optical activations typically want mW/μm² intensities and large footprints. Hybrid optoelectronic activations restore nonlinearity at the cost of latency and O/E/O complexity. "Structural nonlinearity" encodes the input in tunable parameters of a linear device, which maps poorly onto a standard deep net.
Zhou, Kim, Tao and colleagues at Wisconsin-Madison, with Ming Zhou at Stanford and corresponding author Zongfu Yu, take a quantum route. A single emitter saturates after a few photons per lifetime, so the nonlinearity is far stronger than in bulk media. The open question is how much that buys in network expressivity, and at what intensity.
The activation embeds one SiV⁻ color center in an adjoint-optimized GaP-on-diamond structure with a 1.5×0.7 μm² design region. The emitter is a two-level system with Γ₀=2π×94 MHz under a lifetime-limited linewidth. Weak fields see a linear scatterer; strong fields see a nearly transparent object. A two-port interference geometry is set for |Δt|=1. Three-dimensional nonlinear FDFD supplies the input-output curve that becomes the effective activation.
Training is physics-aware. The activation curve is frozen from full-wave simulation. Linear blocks are complex transmission matrices under energy-preserving constraints, trained in PyTorch, then realized as compact photonic blocks by a second adjoint solve. Classification is rechecked with full-wave nonlinear FDFD; the Pong study uses 2D simulation for compute reasons. Expressivity is the per-layer growth factor r of a data-manifold's total curvature. Digital MLP activations typically sit at r≈1.045-1.095 per layer.
At I=0.24 nW/μm² the quantum activation reaches r≈1.14, above the digital baseline. Matching that baseline takes more than 72.6 W/μm² in a 50 μm silicon Kerr waveguide and about 0.02 W/μm² in 15 nm stacked graphene. That is roughly 8.3×10⁷ times better than graphene and 3.0×10¹¹ times better than silicon. Relative to typical conventional optical materials the paper quotes seven orders of magnitude. The scheme still works if quantum efficiency falls to 60%.
| Platform | Intensity to match digital r |
| Silicon Kerr (50 μm) | >72.6 W/μm² |
| Graphene saturable absorber | 0.02 W/μm² |
| Quantum activation | 0.24 nW/μm² |
On supervised tasks, nonlinear versus linear test accuracy is 96.72% versus 88.41% on MNIST and 87.87% versus 80.23% on FashionMNIST. A three-class spiral is separable in full-wave simulation. For RL they train 104 models (52 linear, 52 nonlinear). Linear Pong policies sit under a low ceiling with high variance; nonlinear policies scale toward near-perfect play against a human-level reference reward of 9.3. On HalfCheetah the nonlinear policy runs; the linear policy falls.
Estimating 3 Lseq dmodel optical neurons per Transformer layer at 0.1 μm² each, the nonlinearity-limited optical power stays below 1.2 W from GPT-2 through DeepSeek-V3, scaling as P∝Nparam^0.66. Silicon or graphene routes hit 10⁸ W on the largest models. This is a lower bound from the activation, not a system power number.
The scarce resource in optical computing is nonlinearity that can be stacked inside a chip-scale power budget, not faster linear multiply-adds. The paper translates "emitters are very nonlinear" into an r(I) that can be compared with digital nets, plus a modular inverse-design training loop. The practical filter is the pair 0.24 nW/μm² versus 72.6 W/μm²: a material that cannot match digital r will stay weakly nonlinear no matter how many layers you add.
The work is numerical. There is no fabricated chip. A bare SiV⁻ responds in the sub-GHz band, while optical computing often modulates at 10-50 GHz. The inverse-designed unit already shows a Purcell factor of about 2.74, and cavity literature reaches 2π×4.6 GHz, so GHz operation still needs extra enhancement. Solid-state emitters often need cryogenics to approach lifetime-limited linewidths; cooling power is outside the 1.2 W bound. Inhomogeneous broadening, deterministic placement, and on-chip dimensionality all sit far below LLM scale. Pong uses 2D simulation. The power formula assumes every optical neuron can run at Imin. Near-term large optical LLMs are more plausible as free-space diffractive systems than as integrated photonics.