Gaussian Light Transport
Patrick Attimont, Kartic Subr, Cyril Soler
SIGGRAPH Asia 2026 (Conference T
cs.GR
2026-09-10
INRIA fits 13D Gaussians to the rendering-equation residual. Living Room trains in 11m30s and renders 1920×1080 at 16.5ms/frame, 3.7× faster than Vertex Features.
Global illumination wants a full radiance field that any camera can slice. Path tracing recomputes per view, which is expensive for walkthroughs. Galerkin methods such as radiosity give a view-independent solution and then struggle with caustics and glossy highlights. Neural Radiosity fits the residual of the rendering equation with a network tied to a grid or hash encoding: large memory, slow training, little adaptivity to the actual light field.
The practical question is a radiance field that is not glued to the mesh, fits in a few megabytes, and can be sliced in real time.
Outgoing radiance L is a weighted sum of M Gaussians in 13 dimensions: position x, outgoing direction ω, normal n, albedo a, roughness r. Putting materials and normals into the representation lets disjoint surfaces that look alike share a kernel, instead of laying functions on a grid.
The objective does not fit a path-traced image. It minimizes the residual of the rendering equation: L should equal emission plus the transport operator applied to L. The loss is Monte Carlo over area and direction. Gradients use a semi-gradient trick that drops the inner integral, trading a bit of bias for lower variance. Each outgoing sample also draws 32 incident directions and 32 explicit light samples.
The mixture grows and shrinks with the residual, split, spawned, and pruned much like 3D Gaussian Splatting. Pruning uses relative energy with τ=8×10⁻³. Covariances are block-diagonal across position, direction, material, and normals, dropping cross-subspace correlation for spatial locality and faster evaluation. An anisotropy penalty on position (λ=0.008) stops contact regions from stretching into thin bright streaks. The residual is divided by the current prediction so dark regions are not ignored.
Evaluation bins queries with world-space Morton codes, then culls Gaussians against 13D tile bounds. On Bedroom, about 22K Gaussians, 76 kernels evaluated per pixel on average. Mitsuba owns the scene, PyTorch the optimizer, Taichi the kernels and analytic derivatives. Perfectly specular surfaces are not learned directly; paths continue to the first smooth hit.
Two ways to draw a frame. GS reads L at the primary hit. GS+MC traces one bounce and uses the model as incident radiance.
The baseline is Neural Radiosity with vertex features (Su et al., 2025), RTX 4080 SUPER, 1920×1080, 1 spp. This method runs a fixed 30K iterations; the baseline trains until artifacts clear (40K–100K). Training is 10.0–23.4× faster, rendering 2.2–4.5× faster, and FLIP is lower on five of six scenes. Disk size stays between 1 and 5.9 MB. A hash-grid encoding is a fixed 35.7 MB; vertex features range from 0.9 to 26.4 MB.
| Scene | Train | Model | 1 spp | FLIP (ours / VF) |
| Bedroom | 10m48s vs 2h10m | 4.4 vs 26.4 MB | 18.1 vs 70.4 ms | 0.0824 / 0.0774 |
| Chair | 10m07s vs 1h42m | 4.3 vs 1.9 MB | 16.1 vs 52.3 ms | 0.0486 / 0.0602 |
| Dining Room | 9m17s vs 1h54m | 5.4 vs 6.0 MB | 13.5 vs 60.2 ms | 0.0831 / 0.1012 |
| Living Room | 11m30s vs 4h29m | 5.8 vs 5.3 MB | 16.5 vs 61.5 ms | 0.1177 / 0.1362 |
| Rings | 7m37s vs 1h30m | 1.0 vs 0.9 MB | 19.9 vs 44.0 ms | 0.0425 / 0.0509 |
| Veach Ajar | 12m08s vs 2h25m | 4.1 vs 6.5 MB | 13.0 vs 46.5 ms | 0.0641 / 0.1499 |
Living Room's 16.5 ms is about 63 fps. Bedroom is the one scene where FLIP is slightly worse. Ablations that drop splitting, spawning, or the material/normal axes raise MSE; Bedroom full-model MSE is 0.00636 versus 0.03354 with no split. GS+MC is noisier at 1 spp and overtakes pure GS once it reaches 1024 spp.
This takes the explicit, cullable, adaptive kernels of 3D Gaussian Splatting and uses them to solve the rendering equation, rather than fitting an already rendered light field. For architectural walkthroughs or baked indirect light, a few-megabyte view-independent solution that trains in minutes and slices in milliseconds is closer to something you can ship in a viewer than Neural Radiosity.
The setting is static precomputed GI. It is not a general replacement for path tracing, and it is not built for dynamic scenes. The representation is new; the solver is still residual minimization.
The authors are direct. Residual sampling is uniform over surfaces, which favors spatially large Gaussians; curved glossy regions would rather see curvature- or bandwidth-aware sampling. The implementation is static scenes only. Tile culling assumes queries cluster in space and direction. That assumption weakens during training and GS+MC, and directional culling can slow inference by about 1.5× in the worst case. Gaussians are smooth by construction. Edges stay sharp where albedo, normal, or roughness jumps; discontinuities that live only in position or direction get softened, most visibly on near-specular highlights.
A bounded number of Gaussians biases the solution. Bedroom does not beat vertex features on FLIP. Dynamic objects and participating media are listed as future work.