Luce: Relightable Gaussians for 3D Asset Generation
Mayank Singh, Michele Stoppa, Alvise Memo, Rui Yu, Harsha Kalli, Srimanth Gunturi, Muhammad Ahmed Riaz, Behrooz Shahsavari, Waleed Abdulla, David E. Jacobs
cs.CV, cs.AI, cs.GR
2026-08-25
Luce generates relightable PBR Gaussians from one image; Toys4K FID is 20.99 vs TRELLIS 2's 29.22, and CLIP is 0.8519 vs 0.8299 on 130 AI images.
Single-image 3D generation already produces plausible shapes. Materials are the bottleneck. 3D Gaussian Splatting bakes lighting into appearance, so a new lamp looks wrong. A textured mesh can enter a production renderer, yet surface text, logos, and engravings usually get smoothed away. TRELLIS 2 stores PBR attributes on voxels; LiTo eats view-dependent effects with spherical harmonics. Neither puts geometry and relightable materials into one generatable Gaussian representation.
Luce, from Apple, takes one image and emits an asset with albedo, metallic, roughness, and normals. The asset shades under any environment map and can optionally export a mesh with a tangent-space normal map. The tell is whether lettering on the surface stays legible.
The representation is a voxelized multimodal Gaussian cloud. Each voxel that the surface intersects holds three independent Gaussian sets, for albedo, metallic-roughness, and normals, each with its own positions, scales, rotations, and opacities. Polished wood wants dense albedo for grain; brushed metal wants dense normals and roughness for scratches. One shared set makes the modalities compete for the same primitives. Normals are stored, not derived from the surface, so they can later bake detail finer than the extracted mesh.
There are no spherical harmonics. Each modality is alpha-composited with standard 3DGS, yielding per-pixel albedo, metallic, roughness, and normals. Shading is deferred Cook-Torrance with split-sum image-based lighting. View-dependent highlights are analytic.
The raw cloud is too fat to generate directly. A structured latent VAE compresses a 128³ sparse grid with K=8 Gaussians per modality per voxel into a 64³ latent with 32 channels per voxel. The decoder stays at the coarse resolution and emits K'=32 Gaussians per modality per voxel, trading spatial resolution for density. Reconstruction is per-modality differentiable rendering with ℓ1, LPIPS, and SSIM, plus a KL term. The encoder is a 16-block sparse transformer that almost entirely runs at 128³ and downsamples only in the last block.
Generation follows TRELLIS's structure-then-latent recipe. The sparse-structure flow is TRELLIS 2's pretrained weights, untouched, and decides which voxels are occupied. SLatFlow is a 30-block DiT that then writes a full PBR latent on those voxels. Conditioning concatenates DINOv2 ViT-L/14 features from layers 6, 12, 18, and 24: early layers keep spatial detail, late layers keep semantics. The same latent has two exits: a Gaussian decoder that splats directly, or a FlexiCubes mesh decoder plus four maps (diffuse, metallic, roughness, tangent normals).
Training data is about 500K PBR-filtered assets from Objaverse and Objaverse-XL plus 158K from TexVerse. Each asset is rendered from 150 spherical cameras at 1024² in all three modalities, with authored normal maps applied on the normal pass. The VAE and the flow each run 500K steps on 64 H100s, about 14 days apiece.
The main tests are single-image generation on 412 Toys4K assets and on 130 Gemini 3.1 Flash Image pictures prompted for text, logos, and mixed materials. Times are averages on one H100.
| Method | Params | Time | Toys4K FID↓ | Toys4K CLIP↑ | AI-image CLIP↑ |
| TRELLIS GS | 1.70B | 29.6s | 30.75 | 0.8898 | 0.8299 |
| TRELLIS 2 | 7.48B | 176.7s | 29.22 | 0.8895 | 0.8110 |
| LiTo | 1.84B | 76.4s | 29.76 | 0.8909 | 0.8234 |
| Luce GS | 4.48B | 42.2s | 20.99 | 0.9062 | 0.8519 |
| Luce mesh + tangent normals | 4.58B | 159.2s | 21.10 | 0.9072 | 0.8508 |
Toys4K FID falls from 29.22 for the strongest baseline, TRELLIS 2, to 20.99, about 28% relative, matching the abstract. The mesh with tangent normals slightly leads the Gaussian path on CLIP and SigLIP2. On the in-house AI images, CLIP is 0.8519 against 0.8299 for TRELLIS GS.
Reconstruction on 338 Toys4K assets that have all three PBR maps is supporting evidence for the representation, not the headline. Luce GS reaches 36.1 dB color PSNR and 34.6 dB normals, above TRELLIS 2 at 34.5 / 32.1. TRELLIS 2 wins albedo and metallic-roughness PSNR (40.7 vs 38.6, 42.7 vs 39.1). Baking tangent normals lifts mesh normal PSNR from 29.5 to 33.0 dB at zero extra polygons.
The ablation is clean. With only DINOv2 layer 24, Toys4K FID rises from 20.99 to 25.21 and AI-image CLIP drops from 0.8519 to 0.8081. Multi-layer conditioning is the main switch for keeping text and logos.
Evaluation lighting is not the conditioning lighting. Methods that factor materials (Luce, TRELLIS 2) are relit to the evaluation HDRI; TRELLIS and LiTo carry baked appearance into FID. Luce conditions SLatFlow at 1036²; TRELLIS and LiTo use 518²; TRELLIS 2's latent stage uses DINOv3 at 1024². The numbers lead. The comparison is not fully matched.
This moves generative 3D toward a production asset format. It is not another radiance field that only looks right under fixed lights. The Gaussian path is 42 seconds and 4.5B parameters: lighter than TRELLIS 2 at 177 seconds and 7.5B, slower than TRELLIS GS at 30 seconds, in exchange for relighting and readable surface text. The mesh path sits around 150 seconds, closer to "importable into a DCC" than to an interactive preview.
If the job is an object-centric asset that must change environment lighting and ship PBR maps, this representation is a better fit than spherical-harmonic appearance. If the job is the cheapest plausible 3D, TRELLIS GS is still cheaper. The gain is incremental, and it is aimed at material factorization and high-frequency detail, not at half a FID point.
The authors list four. Finite voxel resolution under-resolves features that span only a few voxels relative to the object's extent. The material model is a standard metallic workflow with Cook-Torrance: no subsurface scattering, anisotropy, translucency, or thin-film interference. Mesh export reuses TRELLIS's FlexiCubes; UV-aware extraction is future work. The pipeline is object-centric; scene-level lighting consistency and object interaction are open.
The comparison protocol is unmatched, as above. The 130 AI images come from one generator with prompts biased toward detail and text, which is exactly the signal multi-layer DINOv2 is built to keep; the generalization radius is untested on an independent set. The paper does not describe public code or weights in the main text; a reproduction budget is on the order of 64 H100s times 14 days times two stages.