2026-08-22
UC Davis and TAMU add latency to the interconnect FoM and propose TSOV 3D photonic chiplets; they project >10 TB/s/mm² vs UCIe-3D and report 496 fJ/bit.
Training compute for generative models has moved from petaFLOPs toward yottaFLOPs. Data movement, not arithmetic, dominates energy. One cited workload study puts interconnects at 62.7% of the power budget. Copper still works at extreme short reach. Stretch the channel and attenuation, crosstalk, and jitter force equalization plus clock-and-data recovery (CDR), and energy per bit climbs. HBM, UCIe, and co-packaged optics (CPO) all pull photonics closer to the package. The electrical run from ASIC to optical engine can still be tens of millimeters. This paper wants the high-speed data plane in optics, running vertically through a 3D stack, with copper kept for power delivery and ultra-short, latency-critical links.
The proposed platform is 3D-EPIC: chiplets stacked on an active optical interposer. Vertical light uses through-silicon optical vias (TSOVs). Electrical TSVs and 2.5D wiring still carry power and cache-coherent short hops. Every high-speed node gets a direct tap into a global optical fabric, compatible with HBM and PCIe and with off-package fiber. Electronics and photonics can be fabricated on different nodes: an older monolithic photonic process for the PIC, a leading CMOS node for the EIC.
They put link latency into the figure of merit: bandwidth efficiency over energy efficiency, times reach over latency. An earlier DARPA PIPES FoM left latency out. CPO marketing that only quotes bandwidth and pJ/bit will not match a real link budget once delay is counted. This paper writes latency in explicitly.
The electrical-versus-optical energy model is simplified: 8 Gbps, 60 fJ/bit at the receiver, 1 dB/cm waveguide loss, 3 dB per coupler, 30% laser wall-plug efficiency. The partition length, where optics undercuts copper, is 15.1 mm in that model. Adding a channel-loss-dependent DFE cost pulls the crossover to 2.5 mm. Published CDR numbers sit around 1900 fJ/bit; optically forwarding the clock is how they hope to drop that term.
Bandwidth density is benchmarked against UCIe-3D. Electrical density starts from bump pitch, then power/ground and repair overheads. Optical density assumes a fraction of bumps become TSOVs (conversion ratio RTSOV), each carrying N WDM wavelengths. A 55 μm pitch matches an HBM-like stack.
On a 1 mm by 1 mm chiplet at 55 μm pitch, hexagonal array, overhead 0.39, realizable electrical density is 925.98 GB/s/mm². A 32-wavelength optical design on the same geometry lands at 768 GB/s/mm². Raising WDM to 39 wavelengths reaches 936 and just beats copper. At denser pitches even 4-wavelength WDM can win, because optical channel rate is not pinned by electrical signal integrity. Push conversion ratio and WDM further and the paper claims >10 TB/s/mm². A 2% TSOV conversion already matches electrical 3D at 55 μm. Area I/O scales with the square of die edge; shoreline I/O scales linearly, so 3D optical vias pull ahead as dies grow.
| Scheme | Setup | Bandwidth density |
| UCIe-3D realizable | 55 μm, hex, OH 0.39 | 925.98 GB/s/mm² |
| 3D optical, 32 λ | same geometry | 768 GB/s/mm² |
| 3D optical, 39 λ | same geometry | 936 GB/s/mm² |
On the device side, FDTD of a full TSOV (two 45° mirrors plus a vertical waveguide) gives 0.7 dB with no offset and 0.42 dB with nanometer offsets, staying under 1 dB across the C-band. They have etched a 20 μm-tall, 370 nm-wide via at 54:1 aspect ratio, targeting 100 μm / 270:1. AuSn20 eutectic bonding closed a 2500-pad daisy chain at low yield; the best sample took 22 kgf of shear and the die cracked before the bond.
A previously reported transceiver, 12 nm EIC plus AIM Photonics PIC bonded with DBI, reached -17.01 dBm receiver OMA sensitivity at 25 Gb/s (among the lowest then published), 496 fJ/bit SerDes at 18 Gb/s, and 191 fJ/bit for the receiver alone at 25 Gb/s. Thirty-two optical and electrical channels, one pair forwarding the clock in optics. If the photodetector and TIA move into a GF45SPCLO monolithic process and input capacitance falls from 28 fF to 3.2 fF, simulated OMA sensitivity is -24.2 dBm. Fold in a more advanced CMOS node and they sketch a path to ≤100 fJ/bit. That number is a projection, not a measurement.
For GPU-cluster and CPO engineers, the paper turns "optics is always cheaper" into a ledger that includes latency, and it flags the in-package electrical run as CPO's hidden term. If TSOVs work, high-speed I/O stops being shoreline-limited and scales with area, on the same bump grid as HBM. The 496 fJ/bit heterogeneous transceiver is the part that exists in silicon. 10 TB/s/mm² and 100 fJ/bit are a roadmap.
TSOVs are still simulation plus a single-via etch. There is no measured 3D-chiplet optical link. Both ≤100 fJ/bit and >10 TB/s/mm² hang on WDM count, node shrink, and conversion ratio; the authors park spectrum planning outside this study. Eutectic yield is low, and the metal composition drifted until they changed the evaporation recipe. The partition-length model omits a full DSP stack, and 8 Gbps is far below today's 100G/lane. Some commercial points on the FoM plot are literature estimates. Putting latency into the FoM is the right instinct, but latency is not defined the same way across standards, so the ranking in Figure 1 is parameter-sensitive.