Huawei's Folded Kirin NPU Uses 66% Less Power at the Same 29 TOPS

Huawei's $τ$ Chip Was Supposed to Melt?

Tingbo He

cs.AR

2026-09-03

Huawei LogicFolding stacks two logic tiers. Kirin 2026 density +55%; at matched 29 TOPS the NPU draws 66% less power than planar Kirin 9030 Pro.

What problem this solves

The first objection to 3D integration is heat. Stack active transistors on active transistors, raise power per square millimeter, and junction temperature hits the throttle. Huawei cannot buy EUV, so geometric shrink is stuck near 7 nm. The company switched to shrinking τ, the characteristic delay of a chip's critical paths, pronounced tau and written 韬 in Chinese. The flagship method is LogicFolding: take a planar circuit, fold it onto two silicon tiers, and turn long horizontal wires into short vertical hops.

When this was first shown at ISCAS in May 2026, the room's sharpest question was still heat. Measurements on Kirin 2026 reversed that objection.

Method

Hybrid bonding is treated as a wafer-level, cross-layer device step, not packaging. Two wafers are aligned, pressed, and annealed. Oxides form covalent bonds; copper pads fuse into continuous metal. Kirin 2026, on 40 nm bonding tools, reaches a 1.5 µm pitch and 50 million vertical interconnects, of which 10% to 15% carry signals. Kirin 2027 silicon is already at 1 µm and more than 100 million links. Within three years the target is a 720 nm pitch that matches top-metal, with more than 200 million links. Wafer-to-wafer is the path because lithography overlay is already nanometers; die pick-and-place bonding still lives at micrometers.

Folding targets the longest, hottest horizontal routes inside the NPU, CPU, GPU, and DSP. Wire length on folded paths falls about 20% on typical cores and up to 70% on some critical paths. On one block, clock wiring shrank 28% and buffer count dropped from 43,600 to 19,000.

Dynamic power is about 90% of a phone's daily bill. In P = αCV²f, C on an advanced node is dominated by wires, not gates. The office analogy: the commute burns more energy than the desk. In a large AI cluster, more than 80% of energy goes to moving data. Folding shortens the commute, so C falls linearly. Extra transistors buy parallel copies that run at lower voltage, so V² falls quadratically. The prior NPU had one big core and two efficiency cores; folding paid for four big cores. That array delivers 70 TOPS at 0.7 V; the predecessor needed 0.85 V for less than half the throughput.

Results

Against the planar Kirin 9030 Pro, at matched performance:

BlockIso-performancePowerOther
NPU29 TOPS-66%clock -63%, 0.85 V to 0.55 V, power density -73%
GPU61 FPS-58%200 mV lower voltage
CPU big coresame HNX score-41%clock -9%, 200 mV lower
DSP, gen-1same work-25%area -40%, power density +24%

Transistor density moved from about 155 million/mm² to 238 million/mm², a 55% jump equal to the prior three years of geometric shrink. At full throttle the NPU hits 70 TOPS (+141%), the GPU +42% frames, the CPU +18% on HNX, and power density in those modes exceeds the planar chip. Kirin 2027 DSP silicon already cuts power 47% and lands below the planar power density. The CPU performance core is back at 3.1 GHz this year, with a 5 GHz-and-up roadmap.

Why it matters

For on-device inference, the hard number is the NPU: same 29 TOPS, two-thirds less power, 0.55 V. This is system-technology co-optimization, not a new transistor. Firms without EUV are buying time instead of area. Firms with EUV may still care, if wires already dominate the power bill.

Cool silicon is a choice, not a free consequence of folding. Turbo modes still exceed planar power density. Kirin ran cool because the time headroom was spent on voltage. τ governs time; energy is a separate discipline.

Limitations

The author is explicit: τ is a time law, not an energy law. A folded chip that runs faster and burns more does not violate τ. The CPU big core is almost serial, so folding buys only a 9% clock cut and little of the quadratic voltage lever; three to five years of work remain. First-generation DSP area shrank faster than power, so power density rose 24%. Heat was not solved in one generation on that block.

This is a Huawei engineer preprint. The baseline is in-house Kirin 9030 Pro, with no third-party retest and no side-by-side against TSMC or Samsung 3D logic stacks. Bonding pitch on 40 nm tools versus circuit-folding gain is not ablated. Junction temperature and hotspot maps are qualitative; no junction-temperature number is given. Read the figures as company-reported silicon, not an independent bake-off.

Terms

Source

What people are saying

Related papers

All paper explainers