A Thread-Register Decoupled GPU Execution Model for Efficient Tensor Computation
matt_d · hn · 2026-08-27
This ArXiv paper proposes a thread-register decoupled GPU execution model designed to improve the efficiency of tensor computation. By decoupling the traditional binding between threads and registers, the method optimizes GPU resource utilization and enhances overall computing performance.
More from Research
- Higher-resolution microscopy can hurt CNNs: downsampling 4x improves U-Net segmentation — bravo_abad · 2026-09-22
- Did OpenAI Solve the Wrong Navier-Stokes Problem? Experts Cry Loophole — joshgans · 2026-09-22
- Bridging LLM Decision Readouts into DuckDB: Zero-Token Probabilistic Classification via LuaJIT UDFs — Shoddy_Telephone9702 · 2026-09-22
- LLM agents fail to converge in double auctions, allocate less efficiently than humans — WillRinehart · 2026-09-22
- Extracting Entities and Relations from 5M Court Decisions Without an Expensive LLM Pass — SignificantZebra5883 · 2026-09-22
- SVEET: streaming video editing with a diffusion model hits 15 FPS on a single H100 — SJTU · 2026-09-22