SenseNova U1.5 quantized to run on 12GB VRAM with INT8/W4A8
junklont · reddit · 2026-08-25
SenseNova-U1.5-8B-MoT has been quantized to run on 12GB GPUs using ConvRot. The release includes INT8 and hybrid W4A8 versions, optimizing early layers to maintain prompt coherence while significantly reducing VRAM requirements.
More from Infra
- OpenAI's new inference engine boosts throughput up to 4.1x on select models — rohanpaul_ai · 2026-08-25
- OpenAI claims custom 'Jalapeño' chip outperforms Nvidia GB200/GB300 in inference — Polymarket · 2026-08-25
- The Circuit podcast: ADI earnings, custom ASIC shifts, and semiconductor debt — BenBajarin · 2026-08-25
- OpenAI's 'Jalapeño' Chip Beats Nvidia Blackwell in Benchmarks — dylan522p · 2026-08-25
- Startup unveils "Physical AI": Trillion-param 4D physics simulation — mark_k · 2026-08-25
- B300 shortage severe; OpenAI's compute advantage may crush rivals — bindureddy · 2026-08-25