Sherry-style 3:4 ternary weights hit 1.375 bits: 1.6MB WebGPU model matches 7.8MB int8
Brilliant-Hall1387 · reddit · 2026-09-30
Precisit ran Sherry-style 3:4 ternary weights (T34: per 4 weights, one zero and three ±1, packed into 5 bits = 1.375 bits/weight) in the browser via WebGPU. A 1.59MB Connect Four scorer matched its 7.8MB int8 counterpart, at 1.1ms per move on an M5 Pro.
Over 200 games, T34 trained with the format in the forward pass scored 0.93/0.91 vs depth-4/depth-6 bots — while post-training conversion collapsed to 0.13, showing ternary quantization must happen during training. Other findings: attention q/k/v matrices are the most sensitive layers, group size barely matters, and seed variance is large (0.945 vs 0.882 on identical recipes). Code, models, demo and write-up are MIT-licensed.
More from Infra
- Ollama 0.35 ships with local Nimble support on Mac — Technovangelist · 2026-09-30
- Nvidia CEO Jensen Huang: Data centers are now 'superintelligence factories' — Polymarket · 2026-09-30
- AMD's Hyperloom and ROCm 10: AI agents tune GPU kernels overnight with accuracy checks — AnushElangovan · 2026-09-30
- Efficient raises $97M Series B to rethink computing from physical AI to data centers — Sethwinterroth · 2026-09-30
- Developer: OpenAI is taking compute from paying users for consumer agents — arthurcolle · 2026-09-30
- Prime Intellect to deploy on NVIDIA's new Vera CPU in first wave — eliebakouch · 2026-09-30