Sherry-style 3:4 ternary weights hit 1.375 bits: 1.6MB WebGPU model matches 7.8MB int8

Brilliant-Hall1387 · reddit · 2026-09-30

Precisit ran Sherry-style 3:4 ternary weights (T34: per 4 weights, one zero and three ±1, packed into 5 bits = 1.375 bits/weight) in the browser via WebGPU. A 1.59MB Connect Four scorer matched its 7.8MB int8 counterpart, at 1.1ms per move on an M5 Pro.

Over 200 games, T34 trained with the format in the forward pass scored 0.93/0.91 vs depth-4/depth-6 bots — while post-training conversion collapsed to 0.13, showing ternary quantization must happen during training. Other findings: attention q/k/v matrices are the most sensitive layers, group size barely matters, and seed variance is large (0.945 vs 0.882 on identical recipes). Code, models, demo and write-up are MIT-licensed.

Original post →

More from Infra

Infra channel →