Activation-aware quantization improves GPQA accuracy but slows Gemma 3 4B QAT
devildip · reddit · 2026-07-21
- The poster is testing activation-aware quantization and sees a clear trade-off: higher accuracy comes with lower throughput.
- On GPQA with Gemma 3 4B QAT, the test setup reached 27.8% accuracy at 331 ms, while the control reached 23.2% at 256 ms.
- The question is practical rather than theoretical: if you had to choose, would you prefer faster quantization or more accurate quantization?
More from Research
- OmniSearch puts text, images, audio, and video into one semantic search space — victorialslocum · 2026-07-21
- Cold Spring Harbor Asia sets a genome biology conference in Suzhou for Oct. 12–16 — jmuiuc · 2026-07-21
- A clean counterexample shows a map can be locally diffeomorphic yet globally fold — Algomancer · 2026-07-21
- Xiaohongshu’s dots-note-3.0 gets a perfect IMO score and becomes the world’s second gold model — 量子位 · 2026-07-21
- Statistical theory paper studies how fast signatures learn in path regression — chaumian · 2026-07-21
- PROWL uses a world model to keep Minecraft agents exploring after failures — nathanbenaich · 2026-07-21