Don't Run This Model in FP16: Activations Exceed Its Dynamic Range

tomaarsen · x · 2026-10-07

The author warns against running the model in FP16 and recommends BF16 or FP32 instead. The model's activation range exceeds FP16's dynamic range, and the model card warns of NaNs or silently degraded embeddings. Use BF16 where natively supported, and FP32 elsewhere, including most CPUs.

The quoted post adds multimodal input defaults: video is sampled at 1 frame per second, audio should be mono at 16 kHz, and the vision budget is configurable from 70 to 1,120 soft tokens per image/frame — more tokens trade latency and context capacity for finer visual detail.

Related event: Google Open-Sources EmbeddingGemma 2, First Native Multimodal Embedding Model(37 posts)→

Original post →

More from coding & agent

coding & agent channel →