Discussing Direct FP8 Weight Loading in Ostris AI Toolkit

iz-Moff · reddit · 2026-07-10

When training a LoRA with limited VRAM using the Ostris AI Toolkit, a developer enabled FP8 quantization. However, the tool loads the full-size Transformer weights first and quantizes them every time. This wastes time and causes peak VRAM usage during loading and quantization, easily triggering OOM (Out of Memory) errors. The poster asks why this process must be repeated and how to make the tool load the already downloaded FP8 weights directly.

Original post →

More from Infra

Infra channel →