Discussing Direct FP8 Weight Loading in Ostris AI Toolkit
iz-Moff · reddit · 2026-07-10
When training a LoRA with limited VRAM using the Ostris AI Toolkit, a developer enabled FP8 quantization. However, the tool loads the full-size Transformer weights first and quantizes them every time. This wastes time and causes peak VRAM usage during loading and quantization, easily triggering OOM (Out of Memory) errors. The poster asks why this process must be repeated and how to make the tool load the already downloaded FP8 weights directly.
More from Infra
- How to build a PostgreSQL-backed semantic search pipeline with pgvector and Ollama — KhuyenTran16 · 2026-07-21
- NeurIPS 2026 workshop calls papers on on-device intelligence — YiMaTweets · 2026-07-21
- Milled from Solid Aluminum: AI Rig Multi-GPU Case for Local Compute — dee_hw · 2026-07-21
- FutureCaribbean’s Buildathon offers $50K, H200 compute, and an NYSE pitch — HeyAmit_ · 2026-07-21
- A new series tests which data-science workflows can run on GPUs today — pandeyparul · 2026-07-21
- Former AWS operator says Bedrock margins can beat SageMaker as agentic AI lifts CPU demand — RihardJarc · 2026-07-21