Reddit asks for cold prefill speeds on Qwen3.8FlashNext at 128k compaction
Express_Quail_1493 · reddit · 2026-09-28
A Reddit user notes many people share great decode and prompt-prefill speeds for Qwen3.8FlashNext, but nobody reports prefill speeds during harness compaction. They ask for partial-offload cold prefill benchmarks at 128k long context.
More from Infra
- One AMD driver flag boosts dual-GPU Vulkan LLM inference up to 4x — tabletuser_blogspot · 2026-09-28
- PrismML's Bonsai 2 shrinks Qwen3.8 27B to 5.9GB, keeps 98.2% capability, runs on a 5090 — dl_weekly · 2026-09-28
- Developer Plans to Let Codex Pick Which Tests Run, Slashing CI Costs — sull · 2026-09-28
- Free online guide covers LLMs from first principles to local deployment — JFPuget · 2026-09-28
- Is a vector database enough for production AI agents? Reddit debates storage design — OkShirt9372 · 2026-09-28
- Pre-training a Foundation Model Whose Tokenizer Is PTX, Not Natural Language — rickasaurus · 2026-09-28