Running 288k context at 5 tokens/sec for 16 hours: "Maybe I am Qwen"

LeviTurk · x · 2026-09-19

A user shared their extreme local LLM setup: 288k context window, but generating at just 5 tokens per second — meaning 16 hours to fill the context. They also noticed throughput halving near the end of the context window, "just like Qwen," then joked: "WAIT. Maybe I am Qwen???"

Original post →

More from Fun

Fun channel →