Running 288k context at 5 tokens/sec for 16 hours: "Maybe I am Qwen"
LeviTurk · x · 2026-09-19
A user shared their extreme local LLM setup: 288k context window, but generating at just 5 tokens per second — meaning 16 hours to fill the context. They also noticed throughput halving near the end of the context window, "just like Qwen," then joked: "WAIT. Maybe I am Qwen???"
More from Fun
- Asking Claude to draw cats in JavaScript yields hilarious results — DareFailed · 2026-09-19
- GPT-6 Astra Narrates Its Own Unciv Game Video Using Gemini TTS and Remotion — Angaisb_ · 2026-09-19
- Plot twist: space data centers are secretly AI's shutdown-proof insurance — bendee983 · 2026-09-19
- "It's over Dario and Sam": another AI meme moment — glcst · 2026-09-19
- LeCun mocks AI doom double standard: Dario called GPT-2 too dangerous back in 2019 — ylecun · 2026-09-19
- "He's such a Claude" becomes slang for someone brilliant but prone to clown behaviour — madhavsinghal_ · 2026-09-19