Kimi K3 at Q1 writes working code in one shot; GLM5.3 API ran 7 hours and delivered
segmond · reddit · 2026-08-19
A Reddit thread collects hands-on reports on the latest open models. The poster reports Kimi K3 at Q1M quantization "generated a lot of code that worked in one shot" and is working toward Q2/Q4; GLM5.3 via API, given full server access, ran 7 hours and completed serious work — building vLLM with custom patches — leaving them very impressed. They're downloading DeepSeek-V4-Pro-0813-Q2 to test coherence at Q2 versus Flash Q8, while Qwen3.8-2.4T has fizzled in discussion.
More from Infra
- Open-Source RL Framework Miles v0.1 Ships: 1,326 Commits, 85 GPU E2E CI Tests, Firecracker Sandbox Rollouts — ying11231 · 2026-08-19
- Onchain AI Gateways Surge, Stripe Eyeing Cheap Inference Integration — 0xJeff · 2026-08-19
- Meta Deploys CXL to Reuse DDR4, Cutting Server Count by 25% — BenBajarin · 2026-08-19
- Fix for llama.cpp 100% CPU single core usage despite full GPU offload — MelodicRecognition7 · 2026-08-19
- NVIDIA Releases PyG 26.07 with TrueQuery: A GNN+LLM Graph RAG Framework — AllThingsApx · 2026-08-19
- Qualcomm CEO outlines AI inference vision post-Modular acquisition — samcharrington · 2026-08-19