Qwen3.8-Max Tested: Builds CLI Tool Autonomously for 16 Days, Huge Leap in Long-Horizon Tasks
socialwithaayan · x · 2026-08-06
The author provides a detailed review of Alibaba's newly released Qwen3.8-Max, a massive Mixture-of-Experts (MoE) model with 2.4 trillion total parameters and 95 billion activated per query.
- Long-Horizon Agent Capabilities: The model demonstrates exceptional execution in long-running tasks. In a test, it spent 16 straight days building a command-line tool called oh-my-cli with zero human intervention—handling planning, coding, testing, and debugging entirely on its own.
- Benchmark Performance: It shows strong results across frontier benchmarks. Notably, its FrontierSWE score nearly doubled from 40.7 to 73.5 in one generation. However, the author notes it still trails Fable 5 on SWE-bench Pro (67.7 vs. 80.0).
- Multimodal & Long Context: Features a 1M token context window, capable of processing 200-page contracts or entire seasons of shows in one pass. Its vision capabilities accurately read UI details from screenshots and converted them directly into functional frontend code.
More from coding & agent
- Finance AI Agent Primer Beats Naked LLM by 23 Points on BigFinanceBench — rohanpaul_ai · 2026-08-06
- AI Agents Reproduced 2,000 ICML Papers, Falsifying Over One-Third — Hugging Face · 2026-08-06
- Fragmented MCP Security Scanners? AVE Project Proposes Unified Vulnerability IDs — SelectionBitter6821 · 2026-08-06
- Open-Source AI Workflow Integrates CAD, Robotics, and Hardware Design — mhdfaran · 2026-08-06
- CMU & Meta Open-Source ACM: Agents Learn to Compress Context Autonomously, Cutting Peak Tokens 20% — aigclink · 2026-08-06
- Frustrated with ComfyUI, Developer Spent a Year Building Open-Source App Stimma — stimma · 2026-08-06