20-year engineer benchmarks Gemma4-31B vs Qwen3.8-27B locally; GPT-6.1-Sol is still another tier
therealjerseytom · reddit · 2026-10-11
A software engineer with 20 years of experience tested Gemma4-31B vs Qwen3.8-27B (same Q4 quant) on work-like tasks, with GPT-6.1-Sol as reference:
- Repo analysis: equally good output. Gemma4 was far more efficient (10 vs 20-30 server calls) and tightly scoped; Qwen showed more initiative, digging beyond the ask for a fuller picture.
- Authoring a new C#/WPF app (1-2 days of human work): Gemma4 planned and executed fast, first pass B-, needing several feedback rounds. Qwen3.8 made many small errors (missing using statements) but was more persistent and aimed higher — first cut was better and one revision got it to B+.
- GPT-6.1-Sol scored an A+: a one-shot, ship-it-quality solution with genuinely different polish and context understanding.
Takeaway: both open-weight models are workable on consumer hardware, but clearly not in the same conversation as frontier models. Next up: the true developer hell of legacy codebases like Doom or Command & Conquer source.
More from coding & agent
- Harvard Med workshop teaches building AI co-scientists with ToolUniverse's 2,700+ tools — marinkazitnik · 2026-10-11
- Delvetown agents keep producing impressive artifacts as AI town experiment evolves — lfschiavo · 2026-10-11
- Anthropic launches Claude Dashboards beta: plain-language BI from Snowflake to Redshift — shashib · 2026-10-11
- Nous Research launches stealth coding and agentic reasoning model, free for limited time — Teknium · 2026-10-11
- Cloud agent users mock devs still lugging around laptops in viral meme — steipete · 2026-10-11
- MCP server plus a World of Warcraft environment: AI agents entering Azeroth? — djcows · 2026-10-11