Lan Hackathon Review: Qwen Excels in Physics/Engineering, K3 in General Intelligence
葬AI · wechat · 2026-09-01
The 'Bench Hackathon' held in an internet cafe featured 30 custom benchmarks, revealing model capabilities in physics simulation, embedded drivers, and poker strategy beyond standard coding tests.
Embedded Driver Skills:
- Qwen 3.8-Max: Ranked 1st in both tasks, demonstrating strong system-level agent capabilities (reading manuals, reverse engineering, closed-loop fixes).
- DeepSeek: Passed simple tasks but failed in complex DMA state machine scenarios.
- Opus 5.0: Best code quality and understanding, but lacked stability under pressure.
Physics Understanding (Drones, Arms, Cutting, etc.):
- Qwen: Highest overall score, the only model to achieve stable drone flight control.
- GPT: Dominated suction cup manipulation; stable in other tasks.
- K3: Strong physical intuition (1st in material grasping) but weak execution discipline (drones crashed).
- Conclusion: No model is omniscient in physical scenarios.
Texas Hold'em (Wall-Breaker Bench):
- K3: Best performance, effectively utilizing behavioral tells without blind calling.
- Qwen: Too conservative; win rate decreased with video modalities added.
Overall, Qwen excels in engineering/physics due to rich training data, while K3 leads in multimodal/general intelligence. LLMs remain 'coding-specialized', far from AGI.
More from Models
- Has anyone tuned a model to operate exclusively in E-prime yet? — cephaloform · 2026-09-01
- Heavy users report Claude quality dropping over the past week: eager to execute, no more clarifying questions — Numerous_Leopard_522 · 2026-09-01
- User seeks best LLM for CLI coding on single 3080 Ti — -samae1- · 2026-09-01
- Polymarket Traders Put 62% Odds on Next Mythos Model Landing by September 2 — Polymarket · 2026-09-01
- Creator finds Opus 5 powerful but confusing and overly complex — k_kool_ruler · 2026-09-01
- Why AI text still reads like a bot: ICLR paper cuts slop by 90% via inference-time bans — ziv_ravid · 2026-09-01