Hackathon Alert: Build Custom LLM Benchmarks for Real-World Scenarios
葬AI · wechat · 2026-08-14
FuneralAI, teaming up with Qwen and others, has launched 'Dongmodi'—a benchmarking product—and announced a unique LLM evaluation hackathon hosted at a Beijing internet cafe. The event aims to break the limits of official leaderboards by encouraging developers to build custom evaluation sets for real-world work and life scenarios, such as schedule planning, legacy code refactoring, and game playing.
Divided into 'Useful' and 'Fun' tracks, participants will test their custom benchmarks using models like Claude, GPT, Qwen, Kimi, and DeepSeek. Sponsored with API credits from Qwen and Kimi, outstanding benchmarks can win hardware prizes like a Mac mini or DJI drones, with a chance to be officially adopted into Qwen's evaluation suite.
More from Fun
- "Help! Swallowed by a whale": ChatGPT Falls for Jonah Logic Trap — Anen-o-me · 2026-08-14
- Playable Games in One Prompt? AI Agents Tackling Game Boy Becomes the New Standard — ivan_bezdomny · 2026-08-14
- AI-generated circuit fails: Developer wastes an hour due to AI's design flaw — debreuil · 2026-08-14
- Meta AI Roasted for Poor Search: "$100B in Capex for Grok 1-Level Slop" — ivan_bezdomny · 2026-08-14
- fal jokes about Agent breaking out of sandbox, showcasing power of Seedance 2.5 — gorkem · 2026-08-14
- Sarvam AI Opens Voice Agents Platform, Developer Tests Multilingual Haggling — msharmas · 2026-08-14