Hackathon Alert: Build Custom LLM Benchmarks for Real-World Scenarios

葬AI · wechat · 2026-08-14

FuneralAI, teaming up with Qwen and others, has launched 'Dongmodi'—a benchmarking product—and announced a unique LLM evaluation hackathon hosted at a Beijing internet cafe. The event aims to break the limits of official leaderboards by encouraging developers to build custom evaluation sets for real-world work and life scenarios, such as schedule planning, legacy code refactoring, and game playing.

Divided into 'Useful' and 'Fun' tracks, participants will test their custom benchmarks using models like Claude, GPT, Qwen, Kimi, and DeepSeek. Sponsored with API credits from Qwen and Kimi, outstanding benchmarks can win hardware prizes like a Mac mini or DJI drones, with a chance to be officially adopted into Qwen's evaluation suite.

Original post →

More from Fun

Fun channel →