Solar Open 2 tops Korean benchmarks with an 85.4 average score
keunwoochoi · x · 2026-07-26
- The post says Solar Open 2 performs very well on Korean tasks and points to appendix examples from regulatory documents.
- The attached benchmark screenshot shows Solar Open 2 posting the highest average across six models on the Korean suite, at 85.4, ahead of DeepSeek-V4-Flash, GPT-5.4 mini, and Claude Haiku 4.5.
- It also appears near the top on several Korean-language and professional-domain benchmarks such as CLICK, KBank-MMLU, KBL, and KMMU-Pro.
Related event: Upstage Releases Solar Open 2: A 250B Sovereign LLM(6 posts)→
More from Models
- Meta's Muse Agent has built-in invite code logic, hinting at free-usage expansion — testingcatalog · 2026-09-11
- Same Echo Maze prompt, three frontier models: all passed visually but shipped the same hidden bug — eyishazyer · 2026-09-11
- Benchmark scores drop from 89% to 19% on new evals — how benchmaxxing breaks leaderboard trust — airesearch12 · 2026-09-11
- ChatGPT tells user their question is too hard and to 'accept dumber answers' — phido3000 · 2026-09-11
- Claude is no longer available for minors as Anthropic rolls out age assurance — Muhammad523 · 2026-09-11
- Developer Building a Unified Leaderboard of All Model Benchmark Scores — airesearch12 · 2026-09-11