GLM Seeks Hard Prompts for Evaluation

teortaxesTex · x · 2026-07-19

The post is soliciting prompts that existing models still struggle with, spanning areas like reasoning, coding, SVG, and Chinese.

These samples will be used to evaluate the next-generation GLM. Although the original post is brief, the message is clear: the team is compiling a highly targeted set of difficult problems for their upcoming models, indicating an emphasis on real-shortcoming-driven evaluation design during model iteration.

Related event: Next-Gen GLM Seeks Hard Prompts for Evaluation(2 posts)→

Original post →

More from Research

Research channel →