GLM flubs random-context test, possibly confusing user and system roles
eliebakouch · x · 2026-09-17
HF researcher eliebakouch shared a model behavior observation: GLM gave a strange answer on a random-context test, which he suspects comes from Astra confusing the "user" and "system" roles — "which seems very dumb." He clarified the test wasn't intentionally out-of-distribution. A scattered but real data point on role-handling robustness.
More from Models
- GPT-6 Astra's Epoch ECI Score Revised Down on Weak Long-Horizon Software Engineering — Jsevillamol · 2026-09-17
- Mystery model 'Union Alpha' tops DeepSWE, claiming GPT-6-class capability at DeepSeek-level prices — daniel_mac8 · 2026-09-17
- Jensen Huang at All-In Summit: AI leadership will be built by everyone, open and closed models both matter — NVIDIAAI · 2026-09-17
- Model 'Jev' shows well-calibrated probabilities: 1.74pp average calibration error across benchmarks — hackgoofer · 2026-09-17
- 500 Dirty Web Pages Benchmarked: 12B Open Model Falls Off a Cliff Where 70Bs Converge — JUSTINWOODS118 · 2026-09-17
- Filtering synthetic envs where Qwen deterministically fails but GLM5.3 solves reveals odd behaviors — kalomaze · 2026-09-17