Open-weights experiment: flipping one activation direction controls whether an LLM answers or stops

rayanpal_ · reddit · 2026-09-10

A researcher trained an open-weight Qwen3-4B model that generates a correct four-digit comparison, then autonomously decides to answer GO or end generation — no external filter involved. With prompt, weights, and reasoning trace held fixed, flipping one internal activation direction fully reversed answer/stop behavior: 40/40 answer→stop, 40/40 stop→answer, 640/640 controls unchanged. Weights, code, raw records and a paper are public. The author previously documented a cross-vendor "Semantic Void" behavior — models returning exactly zero visible output bytes after successful execution — across 31,430 trials spanning OpenAI, Anthropic, Google and Moonshot models.

Original post →

More from Research

Research channel →