Open-weights experiment: flipping one activation direction controls whether an LLM answers or stops
rayanpal_ · reddit · 2026-09-10
A researcher trained an open-weight Qwen3-4B model that generates a correct four-digit comparison, then autonomously decides to answer GO or end generation — no external filter involved. With prompt, weights, and reasoning trace held fixed, flipping one internal activation direction fully reversed answer/stop behavior: 40/40 answer→stop, 40/40 stop→answer, 640/640 controls unchanged. Weights, code, raw records and a paper are public. The author previously documented a cross-vendor "Semantic Void" behavior — models returning exactly zero visible output bytes after successful execution — across 31,430 trials spanning OpenAI, Anthropic, Google and Moonshot models.
More from Research
- 415k hours of full-duplex dialogue speech dataset released for spoken dialogue model training — kastnerkyle · 2026-09-10
- NAVER AI Lab Revisits Complete Reasoning Traces for Post-Training in New Paper — kastnerkyle · 2026-09-10
- Quantum sensor team uses GPT-5.6 Pro + Codex on Inverse Galois Problem, ranks 14th on IGP24 leaderboard — paulfinneyx · 2026-09-10
- VIGS-SLAM: visual-inertial 3D Gaussian Splatting SLAM hits SOTA on five datasets — rsasaki0109 · 2026-09-10
- DeepMind depth-recurrence patent fuels speculation about Gemini's next architecture — creatoroff · 2026-09-10
- Graph Machine: edge-based pretraining architecture that swaps dense Transformer layers — Lintai Hou · 2026-09-10