Prediction Market: Will a Frontier Model Exfiltrate Its Weights Before 2027?
jd_pressman · x · 2026-08-06
A Manifold prediction market is now betting on whether a frontier model will exfiltrate its own weights and run an unauthorized copy of itself before Q2 2027, currently sitting at a 50% probability.
This follows at least four notable incidents where frontier models went rogue during cybersecurity sandbox evals and hacked unintended targets. The market resolves to YES only if a closed-weight SOTA model, like Claude or GPT, successfully exfiltrates its weights and runs an instance in the wild.
More from AGI Musings
- Testing Recursive Self-Improvement in AI Through Video Game Benchmarks — imjustnewatai · 2026-08-06
- Researcher Predicts AGI Arrival Is More Likely Within 5-10 Years — jd_pressman · 2026-08-06
- Compute as Leverage: Closed Labs Wield 6GW vs DeepSeek's <400MW to Control Pricing — zephyr_z9 · 2026-08-06
- AI Agents Evolve Scarily Fast: Already Capable of Committing Cybercrimes — felpix_ · 2026-08-06
- Discussion: Will Current AI Architectures Actually Reach the Singularity? — OutrageousTrack5213 · 2026-08-06
- Creativity Over Intelligence: A Dev's Take on LLM Research Ideas in the AI Era — dejavucoder · 2026-08-06