If AI finds a 1000x better architecture no human can understand, should we build it?

Neurogence · reddit · 2026-10-11

A Reddit poster, riffing on Terence Tao's recent comments on interpretability, poses a thought experiment: suppose an internal model discovers a brand-new AI architecture yielding models 1,000x more capable, dramatically faster, and far cheaper to train — but the reasoning behind it is as incomprehensible to humans as a transformer is to an ape, and cannot be explained even with unlimited analogies.

The question: should humanity refuse to build architectures it cannot understand over interpretability and safety fears, even if they could bring accelerating intelligence, cures for all diseases, and unlimited abundance? The thread argues interpretability may become impractical as we approach true ASI.

Related event: Terence Tao's Thought Experiment Sparks Debate on AI Discoveries Beyond Human Understanding(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →