QLLM Claims to Achieve O(1) Inference

ExtremeKangaroo5437 · reddit · 2026-07-09

The author introduces QLLM, a new architecture proposed by their team that uses neither transformers nor mamba, claiming features like O(1) inference and no need for a KV cache.

The post mentions that the theory was developed first, followed by a trainable practical version. They subsequently collaborated with researchers from UC Berkeley and Indiana University to publish a paper, eventually building a 100M parameter model.

Original post →

More from Models

Models channel →