LOTUS: Looped Transformer with Parallel Latents
burny_tech · x · 2026-07-19
Researchers introduced LOTUS (Looped Transformers with parallel supervision on latents). This method inserts K latent blocks between the model's question and answer, running R loops of computation with the base model to obtain the final latents.
Training uses two simple loss functions: first, a standard cross-entropy loss on the chain-of-thought via the model's own LM head; second, a standard next-token prediction based on these latents to derive the final answer.
More from Research
- Project APE finds verifier reliability drops when papers contain multiple errors — soumitrashukla9 · 2026-07-22
- Project APE says verifier costs fell about 90x in a year as Chinese open models lead — soumitrashukla9 · 2026-07-22
- OpenAI-linked paper says capability RL can make models more reward-seeking — MariusHobbhahn · 2026-07-22
- Project APE builds its verifier benchmark from 100 AI-written papers with injected errors — soumitrashukla9 · 2026-07-22
- Paper proposes a CRED taxonomy and benchmark to measure research-error detectors — soumitrashukla9 · 2026-07-22
- OpenAI says long-horizon models need safety and alignment checks across full action sequences — rhiever · 2026-07-22