BlockSearch: a 0.6B in-context retriever rivals dense retrieval at million-token scale
CShorten30 · x · 2026-10-01
- New Weaviate Podcast episode with Siddharth Gollapudi (UC Berkeley PhD) on Long Context LLMs vs Search/RAG: competing technologies or mutually improving?
- Covers his paper "Can Language Models Actually Retrieve In-Context? Drowning in Documents at Million Token Scale" (with UT Austin and UC Berkeley).
- The team trained BlockSearch, a 0.6B-parameter in-context retriever that competes with single-vector dense retrievers up to million-token scale (20K documents of 500 tokens each).
- Podcast also tackles: is RAG dead, how this compares to generative retrieval, and whether it's just a reranker.
Related event: Long-Context LLMs vs. RAG: Rivals or Allies?(2 posts)→
More from Research
- How an internal 1945 EDVAC draft made von Neumann architecture famous — burny_tech · 2026-10-01
- Stanford's Kundaje calls out RNA model renaming: 'classical fine-tuning isn't a new model' — anshulkundaje · 2026-10-01
- SpikingBrain fuses linear attention with spiking neurons for zero-latency edge LLMs — gekobraa · 2026-10-01
- Ben Recht: utility maximization is inescapable—we must learn when it's misapplied — beenwrekt · 2026-10-01
- SCOPD self-distillation recovers 92% of full-context VLM accuracy with 90% fewer visual tokens — CSProfKGD · 2026-10-01
- Xiaomi MiMo-V2.6 report: one mixed GRPO run across all domains, $2.6M for Pro — SergioPaniego · 2026-10-01