No Evals, Building Blind: Contextual Embedding Models Must Rank Disambiguating Chunks
antoine_chaffin · x · 2026-10-01
Antoine Chaffin argues that without evaluations you are building blind: good evals are hard and require real effort to escape usual biases and assess new capabilities. Discussions with @n0riskn0r3ward made them question what worked and what didn't in their model and method, and he believes the effort pays off long term.
The quoted @n0riskn0r3ward shares lessons from contextual embedding models: since voyage-context-3 launched he has used and benchmarked them, finding that for long documents you need them to rank not just the answer chunk but also the disambiguating context chunks highly. Done well, this lets you or your search agent read only the relevant fraction of a long document instead of every page. He started measuring that too, and offered to share benchmark results with the Perplexity team after their contextual model release. Not all modeling teams want third-party feedback, but Perplexity was eager.
More from coding & agent
- Private AI Proxy for Mac verifies AI services via hardware attestation before agent code leaves your machine — bgmshana · 2026-10-01
- Musk touts new Grok Bot: pair it with Cursor cloud agents to manage coding in Slack — elonmusk · 2026-10-01
- How to generate motion graphics with Claude Code in 6 steps, driven purely by time — aakashgupta · 2026-10-01
- Survey aims to map how professional software engineers actually use coding agents — thegautamkamath · 2026-10-01
- OpenAI engineer runs all her engineering work through her dot agent — jxnlco · 2026-10-01
- 17 open-source GitHub repos to learn AI engineering from scratch — ZabihullahAtal · 2026-10-01