No Evals, Building Blind: Contextual Embedding Models Must Rank Disambiguating Chunks

antoine_chaffin · x · 2026-10-01

Antoine Chaffin argues that without evaluations you are building blind: good evals are hard and require real effort to escape usual biases and assess new capabilities. Discussions with @n0riskn0r3ward made them question what worked and what didn't in their model and method, and he believes the effort pays off long term.

The quoted @n0riskn0r3ward shares lessons from contextual embedding models: since voyage-context-3 launched he has used and benchmarked them, finding that for long documents you need them to rank not just the answer chunk but also the disambiguating context chunks highly. Done well, this lets you or your search agent read only the relevant fraction of a long document instead of every page. He started measuring that too, and offered to share benchmark results with the Perplexity team after their contextual model release. Not all modeling teams want third-party feedback, but Perplexity was eager.

Original post →

More from coding & agent

coding & agent channel →