Ramp benchmarks Jev to replace LLM reranking: 10x lower tail latency at 300ms, 3x cheaper
multiply_matrix · x · 2026-09-25
Ashwin at Ramp benchmarked Jev as a replacement for LLM-based reranking in Ramp's accounting product, with promising results: Jev matches current accuracy on GPT-5.6 Luna while cutting tail latency 10x to 300ms at 3x lower cost. The team plans to roll it out to 70k production customers soon, currently only blocked by rate limits.
More from Infra
- Inference startup Jatevo returns: 124B tokens, 1.66M requests, $147K of inference delivered — toptickcrypto · 2026-09-25
- Calibration-Free Quantization Method TQ Open-Sourced, Hits 92.4% Top-1 on Qwen 27B 4-bit — textclf · 2026-09-25
- Agentic AI changes the CPU-to-GPU ratio: 5% CPU allocation cuts token cost ~3.7% — BenBajarin · 2026-09-25
- PreFT Paper Accepted at NeurIPS: Prefill-Only LoRA Adapters Speed Up Multi-Adapter Serving — aryaman2020 · 2026-09-25
- A 32-GPU motherboard for local AI shows up, and people want one in the closet — TheZachMueller · 2026-09-25
- Modern Microprocessors: A 90-Minute Guide Still the Best Crash Course for Systems Engineers — blaizedsouza · 2026-09-25