Eval: New Model Hits 91.4 NDCG@10, Leading by 6.2 Points
tomaarsen · x · 2026-08-26
Evaluated against 47 model configurations (dense, sparse, lexical, and multi-vector) on 1,000 held-out medical questions and 200,000 passages. The new model achieves a 91.4 NDCG@10, leading the best zero-shot model of any architecture by 6.2 points.
More from Apps
- Indie Game Prism Rules: Real-time Light Tracing with WebGL — measure_plan · 2026-08-27
- Grok Prompt: Generating Gigantic Forced-Perspective Ad Photography — aziz4ai · 2026-08-27
- Forensically: Free browser-based image forensics toolkit for detailed analysis — petewoodbridge · 2026-08-27
- Refine: AI verification tool trusted by world-class researchers — littmath · 2026-08-27
- Open Source AI Peer Review Tool 'coarse' Costs <$2, Outperforms Competitors — littmath · 2026-08-27
- .NET 11 Preview Released: C# 15 Union Types and MAUI CoreCLR — unixterminal · 2026-08-27