Early Jev test: model fails lat-long regression, suggested as a location-understanding eval
MikkoH · x · 2026-09-17
The author tested the new-architecture model Jev on latitude-longitude regression and found it couldn't handle it even with ten criteria. But he suggests the task could make an interesting eval for whether Jev genuinely understands locations—an early probe into the new architecture's capability boundaries.
More from Models
- Linkup open-sources SPARSEUP, a <150M sparse retriever scoring 56+ on BEIR-13 — antoine_chaffin · 2026-09-17
- OpenAI reportedly close to solving Millennium Prize problem, the Hodge Conjecture — Hesamation · 2026-09-17
- Liquid AI's small Longevity models beat GPT-5 and Claude Opus on Cell-published aging benchmarks — helloiamleonie · 2026-09-17
- Polymarket puts just 8% odds on a Chinese company topping the Chatbot Arena leaderboard — Polymarket · 2026-09-17
- Debate: how did Anthropic hire all the alignment talent yet ship models 'worse aligned' than GPT? — almmaasoglu · 2026-09-17
- Pro RL training halted by cluster-grader network issue plus bad patterns from cyber dataset; now resumed — _AndrewZhao · 2026-09-17