Full Leaderboard Released for LLM Backend Coding Benchmark (Vector DB from Scratch)
karminski3 · x · 2026-09-14
The author shares the full leaderboard and charts for the earlier LLM backend-coding benchmark (building a vector database from scratch, scored by DB performance).
Key conclusions as before: Fable-5.1 peaks highest (21191.51) but swings over 50% across three runs, GPT6-Astra is the most stable (12985.28 / 12134.83 / 10676.74), and DeepSeek-V4.1-Flash is the best value, nearly matching Kimi-K3 with <20% variance.
More from Models
- Kokoro TTS ported to Apple Core AI: 54 voices running fully on-device with zero API cost — amos_gyamfi · 2026-09-14
- "Everyone is cheating on AI benchmarks": a hacker's plea to optimize for the real world — hackgoofer · 2026-09-14
- Dev Says 'Astra' Via OpenRouter Burned $30 in About 9 Seconds — haydendevs · 2026-09-14
- Dev Shares Token Split: Open Models Outused Closed 10:1 for Coding and Research — xeophon · 2026-09-14
- Dario Amodei: Claude has spotted medical issues doctors missed, big gains in biology — Olivier__OG · 2026-09-14
- Codex Users Report Capacity Woes as OpenAI's Compute Strain Looks Real — mark_k · 2026-09-14