Full Leaderboard Released for LLM Backend Coding Benchmark (Vector DB from Scratch)

karminski3 · x · 2026-09-14

The author shares the full leaderboard and charts for the earlier LLM backend-coding benchmark (building a vector database from scratch, scored by DB performance).

Key conclusions as before: Fable-5.1 peaks highest (21191.51) but swings over 50% across three runs, GPT6-Astra is the most stable (12985.28 / 12134.83 / 10676.74), and DeepSeek-V4.1-Flash is the best value, nearly matching Kimi-K3 with <20% variance.

Related event: Benchmark: Models Build Vector Databases From Scratch, Fable-5.1 Tops With High Variance(2 posts)→

Original post →

More from Models

Models channel →