LLM vs classical ML across 8 datasets: labeled data still favors SVM and XGBoost
Ok_Juggernaut2187 · reddit · 2026-09-20
The author benchmarked Jev in zero-shot and few-shot settings against 11 classical ML models (SVM, XGBoost, logistic regression, etc.) on 8 classification datasets using balanced accuracy. Takeaways: with labeled data, classical models win — they beat Jev on Banking77 and business tabular data while avoiding per-call API costs; Jev only did well on IMDb; few-shot examples gave no consistent boost; and there's no reason to replace a working ML model. Setup: 3 training seeds, fixed test split, limited tuning budget, with zero-shot predictions cached across seeds. Full results, code and notebook are open-sourced.
More from Models
- Researchers Use Claude Opus 5 to Hijack OpenAI Employee Accounts for Under $3,000 — FinanceYF5 · 2026-09-20
- NetEase Youdao open-sources Confucius4-R2T2 streaming ASR with 200-600ms latency — aigclink · 2026-09-20
- Alibaba open-sources DAMO RADAR medical model detecting cancer and ~150 conditions — emmanuelvivier · 2026-09-20
- Anthropic reportedly testing Claude Money to link bank accounts for spending analysis — emmanuelvivier · 2026-09-20
- Mozilla report: open-weight AI models now just four months behind the closed frontier — mark_k · 2026-09-20
- GPT Pro user finds model browses completely irrelevant sources, questions real retrieval — CompetentRaindeer · 2026-09-20