Ollama's jmorgan: small models now handle most conversational and reasoning use cases
ollama · x · 2026-09-14
Ollama co-founder jmorgan praised the Stanford Hazy Research work by Avanika Narayan, Jon Saad-Falcon and team (with John Hennessy involved), saying small models are "truly capable of the majority of conversational use cases and even a majority of difficult reasoning ones." Their work was featured prominently in the Financial Times, arguing the world is becoming less dependent on centralized cloud AI — a signal of the shift toward local, on-device small models.
More from Infra
- agi-memory: SQLite-only persistent memory MCP server for coding assistants, 32MB RAM — Rude_Gate7599 · 2026-09-14
- $3000 home server with 128GB VRAM runs Qwen3.8-next at 1.3k tps prefill, 70 tps code — Thin_Pollution8843 · 2026-09-14
- Top 10 foundry revenue hits $53.5B in Q2, TSMC holds 72.5% share — Beth_Kindig · 2026-09-14
- Qwen3.8 27B INT4 With 144K Context Runs on a Single RTX 3090 via vLLM — Altruistic_Heat_9531 · 2026-09-14
- Nvidia's $59.7B Quarterly Net Income Works Out to ~$656M Profit Per Day — himanshustwts · 2026-09-14
- Tiny Neural Nets Revival: Transformer Hits ~1500 tok/s on M4 CPU via SME2 Instructions — GregoryDiamos · 2026-09-14