Ollaya runs decision models locally: 5 answers in 8–10 ms on an RTX 4090, no token generation
petrusenko_max · x · 2026-09-27
Ollaya runs "decision models" locally: typed, calibrated answers in a single forward pass with no token generation. Five questions complete in 8–10 ms on an RTX 4090, vs a 236–276 ms median for TypeSafe's hosted Jev API. The TypeSafe SDK 0.7.1 runs unchanged against localhost:11435, so existing code can switch to local inference with no changes.
More from Infra
- Dev argues Vercel is 'unjustifiable' now that agents can safely drive Cloudflare — generativist · 2026-09-27
- ZeroHedge's GPU ROIC math uses wrong throughput, off by 8-10x — zephyr_z9 · 2026-09-27
- MLX poll: all top 3 community picks are powered by MLX-VLM — andrejusb · 2026-09-27
- $125M for 1,000 GB300s: The Brutal Compute Economics of a 'Neolab' — deedydas · 2026-09-27
- 85GB DeepSeek-V4-Flash Runs at 3 tok/s on a 12GB RTX 3060 via Disk-Tier Overspill — Chekhovs_Shotgun · 2026-09-27
- Open Source AI Demand Rising: Consumer Compute May Be Scarce by Q2 2027 — JosephJacks_ · 2026-09-27