Ollaya runs decision models locally: 5 answers in 8–10 ms on an RTX 4090, no token generation

petrusenko_max · x · 2026-09-27

Ollaya runs "decision models" locally: typed, calibrated answers in a single forward pass with no token generation. Five questions complete in 8–10 ms on an RTX 4090, vs a 236–276 ms median for TypeSafe's hosted Jev API. The TypeSafe SDK 0.7.1 runs unchanged against localhost:11435, so existing code can switch to local inference with no changes.

Original post →

More from Infra

Infra channel →