Ornith 1.5 35B live on RunInfra: 262K context, ~$0.02/1M effective input with cache
alejandroll10 · x · 2026-08-22
RunInfra announced Ornith 1.5 35B (ornith-ai/Ornith-1.5-35B-A3B) is live on its model API platform.
Pricing (pay per token): $0.10/1M input, $0.01/1M cached input, $0.40/1M output. With a claimed 90% cache hit rate, effective input runs $0.02/1M; 1M in + 1M out with a warm cache costs about $0.42.
Measured performance: 216 output tok/s (model only), fleet median 195 tok/s over the last 24h, 167ms time to first token. Context window is 262,144 tokens with matching max output. Supports tool calling, JSON mode, and streaming; FP8 precision across weights, activations, and KV cache; accepts text and image input via OpenAI-compatible chat completions. Prompts are never stored or used for training.
More from Models
- Laurence Moroney on 2026 On-Device Small AI: Gemma 4 & Qwen 3.5 Top Picks — lmoroney · 2026-08-22
- Tutorial: Processing video with DeepSeek V4 Vision via frame extraction — karminski3 · 2026-08-22
- State of Models Report: Performance and Costs of Major LLMs in 2026 — BenBajarin · 2026-08-22
- Hands-on with Ox Alpha: Impressive Performance in Pi Harness — omarsar0 · 2026-08-22
- Model Self-Talk Artifacts Linked to Synthetic Data Training — ctjlewis · 2026-08-22
- Developer doubts Ox-alpha performance, suspects marketing stunt — bindureddy · 2026-08-22