Astorias AI launches budget inference service with Qwen at $0.1/M input tokens
me_broke · reddit · 2026-09-22
New startup Astorias AI shipped an inference service aimed at affordability and reliable agent tooling, currently hosting a Qwen-class 27B model at $0.1/M input, $1.6/M output, and $0.05/M cached tokens — among the cheapest tiers around. GLM flash and other flash models are planned next.
Its next feature, "Agent Sandbox," is still in customer research: a serverless container offering (e.g., 8GB RAM / 4 vCPU at $0.015) where agents can run and spin up multiple sandboxes, priced against Novita and Modal. The author is soliciting feedback on what users would actually pay for.
More from coding & agent
- Storewake MCP server puts App Store rankings, RevenueCat revenue and Apple Ads into your AI assistant for $19/mo — Pfernan95 · 2026-09-22
- Free workshop walks through building an LLM Wiki for agent long-term memory — Al_Grigor · 2026-09-22
- Un-fusing a realtime voice stack (STT → LLM → TTS) cut costs 14x — and the real win was text-level guardrails — Cloudsurfer_90 · 2026-09-22
- The last mile of agents: render structured output with Gamma instead of dumping JSON — raw-hit10 · 2026-09-22
- Fan-made Ado chibi pet released for the OpenAI Codex desktop app — secemp9 · 2026-09-22
- Opinion: SaaS becomes the harness and window into agents — guardrails plus visibility — StatisticianKey7858 · 2026-09-22