Astorias AI launches budget inference service with Qwen at $0.1/M input tokens

me_broke · reddit · 2026-09-22

New startup Astorias AI shipped an inference service aimed at affordability and reliable agent tooling, currently hosting a Qwen-class 27B model at $0.1/M input, $1.6/M output, and $0.05/M cached tokens — among the cheapest tiers around. GLM flash and other flash models are planned next.

Its next feature, "Agent Sandbox," is still in customer research: a serverless container offering (e.g., 8GB RAM / 4 vCPU at $0.015) where agents can run and spin up multiple sandboxes, priced against Novita and Modal. The author is soliciting feedback on what users would actually pay for.

Original post →

More from coding & agent

coding & agent channel →