Liquid AI makes the case for edge agents: 350M-2.6B models with no per-token bill

JosephJacks_ · x · 2026-09-28

At the AWS Builder Loft in San Francisco, Liquid AI's ML engineer Tianshu Yu outlined its edge-agent lineup: text models from 350M to 2.6B parameters, a 3B vision-language model, and audio, runnable on-device via llama.cpp or at scale with vLLM/SGLang and OpenRouter.

Four reasons for the edge: free inference (no per-token bill), ultra-low latency, privacy (healthcare partners keep data on their own servers), and offline operation. Three hard parts: fitting on a phone, running fast on CPU/NPU rather than GPU, and excelling at agentic tasks. His takeaway: "nowadays we care more about the model doing things rather than being a chatbot."

Original post →

More from coding & agent

coding & agent channel →