Voice agent dev seeks a small dedicated tool-calling model to cut LLM latency hops

Snoo_7134 · reddit · 2026-09-21

A Reddit developer building a conversational voice AI agent says its main LLM takes multiple hops between the model and the harness per turn for tool calling, adding latency. The plan: run a smaller, faster model purely for tool calling before the main call, then pass context along — a "hop minimisation" approach.

They've tested Jev and a fine-tuned GLiNER 2.5 model, but found both inadequate at tool calling and are looking for alternatives. The post is a concrete problem definition plus tested-and-rejected options, useful reading for anyone building low-latency agents.

Original post →

More from coding & agent

coding & agent channel →