Modly makes llama.cpp its default agent engine for local 3D mesh generation

Lightnig125 · reddit · 2026-10-06

Open-source desktop app Modly v0.4.3 now ships llama.cpp as its default agent engine: it turns images or prompts into 3D meshes using only local models, plus a chat agent that can operate the app itself.

How it works:

Demo: on an RTX 3060 12GB, Qwen 3.5 4B (Q4KM) decimates a 2.6M-triangle mesh to 300k, calling decimatemesh with correct path and target; 9s with the model loaded, 40s for the first call (server startup + weight loading).

Honest limits: small models sometimes fabricate results — one test stopped above the decimation target (UV seams limit simplification) and the model invented a reason instead of reporting the number; multi-step workflow creation is notably less reliable at 4B than single tool calls. Any OpenAI-compatible endpoint still works as an optional backend. The author asks which ≤8B models are most reliable for tool calling on llama.cpp; their pick so far is Qwen 3/3.5 4B.

Original post →

More from coding & agent

coding & agent channel →