Enabling embeddings on Qwen3 27B halves token generation speed in llama.cpp

No_Advance3911 · reddit · 2026-09-15

A user testing Qwen3 27B's built-in embeddings support in llama.cpp reports decent embedding quality, but with a major cost: enabling embeddings cuts token generation speed roughly in half versus running the model normally.

They're asking whether generation performance can be preserved with embeddings on, or whether llama.cpp can switch between tasks without unloading and reloading the model each time — a practical problem for running one local model for multiple workloads.

Original post →

More from Infra

Infra channel →