Enabling embeddings on Qwen3 27B halves token generation speed in llama.cpp
No_Advance3911 · reddit · 2026-09-15
A user testing Qwen3 27B's built-in embeddings support in llama.cpp reports decent embedding quality, but with a major cost: enabling embeddings cuts token generation speed roughly in half versus running the model normally.
They're asking whether generation performance can be preserved with embeddings on, or whether llama.cpp can switch between tasks without unloading and reloading the model each time — a practical problem for running one local model for multiple workloads.
More from Infra
- VMware Private AI Foundation with NVIDIA Tops Private AI Cloud Evaluation at 9.1/10 — DavidLinthicum · 2026-09-15
- SpaceX President: Compute Rental Is a Great Business and AI Demand Shows No Slowdown — XFreeze · 2026-09-15
- Orthrus Study: Lossless Speculative Decoding Holds Only at High Numerical Precision — Ilya Koziev · 2026-09-15
- DeepSeek V4.1 Flash on M3 Ultra nearly doubles decode to 31 t/s with first public DSpark Metal port — IngeniousIdiocy · 2026-09-15
- Trigger.dev's chat.agent turns AI chats into durable tasks that survive crashes and redeploys — CodeByPoonam · 2026-09-15
- Oracle Executes Pre-Dawn Mass Layoffs as AI Data Center Spending Balloons — 量子位 · 2026-09-15