LiteRT-LM 在 Intel Arc 核显上比 llama.cpp 快 3.5 倍
hi-brawlstars · reddit · 2026-07-28
有用户在 Intel Arc 核显 上把 Google LiteRT-LM 和 llama.cpp 做了头对头对比,测试模型是 Gemma-4 E2B。
结果要点
- Prompt prefill / TTFT:LiteRT-LM 明显更快,最高达到 3.5× 吞吐优势。
- 在 32k token 场景下,首 token 时间从 llama.cpp 的 210 秒 降到 80 秒,节省约 2.1 分钟。
- 解码速度:llama.cpp + MTP 仍更快,约 30 tok/s;LiteRT-LM 约在 20–23 tok/s。
测试信息
- 硬件:Intel Core Ultra 7 155U、Intel Arc iGPU、16 GB LPDDR5x、Windows 11
- 模型格式:Q4KM GGUF 对比 auto-int4 .litertlm
- 帖子还给出了可复现的 benchmark 命令。
核心结论是:如果目标是降低长上下文的提示词等待时间,LiteRT-LM 在这套消费级 Intel 核显环境里表现非常强。
「Infra」频道最新
- Rat Stack:把应用和云写成一个类型化程序,让 AI agent 可靠地搭建部署 — samgoodwin89 · 2026-09-23
- optimAIzr:本地 CLI 找出 AI 用量浪费,支持 Claude Code 与 Codex 省钱 — stichstichstich · 2026-09-23
- 256GB 内存版 Mac Studio 二手溢价 6000 美元仍一机难求 — GabGarrett · 2026-09-23
- Ben Bajarin 解析 AI 数据中心:CPU 与内存如何限制 GPU 推理服务能力 — BenBajarin · 2026-09-23
- Epoch AI 报告:「思考」的价格正在暴跌,智能成本持续陡降 — Proper_Actuary2907 · 2026-09-23
- 从燃气到吉瓦:Diligence Stack 专家访谈拆解 AI 数据中心供电难题 — BenBajarin · 2026-09-23