Exllamav3 实测碾压 Llama.cpp,双 3060 推理速度倍增
Ecstatic-Wash-7667 · reddit · 2026-08-17
用户在双 RTX 3060 显卡配置下对比了 Exllamav3 和 Llama.cpp 的性能。测试覆盖 Qwen 3.8 27b 和 Qwen 3.6 35b 两个模型,进行 3 次冷启动基准测试。结果显示 Exllamav3 在 tokens/sec (tg) 指标上显著优于 Llama.cpp,差距在特定测试中极为悬殊。
「Infra」频道最新
- Stripe 70 亿美元收购 OpenRouter,卖铲子的先赢了 — sven_ai · 2026-08-17
- AI 支出泡沫失控:云厂商未来承诺破 3 万亿美元 — GaryMarcus · 2026-08-17
- AI能耗巨大,如何从物理与伦理实现净正面? — AryHHAry · 2026-08-17
- Merge Gateway 接入 Grok 4.6 并提供限时折扣 — shensi · 2026-08-17
- 对比 20 年前,现代开发者的每一次提交都触发庞大算力 — andreisavu · 2026-08-17
- 调查指微软 AI 芯片缺口达数百万颗 — nordicinst · 2026-08-17