Dev discovers vLLM accepts input embeds, speculates finetuning tokens are underused
cephaloform · x · 2026-10-09
开发者 @cephaloform 表示此前因 vLLM 接口变动太快而回避使用,现在才发现 vLLM 支持传入 input embeds,打算实际尝试。他在讨论中提出一个猜想:社区可能严重低估了 prompt tuning 类微调 token(以及 DreamBooth 式高显存配置)的潜力,如果推理框架普遍支持接受 inputsids/embeds,结合进化算法或许能打开新玩法。
More from Infra
- tinygrad runs its GitHub Actions CI on 4 new tinybox machines — AIFlow_ML · 2026-10-09
- OpenAI's 10,000-agent, 130B-token run pushed slime v0.4.0 to rethink RL infrastructure scale — teortaxesTex · 2026-10-09
- TokenRouter Serving System Boosts Token-Level LLM Routing Throughput up to 64x — nics-efc · 2026-10-09
- SparseDecoding: Decoding-Aware Pruning Yields up to 1.48x Faster LLM Inference — encodelab · 2026-10-09
- x86_64 Edge runs with GPU acceleration on Arduino VENTUNO Q via FEX — unixterminal · 2026-10-09
- Edge0: a 35B model runs on an iPhone in airplane mode with just 1.8GB peak memory — FinanceYF5 · 2026-10-09