Dev wires up a vLLM reproduction in minutes with Codex, runs it on DGX Spark
remilouf · x · 2026-09-19
remilouf reports having Codex wire up an unoptimized vLLM version using their type safety library in a couple of minutes, running it on a DGX Spark, with plans to retry on real hardware with a base model.
In the quoted tweet he argues an OSS alternative to Jev that is at least as fast with free output tokens is easy to build; the open question is generality and calibration, likely hinging on post-training with clean data. Use cases combining classification and tool calling could go even faster.
Related event: Dev replicates open-source alternative in minutes on DGX Spark(2 posts)→
More from coding & agent
- Dev hides AI thinking notes to stop spoilers as GLM-5.3 Flash stays on cost Pareto frontier in chess app — MikePFrank · 2026-09-19
- An AI agent called customer service for a user — the rep couldn't tell it wasn't human — armand_ruiz · 2026-09-19
- GuppyLM: Train a 9M-parameter LLM from scratch in 5 minutes with one Colab notebook — tom_doerr · 2026-09-19
- The real headache of browser-controlling agents: sites keep breaking your selectors — SelfZealousidealy · 2026-09-19
- PowerShell MVP builds an AI-driven lab that turns ideas into runnable, validated scripts daily — dfinke · 2026-09-19
- An agent auto-collected 382 Jev demos overnight into a community directory — thisiskp_ · 2026-09-19