Dev wires up a vLLM reproduction in minutes with Codex, runs it on DGX Spark

remilouf · x · 2026-09-19

remilouf reports having Codex wire up an unoptimized vLLM version using their type safety library in a couple of minutes, running it on a DGX Spark, with plans to retry on real hardware with a base model.

In the quoted tweet he argues an OSS alternative to Jev that is at least as fast with free output tokens is easy to build; the open question is generality and calibration, likely hinging on post-training with clean data. Use cases combining classification and tool calling could go even faster.

Related event: Dev replicates open-source alternative in minutes on DGX Spark(2 posts)→

Original post →

More from coding & agent

coding & agent channel →