Perf engineer praises memory-overcommit evals, suggests exposing stats to fleet scheduler

sloppenheimer · x · 2026-09-24

The author highlights the evaluations section of a deployment writeup, noting that memory-overcommit behavior over time looks clean and largely contention-free. He suggests exposing these stats to the scheduler could unlock better fleet-level distribution wins at scales above 160 nodes. He also recalls past struggles with NVIDIA MIG in vGPU/microVM setups, says the FnCall approach sidesteps them, and wonders whether an nvidia-smi-based external reset exists for when an agent trashes a GPU.

Related event: Users Critique Systems Paper's Evaluation, Missing gVisor Benchmarks(2 posts)→

Original post →

More from coding & agent

coding & agent channel →