NInfer fork runs 555k-token context on a single RTX 5090 with custom NVFP4 KV cache and YaRN

Lumpy-Comedian-1027 · reddit · 2026-09-06

A developer released a fork of the NInfer inference engine that dramatically extends local inference for QIn3.8-27B:

Original post →

More from coding & agent

coding & agent channel →