Redditor Compiles Mega List of Open-Source LLM Inference Optimization Projects and Papers

Dramatic-Chard-5105 · reddit · 2026-09-04

A Reddit user compiled a comprehensive list of recent open-source projects and papers improving open LLM efficiency and accessibility. Hardware/inference projects include colibri (unified VRAM+RAM+storage memory hierarchy for MoE), exo (Mac clustering), MLX distributed-inference experiments, Houmo's DRAM-PIM targeting >1TB/s and 3× efficiency, and d-Matrix 3DIMC. Papers cover edge-distributed MoE inference (WDMoE, MDI-LLM), on-demand expert loading (OD-MoE: 99.94% expert-prediction accuracy at 1/3 GPU memory), and speculative decoding with expert prefetching (MoE-SpeQ up to 2.34×, SP-MoE, MoE-Spec).

Original post →

More from Infra

Infra channel →