Redditor Compiles Mega List of Open-Source LLM Inference Optimization Projects and Papers
Dramatic-Chard-5105 · reddit · 2026-09-04
A Reddit user compiled a comprehensive list of recent open-source projects and papers improving open LLM efficiency and accessibility. Hardware/inference projects include colibri (unified VRAM+RAM+storage memory hierarchy for MoE), exo (Mac clustering), MLX distributed-inference experiments, Houmo's DRAM-PIM targeting >1TB/s and 3× efficiency, and d-Matrix 3DIMC. Papers cover edge-distributed MoE inference (WDMoE, MDI-LLM), on-demand expert loading (OD-MoE: 99.94% expert-prediction accuracy at 1/3 GPU memory), and speculative decoding with expert prefetching (MoE-SpeQ up to 2.34×, SP-MoE, MoE-Spec).
More from Infra
- Marin's open 535B-A23B model is 13% trained, funded by Huang Foundation — dlwh · 2026-09-04
- Built a Dual RTX 6000 Pro Rig for Local DeepSeek — Warns Against Influencer Build Hype — HankYeomans · 2026-09-04
- Pinokio 8.2.0 Ships Universal Disk Saver and Nested Folder Support — cocktailpeanut · 2026-09-04
- Inference engines are an underexamined attack surface, self-hosting ops warned — JeremyCMorgan · 2026-09-04
- browser-llm-fit: check if an AI model fits your browser before downloading weights — init0 · 2026-09-04
- YC S26 Demo Day Next Week: Floating Data Centers, Diamond Semiconductors, Bio Computers — ycombinator · 2026-09-04