friend.cpp: experimental local LLM engine adds blue noise sampling and adaptive speculative decoding
amplifiedamp · x · 2026-09-23
A developer released friend.cpp, an experimental local LLM inference engine forked from koboldcpp, featuring adaptive speculative decoding, a tiered prompt cache, per-request adapter hot-swap, and steering. The author also proposes blue noise sampling: generating stable, coherent continuations from models prone to mode collapse (including base models) without reducing temperature. A fork of llama.cpp is available to try it out.
Related event: Blue noise sampling reduces mode collapse in LLMs(2 posts)→
More from Infra
- Qualcomm scales data-center High Bandwidth Compute memory down to phones, laptops and smart glasses — samcharrington · 2026-09-23
- Qualcomm's data-center HBC tech is coming to Snapdragon phones, laptops and glasses — ryanshrout · 2026-09-23
- NVIDIA pitches Confidential Computing for running AI on sensitive enterprise data — nvidia · 2026-09-23
- The AI Flywheel Closes on Chips and Robots, but the Grid Becomes the Real Bottleneck — r0ck3t23 · 2026-09-23
- Dev laments agents built around KV caches, wants inference-first chips — dbreunig · 2026-09-23
- GE Vernova seen hitting $200B backlog by early 2027 as turbine demand outruns guidance — BenBajarin · 2026-09-23