llm-tune: open-source tool finds optimal settings for running local LLMs on your GPU
Distinct-Pie2389 · reddit · 2026-10-09
A developer open-sourced llm-tune, a tool to find the best settings for running local LLMs on your hardware:
- Architectures: Dense, MoE, hybrid MoE/Mamba.
- Hardware: NVIDIA, AMD, Intel Arc GPUs; Apple Silicon via MLX.
- Engines: llama.cpp, Ollama, vLLM.
- Tuning: quantization, context/KV cache, GPU offloading, MTP, sampling, reasoning, and agent harness settings.
- Benchmarks: VRAM/RAM usage, tokens/sec, context recall, output quality.
Currently measured on an RTX 4090; other hardware and backends are documented but need real-world testing. Feedback and benchmark results are welcome.
More from Infra
- Bittensor-based GPU cloud Lium buys back and burns nearly $2.7M of SN51 tokens in six months — markjeffrey · 2026-10-09
- Google open-sources ML Drift: one GPU engine for GLES/OpenCL/Metal/WebGPU, cutting Shorts frame latency 40% — lmoroney · 2026-10-09
- Inference platform doubles providers since launch as competition pushes prices down — AccBalanced · 2026-10-09
- Liquid Inference doubles provider network in one day, undercuts OpenRouter on 30+ models — AccBalanced · 2026-10-09
- SparseEngine: sparse-first inference engine delivers 10x throughput with KV eviction — Jitai Hao · 2026-10-09
- Data centers use ~1 km³ of water a year vs hundreds for households — the clash is location, not volume — AryHHAry · 2026-10-09