Orange Pi 5 Plus LLM benchmarks: core pinning gives +318%, NPU hits 21.55 tok/s
No-Doughnut6532 · reddit · 2026-10-06
A detailed r/LocalLLaMA benchmark of the Orange Pi 5 Plus (RK3588, 16GB LPDDR4x) for running small local LLMs 24/7:
- Core pinning is critical: restricting Ollama to 4 A76 big cores yields up to +318% speedup vs default 8 threads stalling on A55 little cores.
- CPU inference (4T): DeepSeek-Coder 1.3B at 16.9 tok/s, Llama 3.2 1B at 14.6, Qwen 2.5 1.5B at 14.5, Llama 3.2 3B at 7.3, Phi-3 Mini 3.8B at 6.6 tok/s.
- 8B memory wall: Llama 3.1 8B drops to 2.3 tok/s at 85°C; measured LPDDR4x bandwidth (25-30 GB/s) is the hard ceiling.
- NPU: native RKLLM on the 6 TOPS NPU runs Qwen 1.5 0.5B at 21.55 tok/s with sub-100ms TTFT at 0% CPU load.
- Thermals/storage: ships bare-die; idle 52.7°C, 1B-3B at 68-74°C, 8B throttles at 85°C; NVMe reads 2862 MB/s, 1.3GB model loads in 0.6s.
Unit provided free by Orange Pi with no editorial preconditions.
More from Infra
- Google Buys 890 MW of Nuclear Without Building a Single New Reactor — MicahBerkley · 2026-10-06
- Cycle.io Launches DevOps MCP: 3-Node Mongo Replica Set Across 3 Clouds in 15 Minutes — AlexMattoni · 2026-10-06
- Weaviate ships query profiling: a 48ms slow query turned out to be disk reads, not HNSW — CShorten30 · 2026-10-06
- HN Debate: Did Oracle Just Trigger the Implosion of the AI Bubble? — mpweiher · 2026-10-06
- Lambda adopts NVIDIA's AIPerf for model cards showing real-workload inference benchmarks — TheZachMueller · 2026-10-06
- Nvidia nears $6 trillion market value as AI frenzy keeps pushing stocks higher — AryHHAry · 2026-10-06