Ling-3.0-tiny Runs 128K Context on $249 8GB Orin Nano
Puzzleheaded_Base302 · reddit · 2026-08-19
A technical report shows inclusionAI's Ling-3.0-tiny running on an 8GB NVIDIA Jetson Orin Nano Super ($249). Using IQ4NL quantization (4.5 bpw), it achieves the full 131,072 token context window with 33 tok/s decode speed and 7.4GB RAM usage. Details include the required llama.cpp build (BailingMoE V3 support), compilation flags, and performance benchmarks.
More from Infra
- Qwen 3.8 in 4-bit hits 25 tok/s on an M3 Max — fully usable locally — AIandDesign · 2026-08-19
- Hygon Profit Soars; LG Partners with Nvidia for Humanoid Robot — 创业邦 · 2026-08-19
- Building Real Offline AI: Local Agent with Cognitive Loops — HotEstablishment7184 · 2026-08-19
- OpenAI uses ~20% compute for inference during training — eliebakouch · 2026-08-19
- Vercel KMS Lets You Sign JWTs Without Managing Private Keys — cramforce · 2026-08-19
- Apple's Foundation Model Framework: Hybrid AI Routing with Dynamic Profiles — Scobleizer · 2026-08-19