Ling-3.0-tiny Runs 128K Context on $249 8GB Orin Nano

Puzzleheaded_Base302 · reddit · 2026-08-19

A technical report shows inclusionAI's Ling-3.0-tiny running on an 8GB NVIDIA Jetson Orin Nano Super ($249). Using IQ4NL quantization (4.5 bpw), it achieves the full 131,072 token context window with 33 tok/s decode speed and 7.4GB RAM usage. Details include the required llama.cpp build (BailingMoE V3 support), compilation flags, and performance benchmarks.

Original post →

More from Infra

Infra channel →