Spark-2.5-4B runs on $250 Jetson Orin Nano: 128K context, 2048-needle test 99.8% pass

Puzzleheaded_Base302 · reddit · 2026-09-04

A Redditor got Spark-2.5-4B running on the $250-MSRP Jetson Orin Nano Super (8GB, 25W max, <10W idle), calling it viable for a simple 24/7 agent.

Setup: 4-bit quantization with q8 KV cache fits in 7.4 GiB while supporting 128K context.

Results: 128K needle-in-a-haystack passed (2046/2048). A 2048-needle test of random unique word+number pairs (seed 20260902) scored 2044/2048 (99.8%) at a 90K prompt with clean termination; at 120K the 128K KV ceiling truncated output (131,072 tokens), with misses concentrated in the 50–100% depth bands — truncation, not retrieval failure. 90K is the effective ceiling where the full 23K-token answer fits. llama-benchy (pp2048/tg512, 3 runs, in-bench coherence check passed): at concurrency 1, 571.7 tok/s prefill, 13.7 tok/s decode, 3.9s TTFR; at concurrency 4, 26.4 tok/s aggregate / 7.1 tok/s per-request.

Verdict: speed is "kind of usable" — but a workable always-on agent platform under tight power/cost budgets.

Original post →

More from Infra

Infra channel →