iPad + AMD R9700 eGPU Runs Qwen3.8-27B at 159 tok/s via Open-Source LSE Engine

TheOriginalG2 · reddit · 2026-10-09

LemonSeed Studio demos on-device inference on an iPad Pro with an AMD Radeon AI PRO R9700 in a Thunderbolt enclosure, running Qwen3.8-27B Q4 at 131k context. Decode hits 158.5 tok/s with DFlash2 speculative decoding (31.2 baseline), nearly identical across iPadOS/macOS/Linux, while Strix Halo tops out at 64.4. The LSE engine embeds unmodified Linux amdgpu/amdkfd drivers as a PCIDriverKit extension, records forward passes as graphs, fuses ops, and picks the fastest kernels by on-device measurement. Open source, with an OpenAI-compatible server, CLI, and C library.

Original post →

More from Infra

Infra channel →