Modal adds Inkling to Auto Endpoints with DFlash acceleration

Modal has made Thinky Machines’ Inkling available on its Auto Endpoints and said the endpoint uses token-based pricing. What stands out in the available posts is not a new model release, but Modal’s inference setup: according to reposted official claims, the company trained a custom DFlash speculator for Inkling and paired it with SGLang to improve serving performance.

Key details

Multiple posts consistently say Inkling is now usable on Modal. The implementation detail repeated across posts is that Modal trained a custom DFlash speculator specifically for Inkling and runs the setup on SGLang; reposts also say Modal thanked the SGLang team for optimization support.

Performance claims and limits

According to the reposted Modal description, this setup brings 67% higher throughput along with better interactivity. Another repost says Modal described its trained DFlash speculator as faster than MTP for inference. However, the posts do not provide fuller benchmark conditions, baseline configurations, or independent reproduction results, so the current record is limited to Modal’s own stated performance claims.

2026-07-16 ~ 2026-07-16 · 5 related posts