Local Book Translation Pipeline: Gemma Translates, Qwen Edits on Dual P40s

neowisard · reddit · 2026-09-02

The author details a high-efficiency local fiction book translation pipeline running on two Tesla P40s, processing 2-3 books daily. The setup uses a division of labor: Gemma-4-26B-A4B for translation and Qwen3.6-35B-A3B for proofreading. Key optimizations include MTP speculative decoding for speed (40-70 tok/s) and specific llama-server flags.

Original post →

More from coding & agent

coding & agent channel →