Local Book Translation Pipeline: Gemma Translates, Qwen Edits on Dual P40s
neowisard · reddit · 2026-09-02
The author details a high-efficiency local fiction book translation pipeline running on two Tesla P40s, processing 2-3 books daily. The setup uses a division of labor: Gemma-4-26B-A4B for translation and Qwen3.6-35B-A3B for proofreading. Key optimizations include MTP speculative decoding for speed (40-70 tok/s) and specific llama-server flags.
More from coding & agent
- Boomi launches AI Gateway to stop risky agent actions mid-execution — shashib · 2026-09-02
- Anthropic Claims MCP Protocol SDK Hits Hundreds of Millions of Monthly Downloads — emmanuelvivier · 2026-09-02
- Can diagram-driven development make a comeback with LLMs? A Reddit discussion — JBO_76 · 2026-09-02
- AIR Raises $50 Million to Audit AI Agent Skills and Extensions — emmanuelvivier · 2026-09-02
- ICTFax-MCP: An MCP Server to Let AI Agents Send and Track Faxes — ictinnovations · 2026-09-02
- Yegge: Any model will eventually build systems it can no longer maintain — jimmykoppel · 2026-09-02