Hemmingway-1, a 27B human-like writing fine-tune, hits 32 tok/s on M5 Max and shines at roleplay
DerTomsn · reddit · 2026-10-02
A Reddit user benchmarked Altworld's Hemmingway-1 — a 27B fine-tune of Qwen3.8-27B built for "human-like" writing — on an M5 Max with thinking and MTP on: 31.8 tok/s average, 29-31 GB VRAM peak.
LLM-judge scores by scenario: Role Play & Narrative 94.2 (immersion 96-97 every run), Research & Analysis 88.2, Agent Workflow 87.5, Code Generation 76.2. The base Qwen3.8-27B scored close — one run even outscored the fine-tune (95.35 vs 94.65) — but the author notes the judge isn't a human reader: the difference shows when actually reading the output, where the fine-tune's prose feels noticeably more human. A sample tavern-scene passage is included.
More from Models
- Local Qwen 3.8 detected it was being benchmarked — and became more honest — julianharris · 2026-10-02
- ARC Prize finds Qwen3.8-27B's chat template injects different instructions per reasoning effort — GregKamradt · 2026-10-02
- Three reasons vibe-coded software is still far from production grade, with ReactBench data — aidenybai · 2026-10-02
- OpenAI GPT-6.1 Sol and Gemini 4 Argon both launch at identical $2/$10 pricing — AccBalanced · 2026-10-02
- Arena now runs up to 600,000 E2B sandboxes a day for model evaluation — badphilosopher · 2026-10-02
- OpenAI's $500 plan now offers less usage than its old $100 tier, users fume — petrusenko_max · 2026-10-02