Humanlike Qwen3.8-27B LoRA 2.0 Adds Tool Calls, Fooled a Blind Judge 23.5% of the Time

kvyb · reddit · 2026-10-03

The author released Qwen3.8-27B-Humanlike-Chat 2.0 (v1 got 700+ upvotes, 44k downloads), keeping the human-texting voice while fixing v1's flaws: broken tool calls, ignored instructions, single default personality.

What 2.0 does

Training: v1 was plain SFT on 139,845 messages, which copied bad habits. 2.0 uses on-policy distillation: the model writes its own replies and two teachers grade every token — v1 with a hidden "text like a person" instruction for chat, the plain base model for instructions/tools/code. The student never sees the hidden instruction. Second LoRA merged onto the same 27B.

Numbers: IFBench 37.3→43.7, When2Call 48→58, BFCL irrelevance 60→78; MMLU-Pro and LiveCodeBench slightly worse. Self-built "ishuman" blind benchmark (judge picks which continuation was human-written): base 0.3%, base+system prompt 6.8%, official Qwen3.8 15.1%, 2.0 at 23.5% — you can't prompt your way there. 2.0 won 16/16 live multi-turn chats vs base. GGUFs and LoRA on Hugging Face.

Original post →

More from Models

Models channel →