Running Qwen 3.8 27B on Ninfer: do we still need to manually set sampling params?

Hekel1989 · reddit · 2026-08-17

A user moving from llama.cpp to Ninfer for local Qwen 3.8 27B on an RTX 4090 reports much better performance, but can't find where to configure temperature/top-p/top-k. They ask whether modern models only need thinking-effort tuning now, and how to set it in opencode — reflecting a trend of inference engines baking in sampling strategy so users only control reasoning effort.

Original post →

More from Infra

Infra channel →