Running Qwen 3.8 27B on Ninfer: do we still need to manually set sampling params?
Hekel1989 · reddit · 2026-08-17
A user moving from llama.cpp to Ninfer for local Qwen 3.8 27B on an RTX 4090 reports much better performance, but can't find where to configure temperature/top-p/top-k. They ask whether modern models only need thinking-effort tuning now, and how to set it in opencode — reflecting a trend of inference engines baking in sampling strategy so users only control reasoning effort.
More from Infra
- LTX-2.5 vs MiniMax H3 i2v on an RTX 5090: 1080p vs 1344×768 is what 32GB fits — chanteuse_blondinett · 2026-08-17
- GitHub Outage: Web and API Error Rates Exceed 20% — jonathan_wilke · 2026-08-17
- Building AI for quantum on classical computers is like designing jets with sailboat tech — AryHHAry · 2026-08-17
- Stripe's OpenRouter acquisition builds the stack for autonomous machine payments — 0xSammy · 2026-08-17
- PyTorchCon NA Lands Oct 20-21 With a Dedicated Inference Track — PyTorch · 2026-08-17
- Amazon Tracked Destroying Rare Books for AI Training Data: 404 Media — james_mtc · 2026-08-17