Same open weights, dumber outputs: 3 ways inference hosts quietly downgrade you

lucasbennett_1 · reddit · 2026-09-01

A Reddit user details how the same open weights can feel noticeably dumber on one host than another with zero errors, due to three silent downgrades:

Before committing, check whether the model page states precision and served context (e.g. DeepInfra's ds v4 pro listing shows fp4 and 66k), cross-check on the Artificial Analysis leaderboard, and run the same reasoning + tool-call prompt on two endpoints, keeping a local full-precision quant as reference.

Original post →

More from Infra

Infra channel →