Researcher says alleged Phi model behavior examples failed to replicate
BlancheMinerva · x · 2026-09-09
After seeing a thread showing alleged inputs and outputs where Phi and other models behaved in a surprising "by default" way, BlancheMinerva ran the exact prompts on her own machine.
None of the examples replicated the claimed behavior, casting doubt on the evidence in the original claims — a useful reminder to verify such demos before drawing conclusions.
More from Models
- inclusionAI open-sources vision model Ling-3.0-flash-VL with BF16 and FP8 weights — FellMentKE · 2026-09-09
- OpenAI May Already Be Pretraining GPT-6.5 That Beats Astra on Evals — willdepue · 2026-09-09
- After the drama, the real question: is OpenAI's 'internal model' already GPT-6.5 pretraining? — willdepue · 2026-09-09
- Ex-OpenAI researcher disputes Cohere CEO: rewritten ZDR data can't be used for training — willdepue · 2026-09-09
- OpenAI used a model 'significantly more capable' than Astra for Navier-Stokes, per Axios — 141_1337 · 2026-09-09
- Dwarkesh on MagicAILabs' 50x compute claim: RSI may be less compute-bottlenecked than we think — AccBalanced · 2026-09-09