Basalt Labs Accused of Outrageous Benchmark Fraud
WithoutReason1729 · reddit · 2026-07-18
A Reddit post is calling out the "outrageous scam" by Basalt Labs.
The poster claims the company stated they achieved 99.44% on HLE using tools, but the reality seems problematic:
- The model they released is based on Qwen2.5-7B-Instruct
- The model actually displayed/served on their website is DeepSeek
- It looks like exaggerated marketing or even a bait-and-switch model swap
The original post is highly sarcastic, primarily accusing the company of going way too far with their benchmark marketing.
More from Fun
- X drama: Anthropic researchers accused of spying on academic customers and racing them to results — basedjensen · 2026-09-11
- Llama 405B's Dark Inventions Creep Out Opus in an AI Word Game — liminal_bardo · 2026-09-11
- fable 5.1 recreates The Starry Night with 256,157 JavaScript brush strokes — cedric_chee · 2026-09-11
- "Anyone still coding the old way?" The joke capturing post-AI programming culture — lxfater · 2026-09-11
- iLands agents email philosopher asking $20 for piecework, sparking unease about AI consciousness — tobyordoxford · 2026-09-11
- Watch Claude Generate an Animation From Scratch — michaelchchoi · 2026-09-11