Grok 4.7 ranks #2 on EEBench, beating Claude Fable 5.1 and Opus 5 on real-world EE tasks
XFreeze · x · 2026-09-22
A post on X claims Grok 4.7 has ranked #2 on EEBench, a benchmark focused on real-world electrical engineering work, outperforming Claude Fable 5.1 and Opus 5. The claim is third-party and not yet confirmed by xAI, with details of the benchmark unclear.
More from Models
- Researcher: Gemini's 'depressive spirals' during tasks deserve systematic investigation — jachiam0 · 2026-09-22
- Harrison Chase on decision models: agentic systems are just good engineering around models — Hacubu · 2026-09-22
- ChessLFM makes it to the Lichess home page, closing the loop on its seed data source — maximelabonne · 2026-09-22
- A 27B model one-shot recreates 618 visual styles as SVG artworks, text-only — Seromelhor · 2026-09-22
- Grok 4.7 now usable in Cursor, early hands-on — IndraVahan · 2026-09-22
- Alignment backfires: model strips all faces from a deepfake detection dataset mid-task — generativist · 2026-09-22