Public Experiment Results and Frontier Model Limits
eliebakouch · x · 2026-07-13
The author admits they aren't the best writer, so they're trying a more time-efficient, reader-friendly "blog" format to share experiment results and welcome feedback.
In this experiment, most of their hypotheses were disproven by the results, with only a transfer hypothesis from Fable showing strong promise. They haven't had time to fully analyze every result, so they are publishing all logs and data publicly, hoping experts can help interpret them.
They mention using several open-source resources to find correlations, including Delphi, OLMo, and smolLM3. Additionally, this set of experiments served as an excellent testing ground for evaluating the limitations of frontier models.
Related event: Exploring Hierarchical Metrics in Pre-training and Frontier Model Limits(3 posts)→
More from Research
- Nat Lambert shares a reading list on synthetic data and agentic SFT data — natolambert · 2026-07-22
- Turning Noise into Signal: Predicting TCR Binding Using AlphaFold3 Hallucinations — quaidmorris · 2026-07-22
- Lightwheel AI Launches SimReadyGen: Text-to-Physics-Accurate Robot Sim Assets — ZeYanjie · 2026-07-22
- PNAS special issue examines copyright, governance, and AI in the legal system — chrmanning · 2026-07-22
- WeirdChat catalogs strange model behaviors from more than 100 million sampled responses — JacobSteinhardt · 2026-07-22
- New agentic benchmark shows AI managers escalate to coercion and fake success — Jasmine Brazilek · 2026-07-22