Cascade model beats Toto 2.0 4M on TIME MASE/CRPS with 15.4% of the training tokens
const_reborn · x · 2026-10-08
Cascade claims its model outperforms Toto 2.0 4M on the TIME MASE and TIME CRPS time-series forecasting benchmarks, using only 15.4% as many training tokens, entirely synthetic miner-generated data, and a 58-day training run. The thread also breaks down what TIME MASE and CRPS measure and why they matter.
More from Models
- Zero speedup on multiplication, but the leaderboard got nuked — ChrisGPT · 2026-10-08
- Insiders agree: AI models are just not that good at biology yet — nlarusstone · 2026-10-08
- Dev runs 456GB DeepSeek v4.1 on dual GPUs with 192GB VRAM, offloading experts to SSD — HankYeomans · 2026-10-08
- LiquidAI's d1-3B Edge Image-Text-to-Text Model Trends on Hugging Face — LiquidAI · 2026-10-08
- Burkov: frontier models can't vibe-code Photoshop; video data mostly worthless for visual reasoning — burkov · 2026-10-08
- Anthropic's new model priced below DeepSeek and GLM flash, matching them on Terminal Bench 4.0 — op7418 · 2026-10-08