Training from scratch on a single H100 hits 76% on ARC-AGI-1 in ~4 hours

GregKamradt · x · 2026-09-15

Developer kschweig shares that a simple autoregressive transformer with a few twists, trained from scratch and running inference in a bit more than 4 hours on a single H100, scores 76% on ARC-AGI-1. A throughput-optimized recipe reaches 44% (TRM-level performance) in just 17 minutes. Full details are in the attached thread.

Original post →

More from Infra

Infra channel →