Halo post-training framework claims 2.8x TRL throughput, accused of cherry-picking benchmarks

_ScottCondron · x · 2026-09-23

The new open-source post-training framework Halo claims up to 2.8x TRL throughput with lower peak memory while keeping models in native HuggingFace format. But Axolotl author Wing Lian points out that in single-GPU SFT of GLM Flash on B200s, Halo is actually 5% slower than Axolotl — hence why the team only benchmarked against TRL, which he calls cherry-picked data.

Original post →

More from Infra

Infra channel →