Halo post-training framework claims 2.8x TRL throughput, accused of cherry-picking benchmarks
_ScottCondron · x · 2026-09-23
The new open-source post-training framework Halo claims up to 2.8x TRL throughput with lower peak memory while keeping models in native HuggingFace format. But Axolotl author Wing Lian points out that in single-GPU SFT of GLM Flash on B200s, Halo is actually 5% slower than Axolotl — hence why the team only benchmarked against TRL, which he calls cherry-picked data.
More from Infra
- Dev laments agents built around KV caches, wants inference-first chips — dbreunig · 2026-09-23
- GE Vernova seen hitting $200B backlog by early 2027 as turbine demand outruns guidance — BenBajarin · 2026-09-23
- New LLM papers: recursive language models generalize out of domain; XMerge depth compression — burny_tech · 2026-09-23
- DeepSeek report praised: absurdly tiny KV cache, dense infra design — stochasticchasm · 2026-09-23
- Running Qwen 27B locally on RTX 4090: beats pre-2025 coding models, RAM is the wall — julianharris · 2026-09-23
- FP8 Tuning Cuts 42.9ms Per Step: Custom SGLang Kernels Boost Inference 126% — HankYeomans · 2026-09-23