LoRA Training Test: Gradient Accumulation Is Not a Time-Zero Game

traceml-ai · reddit · 2026-08-19

Experiments with Qwen3-1.7B using TRL and LoRA reveal significant time differences for gradient accumulation strategies (1×4, 2×2, 4×1) despite identical effective batch sizes.

Results:

Analysis:

Conclusion: Treat effective batch (optimization behavior) and physical batch (memory/speed) as separate choices. Start with the largest physical batch that fits.

Related event: LoRA Benchmarks: Larger Batch Beats Gradient Accumulation by 17%(2 posts)→

Original post →

More from Infra

Infra channel →