Pipeline Fixes Alone Cut 6.5M NanoForecast MASE 43.8%, Beating TimesFM on 4 of 6

NanoForecast v0.5: Competitive Time Series Forecasting Through Training Pipeline Optimization

Gautam Kishore

cs.LG

2026-09-16

Retraining the same 6.5M NanoForecast with three pipeline fixes cuts overall MASE 43.8% (3.030 to 1.704) and beats 200M TimesFM on four of six benchmarks.

What problem this solves

TimesFM trains a 200M decoder-only transformer on about 100B time points. Chronos adapts T5 to quantized series at up to 710M parameters. PatchTST in the official setup used here sits at 15M+. Accuracy went up. So did the GPU bill. Teams without that stack cannot ship those models.

NanoForecast aims narrower: training that fits on consumer hardware, inference on CPU, scores that still hold on public benchmarks. The released v0.3 checkpoint already is this 6.5M architecture. Under the paper's standard protocol its overall MASE was 3.030, too weak to use. The architecture does not change. Three silent pipeline bugs get fixed (loss scope, tensor shape alignment, augmentation coverage), the same corpus and compute budget are reused, and the new checkpoint is v0.5. The contribution is empirical: get the gradients and the data diet right.

Method

The backbone is from the Reverso family: a depthwise long convolution over the full token sequence, a DeltaNet linear RNN that updates a matrix state with the delta rule, and a gated MLP. A learned router emits one weight triple per window and mixes the three. Eight layers, width 96, patch size 8, context 512, horizon 48, 6.5M parameters. Each window is robust-scaled with median and IQR, cut into non-overlapping patches, and prefixed with a learned frequency embedding.

DeltaNet keeps the matrix state across calls, so a new observation costs one forward pass and history is not re-fed. That is the streaming path. Heads emit a point forecast and monotonic quantiles built from a median plus non-negative softplus offsets. Reported point numbers use the pinball-trained median p50, which matches MAE-based MASE.

Three pipeline bugs.

Loss scope. The dataloader always inserted a horizon key. With multihorizon off, the point loss still backpropagated through the full context-length output. Gradients went into reconstructing history the model had already seen. Curves still converged. Forecast accuracy did not. Truncate predictions and targets to H before the loss.

Tensor shape alignment. The quantile term ran before truncation, scoring context-length activations against (B, H) targets. Gradients in that branch went through the wrong axes. Uncertainty and point accuracy both dropped. Truncate both tensors first, then compute every loss term.

Augmentation coverage. v0.3 only scaled, shifted, and jittered, a thin diet on a fixed corpus. v0.5 samples jitter, scale, shift, mask, and time reversal in the training loop.

The three fixes are scored together. There is no per-bug ablation. The released v0.3 and v0.5 checkpoints bound the joint effect.

Results

Every number, baselines included, uses one protocol: context 512, horizon 48, non-overlapping test windows, MASE scaled by in-sample seasonal-naive MAE (lag 24 hourly, 96 for 15-minute, 7 daily). ETT splits are 70/20/10 in time; the others are 70/10/20. TimesFM is the public 200M checkpoint. PatchTST is retrained per dataset with official hyperparameters for 40 epochs. Chronos-T5-large is omitted because inference did not fit on the large sets.

Datasetv0.5 (6.5M)TimesFM (200M)PatchTST (15M+)
ETTh10.6760.7050.781
ETTh21.1101.3601.467
ETTm10.2870.5450.488
Exchange4.3174.3833.861
Electricity2.0290.9231.347
Traffic1.8050.7651.379
Overall1.7041.4471.554

Against v0.3, overall MASE falls from 3.030 to 1.704, a 43.8% cut. Almost all of that cut sits on Exchange: 11.758 to 4.317 (-63.3%). ETTh2 drops 16.4%, Electricity 8.3%, Traffic 5.7%. ETTh1 and ETTm1 move 0.7% and 0.2%, a wash in practice.

V0.5 has lower MASE than TimesFM on the three ETT sets plus Exchange, at 31x fewer parameters. The gaps that open are ETTm1 (0.287 vs 0.545) and ETTh2 (1.110 vs 1.360). ETTh1 0.676 vs 0.705 and Exchange 4.317 vs 4.383 are thin, closer to a tie. Electricity 2.029 vs 0.923 and Traffic 1.805 vs 0.765 still belong to TimesFM. Against PatchTST, v0.5 wins all three ETT sets and loses Exchange, Electricity, and Traffic. Overall 1.704 still trails TimesFM at 1.447 and PatchTST at 1.554.

Efficiency ratio is MASE inverse over millions of parameters: 0.090 for v0.5, 0.043 for PatchTST, 0.0035 for TimesFM. Quantile calibration moved the wrong way. The nominal 80% band covered 90.3% of held-out points for v0.3 and 51.3% for v0.5. Point forecasts are unaffected. Treat the bands as relative uncertainty.

On an Apple M4 at batch 1: PyTorch FP32 19.5 ms, ONNX FP32 10.7 ms, a streaming step 19.1 ms. INT8 dynamic quantization slows to 33.3 ms. One NVIDIA T4 takes about 12 hours. The validation-best snapshot moves from epoch 147 (v0.3) to epoch 51 (v0.5). ONNX export is 27.9 MB FP32 and 9.2 MB INT8.

Why it matters

Whether the loss is truncated to the horizon, whether the quantile branch sees matching shapes, whether augmentation is only three cheap transforms: those details can look fine on a training curve and still cost more than 40% MASE. Check them before blaming the architecture.

6.5M parameters, 10 to 20 ms on CPU, an ONNX file under 30 MB: that is in range for edge boxes and live analytics. On the three ETT sets a correctly trained small model can stand in for 200M TimesFM. On 321-client electricity and 862-sensor traffic it cannot. Pretraining breadth still pays there. The Exchange edge over TimesFM is 0.066 MASE. Do not sell it as a gap.

This is a pipeline-hygiene result, not an architecture result. Overall accuracy still lags the two larger baselines. The wins cluster on smaller, univariate, ETT-like series.

Limitations

The paper lists the main ones: overall MASE still trails; context is locked at 512; channels are independent; quantile intervals run narrow; only six public datasets; training still wants a GPU for 12 hours; one seed per configuration. Released checkpoints are trained for H=48 only.

There is no per-bug ablation, so 43.8% is a bundle, and Exchange accounts for most of it. TimesFM is a public checkpoint pretrained on about 100B time points; NanoForecast is a 12-hour retrain on a small corpus. Evaluation windows match. Training regimes do not. INT8 is slower on this M4, and Raspberry Pi-class numbers are future work, so the edge story currently stops at a laptop CPU. Chronos, Moirai, and Lag-Llama are absent. The small-versus-foundation framing is really versus TimesFM and PatchTST.

Terms

Source

Related papers

All paper explainers