Quick Answer: NanoForecast v0.5 (6.5M parameters) beats TimesFM (200M parameters) on 4 out of 6 standard time series benchmarks — using 31× fewer parameters. The improvement came entirely from fixing three training pipeline bugs, not from any architecture change.

The Three Bugs That Killed Performance

NanoForecast v0.3 had 3.030 MASE on our standard benchmark protocol. The architecture was sound — DeltaNet + LongConv with zero-shot probabilistic forecasting. The problem was in the pipeline:

  1. Loss scope: The loss was computed over the full sequence including context, not just the forecast horizon. The model was learning to reconstruct history instead of predict the future.
  2. Tensor shape misalignment: A transposition error in the data loader caused subtle shape bugs that only manifested at certain batch sizes. This introduced training instability that looked like model capacity issues.
  3. Augmentation coverage: Time-series augmentation (jitter, scaling, permutation) was applied inconsistently — some splits got no augmentation at all, causing domain gaps between train and eval.

Results After Fixing All Three

Same architecture. Same data. Same compute budget. MASE improved from 3.030 to 1.704 — a 43.8% reduction. On the ETT family:

  • ETTh1: NanoForecast 0.676 vs TimesFM 0.705 ✅
  • ETTh2: NanoForecast 1.110 vs TimesFM 1.360 ✅
  • ETTm1: NanoForecast 0.287 vs TimesFM 0.545 ✅
  • Exchange Rate: NanoForecast 4.317 vs TimesFM 4.383 ✅
  • Electricity: TimesFM leads (high cardinality)
  • Traffic: TimesFM leads (high cardinality)

Edge Deployment

CPU-only inference at 19.5ms per sample on Apple M4. Full ONNX export included. Streaming RNN mode available for IoT deployments where samples arrive one at a time.

Links

ArXiv: 2609.31669 · GitHub: eulogik/NanoForecast · HuggingFace

Want more technical insights?

Follow us on X for real-time updates on KARN and our other ventures.