Skip to content

fix: make naive and snaive forecasts match their definitions - #1611

Open
Nicoló Angileri (nicoloangileri) wants to merge 3 commits into
microsoft:mainfrom
nicoloangileri:fix/naive-forecasters
Open

Nicoló Angileri (nicoloangileri) wants to merge 3 commits into
microsoft:mainfrom
nicoloangileri:fix/naive-forecasters

Conversation

@nicoloangileri

@nicoloangileri Nicoló Angileri (nicoloangileri) commented Sep 25, 2026 •

Copy link
Copy Markdown

Why are these changes needed?

Three of the simple time series forecasters do not forecast what their names say.

  • naive fits SimpleExpSmoothing with smoothing_level=0.0, so the level never moves from the estimated initial level, which is the training mean. It forecasts the same values as avg. Naive.predict(int) also returned params["initial_level"] as the "last observation".
  • snaive uses smoothing_level=1.0, which gives the plain naive forecast: the last observation for every step, whatever season is.
  • savg reads season from the fit kwargs, but AutoML passes the tuned value through the estimator params, so the searched season never reached ar_select_order and maxlag stayed at 1.

With the 30-point linear series from the #1587 regression test (10 to 20) and period=3, and with an 84-day series with a weekly pattern and period=7:

estimator forecast on main forecast with this PR
naive, linear 15, 15, 15 (mean) 20, 20, 20 (last value)
naive, weekly 105.15 (mean) 109.3 (last value)
snaive, weekly, season=7 109.3 for all 7 days 107.7, 112.8, 109.9, 105.0, 116.1, 102.2, 109.3 (last week)
savg, weekly, season=7 104.1 to 105.5, almost flat follows the weekly pattern

On the NYC energy data used in test_extra_models.py, main gives naive the same MAPE as avg (0.1377), and snaive the MAPE that naive gets with this PR (0.1152).

Changes:

  • Naive uses smoothing_level=1.0, so it forecasts the last observation. Its predict override is removed, and the StatsModelsEstimator.predict path handles both int and DataFrame inputs.
  • In-sample predictions keep only the requested timestamps. Statsmodels predicts every period between the first and last requested timestamp, so a sparse request used to return extra rows. This applies to every statsmodels-based estimator.
  • SeasonalNaive keeps the last training season and repeats it (SeasonalRandomWalk), so its forecasts repeat the last season exactly. Its in-sample predictions are the values one season earlier, and the first season uses the first observation. It requires a positive integer season, falls back to the naive model only for season == 1, and raises if the training data is shorter than one season.
  • StatsModelsEstimator.predict(int) builds the forecast frame from train_end_date, so an integer horizon starts right after the fitted training data. It used to start after end_date, which is the end of the held-out test rows when a dataset has them. This also applies to ARIMA, SARIMAX and Holt-Winters.
  • SeasonalAverage uses self.params["season"] when fit does not receive season, so the existing fit(dataset, season=3) call still works.

#1545 also touches SeasonalNaive.predict ([0] to .iloc[0]). This PR removes that method, so whichever lands second will need a small rebase.

Related issue number

None.

Checks

The three new tests in test/automl/test_forecast.py fail on main and pass here. test/automl/test_forecast.py: 19 passed, 2 skipped. I also ran naive, snaive, savg and avg through AutoML on the NYC energy data from test_extra_models.py, with exogenous regressors. All return 180 forecasts with no NaN, but I didn't run that module itself because it needs mlflow and Spark.

Naive fitted simple exponential smoothing with smoothing_level=0, so it
forecast the estimated initial level (the training mean) instead of the
last observation, the same as avg. SeasonalNaive used smoothing_level=1,
so it forecast the last observation and ignored the tuned season.

- Naive uses smoothing_level=1 and forecasts the last observation
- SeasonalNaive fits a seasonal random walk, SARIMA(0,0,0)(0,1,0)[season],
  whose forecasts repeat the last season
- SeasonalAverage reads the tuned season from its params when fit does
  not receive one
@nicoloangileri

Copy link
Copy Markdown
Author

@microsoft-github-policy-service agree

@thinkall Li Jiang (thinkall) left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall review of the complete PR: changes are required.

  1. flaml/automl/time_series/ts_model.py:748-755: invalid seasonal lengths fabricate forecasts. Require a positive integer, reserve naive fallback for season == 1, and reject folds with fewer than season training observations.
  2. flaml/automl/time_series/ts_model.py:336-342: predict(int) enriches the horizon before handling integers, so datasets with held-out test rows can shift the seasonal phase. Handle integer horizons first and forecast directly from the fitted training boundary.
  3. flaml/automl/time_series/ts_model.py:755-757: using SARIMAX for deterministic seasonal repetition is prohibitively expensive for long seasons and repeated CV fits. Cache the final seasonal cycle and implement tiling/slicing directly, including supported in-sample shifts.

Posted by thinkall-agent-auto-reviewer

- Require a positive integer season, keep the naive fallback for season
  1 only, and reject training data shorter than one season
- Start integer prediction horizons right after the fitted training
  data instead of after held-out test rows
- Repeat the last training season directly instead of fitting SARIMAX,
  which was slow for long seasons, and predict in-sample values one
  season back
@nicoloangileri

Copy link
Copy Markdown
Author

Thanks for the review. All three points are addressed in 327a744:

  1. SeasonalNaive now requires a positive integer season and raises ValueError for 0, negative or non-integer values. It falls back to the naive model only for season == 1, and raises when the training data has fewer than season observations.
  2. StatsModelsEstimator.predict(int) now builds the forecast frame from train_end_date before enriching, so an integer horizon starts right after the fitted training data. It used to start after end_date, which is the end of the held-out test rows when a dataset has them. This also changes predict(int) for ARIMA, SARIMAX and Holt-Winters, which had the same starting point.
  3. SeasonalNaive no longer fits SARIMAX. SeasonalRandomWalk keeps the last training season and tiles it for forecasts, and in-sample predictions are the values one season earlier (the first season uses the first observation). With 1,800 daily points and season=180, a fit now takes about 8 ms instead of 4.1 s.

New tests cover each point: invalid seasons, too few observations, predict(7) with 10 held-out rows, and the in-sample shift. They fail on the previous commit and pass now. test/automl/test_forecast.py: 25 passed, 2 skipped.

@thinkall Li Jiang (thinkall) left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall review of the complete current PR: one change is still required.

  • flaml/automl/time_series/ts_model.py:390-394,758-759: sparse in-sample requests are reduced to a start/end interval, and SeasonalRandomWalk.predict() returns every timestamp in that range. A request for only January 10 and January 12 therefore also returns January 11, breaking row-count and positional alignment. Select/reindex fitted values by the exact requested timestamps after validation, and add a sparse in-sample regression test.

The prior seasonal-length, integer-horizon phase, and SARIMAX performance blockers are resolved.

Posted by thinkall-agent-auto-reviewer

Statsmodels estimators predict every period between the first and last
requested in-sample timestamps, so a sparse request such as January 10
and 12 also returned January 11. Select the requested timestamps from
that range, validated the same way as future timestamps.
@nicoloangileri

Copy link
Copy Markdown
Author

Addressed in 7a97eec. StatsModelsEstimator.predict now keeps only the requested timestamps from the in-sample start/end range, with the same unique, increasing and aligned checks as the future path. The range comes from the shared statsmodels path, so this also fixes sparse in-sample requests for ARIMA, SARIMAX, Holt-Winters, avg, savg and naive, which also returned every period in the range.

The in-sample test now also requests three sparse training dates. It fails on the previous commit (11 rows instead of 3) and passes now. test/automl/test_forecast.py: 25 passed, 2 skipped.

@thinkall Li Jiang (thinkall) left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Overall review of the complete current PR: approved.

Sparse in-sample prediction now selects the exact requested timestamps and preserves cardinality, order, index, and values. The earlier seasonal-length validation, integer-horizon phase, and deterministic high-performance seasonal repetition fixes remain correct. Overall tests cover sparse/contiguous in-sample requests, delayed and partial future cycles, held-out test data, holdout/CV folds, timestamp validation, calendar indexes, exogenous model paths, shapes/dtypes, empty requests, and serialization.

No blocking issues found.

Posted by thinkall-agent-auto-reviewer

@nicoloangileri

Copy link
Copy Markdown
Author

Hi Li Jiang (@thinkall), thanks again for the review. CI is green and there are no conflicts, but the PR still shows as blocked on my side. Is anything else needed from me before it can be merged? Happy to rebase or adjust if needed.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants