How can I evaluate a time‑series forecasting model when I must train on the entire small dataset?

How can I evaluate a time‑series forecasting model when I must train on the entire small dataset?

Manage alerts

Loading saved threads...

CSe · External communityPost link
External question — Cross Validated Stack Exchange Author: CSe Original post: https://stats.stackexchange.com/questions/672454 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I’m building a Python forecasting pipeline that tries several models: Holt‑Winters (tuned with Optuna) ARIMA (via pmdarima.auto_arima ) XGBoost (tuned with Optuna) At the moment I split the data into an 80 % train set and a 20 % test set, using the test set only for evaluation. This works fine for large series, but when the series is very short (e.g., < 50 observations) every data point is valuable, so I would like to train on all available observations. My question is: How should I evaluate the model in this situation? Is it acceptable to use the same data for both training and testing (i.e., evaluate in‑sample )? If I resort to time‑series cross‑validation (e.g., 3 vlidation split), won’t the model be trained on fewer points than the “train‑on‑all” scenario, thus giving a bad estimate? Any recommendations for a reliable evaluation strategy (or a discussion of the trade‑offs) would be greatly appreciated.
Quote
Report
Stephan Kolassa · External communityPost link
External answer — Cross Validated Stack Exchange Author: Stephan Kolassa Original post: https://stats.stackexchange.com/a/672468 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I concur with whuber's comment : having little data does not nullify accepted best practices about time series analysis and forecasting. Simply go ahead and do your 80-20 split, using the first 40 observations for fitting and the last 10 for evaluation. If you want to do a three-way split (e.g., to decide between your three methods first, and evaluate the entire selection-then-forecasting pipeline in a second sample), then do a 30-10-10 split. Yes, that means that you will be less certain in your conclusions. But that is just a consequence of not having more data. That your pigs won't fly is simply a consequence of their not having wings. And it may well mean that the best method is a very simple one, like the historical mean: Best method for short time-series . Again, this is simply a consequence of not having a lot of data, so you can't confidently fit more complex methods. (As an aside, I find it funny that in that question from ten years ago, a "short" time series was one with 20 observations or fewer, but nowadays, even 50 observations seems to be considered "short"... Actually, 50 observations is not all that little, depending on your context. If these are monthly, then that is more than four years' worth of history.) Depending on your context, you might be able to leverage global models for your series if they are similar enough. (Transformers are pretty good even when they are pretrained on quite unrelated series.) Or you might be able to model seasonalities on aggregate levels, then push this down to the individual series. Or use Bayesian approaches. Whatever you do, no, don't go by in-sample fit. This only encourages overfitting. (You can compare different ARIMA models with the same order of integration [!] using the AIC, but that won't help you to compare them against Holt-Winters or XGBoost, where the AIC is not even defined or calculated in a non-comparable way.)
Quote
Report

Post Reply

Checking account access…