Is using contemporaneous components to forecast an aggregate a valid method or a form of data leakage?
Is using contemporaneous components to forecast an aggregate a valid method or a form of data leakage?
Loading saved threads...
PSE · External communityPost link
External question — Cross Validated Stack Exchange
Author: PSE
Original post: https://stats.stackexchange.com/questions/670272
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I am in the middle of a deep methodological debate regarding a time series forecasting problem and would appreciate the community's expert opinion.
The Context
I am trying to forecast an aggregate time series, which we can simplify as the
US CPI
(
Aggregate_t
). My dataset is monthly, from Jan 1999 to Aug 2025.
I have three types of variables:
The target series itself:
Aggregate_t
.
Its two main components (for simplicity):
Component_A_t
(e.g., Goods) and
Component_B_t
(e.g., Services).
A set of external, exogenous variables, which we can call
Exogenous_t
.
The crucial point is that the aggregate is a deterministic function of its components:
Aggregate_t
is, by definition, a weighted average of
Component_A_t
and
Component_B_t
.
The Backtesting Framework
My goal is to backtest a model (like LASSO or XGBoost) using an expanding window approach.
My full dataset runs until
August 2025
.
My backtest loop uses expanding windows, with the
data_base
(the last point of information for a given forecast origin) running from, for example, Jan 2017 up to
July 2025
.
I stop the
data_base
loop in July 2025 so that for this final point, the true value for
August 2025
is available in my dataset to evaluate the
h=1
forecast.
The Methodological Debate
The core of my debate is how to correctly build the feature set
s_t
for a given
data_base
t
(e.g.,
July 2025
) to predict the target
t+1
(e.g.,
August 2025
).
My Position:
My argument is that I should use all information available at time
t
. Since the values for
Component_A_July
,
Component_B_July
, and
Exogenous_July
are all known, they are all valid features. My proposed feature set
s_t
would therefore include:
Contemporaneous values of the components (
Component_A_t
,
Component_B_t
).
Contemporaneous values of the exogenous variables (
Exogenous_t
).
Lags of all variables (
Aggregate_{t-1}
,
Component_A_{t-1}
,
Exogenous_{t-1}
, etc.).
The AI's Position:
I've had a long and detailed debate with an AI assistant (Gemini) who strongly argues against this. The AI's position is that including the contemporaneous components (
Component_A_t
,
Component_B_t
) constitutes a form of
identity or proxy leakage
.
The AI's reasoning is that any flexible model will first learn the near-perfect deterministic relationship
Aggregate_t ≈ w1*Component_A_t + w2*Component_B_t
. The model then uses this reconstructed
Aggregate_t
to trivially predict
Aggregate_{t+1}
due to the series' high autocorrelation. The AI claims this contaminates the experiment, as the model ignores more subtle signals from the lags and exogenous variables, leading to inflated backtest metrics and a non-robust model that has not learned any true economic dynamics.
I find it very difficult to accept the AI's position, as I am not leaking any
future
information. I am simply using all available data from the present (
t
) to predict the future (
t+1
).
The Question
For this specific backtesting setup, where the target
y_{t+h}
is known within the historical dataset for every step of the loop:
Is it a methodologically sound practice to include contemporaneous components (Component_A_t, Component_B_t) in the feature set s_t to predict Aggregate_{t+h}? Or does this indeed represent a form of proxy leakage that invalidates the evaluation of the other predictors and leads to overly optimistic results, as the AI suggests?
P.S. Empirical Results from an A/B Test
To add empirical context, I have run the backtest for both scenarios using real data.
Method A (My hypothesis):
Including the contemporaneous components.
Method B (The AI's "clean" recommendation):
Excluding the contemporaneous components.
Here are the results for the
h=1
forecast RMSE:
For LASSO:
The results were surprisingly almost identical. Method A (with the shortcut) had an RMSE of
~0.396
, while Method B (clean) had an RMSE of
~0.395
. The interpretation seems to be that LASSO's L1 regularization effectively handled the multicollinearity and "defended" itself from the shortcut by zeroing out redundant features.
For XGBoost (untuned):
The results were different. Method B (clean) had a relatively high RMSE of
~0.432
. Method A (with the shortcut) improved upon this, with an RMSE of
~0.407
.
This empirical result adds to my confusion. The shortcut seems to exist (it helps the XGBoost model), but it's not as catastrophic as the AI initially predicted, and the simpler LASSO model seems robust to it and performs better overall. This reinforces my question about whether including these features is a truly invalid practice or just a risky one with specific trade-offs depending on the model used.
Quote
Report
Stephan Kolassa · External communityPost link
External answer — Cross Validated Stack Exchange
Author: Stephan Kolassa
Original post: https://stats.stackexchange.com/a/670273
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
You are completely correct, and Gemini is wrong.
The model then uses this reconstructed
Aggregate_t
to trivially predict
Aggregate_{t+1}
due to the series' high autocorrelation.
Well, but that is wonderful! If your goal is to predict
Aggregate
, based on your two components, and the total has such a high autocorrelation that it is easy to forecast
without using the separate components
, then why in the world would you
not
do that?
Disregard Gemini and move on.
(Alternatively, use a hierarchical reconciliation step: forecast the Aggregate by itself, and also calculate the weighted average between the forecasted components, which will give you a
different
forecast of the Aggregate. I strongly suspect that the average of these two forecasts will be better than either one.)
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External answer — Cross Validated Stack Exchange Author: Stephan Kolassa Source score (net votes, not local likes): 3 Original post: https://stats.stackexchange.com/a/670273 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. You are completely correct, and Gemini is wrong. The model then uses this reconstructed Aggregate_t to trivially predict Aggregate_{t+1} due to the series' high autocorrelation. Well, but that is wonderful! If your goal is to predict Aggregate , based on your two components, and the total has such a high autocorrelation that it is easy to forecast without using the separate components , then why in the world would you not do that? Disregard Gemini and move on. (Alternatively, use a hierarchical reconciliation step: forecast the Aggregate by itself, and also calculate the weighted average between the forecasted components, which will give you a different forecast of the Aggregate. I strongly suspect that the average of these two forecasts will be better than either one.)
Checking account access…