How should correlated, regime-dependent return-path forecasts be combined into a trading consensus?
How should correlated, regime-dependent return-path forecasts be combined into a trading consensus?
Loading saved threads...
Russlan Ramdowar · External communityPost link
External question — Quantitative Finance Stack Exchange
Author: Russlan Ramdowar
Original post: https://quant.stackexchange.com/questions/85792
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
This question comes from implementing the consensus engine in iPulse AI, an Open Agentic Investment Research Platform that I'am developing with my team for 3 years now.
I mention the platform purpose only to explain the applied setting and context. By “open,” I mean that the research questions, methodology, evaluation procedures, historical forecasts, limitations, results and unsuccessful approaches are intended to be publicly inspectable and subject to external scrutiny.
Since launch in November 2025 (2 years after development), I have recorded each production forecast as it was originally issued, together with its timestamp, detailed AI agent configuration, specified Investment Evaluation Framework it's instructed to follow, contemporaneous market and fundamental inputs, and the subsequently realized return path.
Everything runs in batches of Deep Analysis, done every two weeks or so. Each agent is intended to produce a structured investment thesis rather than merely classify an asset as going up or down. The current implementation contains 12 distinct agent configurations evaluating 380 assets. One run therefore produces 12 separate theses per asset and 4,560 analyst–asset evaluations overall.
The evaluations are launched as a single time-aligned batch, using the same evaluation timestamp and a common set of inputs (Global events starting from model's knowledge cutoff date till date, and asset specific fundamentals and financials). This minimizes differences caused by news or market conditions changing during the run. However, the agents are not genuinely independent: they may share input data while differing in the model (ChatGPT vs Claude vs Gemini vs Grok), investment framework (Value Analysis Framework, Power Dynamics Framework, AI readiness and visionary disruptor framework etc) and mode (web search enabled or not).
This dependence is what makes the consensus problem difficult. My initial implementation treated consensus as a democratic vote—effectively a board of AI advisors or (AI Parliament) in which every member had equal weight. I am now looking for statistically grounded criticism and established, implementable methods for aggregating these different forecasts while accounting for shared information, correlated errors and changing performance.
Suppose the system contains
$m$
systematic research analysts forecasting the same asset over the same horizon.
An analyst is not just one model. Each analyst is a configuration consisting of:
an LLM or reasoning model;
an investment-analysis framework;
a retrieval mode, with or without external search;
a common set of market and fundamental input data.
The input dataset is largely shared, while the model, framework and retrieval configuration can differ. Consequently, the analysts are neither independent nor equally correlated. For example, two analysts using the same investment framework may make similar errors even if they use different underlying models (ChatGPT vs Gemini vs Claude).
At forecast origin
$t$
, analyst
$i$
produces a path of cumulative excess-return forecasts:
$$
f_{i,t}
=
\left(
f_{i,t}(1),\ldots,f_{i,t}(H)
\right)^\top .
$$
I want to construct a consensus path
$$
c_t(h)
=
\sum_{i=1}^{m} w_{i,t} f_{i,t}(h),
\qquad
w_{i,t}\geq 0,
\qquad
\sum_{i=1}^{m}w_{i,t}=1.
$$
I have considered four possible weighting mechanisms.
1. Equal weighting
$$
w_{i,t}=\frac{1}{m}.
$$
This is simple and difficult to overfit, but treats several highly similar analysts as independent confirmations.
2. Weighting by recent historical performance
For a specified path-level loss
$L_{i,s}$
, one possibility is an exponentially discounted rule such as
$$
w_{i,t}
\propto
\exp\left(
-\eta
\sum_{s<t}\rho^{\,t-s}L_{i,s}
\right).
$$
This adapts to changing performance, but may chase noise and eliminate temporarily weak analysts.
3. Covariance-aware weighting
Let
$e_{i,t}(h)$
denote analyst
$i$
's forecast error at horizon
$h$
. An integrated error-covariance matrix could be estimated as
$$
\Omega_{ij}
=
\sum_{h=1}^{H}
a_h\,
\operatorname{Cov}
\left(
e_{i,t}(h),e_{j,t}(h)
\right),
$$
followed by
$$
w_t
=
\arg\min_{w\in\Delta_m}
w^\top\Omega w,
$$
where
$\Delta_m$
is the probability simplex.
This accounts for average error dependence, but assumes that the estimated relationship remains relevant in the current macroeconomic environment.
4. Regime-dependent weighting
The weights could instead depend on observable or latent state variables
$z_t$
:
$$
w_t=g(z_t),
$$
where
$z_t$
might contain volatility, inflation, liquidity, growth or other macroeconomic-state indicators.
This could recognize that some analytical frameworks perform better in particular environments, but estimating many conditional weights from a limited history creates substantial overfitting risk.
A further complication is the potentially valuable contrarian analyst (The Michael Burry AI for example) . An analyst may have worse unconditional performance than the group but make different errors, or perform particularly well during the relatively rare periods in which the majority suffers a common-mode failure. Weighting analysts only by individual average loss could remove exactly the analyst that provides the most useful conditional diversification.
I have started storing every forecast at its original issue time, together with the analyst configuration, contemporaneous macro state and subsequently realized return path. The weights can therefore be evaluated through chronological walk-forward tests rather than fitted and tested on the same observations.
For this question, assume that an economically appropriate path-level loss has already been selected. Choosing that loss is a related but separate problem: pointwise MSE can, for example, score a forecast of a useful cyclical pattern poorly when its phase is slightly displaced.
Is there an established and practically implementable framework for combining:
unequal and partially shared information;
correlated forecast errors;
regime-dependent analyst skill; and
the conditional value of contrarian forecasts?
Would the statistically defensible approach be a hierarchical or factor model for the analysts' shared error components combined with dynamic model averaging, a contextual prediction-with-expert-advice algorithm, or a regularized rolling optimization of the combination weights?
In particular, how can one distinguish a genuinely diversifying contrarian analyst from a merely noisy one using only chronologically resolved forecasts, while limiting overfitting when the number of market regimes and analyst configurations is large relative to the available history?
Related discussions include
ensemble techniques for return forecasts
,
combining forecasts at different horizons
, and
combining alternative volatility estimates
, but they do not appear to address this combination of structured dependence, regime-dependent skill and path forecasts.
Quote
Report
dikovaxi · External communityPost link
External answer — Quantitative Finance Stack Exchange
Author: dikovaxi
Original post: https://quant.stackexchange.com/a/85820
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
There is an established framework, and it is smaller than the four options suggest — most of what you list collapses into one object, the error-covariance matrix, and one constraint, how little history you have. I'll take your four requirements in turn and then give the recipe I would actually run.
1. Shared information and correlated errors are the same problem, and it has a closed form.
For a fixed loss that is (locally) quadratic, the optimal combination of
$m$
forecasts is the Bates–Granger (1969) solution
$$w^{\ast}=\frac{\Omega^{-1}\mathbf 1}{\mathbf 1^\top\Omega^{-1}\mathbf 1},$$
i.e. the minimum-variance portfolio of the analysts' errors. Two consequences matter for your setting. First, an analyst with worse
individual
loss but errors that are weakly or negatively correlated with the group receives
positive
weight from this formula automatically — that is your contrarian, and you do not need a separate mechanism for him; you need
$\Omega$
to be estimated well enough to see him. Second, shared information caps what any combination can achieve: Clemen & Winkler (1985) show that
$m$
experts with pairwise error correlation
$\rho$
are worth
$m/(1+(m-1)\rho)$
independent ones. With twelve configurations that share inputs and largely share frameworks,
$\rho$
of 0.6–0.8 is realistic, which puts your "AI parliament" at roughly 1.5–2 independent opinions. The parliament is not twelve votes; it is two, repeated.
2. Why equal weights keep winning, and how to beat them without pretending.
The "forecast combination puzzle" (Stock & Watson 2004; Smith & Wallis 2009; Claeskens et al. 2016) is that estimated
$w^{\ast}$
usually loses to
$1/m$
out of sample, because the estimation error in
$\Omega$
costs more than the dependence structure gains. The known remedies are all forms of
shrinkage toward equal weights
(Diebold & Pauly 1990; Ledoit–Wolf on
$\Omega$
), trimmed means, and — the one that fits you best — a
structured
$\Omega$
. Your analysts are built from a small number of shared components (model family, framework, retrieval mode, common data), so a factor model for the errors,
$$e_{i,t}(h)=\lambda_i^\top f_t(h)+\varepsilon_{i,t}(h),\qquad \Omega=\Lambda\Phi\Lambda^\top+\Psi,$$
with factors indexed by framework, model family and retrieval, has a few dozen parameters instead of 78. That
is
the "hierarchical model for shared error components" you mention; it is the defensible version of option 3. The cheap approximation that captures most of it:
cluster, then average
— average forecasts within a framework (or within a model family), then across clusters with shrunk weights. Averaging within clusters removes the double counting; the small number of clusters keeps the weight estimation honest.
3. Regime dependence: count your history before you fit it.
Biweekly batches since November 2025 give ~20 forecast origins. The prediction-with-expert-advice results you cite (Cesa-Bianchi & Lugosi 2006) are the honest lens here: the exponentially weighted forecaster — your option 2 — has regret
$O(\sqrt{T\log m})$
, and fixed-share (Herbster & Warmuth 1998) or dynamic model averaging with a forgetting factor (Raftery, Kárný & Ettler 2010; Koop & Korobilis 2012) pays an additional
$\sqrt{T\,k\log m}$
for
$k$
regime switches. At
$T\approx20$
those bounds are of the same order as the loss differences you are trying to detect, so
any
state-dependent weighting
$g(z_t)$
with more than one or two state variables is fitting noise by construction. What the history does support: (a) one coarse state (a volatility regime, say) with fixed-share style adaptivity, and (b) the
cross-section
. Each origin has 380 assets; after removing the common market component (evaluate excess-over-cross-sectional-mean paths, otherwise all 380 errors share one factor and your effective sample is still ~20), you have real replication for estimating
unconditional
$\Omega$
and for testing skill — not for regime-conditional weights. Giacomini & White (2006) "Tests of conditional predictive ability" is the correct test for "framework A beats framework B
when
$z_t$
is high": it is built for exactly this question and respects small samples.
4. Distinguishing a diversifying contrarian from a noisy one.
Measured, not argued: compute the walk-forward loss of the combination with and without the analyst, under shrunk weights, and look at the loss
conditional on group failure episodes
(the top-decile group-error origins). With 20 origins you have two or three such episodes — enough to see a sign, not to estimate a weight. Treat him the way the factor model does: his loading on the shared factor is negative, his idiosyncratic variance is what makes him "noisy"; shrink both toward zero and let the posterior decide his weight. A contrarian whose value only appears with an unshrunk
$\Omega$
is noise.
One thing your setup allows that classical panels don't.
Shared information can be
corrected
rather than merely modelled if each analyst also reports what it expects the
others
to say. Palley & Soll (2019,
Management Science
, "Extracting the wisdom of crowds when information is shared") show that the "pivot" — moving the mean forecast away from the mean
meta
-forecast — recovers the shared/private information split without estimating any covariance; Prelec, Seung & McCoy (2017,
Nature
) is the discrete version. For LLM agents a meta-prediction costs one extra field in the output schema. I would add it before I added a single regime variable.
What I would run, in order
Loss on cross-sectionally demeaned paths; equal weights as the benchmark that everything must beat out of sample.
Cluster-then-average by framework and model family; report the effective number of independent analysts from
$\rho$
.
Factor model for errors on the walk-forward history; shrunk inverse-covariance weights with the shrinkage intensity chosen by walk-forward, not in-sample.
Meta-predictions and the pivot correction.
At most one regime variable, fixed-share adaptivity, and a Giacomini–White test before you believe any conditional skill.
Diebold–Mariano against the equal-weight benchmark at every step; stop adding structure at the first step that does not win.
From experience combining a dozen specialised models on a different kind of forecast: the largest gains came not from weights at all but from letting members
abstain
— combining by agreement and silence rather than by averaging — which sidesteps most of the covariance estimation. It is worth checking whether a "no forecast" output is allowed in your evaluation framework; if it is, the consensus problem changes shape.
References: Bates & Granger (1969)
Operational Research Quarterly
; Clemen & Winkler (1985)
Management Science
; Diebold & Pauly (1990)
J. Forecasting
; Stock & Watson (2004)
J. Forecasting
; Smith & Wallis (2009)
Oxford Bull. Econ. Stat.
; Claeskens, Magnus, Vasnev & Wang (2016)
Int. J. Forecasting
; Herbster & Warmuth (1998)
Machine Learning
; Cesa-Bianchi & Lugosi (2006)
Prediction, Learning, and Games
; Raftery, Kárný & Ettler (2010)
Technometrics
; Koop & Korobilis (2012)
Int. Econ. Review
; Giacomini & White (2006)
Econometrica
; Diebold & Mariano (1995)
JBES
; Palley & Soll (2019)
Management Science
; Prelec, Seung & McCoy (2017)
Nature
.
Quote
Report
Post Reply
Checking account access…