Time scaling of AR(1) process for modelling financial returns

Time scaling of AR(1) process for modelling financial returns

Manage alerts

Loading saved threads...

skoestlmeier · External communityPost link
External question — Cross Validated Stack Exchange Author: skoestlmeier Original post: https://stats.stackexchange.com/questions/652799 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. Process: Consider an AR(1) process with zero mean, * $\lambda_t = \kappa \cdot \lambda_{t-1} + \omega_t$ , with $\kappa = 0.9$ , $\omega \sim N(0, \sigma_{\omega}^2)$ , and $\sigma_{\omega}^2 = 0.00027$ . The initial value of $\lambda_0$ is drawn from the stationary distribution $N \left(0, \frac{\sigma_{\omega}^2}{(1-\kappa^2)} \right)$ I use this process to generate a sample of length $T=672$ . * I know that a mean of zero is not realistic for modelling stock returns, but my question does not depend on this choice. The above time-series should be interpreted as being monthly returns (stated in percentage points), i.e., 672 months of observations. Question: What are appropriate parameter values for generating a total of 14,112 daily observations (i.e., $672 \times 21$ , if we assume 21 trading days within a month) to be compliant with the above (monthly) process? A monthly return is the cumulative product of all 21 daily returns within a given month. Attempt (in R): DAYS <- 21 T <- 672 kappa <- 0.9 variance_omega <- 0.00027 get_init_lambda <- function(variance) return(rnorm(n = 1, mean = 0, sd = sqrt(variance / (1-(kappa)^2)))) #### monthly analysis set.seed(1234) lambda_T <- vector(mode = "numeric", length = T) lambda_shock <- rnorm(T, mean = 0, sd = sqrt(variance_omega)) lambda_T[1] <- kappa * get_init_lambda(variance_omega) + lambda_shock[1] for(i in 2:T) lambda_T[i] <- kappa * lambda_T[i-1] + lambda_shock[i] acf(lambda_T)$acf[2] # [1] 0.9064949 # as expected var(lambda_T) # [1] 0.001547795 #### daily analysis set.seed(1234) kappa_daily <- 0.90 # ??? how to set kappa_daily, so monthly returns have AC = 0.9? lambda_T_daily <- vector(mode = "numeric", length = T * DAYS) lambda_shock_daily <- rnorm(T * DAYS, mean = 0, sd = sqrt(variance_omega / DAYS)) lambda_T_daily[1] <- kappa_daily * get_init_lambda(variance_omega / DAYS) + lambda_T_daily[1] for(i in 2:(T * DAYS)) lambda_T_daily[i] <- kappa_daily * lambda_T_daily[i-1] + lambda_shock_daily[i] # retrieve begin and end of month index; assume 21 trading days each month begin_month <- seq(1, T * DAYS, 21) end_month <- c(tail(begin_month, -1) - 1, T * DAYS) lambda_monthly_aggregate <- vector(mode = "numeric", length = T) for(i in 1:T){ month_ind <- (begin_month[i]):(end_month[i]) daily_ret_within_month <- 0.01*lambda_T_daily[month_ind] # returns are in % monthly_return <- 100*(cumprod(1+daily_ret_within_month) - 1) # cumulate daily returns lambda_monthly_aggregate[i] <- monthly_return[DAYS] # retrieve cum. ret. after 21 days } # not identical with above values: autocorrelation is too low and variance of mth. returns too high! > acf(lambda_monthly_aggregate)$acf[2] [1] 0.2988249 var(lambda_monthly_aggregate) [1] 0.01494742
Quote
Report
Adam Check · External communityPost link
External answer — Cross Validated Stack Exchange Author: Adam Check Original post: https://stats.stackexchange.com/a/652913 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. For the AR(1) case, there is a clear mapping to the new value of kappa. Under the monthly data generating process (DGP), any deviation from the mean persists at a rate of $\kappa=0.90$ (i.e., it decays by 10% per month). Since there are 21 business days in your month, the rate of daily persistence is $\kappa_{daily} = 0.9^{\left (\frac{1}{21} \right )} = 0.9949954$ . To check this gives the same rate of monthly decay, after one month (21 days) we have: $$E(y_{T+21})= 0.9949954^{21} = 0.90,$$ after two months (42 days) we have: $$E(y_{T+21})= 0.9949954^{42} = 0.81 $$ etc. For the new value of the variance, if you want the unconditional variance of the monthly series to be the same under either the daily or monthly DGP, then rather than setting $\sigma_{daily}^2 \approx \frac{\sigma^2}{21}$ , you need to do the explicit calculation. The monthly data generating process implies an unconditional variance of: $$\text{Var}(Y) = \frac{0.00027}{1-0.9^2}.$$ If you want the daily DGP to provide the same unconditional monthly variance as generated by the monthly DGP, then you should therefore set: $$\frac{\sigma_{daily}^2}{1- \left ( 0.9^{\frac{1}{21}} \right )^2} = \frac{0.00027}{1-0.9^2}, $$ $$\sigma_{daily}^2 = \left (1 - 0.9^{\frac{2}{21}} \right ) \frac{0.00027}{1-0.9^2}. $$ Edit (8/31/24) I believe @mlofton makes an important point. The accepted answer and my answer are both correct for an AR process sampled every $k$ periods. So they are correct if the two series are: (1) daily returns and (2) the daily return sampled every 21 days (e.g. the daily return on the last day of the month). However, if the monthly series is the sum of the daily AR(1) process, they dynamics become more complicated. The simulation code below shows this. It computes a simulated path of daily returns, then does two things. First, it samples that path every 21 days (i.e., gets the daily return on the last day of each month). Second, it sums the daily returns in blocks of size 21 (so each element is the sum of the daily returns in that month). This latter series is typically how we would think of a monthly return. We get the expected results when the data is sliced - in the OP's example, the AR(1) coefficient is roughly 0.90 and the conditional variance is roughly 0.00027. The ACF of the residuals show nearly no autocorrelation in the residuals - it is all captured by the AR(1) model. On the other hand, then the monthly return is computed as the sum of the daily returns in that month, the AR(1) coefficient is roughly 0.93, and the variance is around 0.079. The ACF of the residuals shows that there is still substantial uncaptured autocorrelation through roughly lag 20. Series: Y_summed ARIMA(1,0,0) with zero mean Coefficients: ar1 0.9317 s.e. 0.0011 sigma^2 = 0.07877: log likelihood = -14833.6 AIC=29671.2 AICc=29671.2 BIC=29690.22 Full code here: library(fpp2) set.seed(2024) # Parameters at daily frequency phi <- 0.9949954 sigma <- sqrt(1.418802e-5) sigma_unc <- sqrt( sigma^2 / (1-phi^2)) # Initialize sample N <- 21*100000 xt1 <- rnorm(n=1,mean=0,sd=sigma_unc) X <- rep(NA,N) # Loop to generate sample for (i in 1:N){ xt <- phi*xt1 + rnorm(n=1,mean=0,sd=sigma) X[i] <- xt xt1 <- xt } # Take every 21st value (daily return on last day in each month) Y_sliced <- X[seq(21,N,21)] # Sum in blocks of size 21 (total monthly return) Y_summed <- rep(NA,length(Y_sliced)) for (i in 1:length(Y_sliced)){ start_place <- 21*(i-1)+1 end_place <- i*21 Y_summed[i] <- sum(X[start_place:end_place]) } # Fit ARIMA to daily data fit_X <- Arima(X,order=c(1,0,0),include.constant=FALSE) # Fit ARIMA to daily return on last day in each month fit_Y_sliced <- Arima(Y_sliced,order=c(1,0,0),include.constant=FALSE) # Fit ARIMA to total monthly return fit_Y_summed <- Arima(Y_summed,order=c(1,0,0),include.constant=FALSE) # Print estimation results print(summary(fit_Y_sliced)) print(summary(fit_Y_summed)) # Print ACFs print(ggAcf(residuals(fit_Y_sliced))) print(ggAcf(residuals(fit_Y_summed)) + ggtitle('ACF of residuals from AR(1) model fit to monthly returns')) ggsave('ACF_monthly_returns.png',device = 'png',width = 2000, height=1000,units = 'px')
Quote
Report
Ben · External communityPost link
External answer — Cross Validated Stack Exchange Author: Ben Original post: https://stats.stackexchange.com/a/652992 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. Consider two overlapping AR(1) models for the same stationary process, with the latter expressing the process in terms of increments of $k$ time units: $$\begin{align} X_{t+1} &= \phi X_t + \varepsilon_t, \\[6pt] X_{t+k} &= \kappa X_t + \omega_t. \\[6pt] \end{align}$$ Using the auto-covariance function for a stationary AR(1) process, we can then write the auto-covariance and the variance in terms of either parameterisation as: $$\begin{align} \mathbb{Cov}(X_t,X_{t+k}) &= \frac{\phi^k}{1-\phi^2} \cdot \sigma_\varepsilon^2 \quad \quad \quad \quad \quad \mathbb{V}(X_t) = \frac{1}{1-\phi^2} \cdot \sigma_\varepsilon^2, \\[6pt] \mathbb{Cov}(X_t,X_{t+k}) &= \frac{\kappa}{1-\kappa^2} \cdot \sigma_\omega^2 \quad \quad \quad \quad \quad \mathbb{V}(X_t) = \frac{1}{1-\kappa^2} \cdot \sigma_\omega^2. \\[6pt] \end{align}$$ Equating the variance and auto-covariance in these two expressions of the process gives the transforming equations: $$\begin{align} \kappa &= \phi^k \quad \quad \quad \quad \quad \sigma_\omega^2 = \frac{1-\phi^{2k}}{1-\phi^2} \cdot \sigma_\varepsilon^2, \\[6pt] \phi &= \kappa^{1/k} \quad \quad \quad \quad \ \sigma_\varepsilon^2 = \frac{1-\kappa^{2/k}}{1-\kappa^{2}} \cdot \sigma_\omega^2. \\[6pt] \end{align}$$ This gives you the general parameter correspondence between a stationary AR(1) model measured in single time-units and its corresponding model form when we deal with increments of $k$ time units. In your case you have $k=21$ , $\kappa=0.9$ and $\sigma_\omega^2=0.00027$ which gives the corresponding parameters (for the daily process): $$\phi = 0.9949954 \quad \quad \quad \quad \ \sigma_\varepsilon^2 = 1.418802 \times 10^{-5}.$$
Quote
Report
mlofton · External communityPost link
External answer — Cross Validated Stack Exchange Author: mlofton Original post: https://stats.stackexchange.com/a/653164 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. This is not an answer but I figured out what my problem is-was regarding both Ben's and Adam's answer. They both had great answers but my thick brain kept being bothered by the unconditional variance because I felt like they shouldn't be set equal to each other but I didn't know why. This amorphous bother eventually led to me realizing that when the OP generates monthly returns from the k = 21 daily returns, the returns are essentially being summed. Now, CLEARLY the sum() function is not used in the R code for calculating the monthly from the daily but it's disguised in the form of a cumprod on the arithmetic returns. ( OP could have summed the log differences and the log differences are an approximation to the arithmetic ). So, the transformed series at the aggregated level ( monthly ) , although a monthly process, is really a process representing the sum of the k = 21 daily observations. The daily observations represent the unaggregated process. So, if $x_t$ is the process for the daily, then the first monthly observation: $Y_{(t+k)^{*}} = \sum_{i=1}^{k} x_{(t+i)}$ is what they call the first observation of the "non-overlapping temporal aggregate". Although not expressed in this manner explicitly, the problem really becomes transforming either: $$Y_{(t+k)^{*}} \longrightarrow x_t $$ or $$ x_t \longrightarrow Y_{(t+k)^{*}} $$ depending on what is given. I haven't obtained the paper yet but Amemiya and Wu (1972) is the seminal one. They supposedly show that, if one has a AR(1) process at the unaggregated level (daily), then, the "non-overlapping temporal aggregate" ( which is exactly what we have because the monthly is the sum of the daily and the next monthly has totally new observations ), is an ARMA process with some order for p and q. When I get the paper, I'll see if it says anything regarding the AR(1) when it's "un-aggregated" so $$Y_{(t+k)^{*}} \longrightarrow x_t $$ . #================================================================= 08-26-2024: ADDENDUM TO MY ORIGINAL ANSWER SINCE I FINALLY OBTAINED THE PAPER BY AMEMIYA AND WU. The paper is titled "The Effect Of Aggregation on Prediction in the Autoregressive Model", Amemiya and Wu, Journal of The American Statistical Association", September, 1972, Volume 67, Number 339, Theory and Methods Section. I write the abstract here: This article shows that, if the original variable follows a pth order autoregressive system, then the non-overlapping sum follows a pth order autoregression with at most a pth order moving-average of an independent sequence regardless of the length of the summation. From such an aggregate model, we derive the optimal predictor of the aggregate variable and show that it performs remarkably well compare to the optimal disaggregate predictor. The article contains both theoretical and numerical analysis. In my opinion, the cons to the paper are: I read it very loosely but I don't see a clear explanation on the $p=1$ case which would have been nice. The pth order case is detailed but it's obviously more complicated. It sounded from your question that you wanted to go in the opposite direction from what the paper considers. The paper starts from the consideration of the autoregressive model for the dis-aggregated data ( the daily returns) and then derives the resulting model for the aggregated data (the monthly returns). You wanted to go in from the monthly to the daily, but, in a way, the paper's approach of going from the higher frequency level ( dis-aggregated ) to the lower frequency level seems more natural. So, maybe you can change your problem and model it from that perspective ? There is also a text by B.D Anderson called "Time Series Analysis and Forecasting" that covers some aggregation material on the $p=1$ case. Anderson derives a $p=1$ relation in that book (the last chapter) and then references his masters thesis for details. I tried to get his masters thesis and failed. Anderson is not so well known ( like say T.W. ) but he did a lot in time-series. He was never in a niche academic position so a lot of his material wasn't published or it's not easily available. Anyway, reading the paper should give you a better idea of what you're up against. It's not an easy topic. I'm not familiar with Wu ( I confused Wei and Wu initially ) but Amemiya is viewed as one of the top econometricians ever, atleast in the 1900's. If you feel that the paper leaves you with no increased understanding, then try to get your hands on O.D Anderson's text. He writes more clearly than the authors in the paper but he still skips steps. That's why I looked for his thesis. I wish you luck and, if I had more time, it would be interesting to try to figure out just the $p=1$ case. But I know that, for me, the time it would take and the chances of me figuring out anything useful don't justify the effort right now. I have too many other things going on not even related to stats-time series. Maybe Adam or Ben want to take a stab at it since they did a nice job originally. All the best. Mark P.S: Oh, If you want the paper, it's most likely available on JSTOR but it will cost something depending on your subscription. I recommend JPASS ( subscription for downloading on JSTOR ) for papers even though I didn't get this one that way. Also,check out the Tesler reference at the end of the Amemiya and Wu paper. I have it somewhere and it might be easier to follow than the paper. The reference is: Tesler, Lester G.," Discrete Samples and Moving Sums in Stationary Stochastic Processes, Journal of The American Statistical Association, 62 (June 1967), 484-499. #================================================================ 08-30-2024: ANOTHER ADDENDUM WHICH PROVIDES AN ATTEMPT AT DERIVING THE MODEL FOR THE NON-OVERLAPPING SUM OF AN AR(1). Suppose we have the following AR(1): $x_t = \phi x_{t-1} + \epsilon_t$ and we want to figure out what the resulting model is for $ y_{t+m-1} = \sum_{i=0}^{m-1} x_{t+i}$ . So, we have: $$ x_t$$ $$x_{t+1} = \phi x_{t} + \epsilon_{t+1}$$ $$x_{t+2} = \phi x_{t+1} + \epsilon_{t+2}$$ $$\cdots$$ $$x_{t+m-1} = \phi x_{t+m-2} + \epsilon_{t+m-1}$$ Now, the trick is to write each successive $x_{t+i}$ as a function of the first observation $x_t$ . This results in $$ x_t$$ $$x_{t+1} = \phi x_{t} + \epsilon_{t+1}$$ $$x_{t+2} = \phi^2 x_{t} + \phi \epsilon_{t+1} + \epsilon_{t+2}$$ $$x_{t+3} = \phi^3 x_{t} + \phi^2 \epsilon_{t+1} + \phi \epsilon_{t+2} + \epsilon_{t+3}$$ $$\cdots$$ $$x_{t+m-1} = \phi^{m-1} x_{t} + \phi^{m-2} \epsilon_{t+1} + \ldots + \phi \epsilon_{t+m-2} + \epsilon_{t+m-1}$$ Since we want the resulting model is for $\sum_{i=0}^{m-1} x_{t+i}$ , we need to sum all of the $x_{t+i}$ above. First we will sum all the terms involving $x_t$ . This results in equation (1). (1) $$\sum_{i=0}^{m-1} \rho^{i} x_{t} = \frac{(1-\phi^{m})}{(1-\phi)} x_{t}$$ Next, we need to sum all of the terms involving $\epsilon_t$ and all of the terms involving $\epsilon_{t+1}$ and $\ldots$ all of the terms involving $\epsilon_{m-2}$ and finally all of the terms involving $\epsilon_{t+m-1}$ . Notice that there are $(m-2)$ terms involving $\epsilon_{t+1}$ and $(m-3)$ terms involving $\epsilon_{t+2}$ and $\ldots (m-(m-2))$ terms involving $\epsilon_{t+m-2}$ and finally $\ldots (m-(m-1))$ terms involving $\epsilon_{t+m-1}$ . Next, we create the finite sums for each of terms below: $$ \sum_{i=0}^{m-2} \phi^{i} \epsilon_{t+1}$$ $$ \sum_{i=0}^{m-3} \phi^{i} \epsilon_{t+2}$$ $$\cdots$$ $$ \sum_{i=0}^{(m-(m-1))} \phi^{i} \epsilon_{t+m-2}$$ $$ \sum_{i=0}^{(m-m)} \phi^{i} \epsilon_{t+m-1}$$ Fortunately, each of these sums is a finite geometric series so they can be simplified: $$ \sum_{i=0}^{m-2} \phi^{i} \epsilon_{t+1} = \frac{(1-\phi^{m-1})}{(1-\phi)} \epsilon_{t+1} $$ $$ \sum_{i=0}^{m-3} \phi^{i} \epsilon_{t+2} = \frac{(1-\phi^{(m-2)})}{(1-\phi)} \epsilon_{t+2}$$ $$\cdots$$ $$ \sum_{i=0}^{(m-(m-1))} \phi^{i} \epsilon_{t+m-2} = \frac{(1-\rho^2)}{(1-\phi)} \epsilon_{t+m-2}$$ $$ \sum_{i=0}^{(m-m)} \phi^{i} \epsilon_{t+m-1} = \epsilon_{t+m-1}$$ In order to create the final expression for $y_{t+m-1}$ , we need to add all of the geometric series results to the expression involving $x_t$ in equation (1). This results in: $$y_{t+m-1} = \frac{(1-\phi^{m})}{(1-\phi)} x_{t} + \frac{(1-\phi^{m-1})}{(1-\phi)} \epsilon_{t+1} + \frac{(1-\phi^{(m-2)})}{(1-\phi)} \epsilon_{t+2} + \ldots + \frac{(1-\phi^2)}{(1-\phi)} \epsilon_{t+m-2} + \epsilon_{t+m-1}$$ This looks to me kind of like an ARMA(1,(m-1)) but maybe there is a different and better way to go about this and obtain a simpler or different expression ? The fact that the first term is $x_t$ rather than the lagged dependent variable $y_{t}$ may not really be a problem because $x_t$ is the very first term so, rather than $x_t$ , it would actually be the previous m length sum. That's why I have the 1 in the ARMA. Definitely a sloppy derivation even if it's correct. Any inputs, suggestions or corrections are welcome.
Quote
Report

Post Reply

Quoted from Forex.com.bd-Editorial External answer — Cross Validated Stack Exchange Author: Ben Source score (net votes, not local likes): 3 Original post: https://stats.stackexchange.com/a/652992 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. Consider two overlapping AR(1) models for the same stationary process, with the latter expressing the process in terms of increments of $k$ time units: $$\begin{align} X_{t+1} &= \phi X_t + \varepsilon_t, \\[6pt] X_{t+k} &= \kappa X_t + \omega_t. \\[6pt] \end{align}$$ Using the auto-covariance function for a stationary AR(1) process, we can then write the auto-covariance and the variance in terms of either parameterisation as: $$\begin{align} \mathbb{Cov}(X_t,X_{t+k}) &= \frac{\phi^k}{1-\phi^2} \cdot \sigma_\varepsilon^2 \quad \quad \quad \quad \quad \mathbb{V}(X_t) = \frac{1}{1-\phi^2} \cdot \sigma_\varepsilon^2, \\[6pt] \mathbb{Cov}(X_t,X_{t+k}) &= \frac{\kappa}{1-\kappa^2} \cdot \sigma_\omega^2 \quad \quad \quad \quad \quad \mathbb{V}(X_t) = \frac{1}{1-\kappa^2} \cdot \sigma_\omega^2. \\[6pt] \end{align}$$ Equating the variance and auto-covariance in these two expressions of the process gives the transforming equations: $$\begin{align} \kappa &= \phi^k \quad \quad \quad \quad \quad \sigma_\omega^2 = \frac{1-\phi^{2k}}{1-\phi^2} \cdot \sigma_\varepsilon^2, \\[6pt] \phi &= \kappa^{1/k} \quad \quad \quad \quad \ \sigma_\varepsilon^2 = \frac{1-\kappa^{2/k}}{1-\kappa^{2}} \cdot \sigma_\omega^2. \\[6pt] \end{align}$$ This gives you the general parameter correspondence between a stationary AR(1) model measured in single time-units and its corresponding model form when we deal with increments of $k$ time units. In your case you have $k=21$ , $\kappa=0.9$ and $\sigma_\omega^2=0.00027$ which gives the corresponding parameters (for the daily process): $$\phi = 0.9949954 \quad \quad \quad \quad \ \sigma_\varepsilon^2 = 1.418802 \times 10^{-5}.$$

Cancel quote

Checking account access…