Detrending bimodal data before quantifying variation

Detrending bimodal data before quantifying variation

Manage alerts

Loading saved threads...

Marton Horvath · External communityPost link
External question — Cross Validated Stack Exchange Author: Marton Horvath Original post: https://stats.stackexchange.com/questions/674634 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I have a strictly positive non-stationary time-series data set, showing a positive trend, which results in a bimodal distribution. I aim to quantify the variation in data relative to the mean. However, taking CV = $\sigma_\text{raw} / \mu_\text{raw}$ (or the corresponding expression regarding the log-transformed data) would not accurately represent the variation in data. Therefore, I had the thought to detrend the data, resulting in a zero-centred normal distribution, and take $\sigma_\text{residuals} / \mu_\text{raw}$ for quantifying variation in the data, cleaned from the influence of the linear trend. The timescale of the trend is significantly greater than the "micro" fluctuations I would like to assess. I wonder if: 1) this is a sound approach, 2) there are any established processes for such problems? As an alternative, maybe a rolling CV could be an approach? Maybe DFA (Detrended Fluctuation Analysis) could be another option, but it just feels like overkilling the problem in my case. Edit based on comments: @PeterFlom made a fair observation, writing that the best option one might have is to quantize the data and calculate CV per certain time windows (after a short initial period, the data is quite "flat" before starting to rise). Something I have also been considering. However, for practical reasons, I would prefer trying to describe the data set using a single variable. So perhaps what I am after is rather an optimal trade-off (as I mentioned in one of the comments below, there are other data sets, most of which show considerably simpler behaviour - unimodal normal distribution, no drift). Regarding the variation, I have been thinking about finding a CV alternative. Edit nr.2 after learning from comments: Reflecting upon what I have done so far based on the comments, made me formulate that what should be achieved here is the quantification of variation on two different levels. Firstly, on a "macro" level (e.g., this is represented by the upward trend from ~200 sec in the example), which has not been discussed here in details. Secondly, the "micro" variation has to be quantified, as this holds very important information for me (i.e., the variation which appears to be "noise" in the example below). My reflections lead me to smooth the scatter trace (after excluding the first 10 data points as that "ramp-up" is not meaningful for my analysis) using a moving mean filter (green curve), and now I consider that calculating, e.g., RMSE relative to the smoothened signal, would capture the layer of variation I want. For reference: Raw time-series data (example): Distribution of raw data: Distribution of residuals: To detrend the data, I used a rolling mean approach with a window size of 60 (i.e., $n/10$ ), then calculated res = value - rolling mean.
Quote
Report

Post Reply

Quoted from Forex.com.bd-Editorial External question — Cross Validated Stack Exchange Author: Marton Horvath Source score (net votes, not local likes): 0 Original post: https://stats.stackexchange.com/questions/674634 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I have a strictly positive non-stationary time-series data set, showing a positive trend, which results in a bimodal distribution. I aim to quantify the variation in data relative to the mean. However, taking CV = $\sigma_\text{raw} / \mu_\text{raw}$ (or the corresponding expression regarding the log-transformed data) would not accurately represent the variation in data. Therefore, I had the thought to detrend the data, resulting in a zero-centred normal distribution, and take $\sigma_\text{residuals} / \mu_\text{raw}$ for quantifying variation in the data, cleaned from the influence of the linear trend. The timescale of the trend is significantly greater than the "micro" fluctuations I would like to assess. I wonder if: 1) this is a sound approach, 2) there are any established processes for such problems? As an alternative, maybe a rolling CV could be an approach? Maybe DFA (Detrended Fluctuation Analysis) could be another option, but it just feels like overkilling the problem in my case. Edit based on comments: @PeterFlom made a fair observation, writing that the best option one might have is to quantize the data and calculate CV per certain time windows (after a short initial period, the data is quite "flat" before starting to rise). Something I have also been considering. However, for practical reasons, I would prefer trying to describe the data set using a single variable. So perhaps what I am after is rather an optimal trade-off (as I mentioned in one of the comments below, there are other data sets, most of which show considerably simpler behaviour - unimodal normal distribution, no drift). Regarding the variation, I have been thinking about finding a CV alternative. Edit nr.2 after learning from comments: Reflecting upon what I have done so far based on the comments, made me formulate that what should be achieved here is the quantification of variation on two different levels. Firstly, on a "macro" level (e.g., this is represented by the upward trend from ~200 sec in the example), which has not been discussed here in details. Secondly, the "micro" variation has to be quantified, as this holds very important information for me (i.e., the variation which appears to be "noise" in the example below). My reflections lead me to smooth the scatter trace (after excluding the first 10 data points as that "ramp-up" is not meaningful for my analysis) using a moving mean filter (green curve), and now I consider that calculating, e.g., RMSE relative to the smoothened signal, would capture the layer of variation I want. For reference: Raw time-series data (example): Distribution of raw data: Distribution of residuals: To detrend the data, I used a rolling mean approach with a window size of 60 (i.e., $n/10$ ), then calculated res = value - rolling mean.

Cancel quote

Checking account access…
Detrending bimodal data before quantifying variation | Forex.com.bd