Best Practices for Imputing Missing Data in Trade Data (Linear Interpolation and Random Volume)
Best Practices for Imputing Missing Data in Trade Data (Linear Interpolation and Random Volume)
Loading saved threads...
Mocak · External communityPost link
External question — Cross Validated Stack Exchange
Author: Mocak
Original post: https://stats.stackexchange.com/questions/655747
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I am working on a dataset containing trade data, and my goal is to impute the missing data for a period of around 24 hours. Here's a sample of the trade data I'm working with:
timestamp
symbol
price
quantity
isBuyerMaker
1727682228788
BNBUSDT
582.55000
1.32
True
1727682228837
NEIROUSDT
0.00103
470374.00
False
1727682228982
NEIROUSDT
0.00103
1374.00
True
1727682229035
NEIROUSDT
0.00103
49750.00
True
1727682229035
NEIROUSDT
0.00103
20877.00
True
Unfortunately, I have a missing block of trades spanning nearly 24 hours. While I could fetch the missing trades from the original source, the objective is not to use actual values but to impute the missing data.
Here is an example graph of quantity and price before the imputation:
To impute the missing data, I used Linear Interpolation.
The frequency of the imputed data is based on the mean frequency of trades, excluding any large time gaps.
The quantity is randomly selected from the existing trade volumes.
Here is the graph after filling the missing data:
My question:
Is this approach (using Linear Interpolation for price and random selection of trade volumes) a valid method for handling missing data in trade datasets? Are there any improvements or alternative techniques you would suggest to make the imputation more realistic and reflective of actual market behavior?
Quote
Report
Gijs · External communityPost link
External answer — Cross Validated Stack Exchange
Author: Gijs
Original post: https://stats.stackexchange.com/a/655755
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I don't think it is very realistic or valid, but the answer really depends on your purposes. Let's say you want to use this for backtesting a trading algorithm. Then this method will introduce a large bias. For example, an algorithm that will trade based on the assumption that the price change will be the same as the price change in the previous hour will do really well on this day. And since it will break even on other days (assuming a random walk, not really but bear with me), it might outperform other, more subtle algorithms. Based on your imputed data, you might reach a wrong conclusion on performance. So in this case, I would call this imputation method invalid. And for most purposes really, the outcome of this linear interpolation is really different from a normal day. The question for you is, is it valid enough for your purposes? Why do you need to impute any data?
Quote
Report
Post Reply
Checking account access…