How to generate synthetic FX data for backtesting?
How to generate synthetic FX data for backtesting?
Loading saved threads...
JeremyKun · External communityPost link
External question — Quantitative Finance Stack Exchange
Author: JeremyKun
Original post: https://quant.stackexchange.com/questions/1684
License: CC BY-SA 3.0 — https://creativecommons.org/licenses/by-sa/3.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I want to generate synthetic forex data for the purpose of backtesting my trading algorithms. I have some rough ideas in mind on how to do this:
Start with a curve representing a trend, then randomly generate points around the curve according to a Gaussian or some other distribution. Then take the generated points and somehow generate the bar data (open, high, low, close) around those points; alternatively, add a time factor, and then randomly determine when a tick occurs and collect the data into bars afterward.
My question is: is this far off from the established methods for synthetic data generation? I suppose that raises the more basic question:
are
there any established methods for synthetic data generation? I can't seem to find any writing on this subject, be it a blog post or a research paper.
So in addition to a request for external resources on synthetic data generation, I'd like to know what sorts of distributions best model the relationships between open, high, low, close, or how to generate the appropriate intervals between ticks, the spread between ask and bid prices, etc.
Quote
Report
Peter Cotton · External communityPost link
External answer — Quantitative Finance Stack Exchange
Author: Peter Cotton
Original post: https://quant.stackexchange.com/a/29668
License: CC BY-SA 3.0 — https://creativecommons.org/licenses/by-sa/3.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I've been looking at the same question. I believe the data augmentation literature is relevant. And no doubt some ideas from Good-Turing frequency estimation and descendants can be adapted. Another idea is to transform one exchange rate so that it roughly coincides with another, and use a barrage of off-the-shelf classification algorithms to see if they can determine fake from real data. I think it all depends on which invariants in the data you believe you know well and which residual behavior is common across different time series.
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External question — Quantitative Finance Stack Exchange Author: JeremyKun Source score (net votes, not local likes): 17 Original post: https://quant.stackexchange.com/questions/1684 License: CC BY-SA 3.0 — https://creativecommons.org/licenses/by-sa/3.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I want to generate synthetic forex data for the purpose of backtesting my trading algorithms. I have some rough ideas in mind on how to do this: Start with a curve representing a trend, then randomly generate points around the curve according to a Gaussian or some other distribution. Then take the generated points and somehow generate the bar data (open, high, low, close) around those points; alternatively, add a time factor, and then randomly determine when a tick occurs and collect the data into bars afterward. My question is: is this far off from the established methods for synthetic data generation? I suppose that raises the more basic question: are there any established methods for synthetic data generation? I can't seem to find any writing on this subject, be it a blog post or a research paper. So in addition to a request for external resources on synthetic data generation, I'd like to know what sorts of distributions best model the relationships between open, high, low, close, or how to generate the appropriate intervals between ticks, the spread between ask and bid prices, etc.
Checking account access…