How to generate synthetic FX data for backtesting?

How to generate synthetic FX data for backtesting?

Manage alerts

Loading saved threads...

JeremyKun · External communityPost link
External question — Quantitative Finance Stack Exchange Author: JeremyKun Original post: https://quant.stackexchange.com/questions/1684 License: CC BY-SA 3.0 — https://creativecommons.org/licenses/by-sa/3.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I want to generate synthetic forex data for the purpose of backtesting my trading algorithms. I have some rough ideas in mind on how to do this: Start with a curve representing a trend, then randomly generate points around the curve according to a Gaussian or some other distribution. Then take the generated points and somehow generate the bar data (open, high, low, close) around those points; alternatively, add a time factor, and then randomly determine when a tick occurs and collect the data into bars afterward. My question is: is this far off from the established methods for synthetic data generation? I suppose that raises the more basic question: are there any established methods for synthetic data generation? I can't seem to find any writing on this subject, be it a blog post or a research paper. So in addition to a request for external resources on synthetic data generation, I'd like to know what sorts of distributions best model the relationships between open, high, low, close, or how to generate the appropriate intervals between ticks, the spread between ask and bid prices, etc.
Quote
Report
Peter Cotton · External communityPost link
External answer — Quantitative Finance Stack Exchange Author: Peter Cotton Original post: https://quant.stackexchange.com/a/29668 License: CC BY-SA 3.0 — https://creativecommons.org/licenses/by-sa/3.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I've been looking at the same question. I believe the data augmentation literature is relevant. And no doubt some ideas from Good-Turing frequency estimation and descendants can be adapted. Another idea is to transform one exchange rate so that it roughly coincides with another, and use a barrage of off-the-shelf classification algorithms to see if they can determine fake from real data. I think it all depends on which invariants in the data you believe you know well and which residual behavior is common across different time series.
Quote
Report

Post Reply

Checking account access…