Is there a way to derive a fair price from cryptocurrency trade (no quote) data that is free of bid-ask bounce?

Is there a way to derive a fair price from cryptocurrency trade (no quote) data that is free of bid-ask bounce?

Manage alerts

Loading saved threads...

QMath · External communityPost link
External question — Quantitative Finance Stack Exchange Author: QMath Original post: https://quant.stackexchange.com/questions/83961 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I'm working with Kraken historical ETH-USD trade data from 2017 onward, which includes: Timestamp (Datetime) Trade ID Trade price Trade volume Taker side (buy/sell) Order type (market or marketable limit) My goal is to derive a fair (efficient) price series that's minimally affected by bid-ask bounce with the goal being to use it for execution backtesting and modeling/analysis. I am okay with downsampling to 1 and maybe 5 minute bars, but would like to keep the data as granular as possible. Approaches tried: OHLCV bars (1, 5, 15m): Log returns on close prices still show autocorrelation (≤ -0.1) at lag 1–3. Spread estimation (Roll's, Corwin-Schultz): Often gives NaNs or zero/negative spreads — possibly due to model or implementation issues. Low-order linear ARMA on log returns: Limited success. I haven’t yet tried rolling ARMA. I saw that this question somewhat addresses my issue, but it seems to me that Hasbrouck’s model requires simultaneous estimation of the efficient price and the spread? Thanks for any input/references on this.
Quote
Report
dikovaxi · External communityPost link
External answer — Quantitative Finance Stack Exchange Author: dikovaxi Original post: https://quant.stackexchange.com/a/85818 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. Short version. Because your data has the taker side, you don't need Roll, Corwin–Schultz or a joint Hasbrouck-style estimation at all: the trade sign $q_t\in\{+1,-1\}$ is exactly the latent variable those models exist to infer. With the sign observed, the bounce is a known component and the sign-adjusted price $$\hat m_t = p_t - c\,q_t,\qquad c=\frac{\operatorname{cov}(\Delta p_t,\Delta q_t)}{\operatorname{var}(\Delta q_t)}$$ removes it. That is the Huang–Stoll / Glosten–Harris regression $\Delta p_t = c\,\Delta q_t+\varepsilon_t$ — OLS with one regressor — and $c$ is the effective half-spread. Estimate it per day (or per hour if the spread moves intraday), then sample $\hat m_t$ at the bar close. Two things I would not do: don't build the return series from bar VWAP (averaging a random walk inside the bar produces positive autocorrelation, Working 1960; I measure +0.2 to +0.3 at lag 1 below), and don't use the "last taker-buy price = ask, last taker-sell price = bid" mid as your price series — it is the right tool for simulating fills , because it tells you the touch on each side, but as a mid it goes stale on one side and smooths returns. A test on a market where the truth is available. Binance USD-M perpetuals publish both aggTrades (with taker side) and bookTicker (every top-of-book change) as daily files on data.binance.vision (bookTicker until March 2024), so trade-only estimators can be checked against the real quoted mid. Here is 2024-03-26 for XRPUSDT: tick = 1.56 bp, quoted spread = 1 tick 99% of the day, 130 trades/min, 97% of taker trades execute at the prevailing touch — a clean bounce regime, close to what you have. Lag-1 autocorrelation of log returns, and mean distance of the bar-close level from the true quoted mid: last trade last-buy/last-sell mid $p_t - c\,q_t$ true quoted mid ACF(1), 1 s bars −0.173 +0.095 +0.034 +0.044 ACF(1), 5 s bars −0.041 +0.051 +0.032 +0.029 ACF(1), 60 s bars −0.045 −0.036 −0.035 −0.036 mean |level − true mid| 0.78 bp 0.23 bp 0.15 bp 0 $c$ estimated on the whole day is 0.47 tick = 0.73 bp, and the 0.78 bp error of the last-trade series is, as it should be, the half spread. The sign-adjusted price sits on the quoted mid within 0.15 bp and has the same autocorrelation as the mid at every horizon: the bounce (−0.17 at 1 s) is gone. Note also that by 60 s the bounce is invisible in every series on this venue — whatever negative autocorrelation remains at 1–5 min is present in the true mid too, i.e. it is genuine short-horizon mean reversion, not microstructure. BTCUSDT on the same day teaches the opposite lesson. Quoted spread is 1 tick (0.014 bp), but the touch usually holds only a few thousandths of a BTC, only 40% of taker trades execute at the pre-trade touch (the rest sweep several levels within the same millisecond), and the regression half-spread is 3.35 ticks — 6.7× the quoted half-spread. For an execution backtest that effective spread is the number you want, and you get it from trades + signs; the quotes would have misled you. Why Roll and Corwin–Schultz returned NaNs or nonsense. Roll assumes the negative autocovariance of price changes is caused by the spread . At the trade clock on BTCUSDT it returns 52 ticks (the autocovariance there comes from sweeps that revert, not from the quote). On 1–5-minute bars it simply estimates whatever mean reversion exists at that horizon: on XRPUSDT it gives 1.1 bp at 1 s (≈ the true 1.56), 3.4 bp at 1 min and 7.6 bp at 5 min — growing with the bar while the quoted spread is constant. Corwin–Schultz assumes the high–low range is spread plus diffusion volatility; on minute-scale crypto bars the range is dominated by volatility, the estimator goes negative and you get NaN. Nothing is wrong with your implementation; the estimators are being asked a question the data doesn't answer. About lags 2–3. Bid–ask bounce is an MA(1) effect: it lives at lag 1 only. If −0.1 at lags 2 and 3 survives in $\hat m_t$ , it is a transitory price component (Kraken in 2017 was thin and often slow relative to the larger venues, and large trades pushed price and reverted over minutes). The tool for that is a state-space model with a random-walk efficient price and a stationary transitory term — Hasbrouck (1993), Menkveld, Koopman & Lucas (2007). But note that the "simultaneous estimation of efficient price and spread" you were worried about disappears once $q_t$ is observed: the measurement equation is $p_t = m_t + c\,q_t + u_t$ with $q_t$ known, so the Kalman filter has three or four parameters and $c$ is essentially the OLS coefficient above. Recipe $q_t=+1$ for a taker buy, $-1$ for a taker sell; $c$ by OLS of $\Delta p$ on $\Delta q$ , per day or per hour. Price series: $\hat m_t = p_t - c\,q_t$ , sampled as the last observation in each bar, never averaged. Fills in the backtest: buy at the last taker-buy price (ask proxy), sell at the last taker-sell price (bid proxy), with an age check on that side's last observation. Check: ACF(1) of returns ≈ 0 (or equal to whatever the venue's mid has), and realized variance roughly flat across sampling intervals (the volatility signature plot of Andersen, Bollerslev, Diebold & Labys 2000). If the signature keeps rising as you sample finer, there is noise left. q = np.where(df.taker_side == 'buy', 1, -1) dp, dq = df.price.diff(), pd.Series(q, index=df.index).diff() c = (dp * dq).sum() / (dq * dq).sum() # effective half-spread df['m_hat'] = df.price - c * q df['ask_proxy'] = df.price.where(q == 1).ffill() df['bid_proxy'] = df.price.where(q == -1).ffill() bars = df.set_index('ts').resample('1min').last() # last, not mean References: Roll (1984) J. Finance ; Glosten & Harris (1988) JFE ; Huang & Stoll (1997) RFS ; Hasbrouck (1993) RFS "Assessing the quality of a security market"; Menkveld, Koopman & Lucas (2007) JBES ; Andersen, Bollerslev, Diebold & Labys (2000) "Great realizations"; Working (1960) Econometrica on autocorrelation induced by averaging.
Quote
Report

Post Reply

Checking account access…