Is there a way to derive a fair price from cryptocurrency trade (no quote) data that is free of bid-ask bounce?
Is there a way to derive a fair price from cryptocurrency trade (no quote) data that is free of bid-ask bounce?
Loading saved threads...
QMath · External communityPost link
External question — Quantitative Finance Stack Exchange
Author: QMath
Original post: https://quant.stackexchange.com/questions/83961
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I'm working with Kraken historical ETH-USD trade data from 2017 onward, which includes:
Timestamp (Datetime)
Trade ID
Trade price
Trade volume
Taker side (buy/sell)
Order type (market or marketable limit)
My goal is to derive a fair (efficient) price series that's minimally affected by bid-ask bounce with the goal being to use it for execution backtesting and modeling/analysis. I am okay with downsampling to 1 and maybe 5 minute bars, but would like to keep the data as granular as possible.
Approaches tried:
OHLCV bars (1, 5, 15m): Log returns on close prices still show autocorrelation (≤ -0.1) at lag 1–3.
Spread estimation (Roll's, Corwin-Schultz): Often gives NaNs or zero/negative spreads — possibly due to model or implementation issues.
Low-order linear ARMA on log returns: Limited success. I haven’t yet tried rolling ARMA.
I saw that
this
question somewhat addresses my issue, but it seems to me that Hasbrouck’s model requires simultaneous estimation of the efficient price and the spread?
Thanks for any input/references on this.
Quote
Report
dikovaxi · External communityPost link
External answer — Quantitative Finance Stack Exchange
Author: dikovaxi
Original post: https://quant.stackexchange.com/a/85818
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
Short version.
Because your data has the taker side, you don't need Roll, Corwin–Schultz or a joint Hasbrouck-style estimation at all: the trade sign
$q_t\in\{+1,-1\}$
is exactly the latent variable those models exist to infer. With the sign observed, the bounce is a
known
component and the sign-adjusted price
$$\hat m_t = p_t - c\,q_t,\qquad c=\frac{\operatorname{cov}(\Delta p_t,\Delta q_t)}{\operatorname{var}(\Delta q_t)}$$
removes it. That is the Huang–Stoll / Glosten–Harris regression
$\Delta p_t = c\,\Delta q_t+\varepsilon_t$
— OLS with one regressor — and
$c$
is the effective half-spread. Estimate it per day (or per hour if the spread moves intraday), then sample
$\hat m_t$
at the bar close. Two things I would
not
do: don't build the return series from bar VWAP (averaging a random walk inside the bar produces positive autocorrelation, Working 1960; I measure +0.2 to +0.3 at lag 1 below), and don't use the "last taker-buy price = ask, last taker-sell price = bid" mid as your price series — it is the right tool for
simulating fills
, because it tells you the touch on each side, but as a mid it goes stale on one side and smooths returns.
A test on a market where the truth is available.
Binance USD-M perpetuals publish both
aggTrades
(with taker side) and
bookTicker
(every top-of-book change) as daily files on data.binance.vision (bookTicker until March 2024), so trade-only estimators can be checked against the real quoted mid. Here is 2024-03-26 for XRPUSDT: tick = 1.56 bp, quoted spread = 1 tick 99% of the day, 130 trades/min, 97% of taker trades execute at the prevailing touch — a clean bounce regime, close to what you have. Lag-1 autocorrelation of log returns, and mean distance of the bar-close level from the true quoted mid:
last trade
last-buy/last-sell mid
$p_t - c\,q_t$
true quoted mid
ACF(1), 1 s bars
−0.173
+0.095
+0.034
+0.044
ACF(1), 5 s bars
−0.041
+0.051
+0.032
+0.029
ACF(1), 60 s bars
−0.045
−0.036
−0.035
−0.036
mean |level − true mid|
0.78 bp
0.23 bp
0.15 bp
0
$c$
estimated on the whole day is 0.47 tick = 0.73 bp, and the 0.78 bp error of the last-trade series is, as it should be, the half spread. The sign-adjusted price sits on the quoted mid within 0.15 bp and has the same autocorrelation as the mid at every horizon: the bounce (−0.17 at 1 s) is gone. Note also that by 60 s the bounce is invisible in
every
series on this venue — whatever negative autocorrelation remains at 1–5 min is present in the true mid too, i.e. it is genuine short-horizon mean reversion, not microstructure.
BTCUSDT on the same day teaches the opposite lesson. Quoted spread is 1 tick (0.014 bp), but the touch usually holds only a few thousandths of a BTC, only 40% of taker trades execute at the pre-trade touch (the rest sweep several levels within the same millisecond), and the regression half-spread is 3.35 ticks — 6.7× the quoted half-spread. For an execution backtest that effective spread is the number you want, and you get it from trades + signs; the quotes would have misled you.
Why Roll and Corwin–Schultz returned NaNs or nonsense.
Roll assumes the negative autocovariance of price changes is
caused by the spread
. At the trade clock on BTCUSDT it returns 52 ticks (the autocovariance there comes from sweeps that revert, not from the quote). On 1–5-minute bars it simply estimates whatever mean reversion exists at that horizon: on XRPUSDT it gives 1.1 bp at 1 s (≈ the true 1.56), 3.4 bp at 1 min and 7.6 bp at 5 min — growing with the bar while the quoted spread is constant. Corwin–Schultz assumes the high–low range is spread plus diffusion volatility; on minute-scale crypto bars the range is dominated by volatility, the estimator goes negative and you get NaN. Nothing is wrong with your implementation; the estimators are being asked a question the data doesn't answer.
About lags 2–3.
Bid–ask bounce is an MA(1) effect: it lives at lag 1 only. If −0.1 at lags 2 and 3 survives in
$\hat m_t$
, it is a transitory price component (Kraken in 2017 was thin and often slow relative to the larger venues, and large trades pushed price and reverted over minutes). The tool for that is a state-space model with a random-walk efficient price and a stationary transitory term — Hasbrouck (1993), Menkveld, Koopman & Lucas (2007). But note that the "simultaneous estimation of efficient price and spread" you were worried about disappears once
$q_t$
is observed: the measurement equation is
$p_t = m_t + c\,q_t + u_t$
with
$q_t$
known, so the Kalman filter has three or four parameters and
$c$
is essentially the OLS coefficient above.
Recipe
$q_t=+1$
for a taker buy,
$-1$
for a taker sell;
$c$
by OLS of
$\Delta p$
on
$\Delta q$
, per day or per hour.
Price series:
$\hat m_t = p_t - c\,q_t$
, sampled as the
last
observation in each bar, never averaged.
Fills in the backtest: buy at the last taker-buy price (ask proxy), sell at the last taker-sell price (bid proxy), with an age check on that side's last observation.
Check: ACF(1) of returns ≈ 0 (or equal to whatever the venue's mid has), and realized variance roughly flat across sampling intervals (the volatility signature plot of Andersen, Bollerslev, Diebold & Labys 2000). If the signature keeps rising as you sample finer, there is noise left.
q = np.where(df.taker_side == 'buy', 1, -1)
dp, dq = df.price.diff(), pd.Series(q, index=df.index).diff()
c = (dp * dq).sum() / (dq * dq).sum() # effective half-spread
df['m_hat'] = df.price - c * q
df['ask_proxy'] = df.price.where(q == 1).ffill()
df['bid_proxy'] = df.price.where(q == -1).ffill()
bars = df.set_index('ts').resample('1min').last() # last, not mean
References: Roll (1984)
J. Finance
; Glosten & Harris (1988)
JFE
; Huang & Stoll (1997)
RFS
; Hasbrouck (1993)
RFS
"Assessing the quality of a security market"; Menkveld, Koopman & Lucas (2007)
JBES
; Andersen, Bollerslev, Diebold & Labys (2000) "Great realizations"; Working (1960)
Econometrica
on autocorrelation induced by averaging.
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External question — Quantitative Finance Stack Exchange Author: QMath Source score (net votes, not local likes): 1 Original post: https://quant.stackexchange.com/questions/83961 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I'm working with Kraken historical ETH-USD trade data from 2017 onward, which includes: Timestamp (Datetime) Trade ID Trade price Trade volume Taker side (buy/sell) Order type (market or marketable limit) My goal is to derive a fair (efficient) price series that's minimally affected by bid-ask bounce with the goal being to use it for execution backtesting and modeling/analysis. I am okay with downsampling to 1 and maybe 5 minute bars, but would like to keep the data as granular as possible. Approaches tried: OHLCV bars (1, 5, 15m): Log returns on close prices still show autocorrelation (≤ -0.1) at lag 1–3. Spread estimation (Roll's, Corwin-Schultz): Often gives NaNs or zero/negative spreads — possibly due to model or implementation issues. Low-order linear ARMA on log returns: Limited success. I haven’t yet tried rolling ARMA. I saw that this question somewhat addresses my issue, but it seems to me that Hasbrouck’s model requires simultaneous estimation of the efficient price and the spread? Thanks for any input/references on this.
Checking account access…