Hawkes process MLE calibration diverges (β → bound) on tick data with millisecond-tied timestamps. Is timestamp jittering the standard fix?

Hawkes process MLE calibration diverges (β → bound) on tick data with millisecond-tied timestamps. Is timestamp jittering the standard fix?

Manage alerts

Loading saved threads...

Timcy Arora · External communityPost link
External question — Quantitative Finance Stack Exchange Author: Timcy Arora Original post: https://quant.stackexchange.com/questions/85703 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I'm fitting a univariate Hawkes process with exponential kernel to real tick-level trade data (Binance BTCUSDT, ~1.25M trades/day, millisecond timestamps). The intensity is: $$\lambda(t) = \mu + \sum_{t_i < t} \alpha e^{-\beta(t - t_i)}$$ I'm maximizing the exact log-likelihood via the standard recursive form (Ozaki 1979): $$R_i = \sum_{j<i} e^{-\beta(t_i - t_j)}, \qquad R_i = e^{-\beta(t_i - t_{i-1})}\left(1 + R_{i-1}\right)$$ $$\log L = \sum_i \log\left(\mu + \alpha R_i\right) - \mu T - \frac{\alpha}{\beta}\sum_i\left(1 - e^{-\beta(T - t_i)}\right)$$ optimized with scipy.optimize.minimize (L-BFGS-B), enforcing $\mu, \alpha, \beta > 0$ and $\alpha < \beta$ (stationarity) via bounds/penalty. Problem: without a bound on $\beta$ , the optimizer diverges $\alpha, \beta$ both run to $10^6$ + while $\alpha/\beta$ (branching ratio) stays sensible (~0.6). With a bound on $\beta$ (e.g. $\beta \le 100$ ), the fit simply pins $\beta$ at exactly the bound rather than converging to an interior value indicating the optimizer still wants to go higher, not that 100 is the right answer. I believe the cause is tied event times : many trades share the same millisecond timestamp (one market order filling against several resting limit orders logged separately), which violates the continuous-time assumption underlying the likelihood, effectively rewarding the optimizer for treating simultaneous events as an infinitely fast excitation burst. My questions: Is tie-breaking (e.g. jittering same-timestamp events by a small uniform offset within the timestamp resolution) the standard/correct fix here, or is there a more principled likelihood correction for tied arrivals in Hawkes estimation? Is there a standard reference for handling this in the market-microstructure Hawkes literature (tick data specifically), as opposed to the general point-process literature? Is a hard upper bound on $\beta$ ever appropriate as anything beyond a stopgap, or does its necessity always indicate a data problem (like ties) rather than a legitimate model constraint?
Quote
Report
dikovaxi · External communityPost link
External answer — Quantitative Finance Stack Exchange Author: dikovaxi Original post: https://quant.stackexchange.com/a/85819 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. Short version. No — jittering is not the fix, and the bound is not a stopgap either: with tied timestamps the exponential-kernel MLE does not exist , so any device that makes the optimiser stop (a bound, a jitter width) simply becomes the estimate. The principled fix is to make the timestamp define the event: collapse records that share a timestamp into one event with a mark (fill count or volume), which is also what physically happened — one taker order matched against several resting orders. After that the likelihood has an interior maximum. I reproduced your situation on the same market and public data (Binance USD-M BTCUSDT trades and aggTrades , 2024-03-26, data.binance.vision) so the numbers below can be re-run. Why the likelihood diverges. In the Ozaki recursion a tie ( $t_i=t_{i-1}$ ) gives $R_i = 1+R_{i-1}$ exactly, so for a tied event $\log(\mu+\alpha R_i)\ge\log\alpha$ , while the compensator $\frac{\alpha}{\beta}\sum_i\big(1-e^{-\beta(T-t_i)}\big)\approx \frac{\alpha}{\beta}\,N$ stays bounded when $\alpha,\beta\to\infty$ at fixed $\alpha/\beta$ . The log-likelihood therefore increases without limit along that ray — which is precisely the " $\alpha,\beta\to10^6$ with a sensible branching ratio" you observe. It is not the optimiser; there is no maximum to find. How tied the data really is (whole day, 4.47 M trades , 1.80 M aggTrades ): 84% of trade records share their millisecond with another record, 74% of inter-event gaps are exactly zero (groups average 3.8 fills, p99 = 38, max 371). aggTrades does not solve it — 66% still tied, 56% zero gaps — because a market order that walks several price levels produces one aggregated record per level, all with the same timestamp. Fits on a 20-minute window (08:00–08:20 UTC; exact likelihood, L-BFGS-B on log-parameters, $\beta$ free up to $10^7$ ): data events $\hat\beta$ (1/s) $\hat\alpha/\hat\beta$ raw trades 115,771 runs to the bound 0.73 raw trades, $\beta\le100$ imposed 115,771 100 (pinned) 0.87 raw trades + jitter U(0, 1 ms) 115,771 4,830 0.76 raw trades + jitter U(0, 10 ms) 115,771 623 0.79 aggTrades 44,146 runs to the bound 0.57 one event per ms, mark = fill count 19,105 52.5 (interior) 0.21 same, 3-exponential kernel 19,105 177 / 6.6 / 0.11 0.69 The profile log-likelihood over $\beta$ (with $\mu,\alpha$ re-optimised at each value) tells the same story: for raw trades and for aggTrades it rises monotonically from $\beta=1$ to $10^6$ (by 8.5×10⁵ and 2.6×10⁵ log-likelihood units respectively), whereas for the tie-collapsed marked data it peaks at $\beta\approx10^2$ and falls by 1,200 by $\beta=10^3$ . Two things to read off the table: Jitter sets the answer. $\hat\beta$ is about 5 / (jitter width): 1 ms gives ~5,000, 10 ms gives ~600. You are not estimating a decay rate, you are estimating your own uniform distribution. The same is true of the bound: at $\beta\le100$ you get "0.87" for the branching ratio, a number with no content. Collapsing ties gives a well-posed problem — but note what happens to the branching ratio: 0.21 with one exponential, 0.69 with three (time scales ≈ 6 ms, 150 ms, 9 s). A single exponential fitted to tick data is badly misspecified — real kernels are slowly decaying, roughly power-law over many decades (Bacry, Dayri & Muzy 2012; Hardiman, Bercot & Bouchaud 2013) — and a single scale absorbs whichever part of the kernel the data forces on it. With ties present, that is the zero-lag spike; without them, the sub-second cluster, and the slow mass is missed. If the branching ratio is the quantity of interest, use a sum of exponentials (three or four scales from ms to minutes) or a non-parametric kernel; the exponential value is not comparable across papers or preprocessing choices. Your three questions. Jitter or something more principled? Jitter is used (and is what several toolkits do quietly), but it is a prior on $\beta$ dressed as preprocessing. The principled options are (a) aggregate same-timestamp records into one marked event — the intensity becomes $\lambda(t)=\mu+\sum_{t_j<t}\alpha\,m_j e^{-\beta(t-t_j)}$ and the recursion is $R_i=e^{-\beta(t_i-t_{i-1})}(m_{i-1}+R_{i-1})$ ; or (b) accept that the data are discrete and fit at the resolution you have — bin the timeline and estimate the kernel as an INAR(p) (Kirchner 2017), which handles simultaneity by construction. Sub-millisecond excitation is not identifiable from millisecond stamps under any method. References for tick data specifically: Bowsher (2007, J. Econometrics ) on trades/quotes as multivariate Hawkes with same-timestamp handling; Filimonov & Sornette (2015, Quantitative Finance ) "Apparent criticality and calibration issues in the Hawkes self-excited point process model", which treats timestamp discretisation and the resulting bias in $\hat\beta$ and the branching ratio; Lallouache & Challet (2016, Quantitative Finance ) "The limits of statistical significance of Hawkes processes fitted to financial data", on how the timestamp resolution bounds the shortest identifiable kernel scale; Bacry, Mastromatteo & Muzy (2015) for the review; Bacry, Dayri & Muzy (2012) and Hardiman, Bercot & Bouchaud (2013) for the kernel shape. Is a hard bound ever legitimate? Only as a diagnostic: if $\hat\beta$ sits on any bound you choose, the MLE does not exist for that data/model pair and the estimate is the bound. It always means ties (or a zero-lag spike the exponential cannot represent), never a model constraint. # marked, tie-free recursion (t in seconds, m = fills per timestamp) t, m = np.unique(t_ms, return_counts=True); t = t / 1e3 R = 0.0; ll = 0.0 for i in range(len(t)): if i: R = np.exp(-beta * (t[i] - t[i-1])) * (m[i-1] + R) ll += np.log(mu + alpha * R) ll -= mu * T + (alpha / beta) * np.sum(m * (1 - np.exp(-beta * (T - t))))
Quote
Report

Post Reply

Quoted from Forex.com.bd-Editorial External question — Quantitative Finance Stack Exchange Author: Timcy Arora Source score (net votes, not local likes): 1 Original post: https://quant.stackexchange.com/questions/85703 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I'm fitting a univariate Hawkes process with exponential kernel to real tick-level trade data (Binance BTCUSDT, ~1.25M trades/day, millisecond timestamps). The intensity is: $$\lambda(t) = \mu + \sum_{t_i < t} \alpha e^{-\beta(t - t_i)}$$ I'm maximizing the exact log-likelihood via the standard recursive form (Ozaki 1979): $$R_i = \sum_{j<i} e^{-\beta(t_i - t_j)}, \qquad R_i = e^{-\beta(t_i - t_{i-1})}\left(1 + R_{i-1}\right)$$ $$\log L = \sum_i \log\left(\mu + \alpha R_i\right) - \mu T - \frac{\alpha}{\beta}\sum_i\left(1 - e^{-\beta(T - t_i)}\right)$$ optimized with scipy.optimize.minimize (L-BFGS-B), enforcing $\mu, \alpha, \beta > 0$ and $\alpha < \beta$ (stationarity) via bounds/penalty. Problem: without a bound on $\beta$ , the optimizer diverges $\alpha, \beta$ both run to $10^6$ + while $\alpha/\beta$ (branching ratio) stays sensible (~0.6). With a bound on $\beta$ (e.g. $\beta \le 100$ ), the fit simply pins $\beta$ at exactly the bound rather than converging to an interior value indicating the optimizer still wants to go higher, not that 100 is the right answer. I believe the cause is tied event times : many trades share the same millisecond timestamp (one market order filling against several resting limit orders logged separately), which violates the continuous-time assumption underlying the likelihood, effectively rewarding the optimizer for treating simultaneous events as an infinitely fast excitation burst. My questions: Is tie-breaking (e.g. jittering same-timestamp events by a small uniform offset within the timestamp resolution) the standard/correct fix here, or is there a more principled likelihood correction for tied arrivals in Hawkes estimation? Is there a standard reference for handling this in the market-microstructure Hawkes literature (tick data specifically), as opposed to the general point-process literature? Is a hard upper bound on $\beta$ ever appropriate as anything beyond a stopgap, or does its necessity always indicate a data problem (like ties) rather than a legitimate model constraint?

Cancel quote

Checking account access…