Dealing with Missing Timestamps, but have Delayed Timestamps?
Dealing with Missing Timestamps, but have Delayed Timestamps?
Loading saved threads...
QMath · External communityPost link
External question — Cross Validated Stack Exchange
Author: QMath
Original post: https://stats.stackexchange.com/questions/675138
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
Over an internet connection, I receive live messages about updates to an order book for an asset on a stock exchange. I then record these with the time that I received each message.
For each trade, I then have two timestamps:
The time that the update occurred as recorded by the exchange,
$t$
.
The time that the update message is received by me over said internet connection,
$s$
.
The latency of the internet connection means that for each message,
$s$
is greater than
$t$
. My goal is to analyze these messages as a time series indexed by
$t$
.
My problems are that:
For some messages,
$t$
is missing, but we still have
$s$
.
$t$
is recorded at millisecond precision and
$s$
is recorded at nanosecond precision.
I also do not want to remove any messages from analysis because I think the validity of the order book at any time is dependent on having an unbroken stream of messages.
It seems like an approach that addresses these problems would be to, for each message with a missing
$t$
, impute
$s$
rounded down to the nearest millisecond.
However, does this induce any sort of look-ahead or other bias? and, if so, is there anything “safe” that I can do that doesn’t involve building a model for the delay between occurrence time and time received?
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External question — Cross Validated Stack Exchange Author: QMath Source score (net votes, not local likes): 2 Original post: https://stats.stackexchange.com/questions/675138 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. Over an internet connection, I receive live messages about updates to an order book for an asset on a stock exchange. I then record these with the time that I received each message. For each trade, I then have two timestamps: The time that the update occurred as recorded by the exchange, $t$ . The time that the update message is received by me over said internet connection, $s$ . The latency of the internet connection means that for each message, $s$ is greater than $t$ . My goal is to analyze these messages as a time series indexed by $t$ . My problems are that: For some messages, $t$ is missing, but we still have $s$ . $t$ is recorded at millisecond precision and $s$ is recorded at nanosecond precision. I also do not want to remove any messages from analysis because I think the validity of the order book at any time is dependent on having an unbroken stream of messages. It seems like an approach that addresses these problems would be to, for each message with a missing $t$ , impute $s$ rounded down to the nearest millisecond. However, does this induce any sort of look-ahead or other bias? and, if so, is there anything “safe” that I can do that doesn’t involve building a model for the delay between occurrence time and time received?
Checking account access…