Is a deterministic linkage gate the honest choice when no labelled data exists, and can its false negative rate be bounded at all?
Is a deterministic linkage gate the honest choice when no labelled data exists, and can its false negative rate be bounded at all?
Loading saved threads...
Jak Potvin · External communityPost link
External question — Cross Validated Stack Exchange
Author: Jak Potvin
Original post: https://stats.stackexchange.com/questions/676889
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I am linking two public datasets of vehicle incidents published by the same regulator (NHTSA) through entirely separate channels. One is incident filings submitted by manufacturers under a standing order. The other is consumer complaints filed by vehicle owners. Neither source publishes a key identifying an individual vehicle or an individual event, and no external record states which filing corresponds to which complaint. There is no labelled data and no way to obtain any.
The shared fields
Reporting entity name.
Free text, formatted differently in each source, so it needs normalization and fuzzy comparison.
An 11 character VIN prefix.
WMI, VDS, check digit, model year, plant. Characters 12 to 17, the serial identifying one specific vehicle, are published by neither source. A shared prefix means "same configuration, model year and plant", which thousands of vehicles share.
Incident date.
Filings publish month only (for example "APR-2026"). Complaints carry a full date. The finest shared resolution is the month.
What I built
A deterministic gate. Two records link only if all three hold:
same reporting entity after normalization,
same 11 character VIN prefix,
incident months within 1 of each other.
Entity name similarity is banded at 90 and above for an automatic merge, 80 to 89 for a flagged candidate never merged unattended, below 80 rejected. No similarity score overrides any of the three conditions. The design deliberately biases toward non-linkage, because a false link fabricates an apparent difference between two unrelated events. For scale, against 1,181 independent records the gate produces 44 links in one subpopulation and 0 in the other.
Three weaknesses I can see in my own approach
The thresholds are asserted, not estimated.
The 90/80 bands and the one month tolerance were carried over from earlier work, not fitted to this data. I can defend them as conservative, not as optimal.
The gate gives a binary decision with no calibrated uncertainty.
Every output is labelled low confidence, but that is a word, not a number, and I have no principled route to turning it into one.
I have traded recall for precision without being able to quantify either.
I know the direction of the bias. I do not know its size.
Questions
When no labelled data can exist, is a deterministic gate the defensible choice, or is it just an undeclared probabilistic model with its parameters hidden in threshold constants?
Can a false negative rate be bounded at all in this setting, or is "recall is not estimable here" the honest thing to state and stop?
Does the one month tolerance earn the candidate space it adds? With month only precision on one side, plus or minus one month roughly triples the candidate window, and I cannot tell whether that buys real recall or only noise.
I have posted my own current reasoning as an answer below and would like it argued with.
I am asking about method only. The code built on this linkage computes no match rate as a claim about anyone's behaviour, produces no conduct claims, and names no organization as having done anything. It reports differences between two cited documents and stops. I would rather be told this approach does not hold up than be reassured that it does.
Quote
Report
Jak Potvin · External communityPost link
External answer — Cross Validated Stack Exchange
Author: Jak Potvin
Original post: https://stats.stackexchange.com/a/676890
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
This is my own current reasoning, posted so it can be argued with rather than because I think it settles anything. I am not confident in it.
On the gate versus a probabilistic model
The argument for the gate is that with no labelled data, any probabilistic score here would be a monotone rescaling of my own modelling choices rather than a measurement of anything. EM would still converge and still report m and u, but those values would be driven by my choice of comparison functions and by the class balance I implicitly assumed, not by evidence in the data. A threshold constant at least has the virtue of being visible in a config file, where a reviewer can see it and disagree with it. A fitted parameter that is really an assumption is harder to audit, not easier.
I am aware this is close to an admission that my thresholds are unvalidated assumptions, and I am not sure "at least they are visible" is a sufficient defence. That is the part of question 1 I would most like challenged.
On bounding recall
I think the honest position is that recall is not estimable here. Bounding a false negative rate requires knowing how many true matches exist, and the only route to that quantity in this setting is the linkage itself. Anything I reported would be circular. So I state that recall is unknown rather than estimate it badly, and I accept that this makes the output weaker than a properly validated linkage would be.
On the one month tolerance
My justification is precision honesty rather than performance. One side of the comparison carries month precision only, so a day level gate would assert a precision the data does not have. Plus or minus one month covers the boundary case where an incident near a month end is filed in the adjacent month. I have not tested whether it adds more true links than false candidates, and I do not currently know how I would test that without labels, which is why it is question 3.
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External answer — Cross Validated Stack Exchange Author: Jak Potvin Source score (net votes, not local likes): 2 Original post: https://stats.stackexchange.com/a/676890 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. This is my own current reasoning, posted so it can be argued with rather than because I think it settles anything. I am not confident in it. On the gate versus a probabilistic model The argument for the gate is that with no labelled data, any probabilistic score here would be a monotone rescaling of my own modelling choices rather than a measurement of anything. EM would still converge and still report m and u, but those values would be driven by my choice of comparison functions and by the class balance I implicitly assumed, not by evidence in the data. A threshold constant at least has the virtue of being visible in a config file, where a reviewer can see it and disagree with it. A fitted parameter that is really an assumption is harder to audit, not easier. I am aware this is close to an admission that my thresholds are unvalidated assumptions, and I am not sure "at least they are visible" is a sufficient defence. That is the part of question 1 I would most like challenged. On bounding recall I think the honest position is that recall is not estimable here. Bounding a false negative rate requires knowing how many true matches exist, and the only route to that quantity in this setting is the linkage itself. Anything I reported would be circular. So I state that recall is unknown rather than estimate it badly, and I accept that this makes the output weaker than a properly validated linkage would be. On the one month tolerance My justification is precision honesty rather than performance. One side of the comparison carries month precision only, so a day level gate would assert a precision the data does not have. Plus or minus one month covers the boundary case where an incident near a month end is filed in the adjacent month. I have not tested whether it adds more true links than false candidates, and I do not currently know how I would test that without labels, which is why it is question 3.
Checking account access…