Request: critique my framing of the statistics inference "pipeline" versus ML
Request: critique my framing of the statistics inference "pipeline" versus ML
Loading saved threads...
Chris · External communityPost link
External question — Cross Validated Stack Exchange
Author: Chris
Original post: https://stats.stackexchange.com/questions/664898
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
EDIT
: I got a lot of comments about the length of this post. Let me try to distill it down.
Broadly speaking, here I take "machine learning" to be the process of prediction of outputs for "new" inputs relative to some training data (presumably inputs with the same distribution and support, thus the scare quotes). I take "statistical inference" to be the process of modeling the uncertainty of, among other things, predictions made by models. (Also model fit.)
The earliest historical paradigm seemed to involve generating a "probabilistic dual" for algorithms, e.g. the Gaussian for least squares, and my question was, am I correct in my read of there being a historical 3-step process of
define algorithm,
define "probabilistic dual" distribution that makes it apply,
find conditions for practitioners under which that dual is realistic to the data.
I give two examples of this process below, but only the first one reflects some actual historical digging by me.
I definitely am interested in the opposite direction that some people raised, of intimately understanding a "data-generating process" in terms of finding a probabilistic model for it and then leveraging or deriving an algorithm for fitting a model based on that process.
But my question was, is the three step process I gave effectively a description of how classical statisticians created new model classes? And is it fair to say that step 2 is where modern "machine learning" methods started to veer away, instead keeping just the bias-variance controls on model fit without the UQ?
I have been thinking about the meta/historical processes of statistics, how they differ from ML, and rapprochement between the fields. (This is motivated by interest in uncertainty quantification for "black box" scenarios, like neural or complicated Bayesian models.) I find classical statistics instruction has an "ad-hoc" quality with how it presents new topics and connects algorithmic and probabilistic approaches, and I wonder if I've understood the "duality" there correctly - and also, specifically what statistics exists or has evolved without that dual.
I apologize in advance for the extreme length, but I wanted to try to articulate my understanding and get critique and "wrinkles"/problems in this analysis.
Coming from the ML side, one thing I haven't fully understood for a while is the "pipeline" for statisticians versus ML researchers. Definitionally I'm taking ML as the gamut of prediction techniques, without requiring "inference" via uncertainty quantification or hypothesis testing of the kind that, for specificity, could result in credible/confidence intervals - so ML is then a superset of statistical predictive methods (because some "ML methods" are just direct predictors with little/no uncertainty quantification tooling). This is tricky to be precise about but I am focusing on the lack of a tractable "probabilistic dual" as the defining trait - both to explain the difference and to gesture at what isn't intractable for inference in an "ML" model.
In statistics terms, I'm asking about the soup-to-nuts parametric statistical process as it actually happens, and how it contrasts with ML.
To begin: we know that Gauss
first iterated least squares as one of the techniques he tried for linear regression;
after he decided he liked its performance, he and others worked on defining the Gaussian distribution for the errors as the proper one under which model fitting (here by maximum likelihood with some, today, some information criterion for bias-variance balance, also assuming iid data and errors here - these details I'd like to elide over if possible) coincided with least-squares' answer. So the Gaussian is the "probabilistic dual" to least squares in making that model optimal.
Then he and others conducted research to understand the conditions under which this probabilistic model approximately applied: in particular they found the CLT, a modern form of which helps guarantee things like that betas resulting from least squares follow a normal distribution even when the iid errors assumption is violated. (I need to review exactly what Lindeberg-Levy says.)
So there was a process of:
iterate an algorithm,
define a tractable probabilistic dual and do inference via it,
investigate the circumstances under which that dual was realistic to apply as a modeling assumption, to allow practitioners a scope of confident use
Another example of this, a bit less talked about: logistic regression.
I'm a little unclear on the history but I believe Berkson proposed it, somewhat ad-hoc, as a method for regression on categorical responses;
It was noticed at some point (see Bishop 4.2.4) that there is a "probabilistic dual" in the sense that this model applies, with maximum-likelihood fitting, for linear-in-inputs regression when the class-conditional densities of the data p( x|C_k ) belong to an exponential family;
and then I'm assuming in literature that there were some investigations of how reasonable this assumption was (Bishop motivates a couple of cases)
Now, the ML revolution seems to have thrown this process for a loop by focusing on step 1, but never fulfilling step 2 in the sense of a "tractable" probabilistic model. They realized - SVMs being an early example, but neural networks being quintessential - that there was no need for probabilistic interpretation at all to produce some prediction so long as they kept the aspect of step 2 of handling bias-variance trade-off and finding mechanisms for this; so they defined "loss functions" that they permitted to diverge from tractable probabilistic models or even probabilistic models whatsoever (SVMs).
It turned out that, under the influence of large datasets and with models they were able to endow with huge "capacity," this was enough to get them better predictions than classical models following the 3-step process could have. (How ML researchers quantify goodness of predictions is its own topic I will postpone trying to be precise on.)
Arguably they entered a practically non-parametric framework with their efforts. (The parameters exist only in a weak sense, though far from being a miracle this typically reflects shrewd design choices on what capacity to give.)
Could we comment on and critique this interpretation? I didn't touch either on how ML replaced step 3 - in my experience this can be some brutal trial and error. I'd be happy to try to firm that up.
I do realize a few limitations:
I haven't really touched the nuance of "unsupervised" learners, which don't really exist in classical statistics.
I haven't compared the non-parametric process with ML generally and would like to do so in a later post.
I haven't specifically addressed non-stationary distributions, which would be especially relevant in comparing e.g. Kalman filters and other time series models with recurrent neural models or indeed, still more different, reinforcement learners.
Quote
Report
Sextus Empiricus · External communityPost link
External answer — Cross Validated Stack Exchange
Author: Sextus Empiricus
Original post: https://stats.stackexchange.com/a/665090
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
You have a long post but I can extract a few notions.
You seem to be describing a process like trial and error where an algorithm is designed before it's properties are understood (e.g. least squares or aggregating data by the arithmetic mean came before the Gaussian distribution). So sometimes algorithms are being developed and used in the scientific fields, without them being fully developed yet. That sounds a lot like the development described by Thomas Kuhn in 'structure of scientific revolutions', where development takes some sort of evilutionary path and the intitial steps taken may not be fully sound yet.
I would not characterise this as the difference between machine learning and statistics.
A possible way to describe the difference is with the use of this cycle of improving a model in a cyclic way (google is broken and can't help me with finding a good source)
Create model
Collect data
Analyse data
Improve model (goto 1)
Statistics is important when data collection resources are low and only a single or few loops occur. We test and describe the model performance ideally (but not necessarily) based on assumptions about statistical variations that may occur. The model is fixed (with only some yet unknown parameters that are free) and predesigned by the researcher. The aim is to test whether a model works and characterise it's performance. Only a few missing parts in the knowledge about the model may be tuned.
Machine learning is making multiple loops and does the modeling (partly) by itself. The researcher mostly creates the building blocks but doesn't specify the model in detail. The statistical properties are often too complex to describe in full detail and instead the use of multiple testing on different data is used to verify the properties of the model (which can still be statistics). The aim is to create a (complex) model that works the best, with less focus on the question how (well) it works.
Those two are not really a duality.
It's more that machine learning has a more specific focus (the self-learning part).
Machine learning revolves more around the modeling approach where the algorithm is used to do a considerable part of the model creation. (as an example what it is not: there are
youtubers
and tiktokkers predicting Formula One racing results based on linear regression where the learning step is only the fitting of parameters, that's not machine learning)
So:
A machine learning model, doesn't need to mean that there is no interest in statistical properties. One might be interested in the performance and distribution of output from some complex trained model and make a description of this based on tests with new data. Once we make a characterisation of the distribution of the performance or output, it doesn't stop being machine learning and turn into statistics.
A statistical model is a model that involves the assumptions concerning the distribution of the data. However, not using such assumptions (e.g. applying some technique like least squares without assuming anything about the distribution) doesn't make it directly machine learning.
Quote
Report
Frank Harrell · External communityPost link
External answer — Cross Validated Stack Exchange
Author: Frank Harrell
Original post: https://stats.stackexchange.com/a/665118
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
The framing of the question misses several points. First of all, statistics has always been highly interested in prediction. The distinctions between statistical models (SM) and ML are more related to
SM uses probability models for data. This allows SMs to have a formal and well-performing way to incorporate partial information such as right or interval-censored Y values and to optimally weight the contributions of repeated measurements per subject.
SM have identified parameters that usually are interpretable
Unlike ML, SM places restrictions on the types of effects that are allowed to be present. The most common example of this is the assumption of additivity of predictors (lack of interaction) but you can also place restrictions such as “X vs Y is smooth to 3 orders of continuity” (cubic spline function for X). Restrictions are what makes SM need far smaller sample sizes than ML to have stable results and avoid overfitting.
SM can be used to make inferences about the data generating process or population
Quote
Report
Thomas Speidel · External communityPost link
External answer — Cross Validated Stack Exchange
Author: Thomas Speidel
Original post: https://stats.stackexchange.com/a/665129
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
This is a complex topic, not least because machine learning having found its success largely in commercial applications keeps changing and evolving. I'd encourage you to read some key papers that talk about this topic. In particular:
Breiman's two culture
: Breiman, L. (2001). Statistical modeling: The two cultures (with comments and a rejoinder by the author). Statistical Science, 16(3).
N.B.: you have to read the comments to the paper to better appreciate the other side of the story
.
Donoho's 50 years of data science
Efron's prediction and estimation paper
: Efron, B. (2020). Prediction, Estimation, and Attribution. International Statistical Review, 88(S1), S28–S59.
Berk's book: Berk, R. A. (2020). Statistical Learning from a Regression Perspective. Springer International Publishing.
I want to cite from Berk's book in particular, because chapter 10 in my view eloquently describes some of the fundamental key differences (here I cite from a previous edition of the book - emphases are mine):
Recall Breiman’s distinction between two cultures: a “data modeling culture” and an “algorithmic modeling culture” (2001b). The data modeling culture favors the generalized linear model and its various extensions. A data analysis begins with a mathematical expression meant to represent the mechanisms by which nature works. Estimation serves to fill in the details. The algorithmic modeling culture is concerned solely with linking inputs to outputs. The subject-matter mechanisms connecting the two are not represented and there is, therefore, no a priori vehicle by which inputs are transformed into outputs. A data analysis is undertaken to invent such a vehicle, so that a good fit results.
There is no requirement whatsoever that the vehicle reveals nature’s machinery
.
To clear the terminology confusion, loosely, “
data modeling culture
” refers to statistics and “
algorithmic modeling culture
” refers to machine learning.
But, there is in practice no clear distinction between procedures that belong in the data modeling culture and procedures that belong in the algorithmic modeling culture. In both cultures, information extracted from data is essential. Even for a correct regression model, parameter estimates are obtained from data. Rather, there is a continuum characterized by how much the results depend on substantively informed constraints imposed on the analysis. For conventional regression, at one extreme, there are extensive constraints meant to represent the machinery by which nature proceeds. At the other extreme, random forests and stochastic gradient boosting mine associations in the data with virtually no substantively informed restrictions. Many procedures, such as those within the generalized additive model, fall in between. How then should a data analysis tool be selected? As a first cut, the importance of explicitly representing nature’s machinery should be determined.
If explanation is the dominant data analysis motive, procedures from the data modeling culture should be favored. If prediction is the dominant data analysis motive, procedures from the algorithmic modeling culture should be favored
. If neither is dominant, procedures should be used that are a compromise between the two extremes.
If one is working within the data modeling culture, the choice of procedures is determined primarily by the
correspondence between subjective-matter information available and features of a candidate modeling approach
. The correspondence should substantial. For example,
if nature is known to proceed through a linear combination of causal variables, a form of conventional regression may well be appropriate
. Working within the algorithmic modeling culture, the choice of procedures ideally is primarily determined by out-of-sample performance.
And here's a quote from the Efron paper linked on prediction vs.attribution:
The “weak learners” model of prediction seems dominant in this example. Evidently there are a great many genes weakly correlated with prostate cancer, which can be combined in different combinations to give near-perfect predictions.
This is an advantage if prediction is the only goal,but a disadvantage as far as attribution is concerned. Traditional methods of attribution operate differently, striving as in Table 1 to identify a small set of causal covariates (even if strict causality cannot be inferred)
.
Quote
Report
civilstat · External communityPost link
External answer — Cross Validated Stack Exchange
Author: civilstat
Original post: https://stats.stackexchange.com/a/665137
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
I agree with the perspectives in the other answers, but some are slightly tangential to your questions. I'll add my own attempt at answering your questions as directly as possible.
It sounds like you're asking about the pipeline of how statistical
methods
and ML
methods
are developed, not the pipeline of how data analysts would approach a practical problem. So as I read it, your questions are:
A. Does your 3-step process accurately describe how statisticians usually develop statistical methods?
B. Do developers of ML methods usually skip steps 2 and 3?
A. Does this 3-step process accurately describe how statisticians usually develop statistical methods?
No, this is just one way. In the Statistics community, researchers can start with any of these 3 steps.
Sometimes it works this way. Other times, as noted in Björn's comment, statisticians can start with a distribution for the data first, then find an algorithm that honors these assumptions. (In one example from my own past work, I first realized that I needed a zero-inflated Beta distribution to model my response variable, and then I worked on an algorithm to estimate such a model.)
So, sometimes we go 1, 2, 3. Other times, we start with 2 and/or 3, then work backwards to 1. Also, sometimes we start with an existing combo of 1+2+3 and adapt it all at once to a new 2 (for example, what if the data are not sampled iid but come from a different specific sampling design?) or a new 3 (for example, what if we want this method to be robust to having a few outliers, without necessarily assuming a specific distribution for the outliers?)
B. Do developers of ML methods usually skip steps 2 and 3?
Not always. Maybe this used to be more common in the past, when the ML community was more entrenched in pure CS departments. If early ML papers included theory, they often focused on computational properties (how many steps will this alg take to run?) rather than statistical properties. But over time, many ML researchers have become more interested in statistical assumptions and properties.
In fact, your distinction between "just 1" vs "1+2+3" may be a better description of the ML-to-StatisticalML pipeline, at least a decade or two ago. Traditional-ML researchers would develop an algorithm, then boundary-crossing researchers from the StatML community (like the authors of
Elements of Statistical Learning
) would come along later to figure out what distributional assumptions are a good match for that algorithm.
But since then, my sense is that ESL and related work have helped convince ML researchers to take more interest in statistical properties. It's useful to fill in gaps, prod at hidden assumptions, and bridge connections between different perspectives. For instance, ridge regression was initially motivated as a tweak to linear regression to make the matrix inversion more numerically stable. But since then, some people have found Bayesian priors that also lead to ridge; others have used the ridge penalty with entirely different settings than linear regression; still others have found connections between ridge regularization and neural network dropout. We may not have a complete characterization of deep neural networks in terms of your steps 2 and 3 yet, but bits and pieces are there. Over time, it's become harder than ever to draw a clear line in the sand between statistics and ML.
PS -- Even your example with Gauss and least squares doesn't really work the way you described it. In Step 3, the CLT gives conditions when the
$\hat\beta$
s are approximately normal, but that's different than your Step 2 in which the "probabilistic dual" is the assumption that the errors are normal. We can get approx-normal coefficients without normal errors.
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External answer — Cross Validated Stack Exchange Author: civilstat Source score (net votes, not local likes): 6 Original post: https://stats.stackexchange.com/a/665137 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. I agree with the perspectives in the other answers, but some are slightly tangential to your questions. I'll add my own attempt at answering your questions as directly as possible. It sounds like you're asking about the pipeline of how statistical methods and ML methods are developed, not the pipeline of how data analysts would approach a practical problem. So as I read it, your questions are: A. Does your 3-step process accurately describe how statisticians usually develop statistical methods? B. Do developers of ML methods usually skip steps 2 and 3? A. Does this 3-step process accurately describe how statisticians usually develop statistical methods? No, this is just one way. In the Statistics community, researchers can start with any of these 3 steps. Sometimes it works this way. Other times, as noted in Björn's comment, statisticians can start with a distribution for the data first, then find an algorithm that honors these assumptions. (In one example from my own past work, I first realized that I needed a zero-inflated Beta distribution to model my response variable, and then I worked on an algorithm to estimate such a model.) So, sometimes we go 1, 2, 3. Other times, we start with 2 and/or 3, then work backwards to 1. Also, sometimes we start with an existing combo of 1+2+3 and adapt it all at once to a new 2 (for example, what if the data are not sampled iid but come from a different specific sampling design?) or a new 3 (for example, what if we want this method to be robust to having a few outliers, without necessarily assuming a specific distribution for the outliers?) B. Do developers of ML methods usually skip steps 2 and 3? Not always. Maybe this used to be more common in the past, when the ML community was more entrenched in pure CS departments. If early ML papers included theory, they often focused on computational properties (how many steps will this alg take to run?) rather than statistical properties. But over time, many ML researchers have become more interested in statistical assumptions and properties. In fact, your distinction between "just 1" vs "1+2+3" may be a better description of the ML-to-StatisticalML pipeline, at least a decade or two ago. Traditional-ML researchers would develop an algorithm, then boundary-crossing researchers from the StatML community (like the authors of Elements of Statistical Learning ) would come along later to figure out what distributional assumptions are a good match for that algorithm. But since then, my sense is that ESL and related work have helped convince ML researchers to take more interest in statistical properties. It's useful to fill in gaps, prod at hidden assumptions, and bridge connections between different perspectives. For instance, ridge regression was initially motivated as a tweak to linear regression to make the matrix inversion more numerically stable. But since then, some people have found Bayesian priors that also lead to ridge; others have used the ridge penalty with entirely different settings than linear regression; still others have found connections between ridge regularization and neural network dropout. We may not have a complete characterization of deep neural networks in terms of your steps 2 and 3 yet, but bits and pieces are there. Over time, it's become harder than ever to draw a clear line in the sand between statistics and ML. PS -- Even your example with Gauss and least squares doesn't really work the way you described it. In Step 3, the CLT gives conditions when the $\hat\beta$ s are approximately normal, but that's different than your Step 2 in which the "probabilistic dual" is the assumption that the errors are normal. We can get approx-normal coefficients without normal errors.
Checking account access…