Standard methods of quantifying model fit in sparse linear regression/sparse coding
Standard methods of quantifying model fit in sparse linear regression/sparse coding
Loading saved threads...
Mike Battaglia · External communityPost link
External question — Cross Validated Stack Exchange
Author: Mike Battaglia
Original post: https://stats.stackexchange.com/questions/660439
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
Suppose you have a dictionary of signals
$X$
and an output vector
$Y$
, and you want to quantify how well
$Y$
can be represented as a sparse linear combination of columns of
$X$
. Specifically,
$Y = X\beta + N$
, where
$\beta$
is sparse and
$N \sim \mathcal{N}(0, I)$
. This is a sparse linear regression problem.
While methods like LASSO provide a way to estimate
$\beta$
with a sparsity trade-off parameter
$\lambda$
, my goal is different:
How can we quantify whether any sparse linear relationship exists at all?
This is less about estimating
$\beta$
; instead we are evaluating the plausibility of sparsely representing
$Y$
for a given
$X$
. Some thoughts I've had:
Posterior Density at MAP
: In the Bayesian LASSO, simply use the posterior density at the MAP solution (L1-regularized MLE). This is basically what BIC does - but does this make sense, given that BIC is built on Laplace's method and our prior is non-differentiable?
Marginal Likelihood
: Integrate over
$\beta$
to compute
$p(Y \mid X)$
. This is theoretically appealing, and there are some good techniques to estimate this in
this thread
- but is it a standard approach here?
Hypothesis Testing
: Generalize the usual hypothesis testing method in OLS, where we'd test
$H_0: \beta = 0$
. Instead we would have to test
$H_0: \beta = 0$
or that
$\beta$
is "dense." For
$L_0$
-sparsity, this is combinatorially hard; are there good approximations (e.g., L1-based)?
Other Approaches?
Are there simpler or more standard ways to quantify this?
In my situation, the dictionary
$X$
is fixed (not learned), and this is a question of comparing the fit of different data
$Y$
to one given model
$X$
.
This evolved from
this thread
about computing marginal likelihood in the Bayesian LASSO; here I'm curious if that really is the standard way or if one of these easier metrics (e.g. the posterior mode density value) would be good enough.
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External question — Cross Validated Stack Exchange Author: Mike Battaglia Source score (net votes, not local likes): 1 Original post: https://stats.stackexchange.com/questions/660439 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. Suppose you have a dictionary of signals $X$ and an output vector $Y$ , and you want to quantify how well $Y$ can be represented as a sparse linear combination of columns of $X$ . Specifically, $Y = X\beta + N$ , where $\beta$ is sparse and $N \sim \mathcal{N}(0, I)$ . This is a sparse linear regression problem. While methods like LASSO provide a way to estimate $\beta$ with a sparsity trade-off parameter $\lambda$ , my goal is different: How can we quantify whether any sparse linear relationship exists at all? This is less about estimating $\beta$ ; instead we are evaluating the plausibility of sparsely representing $Y$ for a given $X$ . Some thoughts I've had: Posterior Density at MAP : In the Bayesian LASSO, simply use the posterior density at the MAP solution (L1-regularized MLE). This is basically what BIC does - but does this make sense, given that BIC is built on Laplace's method and our prior is non-differentiable? Marginal Likelihood : Integrate over $\beta$ to compute $p(Y \mid X)$ . This is theoretically appealing, and there are some good techniques to estimate this in this thread - but is it a standard approach here? Hypothesis Testing : Generalize the usual hypothesis testing method in OLS, where we'd test $H_0: \beta = 0$ . Instead we would have to test $H_0: \beta = 0$ or that $\beta$ is "dense." For $L_0$ -sparsity, this is combinatorially hard; are there good approximations (e.g., L1-based)? Other Approaches? Are there simpler or more standard ways to quantify this? In my situation, the dictionary $X$ is fixed (not learned), and this is a question of comparing the fit of different data $Y$ to one given model $X$ . This evolved from this thread about computing marginal likelihood in the Bayesian LASSO; here I'm curious if that really is the standard way or if one of these easier metrics (e.g. the posterior mode density value) would be good enough.
Checking account access…