Choosing a Bayesian Likelihood Model for Baseline Noise
Choosing a Bayesian Likelihood Model for Baseline Noise
Loading saved threads...
ohshitgorillas · External communityPost link
External question — Cross Validated Stack Exchange
Author: ohshitgorillas
Original post: https://stats.stackexchange.com/questions/667737
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
$\newcommand{\ysignal}{y_\text{signal}}$
$\newcommand{\ystray}{y_\text{stray}}$
$\newcommand{\noise}{\varepsilon}$
I'm developing a Bayesian model to perform a baseline correction on mass spec data, and I've run into a classic model specification problem where I would appreciate some expert guidance.
Goal
The goal is to estimate a true, non-negative signal (
$\ysignal$
) from a raw measurement (
$y'$
). My physical model is:
$$y' = \ysignal + \ystray + \noise$$
where:
$\ysignal$
represents the signal from the target ions of interest
$\ystray$
represents the undesired or 'stray' ions also contributing to the signal
$\noise$
represents a normally distributed, zero-centered instrumental noise component
The challenge lies in modeling
$\ystray$
and
$\noise$
to isolate
$\ysignal$
.
My Probabilistic Framework
We can measure
$\ystray$
and
$\noise$
directly by directing the our mass spec at a mass where no gaseous species exist, thereby ensuring that
$\ysignal = 0$
. We denote these measurements
$B$
:
$$B = \ystray + \noise$$
with mean
$\bar{B}$
and standard deviation
$\sigma_B$
.
Traditional baseline correction involves subtracting
$\bar{B}$
(
$\approx \ystray$
), thereby leaving the noise term
$\noise$
in the final value, decreasing precision and accuracy. I wanted to do better.
To solve for
$\ysignal$
given a raw measurement
$y$
and baseline measurements
$B$
, I use Bayesian inference. The framework is as follows:
Prior
$P(\ysignal)$
: I assume a log-uniform prior for
$\ysignal$
, enforcing the physical constraint that
$\ysignal \geq 0$
.
Likelihood
$P(y\mid\ysignal)$
: The core of the model is a Gaussian likelihood function which assumes the raw measurement
$y'$
is a draw from a normal distribution where the mean is the sum of the true signal and the stray ion component, and the variance includes all known error sources:
$$P(y'\mid\ysignal) =
\frac{1}{\sqrt{2\pi \cdot \text{var}}}
\exp{\left(-\frac{(y' - \ysignal - \ystray)^2}{2 \cdot \text{var}}\right)}$$
Posterior
$P(\ysignal\mid y')$
: Calculated via Bayes' theorem:
$$ P(\ysignal\mid y') = \frac{P(y'\mid\ysignal)P(\ysignal)}{P(y')}$$
The final "baseline corrected" value
$y$
was the mean of the posterior distribution:
$$ y = \int_0^\infty\ysignal P(\ysignal\mid y') \, d\ysignal$$
The Problem
My initial approach was to estimate
$\ystray \approx \bar{B}$
and the baseline's contribution to
$\text{var}$
as
$\sigma_B$
.
I now know from extensive characterization runs (
$N=500$
) that
$B$
does
not
follow a Gaussian distribution. It is non-negative and skewed, and is perhaps better described by a Rice distribution, which would at least make physical sense (DC offset
${} = \ystray$
with Gaussian noise component
$\varepsilon$
).
Here's a plot of the distribution showing the median and MAD as the flat line and shaded region, respectively:
The Constraint
The above plot of
$N=500$
is fully sufficient to characterize
$B$
, but we're not measuring
$N=500$
most of the time. We're measuring
$N=12,$
or at most
$N=20$
per analysis. That's plenty for characterizing our gases of interest (helium isotopes) most of the time, but it's not sufficient for modeling a Rice distribution or anything "fancy" like that.
Examples of what we're actually measuring:
Therefore, I seem to be stuck between a rock and a hard place: Do I...
Keep the Gaussian likelihood function as a practical approximation, but with more robust inputs like the median and MAD? Is there perhaps some justification for this that I'm not seeing?
Use a Rice or other distribution with potentially highly uncertain estimates?
(Your answer here)
My Question
What is the most statistically sound and practical approach for defining the likelihood function in this situation? I need guidance on how to navigate the trade-offs between a potentially misspecified but stable versus a correctly specified but highly uncertain model. Or, ideally, a third option that I'm not seeing...
Data
3.457963e-15 2.629714e-15 2.606773e-15 1.411953e-15 1.765874e-15
5.851072e-15 1.457842e-15 3.371775e-16 4.986981e-15 6.955586e-16
6.543158e-15 4.299480e-15 5.413442e-15 8.380951e-15 5.418035e-15
2.564611e-15 6.686383e-15 5.686884e-15 1.673735e-15 5.096976e-15
1.061633e-15 3.325026e-15 8.707826e-16 4.113838e-15 2.356524e-15
4.680685e-15 2.060269e-15 1.357887e-15 4.596478e-15 4.508928e-15
1.560390e-15 1.254340e-15 1.017982e-15 5.021950e-15 3.865316e-16
1.256796e-14 2.365824e-15 4.316716e-16 6.041665e-15 9.432666e-15
2.614707e-15 2.732145e-15 5.087797e-15 1.435640e-15 1.795884e-15
3.649926e-15 1.236356e-16 2.812509e-16 1.488431e-14 5.525050e-15
9.893365e-16 6.510543e-15 3.520711e-15 2.135790e-15 4.433280e-16
6.611483e-15 5.184774e-15 9.191476e-16 5.654639e-15 6.780510e-15
5.927581e-15 1.368056e-15 4.728397e-16 6.219872e-15 6.668158e-15
1.273860e-14 1.596849e-15 2.400918e-15 2.295510e-15 1.047135e-14
1.532863e-15 1.189869e-14 6.932062e-16 6.348340e-15 1.542663e-16
6.731157e-16 7.120550e-16 4.334700e-15 1.084699e-15 1.070808e-15
2.925099e-15 3.738589e-15 5.434030e-15 1.298861e-15 2.530256e-15
3.985367e-15 1.189609e-15 7.935394e-15 1.185269e-15 2.443453e-15
4.344124e-15 1.008929e-15 6.979292e-15 6.054192e-15 1.614831e-15
9.630459e-16 3.269469e-15 3.103547e-15 2.155879e-15 5.732145e-15
5.387776e-15 1.722842e-15 1.251612e-15 3.612103e-15 4.188492e-15
9.973464e-15 3.859001e-15 9.233625e-16 2.180805e-15 1.006451e-15
2.549980e-15 2.034971e-15 5.788712e-16 3.379714e-15 5.006573e-15
5.411586e-15 9.163815e-15 3.222842e-15 1.097855e-14 1.361360e-15
4.953997e-15 1.096852e-15 4.976810e-15 1.441719e-15 1.865062e-16
1.118886e-14 1.017983e-15 7.208955e-15 6.082588e-16 4.363718e-15
7.445934e-15 2.327751e-15 4.973588e-15 2.900552e-16 4.531247e-16
1.134276e-14 3.847949e-16 2.387153e-15 5.422249e-15 2.347967e-15
4.965156e-15 7.164809e-15 5.651789e-15 1.487476e-15 3.922623e-15
7.785593e-15 5.733137e-15 8.522697e-15 1.061537e-16 1.282365e-15
1.199157e-15 3.691965e-15 6.381325e-15 7.855163e-15 8.869046e-16
6.961806e-15 5.868048e-16 2.009301e-15 8.587428e-15 4.440849e-15
4.387153e-15 8.299232e-15 4.977435e-15 8.686759e-16 7.514890e-16
1.499875e-15 5.840782e-16 2.478424e-15 5.287691e-16 1.173735e-15
9.164931e-15 3.741320e-15 7.205112e-15 4.324173e-16 1.060641e-15
9.711192e-15 3.380453e-16 3.862973e-15 1.843750e-15 8.857146e-15
6.454247e-15 3.693330e-15 2.817586e-15 5.082468e-15 7.446058e-15
4.590030e-15 1.159490e-16 5.348587e-15 5.711683e-15 1.105531e-15
4.771824e-15 9.569699e-16 8.263014e-15 3.274183e-15 7.410097e-15
8.179566e-16 1.810516e-15 2.817585e-15 2.840031e-15 1.154515e-16
6.557043e-15 6.347842e-15 2.475196e-16 5.686510e-15 1.319195e-15
4.474223e-16 6.819448e-15 6.020342e-15 1.228300e-15 1.376736e-15
4.218752e-15 3.248142e-15 3.737228e-15 2.684401e-15 2.038072e-15
6.995299e-16 8.690107e-15 4.299482e-15 1.232267e-15 8.488337e-16
1.209995e-14 5.153152e-15 1.935769e-16 1.042659e-14 6.961809e-15
6.587305e-16 4.019121e-16 3.091767e-15 7.692834e-15 8.874013e-15
1.067386e-14 1.407491e-14 6.830359e-15 6.284353e-15 5.872646e-15
3.398192e-15 1.739956e-15 7.488853e-16 2.740700e-15 2.918401e-15
5.861858e-16 1.204364e-15 3.359252e-15 1.217213e-14 1.029006e-14
3.750992e-15 8.405268e-16 3.349457e-15 5.940478e-15 1.729922e-16
5.666791e-15 3.619917e-15 5.939113e-15 2.772570e-15 9.393584e-16
7.227061e-15 9.924607e-15 1.362067e-14 2.492684e-15 2.594496e-15
5.345990e-16 3.411336e-15 4.448537e-15 2.420885e-15 5.301837e-15
6.298736e-15 4.046371e-16 2.118303e-15 2.876240e-15 4.526043e-15
2.098339e-15 1.239955e-15 5.617584e-17 4.353176e-15 1.572793e-15
5.412948e-16 1.530381e-15 7.164934e-15 8.771578e-15 1.100076e-15
3.713791e-15 9.144358e-16 3.049107e-15 2.533235e-15 1.705106e-15
6.516371e-15 5.545606e-16 2.448908e-15 3.132813e-15 6.561633e-15
3.480912e-16 3.738842e-15 8.352558e-15 3.943950e-15 5.095612e-15
4.799977e-15 5.379466e-15 6.875126e-15 2.003100e-15 5.097722e-15
1.754589e-15 1.077630e-15 5.629713e-15 2.236856e-15 1.918414e-16
4.767363e-15 1.070808e-15 3.094991e-15 1.798116e-15 7.203006e-15
1.804812e-15 8.196058e-15 2.216020e-16 2.651291e-15 5.360742e-15
1.941595e-15 3.177582e-15 7.303200e-15 4.386037e-15 5.668280e-15
8.231275e-15 8.919864e-16 8.076888e-15 4.273315e-15 2.218378e-15
2.350324e-15 9.877613e-15 1.615949e-15 5.727920e-16 3.657864e-15
2.144593e-15 2.940231e-15 3.255203e-16 6.820381e-17 1.592635e-15
7.710810e-16 1.030508e-15 8.176219e-15 4.812375e-15 1.077146e-14
5.221975e-15 7.138396e-15 2.597225e-15 7.742442e-15 2.067213e-15
8.611605e-15 1.348339e-15 9.145588e-16 7.733123e-16 3.147693e-15
4.650920e-15 1.251116e-15 3.773065e-15 4.371531e-15 3.432791e-15
7.461312e-15 3.320313e-15 5.895330e-16 2.933780e-15 1.008931e-15
1.390130e-15 2.104167e-15 3.683161e-15 3.964659e-15 4.203127e-15
4.835316e-15 4.762402e-15 2.148933e-15 1.603175e-15 2.185271e-15
6.801828e-16 2.321427e-15 5.894224e-15 6.877731e-15 9.179315e-15
2.416668e-15 1.890131e-15 6.458583e-15 6.713772e-16 2.428083e-16
4.869914e-15 1.908482e-15 9.416792e-15 1.132044e-14 1.050966e-15
6.360490e-15 1.796382e-15 2.286086e-15 5.448784e-15 5.608013e-15
6.629841e-15 5.542907e-15 5.178565e-16 2.029638e-15 1.866939e-15
6.185142e-15 6.207096e-15 4.002978e-15 1.526041e-15 3.301591e-15
5.788074e-15 1.453251e-15 2.057787e-15 2.795140e-15 1.920016e-15
4.960690e-15 4.819571e-15 1.200388e-16 3.828376e-15 6.920633e-15
5.310033e-16 3.236357e-15 3.655383e-15 2.247643e-15 1.048984e-15
1.942337e-15 1.194631e-14 4.705978e-15 4.297002e-15 1.994793e-15
5.984624e-16 7.300341e-16 1.605160e-15 7.159725e-15 8.892619e-16
2.530880e-15 3.591027e-15 4.582220e-15 5.076141e-15 3.457092e-15
5.741197e-15 2.989212e-15 2.287078e-15 8.985370e-15 6.101186e-17
5.233507e-15 2.226440e-15 5.693328e-15 5.609375e-15 5.433904e-15
5.402034e-15 3.038939e-15 1.164832e-14 1.775173e-15 6.148192e-15
5.212056e-16 1.096330e-14 1.035578e-14 1.254588e-15 3.790917e-16
2.016493e-15 1.094085e-14 3.448165e-15 5.385170e-15 2.372272e-16
3.588804e-16 6.117686e-15 8.279024e-15 3.921009e-15 2.219866e-15
3.711679e-15 9.621800e-16 1.227928e-15 5.031375e-15 2.571928e-15
3.981931e-16 2.016370e-15 1.259425e-14 6.290301e-15 5.946309e-15
8.879219e-15 4.307417e-15 7.514759e-15 1.801217e-15 3.676961e-15
4.709822e-15 3.074158e-15 7.632317e-15 6.693950e-15 3.773192e-15
6.741695e-15 3.682296e-15 2.900421e-15 6.787826e-15 2.898315e-15
7.934022e-16 2.506075e-15 4.223698e-16 1.705978e-15 5.868182e-15
9.208587e-15 1.134673e-15 5.132685e-16 1.410838e-15 5.164062e-15
Addendum: The Weibull Distribution
After further investigation, I've discovered that this distribution is most accurately fit by a Weibull distribution,
not
Rician:
Above is the distribution analysis for the provided data. Here's another dataset:
The skew at the top end is pretty common, but I feel that this is as good as I'm going to get for a pre-packaged solution.
And since I have these data as a function of dwell time, let's see how the Weibull parameters behave as a function of dwell time:
Parameter 1 (shape) has no dwell time dependence that I can see, parameter 2 (offset) is effectively zero anyway, and parameter 3 (scale) has the same dwell time dependence (even with a similar power law term, -0.681) as the median and MAD above.
Quote
Report
Cryo · External communityPost link
External answer — Cross Validated Stack Exchange
Author: Cryo
Original post: https://stats.stackexchange.com/a/667746
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
Your signal is strictly non-negative, as you said, and your measurements seem to be non-negative too. So I am not sure normally distributed noise and your model for the measured data is sound. The data you have given certainly does not have normal distribution, or even log-normal.
I would not aim to model noise adaptively that severely restricts the phenomena you can model.
Lets ignore the true/stray components of the signal and talk about
$y_{full}$
for now. Then you need to decide how does your observed signal depend on that. You mentioned Rice distribution, that may be reasonable if it makes sense for your process. I tried fitting Gamma distribution to the snippet of data you gave. It does not quite fit, but with some additional modifications it may do.
Lets say that we have some cumulative distribution function for your observed signal
$F$
and the matching density function
$f$
, both parametrized by your
$y_{full}$
:
$$
Prob\left[Y'<y' \,|\, Y=y_{full}\right]=\int_0^{y'} dz\, f_{Y'}\left(z;\,y_{full}\right)=F_{Y'}\left(y';\,y_{full}\right)
$$
This should capture your noise description. If you are getting your information from a sensor, and you are operating it within the correct operating range I would expect its response to signal to be linear. At least that's how most detectors work, so the expectation value of your detector should be
$$
E\left[Y' \,|\, Y=y_{full}\right]=c+m\cdot y_{full}
$$
Note that
$c$
here appears event with no stray signal. In light detectors this is sometimes called the 'dark count', i.e. detector always produces something, and if its output is non-negative, then it will never average to zero.
You can then try to link it to a specific distribution to build a regression model. The reason regression makes sense here is because your detector, presumably, is built to respond linearly to your signal. I will use
gamma
distribution as an example, though, as I say, it does not quite fit. If
$Y'\,|\,Y$
is Gamma-distributed
$$
Y'\,|\,Y \sim\Gamma\left(a, \lambda\right)
$$
Then
$$
E\left[Y' \,|\, Y=y_{full}\right]=\frac{\alpha}{\lambda}=c+m\cdot y_{full}
$$
This already suggest how you could parametrize your regression model, e.g.
$\lambda=const$
,
$\alpha=\lambda c+\lambda m\cdot y_{full}$
other parametrizations are possible. If you wanted to constrain this further, you could use second moment
$E\left[Y'^2 \,|\, Y=y_{full}\right]$
(if you can measure that). You will then have a full model for your sensor with few constants you need to fit.
Rise distribution can be treated the same way, though equations won't look as nice, but that's not a problem - solve it numerically if need be. You can also try compositions, e.g. when I tried fitting gamma distribution to your data either the right tail was too steep, or the left tail was too shallow. You could however say that
$Y'=\left(\lambda \cdot Z\right)^{\alpha}$
,
$Z\sim \Gamma\left(a\right)$
, where additional power
$\alpha$
can handle the tails. Or use
Generalized gamma distirbution
.
Basically, find the distribution that does fit your data, then build a regression model for its parameters using your knowledge about how the sensor works (is its response linear)? This will give you
$F_{Y'}$
and
$f_{Y'}$
.
Within the current formulation, your parameters/unknowns are
$y_{full}=y_{stray}+y_{signal}$
,
$c$
,
$m$
,
$a$
,
$\lambda$
. You could tread the model as Bayesian, feed it your data, and get the joint likelihood for all parameters. If you then marginalized
$y_{signal}$
you would get the posterior distribution for
$Y_{signal}|Y'$
Quote
Report
ohshitgorillas · External communityPost link
External answer — Cross Validated Stack Exchange
Author: ohshitgorillas
Original post: https://stats.stackexchange.com/a/667760
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
$\newcommand{\ysignal}{y_\mathrm{signal}}$
$\newcommand{\ystray}{y_\mathrm{stray}}$
I have developed a solution and would sincerely appreciate any feedback on its general soundness.
As noted above, I have strong evidence that these baseline measurements follow a slightly top-skewed
Weibull distribution
.
This follows a two-stage approach:
Stage 1: Calibration
Each system must run a similar baseline-dwell time calibration characterization to determine three parameters:
The Weibull shape parameter,
$\beta$
. This is hypothesized to be a system-wide constant, and I use the average
$\beta$
found across all dwell times.
The Weibull location/offset parameter,
$\theta$
. This is a small negative offset that prevents the model from imposing the unphysical constraint that the noise is always additive. This is also treated as a system-wide constant, determined from the average of the characterization data.
The power law exponent
$k_\lambda$
which describes how the Weibull scale
$\lambda$
changes as a function of dwell (data collection/integration) time.
For the system I've characterized, these parameters are approximately:
$\beta = 1.23$
$\theta = -4.86\times10^{-17}$
A
$k_\lambda = -0.681$
Stage 2: Baseline Correction
By fixing
$\beta$
and
$\theta$
as a constant, then for each analysis, we only need to solve for
$\lambda$
. This is a lot more feasible with N=12 to 20. In fact, I can use the method of moments to solve for
$\lambda$
directly by substituting the median of the baseline measurements
$B$
for a given analysis:
$$\lambda = \frac{(\text{median}(B) - \theta)}{(\ln{2})^{1/\beta}}$$
$\lambda$
is then corrected to the dwell time of the mass being baseline corrected using the power law function described above with
$k_\lambda$
:
$$\lambda_\text{corr} = \lambda\left(\frac{t_\text{dwell,sample gas}}{t_\text{dwell,baseline}}\right)^{k_\lambda}$$
This results in a fully defined Weibull distribution (
$\beta$
,
$\theta$
,
$\lambda_\text{corr}$
) describing the baseline for that specific measurement. This distribution can be used in two ways:
Static Correction
: A point-estimate correction is made by subtracting the mean of the fitted Weibull distribution:
$$\ystray + \epsilon = \theta + \lambda_\text{corr} \Gamma\left(1 + 1/\beta\right)$$
$$y = y' - \ystray - \epsilon = \left(\ysignal + \ystray + \epsilon\right) - \ystray - \epsilon = \ysignal$$
Probabilistic Correction
: The same as described above, but with a 3-parameter Weibull PDF in place of the Gaussian PDF for the likelihood function:
$$P(y'|\ysignal) = \frac{\beta}{\lambda_{corr}} \left( \frac{(y' - \ysignal) - \theta}{\lambda_{corr}} \right)^{\beta-1} \exp\left( -\left( \frac{(y' - \ysignal) - \theta}{\lambda_{corr}} \right)^{\beta} \right)$$
The prior probability function (
$1\times10^{-20} \leq \ysignal \leq 1\times10^{-4}$
A, in 100,000 node log-spaced uninformed prior) and the posterior:
$$P(\ysignal|y') = \frac{P(y'|\ysignal)P(\ysignal)}{P(y')}$$
...remain the same as above. The final baseline corrected value,
$y$
, is given by the mean of all possible values:
$$y = \int_0^\infty \ysignal P(\ysignal|y') d\ysignal$$
Quote
Report
cbeleites · External communityPost link
External answer — Cross Validated Stack Exchange
Author: cbeleites
Original post: https://stats.stackexchange.com/a/667777
License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/
Adaptation: HTML converted to plain text; contact email addresses removed.
Some considerations, but please bear in mind that what I write here stems from my experience with Raman data (where we basically count photons instead of ions):
Count data itself is often nicely
Poisson-distributed
.
This is important here since that would imply that the noise on the ion count has variance proportional to the expected ion count. I. e., your model would have more distinct noise terms:
$$y′=y_{signal}+y_{stray} + \varepsilon + \sigma_{signal} + \sigma_{stray}$$
with
$$\sigma^2_{signal} \propto y_{signal}$$
and
$$\sigma^2_{stray} \propto y_{stray}$$
(Poisson distribution has both mean and variance equal to its parameter
$\lambda$
, but current in mA is only [hopefully] linear to the ion count.)
measured data however has additional effects on it, e.g. from the detector, bias voltages, amplifier, and so on. Those may be summed together into the
$\varepsilon$
term.
Also here, more may be known about this distribution - you may want to chat to someone from the technical/develeopment crowd at the instrument manufacturer's*. (And maybe let us know here what the result is.)
Is there perhaps some justification for [Gaussian distribution]
Normal approximation for Poisson distributions with sufficiently large
$\lambda$
(no. of counts > say, 1000). There is also a correction that allows normal approximation for much smaller count numbers, see e.g. the Wiki page.
Central limit theorem: if sufficiently many sources of noise contribute (in practice!). That may be the approximation for the various sources of noise on the detector and electronics side (assuming the instrument is run with good settings: good settings should have no single source of noise dominate the result)
In addition, since for "good" measurement conditions we expect
$\varepsilon$
to be small and independent of the signal, and only few stray ions, you may be able to drop some terms from your model.
Of course, if you aim for the lowest possible lower limit of quantitation, you likey cannot immediately approximate all those terms away, but need careful consideration of what terms to add to your model in practice. (But then your model may give you rather immediate access to LLoQ)
* E.g. for one Raman instrument I was working with, both the dark current itself and its noise were dependent on the position on the camera - because there was some electronics nearby that heated up substantially and dark current and its noise are highly temperature dependent.
Had that instrument software been set up to automatically subtract the dark current, we had only seen the influence on the noise... Fortunately they didn't so the situation was quite obvious.
Quote
Report
Post Reply
Quoted from Forex.com.bd-Editorial External answer — Cross Validated Stack Exchange Author: cbeleites Source score (net votes, not local likes): 2 Original post: https://stats.stackexchange.com/a/667777 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. Some considerations, but please bear in mind that what I write here stems from my experience with Raman data (where we basically count photons instead of ions): Count data itself is often nicely Poisson-distributed . This is important here since that would imply that the noise on the ion count has variance proportional to the expected ion count. I. e., your model would have more distinct noise terms: $$y′=y_{signal}+y_{stray} + \varepsilon + \sigma_{signal} + \sigma_{stray}$$ with $$\sigma^2_{signal} \propto y_{signal}$$ and $$\sigma^2_{stray} \propto y_{stray}$$ (Poisson distribution has both mean and variance equal to its parameter $\lambda$ , but current in mA is only [hopefully] linear to the ion count.) measured data however has additional effects on it, e.g. from the detector, bias voltages, amplifier, and so on. Those may be summed together into the $\varepsilon$ term. Also here, more may be known about this distribution - you may want to chat to someone from the technical/develeopment crowd at the instrument manufacturer's*. (And maybe let us know here what the result is.) Is there perhaps some justification for [Gaussian distribution] Normal approximation for Poisson distributions with sufficiently large $\lambda$ (no. of counts > say, 1000). There is also a correction that allows normal approximation for much smaller count numbers, see e.g. the Wiki page. Central limit theorem: if sufficiently many sources of noise contribute (in practice!). That may be the approximation for the various sources of noise on the detector and electronics side (assuming the instrument is run with good settings: good settings should have no single source of noise dominate the result) In addition, since for "good" measurement conditions we expect $\varepsilon$ to be small and independent of the signal, and only few stray ions, you may be able to drop some terms from your model. Of course, if you aim for the lowest possible lower limit of quantitation, you likey cannot immediately approximate all those terms away, but need careful consideration of what terms to add to your model in practice. (But then your model may give you rather immediate access to LLoQ) * E.g. for one Raman instrument I was working with, both the dark current itself and its noise were dependent on the position on the camera - because there was some electronics nearby that heated up substantially and dark current and its noise are highly temperature dependent. Had that instrument software been set up to automatically subtract the dark current, we had only seen the influence on the noise... Fortunately they didn't so the situation was quite obvious.
Checking account access…