Machine Learning Technique to Estimate Probability Distribution

Machine Learning Technique to Estimate Probability Distribution

Manage alerts

Loading saved threads...

user120010 · External communityPost link
External question — Cross Validated Stack Exchange Author: user120010 Original post: https://stats.stackexchange.com/questions/245148 License: CC BY-SA 3.0 — https://creativecommons.org/licenses/by-sa/3.0/ Adaptation: HTML converted to plain text; contact email addresses removed. After training a model on the resulting value given some set of features, I would like the model to estimate the probability distribution of the resulting value. I am not interested in classification, I would like the target value to be continuous. For example, I might have S&P day close, S&P day open, trade volume, day high, and day low as features and a target being tomorrows day open. I would like the model to learn the probability distribution of tomorrows day open given these features. Be it through representing the parameters of the distribution, or being able to evaluate the probability of a feature set resulting in a specific target value. Bonus points if this technique can be applied to a multi-target system. Edit: So I found a few methods, as below mentioned, there is Gaussian Process Regression. I found a very well documented implementation in scikitlearn I also found Students-T Process Regression which is useful for me because the distribution I was fitting had a slight skew and a fat tail, though I couldn't find any implementations of it anywhere. The most useful solution I have found so far is Mixture Density Networks which allows you to take a set of features and create a mixture distribution which isn't limited to normal. I also found an RNN Based mixed density network- http://blog.otoro.net/2015/12/12/handwriting-generation-demo-in-tensorflow/
Quote
Report
scherm · External communityPost link
External answer — Cross Validated Stack Exchange Author: scherm Original post: https://stats.stackexchange.com/a/245220 License: CC BY-SA 3.0 — https://creativecommons.org/licenses/by-sa/3.0/ Adaptation: HTML converted to plain text; contact email addresses removed. The example problem you gave sounds like a regression problem, since tomorrow's day open is a continuous value you'd like to predict. A typical approach would be to assume a distribution for your target value and estimate its parameters given some input. There are many ways to do this, but Gaussian process regression might be a good fit. Gaussian processes produce an estimate of the mean and variance of the target variable, which together specify the probability distribution. They can be quite flexible for many applications. Here is a good tutorial: https://www.robots.ox.ac.uk/~mebden/reports/GPtutorial.pdf And "thee book" with a link to some GP software: http://www.gaussianprocess.org/gpml/
Quote
Report
calmcc · External communityPost link
External answer — Cross Validated Stack Exchange Author: calmcc Original post: https://stats.stackexchange.com/a/655103 License: CC BY-SA 4.0 — https://creativecommons.org/licenses/by-sa/4.0/ Adaptation: HTML converted to plain text; contact email addresses removed. The classic method for modeling the conditional distribution, if $y$ is continuous, is quantile regression. The essential point is that quantile regression quantizes the conditional distribution into bins, then models the probability that your data falls into each particular bin, given covariates. Instead of quantile regression, you can also use the recently-developed BaltoBot ( bal anced t ree o f bo osted t rees), which comes with open-source code and a pip-installable package. Essentially, this method hierarchically partitions the space of the output variable, with XGBoost binary classifiers trained to find the right path from root to leaf. For example, for a problem with one-dimensional input, the results look like this: Quantile regression quantizes the distribution into quantiles which are independently modeled. This causes problems like cross-over, where a lower quantile is sometimes predicted to exceed the higher quantile. Statistical strength gets worse as you increase the resolution by modeling more quantiles. In contrast, BaltoBot hierarchically models the distribution, modeling the ordered relationships among bins, so it can simultaneously model low-resolution aspects of the distribution very well, while also modeling high-resolution aspects with less accuracy. (Note: I'm the developer of this method.)
Quote
Report

Post Reply

Quoted from Forex.com.bd-Editorial External question — Cross Validated Stack Exchange Author: user120010 Source score (net votes, not local likes): 2 Original post: https://stats.stackexchange.com/questions/245148 License: CC BY-SA 3.0 — https://creativecommons.org/licenses/by-sa/3.0/ Adaptation: HTML converted to plain text; contact email addresses removed. After training a model on the resulting value given some set of features, I would like the model to estimate the probability distribution of the resulting value. I am not interested in classification, I would like the target value to be continuous. For example, I might have S&P day close, S&P day open, trade volume, day high, and day low as features and a target being tomorrows day open. I would like the model to learn the probability distribution of tomorrows day open given these features. Be it through representing the parameters of the distribution, or being able to evaluate the probability of a feature set resulting in a specific target value. Bonus points if this technique can be applied to a multi-target system. Edit: So I found a few methods, as below mentioned, there is Gaussian Process Regression. I found a very well documented implementation in scikitlearn I also found Students-T Process Regression which is useful for me because the distribution I was fitting had a slight skew and a fat tail, though I couldn't find any implementations of it anywhere. The most useful solution I have found so far is Mixture Density Networks which allows you to take a set of features and create a mixture distribution which isn't limited to normal. I also found an RNN Based mixed density network- http://blog.otoro.net/2015/12/12/handwriting-generation-demo-in-tensorflow/

Cancel quote

Checking account access…