Posterior Predictive Distributions

You have a posterior over parameters. You want the distribution of the next observation, which is usually the thing you actually care about.

The posterior predictive
p(xnewD)=p(xnewθ)p(θD)dθp(x_{\text{new}} \mid D) = \int p(x_{\text{new}} \mid \theta)\,p(\theta \mid D)\,d\theta

Average the likelihood over the posterior, so the prediction carries the parameter uncertainty instead of discarding it.

Average the predictive distribution over every parameter value, weighted by how plausible the parameter is.

Why not just plug in the estimate

The tempting shortcut is to use p(xnewθ^)p(x_{\text{new}} \mid \hat{\theta}) with a point estimate. It is wrong, and wrong in a consistent direction.

Plugging in pretends you know the parameter exactly. The posterior predictive accounts for two sources of uncertainty:

Inherent randomness in the next observation, which the plug-in captures.

Parameter uncertainty, which it does not.

The rest of this lesson is for subscribers

Unlock every lesson in Advanced Topics in Probability and Statistics, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Advanced Topics in Probability and Statistics

35 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them