Bayesian Computation

Outside conjugate families, the posterior normalising constant

The evidence integral
P(D)=P(Dθ)P(θ)dθP(D) = \int P(D \mid \theta)P(\theta)\,d\theta

The denominator that makes the posterior a distribution, and the one part that is usually intractable.

is an integral over the whole parameter space, and in more than a few dimensions it is not computable analytically or by quadrature.

The resolution is to stop trying to compute the posterior and instead draw samples from it. With enough samples you can approximate any quantity you want: means, credible intervals, tail probabilities.

MCMC: the central trick

Markov chain Monte Carlo constructs a Markov chain whose stationary distribution is exactly the posterior. Run it long enough and the states it visits are draws from the posterior, even though you never computed the posterior.

The reason it works: the acceptance rule only ever needs ratios of posterior densities, and the intractable P(D)P(D) cancels in a ratio.

The rest of this lesson is for subscribers

Unlock every lesson in Fundamentals of Probability and Statistics, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Fundamentals of Probability and Statistics

41 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them