Mutual Information

Mutual information
I(X;Y)=x,yP(x,y)logP(x,y)P(x)P(y)I(X;Y) = \sum_{x,y}P(x,y)\log\frac{P(x,y)}{P(x)P(y)}

Zero exactly when the two are independent, which is a stronger claim than zero correlation.

How much knowing YY reduces uncertainty about XX. Equivalently:

I(X;Y)=H(X)H(XY)I(X;Y) = H(X) - H(X\mid Y)

Entropy before minus entropy after. Symmetric, non-negative, and zero exactly when XX and YY are independent.

Why this beats correlation

Correlation detects only linear relationships. Mutual information detects any dependence.

The standard counterexample from covariance: let XX be symmetric around zero and Y=X2Y = X^2. Then ρ=0\rho = 0 while YY is completely determined by XX. Mutual information is large, correctly reporting total dependence.

Note also the difference in what zero means. Zero correlation permits strong dependence; zero mutual information means genuine independence. It is the stronger statement.

Key takeaway

Zero correlation is weak evidence; zero mutual information is a real claim of independence. This is why mutual information catches the quadratic relationships that option-like exposures produce.

The rest of this lesson is for subscribers

Unlock every lesson in Advanced Topics in Probability and Statistics, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Advanced Topics in Probability and Statistics

35 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them