Floating Point and Undefined Behaviour

Two kinds of question catch strong C++ candidates out. One is about the number format: why a sum comes out different in another build, or why a sort crashed on data with a missing price. The other is about undefined behaviour: code that looks correct, passes its tests at -O0, and does something else at -O2. Both come from the gap between what the programmer means and what the standard promises.

What a double holds

A double is an IEEE 754 binary64 value: 1 sign bit, 11 exponent bits and 52 fraction bits. That gives 15 to 17 significant decimal digits and exact integers up to 253=9,007,199,254,740,9922^{53} = 9{,}007{,}199{,}254{,}740{,}992. Above that, not every integer can be represented, and 253+12^{53} + 1 rounds back to 2532^{53}.

Most decimal fractions have no exact binary form, so 0.1 + 0.2 == 0.3 is false. Precision also depends on magnitude, which makes addition non-associative:

The rest of this lesson is for subscribers

Unlock every lesson in Systems Programming for Trading, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Systems Programming for Trading

27 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them