The Cache Hierarchy and What a Miss Costs

Earlier lessons introduced the cache line as the unit that moves between memory and the processor, and false sharing as what happens when two threads write to one line. The systems round goes one level further. It asks why the same algorithm runs several times slower on one data layout than on another, and it expects you to reason with numbers: how big each cache is, how long a miss takes, and how large the working set is.

The hierarchy in numbers

Exact sizes and timings vary by processor and generation, so read the figures below as orders of magnitude, not as a benchmark. The shape is the same on every current server processor.

Level Typical size Time to read a line Scope
L1 data cache 32 KB to 48 KB about 1 ns One core
L2 cache 256 KB to 2 MB about 3 to 5 ns One core
L3 cache several MB to tens of MB about 10 to 20 ns Shared by many cores, often a whole socket
Main memory gigabytes about 60 to 100 ns The whole machine

The rest of this lesson is for subscribers

Unlock every lesson in Systems Programming for Trading, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Systems Programming for Trading

27 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them