The Cache Hierarchy and What a Miss Costs
Earlier lessons introduced the cache line as the unit that moves between memory and the processor, and false sharing as what happens when two threads write to one line. The systems round goes one level further. It asks why the same algorithm runs several times slower on one data layout than on another, and it expects you to reason with numbers: how big each cache is, how long a miss takes, and how large the working set is.
The hierarchy in numbers
Exact sizes and timings vary by processor and generation, so read the figures below as orders of magnitude, not as a benchmark. The shape is the same on every current server processor.
| Level | Typical size | Time to read a line | Scope |
|---|---|---|---|
| L1 data cache | 32 KB to 48 KB | about 1 ns | One core |
| L2 cache | 256 KB to 2 MB | about 3 to 5 ns | One core |
| L3 cache | several MB to tens of MB | about 10 to 20 ns | Shared by many cores, often a whole socket |
| Main memory | gigabytes | about 60 to 100 ns | The whole machine |
The rest of this lesson is for subscribers
Unlock every lesson in Systems Programming for Trading, and every other premium course.
Subscribe to continueTest your knowledge
Keep reading Systems Programming for Trading
27 lessons in this course, and every other premium course, on one subscription.
- Every lesson in every course, with the worked examples and interactive simulators
- Graded questions on every lesson, with explanations for the wrong answers as well as the right one
- The trainers, timed assessments and brainteaser library that go with them