Measuring Latency Without Fooling Yourself
Every performance claim in a systems round invites the same follow-up: how would you know? Saying "I would profile it" does not answer that. A good answer names the tool, the number it reports, and the ways that number can mislead you. This lesson covers the three steps interviewers expect in that answer, and the mistakes that make a slow system look fast.
Start with counters
On Linux, perf stat runs a program and reports the processor's hardware performance counters for the whole run.
perf stat -e cycles,instructions,cache-misses,branch-misses,page-faults ./replay capture.bin
Each counter points at a different cause. Instructions per cycle well below 1 means the processor spends much of its time waiting, usually for memory. A high branch miss rate points at unpredictable branches. Many page faults during steady state mean memory is being touched for the first time, or freed and allocated again. Comparing counters before and after a change is how you confirm a cache or layout fix, rather than trusting the timing alone.
The rest of this lesson is for subscribers
Unlock every lesson in Systems Programming for Trading, and every other premium course.
Subscribe to continueTest your knowledge
Keep reading Systems Programming for Trading
27 lessons in this course, and every other premium course, on one subscription.
- Every lesson in every course, with the worked examples and interactive simulators
- Graded questions on every lesson, with explanations for the wrong answers as well as the right one
- The trainers, timed assessments and brainteaser library that go with them