Distributed Systems for the Design Round

Platform, data and research infrastructure teams at trading firms ask the design round in its distributed form. The prompts are about services spread over many machines: a store for tick data that researchers query, a service that holds risk limits for every desk, or a pipeline that delivers fills to the systems that book them. The vocabulary is the same as at any large technology company. What trading adds is a low tolerance for two particular failures: losing a fill, and doubling an order.

Design for the failures first

In a distributed system, machines crash and restart, networks split into parts that cannot reach each other, and messages are delayed, reordered and delivered more than once. One fact shapes most designs: a timeout cannot tell a slow machine from a dead one. A request with no reply may have failed, may have succeeded with the reply lost, or may still be in progress. Every design choice below is a way of living with that uncertainty.

The rest of this lesson is for subscribers

Unlock every lesson in Systems Programming for Trading, and every other premium course.

Subscribe to continue

Test your knowledge

Questions are only available to subscribers.

Keep reading Systems Programming for Trading

27 lessons in this course, and every other premium course, on one subscription.

  • Every lesson in every course, with the worked examples and interactive simulators
  • Graded questions on every lesson, with explanations for the wrong answers as well as the right one
  • The trainers, timed assessments and brainteaser library that go with them