# From Market Data to Order Book

Szymon Kopyciński · 24 September 2026

How do we ensure that the order book we have in memory matches reality?

## Introduction

When you open a trading terminal and look at a limit order book (LOB) ladder, it looks static. You may see something like:

```
BID                     ASK

99.98   500             100.02   300
99.97   700             100.03   450
99.96  1200             100.04   800
```

The natural assumption is that somewhere, an exchange is simply storing this table and sending it to everyone. In practice, that is commonly not how a trading or market data system sees the market.

Instead, exchanges commonly distribute a stream of events describing how the market _changes_ over time. They leave the job of reconstructing the current state of the orderbook to you.

This article aims to explain how orderbooks are less like tables, and more like state machines. 

## The LOB is always moving

At any moment, a LOB represents the outstanding interest to buy and sell an instrument. At the top of the book, you have the best bid and ask (BBO). 

Suppose the current state is:

```
BID                     ASK

99.98   500             100.02   300
99.97   700             100.03   450
99.96  1200             100.04   800
```

The best bid is $99.98$, the best ask is $100.02$, therefore the spread is $0.04$. 

The mid-price is:

$$
\frac{99.98+100.02}{2}=100
$$

That table above is only the current state, whereas a live market is constantly changing. Participants submit new orders, cancel existing orders, amend orders, and trade against resting liquidity.

Rather than repeatedly transmitting the entire book every time something changes, an exchange can send incremental updates.

For example:

```
1001 ADD BUY 99.99 200
1002 ADD SELL 100.01 150
1003 CANCEL BUY 99.97 300
```

Each message describes a change, and your system applies those changes onto its local copy of the market, which is the only thing your strategy sees.

## Snapshots and deltas

A snapshot is a complete view of the book at one moment in time, like the table above. A delta describes only what changed.

Delta semantics vary by venue. Some venues use absolute updates, some use incremental updates. This article uses incremental, signed changes to a level, where each message says how much to add or remove, not what the level should end up as:

```
1001 ADD    BUY  99.99 200
1002 ADD    SELL 100.01 150
1003 CANCEL BUY  99.97 300
1004 TRADE SELL 100.01 100
```

(Exchange semantics differ, this is a demonstrative example.)

In English, 1003 means:

> Take whatever you currently have at 99.97 on the bid, and remove 300 from it

and 1004 means 100 units traded against the resting ask at 100.01, so that quantity is consumed and the level drops by 100. If a level's quantity reaches zero, the level disappears.

This is cheaper than transmitting the entire book thousands of times a second, but it has a consequence: every delta is only meaningful relative to the state before it. Get the base wrong, or lose one message, and every later message inherits the error.

The general pattern is:

```
snapshot -> delta -> delta -> ... -> current state
```

Each delta is interpreted relative to the last known state, ordered by sequence number.

## A message is missing

The exchange sends the four messages above, but your system receives:

```
1001 ADD    BUY  99.99 200
1002 ADD    SELL 100.01 150
1004 TRADE  SELL 100.01 100
```

Message 1003 is missing. What happens? Probably nothing obvious. Your program won't crash, the CPU won't spike, and the book will look plausible. But you still show 700 at 99.97 when the real quantity is 400, and because deltas are incremental, that error is permanent. Nothing later will correct it, and every subsequent change to that level stacks on top of the wrong number.

This failure is dangerous because it's silent, and doesn't self heal.

### Sequence numbers

This is why market data protocols often (and should) include sequence numbers. The sequence numbers should progress as expected. If you skip from 1002 to 1004, you immediately know 1003 is missing.

Next, let's talk about actual recovery.

### Recovering from a gap

Once you know your local state is wrong, continuing to progress blindly is unsafe.

A recovery flow might look like:

1. sequence gap detected
2. mark local book as invalid
3. request or receive fresh snapshot
4. rebuild state from the snapshot, then apply only deltas newer than that snapshot
5. resume normal processing

But the exact mechanism varies by venue. Some systems distribute separate recovery feeds, some publish A and B market data feeds for reconciliation, some allow clients to request snapshots, and some might even provide replay mechanisms for missing messages.

## Duplicate and out-of-order messages

Missing messages aren't the only problem feed handlers and trading systems need to worry about. You may also encounter duplicates:

```
1001
1002
1002
1003
```

If message 1002 is not idempotent, applying it twice will likely corrupt your state.

Messages can also arrive out of order:

```
1001
1003
1002
1004
```

Whether that is possible depends on the transport and protocol. On UDP/multicast feeds, messages can arrive out of order or be lost, on a single TCP connection, ordering is guaranteed by protocol.

A robust feed handler needs clearly defined behaviour for each case. 

The primary challenge in market data engineering is less about simply parsing and applying prices, and more about maintaining trustworthy state which downstream consumers know they can rely on.

## L2 vs L3 data

L2 represents each price level as one aggregate quantity. With only L2, you can't see the individual orders which make up the price level.

A L3 feed can expose individual order-level events. Instead of:

```
ADD BUY 99.98 500
```

we might see:

```
ADD order_649d7054 BUY 99.98 200
ADD order_fdf483b8 BUY 99.98 300
```

Now the system must handle individual order state, otherwise it falls back to being a L2 view.

If `order_fdf483b8` is cancelled:

```
DELETE order_fdf483b8
```

the remaining quantity at 99.98 becomes 200.

L3 makes it possible to reason about things such as queue position, but it also becomes a larger complexity surface, as there is more state to maintain.

## Trades are not always enough

A common mistake is to assume that if you know every trade, you can reconstruct the LOB at any point in time.

You usually cannot, because trades tell you what actually matched and executed. They don't tell you about:

- newly submitted resting orders
- cancelled orders
- amended quantities
- orders that never traded
- changes deeper in the book

## Feed handler guarantees

A market data feed handler should not necessarily be geared towards processing messages as quickly as possible, but rather guaranteeing a set of invariants, such as:

- sequence numbers are contiguous
- quantities are never negative
- invalid deletes are rejected or flagged
- the best bid does not cross the book (exceed the best ask), normally
- timestamps are non-decreasing (do not move backwards)
- the book is marked as invalid or stale during recovery
- downstream consumers know whether state is trustworthy

## Why this matters to a strategy

Suppose a strategy computes simple order-book imbalance:

$$
I =
\frac{Q_b - Q_a}
{Q_b + Q_a}
$$

where:

$Q_b$ is quantity at the best bid;

$Q_a$ is quantity at the best ask.

Imagine the real market is:

```
BID                     ASK

99.99   900             100.01  100
```

Then:

$$
I = \frac{900-100}{900+100} = 0.8
$$

The strategy observes strong bid-side imbalance.

But suppose you missed a message cancelling 800 units from the bid.

The real state is actually:

```
BID                     ASK

99.99   100             100.01  100
```

So:

$$
I = 0
$$

The model itself could be completely correct. The implementation could be fast. None of that matters if the input state is wrong.

This is one of the recurring themes in quantitative trading: correctness problems in one layer can quietly invalidate everything downstream of it. 

## What real systems add

A production feed handler may also need to deal with:

- binary protocol decoding
- UDP multicast
- multiple channels
- packet loss
- redundant feeds
- timestamps
- instrument definitions
- trading-status messages
- auctions
- crossed or locked markets
- snapshot synchronisation
- replay feeds
- network failover
- memory allocation
- CPU affinity
- NIC tuning and kernel bypass
- logging and metrics
- downstream consumer fan-out
- extremely high message rates

But the core objective stays the same, to maintain accurate, timely, and recoverable representation of the market.

## Your turn

Implement a simple L2 LOB yourself, a couple of hints:

- For the simplified version, use a hashmap/dictionary to represent prices and quantities at levels (note this will be unsorted, so consider how you'll track the top of the book or move to a sorted data structure).
- Come up with a mechanism to apply deltas.
- Add helper functions which return metrics like BBO, spread, bid/ask depth, and microprice.
- Think about its algorithmic complexity: how fast are updates versus reading the top of the book?

If that was easy, try adding some L3 capabilities.


Part of [Quant Foundations](/blog/quant-foundations-2627) · Previous: [Pricing one call option two ways](/blog/quant-foundations-2627/pricing-one-call-option-two-ways) · Next: [An Intro to Market Microstructure](/blog/quant-foundations-2627/how-to-trade-orderbooks)
