2026 · C++20 · public repository
orderbook
A C++ matching engine tested with reconstructible executions from ten books maintained during one NASDAQ TotalView-ITCH session.
- Finding
- The engine's visible front-of-queue order agreed for all 286,725 included executions.
- Scope
- One 268.7M-message TotalView-ITCH file from 2019-12-30, with books maintained for ten liquid symbols.
- Boundary
- This is a descriptive consistency test with documented exclusions, not exchange certification. Synthetic latency results are a separate local benchmark.
Browser demo, loaded on request
Press Run and the same C++ engine, compiled to WebAssembly, matches a simulated order flow (stochastic volatility, trends, fat-tailed news jumps) live in this frame. Open the full demo to work a parent order and measure implementation shortfall, or pit an Avellaneda-Stoikov-style market maker against momentum and mean-reversion bots.
Loads the WebAssembly build into this frame.
Comparison with the ITCH feed
The parser read all 268,744,780 messages in one NASDAQ TotalView-ITCH 5.0 file from 2019-12-30 while the engine maintained books for ten liquid symbols. The 21.2-second measurement, or 12.7M messages per second, therefore covers file parsing plus book updates for those ten symbols; it is not a claim that all 268.7M messages became matching-engine operations. Before each included regular-hours execution, the study checks whether the executed order is the same order the engine holds at the head of the best-price FIFO. The comparison agreed on 286,725 of 286,725 included executions.
Cross-path ‘C’ executions trade at non-display prices, and opening-process releases and price-slid redisplays carry original priority that public ITCH does not expose, so those cases are excluded under documented rules. Including executions at those unreconstructible levels produces a different 99.32% descriptive figure, so the headline names the included denominator explicitly.
One-day queue-position study
A phantom 100-share passive order joins the back of the best-bid visible queue every 15 seconds of market time. The simulator tracks the displayed orders ahead of it as executions and cancels arrive. It assumes the phantom order has no market impact and does not change other participants' behavior. The pooled study contains 15,590 overlapping trials across the ten symbols on one day, each with a 120-second horizon.
1-30
9.2 s
100-200
7.3 s
500-800
8.3 s
1.3-1.9k
14.5 s
1.9-3.2k
20.1 s
3.2-25.7k
22.0 s
Values in %, per depth bucket; the line under each bucket is the median time-to-fill.
Pooled fill rate was 85.1% in the thinnest queue-ahead bin and 64.8% in the deepest. Median time to first fill was 9.2 seconds and 22.0 seconds, respectively. Trade-through probability in the deepest bin was 17.1%; the minimum across the displayed bins was 11.7%. Because trials overlap and pool symbols and times of day, they are not independent and the pattern is descriptive rather than causal. Two event-order anomalies occurred among 311k executions.
Separate synthetic latency benchmark
A separate synthetic benchmark runs single-threaded on one Apple M-series core with a warm cache and -O3, using 1M orders over a 65,536-tick book. The exact chip and compiler version were not recorded, so these results are not cross-machine comparisons. An allocation audit measured zero engine heap allocations on the hot path after sized construction, subject to the caller-buffer limits documented in the repository.
rest
48 ns
Throughput
21 M ops/s
P99
167 ns
cancel
79 ns
Throughput
13 M ops/s
P99
333 ns
match
61 ns/fill
Throughput
16 M fills/s
P99
not reported
sweep
32 ns/fill
Throughput
32 M fills/s
P99
not reported
mixed
105 ns/op
Throughput
10 M ops/s
P99
not reported
| op | latency | throughput | p99 |
|---|---|---|---|
| rest | 48 ns | 21 M ops/s | 167 ns |
| cancel | 79 ns | 13 M ops/s | 333 ns |
| match | 61 ns/fill | 16 M fills/s | not reported |
| sweep | 32 ns/fill | 32 M fills/s | not reported |
| mixed | 105 ns/op | 10 M ops/s | not reported |
The two-level occupancy bitmap makes best-level refresh O(ticks/4096) in the worst case. The unpinned benchmark cannot separate algorithm latency from operating-system scheduling in the microsecond tail.
Correctness
Differential fuzzing sends the same stream of limit, market, IOC, FOK, modify, and cancel operations through the fast engine and a deliberately naive std::map reference. About 6% of draws use edge cases such as out-of-range prices, near-2³² quantities, and id collisions. The test compares fills after every operation and deep state (per-order FIFO at every populated level) every 16 operations. CI adds a run-varying seed and ASan/UBSan jobs. Internal review also mutation-tested selected test guards.