Level two shows depth, not your place in line
A limit order book lists bids to buy and offers to sell. Level-two data (L2) usually show the total quantity at each price. They do not show the ordered list of individual orders inside that total. Under price–time priority, often called FIFO, that hidden order matters: older orders trade before newer ones at the same price.
Danait, Zamora, and Boier hold the visible book path and reconciled trades fixed, then change only where cancellations occur inside the unseen queue. Front cancellation removes older quantity ahead of a tagged order; back cancellation removes newer quantity behind it; a seeded random rule samples live background quantity. All three histories reproduce the same L2 totals.
One cancellation creates three queue positions
Consider a 100-share virtual order submitted behind 300 shares, with 200 shares already behind it. The deterministic toy below applies the same five aggregate events under three cancellation rules. Additions always join behind; market orders consume the visible FIFO queue. Random allocation uses the repository's seeded RNG with seed 260913597.
def step_marker(state, event, cancel_ahead=0):
ahead, remaining, filled = state
kind, quantity = event
if kind == 'cancel':
ahead = max(ahead - cancel_ahead, 0)
elif kind == 'market':
new_fill = min(max(quantity - ahead, 0), remaining)
ahead = max(ahead - quantity, 0)
remaining -= new_fill
filled += new_fill
return ahead, remaining, filledThe site executes the fuller TypeScript compiler in core/; the Python, TypeScript, and C++ snippets are line-for-line translations of the marker update at the heart of that compiler.
| Step | Common event | Front: ahead | Random: ahead | Back: ahead |
|---|---|---|---|---|
| 0 | submit marker | 300 | 300 | 300 |
| 1 | cancel 100 | 200 | 243 | 300 |
| 2 | add 100 | 200 | 243 | 300 |
| 3 | market 250 | 0 | 0 | 50 |
| 4 | cancel 100 | 0 | 0 | 50 |
| 5 | market 80 | 0 | 0 | 0 |
Reconstruct the path before judging a policy
For each L2 transition the paper attributes market removal first, then classifies the residual as either an addition or a cancellation. If a level falls from 500 to 460 shares while 30 shares traded, the residual is −10: our core function returns 10 shares of cancellation and 0 shares of addition. Each reconstructed stream must project back to every observed aggregate snapshot.
A virtual marker records quantity ahead and remaining child size without changing the historical book. When a market order of size m arrives, only the part beyond quantity ahead can fill the marker. This is a price-taking counterfactual: it assumes the test order does not alter later market events.
The same book moves completion by eight points
The held-out study contains 1,080 five-minute episodes on each of two Tokyo Stock Exchange instruments, drawn from the same 18 July 2025 dates. The aggressive benchmark is invariant because it consumes a copied book immediately. Passive-at-touch is sensitive to the queue history it cannot observe.
| Instrument | Metric | Front | Random | Back | Front − back |
|---|---|---|---|---|---|
| 1301.T (100 shares) | Passive completion | 18.15% | 15.29% | 10.14% | 8.01 pp |
| 1301.T (100 shares) | Passive IS (bps) ↓ | 8.008 | 8.418 | 9.018 | -1.010 bps |
| 1301.T (100 shares) | Aggressive IS (bps) ↓ | 10.010 | 10.010 | 10.010 | 0.000 bps |
| 7911.T (200 shares) | Passive completion | 97.20% | 95.52% | 89.81% | 7.39 pp |
| 7911.T (200 shares) | Passive IS (bps) ↓ | 0.765 | 0.926 | 1.149 | -0.384 bps |
| 7911.T (200 shares) | Aggressive IS (bps) ↓ | 2.400 | 2.400 | 2.400 | 0.000 bps |
The endpoints are not bounds on reality
The empirical scope is also deliberately narrow: two securities, seven months of data, and held-out evidence from 18 dates. Parent sizes differ; the simulator assumes zero latency, fees, rebates, and market impact; and terminal inventory trades against a copied book. The result is a powerful identification warning, not a guaranteed live-trading edge.
Even inside the declared scenarios, a useful response is minimax selection: choose the policy with the smallest worst-case cost. On the paper's 1301.T mean shortfalls, the core calculation selects passivewith worst-case loss 9.018 bps. That does not prove global robustness; it says only that this choice survives the scenarios actually compiled.
What to probe next
The natural next experiment expands the compiler before expanding the model: vary latent order partitions, cancellation-position laws, latency, parent size, and price changes; then train or evaluate every execution policy across matched realizations of the same aggregate path. A learned fill predictor can still help, but it should report sensitivity to the missing queue state. The site's random-forest page is a useful place to see how flexible predictors partition observed features; this paper explains why no partition of L2 features alone restores unobserved order identities.
References
- Riya Danait, Yuliana Zamora, and Ioana Boier (2026). Same Book, Different Fills: Partial Identification of FIFO Execution from Aggregate Order Books. arXiv preprint
- André F. Perold (1988). The Implementation Shortfall: Paper versus Reality. The Journal of Portfolio Management 14(3)
- M. Dixon (2018). A high-frequency trade execution model for supervised learning. High Frequency 1(1)