A stable learner can stabilize on the wrong action
Ardaiz, Budati, and Habibnia train deep double Q-learning (DDQL) agents to sell a cryptocurrency parent order. Their mixture-of-experts (MoE) variants partition training across specialist networks and eliminate a degenerate policy that sometimes appears in the dense baseline: submit nothing until the episode ends.
Implementation shortfall is the loss against the price observed when the order arrived, expressed here in basis points (one basis point is 0.01%). Lower is better. CVaR95 averages the worst 5% of losses, so it asks a different question from variation across training seeds.
The headline average is almost a tie
| Policy | Mean IS (bps) | SD of run means | CVaR95 (bps) | Collapsed |
|---|---|---|---|---|
| Immediate liquidation | 0.21 | — | 0.8 | — |
| TWAP | 0.39 | — | 74 | — |
| DDQL | 1.08 | 0.5 | 110.8 | 11 / 100 |
| MoE K=2 | 1.06 | 0.45 | 116.2 | 3 / 100 |
| MoE K=4 | 1.15 | 0.52 | 120.6 | 0 / 100 |
| MoE K=8 | 1.05 | 0.37 | 122.2 | 0 / 100 |
| DDQL cap=8 | 0.95 | 0.5 | 104.1 | 4 / 100 |
None of the MoE mean differences versus DDQL is statistically distinguishable from zero. For K=8 the reported difference is −0.03 basis points with a 95% confidence interval from −0.15 to +0.09. The mechanism claim is therefore about failed-run frequency and dispersion, not better average execution.
Two stability metrics move in opposite directions
The router cannot observe the regime it is asked to route
Each episode begins at an independently sampled time. But the causal router chooses an expert from features of the previous episode. Under independent starts, yesterday's selected window contains no designed information about the next window. The experts can still specialize because routing partitions replay and optimization, but the experiment does not establish market-regime conditioning.
Data dilution grows with the number of experts. At K=8, 2,500 training episodes average only 312.5 episodes per expert. That reduction is consistent with the worsening CVaR tail, although it does not by itself prove the cause.
A tiny schedule parameter explains the collapse
The reported exploration probability is multiplied by 0.9995 every 100 actions. Across 25,000 transitions that is only 250 updates, leaving ε at 0.882. The agent remains highly random near the end of training. A different annealing schedule removes all 12 dense-DDQL collapses in the paired 100-seed experiment.
def exact_mcnemar(baseline_only, modified_only):
discordant = baseline_only + modified_only
if discordant == 0:
return 1.0
lower = min(baseline_only, modified_only)
tail = sum(comb(discordant, k) * 0.5**discordant
for k in range(lower + 1))
return min(1.0, 2.0 * tail)The exact paired McNemar p-value for 12 baseline-only collapses and zero annealing-only collapses is 4.883e-4. Pairing matters: it tests which specification collapses on the same seeds rather than treating two rates as unrelated.
The fixes interact instead of adding up
| DDQL specification | Runs | Mean IS (bps) | Collapsed |
|---|---|---|---|
| As reported | 100 | 1.13 | 12 / 100 |
| + reward | 30 | 0.05 | 0 / 30 |
| + annealing | 100 | 0.28 | 0 / 100 |
| + buffer | 30 | 1.19 | 5 / 30 |
| reward + anneal | 30 | 1.08 | 19 / 30 |
| all three | 100 | 0.93 | 48 / 100 |
The easiest policy wins because the market is frictionless
Immediate liquidation scores 0.21 basis points, better than every learned policy. Replay is based on a mean-aggregated five-minute book, there is no market impact, and liquidity beyond the visible depth is effectively cheap. A floor intended to limit replay depth never binds, even under an extreme parameter value tested by the authors.
The never-trade failure is also partly written into the reward. Waiting earns zero step reward because slippage is measured against the current midprice, then unfinished inventory is handled at the terminal step. The resulting policy is bad under implementation shortfall, but it is locally easy for the learner to discover.
What to probe next
Run a balanced factorial design for reward, annealing, and replay; route only on information available for the current independently sampled episode; equalize data per expert; and evaluate on non-overlapping days. Then add fees, depth-aware impact, latency, participation constraints, and a terminal liquidation rule that matches the economic objective.
DDQL and its expert heads are neural value approximators. See the interactive neural-network walkthrough for the from-scratch machinery beneath those value estimates.
Reproduction notes
Tables, confidence intervals, p-values described as reported, and main outcome values come from the paper. Collapse rates, the exact McNemar result, Holm arithmetic, per-expert exposure, and epsilon curves are deterministic calculations from pure core functions. The faster epsilon curve is explicitly illustrative. No ML library, external image, or random draw is used.
References
- Ardaiz, A.; Budati, V.; Habibnia, A. (2026). Mixture-of-Experts for Cryptocurrency Order Execution: Training Stability, Tail Risk, and Failure Modes. ACM ICAIF 2026 / arXiv:2610.03369
- Dietterich, T. G. (1998). Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms. Neural Computation 10(7):1895–1923