← Blog/blog/moe-execution-specification-collapse

The stable trading model that learned to never trade

01

A stable learner can stabilize on the wrong action

Ardaiz, Budati, and Habibnia train deep double Q-learning (DDQL) agents to sell a cryptocurrency parent order. Their mixture-of-experts (MoE) variants partition training across specialist networks and eliminate a degenerate policy that sometimes appears in the dense baseline: submit nothing until the episode ends.

Implementation shortfall is the loss against the price observed when the order arrived, expressed here in basis points (one basis point is 0.01%). Lower is better. CVaR95 averages the worst 5% of losses, so it asks a different question from variation across training seeds.

02

The headline average is almost a tie

PolicyMean IS (bps)SD of run meansCVaR95 (bps)Collapsed
Immediate liquidation0.21—0.8—
TWAP0.39—74—
DDQL1.080.5110.811 / 100
MoE K=21.060.45116.23 / 100
MoE K=41.150.52120.60 / 100
MoE K=81.050.37122.20 / 100
DDQL cap=80.950.5104.14 / 100
Paper-reported 100-run results. SD measures variation in each run's mean; CVaR95 pools severe episode losses within a policy.
Paper-reported mean implementation shortfall, transformed by pure core-fed chart data. Immediate liquidation and TWAP are deterministic baselines; lower is better.

None of the MoE mean differences versus DDQL is statistically distinguishable from zero. For K=8 the reported difference is −0.03 basis points with a 95% confidence interval from −0.15 to +0.09. The mechanism claim is therefore about failed-run frequency and dispersion, not better average execution.

03

Two stability metrics move in opposite directions

SD of run means
Paper-reported across-run dispersion for DDQL, MoE K=2, MoE K=4, and MoE K=8 (listed order). Lower means training seeds agree more closely.
within-policy CVaR95
Paper-reported within-policy CVaR95 for the same four policies and order. Higher means worse severe execution losses.
04

The router cannot observe the regime it is asked to route

Each episode begins at an independently sampled time. But the causal router chooses an expert from features of the previous episode. Under independent starts, yesterday's selected window contains no designed information about the next window. The experts can still specialize because routing partitions replay and optimization, but the experiment does not establish market-regime conditioning.

Data dilution grows with the number of experts. At K=8, 2,500 training episodes average only 312.5 episodes per expert. That reduction is consistent with the worsening CVaR tail, although it does not by itself prove the cause.

05

A tiny schedule parameter explains the collapse

The reported exploration probability is multiplied by 0.9995 every 100 actions. Across 25,000 transitions that is only 250 updates, leaving ε at 0.882. The agent remains highly random near the end of training. A different annealing schedule removes all 12 dense-DDQL collapses in the paired 100-seed experiment.

reported decayillustrative faster decay
Core-computed epsilon paths at transitions 0, 5k, …, 25k. Red is the reported 0.9995-per-100 schedule; green is a clearly labeled sensitivity using 0.99 per 100 with a 0.05 floor, not a paper result.
def exact_mcnemar(baseline_only, modified_only):
    discordant = baseline_only + modified_only
    if discordant == 0:
        return 1.0
    lower = min(baseline_only, modified_only)
    tail = sum(comb(discordant, k) * 0.5**discordant
               for k in range(lower + 1))
    return min(1.0, 2.0 * tail)

The exact paired McNemar p-value for 12 baseline-only collapses and zero annealing-only collapses is 4.883e-4. Pairing matters: it tests which specification collapses on the same seeds rather than treating two rates as unrelated.

06

The fixes interact instead of adding up

DDQL specificationRunsMean IS (bps)Collapsed
As reported1001.1312 / 100
+ reward300.050 / 30
+ annealing1000.280 / 100
+ buffer301.195 / 30
reward + anneal301.0819 / 30
all three1000.9348 / 100
Paper-reported CPU specification decomposition. Run counts differ, so the collapse chart converts counts to rates without pretending the uncertainty is equal.
Core-computed DDQL collapse percentages from the paper-reported counts. Annealing alone removes collapse, while reward plus annealing restores it in 19 of 30 runs.
07

The easiest policy wins because the market is frictionless

Immediate liquidation scores 0.21 basis points, better than every learned policy. Replay is based on a mean-aggregated five-minute book, there is no market impact, and liquidity beyond the visible depth is effectively cheap. A floor intended to limit replay depth never binds, even under an extreme parameter value tested by the authors.

The never-trade failure is also partly written into the reward. Waiting earns zero step reward because slippage is measured against the current midprice, then unfinished inventory is handled at the terminal step. The resulting policy is bad under implementation shortfall, but it is locally easy for the learner to discover.

08

What to probe next

Run a balanced factorial design for reward, annealing, and replay; route only on information available for the current independently sampled episode; equalize data per expert; and evaluate on non-overlapping days. Then add fees, depth-aware impact, latency, participation constraints, and a terminal liquidation rule that matches the economic objective.

DDQL and its expert heads are neural value approximators. See the interactive neural-network walkthrough for the from-scratch machinery beneath those value estimates.

09

Reproduction notes

Tables, confidence intervals, p-values described as reported, and main outcome values come from the paper. Collapse rates, the exact McNemar result, Holm arithmetic, per-expert exposure, and epsilon curves are deterministic calculations from pure core functions. The faster epsilon curve is explicitly illustrative. No ML library, external image, or random draw is used.

References

  1. Ardaiz, A.; Budati, V.; Habibnia, A. (2026). Mixture-of-Experts for Cryptocurrency Order Execution: Training Stability, Tail Risk, and Failure Modes. ACM ICAIF 2026 / arXiv:2610.03369
  2. Dietterich, T. G. (1998). Approximate Statistical Tests for Comparing Supervised Classification Learning Algorithms. Neural Computation 10(7):1895–1923