The optimizer can honor every constraint and still miss the bill
Mean–variance optimization chooses portfolio weights by balancing expected return against covariance risk. Nosaka, Ikeda, and Takano train the return predictor for the portfolio it induces, rather than for squared prediction error alone. Their KKT reformulation preserves the full-investment rule and the ban on short sales inside learning.
That closes an important prediction–decision gap. But the upper-level objective contains variance and realized return only. It does not charge for moving from last month's weights to this month's weights. The paper reports turnover as a separate metric, and its winning DFL-KKT strategy turns over 0.991 per month in the international universe.
KKT conditions put the deployed portfolio inside training
The lower-level problem minimizes δwᵀVw/2 − (1−δ)r̂ᵀw subject to weights summing to one and remaining nonnegative. KKT conditions—stationarity, feasibility, and complementary slackness—are necessary and sufficient because the covariance estimate is positive definite and the feasible set is convex. Replacing the nested optimizer with those conditions creates one nonlinear training problem.
In two assets, the mechanism is visible without a solver. The unconstrained optimum is a linear function of the predicted return spread; the no-short constraint clips it into [0,1]. Small forecast changes near either boundary can therefore switch an asset on or off, even when their squared prediction errors look similar.
The headline winner also trades much more
| Method | Sharpe | Final wealth | Cum. decision loss | Turnover |
|---|---|---|---|---|
| DFL-KKT | 1.101 | 4.590 | 1.644 | 0.991 |
| IPO-GRAD | 0.984 | 3.869 | 1.719 | 0.299 |
| IPO-CF | 0.634 | 2.526 | 1.924 | 1.088 |
| SPO+ | 0.633 | 3.027 | 1.846 | 0.173 |
| PFL | 0.785 | 3.197 | 1.820 | 1.138 |
| 1/N | 0.747 | 2.774 | 1.902 | 0.016 |
The result is not that DFL-KKT necessarily loses after costs. The paper's gross final wealth advantage is real within its stated experiment. The narrower point is that its training loss and reported wealth do not establish the net ranking at any implementable fee, spread, or market-impact model.
A 20-basis-point assumption can flip the ranking
A basis point is one hundredth of a percentage point. For a transparent sensitivity check, suppose each month retains 1−cτ of wealth, where c is cost per unit turnover and τ is the paper's average turnover. Applying that constant-turnover overlay for the 120 test months gives the code below. It is an audit calculation, not a reconstruction of the unreported month-by-month path.
def cost_adjusted_terminal_wealth(
gross_terminal_wealth, average_turnover, periods, cost_rate
):
if gross_terminal_wealth <= 0 or average_turnover < 0 or cost_rate < 0:
raise ValueError('invalid wealth or cost')
if periods < 0 or int(periods) != periods:
raise ValueError('periods must be a non-negative integer')
retention = 1 - average_turnover * cost_rate
if retention < 0:
raise ValueError('cost exceeds wealth in a period')
return gross_terminal_wealth * retention ** periodsAt 25 bps, the overlay gives DFL-KKT terminal wealth 3.408 and IPO-GRAD 3.537. That reversal follows from reported endpoints and a declared cost rule, not fabricated returns. It is exactly the kind of threshold the frictionless objective leaves unidentified.
Turnover is a state variable, not a footer metric
Let wt−1 be the holdings before rebalancing and wt the new target. A simple turnover measure is Σ|wt,i−wt−1,i|. Moving a two-asset portfolio from [0.75,0.25] to [0.25,0.75] produces turnover 1.0. With the illustrative realized returns and covariance used here, gross mean–variance cost is -0.00328; charging 25 bps per unit turnover changes it to -0.00078.
Adding cΣ|wt−wt−1| to the upper objective changes the learning problem conceptually: yesterday's portfolio becomes part of today's state, and the optimal action develops a no-trade region. A prediction improvement must now be large enough to pay for the rebalance it triggers.
What to probe next
Re-run the rolling test with actual monthly weights and returns; deduct bid–ask spread, commissions, and a size-dependent impact curve before computing Sharpe and wealth; train with the same cost model used at evaluation; and report the ranking across a cost grid rather than at one chosen fee. The international and sector universes should be audited separately because their correlation and liquidity structures differ.
Also separate solver accuracy from economic accuracy. Residuals below 6.6×10⁻⁹ show that the reported KKT system was solved tightly; they cannot show that the system contains every cost that matters. For the predictive layer, continue with the linear-regression walkthrough; for the risk geometry, compare the hierarchical risk parity walkthrough.
References
- K. Nosaka, S. Ikeda, and Y. Takano (2026). Decision-Focused Learning for Mean-Variance Portfolio Optimization via KKT-Based Reformulation. PRICAI 2026; arXiv:2609.21427