A good forecast can imply the wrong intervention
A time-series world model predicts what happens next from recent state history and a future action plan. Low error says its forecasts resemble held-out observations. It does not say the model understands what would happen if we deliberately changed one action.
That distinction matters whenever a forecast becomes a simulator. A clinician may ask whether more norepinephrine raises blood pressure; a greenhouse controller may ask whether opening a windward vent lowers temperature. A model can predict the recorded trajectory accurately while answering either directional question backward.
Treatment assignment can reverse the sign
Observational confounding appears when an action responds to the state it is meant to change. Vasopressors are more likely to be given when blood pressure is already low. In the record, higher dose can therefore coincide with lower pressure even if increasing the dose would raise it.
The core calculation recovers an observational slope of -0.6. A sequence model can exploit that stable association to reduce forecast error—and still learn the wrong response to an action shift.
Mechanism consistency asks a counterfactual sign question
The paper perturbs one future action upward or downward, keeps the history fixed, and averages the change in the target forecast. The score is the share of eligible examples whose predicted change has the known direction. It tests a sign, not whether the effect size is correct.
In a four-step synthetic example, shifting the action raises the average forecast by 0.4. Across four synthetic cases, three signs agree, so consistency is 75%. This illustration uses no randomness and is not a reproduction of a paper experiment.
A scale-free penalty pushes responses across zero
Directional supervision compares the baseline and shifted forecasts, penalizes only response mass with the wrong sign, then divides by the average absolute response. In training, the paper stops gradients through that denominator and adds the result to forecast loss with weight ρ = 0.1.
def directional_penalty(differences, mechanism_sign, epsilon=1e-6):
wrong = sum(max(0.0, -mechanism_sign * d) for d in differences) / len(differences)
magnitude = sum(abs(d) for d in differences) / len(differences)
return wrong / (magnitude + epsilon)
def total_loss(forecast_loss, differences, mechanism_sign, weight=0.1):
direction_loss = directional_penalty(differences, mechanism_sign)
return forecast_loss + weight * direction_lossDeclared signs are repaired almost perfectly
For the mechanisms supplied during training, the intervention works. The paper raises consistency above 0.99 while changing validation MAE by a median of −0.05%. These are paper-reported results, not values generated by the toy example above.
| Supervised mechanism | Gate | + directional supervision |
|---|---|---|
| Windward vent → air temperature | 0.196 | 1 |
| Propofol → BIS | 0.218 | 1 |
| Norepinephrine → mean pressure | 0.15 | 1 |
| Dobutamine → heart rate | 0.864 | 1 |
The repair stops at the edge of the sign list
The important buried result is what happens to mechanisms withheld from supervision. One improves slightly, one stays near perfect, and three become worse. Directional supervision is not evidence that the model has discovered a causal system; it is evidence that the supplied sign constraints can be enforced.
| Held-out mechanism | Gate | After other signs are supervised |
|---|---|---|
| CO₂ dosing → CO₂ | 0.097 | 0.154 |
| Cooling setpoint → temperature | 0.373 | 0 |
| Remifentanil → mean pressure | 0.284 | 0.156 |
| Vasopressin → mean pressure | 0.531 | 0.182 |
| Dopamine → heart rate | 0.999 | 0.978 |
What to probe next
The next evaluation should withhold whole mechanism families, vary perturbation size, and report response magnitude and uncertainty—not only sign accuracy. In healthcare, prospective or carefully designed quasi-experimental validation must come before using a model as a treatment simulator.
Rebuild the confounded association on the linear-regression page, then compare it with a sequence model on the Transformer page. More expressive prediction does not, by itself, change what the data identify.
References
- Haochen Zhang, Jiaheng Guo, Zhen Xu, Zachary Plotkin, Nicholas Konz, Zhen Tan, Tianlong Chen (2026). On the Divergence of Accuracy and Mechanism Consistency in Time Series World Models. arXiv:2610.01842
- Miguel A. Hernán and James M. Robins (2020). Causal Inference: What If. Chapman & Hall/CRC