Correlation is not information left to share
Suppose one model predicts several outputs at once: a robot's joint torques, multiple chemical concentrations, or tomorrow's electricity loads. The outputs can be strongly correlated simply because the input already explains the same common cause. A joint model cannot learn that information twice.
Zhou and Vanschoren test this distinction on 16 multi-target regression datasets. Their coupled model is an intrinsic coregionalisation model (ICM): a Gaussian process whose covariance links both inputs and targets. They compare it with independently fitted Gaussian processes. The important score is joint negative log likelihood (NLL), which rewards a model for assigning high probability to the whole output vector—not merely putting each mean close to its target.
Ask the standardized errors instead
Fit one probabilistic predictor per target using inner out-of-fold predictions. For target t, standardise each error as z = (y − μ) / σ, where μ is the predicted mean and σ the predicted standard deviation. The correlation matrix of these z values, R, measures dependence that remains after both the mean and each target's uncertainty have been accounted for.
The paper compresses R into D = −½ log det(R). For two targets with residual correlation ρ, the determinant is 1 − ρ². Zero means no linear Gaussian dependence remains; the score grows as the residuals approach the same line.
import math
def two_target_diagnostic(rho):
if not math.isfinite(rho) or abs(rho) >= 1:
raise ValueError('rho must be finite and strictly between -1 and 1')
return -0.5 * math.log(1 - rho * rho)It predicts uncertainty gains, not better means
| Dataset | Targets | Raw |corr| | D | Δ joint NLL | Δ R² |
|---|---|---|---|---|---|
| scm20d | 16 | 0.58 | 6.83 | −3.590 | −0.002 |
| atp7d | 6 | 0.63 | 2.02 | −1.446 | −0.000 |
| sarcos | 7 | 0.37 | 1.01 | +0.437 | −0.006 |
| tecator | 3 | 0.91 | 0.47 | +0.778 | −0.008 |
| energy | 2 | 0.82 | 0.03 | −0.073 | +0.005 |
A small score does not mean independence
This is the paper's most important boundary. D sees one global linear Gaussian covariance. In a four-target stress test, residual correlation is +ρ on one half of input space and −ρ on the other. The global correlations cancel: D falls from 1.76 to 0.08, although an oracle region-aware covariance still gains 1.18 nats.
A second construction sets one residual to a centred square of another. Pearson correlation—and therefore D—is zero. A Gaussian covariance layer correctly gains nothing, but a nonparametric conditional model gains 0.30 nats because nonlinear dependence remains.
| Stress test | Global D | Stronger model gain | What the global matrix misses |
|---|---|---|---|
| sign-changing covariance | 0.08 | 1.18 nats | +ρ and −ρ cancel |
| nonlinear residual dependence | 0.00 | 0.30 nats | Pearson correlation is zero |
Couple the covariance without moving the means
The paper's cheap residual-ICM baseline leaves each independent model's mean μ and marginal scale σ untouched, then forms Σ = diag(σ) R diag(σ). Estimating R costs O(mT²), its factorisation O(T³), and each joint score O(T²). That isolates the claimed benefit: calibrated co-movement among errors, without pretending the joint model discovered better point predictions.
This pattern also applies after a frozen shared representation. For example, atransformer can produce features while separate probabilistic heads produce μ and σ. Cross-validated standardized residuals then test whether an additional output-covariance layer is justified by information the representation did not already absorb.
What to probe next
Plot residual dependence against the input before trusting a single matrix. Compare global D with regionwise estimates, distance correlation, and held-out copula or conditional-density scores. Repeat the gate under different marginal predictors: a large D may expose genuine target coupling, or merely a weak independent model that left predictable structure in every error.
Most importantly, choose the evaluation target first. If downstream decisions need only separate means, joint NLL gains may have no operational value. If they need a scenario vector—portfolio risk, coordinated control, or simultaneous thresholds— the covariance is part of the prediction.
References
- F. Zhou and J. Vanschoren (2026). Residual Correlation as a Diagnostic for Joint-Uncertainty Gains from GP Coregionalisation. Asian Conference on Machine Learning; arXiv:2609.30085 [cs.LG]
- E. V. Bonilla, K. M. A. Chai, and C. Williams (2008). Multi-task Gaussian Process Prediction. Advances in Neural Information Processing Systems 20
- C. E. Rasmussen and C. K. I. Williams (2006). Gaussian Processes for Machine Learning. MIT Press