A gradient is a message, not exhaust
Picture a household robot learning from camera frames. Its device keeps the images local and sends gradients to a coordinator. A gradient is the vector of tiny parameter changes requested by one training example. It feels safer than uploading pixels, but it was computed from those pixels—and can retain their structure.
Bhujel and colleagues study that leak for embodied reinforcement learning (RL), where an agent acts inside an environment. Their attacker, TRACE, receives an ordered, per-step stream of policy gradients and reconstructs both RGB frames and actions. The coordinator is honest-but-curious: it follows the protocol but knows the policy architecture, checkpoint, action space, and learning rule, and owns auxiliary trajectories from the same task family.
One column points to the action
Let πk be the policy probability of action k, A theadvantage (how much better the sampled action looked than the baseline), and h the non-negative hidden representation feeding the policy head. Summing gradient column k gives σk = A‖h‖₁(πk − 1[k = a]). If A is positive, the sampled action is the smallest column sum; if A is negative, it is the largest.
def recover_action(probabilities, action, advantage, hidden_l1):
columns = [
advantage * hidden_l1 * (p - int(k == action))
for k, p in enumerate(probabilities)
]
if advantage > 0:
recovered = min(range(len(columns)), key=columns.__getitem__)
else:
recovered = max(range(len(columns)), key=columns.__getitem__)
others = [p for k, p in enumerate(probabilities) if k != action]
gap = abs(advantage) * hidden_l1 * (
1 - probabilities[action] + min(others)
)
return columns, recovered, gapWith three or more actions, the paper also proves a sign-free rule: choose the largest absolute column sum. With exactly two actions, the magnitudes tie, so the sign of A must come from another signal such as the value head. The experiments use five actions.
Certainty erases the evidence
The theorem's separation is Δ = |A|‖h‖₁(1 − πa + mink≠a πk). As the policy becomes deterministic, πa approaches one and every alternative approaches zero. Both Δ and the entire policy-gradient block vanish. The same formula that exposes the action also states exactly when that evidence disappears.
Past gradients turn frames into a trajectory
TRACE encodes each gradient, interleaves it with causal context tokens, and uses a causal transformer to decode images and actions autoregressively. “Causal” means frame t may use gradients up to t, never future ones. On 100 held-out eight-step trajectories from AI2-THOR, it beats single-frame and optimization-based attacks under the paper's adapted RL protocol.
| Method | PPO PSNR (dB) | Action accuracy | Inference / frame |
|---|---|---|---|
| DLG | 5.47 | 16.6% | 31.0 s |
| IG | 8.73 | 31.5% | 39.3 s |
| LtI | 16.79 | 100.0% | 1.2 ms |
| TRACE | 18.77 | 100.0% | 4.5 ms |
Compression is not automatically privacy
Mild compression barely changes the attack. Keeping only the largest 10% of PPO gradient entries leaves 99.9% action accuracy, and eight-bit quantization does the same. Coarse quantization, strong noise, and differentially private SGD (DP-SGD) do much more—but the paper does not measure how much navigation quality they sacrifice.
| PPO defense | Setting | PSNR (dB) | Action accuracy |
|---|---|---|---|
| none | — | 18.9 ± 3.7 | 99.9% |
| quantization | 8-bit | 18.5 ± 3.9 | 99.9% |
| quantization | 4-bit | 12.4 ± 3.2 | 37.9% |
| quantization | 2-bit | 11.3 ± 2.8 | 19.4% |
| pruning | keep 10% | 18.8 ± 3.8 | 99.9% |
| Gaussian noise | σ = 0.1 | 14.2 ± 3.3 | 59.9% |
| DP-SGD | ε = 1 | 12.0 ± 2.2 | 17.6% |
Averaging changes the question
The appendix also tries gradients averaged over four or eight steps. Once order is lost, the action target becomes a histogram, not a sequence, and image quality is scored only on the last frame. For four-step windows, a from-scratch attacker reports 17.03 dB PSNR and total-variation distance 0.296; at eight steps, 16.03 dB and 0.268. Those are single runs under a changed target, not evidence that ordinary local training is broadly invertible.
What to probe next
First, sweep policy entropy and advantage magnitude while measuring action error against Δ; that directly tests the theorem's predicted failure surface. Then separate policy, value, and shared-backbone gradients to learn which block leaks pixels after the action block goes quiet. Finally, replace ordered per-step gradients with realistic local updates: shuffled batches, multiple epochs, clipping, secure aggregation, and an explicit control-utility budget.
The engineering lesson is not “never share gradients.” It is that gradients need a threat model. A neural networkcan expose behavior through a few signed column sums and perception through thousands of correlated coordinates. Audit the exact artifact the server receives, not the fact that raw data stayed on device.
References
- S. Bhujel, S. Shi, R. Huang, N. Zhang, and Y. Xiao (2026). Temporal Gradient Inversion for Private Trajectory Reconstruction in Embodied Reinforcement Learning. Advances in Neural Information Processing Systems; arXiv:2609.30258
- L. Zhu, Z. Liu, and S. Han (2019). Deep Leakage from Gradients. Advances in Neural Information Processing Systems 32
- J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov (2017). Proximal Policy Optimization Algorithms. arXiv:1707.06347