← Blog/blog/deterministic-policy-gradient-leakage

The policy gradient that leaks an action—until it vanishes

01

A gradient is a message, not exhaust

Picture a household robot learning from camera frames. Its device keeps the images local and sends gradients to a coordinator. A gradient is the vector of tiny parameter changes requested by one training example. It feels safer than uploading pixels, but it was computed from those pixels—and can retain their structure.

Bhujel and colleagues study that leak for embodied reinforcement learning (RL), where an agent acts inside an environment. Their attacker, TRACE, receives an ordered, per-step stream of policy gradients and reconstructs both RGB frames and actions. The coordinator is honest-but-curious: it follows the protocol but knows the policy architecture, checkpoint, action space, and learning rule, and owns auxiliary trajectories from the same task family.

02

One column points to the action

Let πk be the policy probability of action k, A theadvantage (how much better the sampled action looked than the baseline), and h the non-negative hidden representation feeding the policy head. Summing gradient column k gives σk = A‖h‖₁(πk − 1[k = a]). If A is positive, the sampled action is the smallest column sum; if A is negative, it is the largest.

def recover_action(probabilities, action, advantage, hidden_l1):
    columns = [
        advantage * hidden_l1 * (p - int(k == action))
        for k, p in enumerate(probabilities)
    ]
    if advantage > 0:
        recovered = min(range(len(columns)), key=columns.__getitem__)
    else:
        recovered = max(range(len(columns)), key=columns.__getitem__)
    others = [p for k, p in enumerate(probabilities) if k != action]
    gap = abs(advantage) * hidden_l1 * (
        1 - probabilities[action] + min(others)
    )
    return columns, recovered, gap

With three or more actions, the paper also proves a sign-free rule: choose the largest absolute column sum. With exactly two actions, the magnitudes tie, so the sign of A must come from another signal such as the value head. The experiments use five actions.

03

Certainty erases the evidence

The theorem's separation is Δ = |A|‖h‖₁(1 − πa + mink≠a πk). As the policy becomes deterministic, πa approaches one and every alternative approaches zero. Both Δ and the entire policy-gradient block vanish. The same formula that exposes the action also states exactly when that evidence disappears.

identifiability gapgradient L1 magnitude
Fig 1. Exact five-action illustrative sweep from the core implementation (A = ‖h‖₁ = 1; non-chosen probability split evenly). Both separation and gradient magnitude collapse in the one-hot limit. No paper measurements are used here.
04

Past gradients turn frames into a trajectory

TRACE encodes each gradient, interleaves it with causal context tokens, and uses a causal transformer to decode images and actions autoregressively. “Causal” means frame t may use gradients up to t, never future ones. On 100 held-out eight-step trajectories from AI2-THOR, it beats single-frame and optimization-based attacks under the paper's adapted RL protocol.

MethodPPO PSNR (dB)Action accuracyInference / frame
DLG5.4716.6%31.0 s
IG8.7331.5%39.3 s
LtI16.79100.0%1.2 ms
TRACE18.77100.0%4.5 ms
Paper-reported PPO baseline comparison at T = 8. PSNR is peak signal-to-noise ratio: higher means lower pixel distortion. Times exclude training the learned attacker.
Fig 2. PSNR recomputed from each method's paper-reported mean MSE using 10 log10(1/MSE). These transformed means differ slightly from the paper's mean per-image PSNR because averaging and the logarithm do not commute.
paper-reported PPO PSNR
Fig 3. Paper-reported PPO ablation. Temporal context supplies the gain: T = 1 reaches 12.85 dB, below the single-frame LtI baseline, while T = 8 reaches 18.74 dB. Runs at different T do not hold total frame count fixed, and T = 64 switches to RoPE.
05

Compression is not automatically privacy

Mild compression barely changes the attack. Keeping only the largest 10% of PPO gradient entries leaves 99.9% action accuracy, and eight-bit quantization does the same. Coarse quantization, strong noise, and differentially private SGD (DP-SGD) do much more—but the paper does not measure how much navigation quality they sacrifice.

PPO defenseSettingPSNR (dB)Action accuracy
none—18.9 ± 3.799.9%
quantization8-bit18.5 ± 3.999.9%
quantization4-bit12.4 ± 3.237.9%
quantization2-bit11.3 ± 2.819.4%
pruningkeep 10%18.8 ± 3.899.9%
Gaussian noiseσ = 0.114.2 ± 3.359.9%
DP-SGDε = 112.0 ± 2.217.6%
Selected paper-reported zero-shot defense results. Random chance is 20% for the five-action policy. DP-SGD uses δ = 10⁻⁵.
Fig 4. Selected paper-reported PPO action recovery. Aggressive quantization and DP-SGD approach five-way random chance; pruning 90% of entries does not. Privacy and control utility were not jointly evaluated.
06

Averaging changes the question

The appendix also tries gradients averaged over four or eight steps. Once order is lost, the action target becomes a histogram, not a sequence, and image quality is scored only on the last frame. For four-step windows, a from-scratch attacker reports 17.03 dB PSNR and total-variation distance 0.296; at eight steps, 16.03 dB and 0.268. Those are single runs under a changed target, not evidence that ordinary local training is broadly invertible.

07

What to probe next

First, sweep policy entropy and advantage magnitude while measuring action error against Δ; that directly tests the theorem's predicted failure surface. Then separate policy, value, and shared-backbone gradients to learn which block leaks pixels after the action block goes quiet. Finally, replace ordered per-step gradients with realistic local updates: shuffled batches, multiple epochs, clipping, secure aggregation, and an explicit control-utility budget.

The engineering lesson is not “never share gradients.” It is that gradients need a threat model. A neural networkcan expose behavior through a few signed column sums and perception through thousands of correlated coordinates. Audit the exact artifact the server receives, not the fact that raw data stayed on device.

References

  1. S. Bhujel, S. Shi, R. Huang, N. Zhang, and Y. Xiao (2026). Temporal Gradient Inversion for Private Trajectory Reconstruction in Embodied Reinforcement Learning. Advances in Neural Information Processing Systems; arXiv:2609.30258
  2. L. Zhu, Z. Liu, and S. Han (2019). Deep Leakage from Gradients. Advances in Neural Information Processing Systems 32
  3. J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov (2017). Proximal Policy Optimization Algorithms. arXiv:1707.06347