Hello,
I have found an issue with the trust region when adavantage is negative. Take a look at this code snippet.
surr = torch.min(ratio, ratio.clamp(eps_clip_log_min, eps_clip_log_max))
loss_surr += ... adv[:,i] * -(surr / act_num) ...
min(r, clamp(r, lo, hi)) is min(r, hi) for any lo β the lower bound is ignored. And the advantage is multiplied in after the min, whereas PPO takes the min over the products, so the sign of the advantage decides which side should bind.
Why it matters
For a positive advantage PPO stops pushing once the ratio exceeds 1+Ξ΅. For a negative advantage it should stop pushing once the ratio drops below 1βΞ΅. The old code applied the upper bound to both cases, so negative-advantage samples had no trust region at all β unbounded descent.
It stays harmless only while MinimumAdvantage = 0, which zeroes every negative advantage before the surrogate sees it. The editor slider goes to β10 and its tooltip invites negative values. Anecdotally for my environments training would completely fall apart when MinimumAdvantage < 0 which is what lead me to look into this.
Example
A three-condition experiment on the my task. All conditions shared one observation configuration, one set of hyperparameters and AdvantageMin = β3; the only variable was whether the trainer carried the fix below. Condition 3 ran the original trainer with extra diagnostics enabled, so the collapse could be watched directly.
| per-gather | condition 2 Β· fixed | condition 3 Β· original |
|---|---|---|
log_std_mean at gather 7 |
0.154 | β0.542 |
log_std_min, worst |
β0.13 | β4.25 (Ο = 0.014) |
approx_kl, peak |
0.009 | 8.55 |
grads/policy_preclip, peak |
0.05 | 679 |
clip_fraction |
0.05 | 0.18 |
critic/explained_variance |
0.21 | 0.53 β 0.20 |
Both fixed arms held log_std_mean between 0.15 and 0.23 across the same window. Measured frac_positive in the collapsing run ran 0.25 β 0.48, so 52β75% of every batch was negative-advantage and therefore unconstrained.
Clip the side that the advantage sign actually makes binding, staying in Epicβs log-space formulation.
surr = torch.where(
adv[:,i] >= 0.0,
torch.min(ratio, torch.full_like(ratio, eps_clip_log_max)),
torch.max(ratio, torch.full_like(ratio, eps_clip_log_min)))
There is one other issue. The log_std is only clamped on the positive side.
log_std_clipped = torch.clip(log_std, None, 10.0)
std = torch.exp(log_std_clipped)
logp = -((value - mean)**2) / (2 * std**2) - log_std_clipped - log_sqrt_2pi
If log_std is pushed very negative (below -54) then std becomes a 0 in the denominator of logp leading to a NaN and training collapse. I saw this occur in my training runs in the current version but not with the fix described above. I was able to prevent this by clamping the negative log_std as well.
log_std_clipped = torch.clip(log_std, -10.0, 10.0)