Summary
When an agent is removed from a LearningAgentsManager via RemoveAgent (to temporarily pause it) and later re-added via AddAgent, the re-added agent permanently desyncs from the trainer’s internal step counters. This happens regardless of whether a new Agent ID is assigned, and does not self-recover over subsequent training iterations. The result is a warning that repeats forever, and the affected agent’s experience is silently dropped from training every step.
LogLearning: Warning: PPOTrainer_X: Agent with id N has non-matching iteration numbers
(observation: 1, action: 1, action modifiers: 1, reward: 2, completion: 2).
Experience will not be processed for it.
What Type of Bug are you experiencing?
AI
Steps to Reproduce
Create a LearningAgentsManager with a PPOTrainer, Interactor, and Policy.
AddAgent for N agents, call BeginTraining/RunTraining in Tick. Confirm training runs normally with no warnings.
After several successful training iterations, call RemoveAgent (or RemoveAgents) on one or all agents from within the Tick event, immediately after RunTraining returns (i.e., not mid-callback — timing confirmed correct via logging).
Wait an arbitrary number of ticks (agent(s) sit outside the manager).
Call AddAgent again for the same agent(s), also placed immediately in Tick, before RunTraining.
Observe the log: PPOTrainer_X: Agent with id N has non-matching iteration numbers (…). Experience will not be processed for it.
Expected Result
The re-added agent should resume gathering valid training experience on the next step, the same way it did when it was first added before training started.
Observed Result
The re-added agent permanently fails the “non-matching iteration numbers” check and its experience is silently dropped from training on every subsequent step, indefinitely (does not self-recover).
Affects Versions
5.8
Platform(s)
Windows