Hello,
I think there is an issue with filling the replay buffer when an episode is longer than the remaining steps.
If you look here:
and here:
InEpisodeFinalObservations[Index][InstanceIdx]) and InEpisodeFinalMemoryStates[Index][InstanceIdx]) are the final states of the episode but because the buffer could not fit an entire episode they are not the correct values to fill in.
For example if the episode was 100 steps long but PartialStepNum is 10. This will take the final values at step 100 giving you values from steps [0, 1 , 2 , 3 , 4 , 5, 6, 7, 8, 100]. This is particularly bad for learning because the episode is marked as truncated so it will bootstrap the value estimate using the critic with the incorrect final observation. Advantages will also be computed from the wrong final observation.
Luckily this can only affect one value per gather. In my training runs 0.4%-0.8% of each gather carried an incorrect advantage and critic return label.
The correct copies are:
Array::Copy(
EpisodeFinalObservations[Index][EpisodeNum],
EpisodeBuffer.ObservationArrays[Index][InstanceIdx][PartialStepNum]);
Array::Copy(
EpisodeFinalMemoryStates[Index][EpisodeNum],
EpisodeBuffer.MemoryStateArrays[Index][InstanceIdx][PartialStepNum]);