Worse GPU performance on RTX3060 after migration from UE5.5 to UE5.7

Hello Epic Games Support Team!

After migration from UE5.5 to UE5.7 we noticed that most of our scenes have worse GPU performance on low spec machines with RTX3050 and RTX3060 in both Test and Shipping versions.

During investigation we confirmed that outside of few minor differences in features due to version upgrade the draw call count, vertex count, material complexity(VGPR count) is the same as in UE5.5, sometimes UE5.7 showed even better metrics due to improved culling. Yet performance for the same scene is usually worse in UE5.7.

In Nvidia Nsight Graphics we can find a difference that small passes are running in multiple small command lists while in UE5.5 they are merged in larger command lists. Images with example are attached. Source Nvidia Nsight GPU traces from RTX3060 are attached as well.

Please, advice us on improving performance for low spec targets like RTX3060.

UE5.5 (large command list)

[Image Removed]

UE5.7 (small command list for every pass)

[Image Removed]

These Cvars that we use might be relevant for the question.

r.NumBufferedOcclusionQueries=2

r.RHIValidation.DebugBreak.Transitions=1

r.RHICmdMinDrawsPerParallelCmdList=512

r.RHICmd.ParallelTranslate.MaxCommandsPerTranslate=2048

r.RHICmd.ParallelTranslate.CombineSingleAndParallel=1

r.D3D12.AllowAsyncCompute=0

[Attachment Removed]

Steps to Reproduce
N/A

[Attachment Removed]

UPD: I run the game with r.RDG.ParallelExecute=0 on RTX3050 and RTX3060. We see GPU time improvement from 0.5ms to 2ms which confirms that the GPU work split to small command lists leads to worse GPU time. Executing all RDG jobs in RHIThread added +5ms making it a bottleneck for most of the scenes. My clarification to this question is if it possible to use ParallelExecute only for big GPU workloads and batch small RDG passes to single command lists to keep the best of both cases. I have tried also r.RDG.ParallelExecute.PassMin=16, but GPU time improvement is not as good as r.RDG.ParallelExecute=0 variant.

[Attachment Removed]

Hello, I am colleague of Oleksii.

We did further investigation into the issue. As it was mentioned above, we confirmed UE 5.7 splits commands into more commandlists when compared to UE 5.5.

Good example are command lists containing parallel draws but also various other commandlists for passes like Atlas or at the beginning of the frame.

These translation chains are closed because of bShouldSplitForParentChild=1, not because MaxCommandsPerTranslate threshold was reached.

As you can see on the following image, this makes MaxCommandsPerTranslate innefective after some threshold because command lists are split because of bShouldSplitForParentChild.

[Image Removed]

Tracking down the introduction of this reason for splitting command lists, it was introduced in UE 5.6 with the following change.

https://github.com/EpicGames/UnrealEngine/commit/739d712a8ff6ccb5828d47f393c67fa849fc2217\#diff\-32797d4deb3f571dd5122f6546f03dc3e0cfc88f902876ef4e5e6fedc9148f63

Changing GRHISupportsParallelRenderPasses = false for D3D12 gets rid of these newly introduced splits and we get to state similar to UE 5.5.

I am curious if this splitting is intended behavior of the engine and whether you are not worried of additional overhead on lower end devices introduced by it. Given that there is no CVar to control this, it does not seem like something users are expected to simply disable.

[Attachment Removed]

Hi Ondrej, Oleksii,

My apologies for the delay on this.

You’ve found the correct changelist where the parallel render pass support was added, around the UE 5.6 timeframe (originally in CL 41638143 in UE5/Main). I’d imagine we don’t have a cvar to disable this as it wasn’t something we noticed in our own profiling, and isn’t affecting us today in later engine versions. We did do some additional work in UE 5.8 to submit GPU work more proactively on the CPU to reduce latency (CL 51908843), which can also lead to smaller batches of work arriving at the GPU.

Having said that, the RHI submission pipeline should be fairly aggressive at batching work. The submission thread within the D3D12 RHI will group available payloads together, so long as there isn’t a fence signal between two of them. Maybe this batching isn’t working well for you? Have you tried instrumenting the number of ExecuteCommandLists calls in D3D12Submission.cpp before/after the engine update?

It would be good to know if the number of ExecuteCommandLists calls has increased in your UE 5.7 build, and whether turning off GRHISupportsParallelRenderPasses drops that back down to UE 5.5 levels. The parallel RP support was added primarily for Vulkan and Metal which have strict requirements about how the parallel command lists can be built. It was later implemented in D3D12 using standard RHI contexts, so its possible we actually don’t need the split on that platform. I’d have to check with the original author.

Cheers,

Luke

[Attachment Removed]