AMD GPU Hang

Hi we are currently investigating a gpu hang that only occurs on amd pc gpus in a specific area of the game.

We are in discussion with amd as to the cause as well.

The issue manifests in a specific area of the game and so far we’ve not been able to reproduce it with hardware raytracing disabled.

The crash dumps we get don’t point to a culprit but often the breadcrumbs point to LumenReflections pass or the UpdateRadianceCaches pass being active at the time.

We have speculated that perhaps a high detail mesh in the acceleration structure was causing one of the passes to exceed the tdr theshold but with nothing pointing towards a culprit it’s been difficult to find a cause and we would’ve expected to see that manifest on the weaker ps5 hardware.

We don’t get a lot from DRED, we get a mix of hung and removed devices but that’s consistent in my experience depending on which thread the error is first picked up.

Catching the hang in the debugger usually shows a bunch of threads sitting blocked on a gpu sync fence.

I think all of this probably points to a hung device but don’t want to rule out a thread sync issue. AMD have pointed us towards AMD Radeon GPU Detective to try and get more information on what is going wrong, I will update the thread when we have results from that test.

[Attachment Removed]

Hi Andrew,

Thanks for reaching out. Please update us as soon as you can. Are you saying this hang only occurs on PS5, or are you saying that PC platforms with AMD GPUs could also be affected?

Cheers,

Tim

[Attachment Removed]

Ok, got it. Well, please let me know when you have some data, and we can take a closer look.

[Attachment Removed]

We caught the crash in Radeon GPU Detective and it detected the crash, unfortunately it wasn’t able to collect any information.

[Attachment Removed]

Oh, that’s unfortunate. We really need some data from the crash to help you out. It’s strange that GPU Detective didn’t collect any information. Was there anything mentioned in the engine logs about that?

[Attachment Removed]

Hi Tim,

AMD asked if I could reproduce the crash while running various ‘driver experiments’ in RGD to see if it would shine any more light on the issue. Unfortunately, most of the driver experiments would cause the game to crash during pso compilation so I wasn’t able to test the crash against those experiments. I was able to test it with the “Force NonUniformResourceIndex” experiment enabled though. It took a lot longer to trigger a crash than usual but I was able to produce one and RGD was able to capture the dump. Preliminary analysis seems to show a page fault accessing FSR4UPSCALER_ScratchBuffer. I’ve attached the crash summary below.

[Attachment Removed]

Hi Andrew,

The log file contains a lot of info, but some crucial pieces are missing. Did you have shader debug symbols enabled at the time of the crash? Also, does the 0xa790a0928e51838 PSO hash match any shader in your build? If you run your project with r.Shaders.Symbols, r.Shaders.ExtraData, and r.DumpShaderDebugInfo, you should have all of the necessary files in your Saved/ folder.

[Attachment Removed]

We currently have our build machine generate symbols via the "Example: Have Only the Build Machine Write Symbols to a Zip File

" section here - https://dev.epicgames.com/documentation/unreal\-engine/shader\-debugging\-workflows\-unreal\-engine?lang\=en\-US

I couldn’t find the hash anywhere in the Saved\Shaders\PCD3D_SM6-PCD3D_SM6 folder and the zip file in Saved\ShaderSymbols is empty.

I have been running various tests to try and narrow down the issue.

We saw some correlation with FSR4, so we tried with FSR disabled but the issue persisted.

I tried to limit the number of rt scene build job to no avail. - r.RayTracing.Geometry.MaxBuiltPrimitivesPerFrame 10000

I tried to disable inline raytracing but the issue persisted - r.RayTracing.AllowInline

Disabling pipeline rt however did prevent the issue from manifesting, so I believe I’ve narrowed it down - r.RayTracing.AllowPipeline

This could correlate with that pso compliation error but we don’t always see that in the logs and I’m not certain it’s the same pso that fails when it does manifest.

I will try and compile a build with the shader debugging flags listed above enabled and see if that lets us track down the shader hash.

[Attachment Removed]

Ok, let me know what you find out. I do want to add that although we have added support for pipelined hardware ray tracing mode, we are still working to make it more stable, and I suspect you are just running into one of these issues. Typically, we recommend users stick with inline ray tracing, which should still give you good results and better stability.

[Attachment Removed]

Because of the low repro rate for the bug, it’s taken time to verify whether disabling pipeline rt did in fact prevent it. Unfortunately we have since managed to reproduce the crash with pipeline rt disabled.

[Attachment Removed]

Hi again, I am sorry to hear that turning off pipeline ray tracing has not helped with your issue. Unfortunately, without a more reliable way to reproduce this issue, it will be tough to narrow down the root cause. Have you gotten any further feedback from AMD about this issue? I know this might be redundant, but have you upgraded your drivers recently? Oftentimes, driver upgrades can fix these kinds of issues on their own, though it is just a shot in the dark.

[Attachment Removed]

Driver version has had no impact on the crash so far no, we are still in discussion with AMD on the issue. We are sending them a build which should hopefully let them shed more light on the issue. In the mean time, in the absence of any other ideas I am going to continue to try and correlate the shader hash of the failed PSO.

[Attachment Removed]

Ok, that sounds like a good idea. Please let me know once you can determine which shader is causing the hang, or if AMD comes back to you with any updates that require action from my end.

[Attachment Removed]

I’ve not been able to get Radeon GPU Detective to catch the crash but AMD did have some success here. Unfortunately the build they caught it in didn’t have gpu markers and without the markers it’s difficult to ascertain the cause. Would you know what I would need to do to enable gpu markers to allow us to diagnose the issue?

[Attachment Removed]

ohh AMD did share the pso hash that crashed and it matched the one we already had suspicions on so I am definitely going to find the corresponding shader for that hash as soon as I can - 790a0928e51838

[Attachment Removed]

I can see from the rgd output that it seems to be faulting on a image_bvh8_intersect_ray with a base address of 0 so I think it’s probably pointing towards the bvh being unloaded while still in use.

Shader Resource Descriptor (SRD) Analysis
=========================================
Instruction: image_bvh8_intersect_ray v[36:45], [v[28:29], v[10:11], v[33:35], v[13:15], v23], s[32:35]// 000000000D20: D3E04010 17004024 0D210A1C
 
    Wave coordinate ID: 0x90508
    BVH (RDNA4):
      Base_address: 0x0000000000000000
      Sort_triangles_first: true
      Box_sorting_heuristic: ClosestMidPoint
      Box_grow_value: 6
      Box_sort_en: true
      Size: 0x0000040000000000 bytes
      Compressed format enable: true
      Box_node_64B: false
      Wide_sort_en: true
      Instance_en: true
      Pointer_flags: true
      Triangle_return_mode: true
      Type: BVH

[Attachment Removed]

Hi Andrew, yeah, that sounds plausible. Unfortunately, without knowing more about which shader we are crashing in, it is hard to say what is actually going on. We could make this question private so that you could share more about your investigation with AMD. Let me know if that works for you.

[Attachment Removed]

I built the game with the following flags -

r.Shaders.Symbols=1

r.Shaders.ExtraData=1

r.Shaders.GenerateSymbols=1

r.Shaders.WriteSymbols=1

I also tried by adding these flags as well -

r.ShaderDevelopmentMode=1

r.DumpShaderDebugInfo=1

r.Shaders.Optimize=0

r.Shaders.SkipCompression=1

but neither build generated a shader with the corresponding hash, I couldn’t find any reference to either the pso hash a790a0928e51838b or the shader hash 96940e8ae64e5124e6fafefc369faf45 of the offending shader in either Townfall\Saved\ShaderDebugInfo\PCD3D_SM6 or Townfall\Saved\ShaderSymbols\PCD3D_SM6 (both had over 50 gigs of shader data so I do believe the flags are working as intended)

[Attachment Removed]

Hi Andrew, those look like the correct settings. Could you upload the RGD dump so we can take a closer look? It might be that those hashes you have seen in the RDG output are not the ones we generate through DXC, so you cannot use them to find the actual shaders.

[Attachment Removed]

Soon after my last email we’ve found we were no longer able to reproduce the crash. There was a fairly large skeletal mesh that had been optimised down to a more reasonable vertex count the day before. We’re still not sure why it was crashing but I currently suspect that the mesh was too big for the scratch buffer unreal uses to build the acceleration structure. The strange thing about that is that the engine does use the GetRaytracingAccelerationStructurePrebuildInfo api to query the driver in order to allocate the correct sized buffer.

I’ve attached the RGD if you want to continue investigating but since it seems we’ve managed to resolve the issue I don’t think we really need further investigation on our end.

Thanks for all your help on this issue.

[Attachment Removed]