[UE 5.8.2][DX12] Repeated DEVICE_HUNG crashes in Nanite/VSM shortly after map load on RTX 4080 SUPER

Summary

Unreal Engine 5.8.2 repeatedly encounters DXGI_ERROR_DEVICE_HUNG shortly after the opening map becomes available, usually two or three frames after map loading completes. The issue reproduces on an RTX 4080 SUPER with both NVIDIA driver 610.60 and 617.14. NVIDIA Aftermath evidence identifies faulting reads in Nanite NodeAndClusterCull and Virtual Shadow Map culling shaders. We are requesting help determining whether this may be related to UE-230827 or another Unreal Engine rendering-path issue.

What Type of Bug are you experiencing?

Rendering (Graphics / Niagara)

Steps to Reproduce

  1. Launch the affected Unreal Engine 5.8.2 project in DirectX 12 -game mode on Windows 11 with Nanite and Virtual Shadow Maps enabled.
  2. From the title screen, select New Game.
  3. Allow the project to load the opening world.
  4. Wait for the map-loading/world-available log point and the first several frames of the new world.
  5. Intermittently, usually two or three frames after map loading completes, the GPU is removed with DXGI_ERROR_DEVICE_HUNG.

The issue is intermittent rather than deterministic. It has reproduced repeatedly on an NVIDIA GeForce RTX 4080 SUPER using both driver 610.60 and 617.14. We do not currently have a minimal reproduction project; these steps describe reproduction in our production project.

Expected Result

The opening world should render normally and the game session should continue without GPU device removal or a process crash.

Observed Result

The GPU is removed with DXGI_ERROR_DEVICE_HUNG and the game process terminates. NVIDIA Aftermath reports faults in Nanite NodeAndClusterCull or Virtual Shadow Map CullPerPageDrawCommands compute shaders.

Six representative events have been independently mapped to faulting structured-buffer reads across Nanite and VSM shader permutations. No matching range was reported by the engine’s live-allocation, live-heap and retained-release lookups for the applicable fault addresses. Aftermath resource tracking also reported no associated resource information.

The same shader families have failed on NVIDIA driver 610.60 and 617.14. Disabling RDG async-compute transient aliasing did not prevent the issue. Disabling RDG async compute and changing one target buffer’s transient allocation path produced zero events in limited 30-launch batches, but those results are not sufficient to establish a fix or root cause.

Affects Versions

5.8

Platform(s)

Windows

For crash reports, include your callstack

This is a GPU device-removal event rather than a conventional CPU exception. The available CPU crash path ends in Unreal Engine’s D3D12 device-removal handling and does not identify the faulting shader operation.

NVIDIA Aftermath GPU crash dumps, matching .nvdbg files, exported DXIL, decoded shader locations and complete engine logs are available. Because the binary dumps may contain embedded machine, account or project information, they have not been attached to this public post and can be provided through an appropriate private support channel.

Additional Notes

Environment:

  • Unreal Engine 5.8.2 Launcher build, CL 56702186
  • Windows 11 64-bit
  • NVIDIA GeForce RTX 4080 SUPER
  • DirectX 12, Nanite and Virtual Shadow Maps enabled
  • Reproduced on NVIDIA driver 610.60 and 617.14

Investigation summary:

  • Six representative GPU crash events have been mapped from Aftermath program counters to DXIL instructions.
  • The mapped faults are structured-buffer reads across Nanite NodeAndClusterCull and VSM CullPerPageDrawCommands shader permutations.
  • Source and uniform-buffer layout analysis associates the reads with four small per-frame buffers referenced through the Scene and VirtualShadowMap uniform buffers. These resource names are source associations, not direct identifications reported by Aftermath.
  • For the applicable fault addresses, no matching range was reported by the engine’s live-allocation, live-heap or retained-release lookups. Aftermath resource tracking also reported no associated resource.
  • A retained-release diagnostic sample used a retention window of 100000, greater than that process’s frame-fence value of 1260. This was a single sample, and the internal retained-array size was not directly observed, so the result is not generalized to all crash signatures.
  • Controlled tests of RDG async compute, async-compute transient aliasing and one buffer’s transient allocation path did not establish a root cause or a reliable fix.
  • We do not currently have a minimal reproduction project.

The symptoms appear related to UE-230827, which currently lists UE 5.8.2 among its affected versions, but we do not have evidence that this project is encountering the same root cause:

We have also submitted a technical summary to NVIDIA. A sanitized full report, detailed reproduction record, six paired Aftermath dump/.nvdbg samples, matching DXIL, decode output and evidence hashes are available upon request through a private channel.