GPU crashes w/o Aftermath dumps follow-up

This question was created in reference to: [Aftermath GPU crash/hang dumps failing on standalone cooked [Content removed]

We are still getting these sorts of crashes despite a number of attempts to fix various issues. I happened to run across a reply within the thread [GPU crash in Nanite::InitContext or [Content removed] that looked potentially relevant:

[*Alex [Content removed] (Epic Games)

7 months ago

Hi,

Putting an update in here for anyone else following along. After further investigation this GPU hang does appear to be similar to what was happening in the case where PSOs were not released in UE 5.5 and the driver was running out of memory and the GPU was crashing. The largest similarity here, as in that case, was the absence of an Aftermath dump file, and the other symptoms included random breadcrumbs and in passes that are simple and shouldn’t crash, as well as the GPU crash taking a long time (30+mins to hours) before crashing and the local PSO cache size becoming large.

Currently, there isn’t a known workaround for this issue, but it appears to be rare, and we’re hopeful a future driver update will address the underlying issue.

The symptoms listed here (lack of Aftermath dump, random breadcrumbs on each instance) sounds almost exactly like what we are seeing. Is there a way that we can test to confirm that we are hitting this issue, or alternatively rule it out? Later in that thread, someone mentions that Nvidia made a custom driver exception for their project, but that’s not something we can really do right now for testing.

[Attachment Removed]

Steps to Reproduce
The crash just happens arbitrarily after playing the game for some period of time. There does not seem to be any consistent reproduction steps, although we are seeing some people/PCs being more susceptible than others.

[Attachment Removed]

Hey there!

Currently, there isn’t a known workaround for this issue, but it appears to be rare, and we’re hopeful a future driver update will address the underlying issue.

This was the issue I mentioned in my earlier post that was later fixed in driver versions 595+ and you mentioned you are seeing GPU crashes in later driver versions. The way this issue was diagnosed was the licensee provided a repro to Nvidia to debug with. The repro was an .exe that had a level with some of their static level content and a camera that flew over that content over and over until it crashed. Nvidia was able to use this to isolate the driver issue. However, Nvidia was able to deploy an immediate workaround for the issue which I think involved increasing the heap limit for that game, so that may be something you can also work with Nvidia to do - at least temporarily to see if it alleviates the crashes. The other thing is to look at how many PSOs/shaders you’re using - in the case above, the PSO/shader usage stabilized after 5-10 minutes - and that made it easier to see that it was the driver optimizations causing the issue. You may also be able to disable those kind of optimizations with NvidiaSettings.exe, but that likely requires special flags that Nvidia could confirm.

Out of curiosity can you provide a sample of some of the breadcrumbs you’re seeing from logs where you get GPU crashes that don’t have active Raytracing passes?

[Attachment Removed]

We had one of our partner studios confirm that they had seen similar issues and that Nvidia had update the proflie for their title to increase the size of the internal shader heap. We threw together a script to apply the same update for internal testing within the studio, and we’ve seen promising results so far. We didn’t go down to 0 GPU crashes reported, but several people who were frequently hitting unexplained crashes have reported much better stability after applying the change. I happened to have a 100% repro case myself when trying to load into the main game map with -gpuvalidation enabled, which went away when I applied the profile change locally.

[Image Removed]

I’m not sure what the other driver optimization changes would be, but we have a support email thread communication open with Nvidia right now in parallel to this one.

As far as breadcrumbs go, we had a full-day studio-wide playthrough recently and I was able to scrape and collate 125 separate log files. I’m sure not all of these are the same thing, but here’s the breakdown of what the RHI breadcrumbs said were active on the GPU across those logs. Of these 125 logs, I believe exactly ONE came with an attached .nv-gpudmp file, and that one said that no GPU fault was found.

=== Top Active Graphics Leaf Breadcrumb Paths ===
   27  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\RayTracingScene
   13  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\HZB\BuildHZB(ViewId=0)\BuildHZB
   10  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\ShadowDepths\RenderVirtualShadowMaps(Nanite)\Shadow Depths - VSM (Nanite)\Nanite::DrawGeometry\MainPass\CalculateSafeRasterizerArgs
    8  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\BasePass\NaniteBasePass\Nanite::BasePass\Nanite::ShadeBinning
    7  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\RenderDeferredLighting\Lights\DirectLighting\UnbatchedLights\Evergreen.BP_UL_Lighting_Runtime_Weather\ShadowProjectionOnOpaque
    5  SceneRender - ViewFamilies\RenderGraphExecute - %s
    5  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\Translucency\RenderTranslucency\ParallelDraw (Index: 0, Num: 1)
    4  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\VirtualTextureFinalizeRequests\BuildRenderingCommands(Culling=Off)
    4  SceneRender - ViewFamilies\NiagaraGpuComputeDispatch
    3  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\RenderDeferredLighting\InitTranslucencyLightingVolumeTextures
    3  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\TC Fog Scattering\TC Fog Scattering\TAA
    3  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\PostProcessing\ScreenSpaceFogScattering 1280x720 (PassAmount=8)\SSFS Downsample
    3  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\PostProcessing\TemporalSuperResolution(sg.AntiAliasingQuality=2) 1280x720 -> 2560x1440
    3  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\ShadowDepths\RenderVirtualShadowMaps(Non-Nanite)\Shadow Depths - VSM (Non-Nanite)\Batched\CullingPasses
    2  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\PostRenderOpsFX\FXSystemPostRenderOpaque\GPUParticles_PostRenderOpaque
    2  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\PostRenderOpsFX\NiagaraGpuComputeDispatch
    2  DisplacerCapturer\SceneRender\RenderGraphExecute - %s\SceneCapture::DisplacerCapturer\Scene\GPUSceneUpdate\NaniteSkinning
    2  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\PostProcessing\MotionBlur
    2  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\ShadowDepths\RenderVirtualShadowMaps(Nanite)\Shadow Depths - VSM (Nanite)\Nanite::DrawGeometry\PostPass\CalculateSafeRasterizerArgs
    2  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\Nanite::VisBuffer\Nanite::DrawGeometry\MainPass\CalculateSafeRasterizerArgs
    2  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\Nanite::VisBuffer\Nanite::DrawGeometry\MainPass\PatchSplit
    2  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\Nanite::VisBuffer\Nanite::DrawGeometry\PostPass\CalculateSafeRasterizerArgs
    2  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\Nanite::VisBuffer\Nanite::DrawGeometry\PostPass\PatchSplit
    2  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\NaniteBasePass\Nanite::EmitDepthTargets
    2  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\PostProcessing\ScreenSpaceFogScattering 1280x720 (PassAmount=8)\SSFS Upsample
 
=== Top Active Compute Leaf Breadcrumb Paths ===
   19  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\RayTracingDynamicGeometry\RayTracingDynamicGeometryUpdate
   17  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\PrepareImageBasedVRS\PrepareImageBasedVRS\ContrastAdaptiveShading
   16  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\ComputeVolumetricFog\VolumetricFog
   16  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\PostProcessing\TemporalSuperResolution(sg.AntiAliasingQuality=2) 1280x720 -> 2560x1440
   16  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\DiffuseIndirectAndAO\LumenReflections
   13  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\LumenSceneLighting\DirectLighting\Light Attenuation ShadowMask
   13  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\LumenSceneLighting\DirectLighting\Offscreen shadows
   12  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\LumenSceneLighting\DirectLighting\Lights
   11  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\LumenSceneLighting\BuildCardUpdateContext
   11  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\LumenSceneLighting\Radiosity
   10  DisplacerCapturer\SceneRender\RenderGraphExecute - %s\SceneCapture::DisplacerCapturer\Scene
    9  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene
    8  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\DiffuseIndirectAndAO\TranslucencyVolumeLighting
    7  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\SkyAtmosphereLUTs
    6  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\DiffuseIndirectAndAO\LumenScreenProbeGather
    5  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\DiffuseIndirectAndAO\LumenScreenProbeGather\UpdateRadianceCaches
    5  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\LumenSceneLighting\DirectLighting\CullTiles 1 lights
    4  WeatherOcclusionSceneCapturer\SceneRender\RenderGraphExecute - %s\SceneCapture::WeatherOcclusionSceneCapturer\Scene
    3  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\PostProcessing\TemporalSuperResolution(sg.AntiAliasingQuality=2) 960x600 -> 1920x1200
    3  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\PostProcessing\TemporalSuperResolution(sg.AntiAliasingQuality=2) 1024x576 -> 2048x1152
    3  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\FXSystemPreRender
    2  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\LumenSceneLighting\DirectLighting\CullTiles 71 lights
    2  TexturePoolCopyOps
    2  WeatherOcclusionSceneCapturer\SceneRender\RenderGraphExecute - %s\SceneCapture::WeatherOcclusionSceneCapturer\Scene\Nanite::EndAsyncUpdate\Nanite::Transcode
    2  SceneRender - ViewFamilies\RenderGraphExecute - %s\Scene\VolumetricCloudShadow

[Attachment Removed]

Thanks, you’re right that running with more PSOs in memory runs a higher risk of the GPU crash related to shader compilation. The defaults are

r.PSOPrecache.KeepInMemoryGraphicsMaxNum = 2000
r.PSOPrecache.KeepInMemoryComputeMaxNum = 1000

I’m sure you’re trying to strike a good balance between number of shaders, shader complexity and quality/content so I’m assuming you’re already working on the content side optimizations but I will add a couple other thoughts.

In 5.7+ we added mode 2 for r.PSOPrecache.KeepInMemoryUntilUsed and made it the default which may be worth integrating:

CL#43263054 (c721a5) Cleanup used precached PSOs, when they are kept in memory to workaround driver cache inefficiencies.

Once a PSO is needed for rendering, the PSO precache system sets a flag that marks the precached PSO (if existing) as used, and enqueues it for cleanup if it is currently being held in memory. This is enabled with r.PSOPrecache.KeepInMemoryUntilUsed=2.

Also remove some unnecessary duplicated code now that we have compute initializers, and clearly separate the implementation for the compile task between graphics and compute. This clean up work should be a “no fuctional change” cleanup but comes with opportunities to simplify the code further down the line.

int32 GPSOPrecacheKeepInMemoryUntilUsed = 2;
static FAutoConsoleVariableRef CVarPSOPrecacheKeepInMemoryUntilUsed(
    TEXT("r.PSOPrecache.KeepInMemoryUntilUsed"),
    GPSOPrecacheKeepInMemoryUntilUsed,
    TEXT("If enabled and if the underlying GPU vendor is NVIDIA or Qualcomm, precached PSOs will be kept in memory instead of being deleted immediately after creation, and will only be deleted once they are actually used for rendering or when the maximum limit of in-memory PSOs is reached.\n")
    TEXT("This can speed up the re-creation of precached PSOs for NVIDIA and Qualcomm drivers and avoid small hitches, at the cost of memory.\n")
    TEXT("It's recommended to set r.PSOPrecache.KeepInMemoryGraphicsMaxNum and r.PSOPrecache.KeepInMemoryComputeMaxNum to a non-zero value to ensure the number of in-memory PSOs is bounded.\n")
    TEXT("Valid options:\n")
    TEXT("0 = off (default)\n")
    TEXT("1 = PSOs are kept in memory after precaching but not deleted immediately when used by the renderer. They are instead only evicted when the limit of in-memory PSOs is reached\n")
    TEXT("2 = PSOs are kept in memory after precaching and deleted when used by the renderer"),
    ECVF_ReadOnly);

In 5.8 we reduced number of PSOs in general in some areas with these changelists:

CL#45891754 (72f889) Add special Slate shader version that uses dynamic branching instead of static specializations. This new variant will be the only path, replacing most of the above after performance and correctness testing.

The purpose of the change is to reduce the…

CL#46233751 (f87999) Add project setting to switch Slate material shaders between fully static specializations, hybrid and fully dynamic. This allows for a balance between GPU performance and the number of shaders compiled / PSOs created.

slate.MaterialShaderMode=1

CL#47605509 (7d7454) Translucent Materials no long compile the non-skylight version of base pass pixel shader to reduce shader permutations and instead always use the skylight version w/ a dynamic branch.

CL#47744412 (2b50c5) Refactor and combine some ShadowDepth permutations to reduce permutation explosion

- Standardize and move the relevant shader logic under “SUPPORTS_XYZ” defines, and dynamic branches on ShadowDepthType

- Renaming some functions to make it clear whi…

CL#45895072 (5ea2c0) Strip Nanite vertex shader HW rasterization from certain platforms by default (reducing 50% of HW raster shader counts). Re-enable with r.Nanite.StripVertexShaders=0 and recooking

CL#48654180 (06ad62) Add a VS specific material Uniform Buffer

Use r.ShaderCompiler.UseMaterialVSUniformBuffer=0 and r.MaterialEditor.AllowSkipAlphaComputationForMaterialBlend=0 to restore the previous behaviour

CL#48668307 (0308b1) Conditionally removed Instanced Static Mesh Vertex Factory for platforms supporting full GPU-Scene (r.InstancedStaticMeshes.UseInstancedStaticMeshVertexFactory, on by default).

CL#50053320 (20a6c6) Allow for stripping Nanite mesh shaders from the build (if desired by the title)

To do this set r.Nanite.AllowMeshShaders=0

CL#49201554 (1de844) Allow shader deduplication for frequencies that access only vertex material properties.

This includes Mesh shaders, and some Nanite rasterization Compute shaders. These last ones must not use the programmable pixel path (used by masked materials and…

CL#47316522 (849b80) Add Material Instance Usage Flag overrides.Moved usage flag setter/getter functions from UMaterial to virtual on UMaterialInterface.

UMaterialInstance versions are based on the contained FMaterialInstanceBasePropertyOverrides.

Added a couple of missing EMaterialUsage entries for bUsedWithEditorCompositing and bUsedWithNeuralNetworks.

Marked bUsedWith… properties as deprecated on UMaterial since any direct access to them is a potential bug as it ignores per instance overrides. After a deprecation period we should probably convert these to a private uint32 bitmask.

Fixed up places in code that accessed these properties directly.

[Attachment Removed]

I accidentally posted my reply as a new answer instead of a reply to this one, so not sure if you get pinged about it normally. Please see below.

Also, while I’ve got you, in another part of that same long thread, they mentioned hitting this D3D debug layer error, which I am also seeing occasionally on the same resource when Async Compute is enabled:

LogD3D12RHI: Error*: [D3DDebug] [ID: 1047] ID3D12CommandQueue1::UpdateTileMappings: Simultaneous-access or Buffer Resource (0x000002CEC7AC1500:‘Nanite.StreamingManager.ClusterPageData’) is still referenced by write|transition_barrier GPU operations in-flight on another Command Queue (0x000002CC6CAF94D0:‘3D Queue (GPU 0)’). It is not safe to start tilemapping GPU operations now on this Command Queue (0x000002CC6CAF8360:‘Compute Queue (GPU 0)’). This can result in race conditions and application instability.*

I don’t see any follow-up to that particular subthread of discussion, but this sounds like the sort of thing that could cause random GPU crashes down the line. Do you know if any fixes for this have been submitted in the interim, or alternatively if this can be safely ignored?

In our case, this could be because we have Scene Capture passes running before the main render pass. I know we’ve had Nanite streaming issues with this before, which required some code changes: [Content removed]

It’s possible that this is also related. I’ll try running with those disabled and see if I still see the D3D errors.

[Attachment Removed]

UPDATE: We decided to address this warning in CL#55750673 and reported it to MS as a debug layer bug.

We’re also seeing this occasionally, likely from:

30495370 (22bc8f7) Perform tile mapping updates on the direct queue when the current queue does not support the operation

We’re considering changing FD3D12DynamicRHI::UpdateReservedResources You’re welcome to test with that, however, we haven’t had time to do much testing and consider the implications of it - I’ve let the devs involved know you’re looking at this too.

That said - I’ve also found discussions from a year ago where we looked into this and determined it was a debug layer error and the transitions from the engine were OK. We used -rhivalidation -rhivalidationlog=Nanite.StreamingManager.ClusterPageData, custom logging of barriers/transitions and PIX captures of the repro case (was happening in City Sample, the MegaLights demo, and Nanite engine test levels).

What’s happening is we execute a command list on the graphics pipe that does SRV -> UAV. We signal Gfx -> Async Compute. The tile commit happens on async. We then fence back from Async -> Graphics. Then do the UAV -> SRV transition, and then the resource gets read by graphics.

[Attachment Removed]

Thanks for providing those extra breadcrumbs. Seeing crashes in RT and BuildHZB and knowing that a larger shader heap helps, suggests the RT shaders are pushing it over the limit or driver shader optimization (constant folding) like we were seeing before - seeing the crashes in other areas like post process effects doesn’t seem as likely to be related.

I suspect your repro with -gpuvalidations increased the shader heap usage enough to cause the crash locally, and increasing the heap size alleviates the issue. Do you have a setup/replay you can run in a loop and monitor PSOs in memory (PSO_TRACK_CACHE_STATS 1) to see if they’re leaking? If those aren’t leaking then it’s likely the driver optimized shaders that are causing the heap overflow, or there’s some place where you get the perfect storm of number of shaders, shader complexity and heap fragmentation that causes it to exceed the limit. If your PSO in memory count does keep climbing then there might be an missing engine fix.

Do you intend to upgrade to a newer engine version soon?

If you provide Nvidia you’re repro they can dump shader heap usage at runtime and see if your PSO count stays constant and the optimized shader count keeps going up causing the crash. Unfortunately, we don’t get much insight into the shader heap usage.

It sure would be nice to get better error reporting when there’s a shader heap related crash, right now the closest clue we have is the absence of an Aftermath dump.

[Attachment Removed]

I can definitely check if we are leaking PSOs. We may not be, however - we have an unhealthy number of shaders, and we try to keep as many PSOs in memory as we can to avoid runtime PSO compilation hitches, which can be significant.

I talked with our PSO system owner ([mention removed]​) and he let me know that we have the following settings enabled when we detect Nvidia GPUs, which likely greatly increase our PSO memory load over a default project. It’s on Nvidia GPU/driver that we’re seeing this issue, so that tracks.

r.PSOPrecache.KeepInMemoryUntilUsed = 1
r.PSOPrecache.KeepInMemoryGraphicsMaxNum = 2500
r.PSOPrecache.KeepInMemoryComputeMaxNum = 5000

[Attachment Removed]

Thanks for the quick response and CL! Just integrated along with a change from a separate EPS thread that also got answered yesterday, and I’ll test this out now.

[Attachment Removed]

Thanks! I’ve forwarded this on internally, and we’ll check it out.

[Attachment Removed]