There are a couple conversations about an elusive Lumen Reflections crash in the DXR ray traversal. I found a cause and fix, I’m posting here for validation and in case it helps.
Setup
AMD RX 7000 line (RDNA3), Adrenalin 26.8.1.
Could not repro on Nvidia, RDNA4.
This was on 5.6.1, but there is no evidence of a fix on 5.8.3 (in perforce or in UDN).
Failure
We kept getting trash in the SBT at record 0 (0s, unrelated data, stale bytes from previous SBTs).
Lumen Reflections trying to perform hit lighting would fail trying to run any-hit shader for record 0 with an instruction-fetch page fault. On our setup hit lighting is only used for reflections, hence we got the crash only there.
Diagnostic
Sometimes, SBT copy-queue upload is incorrect on the 24 first bytes if the buffer is not aligned on cache lines.
- The CPU buffer location was always 104 in a 256 bytes line or 40 bytes in a 64 bytes line. Meaning, problematic bytes sat on the last 24 bytes of a 64 bytes boundary.
- CPU side copy of the buffer remained correct.
- I validated the copy completed (all bytes but the 24 first bytes always matched exactly). I also added a fence between the CPU Memcpy to source buffer and the CopyBufferRegion, this did not fix the issue.
- At times, the GPU SBT copy showed previous frame’s SBT head data on the 24 first bytes (I added a unique ID for testing).
Underlying Cause
I suspect a bug from AMD.
- Stale bytes sometimes showing previous SBT data points away from neighbour memory trashing.
- Not a completion race : only the first 24 bytes were incorrect.
- Follows the source buffer memory layout. The GPU destination buffer is 64 byte aligned, yet the 24 bytes corruption matches the alignment of the source CPU buffer which sat on byte 40 of a 64 line.
- RDNA 4 vs RDNA 3 difference in behaviour.
- Using the graphics queue instead of the copy queue fixes the issue as well.
This points toward an AMD copy-queue issue.
Fix
More of a work around. I align the buffer on 256 bytes in FD3D12Buffer::UploadResourceDataViaCopyQueue
void* pData = GetParentDevice()->GetDefaultFastAllocator().Allocate(BufferSize, 256UL, &SrcResourceLoc);
Crash is gone.
Questions
- Can you confirm this is not fixed as of 5.8.3?
I would be able to test in 5.8.3 quite soon if you can’t answer easily. - Can you validate that this work around is correct?
It is my understanding that I don’t waste significant memory, and that generally speaking it is a faster transfer when aligned. - Do you think my hypothesis of a AMD issue is valid?
Thanks!
Previous discussions without definitive answers.
They are closed, I’m not sure they will get notified by mentioning them here.