UE 5.8 editor FPS regression vs 5.3. when r.Lumen.DiffuseIndirect.Allow=1 (ISM mannequin grid; Game thread!44ms)

We are seeing a large editor Lit viewport performance regression in UE 5.8 vs UE 5.3 when Lumen GI is enabled via r.Lumen.DiffuseIndirect.Allow=1.

What we tested

Same Instanced Static Mesh mannequin-grid Lit scene in both versions (projects attached). Viewport Lit + realtime. Toggle only:

  • r.Lumen.DiffuseIndirect.Allow 1
    • [Image Removed]
  • r.Lumen.DiffuseIndirect.Allow 0
    • [Image Removed]

Ask

Is this expected 5.3→5.8 Lumen cost increase, or a regression? We need Lumen on but editor FPS closer to 5.3 for dense ISM/crowd content. Please advise a recommended scalability/CVar profile, and whether there are known ISM + Lumen Game-thread issues in 5.8.

Attachments

  • UE 5.3 + UE 5.8 repro projects (Config + Content + .uproject)

[Attachment Removed]

Steps to Reproduce
Steps:

  1. Open the attached UE 5.3 project (Lumen_53.ISM) in Unreal Engine 5.3.2.
  2. Open the level Lv_ISM_LUMEN_Test in Test Folder
  3. In the console, Run: r.Lumen.DiffuseIndirect.Allow 1
    1. Note FPS/Frame/Game/Draw/GPU.
  4. Run: r.Lumen.DiffuseIndirect.Allow 0
    1. Note the same stats again.
  5. Repeat steps 1-4 in Unreal Engine 5.8 with the attached UE5.8 project (same level / same ISM mannequin grid content).

Results Tested:

  • In 5.8, Allow = 1 -> drops to ~22 FPS with Game ~44ms; Allow = 0 recovers to ~92 FPS. (Increase grid count)
  • In 5.3 Allow = 1 -> cost around 15 FPS
 r.Lumen.DiffuseIndirect.Allow 1

[Attachment Removed]

Hi there,

It’s difficult to determine what might have changed in Lumen between 5.3 and 5.8, as that would require reviewing years’ worth of changes. Overall, we have made great progress in making Lumen more performant over time. From the screenshot you shared, the main difference in performance is the time spent on the Render Thread (Draw) and the GPU (GPU Time). The FPS counter is also not a reliable way to measure perf, and I suggest you profile your title with Unreal Insights, the CSV profiler, or your preferred profiler. You can turn off async compute (r.RDG.AsyncCompute=0) and profiling both the 5.3 and 5.8 projects. That should give you a good overview of the total per-frame cost of your frame for either engine version. You will then be able to determine whether Lumen is the main bottleneck and tweak your cvars to restore performance to your desired levels. I noticed that your project configurations have different cvar setups and I have called out a few Lumen-related ones which you could tweak to get your numbers down:

r.LumenScene.SurfaceCache.MeshCardsMergeInstances=1
r.LumenScene.SurfaceCache.MeshCardsMinSize=<raise>   ; drop small crowd cards
; step down from editor-default Epic(3) to High(2)
sg.GlobalIlluminationQuality=2
r.Lumen.ScreenProbeGather.StochasticInterpolation=2  ; half-res gather
r.Lumen.ScreenProbeGather.DownsampleFactor=32
; slower scene relighting = cheaper
r.LumenScene.DirectLighting.UpdateFactor=64
r.LumenScene.Radiosity.UpdateFactor=128
; HWRT crowd culling
r.RayTracing.Culling=3
r.RayTracing.Culling.Radius=15000
MeshCardsMergeInstances=1

I hope that helps you get started, but please feel free to let me know if you have any more questions.

Cheers,

Tim

[Attachment Removed]

Hi Tim,

I ran more targeted test and found a very specific discrpancy between two versions that I would like your help understanding. We are only using unreal for rendering so we are not trying to get the highest frame here.

Test: Identical level, Identical Blueprint spawning 900 ISM mannequin instances. Measure with `stat rendering`.

Unreal 5.3:

  • r.LumenScene.SurfaceCache.MeshCardsMergeInstances 0 -> Baseline FPS
  • [Image Removed]
  • r.LumenScene.SurfaceCache.MeshCardsMergeInstances 1 -> No measurable changes in FPS or frame time, but Mesh draw calls is halfed.
  • [Image Removed]

Unreal 5.8:

  • r.LumenScene.SurfaceCache.MeshCardsMergeInstances 0 -> ~10 fps / 95ms

  • [Image Removed]

  • r.LumenScene.SurfaceCache.MeshCardsMergeInstances 1 -> ~29fps / 35ms

  • [Image Removed]

Key observation:

Mesh draw calls remain identical in all cases. The performance difference is entirely within Lumen’s CPU-side Surface Cache update.

This seems to show that in 5.8, Lumen generated per-instance mesh cards for ISM actually get considered while in unreal 5.3 toggle on and off MeshCardMergeInstances doesnt really impact on the fps or farmes. We are not nanite with those mannequins for testing because they are replying on Vertex Animation Texture to playback those animations.

My Questions:

  1. Was MeshCardMergeInstances=1 has any impact on the visual of the instance static meshes, should this only be used to get better performance in editor and toggle it off for final pixel rendering?
  2. Is this considered a known performance regression for dense non-Nanite ISM contnet?
  3. Are the observations by deisgn, or is there a recommanded way for this kind of work in editor, if PCG also uses instance static mesh component to spawn meshes, wont that cause issues. And I thought instance static meshc compone would only load the mesh once and being very light?

Cheers

Wei

[Attachment Removed]

Hi Wei,

Thanks for the extra information, which helped a lot. I reproduced your setup locally on 5.8 with the 900 instance grid, and I want to correct one thing before answering your questions, because it changes what to tune.

The cost is on the GPU, not in Lumen’s CPU-side surface cache update. On my machine (RTX A5000, 1280x720, 900 instances, MeshCardsMergeInstances 0), the median frame is 77.1 ms, render thread 77.1 ms, GPU 76.3 ms, game thread 7.0 ms. The render thread time is tracking GPU time almost exactly, which means the render thread is blocking on the GPU rather than doing work. The InitViews number you are reading in stat rendering is that stall, not CPU work inside the surface cache update. It is worth noting too that your screenshots are from a PIE session rather than the editor Lit viewport, and the game thread in them is not the bottleneck either.

One more thing that affects your comparison: the two projects are not configured the same. Your 5.8 project sets r.Substrate=True in DefaultEngine.ini and your 5.3 project does not, and 5.3 has no Substrate at all. Substrate changes the material and GBuffer path that Lumen uses to capture surface cache cards, which is exactly the work that scales with the number of mesh card sets in your scene. Before we treat any of the remaining gaps as a regression, could you re-run the A/B with Substrate matched, either off in the 5.8 project or on in both, and send a GPU profile of the 5.8 case? That will tell us how much of the delta is Substrate and how much is Lumen.

On your questions.

MeshCardsMergeInstances=1 does change your visuals, so please do not treat it as an editor-only switch. It replaces the per-instance cards with a single card set covering the whole component’s bounds, and it applies a 0.3x resolution scale on top (r.LumenScene.SurfaceCache.MeshCardsMergedResolutionScale). Bounce lighting and reflections from the crowd become much coarser and more prone to leaking. It also only kicks in when the merged bounds are under 10000 units on any axis (r.LumenScene.SurfaceCache.MeshCardsMergedMaxWorldSize) and the combined-to-summed surface area ratio is under 1.7 (r.LumenScene.SurfaceCache.MeshCardsMergeInstancesMaxSurfaceAreaRatio), so it will quietly stop helping as your grid spreads out. If the quality is acceptable for your final pixels, it is fine to leave it on for both, and that is generally the better choice than flipping it between preview and final.

That said, the behavior you are seeing is by design in one important respect: Lumen’s surface cache is per-instance for Instanced Static Meshes in both 5.3 and 5.8. Your intuition about ISM is right for rasterization, where the mesh is shared and instances are cheap, but Lumen’s surface cache does not work that way. 900 instances mean 900 mesh card sets and 900 sets of card captures, and that is the cost you are paying. MeshCardsMergeInstances is the lever that collapses them, which is why it is the only cvar that makes this step cheaper for you.

What changed in 5.8 is that we added surface cache card sharing (r.LumenScene.SurfaceCache.AllowCardSharing, enabled by default), which is exactly the optimization for large numbers of identical instances: matching cards share a single atlas allocation instead of N captures. Unfortunately, your content cannot use it for two separate reasons. It requires a Nanite proxy because the sharing key is built from the Nanite resource ID and shading bins, and it is additionally disabled for any material using PerInstanceRandom, PerInstanceCustomData, or WorldPosition (r. LumenScene.SurfaceCache.DetectCardSharingCompatibility). AnimToTexture drives the animation frame through PerInstanceCustomData, so your VAT mannequins fail both gates. That is why 5.8 does not give you a win here, even though the machinery for it exists.

As for PCG, since PCG normally scatters regular Nanite static meshes, and those do get card sharing. The pattern that hurts is specifically dense non-Nanite instances with per-instance material data, which is what VAT crowds are.

So for your case the practical options are to keep MeshCardsMergeInstances=1 if the quality holds up, raise r.LumenScene.SurfaceCache.MeshCardsMinSize so small crowd instances drop out of the surface cache entirely, and lower r.LumenScene.SurfaceCache.CardCapturesPerFrame from its default of 300 to spread the capture cost over more frames. Longer term, if any of those characters can move to Nanite with the animation driven by something other than per-instance custom data, card sharing would do this properly instead of you trading away quality for it.

Cheers,

Tim

[Attachment Removed]