Hi Wei,
Thanks for the extra information, which helped a lot. I reproduced your setup locally on 5.8 with the 900 instance grid, and I want to correct one thing before answering your questions, because it changes what to tune.
The cost is on the GPU, not in Lumen’s CPU-side surface cache update. On my machine (RTX A5000, 1280x720, 900 instances, MeshCardsMergeInstances 0), the median frame is 77.1 ms, render thread 77.1 ms, GPU 76.3 ms, game thread 7.0 ms. The render thread time is tracking GPU time almost exactly, which means the render thread is blocking on the GPU rather than doing work. The InitViews number you are reading in stat rendering is that stall, not CPU work inside the surface cache update. It is worth noting too that your screenshots are from a PIE session rather than the editor Lit viewport, and the game thread in them is not the bottleneck either.
One more thing that affects your comparison: the two projects are not configured the same. Your 5.8 project sets r.Substrate=True in DefaultEngine.ini and your 5.3 project does not, and 5.3 has no Substrate at all. Substrate changes the material and GBuffer path that Lumen uses to capture surface cache cards, which is exactly the work that scales with the number of mesh card sets in your scene. Before we treat any of the remaining gaps as a regression, could you re-run the A/B with Substrate matched, either off in the 5.8 project or on in both, and send a GPU profile of the 5.8 case? That will tell us how much of the delta is Substrate and how much is Lumen.
On your questions.
MeshCardsMergeInstances=1 does change your visuals, so please do not treat it as an editor-only switch. It replaces the per-instance cards with a single card set covering the whole component’s bounds, and it applies a 0.3x resolution scale on top (r.LumenScene.SurfaceCache.MeshCardsMergedResolutionScale). Bounce lighting and reflections from the crowd become much coarser and more prone to leaking. It also only kicks in when the merged bounds are under 10000 units on any axis (r.LumenScene.SurfaceCache.MeshCardsMergedMaxWorldSize) and the combined-to-summed surface area ratio is under 1.7 (r.LumenScene.SurfaceCache.MeshCardsMergeInstancesMaxSurfaceAreaRatio), so it will quietly stop helping as your grid spreads out. If the quality is acceptable for your final pixels, it is fine to leave it on for both, and that is generally the better choice than flipping it between preview and final.
That said, the behavior you are seeing is by design in one important respect: Lumen’s surface cache is per-instance for Instanced Static Meshes in both 5.3 and 5.8. Your intuition about ISM is right for rasterization, where the mesh is shared and instances are cheap, but Lumen’s surface cache does not work that way. 900 instances mean 900 mesh card sets and 900 sets of card captures, and that is the cost you are paying. MeshCardsMergeInstances is the lever that collapses them, which is why it is the only cvar that makes this step cheaper for you.
What changed in 5.8 is that we added surface cache card sharing (r.LumenScene.SurfaceCache.AllowCardSharing, enabled by default), which is exactly the optimization for large numbers of identical instances: matching cards share a single atlas allocation instead of N captures. Unfortunately, your content cannot use it for two separate reasons. It requires a Nanite proxy because the sharing key is built from the Nanite resource ID and shading bins, and it is additionally disabled for any material using PerInstanceRandom, PerInstanceCustomData, or WorldPosition (r. LumenScene.SurfaceCache.DetectCardSharingCompatibility). AnimToTexture drives the animation frame through PerInstanceCustomData, so your VAT mannequins fail both gates. That is why 5.8 does not give you a win here, even though the machinery for it exists.
As for PCG, since PCG normally scatters regular Nanite static meshes, and those do get card sharing. The pattern that hurts is specifically dense non-Nanite instances with per-instance material data, which is what VAT crowds are.
So for your case the practical options are to keep MeshCardsMergeInstances=1 if the quality holds up, raise r.LumenScene.SurfaceCache.MeshCardsMinSize so small crowd instances drop out of the surface cache entirely, and lower r.LumenScene.SurfaceCache.CardCapturesPerFrame from its default of 300 to spread the capture cost over more frames. Longer term, if any of those characters can move to Nanite with the animation driven by something other than per-instance custom data, card sharing would do this properly instead of you trading away quality for it.
Cheers,
Tim
[Attachment Removed]