Server memory leak in 5.6: ~FRepChangelistState calls StaticBuffer.Empty() before member destruction, causing FRepStateStaticBuffer to skip DestructProperties and leak nested heap allocations of shadowed replicated properties (FastArrays most visibly)

Summary

FRepChangelistState::~FRepChangelistState() calls StaticBuffer.Empty() in its body. FRepStateStaticBuffer::Empty() only empties the underlying byte buffer (sets Buffer.Num() to 0) and does NOT destruct the shadowed properties. The subsequent member destruction of FRepStateStaticBuffer (~FRepStateStaticBuffer) is guarded by "if (Buffer.Num() > 0)" and therefore skips FRepLayout::DestructProperties. As a result, the nested heap allocations owned by the shadowed property copies (nested TArray/FString/TMap, FastArray element structs) are never freed and leak.

Mechanistically this affects every replicated property whose shadow copy owns nested heap allocations, not only FastArrays — the shadow buffer is deep-copied for all parents (CopyCompleteValue) and DestructProperties iterates all parents. FastArrays dominate because their custom-delta shadow holds per-element copies, elements are numerous, and element structs frequently contain nested containers.

Regression origin

git blame attributes the destructor's StaticBuffer.Empty(), the new FRepStateStaticBuffer::Empty() method, and the trailing Buffer = {} in ~FRepStateStaticBuffer all to CL 905d6e274a57d ("Protect against FRepChangelistState double-free"). Before that change, ~FRepChangelistState did not touch StaticBuffer in its body and member destruction ran DestructProperties normally. The intent appears to be double-free hardening; our reading is that StaticBuffer.Empty() was expected to clear the buffer the way Empty() conventionally does in the engine (e.g. FScriptArrayHelper::EmptyValues destructs then empties), but this Empty() does not destruct.

Proposed fix (PR https://github.com/EpicGames/UnrealEngine/pull/15106)

Remove the StaticBuffer.Empty() call from ~FRepChangelistState, letting member destruction take its normal path (~FRepStateStaticBuffer sees Buffer.Num() > 0 and calls DestructProperties). Double-free resilience is preserved by the Buffer = {} in ~FRepStateStaticBuffer (on a second destruction Buffer.Num() is 0, so DestructProperties is naturally skipped and Buffer = {} is idempotent).

Verified still present on latest release and ue5-main.

Questions

1. Can you confirm this is a regression from CL 905d6e274a57d, and that FRepStateStaticBuffer::Empty() intentionally does not destruct properties?
2. Is the intended fix to drop the StaticBuffer.Empty() call, to make Empty() destruct, or something else? We want to align to minimize future merge friction.
3. Any guidance on the double-free scenario the original CL guarded against, so we can regression-test our change does not reintroduce it?

figure.zip(1.14 MB)
memoryinsights.zip(82 MB)

Steps to Reproduce

Repro (dedicated server, legacy replication, non-Iris; happens with or without ReplicationGraph):

1. Run a dedicated server that replicates objects using FastArrays with elements that own nested heap allocations (nested TArray/FString/TMap, or nested FastArrays). In our project these are inventory/loadout-style components.
2. Have players repeatedly log in and out (each brings up and tears down their FastArray-backed components).
3. Watch server memory over time (LLM tag Networking/NetObjReplicator, or Memory Insights).

Observed:
- Memory grows roughly monotonically with uptime / total sessions, independent of concurrent player count. It does NOT recede when players leave, and does NOT recede even with the server empty after a forced GC. (Fig 1)

- LLM localizes the growth to Networking/NetObjReplicator/Replicate/CustomDelta/SharedShadow — the FastArray custom-delta shadow state. (Fig 2)
- Memory Insights (Memory Leaks rule) shows the leaked allocations on the FastArray delta path: FScriptArrayHelper::AddValues -> ResizeGrow -> ResizeAllocation, under CompareProperties_Array_r -> FRepLayout::DeltaSerializeFastArrayProperty -> FNetSerializeCB::NetDeltaSerializeForFastArray -> FFastArraySerializer::FastArrayDeltaSerialize. (Fig 3)
- ~50-60 KB leaks per player login/logout cycle; spans multiple FastArray types, not just one.

Root cause confirmed locally with a debugger: on teardown, the DestructProperties() call in FRepStateStaticBuffer::~FRepStateStaticBuffer() is never hit, because Buffer.Num() is already 0.

A diagnostic dump walking all live replication containers from the NetDriver (AllOwnedReplicators, ReplicationChangeListMap, ...) and summing CountBytes stays around ~10 MB while the LLM tag is multi-GB — a several-hundred-fold divergence at the same instant. Process RSS confirms it is real memory (Untracked ~0), i.e. orphaned allocations no live container points at anymore.

One note on scope, in case it helps prioritization:

This is a legacy-replication-only issue. It lives in FRepChangelistState / FRepStateStaticBuffer / DeltaSerializeFastArrayProperty, which Iris does not use — under Iris, FastArrays go through IrisFastArraySerializer / FastArrayReplicationFragment and never touch this shadow buffer, so Iris is unaffected. The same is true for other content that doesn’t rely on custom-delta FastArrays.

We fully understand Iris is the future of UE networking and that Legacy/ReplicationGraph will eventually be phased out — we’re moving toward Iris ourselves. That said, there are still a large number of shipping projects and live dedicated servers running on the legacy path today, and for them this leak accumulates unbounded over uptime. Given the fix is a single-line change with the double-free protection preserved, we hope it can still be merged to help those legacy deployments.

Hi,

Thank you for the report, pull request, and all the detailed information!

I can confirm your PR has been received in our internal tracker, and the PR will be routed to the appropriate subject matter expert for review. However, I cannot provide an estimate as to when they will be able to review the change. In the meanwhile, I’ve linked to the info you’ve provided here on our internal tracker for visibility, and if we have further questions, someone may reach out either on this thread or on the PR itself.

To answer your questions:

1) Yes, I believe it is intentional that FRepStateStaticBuffer::Empty() does not destroy properties, as any dynamic memory allocated by the properties is allocated elsewhere, with the buffer just holding pointers to this memory (see comments above FRepStateStaticBuffer).

2) We’ll review your PR as soon as possible to determine if this is the correct fix for this issue.

3) The original issue was found due to a rare crash in an internal project, and the underlying cause of that double-free was not able to be identified. This unfortunately does make it difficult to provide guidance on how to regression test your proposed fix. However, if you do run into any issues with your change, please let us know either here or on the PR.

Thanks,

Alex

Thanks Alex, that’s very helpful — especially the context on the original double-free.

On point 1: agreed that FRepStateStaticBuffer::Empty() intentionally only releases the byte buffer, since the properties’ dynamic memory lives elsewhere and the buffer just holds pointers.

To be clear, our PR does not change Empty()'s behavior. The change is in ~FRepChangelistState:

it calls StaticBuffer.Empty(), which sets Buffer.Num() to 0, so the subsequent member destruction ~FRepStateStaticBuffer() — guarded by if (Buffer.Num() > 0) — skips DestructProperties(). That’s exactly the path that would free the “elsewhere” allocations held by the shadowed properties (e.g. nested TArrays inside FastArray items), so they leak. Removing the Empty() call lets member destruction run DestructProperties() normally; the double-free hardening is still provided by the Buffer = {} in ~FRepStateStaticBuffer().

One more update that may be useful: we’ve applied the same change locally and are rolling it out to our production environment over the next couple of days to validate it at scale. If it’s helpful, we’re happy to keep reporting back on the results — we’ll be watching both the live LLM tags and the DS process-level memory (RSS) on our pods over time.

The following shows the actual memory leak status of the previous DS process in a production environment, and the tags in Figure 2 were added by us with custom refinements.

Thanks again for routing this internally.

Hi,

Thank you for the additional information!

It would definitely be helpful to know the results of this change in a production environment, as this can provide info as to whether it addresses the problem as well as if it introduces any new problems.

Thanks,

Alex

Hi Alex,

Thanks again for the earlier response. You mentioned production results would help show both whether the change addresses the problem and whether it introduces any new problems — so here’s what we’ve seen after deploying it to our production environment.

Does it address the problem — yes.

Before the fix, the

Networking/NetObjReplicator

LLM tag grew monotonically to ~16 GB over ~7 days of uptime and never receded, even with the server empty. After the fix, in the same production environment, the tag stays in the low tens of MB and rises/falls with player login/logout as expected — the shadow-buffer memory is now released correctly on teardown. We’re also watching the DS process RSS at the pod level, and it has stabilized accordingly.

Does it introduce new problems — none observed so far.

No new crashes or instability on the object-teardown / disconnect / travel paths (where the original double-free concern would live). You noted the underlying cause of that original double-free was never identified, so we’ve been paying particular attention there, and haven’t reproduced any issue. Just as importantly, there are no anomalies in actual gameplay or business logic — players’ inventory/loadout/mod data and overall experience behave exactly as before. The change is purely a memory-lifetime fix with no observable behavioral side effects.

We’ll keep it running and monitoring over a longer window in production, and will report back if anything changes.

Hi,

Thank you for the update, and I’m glad the fix is working for you!

This will be useful information to have when reviewing the pull request, and I’ve included this on our internal tracker for the PR.

Thanks,

Alex