Unpredictable nDisplay + Switchboard node startup behavior in UE5.6

Hi everyone,

I’m running into an issue with an nDisplay setup in Unreal Engine 5.6 and was hoping someone might have run into something similar.

I’m working on a custom project using nDisplay with Switchboard, and when I try to launch multiple nodes, the behavior is inconsistent:

  • Sometimes all nodes start correctly

  • Other times, one or more nodes freeze during startup and show a black window

When the nodes get stuck, only some of them show the following as last log line:

[0] LogDisplayClusterEngine: Display: CheckGameStartBarrier - we are no longer out of sync. Restoring Play.

A few additional observations:

  • The more nodes I attempt to launch simultaneously, the less likely they are to start successfully

  • This happens both when Switchboard’s “Launch As” parameter is set to Standalone or Packaged Game

  • It also doesn’t seem to matter whether I use a local installation of Switchboard or a remote one

  • Sync Policy does not affect the outcome (currently set to None)

So far, I haven’t been able to identify a clear pattern or bottleneck.

Has anyone experienced similar synchronization/startup issues with nDisplay in 5.6? Any ideas on where to start debugging this would be greatly appreciated.

Thanks in advance!

that log line is the cluster barrier — nodes exchange sync messages and the barrier decides when everyone proceeds. some nodes printing “restoring play” while others hang black means the group split: part of the cluster is past the barrier, the rest never joined, and with your current settings nothing times out to recover. the pattern (worse with more nodes, random which one hangs) matches startup races, not a config error.

what helps on similar setups, in order: (1) prewarm ddc/shaders — first launches on fresh machines compile shaders and miss ddc hard; one node stalls minutes in compile while the others sit at the barrier. run every node once manually so derived data exists, then switchboard launches become uniform. (2) stagger launches instead of starting all nodes in the same instant — startup saturates disk/network and some nodes miss the join window; two or three waves tells you instantly if it was a race. (3) when a node freezes, read that node’s own log, not the primary’s — the last lines show whether it is still in barrier wait (sync problem) or died in compile/driver (different fix). (4) in the ndisplay config, set the primary explicitly rather than auto, and look at the cluster sync timeout — too tight on a lan with many nodes produces exactly this randomness. (5) boring but real: identical gpu drivers on all nodes, fullscreen optimizations off, cluster ports allowed in firewall.

if it still hangs after prewarming + staggering, capture the stuck node’s log at the freeze — the last 20 lines decide which of these two problems you actually have.