Performance of passing data via Niagara user parameters?

Hi :waving_hand: ,

I’m building a GPU Niagara system using custom HLSL in a Niagara scratch pad, which is doing a bit math per particle using data from five vector arrays I pass to Niagara from c++. The arrays contain data for each particle that doesn’t change during the simulation (loaded and calculated in c++). There can be many thousands (even over a million on PC) of particles, so I want it optimised as possible as it needs to run on lower end hardware like a Quest 3. It runs quite well, but I can’t find much documentation on how User parameters are processed and I’m wondering if this could be slowing it down. Does anyone know if I would get performance gains using a different method to pass and read data, such as niagara data interfaces and structured buffers or texture lookups? Here is a simplified example of what I’m doing:

TArray<FVector> SomeData; // per particle data saved as vector
UNiagaraDataInterfaceArrayFunctionLibrary::SetNiagaraArrayVector(NiagaraComponent, FName("ArrayVector"), SomeData); // pass data to niagara user paramter

thanks!

Hi there @Evman!

I did a lot of digging in the Cpp files to double check, and while my Cpp understanding is not perfect, it looks like you’re good.

It looks like while the Niagara parameters are technically their own class, it’s really just some metadata and a variable.

Fron NiagaraTypes.h:

I would suggest checking out this header for yourself, as you will probably be able to understand it better than me.

Overall, if you want to save performance, there are probably better places to be looking. I suggest starting with this page:

the cost model for user array parameters into a GPU sim, roughly:

per frame array upload: every frame the sim runs, the array parameter data is uploaded to the GPU (unless the system and parameter support data aging, which only skips when nothing changed). five arrays over a million floats is about 20MB per frame at 4 bytes per float, which saturates the upload budget and shows up as copy stalls. that dominates everything else in your loop.

what does not cost: particle reads of the array inside scratch are cheap once resident (random access into a GPU buffer), so the sim side is fine at any count. the transfer is the bottleneck, not the math.

ways to cut the transfer:

only set when changed: do not set the array parameter every tick from c++. set it on load and whenever the source data actually changes.

data interface instead of user param: for static data, a DI backed by a structured buffer uploads once. if your five arrays derive from one source (a mesh, a texture), sample that directly in scratch and skip arrays entirely. a common trick: bake the five streams into a volume texture or render targets once at load, then the sim samples textures, zero per-frame uploads.

half precision: if values tolerate it, packing into half2/half4 textures halves the transfer versus float arrays.

array size caps: the whole array is uploaded even if the sim reads a slice, so per-frame growing arrays are the worst case. for dynamic counts, keep the max fixed and write a count parameter.

so: static data = DI or texture, set once. dynamic per-frame data = minimize bytes (halfs, fewer streams) and expect the upload to be the cost center, not the sim.