@woosh/meep-engine 3.18.0 → 3.19.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (120) hide show
  1. package/build/bundle-worker-image-decoder.js +1 -1
  2. package/editor/view/particles/effect/ParticleCurveEditorView.d.ts.map +1 -1
  3. package/editor/view/particles/effect/ParticleCurveEditorView.js +303 -120
  4. package/editor/view/particles/effect/ParticleGradientEditorView.d.ts.map +1 -1
  5. package/editor/view/particles/effect/ParticleGradientEditorView.js +108 -63
  6. package/editor/view/particles/effect/ParticleGraphEditorView.js +1 -1
  7. package/editor/view/particles/effect/ParticleNodeParametersView.d.ts.map +1 -1
  8. package/editor/view/particles/effect/ParticleNodeParametersView.js +3 -1
  9. package/editor/view/particles/effect/particle-editor.css +60 -30
  10. package/package.json +1 -2
  11. package/samples/engine/README.md +1 -1
  12. package/src/core/binary/compression/decompress_bytes.d.ts +13 -0
  13. package/src/core/binary/compression/decompress_bytes.d.ts.map +1 -0
  14. package/src/core/binary/compression/decompress_bytes.js +28 -0
  15. package/src/core/model/node-graph/visual/layout/layout_assign_coordinates.js +67 -22
  16. package/src/engine/asset/loaders/image/ImageDecoderWorker.js +12 -27
  17. package/src/engine/asset/loaders/image/prototypePNG.js +8 -7
  18. package/src/engine/ecs/storage/populateEngineSerializationRegistry.d.ts.map +1 -1
  19. package/src/engine/ecs/storage/populateEngineSerializationRegistry.js +4 -0
  20. package/src/engine/graphics/ecs/particles/ParticleEffect.d.ts +266 -0
  21. package/src/engine/graphics/ecs/particles/ParticleEffect.d.ts.map +1 -0
  22. package/src/engine/graphics/ecs/particles/ParticleEffect.js +455 -0
  23. package/src/engine/graphics/ecs/particles/ParticleEffectSerializationAdapter.d.ts +58 -0
  24. package/src/engine/graphics/ecs/particles/ParticleEffectSerializationAdapter.d.ts.map +1 -0
  25. package/src/engine/graphics/ecs/particles/ParticleEffectSerializationAdapter.js +219 -0
  26. package/src/engine/graphics3/GPUParticleEmitterSystem.d.ts +178 -25
  27. package/src/engine/graphics3/GPUParticleEmitterSystem.d.ts.map +1 -1
  28. package/src/engine/graphics3/GPUParticleEmitterSystem.js +809 -310
  29. package/src/format/image/png/PNGReader.d.ts +7 -6
  30. package/src/format/image/png/PNGReader.d.ts.map +1 -1
  31. package/src/format/image/png/PNGReader.js +13 -12
  32. package/src/format/image/png/chunk/png_chunk_decode_iTXt.d.ts +3 -2
  33. package/src/format/image/png/chunk/png_chunk_decode_iTXt.d.ts.map +1 -1
  34. package/src/format/image/png/chunk/png_chunk_decode_iTXt.js +5 -4
  35. package/src/format/image/png/chunk/png_chunk_decode_zTXt.d.ts +3 -2
  36. package/src/format/image/png/chunk/png_chunk_decode_zTXt.d.ts.map +1 -1
  37. package/src/format/image/png/chunk/png_chunk_decode_zTXt.js +5 -4
  38. package/src/format/image/png/png_inflate.d.ts +3 -3
  39. package/src/format/image/png/png_inflate.d.ts.map +1 -1
  40. package/src/format/image/png/png_inflate.js +29 -39
  41. package/src/format/texture/ktx2/ktx2_read.d.ts +4 -4
  42. package/src/format/texture/ktx2/ktx2_read.d.ts.map +1 -1
  43. package/src/format/texture/ktx2/ktx2_read.js +18 -21
  44. package/src/shade/playground/particle_ecs/README.md +203 -0
  45. package/src/shade/playground/particle_ecs/bonfire_editor.d.ts +39 -0
  46. package/src/shade/playground/particle_ecs/bonfire_editor.d.ts.map +1 -0
  47. package/src/shade/playground/particle_ecs/bonfire_editor.js +315 -0
  48. package/src/shade/playground/particle_ecs/bonfire_effects.d.ts +145 -0
  49. package/src/shade/playground/particle_ecs/bonfire_effects.d.ts.map +1 -0
  50. package/src/shade/playground/particle_ecs/bonfire_effects.js +202 -0
  51. package/src/shade/playground/particle_ecs/bonfire_sprites.d.ts +23 -0
  52. package/src/shade/playground/particle_ecs/bonfire_sprites.d.ts.map +1 -0
  53. package/src/shade/playground/particle_ecs/bonfire_sprites.js +315 -0
  54. package/src/shade/playground/particle_ecs/bonfire_world.d.ts +86 -0
  55. package/src/shade/playground/particle_ecs/bonfire_world.d.ts.map +1 -0
  56. package/src/shade/playground/particle_ecs/bonfire_world.js +303 -0
  57. package/src/shade/playground/particle_ecs/effects/embers.json +1634 -0
  58. package/src/shade/playground/particle_ecs/effects/flame.json +1882 -0
  59. package/src/shade/playground/particle_ecs/effects/smoke.json +1860 -0
  60. package/src/shade/playground/particle_ecs/effects/soot.json +1606 -0
  61. package/src/shade/playground/particle_ecs/index.html +330 -0
  62. package/src/shade/playground/particle_ecs/main.d.ts +2 -0
  63. package/src/shade/playground/particle_ecs/main.d.ts.map +1 -0
  64. package/src/shade/playground/particle_ecs/main.js +566 -0
  65. package/src/shade/playground/particle_ecs/moonlit_environment.d.ts +16 -0
  66. package/src/shade/playground/particle_ecs/moonlit_environment.d.ts.map +1 -0
  67. package/src/shade/playground/particle_ecs/moonlit_environment.js +144 -0
  68. package/src/shade/playground/particle_editor/README.md +5 -4
  69. package/src/shade/playground/profile_hotkey.d.ts +58 -0
  70. package/src/shade/playground/profile_hotkey.d.ts.map +1 -0
  71. package/src/shade/playground/profile_hotkey.js +325 -0
  72. package/src/shade/playground/ssr_variance/README.md +105 -0
  73. package/src/shade/playground/ssr_variance/capture.d.ts +16 -0
  74. package/src/shade/playground/ssr_variance/capture.d.ts.map +1 -0
  75. package/src/shade/playground/ssr_variance/capture.js +96 -0
  76. package/src/shade/playground/ssr_variance/index.html +22 -0
  77. package/src/shade/playground/ssr_variance/main.d.ts +2 -0
  78. package/src/shade/playground/ssr_variance/main.d.ts.map +1 -0
  79. package/src/shade/playground/ssr_variance/main.js +246 -0
  80. package/src/shade/playground/ssr_variance/reference.d.ts +18 -0
  81. package/src/shade/playground/ssr_variance/reference.d.ts.map +1 -0
  82. package/src/shade/playground/ssr_variance/reference.js +78 -0
  83. package/src/shade/playground/ssr_variance/scene.d.ts +11 -0
  84. package/src/shade/playground/ssr_variance/scene.d.ts.map +1 -0
  85. package/src/shade/playground/ssr_variance/scene.js +61 -0
  86. package/src/shade/playground/ssr_variance/statistics.d.ts +44 -0
  87. package/src/shade/playground/ssr_variance/statistics.d.ts.map +1 -0
  88. package/src/shade/playground/ssr_variance/statistics.js +51 -0
  89. package/src/shade/renderer/loader/gltf/tiny-gltf.d.ts.map +1 -1
  90. package/src/shade/renderer/loader/gltf/tiny-gltf.js +9 -3
  91. package/src/shade/renderer/loader/usd/usd_decode_image.d.ts +4 -4
  92. package/src/shade/renderer/loader/usd/usd_decode_image.d.ts.map +1 -1
  93. package/src/shade/renderer/loader/usd/usd_decode_image.js +4 -4
  94. package/src/shade/renderer/particles/DESIGN.md +748 -703
  95. package/src/shade/renderer/particles/runtime/ParticleEmitter.d.ts +11 -4
  96. package/src/shade/renderer/particles/runtime/ParticleEmitter.d.ts.map +1 -1
  97. package/src/shade/renderer/particles/runtime/ParticleEmitter.js +431 -424
  98. package/src/shade/renderer/postprocess/ssr/SSR.d.ts.map +1 -1
  99. package/src/shade/renderer/postprocess/ssr/SSR.js +2 -1
  100. package/src/shade/renderer/postprocess/ssr/reproject/chunk_ssr_reprojection.d.ts.map +1 -1
  101. package/src/shade/renderer/postprocess/ssr/reproject/chunk_ssr_reprojection.js +0 -6
  102. package/src/shade/renderer/postprocess/ssr/reproject/chunk_ssr_sample_history.d.ts.map +1 -1
  103. package/src/shade/renderer/postprocess/ssr/reproject/chunk_ssr_sample_history.js +19 -5
  104. package/src/shade/renderer/postprocess/ssr/reproject/chunk_ssr_temporal_accumulate.d.ts +4 -0
  105. package/src/shade/renderer/postprocess/ssr/reproject/chunk_ssr_temporal_accumulate.d.ts.map +1 -0
  106. package/src/shade/renderer/postprocess/ssr/reproject/chunk_ssr_temporal_accumulate.js +133 -0
  107. package/src/shade/renderer/postprocess/ssr/ssr_reproject_shader.d.ts +0 -8
  108. package/src/shade/renderer/postprocess/ssr/ssr_reproject_shader.d.ts.map +1 -1
  109. package/src/shade/renderer/postprocess/ssr/ssr_reproject_shader.js +16 -113
  110. package/src/shade/renderer/texture/source/texel_data_from_ktx2.d.ts +2 -2
  111. package/src/shade/renderer/texture/source/texel_data_from_ktx2.d.ts.map +1 -1
  112. package/src/shade/renderer/texture/source/texel_data_from_ktx2.js +3 -3
  113. package/src/shade/renderer/particles/shaders/chunk_particle_emitter_warmup.d.ts +0 -20
  114. package/src/shade/renderer/particles/shaders/chunk_particle_emitter_warmup.d.ts.map +0 -1
  115. package/src/shade/renderer/particles/shaders/chunk_particle_emitter_world_sphere.d.ts +0 -14
  116. package/src/shade/renderer/particles/shaders/chunk_particle_emitter_world_sphere.d.ts.map +0 -1
  117. package/src/shade/renderer/particles/shaders/shader_particle_reclaim.d.ts +0 -50
  118. package/src/shade/renderer/particles/shaders/shader_particle_reclaim.d.ts.map +0 -1
  119. package/src/shade/renderer/postprocess/ssr/reproject/shader_ffx_denoiser_reflections_reproject.d.ts +0 -17
  120. package/src/shade/renderer/postprocess/ssr/reproject/shader_ffx_denoiser_reflections_reproject.d.ts.map +0 -1
@@ -1,703 +1,748 @@
1
- # GPU-Driven Particle System — Design
2
-
3
- A specialized, deeply-integrated, GPU-driven particle engine for the v2 renderer.
4
- Destiny-style: **one** simulation dispatch runs **all** particles of **all** emitters by
5
- interpreting a compact per-emitter bytecode program (a shader VM). Effects are authored as
6
- node graphs and lowered to that bytecode by a testable JS compiler. Emitters are scene nodes, and
7
- the decision to spawn is the GPU's.
8
-
9
- ## Goals (from the brief)
10
-
11
- - Single simulation execution shader (data-driven VM, not one-pipeline-per-effect).
12
- - Projection modes: camera-facing billboard (+ velocity-stretched, axis-locked, world-oriented).
13
- - Depth sorting for correct transparency, built on the `csdldf` prefix scan.
14
- - Node-based composition → shader VM (compile graph → bytecode).
15
- - Frustum culling of spawning (optional, flagged, sleep-when-culled + bounded catch-up).
16
- - Flipbooks (atlas sub-rect + frame grid).
17
- - Texture lookup via atlas (region table; mirrors meep `AtlasPatch` uv offset/scale).
18
- - Soft-depth fade (flagged).
19
- - `GPUDatabase` for CPU-managed, low-churn tables (emitters).
20
- - Scene-node / bone attachment: an emitter **is** a `Node3D`; its world matrix is its row of the
21
- scene's `transforms` table, composed by the GPU hierarchy pass.
22
- - Distinct per-particle lifecycle: **no** built-in age/position — the schema + program define state.
23
- - Billboard renderer reuses engine shader chunks incl. lighting (forward-shaded, optional per emitter).
24
- - Blend modes (alpha + additive) via a **single** premultiplied blend setup.
25
- - `AnimationCurve` sampling during simulation (reuse `chunk_animation_curve_evaluate`).
26
- - Spawning decided on the GPU, per emitter, with the CPU able to ask for bursts through the same path.
27
- - No per-frame allocation on the CPU beyond the frame graph's own records.
28
-
29
- ## Layered architecture
30
-
31
- ```
32
- authoring: NodeGraph (nodes+ports) ──compile──▶ Program (bytecode: INIT + UPDATE)
33
- standard node library │
34
- ▼
35
- runtime CPU: ParticleEmitter (a Node3D) ── Scene ──▶ GPUSceneContext (transforms table)
36
- EmitterRegistry ── GPUDatabase(emitters + emitter_state) ─┤ ProgramHeap ([header][constants][code])
37
- GPUParticleSystem (owns every buffer, records the frame)
38
- ▼
39
- runtime GPU: particle pool (array<u32>, uniform stride) + alive/dead lists + counters
40
- emitter state table (GPU-owned: generation, accumulator, age, sleep, pending)
41
- ┌─ spawn commands (CPU bursts → pending)
42
- ├─ emitter tick (per row: cull on the published box, integrate rate, take budget → count)
43
- ├─ scan (csdldf) → emit dispatch
44
- ├─ emit (VM INIT program) ── per newborn: find emitter in scan, pop dead → init → append alive_new
45
- ├─ simulate (VM UPDATE program) ── update → (kill|orphan→dead | keep→alive_new)
46
- ├─ bucket (histogram → csdldf scan → scatter) ── alive → program-coherent sim list
47
- ├─ bounds (per particle → per emitter box, published to the state row)
48
- ├─ build indirect (dispatch + draw args)
49
- ├─ sort (histogram → csdldf scan → scatter) [render order]
50
- └─ render (instanced billboard quads, forward + optional lighting + soft depth)
51
- ```
52
-
53
- ### Particle storage — single shared pool, uniform stride
54
-
55
- All particles live in one `array<u32>` pool with a fixed `RECORD_WORDS` stride (config, default 32
56
- words = 128 B). Word 0 is a reserved header holding the owning **emitter row** and that row's
57
- **generation** (`data/particle_header.js`); user attributes occupy words `1..RECORD_WORDS`. A single
58
- pool + single alive/dead lists means **one** emit dispatch and **one** simulate dispatch cover every
59
- emitter — the Destiny single-shader property.
60
-
61
- Lifecycle uses the canonical GPU dead-list / alive-list ping-pong:
62
- - `dead_list`: stack of free slot indices (+ atomic count).
63
- - `alive_list[2]`: ping-pong compact lists of live slot indices (+ counts).
64
- - `counters`: atomics {alive_count, dead_count, spawn budget, …} + derived indirect dispatch/draw args.
65
-
66
- Emit: pop dead → run INIT → append `alive_new`. Simulate: read the **coherence list** last frame's
67
- bucketing pass built from `alive_current` → run UPDATE → kill pushes to `dead`, survival appends to
68
- `alive_new`. Swap lists each frame.
69
-
70
- **The loop is strictly GPU-driven.** The build-indirect pass snapshots the padded length of the
71
- coherence list into `counters[ALIVE_IN]` and writes the indirect dispatch args next frame's simulate
72
- consumes (`dispatchWorkgroupsIndirect`) plus the indirect draw args this frame's render consumes
73
- (`drawIndirect`, sized by the *unpadded* alive count); simulate bound-checks against the ALIVE_IN
74
- snapshot. The sort's per-particle passes take the dispatch the bucket reset writes from the alive
75
- count. Sizing any of this from a CPU readback is a correctness bug, not a shortcut: the readback is
76
- inevitably stale, so a growing population orphans the tail of the alive list (slots leak into limbo)
77
- and a shrinking one re-simulates stale entries — double-freeing slots so two live particles later
78
- share one record. The emit pass's dead-list pop guards against wrapped counters
79
- (`old == 0 || old > capacity`), so racing an exhausted pool drops spawns cleanly instead of handing
80
- out garbage slots.
81
-
82
- ### Spawning is the GPU's decision
83
-
84
- The CPU never walks the emitters per frame and never counts particles. Per frame:
85
-
86
- 1. **spawn commands** (`shaders/shader_particle_spawn_commands.js`, only when the host queued any):
87
- one thread per `[row, count]` pair the host authored, `atomicAdd` into that row's `PENDING` in
88
- the GPU-owned emitter state. This is `GPUParticleSystem#spawn` — the CPU's facade over the GPU
89
- path, not a second path: a burst goes through the same budget, expansion and INIT program as the
90
- emitter's rate. It is also the channel a VM opcode would use to spawn from the GPU (a sub-emitter
91
- or trail would `atomicAdd` its target's `PENDING`).
92
- 2. **emitter tick** (`shaders/shader_particle_emitter_tick.js`): one thread per emitter table row,
93
- walked with the table's page iterator so an unallocated page costs one word and a freed row one
94
- bit. Per live row: recognise the row by generation (a mismatch starts the state over — new
95
- emitter, or a reused row, with nothing for the CPU to clear); take `PENDING`; publish the
96
- bounding box emit and simulate measured last frame into the state row and reset the accumulator
97
- (see **Bounds**); if the emitter opts into culling, test that box against the camera frustum and
98
- sleep while culled, paying back a bounded catch-up on wake; integrate
99
- `spawn_rate * dt`; turn the whole part of the accumulator into this frame's count, clamped by the
100
- **spawn budget** the reset pass seeded from the dead count — so an emitter is only ever debited
101
- for particles the pool can hold; write the count to `spawn_counts.elements[row]`.
102
- 3. **scan**: `graph_prefix_scan_csdldf` over the counts — each row's batch ends at its inclusive sum,
103
- and the last sum is the number of particles born.
104
- 4. **spawn args**: `shader_prefix_sum_to_command` (the rasteriser's own) turns that total into the
105
- emit dispatch.
106
- 5. **emit**: one thread per newborn, locating its emitter by `lower_bound_branchless` over the sums
107
- and its batch fraction from the sum before — the meshes-to-meshlets expansion applied to
108
- particles. No per-emitter cap, no thread loops over its particles.
109
-
110
- On a frame that registered emitters with a `prewarm`, and only then, a loop of the warm-up pipeline
111
- (`warmup/`) runs after the frame's own simulate: the same decision-scan-emit shape per tick over a
112
- transient population, plus an advance of its own, then a migration into the main pool. See
113
- **Pre-warm**.
114
-
115
- The host contributes two numbers to all of that: the emitter table's element capacity (the scan's
116
- length — static topology) and the length of the command list it wrote itself.
117
-
118
- ### The emitter state (`data/PARTICLE_EMITTER_STATE.js`)
119
-
120
- What the GPU integrates per emitter — generation, accumulator, age, sleep, pending — is the
121
- `emitter_state` table of the same database, one row per emitter at the emitter's row, **not**
122
- fields of the emitter record. The record is CPU-shadowed and restaged whole whenever the author
123
- changes a field, and that restage would overwrite an accumulator the GPU has been adding to for a
124
- hundred frames. The scene's transforms table mixes CPU fields with a GPU-written `global` only
125
- because `global` is recomputed from scratch every frame; nothing here is. In its own table the CPU
126
- touches a row exactly twice — a zero record when the emitter's row is claimed, the table's own
127
- zero-fill when it is released — and the database's resize copies pages on the GPU, so growth keeps
128
- what the tick has written. Nothing in it is atomic: each row is one tick thread's, and the
129
- spawn-command pass is handed one command per row (the system merges bursts on the CPU). The
130
- emitter's `age` here is the VM's `EMITTER_AGE` builtin — and the one field the warm-up pipeline
131
- writes, as it runs an emitter's clock forward (see **Pre-warm**).
132
-
133
- ### Eight storage buffers
134
-
135
- A compute stage gets eight storage buffers by default, and the design does not ask for more. Two
136
- things keep every pass under it: the emitter record and the GPU-owned state are two tables of one
137
- database buffer, read through one binding; and the VM's constants and code are one packed buffer —
138
- `[header][constants][code]`, the header's first word saying where the code starts
139
- (`ParticleConstants.PARTICLE_PROGRAM_HEADER_WORDS`). Emit and simulate bind exactly eight; their
140
- wide-register variants bind nine, which is the one place the system spends past the default and is
141
- still inside the ten the engine demands of an adapter. `particle_shaders_wgsl_validity.spec` pins
142
- the count for every pass.
143
-
144
- ### The VM (`ParticleVM`)
145
-
146
- Register machine, emulator-safe (bounded instruction loop + big `switch`, no subgroup ops, no
147
- recursion, no data-dependent jumps — conditionals are predicated via CMP/SELECT, destruction via
148
- `KILL`).
149
-
150
- - Register file: `array<f32>`, a register being one slot and a value of width *w* occupying *w*
151
- consecutive ones. The interpreter reaches it through `vm_reg_read` / `vm_reg_write`, which a
152
- **backing chunk** supplies, and the two backings are the two interpreters:
153
- `vm/chunk_particle_vm_registers_fast.js` puts the file in `var<private>`
154
- (`PARTICLE_VM_FAST_REGISTER_SLOTS`, dynamically indexed, so its size is the shader's register
155
- footprint), and `vm/chunk_particle_vm_registers_wide.js` puts it in a storage buffer
156
- (`PARTICLE_VM_REGISTER_SLOTS` — every slot an 8-bit operand base can name). See **Two
157
- interpreters** below.
158
- - Instruction: fixed **4×u32** words `[op | width << 8 | kinds << 16 | dst << 24, s0, s1, s2]`.
159
- Word 0 packs the opcode (bits 0..7), the lane count (8..10), each source's `VM_KIND` (two bits per
160
- slot at 16..21) and the destination slot (24..31). A source word means what its kind says: `REG`
161
- is a base slot plus a swizzle, `IMM` is an `f32` bit pattern read in every lane, and `NONE` is a
162
- raw field the opcode reads for itself (attribute word-offset+comp / curve handle / builtin id /
163
- pool index).
164
- - Constant pool: `array<vec4f>` per program, referenced by `LOAD_CONST`.
165
- - Two entry sections per program: **INIT** and **UPDATE** (offsets in the emitter record). INIT
166
- gives an already-allocated record its state; it has no say in whether the particle exists.
167
- - Builtins (`LOAD_BUILTIN`): DT, EMITTER_AGE, EMITTER_POSITION, EMITTER_DIRECTION, EMITTER_UP,
168
- PARTICLE_INDEX, SPAWN_FRACTION, TIME … assembled per-invocation by the host shader. The emitter's
169
- position and axes are its node's `global` out of the scene database.
170
- - `CURVE dst, handle, s0`: `dst.x = animation_curve_evaluate(curves[handle], s0.x)` — reuses the
171
- engine's GPU AnimationCurve chunk + database.
172
-
173
- #### Two interpreters
174
-
175
- The register file normally lives in shader registers, which is what makes it fast and what makes it
176
- small. An effect whose compiled program wants more slots than it holds is not an error and not an
177
- authoring ceiling: the whole system swaps emit and simulate for variants whose file is a storage
178
- buffer, and everything runs, slowly.
179
-
180
- - **The trigger is the set, not the program.** `ProgramHeap` reports `max_register_count` over the
181
- programs it holds, and `GPUParticleSystem` is where that number is compared against
182
- `PARTICLE_VM_FAST_REGISTER_SLOTS` — the store has no opinion about what a register budget means,
183
- and which backing a pass is compiled against is the VM's business. One hungry effect moves every
184
- emitter onto the wide pair; there is one pipeline per pass and it is chosen once, per frame, in
185
- `graph_particles`. It moves back when that effect is unregistered: the heap knows what it holds,
186
- so the number is a maximum over the live set rather than a high-water mark.
187
- - **The launch is bounded.** The wide passes need a register file per invocation in flight, so they
188
- do not run one lane per particle — that would size the buffer by pool capacity. They launch
189
- `PARTICLE_VM_WIDE_LAUNCH_LANES` lanes whatever the workload and walk their input with a grid-stride
190
- loop, so the file is `LANES * SLOTS * 4` bytes and nothing else. The stride is a whole number of
191
- workgroups, which is what keeps a wave inside one coherence bucket and the wave-uniform section
192
- decode sound.
193
- - **Slot-major, and never cleared.** `vm_registers[slot * LANES + lane]`, so a wave reading one slot
194
- reads a coalesced run. The buffer is transient and no pass zeroes it: a register is undefined until
195
- written, and a program that reads one first is corrupt. That is the one place the two backings are
196
- not interchangeable — the fast file zeroes, so the reference executor and the fast VM agree on a
197
- value the wide VM has no opinion about.
198
- - **Held to the same bar.** `vm/chunk_particle_vm.spec.js` runs every parity case, fuzz included,
199
- against BOTH backings; `shaders/particle_wide_register_passes.spec.js` runs the wide passes against
200
- the fast ones under the emulator and exercises the grid-stride loop.
201
-
202
- **Parity strategy.** The opcode numbers + operand semantics live in one source of truth
203
- (`isa/ParticleVMISA.js`). Two executors implement it: a JS reference VM (`vm/ParticleVMReference.js`)
204
- and a hand-written WGSL interpreter (`vm/chunk_particle_vm.js`). A spec runs identical programs
205
- through both (WGSL via the emulator) and asserts bit-for-bit-close equality. A meta-test asserts both
206
- executors cover every opcode in the ISA table.
207
-
208
- ### Node graph → bytecode
209
-
210
- A core `NodeGraph` holds particle nodes (ports + parameters) for INIT and UPDATE. `compile()`:
211
- 1. topologically orders reachable nodes,
212
- 2. lowers each node to VM ops (nodes emit via a small builder that allocates result registers),
213
- 3. linear-scan register allocation with liveness,
214
- 4. interns constants into the pool,
215
- 5. resolves curve/attribute references,
216
- 6. emits `Program { init, update, constants, reg_count, attributes, curves }`.
217
-
218
- Standard node library: constants, attribute get/set, arithmetic/vector/math, random (uniform/sphere/
219
- disk/cone), curve-sample, noise (curl), integrate, gravity, emit-shapes, compare/select, kill-if.
220
-
221
- ### Emitters are scene nodes
222
-
223
- `ParticleEmitter extends Node3D`. Add one to a `Scene` — on its own or under a joint — and the
224
- scene context's membership sweep gives it a `transforms` row like any node; the GPU hierarchy pass
225
- composes its `global` from its own transform and its parents' every frame, animation included.
226
- `GPUParticleSystem` sweeps the scene's node list when the scene's membership version moves, gives
227
- each emitter node a row in the emitter table pointed at its `transforms` row, and retires the row
228
- when the node leaves. Attaching an emitter to a bone is `joint.addChild(emitter)`; there is no
229
- per-frame transform copy, and an emitter under a GPU-animated joint follows it — which the old
230
- CPU-side copy of `transform_global` could not do, since the CPU never updates that matrix for a
231
- node the GPU owns.
232
-
233
- **The node knows nothing about the GPU.** A `ParticleEmitter` is a transform, an attribute
234
- `layout`, a compiled `program`, a `texture` URL, and the emission parameters — all of it the
235
- author's, all of it meaning the same thing with no device present, all of it what a saved effect is
236
- made of. Everything the system assigns is a `GPUParticleEmitterContext`, held by the
237
- `EmitterRegistry` against the emitter and reached through `registry.context(emitter)`: the table
238
- row, the generation, the program placement, the node's `transforms` row, the atlas patch. So an
239
- emitter cannot be saved with a row in it, cannot be given one by hand, and "is this emitter live?"
240
- is asked of the registry rather than read off a `-1` that anything could have written. The record
241
- writer is where the two meet — `write_emitter_record(record, emitter, context)`, next to the struct
242
- it fills.
243
-
244
- Restaging is driven by `Node3D#version`, the node's own change counter, which the renderer already
245
- reads to decide whether a transform needs re-uploading; the context remembers the version its row
246
- was staged from. An author who edits an emitter says so the way they say it for any node —
247
- `needsUpdate = true` — and there is no second flag to forget. A change to the context itself (a
248
- `transforms` row, a repacked atlas patch) moves no version, so the context marks itself instead.
249
-
250
- ### Emitter table (`GPUDatabase`)
251
-
252
- One `emitters` table (CPU-managed, low churn). Row = `PARTICLE_EMITTER_STRUCT`: program
253
- offsets+lengths, constant-pool offset, seed, spawn rate, **node** (transforms row), **generation**,
254
- flags (lighting/soft/sort/cull/blend/projection), render binding (attr offsets for
255
- position/size/color/rotation/frame/velocity), atlas region + flipbook grid, and the program id
256
- the coherence pass buckets on. The atlas region is resolved, not authored: the emitter names an
257
- image and the system packs it, the way `ShaderManager` packs the old engine's sprites, so a repack
258
- moves every patch without any emitter changing. The seed is resolved too — an emitter left at 0 gets
259
- its registration's generation, which no other emitter of the registry shares, rather than being
260
- written to. Programs+constants live in a packed `ProgramHeap` buffer
261
- (variable length → not a fixed-stride table).
262
-
263
- The struct is the ONLY statement of that layout. The JS writer is the database's
264
- `write_wgsl_type_value`, and the WGSL reader — the `ParticleEmitter` struct declaration plus
265
- `database_read_..._element` — is generated from the same struct by the table descriptor. There is no
266
- hand-packed word offset table to drift out of step with the shaders, and no hand-written reader.
267
- The bucket histogram, which runs once per live particle and wants only `program_id`, reads it
268
- through `chunk_read_field` rather than decoding a whole row.
269
-
270
- **Editing the set at runtime is the point of using a table.** `EmitterRegistry.add` stages one
271
- record and `remove` frees one row; neither touches the others. Rows are stable for an emitter's
272
- lifetime because a live particle's record header holds the row of the emitter that spawned it —
273
- renumbering a survivor would re-point its particles at somebody else. A removed row is handed back
274
- to the next `add`, so the id space does not grow without bound.
275
-
276
- **Generations make that reuse safe.** Every add takes the next value of one running counter, never
277
- zero, and the particles of an emitter carry its generation in their header. The simulate pass kills
278
- any particle whose generation is not its row's current one — the row is zero-filled (generation 0)
279
- or a later tenant's — so a hard removal frees the emitter's particles within a frame, and a reused
280
- row cannot adopt a stranger's. Graceful retirement (rate to zero, remove once the tail has died)
281
- remains an author's choice rather than a correctness rule.
282
-
283
- **Programs are edited the same way.** `ProgramHeap` is a pair of range allocators
284
- (`OffsetAllocator`) over the two sections of the packed buffer, keyed by a `HashMap` on the bytecode
285
- itself — `ParticleProgram` answers `equals`/`hash`, so content deduplication is the map's doing
286
- rather than a stringified key's. `acquire` takes a reference and `release` gives one back, so the
287
- last emitter of an effect is what unpacks it, and `EmitterRegistry.set_program` swaps an emitter onto
288
- different bytecode while it is live. What makes that possible is that a program is **relocatable**:
289
- instructions carry no absolute addresses and constant references are program-local, so the heap can
290
- move one by copying its words and rewriting the placement.
291
-
292
- Two things are deliberately NOT reclaimed. **Program ids** are reused rather than renumbered — they
293
- are baked into every emitter record, so compacting them would mean rewriting the whole table — and
294
- `bucketCount()` is sized from the id capacity, the set's peak, so it does not move every time an
295
- effect comes or goes. An id below the peak with nothing in it is an empty bin, which rounds up to
296
- zero workgroups and shifts nothing after it. **Section capacity** only rises: a repack (which happens
297
- when a section has no contiguous run long enough) compacts what is live and grows only if it must,
298
- so a set that peaked at a hundred effects does not pay for a resize every time it returns to three.
299
-
300
- A repack is the one change no emitter announces, since it is not an edit to any of them.
301
- `EmitterRegistry.flush` asks the heap once per frame — `placement_version` — and restages every row
302
- when it has moved.
303
-
304
- Programs are immutable after compilation. Both `registry.set_program(emitter, program)` and replacing
305
- `emitter.program` followed by `needsUpdate = true` are supported. Flush reconciles replacements
306
- before inspecting heap relocation, since acquiring a replacement may move other programs too.
307
- The layout must remain compatible with particles already in flight.
308
-
309
- A live program assignment increments `program_assignment_version`. On the next frame, before reset
310
- or simulation, the previous compact alive list is bucketed again and its padded count and dispatch
311
- are rebuilt. This prevents a previously shared wave from executing one emitter's newly edited program
312
- over another emitter's particles. Ordinary frames keep the existing end-of-frame bucketing only;
313
- content-identical replacements need no repair. Both bucket runs on an edit frame reuse the latest
314
- graph handles for their scratch buffers, so each persistent buffer is imported once.
315
-
316
- The coherence list is persistent GPU state, including its padding. Growing the bucket capacity copies
317
- it to the replacement buffer before destroying the old one. The old count and indirect dispatch
318
- remain valid across growth; only the per-frame histogram/cursor/key scratch may be discarded.
319
-
320
- Every row of a frame is staged through one module-static record (`EmitterRegistry`), so a flush
321
- allocates nothing however many rows it restages.
322
-
323
- ### Bounds
324
-
325
- An emitter has no bounds it could declare. Its particles are simulated in world space by arbitrary
326
- bytecode: they go where the program sends them, at whatever size it gives them, for as long as it
327
- keeps them alive. The only bound an author could write down and still be right is an infinite one,
328
- which culls nothing and reserves every depth slice. So the box is **measured**, on the GPU, every
329
- frame.
330
-
331
- Two passes of its own (`bounds/`), at the end of the frame, after emit and simulate have both
332
- finished appending to the alive list:
333
-
334
- - **accumulate** — one thread per live particle: read where its quad is and merge it into its
335
- emitter's accumulator. The centre is the record's `render_position`; the dilation is
336
- `particle_billboard_radius`, the largest offset `particle_billboard_offset` can produce for that
337
- size, velocity and projection, over every corner, every roll and any camera. The box has to hold
338
- the *quads*, not the centres, and it has to be camera-independent, because it is read on a later
339
- frame. The merge is `aabb_atomic_min_f32` / `aabb_atomic_max_f32`, the same CAS-spin the
340
- skinned-mesh bounds refresh uses; both load first and return without a CAS once the box holds the
341
- contribution, so what survives is load traffic, not exchanges. Dispatched indirectly from
342
- `sim_dispatch_args`, the args the coherence reset wrote from the GPU's own alive count.
343
- - **publish** — one thread per emitter row: write the accumulation to `bounds_min` / `bounds_max` of
344
- the state row (`data/PARTICLE_EMITTER_STATE.js` — measured state is GPU-owned state, and on the
345
- record it would be clobbered by the next restage), and reset the accumulator. The reset is
346
- unconditional, which is what lets a row that changes hands start clean.
347
-
348
- **Why a pass and not a tail on emit and simulate.** It was that first. Booking from inside those two
349
- measures no single population — simulate sees last frame's survivors and emit sees this frame's
350
- newborns, so the union spans two frames and is looser than either — and it pays six atomics per
351
- particle inside the pass whose register pressure decides the system's occupancy, to compute
352
- something that needs none of the VM, none of the scene, and five words of the record. Over the alive
353
- list instead, the measurement is of exactly the set that is alive and about to be drawn.
354
-
355
- **Where the accumulator lives.** Six words per emitter row, in the **counters** buffer, after the
356
- counters themselves (`data/PARTICLE_COUNTERS.js`). It has to be an `array<atomic<u32>>` binding, and
357
- the emitter database cannot be it — the passes that read a record read it out of an `array<u32>`
358
- binding, and one buffer under two access types is aliasing. The counters buffer is already bound in
359
- both bounds passes, already atomic, already the system's cross-pass scratch, so a buffer of its own
360
- would add a binding and a growth path to say the same thing. It therefore grows with the emitter
361
- table, by copy — the free-list length in it is GPU-owned and unreconstructable — and the rows the
362
- growth adds are written to the empty interval, because a zeroed accumulator reads back as a
363
- degenerate box at the world origin rather than as an absence.
364
-
365
- **The empty case.** An emitter with no measured particles has infinite, unknown bounds:
366
- `bounds_min = [-Infinity, -Infinity, -Infinity]`, `bounds_max = [Infinity, Infinity, Infinity]`.
367
- The CPU stages this on registration and the publish pass restores it whenever the population is
368
- empty. The tick treats unknown bounds as visible, so a slow emitter can accumulate its first spawn
369
- without prewarm even when its node is offscreen. A depleted population can restart the same way.
370
- AVBOIT occupancy skips unknown bounds because there are no live particles to splat. Both readers
371
- test the sentinel before geometric arithmetic, avoiding NaNs from operations on infinities.
372
- The sentinel check compares integer bits. The publish pass swaps the empty accumulator's loaded
373
- endpoints; neither path constructs an infinite floating-point constant, which WGSL rejects at
374
- shader creation time.
375
- Measured, nonempty populations continue to use their finite bounds for culling. No bounds are authored.
376
- Prewarm controls the initial age distribution only; it is not required for culling correctness.
377
-
378
- ### Pre-warm
379
-
380
- A continuous effect that starts empty grows into its steady state over one particle lifetime, in
381
- front of whoever is watching. `ParticleEmitter.prewarm` — seconds, `0` for off — buys that lifetime
382
- back: an emitter registered with one is run forward by that long before it is first drawn.
383
-
384
- **A warm-up is the ordinary loop, run over a population of its own.** Spawn, init, advance, per
385
- tick, until every emitter in the batch has been running for its own `prewarm`. It is *not* a batch
386
- of particles dumped in one tick and aged: the emitter keeps spawning throughout, what dies is
387
- replaced as it is in a live scene, and whatever drives the rate — a spawn curve, `EMITTER_AGE`, the
388
- program itself — drives it here, because it is the same program being run. What the loop leaves is
389
- the population the emitter would have had.
390
-
391
- **It is a pipeline apart from the main one** (`warmup/`), because the main one is tuned and must not
392
- pay for this. Emit and simulate are the passes whose register pressure sets the system's occupancy,
393
- and their waves are kept program-uniform by the bucketing; a warm-up folded into either would spend
394
- registers on every frame for work that happens on one, and a loop inside a lane would break the
395
- uniform flow across a wave. So the warm-up shares code with the main frame and integrates nothing
396
- into it: the main **emit pass is reused unchanged**, bound to the warm-up's buffers; the tick is a
397
- pass of the warm-up's own over a host-written queue (`shader_particle_warmup_tick`) rather than over
398
- the emitter table; the **advance** is a pass of the warm-up's own (`chunk_particle_warmup_step.js`),
399
- decoding the program section per lane — it reads the raw alive list, so it cannot lean on the
400
- bucketing's wave-uniformity, and its occupancy is its own business. The main passes' compiled code
401
- is byte-identical whether `warmup/` exists or not.
402
-
403
- **The population is transient.** Pool, both alive lists, free list and counters are frame-graph
404
- transients that exist for the loop and are then moved, record by record, into the main pool by the
405
- **migrate** pass — a pop from the main free list per particle, and never a push, so it is the one
406
- point of contact and it is safe by construction. The loop's own free-list traffic never touches the
407
- main list, and the main passes never see a half-warmed particle.
408
-
409
- **Bounded.** `warmup/warmup_schedule.js`: at most `PARTICLE_WARMUP_MAX_TICKS` ticks, no shorter than
410
- `PARTICLE_WARMUP_TICK_SECONDS`, the tick *length* stretching to cover the longest pre-warm in the
411
- batch rather than the count growing to cover it — the count is what the frame pays for, at six or
412
- seven dispatches a tick. An emitter takes part in a tick once the loop's clock is within its own
413
- pre-warm of now, so two emitters queued with different durations each end having run for exactly
414
- their own; the shorter one sat out the start.
415
-
416
- **Jittered, so it does not arrive in bands.** Every particle born in one tick would otherwise reach
417
- the end with exactly that tick's age, and N ticks would read as N cohorts. The advance jitters the
418
- tick per particle (`PARTICLE_WARMUP_DT_JITTER`, hashed from slot and tick — a hash, not a draw, so
419
- the program's own random stream is untouched), and the host jitters each emitter's clock and starting
420
- accumulator (`PARTICLE_WARMUP_EMITTER_JITTER`), so emitters in one batch do not spawn on the same
421
- beat either. A program does not care what `DELTA_TIME` it is handed; the smear is free.
422
-
423
- **Where it sits.** After the frame's own emit and simulate, before the bucketing. Migrate appends to
424
- the same alive list they did, so the bucketing puts the warmed particles in next frame's simulate
425
- list, the bounds pass measures them into their emitter's box this frame, and the draw draws them
426
- this frame. The tick also writes the emitter's
427
- `age` into its state row as it goes, so `EMITTER_AGE` reads as it would have, and the next real tick
428
- adds its `dt` to a row that says `prewarm`.
429
-
430
- **Only on the frame that needs it.** The host knows: an emitter is queued when it is registered with
431
- a `prewarm` (`GPUParticleSystem#sync_membership`), the queue is uploaded and consumed by the frame
432
- that records the loop, and an ordinary frame records nothing of this.
433
-
434
- The editor's CPU preview (`editor/particles/effect/ParticleReferenceSimulation.js`) runs the same
435
- loop over the same schedule before its first frame, and restarts the effect when the duration
436
- changes: a pre-warm is a statement about how an effect starts, so the only way to show one is to
437
- start it.
438
-
439
- ### Culling & attach
440
-
441
- An emitter that sets `EMITTER_FLAG.CULL` has the box on its state row tested against the camera
442
- frustum by the tick pass, with the renderer's shared `chunk_aabb3_intersects_frustum`. The tick does
443
- not measure it — it reads what the bounds passes published at the end of the previous frame — so all
444
- that is in the spawn path is one frustum test. Because the box is where the particles are rather
445
- than where the node is, an effect whose particles have drifted on screen keeps spawning even when
446
- its emitter has not. Culled emitters spawn nothing and accrue sleep; on wake the tick adds a
447
- catch-up of at most `max_catchup_seconds` of rate. Live particles keep simulating while their
448
- emitter is culled, so its box keeps moving and keeps being published. HiZ occlusion is not
449
- written.
450
-
451
- ### Two orderings, and why they are separate
452
-
453
- There are two independent permutations of the live particles each frame. They are produced by
454
- different passes into different buffers and must not be conflated or derived from one another.
455
-
456
- **Simulation order** (`coherence/`) — camera-independent, changes only when emitters or their
457
- programs change. The alive list comes out of compaction (`atomicAdd(&counters[ALIVE])`), so its
458
- order says nothing about which emitter a particle belongs to and a wave routinely holds particles
459
- running different VM programs. That costs twice: the interpreter's ~40-case opcode `switch`
460
- diverges, so a wave serially executes the union of the opcodes its lanes want; and `vm_reg[dst]` /
461
- `vm_reg[s0]` index a VGPR-resident register file dynamically, which AMD reaches with
462
- `v_movrels`/`v_movreld` behind a waterfall loop whose trip count is the number of *distinct*
463
- register indices live in the wave. Grouping the wave by program collapses both at once.
464
-
465
- The grouping is also **load-bearing**: the simulate shader reads its section parameters (code offset,
466
- instruction count, constant offset) from the wave's first lane through `subgroupBroadcastFirst`
467
- (`particle_wave_uniform`), which is what lets the driver fetch instruction words through the scalar
468
- unit, branch the opcode switch on a scalar and index the register file through M0 without a
469
- waterfall. A wave holding two programs would run one over the other's particles, so the coherence
470
- pass is part of simulate's correctness. Emit keeps a per-lane decode: its waves span emitters.
471
-
472
- The bucket key is the **program**, not the emitter: `ProgramHeap` deduplicates by content and hands
473
- back a `program_id` the emitter record carries, so emitters compiled from the same graph share
474
- one bucket. Each bucket is padded out to `PARTICLE_SIMULATE_WORKGROUP_SIZE` (64 — a multiple of both
475
- plausible subgroup widths, so no shader needs to query `subgroup_size`), which is what makes the
476
- grouping hold at wave granularity: an unpadded boundary lands mid-wave and that wave spans two
477
- programs again. Padding entries carry `PARTICLE_SIM_PAD_SLOT` and the simulate shader drops them
478
- *before* the compaction atomics, so they never enter the alive or dead list.
479
-
480
- It is an **indirection** list. Particle records never move: slots are referenced by the dead list and
481
- across frames, so relocating a record would leave both pointing at the wrong particle.
482
-
483
- The three-phase shape is the depth sort's, with the padding added: histogram (atomic per bucket) →
484
- round each bin up to a whole workgroup → `graph_prefix_scan_csdldf` (reused as-is — the one scan here
485
- whose only cross-workgroup wait is CSDLDF's bounded spin-then-fallback) → scatter → fill the tails
486
- with the sentinel. Nothing reads back to the CPU; the only host-supplied number is the bucket count,
487
- which is static topology.
488
-
489
- **Render order** (`sort/`) — per-frame, particles sorted back-to-front by view depth so
490
- premultiplied transparency composites correctly. Counting sort on a quantised view depth
491
- (`PARTICLE_SORT_BUCKET_COUNT` bins over `GPUParticleSystem#sort_depth_range`): histogram from the
492
- camera uniform's view matrix, `graph_prefix_scan_csdldf` the bin counts to offsets, scatter (index
493
- payload). The sorted index list feeds the instanced draw. Both per-particle passes are bounded by
494
- the alive counter and dispatched by the args the bucket reset wrote from it.
495
-
496
- Reordering the render list for program coherence would break transparency; depth-sorting the
497
- simulation list would destroy the coherence. They share no buffer and no counter slot —
498
- `PARTICLE_COUNTER.SORTED` belongs to the render order, `SIM_PADDED` to the simulation order.
499
-
500
- ### Rendering
501
-
502
- One render pipeline, one blend state = premultiplied: `color: (one, one-minus-src-alpha)`,
503
- `alpha: (one, one-minus-src-alpha)`. Additive vs. alpha is chosen by the shader emitting premultiplied
504
- RGB with alpha 0 (additive) vs. alpha a (normal) — GPU-driven, no pipeline switch. A 6-vertex quad is
505
- expanded per instance (billboard/velocity-stretch/axis-lock/world modes). Fragment: sample atlas
506
- (region + flipbook frame), optional forward lighting via `chunk_shade_standard_material_direct` +
507
- IBL, optional soft-depth fade. Depth-tested, no depth write. Draws into the HDR color target.
508
- Soft-depth fading scales both RGB and alpha for premultiplied inputs; straight-alpha and additive
509
- inputs have alpha scaled before premultiplication. Both draw paths share `particle_fade_color`.
510
-
511
- ### Drawing under the renderer: an AVBOIT side channel
512
-
513
- Under `Renderer` the particles are not drawn by the billboard pass above. They go through the AVBOIT
514
- transparency pipeline (`rasterize/native/avboit/`) as a **side channel** — the contract in
515
- `AVBOITSideChannel.js` — so they composite order-independently with the transparent meshes and
516
- against the opaque scene, through the same volume and the same accumulators. Quads are not packed
517
- into meshlets; the orchestrator asks the channel for its contribution at the three points where the
518
- meshlet buckets have just made theirs:
519
-
520
- 1. **occupancy** (`shader_particle_avboit_occupancy`): one thread per emitter row, the emitter's
521
- measured box (the culling box, off the state row) as a view-depth interval, dilated and OR-ed
522
- into the occupancy bits. The interval is the box's centre depth plus or minus its extent along
523
- the view axis — a support projection, not eight transformed corners. Per emitter, not per
524
- particle — no contention, and a particle that strays lands on a run boundary rather than anywhere
525
- wrong. Additive emitters mark nothing. The box matters here even for an emitter that never opted
526
- into culling, which is the second reason it is measured for every emitter rather than only for
527
- the ones that cull.
528
- 2. **splat** (`shader_particle_avboit_splat`): the billboard quads rasterized at the volume's
529
- resolution, coverage from the atlas alpha, extinction through `avboit_splat_slice` — the same
530
- write path as a surface. Additive particles write nothing. A soft-depth particle is faded here
531
- against the same scene depth the draw fades it against, one sample per voxel: what the volume
532
- records has to be the coverage the accumulators are weighted by, or a card that fades out where
533
- it meets the floor goes on darkening the floor through itself.
534
- 3. **draw** (`shader_particle_avboit_draw`): the quads at internal resolution, depth-tested read-only
535
- against the scene (which lets the soft fade read the same depth), fog composited, transmittance
536
- from the integral. Which of the two it is, is decided by coverage rather than by the blend enum,
537
- so an additive emitter and a premultiplied sprite with no alpha take the same path — as they do
538
- in the splat. A covering particle writes `color / norm / extinction` as a surface does, in-scatter
539
- included, weighted by the coverage it took. An emitting one writes the fourth accumulator,
540
- **emission**: light attenuated by the fog and by the volume in front of it, and carrying none of
541
- the fog's own in-scatter, because the resolve adds emission outside the normalization and the
542
- scene behind already has it. That accumulator is the one addition to AVBOIT's resolve for this.
543
- The whole decision is `chunk_particle_avboit_contribution`, and it is spec'd on the emulator.
544
-
545
- No sort: the volume makes order irrelevant, so this path reads the alive list and the draw args
546
- straight from `simulate`. The billboard vertex stage is one chunk
547
- (`chunk_particle_billboard_vertex`) shared by the standalone renderer and both AVBOIT passes.
548
- `GPUParticleSystem` implements the three methods; `Renderer.feature_particles_enabled` runs
549
- `simulate` before the transparency pass and hands the system to the orchestrator. MBOIT is legacy
550
- and gets no particle path.
551
-
552
- ### Texture ownership: ECS
553
-
554
- `engine/graphics3/GPUParticleEmitterSystem.js` owns the sprite atlas for one Shade scene. Construct it
555
- with `(graphicsEngine, scene, assetManager)` and add it to the EntityManager. It observes the scene's
556
- existing `ParticleEmitter` nodes; there is no second component or transform to synchronize. It is
557
- opt-in and can coexist with the legacy CPU `ParticleEmitterSystem`.
558
-
559
- ```js
560
- entityManager.addSystem(new GPUParticleEmitterSystem(graphics, scene, assetManager));
561
- ```
562
-
563
- The system loads image URLs through `AssetManager` (`GameAssetType.Image`) and shares one patch per
564
- URL. It reuses `TextureAtlas`, whose packer is `MaxRectanglesPacker`. When the last emitter releases
565
- an image, its patch is removed. Loads completing after a URL change, removal or shutdown are ignored.
566
- Failed images are reported and remain transparent until the URL changes or the emitter is re-added.
567
- Untextured emitters use a white texel; pending images use a transparent texel.
568
-
569
- At `FrameStart`, after the scene context has established transform rows, the system calls
570
- `GPUParticleSystem.sync_membership()`, uploads changed atlas pixels, and refreshes all patch regions.
571
- This occurs before particle simulation stages emitter rows. Every repack or resize therefore updates
572
- existing emitters too. The extension enables the renderer's GPU particle feature for its scene;
573
- rendering uses the existing AVBOIT path. The system owns a `GPUTextureContext` for the atlas, using
574
- its resize and cached-view handling and advancing its version after each pixel upload. Shutdown
575
- restores the default atlas and destroys the owned context. A replacement device gets a fresh upload
576
- from the retained CPU atlas.
577
-
578
- `GPUParticleSystem` itself does no asset loading or packing. Standalone callers can still supply
579
- `atlas`/`atlas_sampler` and call `registry.set_atlas_region()` themselves. The `atlas` binding remains
580
- a `GPUTextureView`, obtained from its owner's context; the default white texture is also owned
581
- through `GPUTextureContext`.
582
-
583
- ### Ownership: `GPUParticleSystem`
584
-
585
- Constructed once against a `GPUSceneContext`, in the mould of `ReSTIRDI`. It owns every persistent
586
- buffer — pool, alive lists, dead list, counters, both indirect-args buffers, the coherence and sort
587
- scratch, the emitter database (both tables), the packed program buffer, the spawn-command ring — and grows the ones that depend on the emitter set inside `execute`. The
588
- default atlas is a 1x1 white texel. `execute({graph, camera, scene_database, color, depth, dt})`
589
- sweeps the scene, pushes what changed (on a command context opened only when something needs
590
- one), and records the whole frame through `graph_particles`, returning the colour handle after
591
- the draw. An integrator does not encode a particle pass by hand.
592
-
593
- ## File layout
594
-
595
- ```
596
- src/shade/renderer/particles/
597
- DESIGN.md
598
- ParticleConstants.js pool stride, reg count, workgroup sizes, limits
599
- isa/ParticleVMISA.js opcode table (single source of truth)
600
- isa/ParticleProgram.js Program value type + (de)serialization to u32
601
- isa/ParticleAssembler.js tiny assembler (build programs in tests/library)
602
- vm/ParticleVMReference.js JS reference executor
603
- vm/chunk_particle_vm.js WGSL interpreter CodeChunk (parity target)
604
- vm/chunk_particle_vm_registers_fast.js register file in var<private> — the interpreter that runs
605
- vm/chunk_particle_vm_registers_wide.js register file in storage — the fallback
606
- shaders/chunk_particle_simulate_step.js one particle's simulate, shared by both simulate shaders
607
- shaders/chunk_particle_emit_step.js one newborn's emit, shared by both emit shaders
608
- graph/ParticleNodeDescription.js node type base (ports + lowering)
609
- graph/ParticleNodeRegistry.js standard node library (a NodeRegistry)
610
- graph/particle_graph_authoring.js authoring sugar over core/model/node-graph
611
- graph/compile_particle_graph.js graph → Program (register alloc, const intern)
612
- layout/ParticleLayout.js attribute schema ↔ record word offsets
613
- data/PARTICLE_EMITTER_STRUCT.js the emitter row + flag/blend/projection encoding
614
- data/particle_emitter_record.js emitter + context -> one row of that struct
615
- data/PARTICLE_EMITTER_STATE.js the GPU-owned per-emitter state row (second table of the database)
616
- data/PARTICLE_DATABASE_SPEC.js the GPUDatabase schema, and the emitter table's descriptor
617
- data/PARTICLE_COUNTERS.js counter slot indices
618
- data/particle_header.js record header encoding (row + generation), JS + WGSL
619
- coherence/particle_program_buckets.js CPU reference for the simulation-order bucketing
620
- coherence/shader_particle_bucket.js reset / histogram / pad / scatter / fill
621
- coherence/graph_particle_bucket.js frame-graph wiring, around graph_prefix_scan_csdldf
622
- runtime/ProgramHeap.js the one packed [header][constants][code] buffer, allocator-backed
623
- runtime/ProgramPlacement.js where one program sits in it — what an emitter record carries
624
- runtime/ParticleEmitter.js one emitter — a Node3D and an effect, with no GPU in it
625
- runtime/GPUParticleEmitterContext.js what a registry assigns one: row, generation, placement, patch
626
- runtime/EmitterRegistry.js GPUDatabase(emitters) + the live emitters + their contexts + the program heap
627
- runtime/create_particle_effect.js layout + INIT/UPDATE graphs → compiled effect
628
- shaders/shader_particle_spawn_commands.js CPU bursts → pending spawns
629
- shaders/shader_particle_emitter_tick.js per-emitter spawn decision
630
- shaders/shader_particle_emit.js per-newborn INIT, sized by the scan
631
- shaders/shader_particle_simulate.js per-particle UPDATE + compaction + orphan kill
632
- shaders/shader_particle_finalize.js reset counters (+ budget) / build indirect args
633
- shaders/shader_particle_render.js billboard pipeline
634
- shaders/chunk_particle_setup_context.js VM builtins from the emitter's node, + the RNG seed
635
- shaders/chunk_particle_emitter_age.js the tick pass's running age, off the emitter_state table
636
- shaders/chunk_particle_spawn_budget.js take from the frame's spawn budget; both ticks spend it
637
- warmup/warmup_schedule.js ticks and tick length for a batch of pre-warms
638
- warmup/shader_particle_warmup.js init / reset / tick over the host's queue / migrate
639
- warmup/chunk_particle_warmup_step.js one particle's advance: per-lane decode, jittered tick
640
- warmup/shader_particle_warmup_advance.js the advance over the fast file (+ _wide.js)
641
- warmup/graph_particle_warmup.js the loop, recorded on a frame that registered a pre-warm
642
- shaders/chunk_particle_curve_disabled.js the no-curves-bound CURVE provider
643
- shaders/chunk_particle_billboard_vertex.js the billboard vertex stage, shared by every draw
644
- shaders/shader_particle_avboit_*.js occupancy / splat / draw through AVBOIT
645
- bounds/* the measured emitter box: accumulate, publish
646
- sort/* render-order depth sort
647
- graph_particles.js simulate (to the draw args) + the standalone sort and draw
648
- graph_particles_avboit.js the AVBOIT side channel: occupancy, splat, draw
649
- GPUParticleSystem.js the feature: owns the buffers, sweeps the scene, execute()
650
- *.spec.js co-located tests (JS + emulator + software device)
651
- ```
652
-
653
- ## Testing
654
-
655
- - Pure JS: layout packing, assembler, register allocator, node compiler, heap/emitter/registry
656
- bookkeeping (via `SoftwareGPUDevice`), the header encoding.
657
- - WGSL via emulator: VM opcode parity (JS vs WGSL), curve sampling, sort key/depth, atlas+flipbook
658
- UV math, billboard expansion, soft-depth fade, and every compute pass on small fixtures — the tick
659
- (rate integration, generations, pending, budget, culling, page-group walk), emit (expansion over
660
- the scan, node position, orphans), simulate, spawn commands, finalize, bucket, sort. Each such spec
661
- carries its own copy of the scaffolding: a real scene with a `GPUSceneContext` and a real registry,
662
- laid out by `gpu_database_words` the way the GPU holds both databases, so the shaders read rows the
663
- real writers wrote.
664
- - Multi-frame capstone (`particle_lifecycle.spec`): reset → commands → tick → (JS scan) → emit →
665
- simulate across frames, with the GPU deciding to spawn, a CPU burst, the budget, and death.
666
- - Orchestration on `SoftwareGPUDevice`: `GPUParticleSystem.spec` (scene sweep, growth, uploads only
667
- when needed, the ring), `graph_particles.spec` (pass sequence, indirect sizing, the two orderings,
668
- every persistent buffer imported exactly once), and the playground's `particle_scene.spec` (the
669
- integration: geometry → hierarchy → particles → draw).
670
- - Cross-parity meta-tests: every ISA opcode implemented by both executors.
671
-
672
- Regression checks use the existing tools: seeded software-device buffer copies for list growth,
673
- recorded pass ordering for program edits, emulator math for blend/fade and unknown bounds, and real
674
- AssetManager/TextureAtlas bookkeeping with a stub image loader for asynchronous atlas lifetimes.
675
- No full software WebGPU implementation or cross-frame shader simulator is added. The emulator's
676
- `subgroupBroadcastFirst` is identity, so grouping correctness is checked at its input contract.
677
-
678
- End-to-end GPU is out of scope for node (no binary GPU dep); emulator + JS references give the
679
- coverage the brief asks for.
680
-
681
- ### Emulator caveats discovered (all correct on real hardware)
682
-
683
- The in-repo WGSL emulator (unit-test tool, not GPU parity) has limitations the shaders are written
684
- around, documented in the memory note: `switch` case bodies are dropped (use if/else); repeated
685
- read-modify-write of one location across a loop is corrupted (RNG is verified in isolation, not
686
- through the interpreter loop); u32 multiply is imprecise for large operands (PCG can't be checked
687
- through it); `/` is FLOAT division, so integer division only truncates where the result is stored
688
- straight into a typed array — `(n + 63u) / 64u * 64u` reads as 66 there and 64 on a device, and a
689
- binary search's `mid` is a shift, not a divide; matrix column indexing (`m[0]`) is not evaluated,
690
- so an axis scale is read as `m * vec4(1,0,0,0)`. `array<atomic<u32>>` bindings are modeled as
691
- `{value}` objects in tests; a pass that binds the same buffer as plain `array<u32>` gets a
692
- `Uint32Array`. The scene context writes a row's `global` one publish behind (the hierarchy pass
693
- recomposes it on a device), so a node's matrix under the emulator is the one it had when its row
694
- was seeded.
695
-
696
- ### Renderer integration
697
-
698
- Done: `Renderer.feature_particles_enabled` (off by default), `Renderer.particles(scene_ctx)` creating
699
- one system per scene context on first use, `simulate` recorded before the transparency pass and the
700
- system passed to the AVBOIT orchestrator as its side channel. Remaining:
701
- 1. For lit emitters, compile a simulate/render variant that swaps the disabled curve/lighting hooks
702
- for `chunk_particle_curve_animation` (+ bind `GPUAnimationManager.database.buffer`) and a
703
- shading-chunk-backed lighting hook (+ the `SHADING_LIGHT_RESOURCE_GROUP`).
1
+ # GPU-Driven Particle System — Design
2
+
3
+ A specialized, deeply-integrated, GPU-driven particle engine for the v2 renderer.
4
+ Destiny-style: **one** simulation dispatch runs **all** particles of **all** emitters by
5
+ interpreting a compact per-emitter bytecode program (a shader VM). Effects are authored as
6
+ node graphs and lowered to that bytecode by a testable JS compiler. Emitters are scene nodes, and
7
+ the decision to spawn is the GPU's.
8
+
9
+ ## Goals (from the brief)
10
+
11
+ - Single simulation execution shader (data-driven VM, not one-pipeline-per-effect).
12
+ - Projection modes: camera-facing billboard (+ velocity-stretched, axis-locked, world-oriented).
13
+ - Depth sorting for correct transparency, built on the `csdldf` prefix scan.
14
+ - Node-based composition → shader VM (compile graph → bytecode).
15
+ - Frustum culling of spawning (optional, flagged, sleep-when-culled + bounded catch-up).
16
+ - Flipbooks (atlas sub-rect + frame grid).
17
+ - Texture lookup via atlas (region table; mirrors meep `AtlasPatch` uv offset/scale).
18
+ - Soft-depth fade (flagged).
19
+ - `GPUDatabase` for CPU-managed, low-churn tables (emitters).
20
+ - Scene-node / bone attachment: an emitter **is** a `Node3D`; its world matrix is its row of the
21
+ scene's `transforms` table, composed by the GPU hierarchy pass.
22
+ - Distinct per-particle lifecycle: **no** built-in age/position — the schema + program define state.
23
+ - Billboard renderer reuses engine shader chunks incl. lighting (forward-shaded, optional per emitter).
24
+ - Blend modes (alpha + additive) via a **single** premultiplied blend setup.
25
+ - `AnimationCurve` sampling during simulation (reuse `chunk_animation_curve_evaluate`).
26
+ - Spawning decided on the GPU, per emitter, with the CPU able to ask for bursts through the same path.
27
+ - No per-frame allocation on the CPU beyond the frame graph's own records.
28
+
29
+ ## Layered architecture
30
+
31
+ ```
32
+ authoring: NodeGraph (nodes+ports) ──compile──▶ Program (bytecode: INIT + UPDATE)
33
+ standard node library │
34
+ ▼
35
+ runtime CPU: ParticleEmitter (a Node3D) ── Scene ──▶ GPUSceneContext (transforms table)
36
+ EmitterRegistry ── GPUDatabase(emitters + emitter_state) ─┤ ProgramHeap ([header][constants][code])
37
+ GPUParticleSystem (owns every buffer, records the frame)
38
+ ▼
39
+ runtime GPU: particle pool (array<u32>, uniform stride) + alive/dead lists + counters
40
+ emitter state table (GPU-owned: generation, accumulator, age, sleep, pending)
41
+ ┌─ spawn commands (CPU bursts → pending)
42
+ ├─ emitter tick (per row: cull on the published box, integrate rate, take budget → count)
43
+ ├─ scan (csdldf) → emit dispatch
44
+ ├─ emit (VM INIT program) ── per newborn: find emitter in scan, pop dead → init → append alive_new
45
+ ├─ simulate (VM UPDATE program) ── update → (kill|orphan→dead | keep→alive_new)
46
+ ├─ bucket (histogram → csdldf scan → scatter) ── alive → program-coherent sim list
47
+ ├─ bounds (per particle → per emitter box, published to the state row)
48
+ ├─ build indirect (dispatch + draw args)
49
+ ├─ sort (histogram → csdldf scan → scatter) [render order]
50
+ └─ render (instanced billboard quads, forward + optional lighting + soft depth)
51
+ ```
52
+
53
+ ### Particle storage — single shared pool, uniform stride
54
+
55
+ All particles live in one `array<u32>` pool with a fixed `RECORD_WORDS` stride (config, default 32
56
+ words = 128 B). Word 0 is a reserved header holding the owning **emitter row** and that row's
57
+ **generation** (`data/particle_header.js`); user attributes occupy words `1..RECORD_WORDS`. A single
58
+ pool + single alive/dead lists means **one** emit dispatch and **one** simulate dispatch cover every
59
+ emitter — the Destiny single-shader property.
60
+
61
+ Lifecycle uses the canonical GPU dead-list / alive-list ping-pong:
62
+ - `dead_list`: stack of free slot indices (+ atomic count).
63
+ - `alive_list[2]`: ping-pong compact lists of live slot indices (+ counts).
64
+ - `counters`: atomics {alive_count, dead_count, spawn budget, …} + derived indirect dispatch/draw args.
65
+
66
+ Emit: pop dead → run INIT → append `alive_new`. Simulate: read the **coherence list** last frame's
67
+ bucketing pass built from `alive_current` → run UPDATE → kill pushes to `dead`, survival appends to
68
+ `alive_new`. Swap lists each frame.
69
+
70
+ **The loop is strictly GPU-driven.** The build-indirect pass snapshots the padded length of the
71
+ coherence list into `counters[ALIVE_IN]` and writes the indirect dispatch args next frame's simulate
72
+ consumes (`dispatchWorkgroupsIndirect`) plus the indirect draw args this frame's render consumes
73
+ (`drawIndirect`, sized by the *unpadded* alive count); simulate bound-checks against the ALIVE_IN
74
+ snapshot. The sort's per-particle passes take the dispatch the bucket reset writes from the alive
75
+ count. Sizing any of this from a CPU readback is a correctness bug, not a shortcut: the readback is
76
+ inevitably stale, so a growing population orphans the tail of the alive list (slots leak into limbo)
77
+ and a shrinking one re-simulates stale entries — double-freeing slots so two live particles later
78
+ share one record. The emit pass's dead-list pop guards against wrapped counters
79
+ (`old == 0 || old > capacity`), so racing an exhausted pool drops spawns cleanly instead of handing
80
+ out garbage slots.
81
+
82
+ ### Spawning is the GPU's decision
83
+
84
+ The CPU never walks the emitters per frame and never counts particles. Per frame:
85
+
86
+ 1. **spawn commands** (`shaders/shader_particle_spawn_commands.js`, only when the host queued any):
87
+ one thread per `[row, count]` pair the host authored, `atomicAdd` into that row's `PENDING` in
88
+ the GPU-owned emitter state. This is `GPUParticleSystem#spawn` — the CPU's facade over the GPU
89
+ path, not a second path: a burst goes through the same budget, expansion and INIT program as the
90
+ emitter's rate. It is also the channel a VM opcode would use to spawn from the GPU (a sub-emitter
91
+ or trail would `atomicAdd` its target's `PENDING`).
92
+ 2. **emitter tick** (`shaders/shader_particle_emitter_tick.js`): one thread per emitter table row,
93
+ walked with the table's page iterator so an unallocated page costs one word and a freed row one
94
+ bit. Per live row: recognise the row by generation (a mismatch starts the state over — new
95
+ emitter, or a reused row, with nothing for the CPU to clear); take `PENDING`; publish the
96
+ bounding box emit and simulate measured last frame into the state row and reset the accumulator
97
+ (see **Bounds**); if the emitter opts into culling, test that box against the camera frustum and
98
+ sleep while culled, paying back a bounded catch-up on wake; integrate
99
+ `spawn_rate * dt`; turn the whole part of the accumulator into this frame's count, clamped by the
100
+ **spawn budget** the reset pass seeded from the dead count — so an emitter is only ever debited
101
+ for particles the pool can hold; write the count to `spawn_counts.elements[row]`.
102
+ 3. **scan**: `graph_prefix_scan_csdldf` over the counts — each row's batch ends at its inclusive sum,
103
+ and the last sum is the number of particles born.
104
+ 4. **spawn args**: `shader_prefix_sum_to_command` (the rasteriser's own) turns that total into the
105
+ emit dispatch.
106
+ 5. **emit**: one thread per newborn, locating its emitter by `lower_bound_branchless` over the sums
107
+ and its batch fraction from the sum before — the meshes-to-meshlets expansion applied to
108
+ particles. No per-emitter cap, no thread loops over its particles.
109
+
110
+ On a frame that registered emitters with a `prewarm`, and only then, a loop of the warm-up pipeline
111
+ (`warmup/`) runs after the frame's own simulate: the same decision-scan-emit shape per tick over a
112
+ transient population, plus an advance of its own, then a migration into the main pool. See
113
+ **Pre-warm**.
114
+
115
+ The host contributes two numbers to all of that: the emitter table's element capacity (the scan's
116
+ length — static topology) and the length of the command list it wrote itself.
117
+
118
+ ### The emitter state (`data/PARTICLE_EMITTER_STATE.js`)
119
+
120
+ What the GPU integrates per emitter — generation, accumulator, age, sleep, pending — is the
121
+ `emitter_state` table of the same database, one row per emitter at the emitter's row, **not**
122
+ fields of the emitter record. The record is CPU-shadowed and restaged whole whenever the author
123
+ changes a field, and that restage would overwrite an accumulator the GPU has been adding to for a
124
+ hundred frames. The scene's transforms table mixes CPU fields with a GPU-written `global` only
125
+ because `global` is recomputed from scratch every frame; nothing here is. In its own table the CPU
126
+ touches a row exactly twice — a zero record when the emitter's row is claimed, the table's own
127
+ zero-fill when it is released — and the database's resize copies pages on the GPU, so growth keeps
128
+ what the tick has written. Nothing in it is atomic: each row is one tick thread's, and the
129
+ spawn-command pass is handed one command per row (the system merges bursts on the CPU). The
130
+ emitter's `age` here is the VM's `EMITTER_AGE` builtin — and the one field the warm-up pipeline
131
+ writes, as it runs an emitter's clock forward (see **Pre-warm**).
132
+
133
+ ### Eight storage buffers
134
+
135
+ A compute stage gets eight storage buffers by default, and the design does not ask for more. Two
136
+ things keep every pass under it: the emitter record and the GPU-owned state are two tables of one
137
+ database buffer, read through one binding; and the VM's constants and code are one packed buffer —
138
+ `[header][constants][code]`, the header's first word saying where the code starts
139
+ (`ParticleConstants.PARTICLE_PROGRAM_HEADER_WORDS`). Emit and simulate bind exactly eight; their
140
+ wide-register variants bind nine, which is the one place the system spends past the default and is
141
+ still inside the ten the engine demands of an adapter. `particle_shaders_wgsl_validity.spec` pins
142
+ the count for every pass.
143
+
144
+ ### The VM (`ParticleVM`)
145
+
146
+ Register machine, emulator-safe (bounded instruction loop + big `switch`, no subgroup ops, no
147
+ recursion, no data-dependent jumps — conditionals are predicated via CMP/SELECT, destruction via
148
+ `KILL`).
149
+
150
+ - Register file: `array<f32>`, a register being one slot and a value of width *w* occupying *w*
151
+ consecutive ones. The interpreter reaches it through `vm_reg_read` / `vm_reg_write`, which a
152
+ **backing chunk** supplies, and the two backings are the two interpreters:
153
+ `vm/chunk_particle_vm_registers_fast.js` puts the file in `var<private>`
154
+ (`PARTICLE_VM_FAST_REGISTER_SLOTS`, dynamically indexed, so its size is the shader's register
155
+ footprint), and `vm/chunk_particle_vm_registers_wide.js` puts it in a storage buffer
156
+ (`PARTICLE_VM_REGISTER_SLOTS` — every slot an 8-bit operand base can name). See **Two
157
+ interpreters** below.
158
+ - Instruction: fixed **4×u32** words `[op | width << 8 | kinds << 16 | dst << 24, s0, s1, s2]`.
159
+ Word 0 packs the opcode (bits 0..7), the lane count (8..10), each source's `VM_KIND` (two bits per
160
+ slot at 16..21) and the destination slot (24..31). A source word means what its kind says: `REG`
161
+ is a base slot plus a swizzle, `IMM` is an `f32` bit pattern read in every lane, and `NONE` is a
162
+ raw field the opcode reads for itself (attribute word-offset+comp / curve handle / builtin id /
163
+ pool index).
164
+ - Constant pool: `array<vec4f>` per program, referenced by `LOAD_CONST`.
165
+ - Two entry sections per program: **INIT** and **UPDATE** (offsets in the emitter record). INIT
166
+ gives an already-allocated record its state; it has no say in whether the particle exists.
167
+ - Builtins (`LOAD_BUILTIN`): DT, EMITTER_AGE, EMITTER_POSITION, EMITTER_DIRECTION, EMITTER_UP,
168
+ PARTICLE_INDEX, SPAWN_FRACTION, TIME … assembled per-invocation by the host shader. The emitter's
169
+ position and axes are its node's `global` out of the scene database.
170
+ - `CURVE dst, handle, s0`: `dst.x = animation_curve_evaluate(curves[handle], s0.x)` — reuses the
171
+ engine's GPU AnimationCurve chunk + database.
172
+
173
+ #### Two interpreters
174
+
175
+ The register file normally lives in shader registers, which is what makes it fast and what makes it
176
+ small. An effect whose compiled program wants more slots than it holds is not an error and not an
177
+ authoring ceiling: the whole system swaps emit and simulate for variants whose file is a storage
178
+ buffer, and everything runs, slowly.
179
+
180
+ - **The trigger is the set, not the program.** `ProgramHeap` reports `max_register_count` over the
181
+ programs it holds, and `GPUParticleSystem` is where that number is compared against
182
+ `PARTICLE_VM_FAST_REGISTER_SLOTS` — the store has no opinion about what a register budget means,
183
+ and which backing a pass is compiled against is the VM's business. One hungry effect moves every
184
+ emitter onto the wide pair; there is one pipeline per pass and it is chosen once, per frame, in
185
+ `graph_particles`. It moves back when that effect is unregistered: the heap knows what it holds,
186
+ so the number is a maximum over the live set rather than a high-water mark.
187
+ - **The launch is bounded.** The wide passes need a register file per invocation in flight, so they
188
+ do not run one lane per particle — that would size the buffer by pool capacity. They launch
189
+ `PARTICLE_VM_WIDE_LAUNCH_LANES` lanes whatever the workload and walk their input with a grid-stride
190
+ loop, so the file is `LANES * SLOTS * 4` bytes and nothing else. The stride is a whole number of
191
+ workgroups, which is what keeps a wave inside one coherence bucket and the wave-uniform section
192
+ decode sound.
193
+ - **Slot-major, and never cleared.** `vm_registers[slot * LANES + lane]`, so a wave reading one slot
194
+ reads a coalesced run. The buffer is transient and no pass zeroes it: a register is undefined until
195
+ written, and a program that reads one first is corrupt. That is the one place the two backings are
196
+ not interchangeable — the fast file zeroes, so the reference executor and the fast VM agree on a
197
+ value the wide VM has no opinion about.
198
+ - **Held to the same bar.** `vm/chunk_particle_vm.spec.js` runs every parity case, fuzz included,
199
+ against BOTH backings; `shaders/particle_wide_register_passes.spec.js` runs the wide passes against
200
+ the fast ones under the emulator and exercises the grid-stride loop.
201
+
202
+ **Parity strategy.** The opcode numbers + operand semantics live in one source of truth
203
+ (`isa/ParticleVMISA.js`). Two executors implement it: a JS reference VM (`vm/ParticleVMReference.js`)
204
+ and a hand-written WGSL interpreter (`vm/chunk_particle_vm.js`). A spec runs identical programs
205
+ through both (WGSL via the emulator) and asserts bit-for-bit-close equality. A meta-test asserts both
206
+ executors cover every opcode in the ISA table.
207
+
208
+ ### Node graph → bytecode
209
+
210
+ A core `NodeGraph` holds particle nodes (ports + parameters) for INIT and UPDATE. `compile()`:
211
+ 1. topologically orders reachable nodes,
212
+ 2. lowers each node to VM ops (nodes emit via a small builder that allocates result registers),
213
+ 3. linear-scan register allocation with liveness,
214
+ 4. interns constants into the pool,
215
+ 5. resolves curve/attribute references,
216
+ 6. emits `Program { init, update, constants, reg_count, attributes, curves }`.
217
+
218
+ Standard node library: constants, attribute get/set, arithmetic/vector/math, random (uniform/sphere/
219
+ disk/cone), curve-sample, noise (curl), integrate, gravity, emit-shapes, compare/select, kill-if.
220
+
221
+ ### Emitters are scene nodes
222
+
223
+ `ParticleEmitter extends Node3D`. Add one to a `Scene` — on its own or under a joint — and the
224
+ scene context's membership sweep gives it a `transforms` row like any node; the GPU hierarchy pass
225
+ composes its `global` from its own transform and its parents' every frame, animation included.
226
+ `GPUParticleSystem` sweeps the scene's node list when the scene's membership version moves, gives
227
+ each emitter node a row in the emitter table pointed at its `transforms` row, and retires the row
228
+ when the node leaves. Attaching an emitter to a bone is `joint.addChild(emitter)`; there is no
229
+ per-frame transform copy, and an emitter under a GPU-animated joint follows it — which the old
230
+ CPU-side copy of `transform_global` could not do, since the CPU never updates that matrix for a
231
+ node the GPU owns.
232
+
233
+ **The node knows nothing about the GPU.** A `ParticleEmitter` is a transform, an attribute
234
+ `layout`, a compiled `program`, a `texture` URL, and the emission parameters — all of it the
235
+ author's, all of it meaning the same thing with no device present, all of it what a saved effect is
236
+ made of. Everything the system assigns is a `GPUParticleEmitterContext`, held by the
237
+ `EmitterRegistry` against the emitter and reached through `registry.context(emitter)`: the table
238
+ row, the generation, the program placement, the node's `transforms` row, the atlas patch. So an
239
+ emitter cannot be saved with a row in it, cannot be given one by hand, and "is this emitter live?"
240
+ is asked of the registry rather than read off a `-1` that anything could have written. The record
241
+ writer is where the two meet — `write_emitter_record(record, emitter, context)`, next to the struct
242
+ it fills.
243
+
244
+ Restaging is driven by `Node3D#version`, the node's own change counter, which the renderer already
245
+ reads to decide whether a transform needs re-uploading; the context remembers the version its row
246
+ was staged from. An author who edits an emitter says so the way they say it for any node —
247
+ `needsUpdate = true` — and there is no second flag to forget. A change to the context itself (a
248
+ `transforms` row, a repacked atlas patch) moves no version, so the context marks itself instead.
249
+
250
+ ### Emitter table (`GPUDatabase`)
251
+
252
+ One `emitters` table (CPU-managed, low churn). Row = `PARTICLE_EMITTER_STRUCT`: program
253
+ offsets+lengths, constant-pool offset, seed, spawn rate, **node** (transforms row), **generation**,
254
+ flags (lighting/soft/sort/cull/blend/projection), render binding (attr offsets for
255
+ position/size/color/rotation/frame/velocity), atlas region + flipbook grid, and the program id
256
+ the coherence pass buckets on. The atlas region is resolved, not authored: the emitter names an
257
+ image and the system packs it, the way `ShaderManager` packs the old engine's sprites, so a repack
258
+ moves every patch without any emitter changing. The seed is resolved too — an emitter left at 0 gets
259
+ its registration's generation, which no other emitter of the registry shares, rather than being
260
+ written to. Programs+constants live in a packed `ProgramHeap` buffer
261
+ (variable length → not a fixed-stride table).
262
+
263
+ The struct is the ONLY statement of that layout. The JS writer is the database's
264
+ `write_wgsl_type_value`, and the WGSL reader — the `ParticleEmitter` struct declaration plus
265
+ `database_read_..._element` — is generated from the same struct by the table descriptor. There is no
266
+ hand-packed word offset table to drift out of step with the shaders, and no hand-written reader.
267
+ The bucket histogram, which runs once per live particle and wants only `program_id`, reads it
268
+ through `chunk_read_field` rather than decoding a whole row.
269
+
270
+ **Editing the set at runtime is the point of using a table.** `EmitterRegistry.add` stages one
271
+ record and `remove` frees one row; neither touches the others. Rows are stable for an emitter's
272
+ lifetime because a live particle's record header holds the row of the emitter that spawned it —
273
+ renumbering a survivor would re-point its particles at somebody else. A removed row is handed back
274
+ to the next `add`, so the id space does not grow without bound.
275
+
276
+ **Generations make that reuse safe.** Every add takes the next value of one running counter, never
277
+ zero, and the particles of an emitter carry its generation in their header. The simulate pass kills
278
+ any particle whose generation is not its row's current one — the row is zero-filled (generation 0)
279
+ or a later tenant's — so a hard removal frees the emitter's particles within a frame, and a reused
280
+ row cannot adopt a stranger's. Graceful retirement (rate to zero, remove once the tail has died)
281
+ remains an author's choice rather than a correctness rule.
282
+
283
+ **Programs are edited the same way.** `ProgramHeap` is a pair of range allocators
284
+ (`OffsetAllocator`) over the two sections of the packed buffer, keyed by a `HashMap` on the bytecode
285
+ itself — `ParticleProgram` answers `equals`/`hash`, so content deduplication is the map's doing
286
+ rather than a stringified key's. `acquire` takes a reference and `release` gives one back, so the
287
+ last emitter of an effect is what unpacks it, and `EmitterRegistry.set_program` swaps an emitter onto
288
+ different bytecode while it is live. What makes that possible is that a program is **relocatable**:
289
+ instructions carry no absolute addresses and constant references are program-local, so the heap can
290
+ move one by copying its words and rewriting the placement.
291
+
292
+ Two things are deliberately NOT reclaimed. **Program ids** are reused rather than renumbered — they
293
+ are baked into every emitter record, so compacting them would mean rewriting the whole table — and
294
+ `bucketCount()` is sized from the id capacity, the set's peak, so it does not move every time an
295
+ effect comes or goes. An id below the peak with nothing in it is an empty bin, which rounds up to
296
+ zero workgroups and shifts nothing after it. **Section capacity** only rises: a repack (which happens
297
+ when a section has no contiguous run long enough) compacts what is live and grows only if it must,
298
+ so a set that peaked at a hundred effects does not pay for a resize every time it returns to three.
299
+
300
+ A repack is the one change no emitter announces, since it is not an edit to any of them.
301
+ `EmitterRegistry.flush` asks the heap once per frame — `placement_version` — and restages every row
302
+ when it has moved.
303
+
304
+ Programs are immutable after compilation. Both `registry.set_program(emitter, program)` and replacing
305
+ `emitter.program` followed by `needsUpdate = true` are supported. Flush reconciles replacements
306
+ before inspecting heap relocation, since acquiring a replacement may move other programs too.
307
+ The layout must remain compatible with particles already in flight.
308
+
309
+ A live program assignment increments `program_assignment_version`. On the next frame, before reset
310
+ or simulation, the previous compact alive list is bucketed again and its padded count and dispatch
311
+ are rebuilt. This prevents a previously shared wave from executing one emitter's newly edited program
312
+ over another emitter's particles. Ordinary frames keep the existing end-of-frame bucketing only;
313
+ content-identical replacements need no repair. Both bucket runs on an edit frame reuse the latest
314
+ graph handles for their scratch buffers, so each persistent buffer is imported once.
315
+
316
+ The coherence list is persistent GPU state, including its padding. Growing the bucket capacity copies
317
+ it to the replacement buffer before destroying the old one. The old count and indirect dispatch
318
+ remain valid across growth; only the per-frame histogram/cursor/key scratch may be discarded.
319
+
320
+ Every row of a frame is staged through one module-static record (`EmitterRegistry`), so a flush
321
+ allocates nothing however many rows it restages.
322
+
323
+ ### Bounds
324
+
325
+ An emitter has no bounds it could declare. Its particles are simulated in world space by arbitrary
326
+ bytecode: they go where the program sends them, at whatever size it gives them, for as long as it
327
+ keeps them alive. The only bound an author could write down and still be right is an infinite one,
328
+ which culls nothing and reserves every depth slice. So the box is **measured**, on the GPU, every
329
+ frame.
330
+
331
+ Two passes of its own (`bounds/`), at the end of the frame, after emit and simulate have both
332
+ finished appending to the alive list:
333
+
334
+ - **accumulate** — one thread per live particle: read where its quad is and merge it into its
335
+ emitter's accumulator. The centre is the record's `render_position`; the dilation is
336
+ `particle_billboard_radius`, the largest offset `particle_billboard_offset` can produce for that
337
+ size, velocity and projection, over every corner, every roll and any camera. The box has to hold
338
+ the *quads*, not the centres, and it has to be camera-independent, because it is read on a later
339
+ frame. The merge is `aabb_atomic_min_f32` / `aabb_atomic_max_f32`, the same CAS-spin the
340
+ skinned-mesh bounds refresh uses; both load first and return without a CAS once the box holds the
341
+ contribution, so what survives is load traffic, not exchanges. Dispatched indirectly from
342
+ `sim_dispatch_args`, the args the coherence reset wrote from the GPU's own alive count.
343
+ - **publish** — one thread per emitter row: write the accumulation to `bounds_min` / `bounds_max` of
344
+ the state row (`data/PARTICLE_EMITTER_STATE.js` — measured state is GPU-owned state, and on the
345
+ record it would be clobbered by the next restage), and reset the accumulator. The reset is
346
+ unconditional, which is what lets a row that changes hands start clean.
347
+
348
+ **Why a pass and not a tail on emit and simulate.** It was that first. Booking from inside those two
349
+ measures no single population — simulate sees last frame's survivors and emit sees this frame's
350
+ newborns, so the union spans two frames and is looser than either — and it pays six atomics per
351
+ particle inside the pass whose register pressure decides the system's occupancy, to compute
352
+ something that needs none of the VM, none of the scene, and five words of the record. Over the alive
353
+ list instead, the measurement is of exactly the set that is alive and about to be drawn.
354
+
355
+ **Where the accumulator lives.** Six words per emitter row, in the **counters** buffer, after the
356
+ counters themselves (`data/PARTICLE_COUNTERS.js`). It has to be an `array<atomic<u32>>` binding, and
357
+ the emitter database cannot be it — the passes that read a record read it out of an `array<u32>`
358
+ binding, and one buffer under two access types is aliasing. The counters buffer is already bound in
359
+ both bounds passes, already atomic, already the system's cross-pass scratch, so a buffer of its own
360
+ would add a binding and a growth path to say the same thing. It therefore grows with the emitter
361
+ table, by copy — the free-list length in it is GPU-owned and unreconstructable — and the rows the
362
+ growth adds are written to the empty interval, because a zeroed accumulator reads back as a
363
+ degenerate box at the world origin rather than as an absence.
364
+
365
+ **The empty case.** An emitter with no measured particles has infinite, unknown bounds:
366
+ `bounds_min = [-Infinity, -Infinity, -Infinity]`, `bounds_max = [Infinity, Infinity, Infinity]`.
367
+ The CPU stages this on registration and the publish pass restores it whenever the population is
368
+ empty. The tick treats unknown bounds as visible, so a slow emitter can accumulate its first spawn
369
+ without prewarm even when its node is offscreen. A depleted population can restart the same way.
370
+ AVBOIT occupancy skips unknown bounds because there are no live particles to splat. Both readers
371
+ test the sentinel before geometric arithmetic, avoiding NaNs from operations on infinities.
372
+ The sentinel check compares integer bits. The publish pass swaps the empty accumulator's loaded
373
+ endpoints; neither path constructs an infinite floating-point constant, which WGSL rejects at
374
+ shader creation time.
375
+ Measured, nonempty populations continue to use their finite bounds for culling. No bounds are authored.
376
+ Prewarm controls the initial age distribution only; it is not required for culling correctness.
377
+
378
+ ### Pre-warm
379
+
380
+ A continuous effect that starts empty grows into its steady state over one particle lifetime, in
381
+ front of whoever is watching. `ParticleEmitter.prewarm` — seconds, `0` for off — buys that lifetime
382
+ back: an emitter registered with one is run forward by that long before it is first drawn.
383
+
384
+ **A warm-up is the ordinary loop, run over a population of its own.** Spawn, init, advance, per
385
+ tick, until every emitter in the batch has been running for its own `prewarm`. It is *not* a batch
386
+ of particles dumped in one tick and aged: the emitter keeps spawning throughout, what dies is
387
+ replaced as it is in a live scene, and whatever drives the rate — a spawn curve, `EMITTER_AGE`, the
388
+ program itself — drives it here, because it is the same program being run. What the loop leaves is
389
+ the population the emitter would have had.
390
+
391
+ **It is a pipeline apart from the main one** (`warmup/`), because the main one is tuned and must not
392
+ pay for this. Emit and simulate are the passes whose register pressure sets the system's occupancy,
393
+ and their waves are kept program-uniform by the bucketing; a warm-up folded into either would spend
394
+ registers on every frame for work that happens on one, and a loop inside a lane would break the
395
+ uniform flow across a wave. So the warm-up shares code with the main frame and integrates nothing
396
+ into it: the main **emit pass is reused unchanged**, bound to the warm-up's buffers; the tick is a
397
+ pass of the warm-up's own over a host-written queue (`shader_particle_warmup_tick`) rather than over
398
+ the emitter table; the **advance** is a pass of the warm-up's own (`chunk_particle_warmup_step.js`),
399
+ decoding the program section per lane — it reads the raw alive list, so it cannot lean on the
400
+ bucketing's wave-uniformity, and its occupancy is its own business. The main passes' compiled code
401
+ is byte-identical whether `warmup/` exists or not.
402
+
403
+ **The population is transient.** Pool, both alive lists, free list and counters are frame-graph
404
+ transients that exist for the loop and are then moved, record by record, into the main pool by the
405
+ **migrate** pass — a pop from the main free list per particle, and never a push, so it is the one
406
+ point of contact and it is safe by construction. The loop's own free-list traffic never touches the
407
+ main list, and the main passes never see a half-warmed particle.
408
+
409
+ **Bounded.** `warmup/warmup_schedule.js`: at most `PARTICLE_WARMUP_MAX_TICKS` ticks, no shorter than
410
+ `PARTICLE_WARMUP_TICK_SECONDS`, the tick *length* stretching to cover the longest pre-warm in the
411
+ batch rather than the count growing to cover it — the count is what the frame pays for, at six or
412
+ seven dispatches a tick. An emitter takes part in a tick once the loop's clock is within its own
413
+ pre-warm of now, so two emitters queued with different durations each end having run for exactly
414
+ their own; the shorter one sat out the start.
415
+
416
+ **Jittered, so it does not arrive in bands.** Every particle born in one tick would otherwise reach
417
+ the end with exactly that tick's age, and N ticks would read as N cohorts. The advance jitters the
418
+ tick per particle (`PARTICLE_WARMUP_DT_JITTER`, hashed from slot and tick — a hash, not a draw, so
419
+ the program's own random stream is untouched), and the host jitters each emitter's clock and starting
420
+ accumulator (`PARTICLE_WARMUP_EMITTER_JITTER`), so emitters in one batch do not spawn on the same
421
+ beat either. A program does not care what `DELTA_TIME` it is handed; the smear is free.
422
+
423
+ **Where it sits.** After the frame's own emit and simulate, before the bucketing. Migrate appends to
424
+ the same alive list they did, so the bucketing puts the warmed particles in next frame's simulate
425
+ list, the bounds pass measures them into their emitter's box this frame, and the draw draws them
426
+ this frame. The tick also writes the emitter's
427
+ `age` into its state row as it goes, so `EMITTER_AGE` reads as it would have, and the next real tick
428
+ adds its `dt` to a row that says `prewarm`.
429
+
430
+ **Only on the frame that needs it.** The host knows: an emitter is queued when it is registered with
431
+ a `prewarm` (`GPUParticleSystem#sync_membership`), the queue is uploaded and consumed by the frame
432
+ that records the loop, and an ordinary frame records nothing of this.
433
+
434
+ The editor's CPU preview (`editor/particles/effect/ParticleReferenceSimulation.js`) runs the same
435
+ loop over the same schedule before its first frame, and restarts the effect when the duration
436
+ changes: a pre-warm is a statement about how an effect starts, so the only way to show one is to
437
+ start it.
438
+
439
+ ### Culling & attach
440
+
441
+ An emitter that sets `EMITTER_FLAG.CULL` has the box on its state row tested against the camera
442
+ frustum by the tick pass, with the renderer's shared `chunk_aabb3_intersects_frustum`. The tick does
443
+ not measure it — it reads what the bounds passes published at the end of the previous frame — so all
444
+ that is in the spawn path is one frustum test. Because the box is where the particles are rather
445
+ than where the node is, an effect whose particles have drifted on screen keeps spawning even when
446
+ its emitter has not. Culled emitters spawn nothing and accrue sleep; on wake the tick adds a
447
+ catch-up of at most `max_catchup_seconds` of rate. Live particles keep simulating while their
448
+ emitter is culled, so its box keeps moving and keeps being published. HiZ occlusion is not
449
+ written.
450
+
451
+ ### Two orderings, and why they are separate
452
+
453
+ There are two independent permutations of the live particles each frame. They are produced by
454
+ different passes into different buffers and must not be conflated or derived from one another.
455
+
456
+ **Simulation order** (`coherence/`) — camera-independent, changes only when emitters or their
457
+ programs change. The alive list comes out of compaction (`atomicAdd(&counters[ALIVE])`), so its
458
+ order says nothing about which emitter a particle belongs to and a wave routinely holds particles
459
+ running different VM programs. That costs twice: the interpreter's ~40-case opcode `switch`
460
+ diverges, so a wave serially executes the union of the opcodes its lanes want; and `vm_reg[dst]` /
461
+ `vm_reg[s0]` index a VGPR-resident register file dynamically, which AMD reaches with
462
+ `v_movrels`/`v_movreld` behind a waterfall loop whose trip count is the number of *distinct*
463
+ register indices live in the wave. Grouping the wave by program collapses both at once.
464
+
465
+ The grouping is also **load-bearing**: the simulate shader reads its section parameters (code offset,
466
+ instruction count, constant offset) from the wave's first lane through `subgroupBroadcastFirst`
467
+ (`particle_wave_uniform`), which is what lets the driver fetch instruction words through the scalar
468
+ unit, branch the opcode switch on a scalar and index the register file through M0 without a
469
+ waterfall. A wave holding two programs would run one over the other's particles, so the coherence
470
+ pass is part of simulate's correctness. Emit keeps a per-lane decode: its waves span emitters.
471
+
472
+ The bucket key is the **program**, not the emitter: `ProgramHeap` deduplicates by content and hands
473
+ back a `program_id` the emitter record carries, so emitters compiled from the same graph share
474
+ one bucket. Each bucket is padded out to `PARTICLE_SIMULATE_WORKGROUP_SIZE` (64 — a multiple of both
475
+ plausible subgroup widths, so no shader needs to query `subgroup_size`), which is what makes the
476
+ grouping hold at wave granularity: an unpadded boundary lands mid-wave and that wave spans two
477
+ programs again. Padding entries carry `PARTICLE_SIM_PAD_SLOT` and the simulate shader drops them
478
+ *before* the compaction atomics, so they never enter the alive or dead list.
479
+
480
+ It is an **indirection** list. Particle records never move: slots are referenced by the dead list and
481
+ across frames, so relocating a record would leave both pointing at the wrong particle.
482
+
483
+ The three-phase shape is the depth sort's, with the padding added: histogram (atomic per bucket) →
484
+ round each bin up to a whole workgroup → `graph_prefix_scan_csdldf` (reused as-is — the one scan here
485
+ whose only cross-workgroup wait is CSDLDF's bounded spin-then-fallback) → scatter → fill the tails
486
+ with the sentinel. Nothing reads back to the CPU; the only host-supplied number is the bucket count,
487
+ which is static topology.
488
+
489
+ **Render order** (`sort/`) — per-frame, particles sorted back-to-front by view depth so
490
+ premultiplied transparency composites correctly. Counting sort on a quantised view depth
491
+ (`PARTICLE_SORT_BUCKET_COUNT` bins over `GPUParticleSystem#sort_depth_range`): histogram from the
492
+ camera uniform's view matrix, `graph_prefix_scan_csdldf` the bin counts to offsets, scatter (index
493
+ payload). The sorted index list feeds the instanced draw. Both per-particle passes are bounded by
494
+ the alive counter and dispatched by the args the bucket reset wrote from it.
495
+
496
+ Reordering the render list for program coherence would break transparency; depth-sorting the
497
+ simulation list would destroy the coherence. They share no buffer and no counter slot —
498
+ `PARTICLE_COUNTER.SORTED` belongs to the render order, `SIM_PADDED` to the simulation order.
499
+
500
+ ### Rendering
501
+
502
+ One render pipeline, one blend state = premultiplied: `color: (one, one-minus-src-alpha)`,
503
+ `alpha: (one, one-minus-src-alpha)`. Additive vs. alpha is chosen by the shader emitting premultiplied
504
+ RGB with alpha 0 (additive) vs. alpha a (normal) — GPU-driven, no pipeline switch. A 6-vertex quad is
505
+ expanded per instance (billboard/velocity-stretch/axis-lock/world modes). Fragment: sample atlas
506
+ (region + flipbook frame), optional forward lighting via `chunk_shade_standard_material_direct` +
507
+ IBL, optional soft-depth fade. Depth-tested, no depth write. Draws into the HDR color target.
508
+ Soft-depth fading scales both RGB and alpha for premultiplied inputs; straight-alpha and additive
509
+ inputs have alpha scaled before premultiplication. Both draw paths share `particle_fade_color`.
510
+
511
+ ### Drawing under the renderer: an AVBOIT side channel
512
+
513
+ Under `Renderer` the particles are not drawn by the billboard pass above. They go through the AVBOIT
514
+ transparency pipeline (`rasterize/native/avboit/`) as a **side channel** — the contract in
515
+ `AVBOITSideChannel.js` — so they composite order-independently with the transparent meshes and
516
+ against the opaque scene, through the same volume and the same accumulators. Quads are not packed
517
+ into meshlets; the orchestrator asks the channel for its contribution at the three points where the
518
+ meshlet buckets have just made theirs:
519
+
520
+ 1. **occupancy** (`shader_particle_avboit_occupancy`): one thread per emitter row, the emitter's
521
+ measured box (the culling box, off the state row) as a view-depth interval, dilated and OR-ed
522
+ into the occupancy bits. The interval is the box's centre depth plus or minus its extent along
523
+ the view axis — a support projection, not eight transformed corners. Per emitter, not per
524
+ particle — no contention, and a particle that strays lands on a run boundary rather than anywhere
525
+ wrong. Additive emitters mark nothing. The box matters here even for an emitter that never opted
526
+ into culling, which is the second reason it is measured for every emitter rather than only for
527
+ the ones that cull.
528
+ 2. **splat** (`shader_particle_avboit_splat`): the billboard quads rasterized at the volume's
529
+ resolution, coverage from the atlas alpha, extinction through `avboit_splat_slice` — the same
530
+ write path as a surface. Additive particles write nothing. A soft-depth particle is faded here
531
+ against the same scene depth the draw fades it against, one sample per voxel: what the volume
532
+ records has to be the coverage the accumulators are weighted by, or a card that fades out where
533
+ it meets the floor goes on darkening the floor through itself.
534
+ 3. **draw** (`shader_particle_avboit_draw`): the quads at internal resolution, depth-tested read-only
535
+ against the scene (which lets the soft fade read the same depth), fog composited, transmittance
536
+ from the integral. Which of the two it is, is decided by coverage rather than by the blend enum,
537
+ so an additive emitter and a premultiplied sprite with no alpha take the same path — as they do
538
+ in the splat. A covering particle writes `color / norm / extinction` as a surface does, in-scatter
539
+ included, weighted by the coverage it took. An emitting one writes the fourth accumulator,
540
+ **emission**: light attenuated by the fog and by the volume in front of it, and carrying none of
541
+ the fog's own in-scatter, because the resolve adds emission outside the normalization and the
542
+ scene behind already has it. That accumulator is the one addition to AVBOIT's resolve for this.
543
+ The whole decision is `chunk_particle_avboit_contribution`, and it is spec'd on the emulator.
544
+
545
+ No sort: the volume makes order irrelevant, so this path reads the alive list and the draw args
546
+ straight from `simulate`. The billboard vertex stage is one chunk
547
+ (`chunk_particle_billboard_vertex`) shared by the standalone renderer and both AVBOIT passes.
548
+ `GPUParticleSystem` implements the three methods; `Renderer.feature_particles_enabled` runs
549
+ `simulate` before the transparency pass and hands the system to the orchestrator. MBOIT is legacy
550
+ and gets no particle path.
551
+
552
+ ### The ECS: `ParticleEffect` + `Transform64`
553
+
554
+ `engine/graphics3/GPUParticleEmitterSystem.js` is how a game reaches this system. It is an ordinary
555
+ ECS system over the pair `[ParticleEffect, Transform64]` — the component says what an entity emits,
556
+ the transform says where — in exactly the shape `DecalSystem` and `MeshSystem` use for their own
557
+ content. Construct it with `(graphicsEngine, scene, assetManager)` and add it to the EntityManager.
558
+ It is opt-in and coexists with the legacy CPU `ParticleEmitterSystem`.
559
+
560
+ ```js
561
+ entityManager.addSystem(new GPUParticleEmitterSystem(graphics, scene, assetManager));
562
+
563
+ new Entity()
564
+ .add(new Transform64())
565
+ .add(ParticleEffect.from({ ...create_particle_effect({ layout, init, update }), spawn_rate: 400 }))
566
+ .build(dataset);
567
+ ```
568
+
569
+ `engine/graphics/ecs/particles/ParticleEffect.js` carries everything a `ParticleEmitter` node carries
570
+ except the placement: the layout and compiled program, the texture URL, the rate, `emitting`, the
571
+ pre-warm, the flags, the render bindings and the flipbook grid. Nothing on it depends on a device,
572
+ which is what makes it the unit an editor inspects and a prefab holds.
573
+
574
+ **All of it serializes, the compiled program included** — `ParticleEffectSerializationAdapter`,
575
+ registered in `populateEngineSerializationRegistry`. The instruction set was designed to be a durable
576
+ format, so a level carries bytecode the way it carries a mesh's indices and a shipped build needs no
577
+ compiler: authoring is an editor's job and it is not in the runtime. Two things are consequently wire
578
+ format and cannot be renumbered without a version bump and an upgrader — the **opcode numbers** in
579
+ `isa/ParticleVMISA.js`, and the **order of the attributes** in a layout, because a layout assigns its
580
+ word offsets from that order and the record's render bindings are those offsets. The constant pool
581
+ goes out as `Uint32` bit patterns rather than floats, because that is the form the pool is already
582
+ compared and hashed in (`ParticleProgram#constant_words`): `-0` and `+0` are the same number and
583
+ different constants, and whether a NaN's payload survives a round trip through a JS `number` is
584
+ implementation-defined even though V8 happens to preserve it.
585
+
586
+ The system builds one `ParticleEmitter` scene node per linked entity and owns it — placed from the
587
+ entity's transform on `TRANSFORM64_EVENT_CHANGE`, added to the scene when the component has a
588
+ compiled effect, and retired with the entity. `node_of(entity)` hands the node out for the things an
589
+ entity reference cannot express (parenting to the emitter, reading the composed world matrix);
590
+ `burst(entity, count)` is the entity-addressed one-shot, queued and drained after the frame's
591
+ membership sweep so it works before the first frame has run.
592
+
593
+ Component changes are found by **comparison** rather than announced: once a frame each entity's
594
+ fields are compared against what was last written to its node, and the node is marked dirty only
595
+ where they differ, so a component nobody touched costs a run of integer comparisons and no upload.
596
+ That is what lets gameplay write `effect.spawn_rate = 0` from code that has never heard of this
597
+ system.
598
+
599
+ `GPUParticleSystem` itself still finds emitters by sweeping the scene's node list, so a caller with
600
+ no ECS — a playground, a tool — can build `ParticleEmitter` nodes and add them to a `Scene` directly.
601
+ Such an emitter has nothing in the ECS system's atlas and is given the white texel; see below.
602
+
603
+ `shade/playground/particle_ecs/` is the worked example: a bonfire whose flame, smoke, embers and soot
604
+ are four entities, alongside `ShadedGeometry` and `Light` entities in the same scene.
605
+
606
+ The system loads image URLs through `AssetManager` (`GameAssetType.Image`) and shares one patch per
607
+ URL. It reuses `TextureAtlas`, whose packer is `MaxRectanglesPacker`. When the last entity releases
608
+ an image, its patch is removed. Loads completing after a URL change, removal or shutdown are ignored.
609
+ Failed images are reported and remain transparent until the URL changes or the emitter is re-added.
610
+ Untextured emitters use a white texel; pending images use a transparent texel.
611
+
612
+ At `FrameStart`, after the scene context has established transform rows, the system calls
613
+ `GPUParticleSystem.sync_membership()`, uploads changed atlas pixels, and refreshes the patch region of
614
+ every emitter in the scene — the ones it built get their entity's image, and any other gets the white
615
+ texel, because the atlas is the scene's and an emitter it did not build has nothing in it.
616
+ This occurs before particle simulation stages emitter rows. Every repack or resize therefore updates
617
+ existing emitters too. The extension enables the renderer's GPU particle feature for its scene;
618
+ rendering uses the existing AVBOIT path. The system owns a `GPUTextureContext` for the atlas, using
619
+ its resize and cached-view handling and advancing its version after each pixel upload. Shutdown
620
+ restores the default atlas and destroys the owned context. A replacement device gets a fresh upload
621
+ from the retained CPU atlas.
622
+
623
+ `GPUParticleSystem` itself does no asset loading or packing. Standalone callers can still supply
624
+ `atlas`/`atlas_sampler` and call `registry.set_atlas_region()` themselves. The `atlas` binding remains
625
+ a `GPUTextureView`, obtained from its owner's context; the default white texture is also owned
626
+ through `GPUTextureContext`.
627
+
628
+ ### Ownership: `GPUParticleSystem`
629
+
630
+ Constructed once against a `GPUSceneContext`, in the mould of `ReSTIRDI`. It owns every persistent
631
+ buffer — pool, alive lists, dead list, counters, both indirect-args buffers, the coherence and sort
632
+ scratch, the emitter database (both tables), the packed program buffer, the spawn-command ring — and grows the ones that depend on the emitter set inside `execute`. The
633
+ default atlas is a 1x1 white texel. `execute({graph, camera, scene_database, color, depth, dt})`
634
+ sweeps the scene, pushes what changed (on a command context opened only when something needs
635
+ one), and records the whole frame through `graph_particles`, returning the colour handle after
636
+ the draw. An integrator does not encode a particle pass by hand.
637
+
638
+ ## File layout
639
+
640
+ ```
641
+ src/shade/renderer/particles/
642
+ DESIGN.md
643
+ ParticleConstants.js pool stride, reg count, workgroup sizes, limits
644
+ isa/ParticleVMISA.js opcode table (single source of truth)
645
+ isa/ParticleProgram.js Program value type + (de)serialization to u32
646
+ isa/ParticleAssembler.js tiny assembler (build programs in tests/library)
647
+ vm/ParticleVMReference.js JS reference executor
648
+ vm/chunk_particle_vm.js WGSL interpreter CodeChunk (parity target)
649
+ vm/chunk_particle_vm_registers_fast.js register file in var<private> — the interpreter that runs
650
+ vm/chunk_particle_vm_registers_wide.js register file in storage — the fallback
651
+ shaders/chunk_particle_simulate_step.js one particle's simulate, shared by both simulate shaders
652
+ shaders/chunk_particle_emit_step.js one newborn's emit, shared by both emit shaders
653
+ graph/ParticleNodeDescription.js node type base (ports + lowering)
654
+ graph/ParticleNodeRegistry.js standard node library (a NodeRegistry)
655
+ graph/particle_graph_authoring.js authoring sugar over core/model/node-graph
656
+ graph/compile_particle_graph.js graph → Program (register alloc, const intern)
657
+ layout/ParticleLayout.js attribute schema ↔ record word offsets
658
+ data/PARTICLE_EMITTER_STRUCT.js the emitter row + flag/blend/projection encoding
659
+ data/particle_emitter_record.js emitter + context -> one row of that struct
660
+ data/PARTICLE_EMITTER_STATE.js the GPU-owned per-emitter state row (second table of the database)
661
+ data/PARTICLE_DATABASE_SPEC.js the GPUDatabase schema, and the emitter table's descriptor
662
+ data/PARTICLE_COUNTERS.js counter slot indices
663
+ data/particle_header.js record header encoding (row + generation), JS + WGSL
664
+ coherence/particle_program_buckets.js CPU reference for the simulation-order bucketing
665
+ coherence/shader_particle_bucket.js reset / histogram / pad / scatter / fill
666
+ coherence/graph_particle_bucket.js frame-graph wiring, around graph_prefix_scan_csdldf
667
+ runtime/ProgramHeap.js the one packed [header][constants][code] buffer, allocator-backed
668
+ runtime/ProgramPlacement.js where one program sits in it — what an emitter record carries
669
+ runtime/ParticleEmitter.js one emitter — a Node3D and an effect, with no GPU in it
670
+ runtime/GPUParticleEmitterContext.js what a registry assigns one: row, generation, placement, patch
671
+ runtime/EmitterRegistry.js GPUDatabase(emitters) + the live emitters + their contexts + the program heap
672
+ runtime/create_particle_effect.js layout + INIT/UPDATE graphs → compiled effect
673
+ shaders/shader_particle_spawn_commands.js CPU bursts → pending spawns
674
+ shaders/shader_particle_emitter_tick.js per-emitter spawn decision
675
+ shaders/shader_particle_emit.js per-newborn INIT, sized by the scan
676
+ shaders/shader_particle_simulate.js per-particle UPDATE + compaction + orphan kill
677
+ shaders/shader_particle_finalize.js reset counters (+ budget) / build indirect args
678
+ shaders/shader_particle_render.js billboard pipeline
679
+ shaders/chunk_particle_setup_context.js VM builtins from the emitter's node, + the RNG seed
680
+ shaders/chunk_particle_emitter_age.js the tick pass's running age, off the emitter_state table
681
+ shaders/chunk_particle_spawn_budget.js take from the frame's spawn budget; both ticks spend it
682
+ warmup/warmup_schedule.js ticks and tick length for a batch of pre-warms
683
+ warmup/shader_particle_warmup.js init / reset / tick over the host's queue / migrate
684
+ warmup/chunk_particle_warmup_step.js one particle's advance: per-lane decode, jittered tick
685
+ warmup/shader_particle_warmup_advance.js the advance over the fast file (+ _wide.js)
686
+ warmup/graph_particle_warmup.js the loop, recorded on a frame that registered a pre-warm
687
+ shaders/chunk_particle_curve_disabled.js the no-curves-bound CURVE provider
688
+ shaders/chunk_particle_billboard_vertex.js the billboard vertex stage, shared by every draw
689
+ shaders/shader_particle_avboit_*.js occupancy / splat / draw through AVBOIT
690
+ bounds/* the measured emitter box: accumulate, publish
691
+ sort/* render-order depth sort
692
+ graph_particles.js simulate (to the draw args) + the standalone sort and draw
693
+ graph_particles_avboit.js the AVBOIT side channel: occupancy, splat, draw
694
+ GPUParticleSystem.js the feature: owns the buffers, sweeps the scene, execute()
695
+ *.spec.js co-located tests (JS + emulator + software device)
696
+ ```
697
+
698
+ ## Testing
699
+
700
+ - Pure JS: layout packing, assembler, register allocator, node compiler, heap/emitter/registry
701
+ bookkeeping (via `SoftwareGPUDevice`), the header encoding.
702
+ - WGSL via emulator: VM opcode parity (JS vs WGSL), curve sampling, sort key/depth, atlas+flipbook
703
+ UV math, billboard expansion, soft-depth fade, and every compute pass on small fixtures — the tick
704
+ (rate integration, generations, pending, budget, culling, page-group walk), emit (expansion over
705
+ the scan, node position, orphans), simulate, spawn commands, finalize, bucket, sort. Each such spec
706
+ carries its own copy of the scaffolding: a real scene with a `GPUSceneContext` and a real registry,
707
+ laid out by `gpu_database_words` the way the GPU holds both databases, so the shaders read rows the
708
+ real writers wrote.
709
+ - Multi-frame capstone (`particle_lifecycle.spec`): reset → commands → tick → (JS scan) → emit →
710
+ simulate across frames, with the GPU deciding to spawn, a CPU burst, the budget, and death.
711
+ - Orchestration on `SoftwareGPUDevice`: `GPUParticleSystem.spec` (scene sweep, growth, uploads only
712
+ when needed, the ring), `graph_particles.spec` (pass sequence, indirect sizing, the two orderings,
713
+ every persistent buffer imported exactly once), and the playground's `particle_scene.spec` (the
714
+ integration: geometry → hierarchy → particles → draw).
715
+ - Cross-parity meta-tests: every ISA opcode implemented by both executors.
716
+
717
+ Regression checks use the existing tools: seeded software-device buffer copies for list growth,
718
+ recorded pass ordering for program edits, emulator math for blend/fade and unknown bounds, and real
719
+ AssetManager/TextureAtlas bookkeeping with a stub image loader for asynchronous atlas lifetimes.
720
+ No full software WebGPU implementation or cross-frame shader simulator is added. The emulator's
721
+ `subgroupBroadcastFirst` is identity, so grouping correctness is checked at its input contract.
722
+
723
+ End-to-end GPU is out of scope for node (no binary GPU dep); emulator + JS references give the
724
+ coverage the brief asks for.
725
+
726
+ ### Emulator caveats discovered (all correct on real hardware)
727
+
728
+ The in-repo WGSL emulator (unit-test tool, not GPU parity) has limitations the shaders are written
729
+ around, documented in the memory note: `switch` case bodies are dropped (use if/else); repeated
730
+ read-modify-write of one location across a loop is corrupted (RNG is verified in isolation, not
731
+ through the interpreter loop); u32 multiply is imprecise for large operands (PCG can't be checked
732
+ through it); `/` is FLOAT division, so integer division only truncates where the result is stored
733
+ straight into a typed array — `(n + 63u) / 64u * 64u` reads as 66 there and 64 on a device, and a
734
+ binary search's `mid` is a shift, not a divide; matrix column indexing (`m[0]`) is not evaluated,
735
+ so an axis scale is read as `m * vec4(1,0,0,0)`. `array<atomic<u32>>` bindings are modeled as
736
+ `{value}` objects in tests; a pass that binds the same buffer as plain `array<u32>` gets a
737
+ `Uint32Array`. The scene context writes a row's `global` one publish behind (the hierarchy pass
738
+ recomposes it on a device), so a node's matrix under the emulator is the one it had when its row
739
+ was seeded.
740
+
741
+ ### Renderer integration
742
+
743
+ Done: `Renderer.feature_particles_enabled` (off by default), `Renderer.particles(scene_ctx)` creating
744
+ one system per scene context on first use, `simulate` recorded before the transparency pass and the
745
+ system passed to the AVBOIT orchestrator as its side channel. Remaining:
746
+ 1. For lit emitters, compile a simulate/render variant that swaps the disabled curve/lighting hooks
747
+ for `chunk_particle_curve_animation` (+ bind `GPUAnimationManager.database.buffer`) and a
748
+ shading-chunk-backed lighting hook (+ the `SHADING_LIGHT_RESOURCE_GROUP`).