@woosh/meep-engine 3.17.1 → 3.18.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (265) hide show
  1. package/editor/actions/concrete/ComponentAddAction.d.ts +1 -0
  2. package/editor/actions/concrete/ComponentAddAction.d.ts.map +1 -1
  3. package/editor/actions/concrete/ComponentAddAction.js +4 -0
  4. package/editor/actions/concrete/ComponentRemoveAction.d.ts +1 -0
  5. package/editor/actions/concrete/ComponentRemoveAction.d.ts.map +1 -1
  6. package/editor/actions/concrete/ComponentRemoveAction.js +4 -0
  7. package/editor/actions/concrete/EntityCreateAction.d.ts +1 -0
  8. package/editor/actions/concrete/EntityCreateAction.d.ts.map +1 -1
  9. package/editor/actions/concrete/EntityCreateAction.js +7 -0
  10. package/editor/actions/concrete/EntityRemoveAction.d.ts +1 -0
  11. package/editor/actions/concrete/EntityRemoveAction.d.ts.map +1 -1
  12. package/editor/actions/concrete/EntityRemoveAction.js +4 -0
  13. package/editor/actions/concrete/PatchTerrainHeightAction.d.ts +10 -0
  14. package/editor/actions/concrete/PatchTerrainHeightAction.d.ts.map +1 -1
  15. package/editor/actions/concrete/PatchTerrainHeightAction.js +107 -85
  16. package/editor/actions/concrete/TransformModifyAction.d.ts +1 -0
  17. package/editor/actions/concrete/TransformModifyAction.d.ts.map +1 -1
  18. package/editor/actions/concrete/TransformModifyAction.js +4 -0
  19. package/editor/particles/effect/ParticleEffectDocument.d.ts.map +1 -1
  20. package/editor/particles/effect/ParticleEffectDocument.js +10 -6
  21. package/editor/particles/effect/ParticleReferenceSimulation.d.ts.map +1 -1
  22. package/editor/particles/effect/ParticleReferenceSimulation.js +98 -4
  23. package/editor/particles/effect/make_starter_particle_effect.d.ts.map +1 -1
  24. package/editor/particles/effect/make_starter_particle_effect.js +0 -2
  25. package/editor/tools/engine/ToolEngine.d.ts +30 -0
  26. package/editor/tools/engine/ToolEngine.d.ts.map +1 -1
  27. package/editor/tools/engine/ToolEngine.js +257 -188
  28. package/editor/view/ecs/EntityList.d.ts.map +1 -1
  29. package/editor/view/ecs/EntityList.js +6 -5
  30. package/editor/view/particles/effect/ParticleEmitterInspectorView.d.ts.map +1 -1
  31. package/editor/view/particles/effect/ParticleEmitterInspectorView.js +17 -9
  32. package/package.json +1 -1
  33. package/src/CODEC_LAYOUT_VERIFICATION.md +1 -2
  34. package/src/core/binary/meshopt/MESHOPT_VERTEX_BLOCK.d.ts +56 -0
  35. package/src/core/binary/meshopt/MESHOPT_VERTEX_BLOCK.d.ts.map +1 -0
  36. package/src/core/binary/meshopt/MESHOPT_VERTEX_BLOCK.js +73 -0
  37. package/src/core/binary/meshopt/meshopt_decode_vertex_buffer.d.ts.map +1 -1
  38. package/src/core/binary/meshopt/meshopt_decode_vertex_buffer.js +18 -63
  39. package/src/core/binary/meshopt/meshopt_encode_byte_group.d.ts +53 -0
  40. package/src/core/binary/meshopt/meshopt_encode_byte_group.d.ts.map +1 -0
  41. package/src/core/binary/meshopt/meshopt_encode_byte_group.js +140 -0
  42. package/src/core/binary/meshopt/meshopt_encode_data_block.d.ts +32 -0
  43. package/src/core/binary/meshopt/meshopt_encode_data_block.d.ts.map +1 -0
  44. package/src/core/binary/meshopt/meshopt_encode_data_block.js +77 -0
  45. package/src/core/binary/meshopt/meshopt_encode_index_sequence.d.ts +46 -0
  46. package/src/core/binary/meshopt/meshopt_encode_index_sequence.d.ts.map +1 -0
  47. package/src/core/binary/meshopt/meshopt_encode_index_sequence.js +146 -0
  48. package/src/core/binary/meshopt/meshopt_encode_vertex_buffer.d.ts +40 -0
  49. package/src/core/binary/meshopt/meshopt_encode_vertex_buffer.d.ts.map +1 -0
  50. package/src/core/binary/meshopt/meshopt_encode_vertex_buffer.js +700 -0
  51. package/src/core/binary/meshopt/meshopt_write_uint_var.d.ts +24 -0
  52. package/src/core/binary/meshopt/meshopt_write_uint_var.d.ts.map +1 -0
  53. package/src/core/binary/meshopt/meshopt_write_uint_var.js +39 -0
  54. package/src/core/math/random/seededRandom_Mulberry32.d.ts.map +1 -1
  55. package/src/core/math/random/seededRandom_Mulberry32.js +4 -2
  56. package/src/core/process/undo/Action.d.ts +19 -0
  57. package/src/core/process/undo/Action.d.ts.map +1 -1
  58. package/src/core/process/undo/Action.js +17 -0
  59. package/src/engine/graphics3/GPUParticleEmitterSystem.d.ts +78 -0
  60. package/src/engine/graphics3/GPUParticleEmitterSystem.d.ts.map +1 -0
  61. package/src/engine/graphics3/GPUParticleEmitterSystem.js +310 -0
  62. package/src/engine/graphics3/GraphicsEngine.d.ts +17 -0
  63. package/src/engine/graphics3/GraphicsEngine.d.ts.map +1 -1
  64. package/src/engine/graphics3/GraphicsEngine.js +30 -10
  65. package/src/engine/graphics3/particles/particle_gpu_records.d.ts.map +1 -1
  66. package/src/engine/graphics3/particles/particle_gpu_records.js +0 -2
  67. package/src/engine/interpolation/InterpolationSystem.d.ts +7 -0
  68. package/src/engine/interpolation/InterpolationSystem.d.ts.map +1 -1
  69. package/src/engine/interpolation/InterpolationSystem.js +10 -0
  70. package/src/engine/network/NetworkSession.d.ts.map +1 -1
  71. package/src/engine/network/NetworkSession.js +14 -34
  72. package/src/engine/network/time/RenderPlayout.d.ts +76 -0
  73. package/src/engine/network/time/RenderPlayout.d.ts.map +1 -0
  74. package/src/engine/network/time/RenderPlayout.js +163 -0
  75. package/src/engine/network/transport/adapters/WebSocketTransport.d.ts.map +1 -1
  76. package/src/engine/network/transport/adapters/WebSocketTransport.js +11 -4
  77. package/src/engine/physics/ecs/PhysicsSystem.d.ts +4 -12
  78. package/src/engine/physics/ecs/PhysicsSystem.d.ts.map +1 -1
  79. package/src/engine/physics/ecs/PhysicsSystem.js +39 -13
  80. package/src/engine/physics/fluid/ecs/FluidObstacleSystem.d.ts +4 -4
  81. package/src/engine/sound/simulation/core/VolumeField.d.ts.map +1 -1
  82. package/src/engine/sound/simulation/core/VolumeField.js +4 -1
  83. package/src/format/scene/gltf/MESHOPT_COMPRESSION_PLAN.md +23 -2
  84. package/src/shade/playground/particle_system/README.md +139 -139
  85. package/src/shade/playground/particle_system/particle_prototype.d.ts.map +1 -1
  86. package/src/shade/playground/particle_system/particle_prototype.js +0 -2
  87. package/src/shade/playground/particle_system/particle_scene.d.ts +6 -9
  88. package/src/shade/playground/particle_system/particle_scene.d.ts.map +1 -1
  89. package/src/shade/playground/particle_system/particle_scene.js +6 -13
  90. package/src/shade/playground/ssr_reprojection/README.md +103 -0
  91. package/src/shade/playground/ssr_reprojection/SpatialFilterProbe.d.ts +17 -0
  92. package/src/shade/playground/ssr_reprojection/SpatialFilterProbe.d.ts.map +1 -0
  93. package/src/shade/playground/ssr_reprojection/SpatialFilterProbe.js +68 -0
  94. package/src/shade/playground/ssr_reprojection/index.html +79 -0
  95. package/src/shade/playground/ssr_reprojection/main.d.ts +2 -0
  96. package/src/shade/playground/ssr_reprojection/main.d.ts.map +1 -0
  97. package/src/shade/playground/ssr_reprojection/main.js +245 -0
  98. package/src/shade/playground/ssr_reprojection/scene.d.ts +18 -0
  99. package/src/shade/playground/ssr_reprojection/scene.d.ts.map +1 -0
  100. package/src/shade/playground/ssr_reprojection/scene.js +51 -0
  101. package/src/shade/renderer/Renderer.d.ts +2 -23
  102. package/src/shade/renderer/Renderer.d.ts.map +1 -1
  103. package/src/shade/renderer/Renderer.js +185 -218
  104. package/src/shade/renderer/animation/GPUAnimationManager.d.ts +5 -8
  105. package/src/shade/renderer/animation/GPUAnimationManager.d.ts.map +1 -1
  106. package/src/shade/renderer/animation/GPUAnimationManager.js +12 -12
  107. package/src/shade/renderer/buffer/table/GPUTypedTable.d.ts.map +1 -1
  108. package/src/shade/renderer/buffer/table/GPUTypedTable.js +6 -0
  109. package/src/shade/renderer/buffer/table/single/GPUSingleTypeTable.d.ts.map +1 -1
  110. package/src/shade/renderer/buffer/table/single/GPUSingleTypeTable.js +7 -0
  111. package/src/shade/renderer/global_illumination/brick4/gpu/draw/chunk_brick4_sample_specular.d.ts +4 -0
  112. package/src/shade/renderer/global_illumination/brick4/gpu/draw/chunk_brick4_sample_specular.d.ts.map +1 -0
  113. package/src/shade/renderer/global_illumination/brick4/gpu/draw/chunk_brick4_sample_specular.js +51 -0
  114. package/src/shade/renderer/global_illumination/brick4/gpu/draw/shader_brick4_specular.d.ts.map +1 -1
  115. package/src/shade/renderer/global_illumination/brick4/gpu/draw/shader_brick4_specular.js +5 -35
  116. package/src/shade/renderer/particles/DESIGN.md +703 -514
  117. package/src/shade/renderer/particles/GPUParticleSystem.d.ts +16 -4
  118. package/src/shade/renderer/particles/GPUParticleSystem.d.ts.map +1 -1
  119. package/src/shade/renderer/particles/GPUParticleSystem.js +1189 -991
  120. package/src/shade/renderer/particles/ParticleConstants.d.ts +64 -0
  121. package/src/shade/renderer/particles/ParticleConstants.d.ts.map +1 -1
  122. package/src/shade/renderer/particles/ParticleConstants.js +71 -0
  123. package/src/shade/renderer/particles/bounds/chunk_particle_bounds.d.ts +57 -0
  124. package/src/shade/renderer/particles/bounds/chunk_particle_bounds.d.ts.map +1 -0
  125. package/src/shade/renderer/particles/bounds/chunk_particle_bounds.js +181 -0
  126. package/src/shade/renderer/particles/bounds/chunk_particle_bounds_unknown.d.ts +7 -0
  127. package/src/shade/renderer/particles/bounds/chunk_particle_bounds_unknown.d.ts.map +1 -0
  128. package/src/shade/renderer/particles/bounds/chunk_particle_bounds_unknown.js +12 -0
  129. package/src/shade/renderer/particles/bounds/graph_particle_bounds.d.ts +33 -0
  130. package/src/shade/renderer/particles/bounds/graph_particle_bounds.d.ts.map +1 -0
  131. package/src/shade/renderer/particles/bounds/graph_particle_bounds.js +67 -0
  132. package/src/shade/renderer/particles/bounds/shader_particle_bounds.d.ts +40 -0
  133. package/src/shade/renderer/particles/bounds/shader_particle_bounds.d.ts.map +1 -0
  134. package/src/shade/renderer/particles/bounds/shader_particle_bounds.js +187 -0
  135. package/src/shade/renderer/particles/coherence/graph_particle_bucket.d.ts +15 -2
  136. package/src/shade/renderer/particles/coherence/graph_particle_bucket.d.ts.map +1 -1
  137. package/src/shade/renderer/particles/coherence/graph_particle_bucket.js +14 -8
  138. package/src/shade/renderer/particles/data/PARTICLE_COUNTERS.d.ts +58 -4
  139. package/src/shade/renderer/particles/data/PARTICLE_COUNTERS.d.ts.map +1 -1
  140. package/src/shade/renderer/particles/data/PARTICLE_COUNTERS.js +79 -4
  141. package/src/shade/renderer/particles/data/PARTICLE_EMITTER_STATE.d.ts +23 -1
  142. package/src/shade/renderer/particles/data/PARTICLE_EMITTER_STATE.d.ts.map +1 -1
  143. package/src/shade/renderer/particles/data/PARTICLE_EMITTER_STATE.js +35 -2
  144. package/src/shade/renderer/particles/data/PARTICLE_EMITTER_STRUCT.d.ts +4 -2
  145. package/src/shade/renderer/particles/data/PARTICLE_EMITTER_STRUCT.d.ts.map +1 -1
  146. package/src/shade/renderer/particles/data/PARTICLE_EMITTER_STRUCT.js +5 -8
  147. package/src/shade/renderer/particles/data/particle_emitter_record.d.ts.map +1 -1
  148. package/src/shade/renderer/particles/data/particle_emitter_record.js +0 -6
  149. package/src/shade/renderer/particles/graph_particles.d.ts +25 -3
  150. package/src/shade/renderer/particles/graph_particles.d.ts.map +1 -1
  151. package/src/shade/renderer/particles/graph_particles.js +108 -31
  152. package/src/shade/renderer/particles/graph_particles_avboit.d.ts +6 -4
  153. package/src/shade/renderer/particles/graph_particles_avboit.d.ts.map +1 -1
  154. package/src/shade/renderer/particles/graph_particles_avboit.js +5 -5
  155. package/src/shade/renderer/particles/runtime/EmitterRegistry.d.ts +5 -0
  156. package/src/shade/renderer/particles/runtime/EmitterRegistry.d.ts.map +1 -1
  157. package/src/shade/renderer/particles/runtime/EmitterRegistry.js +27 -6
  158. package/src/shade/renderer/particles/runtime/GPUParticleEmitterContext.d.ts +4 -0
  159. package/src/shade/renderer/particles/runtime/GPUParticleEmitterContext.d.ts.map +1 -1
  160. package/src/shade/renderer/particles/runtime/GPUParticleEmitterContext.js +6 -1
  161. package/src/shade/renderer/particles/runtime/ParticleEmitter.d.ts +35 -27
  162. package/src/shade/renderer/particles/runtime/ParticleEmitter.d.ts.map +1 -1
  163. package/src/shade/renderer/particles/runtime/ParticleEmitter.js +424 -425
  164. package/src/shade/renderer/particles/shaders/chunk_particle_emitter_warmup.d.ts +20 -0
  165. package/src/shade/renderer/particles/shaders/chunk_particle_emitter_warmup.d.ts.map +1 -0
  166. package/src/shade/renderer/particles/shaders/chunk_particle_render_math.d.ts +4 -0
  167. package/src/shade/renderer/particles/shaders/chunk_particle_render_math.d.ts.map +1 -1
  168. package/src/shade/renderer/particles/shaders/chunk_particle_render_math.js +38 -2
  169. package/src/shade/renderer/particles/shaders/chunk_particle_spawn_budget.d.ts +22 -0
  170. package/src/shade/renderer/particles/shaders/chunk_particle_spawn_budget.d.ts.map +1 -0
  171. package/src/shade/renderer/particles/shaders/chunk_particle_spawn_budget.js +42 -0
  172. package/src/shade/renderer/particles/shaders/shader_particle_avboit_draw.js +1 -1
  173. package/src/shade/renderer/particles/shaders/shader_particle_avboit_occupancy.d.ts.map +1 -1
  174. package/src/shade/renderer/particles/shaders/shader_particle_avboit_occupancy.js +53 -22
  175. package/src/shade/renderer/particles/shaders/shader_particle_emitter_tick.d.ts.map +1 -1
  176. package/src/shade/renderer/particles/shaders/shader_particle_emitter_tick.js +213 -216
  177. package/src/shade/renderer/particles/shaders/shader_particle_reclaim.d.ts +50 -0
  178. package/src/shade/renderer/particles/shaders/shader_particle_reclaim.d.ts.map +1 -0
  179. package/src/shade/renderer/particles/shaders/shader_particle_render.js +1 -1
  180. package/src/shade/renderer/particles/warmup/chunk_particle_warmup_step.d.ts +36 -0
  181. package/src/shade/renderer/particles/warmup/chunk_particle_warmup_step.d.ts.map +1 -0
  182. package/src/shade/renderer/particles/warmup/chunk_particle_warmup_step.js +102 -0
  183. package/src/shade/renderer/particles/warmup/graph_particle_warmup.d.ts +102 -0
  184. package/src/shade/renderer/particles/warmup/graph_particle_warmup.d.ts.map +1 -0
  185. package/src/shade/renderer/particles/warmup/graph_particle_warmup.js +329 -0
  186. package/src/shade/renderer/particles/warmup/shader_particle_warmup.d.ts +112 -0
  187. package/src/shade/renderer/particles/warmup/shader_particle_warmup.d.ts.map +1 -0
  188. package/src/shade/renderer/particles/warmup/shader_particle_warmup.js +318 -0
  189. package/src/shade/renderer/particles/warmup/shader_particle_warmup_advance.d.ts +35 -0
  190. package/src/shade/renderer/particles/warmup/shader_particle_warmup_advance.d.ts.map +1 -0
  191. package/src/shade/renderer/particles/warmup/shader_particle_warmup_advance.js +108 -0
  192. package/src/shade/renderer/particles/warmup/shader_particle_warmup_advance_wide.d.ts +3 -0
  193. package/src/shade/renderer/particles/warmup/shader_particle_warmup_advance_wide.d.ts.map +1 -0
  194. package/src/shade/renderer/particles/warmup/shader_particle_warmup_advance_wide.js +70 -0
  195. package/src/shade/renderer/particles/warmup/warmup_schedule.d.ts +62 -0
  196. package/src/shade/renderer/particles/warmup/warmup_schedule.d.ts.map +1 -0
  197. package/src/shade/renderer/particles/warmup/warmup_schedule.js +55 -0
  198. package/src/shade/renderer/postprocess/PostProcess.d.ts.map +1 -1
  199. package/src/shade/renderer/postprocess/PostProcess.js +2 -1
  200. package/src/shade/renderer/postprocess/ssr/REPROJECTION_PROPOSAL.md +286 -0
  201. package/src/shade/renderer/postprocess/ssr/SSR.d.ts +57 -21
  202. package/src/shade/renderer/postprocess/ssr/SSR.d.ts.map +1 -1
  203. package/src/shade/renderer/postprocess/ssr/SSR.js +230 -129
  204. package/src/shade/renderer/postprocess/ssr/SSRReprojectionMode.d.ts +11 -0
  205. package/src/shade/renderer/postprocess/ssr/SSRReprojectionMode.d.ts.map +1 -0
  206. package/src/shade/renderer/postprocess/ssr/SSRReprojectionMode.js +6 -0
  207. package/src/shade/renderer/postprocess/ssr/chunk_ssr_metadata.d.ts +4 -0
  208. package/src/shade/renderer/postprocess/ssr/chunk_ssr_metadata.d.ts.map +1 -0
  209. package/src/shade/renderer/postprocess/ssr/chunk_ssr_metadata.js +32 -0
  210. package/src/shade/renderer/postprocess/ssr/chunk_ssr_spatial_variance.d.ts +4 -0
  211. package/src/shade/renderer/postprocess/ssr/chunk_ssr_spatial_variance.d.ts.map +1 -0
  212. package/src/shade/renderer/postprocess/ssr/chunk_ssr_spatial_variance.js +41 -0
  213. package/src/shade/renderer/postprocess/ssr/reproject/chunk_ssr_reprojection.d.ts +3 -0
  214. package/src/shade/renderer/postprocess/ssr/reproject/chunk_ssr_reprojection.d.ts.map +1 -0
  215. package/src/shade/renderer/postprocess/ssr/reproject/chunk_ssr_reprojection.js +30 -0
  216. package/src/shade/renderer/postprocess/ssr/reproject/chunk_ssr_sample_history.d.ts +4 -0
  217. package/src/shade/renderer/postprocess/ssr/reproject/chunk_ssr_sample_history.d.ts.map +1 -0
  218. package/src/shade/renderer/postprocess/ssr/reproject/chunk_ssr_sample_history.js +70 -0
  219. package/src/shade/renderer/postprocess/ssr/resolve/build_ssr_resolve_shader.d.ts +21 -0
  220. package/src/shade/renderer/postprocess/ssr/resolve/build_ssr_resolve_shader.d.ts.map +1 -0
  221. package/src/shade/renderer/postprocess/ssr/resolve/build_ssr_resolve_shader.js +363 -0
  222. package/src/shade/renderer/postprocess/ssr/resolve/chunk_get_neighbour_weight.js +1 -1
  223. package/src/shade/renderer/postprocess/ssr/resolve/ssr_resolve_shader_Brick4.d.ts +2 -0
  224. package/src/shade/renderer/postprocess/ssr/resolve/ssr_resolve_shader_Brick4.d.ts.map +1 -0
  225. package/src/shade/renderer/postprocess/ssr/resolve/ssr_resolve_shader_Brick4.js +24 -0
  226. package/src/shade/renderer/postprocess/ssr/resolve/ssr_resolve_shader_IBL.d.ts +1 -6
  227. package/src/shade/renderer/postprocess/ssr/resolve/ssr_resolve_shader_IBL.d.ts.map +1 -1
  228. package/src/shade/renderer/postprocess/ssr/resolve/ssr_resolve_shader_IBL.js +16 -332
  229. package/src/shade/renderer/postprocess/ssr/resolve/ssr_resolve_shader_LPV.d.ts +1 -6
  230. package/src/shade/renderer/postprocess/ssr/resolve/ssr_resolve_shader_LPV.d.ts.map +1 -1
  231. package/src/shade/renderer/postprocess/ssr/resolve/ssr_resolve_shader_LPV.js +25 -358
  232. package/src/shade/renderer/postprocess/ssr/ssr_reproject_shader.d.ts +0 -5
  233. package/src/shade/renderer/postprocess/ssr/ssr_reproject_shader.d.ts.map +1 -1
  234. package/src/shade/renderer/postprocess/ssr/ssr_reproject_shader.js +201 -276
  235. package/src/shade/renderer/postprocess/ssr/ssr_spatial_denoise_shader.d.ts +1 -0
  236. package/src/shade/renderer/postprocess/ssr/ssr_spatial_denoise_shader.d.ts.map +1 -1
  237. package/src/shade/renderer/postprocess/ssr/ssr_spatial_denoise_shader.js +167 -143
  238. package/src/shade/renderer/shader/graph/graph_create_buffer.d.ts +15 -0
  239. package/src/shade/renderer/shader/graph/graph_create_buffer.d.ts.map +1 -0
  240. package/src/shade/renderer/shader/graph/graph_create_buffer.js +20 -0
  241. package/src/shade/renderer/view/GPUViewContext.d.ts +3 -4
  242. package/src/shade/renderer/view/GPUViewContext.d.ts.map +1 -1
  243. package/src/shade/renderer/view/GPUViewContext.js +5 -5
  244. package/src/core/binary/meshopt/__meshopt_test_streams.d.ts +0 -24
  245. package/src/core/binary/meshopt/__meshopt_test_streams.d.ts.map +0 -1
  246. package/src/core/binary/meshopt/__meshopt_test_streams.js +0 -237
  247. package/src/core/geom/3d/atlas/atlas_test_fixtures.d.ts +0 -84
  248. package/src/core/geom/3d/atlas/atlas_test_fixtures.d.ts.map +0 -1
  249. package/src/core/geom/3d/atlas/atlas_test_fixtures.js +0 -266
  250. package/src/engine/asset/loaders/gltf_test_fixtures.d.ts +0 -97
  251. package/src/engine/asset/loaders/gltf_test_fixtures.d.ts.map +0 -1
  252. package/src/engine/asset/loaders/gltf_test_fixtures.js +0 -359
  253. package/src/engine/sound/simulation/core/OcclusionSolver.d.ts +0 -29
  254. package/src/engine/sound/simulation/core/OcclusionSolver.d.ts.map +0 -1
  255. package/src/shade/renderer/animation/skin_test_fixtures.d.ts +0 -86
  256. package/src/shade/renderer/animation/skin_test_fixtures.d.ts.map +0 -1
  257. package/src/shade/renderer/animation/skin_test_fixtures.js +0 -250
  258. package/src/shade/renderer/geometry/virtual/format/attribute/vgeo_layers_test_fixtures.d.ts +0 -16
  259. package/src/shade/renderer/geometry/virtual/format/attribute/vgeo_layers_test_fixtures.d.ts.map +0 -1
  260. package/src/shade/renderer/geometry/virtual/format/attribute/vgeo_layers_test_fixtures.js +0 -27
  261. package/src/shade/renderer/particles/particle_test_fixtures.d.ts +0 -151
  262. package/src/shade/renderer/particles/particle_test_fixtures.d.ts.map +0 -1
  263. package/src/shade/renderer/particles/particle_test_fixtures.js +0 -229
  264. package/src/shade/renderer/particles/shaders/chunk_particle_emitter_world_sphere.js +0 -23
  265. package/src/shade/renderer/postprocess/ssr/reproject/shader_ffx_denoiser_reflections_reproject.js +0 -457
@@ -1,514 +1,703 @@
1
- # GPU-Driven Particle System — Design
2
-
3
- A specialized, deeply-integrated, GPU-driven particle engine for the v2 renderer.
4
- Destiny-style: **one** simulation dispatch runs **all** particles of **all** emitters by
5
- interpreting a compact per-emitter bytecode program (a shader VM). Effects are authored as
6
- node graphs and lowered to that bytecode by a testable JS compiler. Emitters are scene nodes, and
7
- the decision to spawn is the GPU's.
8
-
9
- ## Goals (from the brief)
10
-
11
- - Single simulation execution shader (data-driven VM, not one-pipeline-per-effect).
12
- - Projection modes: camera-facing billboard (+ velocity-stretched, axis-locked, world-oriented).
13
- - Depth sorting for correct transparency, built on the `csdldf` prefix scan.
14
- - Node-based composition → shader VM (compile graph → bytecode).
15
- - Frustum culling of spawning (optional, flagged, sleep-when-culled + bounded catch-up).
16
- - Flipbooks (atlas sub-rect + frame grid).
17
- - Texture lookup via atlas (region table; mirrors meep `AtlasPatch` uv offset/scale).
18
- - Soft-depth fade (flagged).
19
- - `GPUDatabase` for CPU-managed, low-churn tables (emitters).
20
- - Scene-node / bone attachment: an emitter **is** a `Node3D`; its world matrix is its row of the
21
- scene's `transforms` table, composed by the GPU hierarchy pass.
22
- - Distinct per-particle lifecycle: **no** built-in age/position — the schema + program define state.
23
- - Billboard renderer reuses engine shader chunks incl. lighting (forward-shaded, optional per emitter).
24
- - Blend modes (alpha + additive) via a **single** premultiplied blend setup.
25
- - `AnimationCurve` sampling during simulation (reuse `chunk_animation_curve_evaluate`).
26
- - Spawning decided on the GPU, per emitter, with the CPU able to ask for bursts through the same path.
27
- - No per-frame allocation on the CPU beyond the frame graph's own records.
28
-
29
- ## Layered architecture
30
-
31
- ```
32
- authoring: NodeGraph (nodes+ports) ──compile──▶ Program (bytecode: INIT + UPDATE)
33
- standard node library │
34
- ▼
35
- runtime CPU: ParticleEmitter (a Node3D) ── Scene ──▶ GPUSceneContext (transforms table)
36
- EmitterRegistry ── GPUDatabase(emitters + emitter_state) ─┤ ProgramHeap ([header][constants][code])
37
- GPUParticleSystem (owns every buffer, records the frame)
38
- ▼
39
- runtime GPU: particle pool (array<u32>, uniform stride) + alive/dead lists + counters
40
- emitter state table (GPU-owned: generation, accumulator, age, sleep, pending)
41
- ┌─ spawn commands (CPU bursts → pending)
42
- ├─ emitter tick (per row: cull, integrate rate, take from budget → count)
43
- ├─ scan (csdldf) → emit dispatch
44
- ├─ emit (VM INIT program) ── per newborn: find emitter in scan, pop dead → init → append alive_new
45
- ├─ simulate (VM UPDATE program) ── update → (kill|orphan→dead | keep→alive_new)
46
- ├─ bucket (histogram → csdldf scan → scatter) ── alive → program-coherent sim list
47
- ├─ build indirect (dispatch + draw args)
48
- ├─ sort (histogram → csdldf scan → scatter) [render order]
49
- └─ render (instanced billboard quads, forward + optional lighting + soft depth)
50
- ```
51
-
52
- ### Particle storage — single shared pool, uniform stride
53
-
54
- All particles live in one `array<u32>` pool with a fixed `RECORD_WORDS` stride (config, default 32
55
- words = 128 B). Word 0 is a reserved header holding the owning **emitter row** and that row's
56
- **generation** (`data/particle_header.js`); user attributes occupy words `1..RECORD_WORDS`. A single
57
- pool + single alive/dead lists means **one** emit dispatch and **one** simulate dispatch cover every
58
- emitter — the Destiny single-shader property.
59
-
60
- Lifecycle uses the canonical GPU dead-list / alive-list ping-pong:
61
- - `dead_list`: stack of free slot indices (+ atomic count).
62
- - `alive_list[2]`: ping-pong compact lists of live slot indices (+ counts).
63
- - `counters`: atomics {alive_count, dead_count, spawn budget, …} + derived indirect dispatch/draw args.
64
-
65
- Emit: pop dead → run INIT → append `alive_new`. Simulate: read the **coherence list** last frame's
66
- bucketing pass built from `alive_current` → run UPDATE → kill pushes to `dead`, survival appends to
67
- `alive_new`. Swap lists each frame.
68
-
69
- **The loop is strictly GPU-driven.** The build-indirect pass snapshots the padded length of the
70
- coherence list into `counters[ALIVE_IN]` and writes the indirect dispatch args next frame's simulate
71
- consumes (`dispatchWorkgroupsIndirect`) plus the indirect draw args this frame's render consumes
72
- (`drawIndirect`, sized by the *unpadded* alive count); simulate bound-checks against the ALIVE_IN
73
- snapshot. The sort's per-particle passes take the dispatch the bucket reset writes from the alive
74
- count. Sizing any of this from a CPU readback is a correctness bug, not a shortcut: the readback is
75
- inevitably stale, so a growing population orphans the tail of the alive list (slots leak into limbo)
76
- and a shrinking one re-simulates stale entries — double-freeing slots so two live particles later
77
- share one record. The emit pass's dead-list pop guards against wrapped counters
78
- (`old == 0 || old > capacity`), so racing an exhausted pool drops spawns cleanly instead of handing
79
- out garbage slots.
80
-
81
- ### Spawning is the GPU's decision
82
-
83
- The CPU never walks the emitters per frame and never counts particles. Per frame:
84
-
85
- 1. **spawn commands** (`shaders/shader_particle_spawn_commands.js`, only when the host queued any):
86
- one thread per `[row, count]` pair the host authored, `atomicAdd` into that row's `PENDING` in
87
- the GPU-owned emitter state. This is `GPUParticleSystem#spawn` — the CPU's facade over the GPU
88
- path, not a second path: a burst goes through the same budget, expansion and INIT program as the
89
- emitter's rate. It is also the channel a VM opcode would use to spawn from the GPU (a sub-emitter
90
- or trail would `atomicAdd` its target's `PENDING`).
91
- 2. **emitter tick** (`shaders/shader_particle_emitter_tick.js`): one thread per emitter table row,
92
- walked with the table's page iterator so an unallocated page costs one word and a freed row one
93
- bit. Per live row: recognise the row by generation (a mismatch starts the state over — new
94
- emitter, or a reused row, with nothing for the CPU to clear); take `PENDING`; if the emitter opts
95
- into culling, test its bounds sphere in world space (through its node's `global`) against the
96
- camera frustum and sleep while culled, paying back a bounded catch-up on wake; integrate
97
- `spawn_rate * dt`; turn the whole part of the accumulator into this frame's count, clamped by the
98
- **spawn budget** the reset pass seeded from the dead count — so an emitter is only ever debited
99
- for particles the pool can hold; write the count to `spawn_counts.elements[row]`.
100
- 3. **scan**: `graph_prefix_scan_csdldf` over the counts — each row's batch ends at its inclusive sum,
101
- and the last sum is the number of particles born.
102
- 4. **spawn args**: `shader_prefix_sum_to_command` (the rasteriser's own) turns that total into the
103
- emit dispatch.
104
- 5. **emit**: one thread per newborn, locating its emitter by `lower_bound_branchless` over the sums
105
- and its batch fraction from the sum before — the meshes-to-meshlets expansion applied to
106
- particles. No per-emitter cap, no thread loops over its particles.
107
-
108
- The host contributes two numbers to all of that: the emitter table's element capacity (the scan's
109
- length — static topology) and the length of the command list it wrote itself.
110
-
111
- ### The emitter state (`data/PARTICLE_EMITTER_STATE.js`)
112
-
113
- What the GPU integrates per emitter — generation, accumulator, age, sleep, pending — is the
114
- `emitter_state` table of the same database, one row per emitter at the emitter's row, **not**
115
- fields of the emitter record. The record is CPU-shadowed and restaged whole whenever the author
116
- changes a field, and that restage would overwrite an accumulator the GPU has been adding to for a
117
- hundred frames. The scene's transforms table mixes CPU fields with a GPU-written `global` only
118
- because `global` is recomputed from scratch every frame; nothing here is. In its own table the CPU
119
- touches a row exactly twice — a zero record when the emitter's row is claimed, the table's own
120
- zero-fill when it is released — and the database's resize copies pages on the GPU, so growth keeps
121
- what the tick has written. Nothing in it is atomic: each row is one tick thread's, and the
122
- spawn-command pass is handed one command per row (the system merges bursts on the CPU). The
123
- emitter's `age` here is the VM's `EMITTER_AGE` builtin.
124
-
125
- ### Eight storage buffers
126
-
127
- A compute stage gets eight storage buffers by default, and the design does not ask for more. Two
128
- things keep every pass under it: the emitter record and the GPU-owned state are two tables of one
129
- database buffer, read through one binding; and the VM's constants and code are one packed buffer —
130
- `[header][constants][code]`, the header's first word saying where the code starts
131
- (`ParticleConstants.PARTICLE_PROGRAM_HEADER_WORDS`). Emit and simulate bind exactly eight; their
132
- wide-register variants bind nine, which is the one place the system spends past the default and is
133
- still inside the ten the engine demands of an adapter. `particle_shaders_wgsl_validity.spec` pins
134
- the count for every pass.
135
-
136
- ### The VM (`ParticleVM`)
137
-
138
- Register machine, emulator-safe (bounded instruction loop + big `switch`, no subgroup ops, no
139
- recursion, no data-dependent jumps — conditionals are predicated via CMP/SELECT, destruction via
140
- `KILL`).
141
-
142
- - Register file: `array<f32>`, a register being one slot and a value of width *w* occupying *w*
143
- consecutive ones. The interpreter reaches it through `vm_reg_read` / `vm_reg_write`, which a
144
- **backing chunk** supplies, and the two backings are the two interpreters:
145
- `vm/chunk_particle_vm_registers_fast.js` puts the file in `var<private>`
146
- (`PARTICLE_VM_FAST_REGISTER_SLOTS`, dynamically indexed, so its size is the shader's register
147
- footprint), and `vm/chunk_particle_vm_registers_wide.js` puts it in a storage buffer
148
- (`PARTICLE_VM_REGISTER_SLOTS` — every slot an 8-bit operand base can name). See **Two
149
- interpreters** below.
150
- - Instruction: fixed **4×u32** words `[op | width << 8 | kinds << 16 | dst << 24, s0, s1, s2]`.
151
- Word 0 packs the opcode (bits 0..7), the lane count (8..10), each source's `VM_KIND` (two bits per
152
- slot at 16..21) and the destination slot (24..31). A source word means what its kind says: `REG`
153
- is a base slot plus a swizzle, `IMM` is an `f32` bit pattern read in every lane, and `NONE` is a
154
- raw field the opcode reads for itself (attribute word-offset+comp / curve handle / builtin id /
155
- pool index).
156
- - Constant pool: `array<vec4f>` per program, referenced by `LOAD_CONST`.
157
- - Two entry sections per program: **INIT** and **UPDATE** (offsets in the emitter record). INIT
158
- gives an already-allocated record its state; it has no say in whether the particle exists.
159
- - Builtins (`LOAD_BUILTIN`): DT, EMITTER_AGE, EMITTER_POSITION, EMITTER_DIRECTION, EMITTER_UP,
160
- PARTICLE_INDEX, SPAWN_FRACTION, TIME … assembled per-invocation by the host shader. The emitter's
161
- position and axes are its node's `global` out of the scene database.
162
- - `CURVE dst, handle, s0`: `dst.x = animation_curve_evaluate(curves[handle], s0.x)` — reuses the
163
- engine's GPU AnimationCurve chunk + database.
164
-
165
- #### Two interpreters
166
-
167
- The register file normally lives in shader registers, which is what makes it fast and what makes it
168
- small. An effect whose compiled program wants more slots than it holds is not an error and not an
169
- authoring ceiling: the whole system swaps emit and simulate for variants whose file is a storage
170
- buffer, and everything runs, slowly.
171
-
172
- - **The trigger is the set, not the program.** `ProgramHeap` reports `max_register_count` over the
173
- programs it holds, and `GPUParticleSystem` is where that number is compared against
174
- `PARTICLE_VM_FAST_REGISTER_SLOTS` — the store has no opinion about what a register budget means,
175
- and which backing a pass is compiled against is the VM's business. One hungry effect moves every
176
- emitter onto the wide pair; there is one pipeline per pass and it is chosen once, per frame, in
177
- `graph_particles`. It moves back when that effect is unregistered: the heap knows what it holds,
178
- so the number is a maximum over the live set rather than a high-water mark.
179
- - **The launch is bounded.** The wide passes need a register file per invocation in flight, so they
180
- do not run one lane per particle — that would size the buffer by pool capacity. They launch
181
- `PARTICLE_VM_WIDE_LAUNCH_LANES` lanes whatever the workload and walk their input with a grid-stride
182
- loop, so the file is `LANES * SLOTS * 4` bytes and nothing else. The stride is a whole number of
183
- workgroups, which is what keeps a wave inside one coherence bucket and the wave-uniform section
184
- decode sound.
185
- - **Slot-major, and never cleared.** `vm_registers[slot * LANES + lane]`, so a wave reading one slot
186
- reads a coalesced run. The buffer is transient and no pass zeroes it: a register is undefined until
187
- written, and a program that reads one first is corrupt. That is the one place the two backings are
188
- not interchangeable — the fast file zeroes, so the reference executor and the fast VM agree on a
189
- value the wide VM has no opinion about.
190
- - **Held to the same bar.** `vm/chunk_particle_vm.spec.js` runs every parity case, fuzz included,
191
- against BOTH backings; `shaders/particle_wide_register_passes.spec.js` runs the wide passes against
192
- the fast ones under the emulator and exercises the grid-stride loop.
193
-
194
- **Parity strategy.** The opcode numbers + operand semantics live in one source of truth
195
- (`isa/ParticleVMISA.js`). Two executors implement it: a JS reference VM (`vm/ParticleVMReference.js`)
196
- and a hand-written WGSL interpreter (`vm/chunk_particle_vm.js`). A spec runs identical programs
197
- through both (WGSL via the emulator) and asserts bit-for-bit-close equality. A meta-test asserts both
198
- executors cover every opcode in the ISA table.
199
-
200
- ### Node graph → bytecode
201
-
202
- A core `NodeGraph` holds particle nodes (ports + parameters) for INIT and UPDATE. `compile()`:
203
- 1. topologically orders reachable nodes,
204
- 2. lowers each node to VM ops (nodes emit via a small builder that allocates result registers),
205
- 3. linear-scan register allocation with liveness,
206
- 4. interns constants into the pool,
207
- 5. resolves curve/attribute references,
208
- 6. emits `Program { init, update, constants, reg_count, attributes, curves }`.
209
-
210
- Standard node library: constants, attribute get/set, arithmetic/vector/math, random (uniform/sphere/
211
- disk/cone), curve-sample, noise (curl), integrate, gravity, emit-shapes, compare/select, kill-if.
212
-
213
- ### Emitters are scene nodes
214
-
215
- `ParticleEmitter extends Node3D`. Add one to a `Scene` — on its own or under a joint — and the
216
- scene context's membership sweep gives it a `transforms` row like any node; the GPU hierarchy pass
217
- composes its `global` from its own transform and its parents' every frame, animation included.
218
- `GPUParticleSystem` sweeps the scene's node list when the scene's membership version moves, gives
219
- each emitter node a row in the emitter table pointed at its `transforms` row, and retires the row
220
- when the node leaves. Attaching an emitter to a bone is `joint.addChild(emitter)`; there is no
221
- per-frame transform copy, and an emitter under a GPU-animated joint follows it — which the old
222
- CPU-side copy of `transform_global` could not do, since the CPU never updates that matrix for a
223
- node the GPU owns.
224
-
225
- **The node knows nothing about the GPU.** A `ParticleEmitter` is a transform, an attribute
226
- `layout`, a compiled `program`, a `texture` URL, and the emission parameters — all of it the
227
- author's, all of it meaning the same thing with no device present, all of it what a saved effect is
228
- made of. Everything the system assigns is a `GPUParticleEmitterContext`, held by the
229
- `EmitterRegistry` against the emitter and reached through `registry.context(emitter)`: the table
230
- row, the generation, the program placement, the node's `transforms` row, the atlas patch. So an
231
- emitter cannot be saved with a row in it, cannot be given one by hand, and "is this emitter live?"
232
- is asked of the registry rather than read off a `-1` that anything could have written. The record
233
- writer is where the two meet — `write_emitter_record(record, emitter, context)`, next to the struct
234
- it fills.
235
-
236
- Restaging is driven by `Node3D#version`, the node's own change counter, which the renderer already
237
- reads to decide whether a transform needs re-uploading; the context remembers the version its row
238
- was staged from. An author who edits an emitter says so the way they say it for any node —
239
- `needsUpdate = true` — and there is no second flag to forget. A change to the context itself (a
240
- `transforms` row, a repacked atlas patch) moves no version, so the context marks itself instead.
241
-
242
- ### Emitter table (`GPUDatabase`)
243
-
244
- One `emitters` table (CPU-managed, low churn). Row = `PARTICLE_EMITTER_STRUCT`: program
245
- offsets+lengths, constant-pool offset, seed, spawn rate, **node** (transforms row), **generation**,
246
- bounds sphere, flags (lighting/soft/sort/cull/blend/projection), render binding (attr offsets for
247
- position/size/color/rotation/frame/velocity), atlas region + flipbook grid, and the program id
248
- the coherence pass buckets on. The atlas region is resolved, not authored: the emitter names an
249
- image and the system packs it, the way `ShaderManager` packs the old engine's sprites, so a repack
250
- moves every patch without any emitter changing. The seed is resolved too — an emitter left at 0 gets
251
- its registration's generation, which no other emitter of the registry shares, rather than being
252
- written to. Programs+constants live in a packed `ProgramHeap` buffer
253
- (variable length → not a fixed-stride table).
254
-
255
- The struct is the ONLY statement of that layout. The JS writer is the database's
256
- `write_wgsl_type_value`, and the WGSL reader — the `ParticleEmitter` struct declaration plus
257
- `database_read_..._element` — is generated from the same struct by the table descriptor. There is no
258
- hand-packed word offset table to drift out of step with the shaders, and no hand-written reader.
259
- The bucket histogram, which runs once per live particle and wants only `program_id`, reads it
260
- through `chunk_read_field` rather than decoding a whole row.
261
-
262
- **Editing the set at runtime is the point of using a table.** `EmitterRegistry.add` stages one
263
- record and `remove` frees one row; neither touches the others. Rows are stable for an emitter's
264
- lifetime because a live particle's record header holds the row of the emitter that spawned it —
265
- renumbering a survivor would re-point its particles at somebody else. A removed row is handed back
266
- to the next `add`, so the id space does not grow without bound.
267
-
268
- **Generations make that reuse safe.** Every add takes the next value of one running counter, never
269
- zero, and the particles of an emitter carry its generation in their header. The simulate pass kills
270
- any particle whose generation is not its row's current one — the row is zero-filled (generation 0)
271
- or a later tenant's — so a hard removal frees the emitter's particles within a frame, and a reused
272
- row cannot adopt a stranger's. Graceful retirement (rate to zero, remove once the tail has died)
273
- remains an author's choice rather than a correctness rule.
274
-
275
- **Programs are edited the same way.** `ProgramHeap` is a pair of range allocators
276
- (`OffsetAllocator`) over the two sections of the packed buffer, keyed by a `HashMap` on the bytecode
277
- itself — `ParticleProgram` answers `equals`/`hash`, so content deduplication is the map's doing
278
- rather than a stringified key's. `acquire` takes a reference and `release` gives one back, so the
279
- last emitter of an effect is what unpacks it, and `EmitterRegistry.set_program` swaps an emitter onto
280
- different bytecode while it is live. What makes that possible is that a program is **relocatable**:
281
- instructions carry no absolute addresses and constant references are program-local, so the heap can
282
- move one by copying its words and rewriting the placement.
283
-
284
- Two things are deliberately NOT reclaimed. **Program ids** are reused rather than renumbered — they
285
- are baked into every emitter record, so compacting them would mean rewriting the whole table — and
286
- `bucketCount()` is sized from the id capacity, the set's peak, so it does not move every time an
287
- effect comes or goes. An id below the peak with nothing in it is an empty bin, which rounds up to
288
- zero workgroups and shifts nothing after it. **Section capacity** only rises: a repack (which happens
289
- when a section has no contiguous run long enough) compacts what is live and grows only if it must,
290
- so a set that peaked at a hundred effects does not pay for a resize every time it returns to three.
291
-
292
- A repack is the one change no emitter announces, since it is not an edit to any of them.
293
- `EmitterRegistry.flush` asks the heap once per frame — `placement_version` — and restages every row
294
- when it has moved.
295
-
296
- Every row of a frame is staged through one module-static record (`EmitterRegistry`), so a flush
297
- allocates nothing however many rows it restages.
298
-
299
- ### Culling & attach
300
-
301
- An emitter that sets `EMITTER_FLAG.CULL` has its bounds sphere — centre through its node's matrix,
302
- radius scaled by the matrix's largest axis scale — tested against the camera frustum by the tick
303
- pass, with the renderer's shared `chunk_sphere_intersects_frustum`. Culled emitters spawn nothing
304
- and accrue sleep; on wake the tick adds a catch-up of at most `max_catchup_seconds` of rate. Live
305
- particles keep simulating while their emitter is culled. HiZ occlusion is not written.
306
-
307
- ### Two orderings, and why they are separate
308
-
309
- There are two independent permutations of the live particles each frame. They are produced by
310
- different passes into different buffers and must not be conflated or derived from one another.
311
-
312
- **Simulation order** (`coherence/`) — camera-independent, changes only when emitters or their
313
- programs change. The alive list comes out of compaction (`atomicAdd(&counters[ALIVE])`), so its
314
- order says nothing about which emitter a particle belongs to and a wave routinely holds particles
315
- running different VM programs. That costs twice: the interpreter's ~40-case opcode `switch`
316
- diverges, so a wave serially executes the union of the opcodes its lanes want; and `vm_reg[dst]` /
317
- `vm_reg[s0]` index a VGPR-resident register file dynamically, which AMD reaches with
318
- `v_movrels`/`v_movreld` behind a waterfall loop whose trip count is the number of *distinct*
319
- register indices live in the wave. Grouping the wave by program collapses both at once.
320
-
321
- The grouping is also **load-bearing**: the simulate shader reads its section parameters (code offset,
322
- instruction count, constant offset) from the wave's first lane through `subgroupBroadcastFirst`
323
- (`particle_wave_uniform`), which is what lets the driver fetch instruction words through the scalar
324
- unit, branch the opcode switch on a scalar and index the register file through M0 without a
325
- waterfall. A wave holding two programs would run one over the other's particles, so the coherence
326
- pass is part of simulate's correctness. Emit keeps a per-lane decode: its waves span emitters.
327
-
328
- The bucket key is the **program**, not the emitter: `ProgramHeap` deduplicates by content and hands
329
- back a `program_id` the emitter record carries, so emitters compiled from the same graph share
330
- one bucket. Each bucket is padded out to `PARTICLE_SIMULATE_WORKGROUP_SIZE` (64 — a multiple of both
331
- plausible subgroup widths, so no shader needs to query `subgroup_size`), which is what makes the
332
- grouping hold at wave granularity: an unpadded boundary lands mid-wave and that wave spans two
333
- programs again. Padding entries carry `PARTICLE_SIM_PAD_SLOT` and the simulate shader drops them
334
- *before* the compaction atomics, so they never enter the alive or dead list.
335
-
336
- It is an **indirection** list. Particle records never move: slots are referenced by the dead list and
337
- across frames, so relocating a record would leave both pointing at the wrong particle.
338
-
339
- The three-phase shape is the depth sort's, with the padding added: histogram (atomic per bucket) →
340
- round each bin up to a whole workgroup → `graph_prefix_scan_csdldf` (reused as-is — the one scan here
341
- whose only cross-workgroup wait is CSDLDF's bounded spin-then-fallback) → scatter → fill the tails
342
- with the sentinel. Nothing reads back to the CPU; the only host-supplied number is the bucket count,
343
- which is static topology.
344
-
345
- **Render order** (`sort/`) — per-frame, particles sorted back-to-front by view depth so
346
- premultiplied transparency composites correctly. Counting sort on a quantised view depth
347
- (`PARTICLE_SORT_BUCKET_COUNT` bins over `GPUParticleSystem#sort_depth_range`): histogram from the
348
- camera uniform's view matrix, `graph_prefix_scan_csdldf` the bin counts to offsets, scatter (index
349
- payload). The sorted index list feeds the instanced draw. Both per-particle passes are bounded by
350
- the alive counter and dispatched by the args the bucket reset wrote from it.
351
-
352
- Reordering the render list for program coherence would break transparency; depth-sorting the
353
- simulation list would destroy the coherence. They share no buffer and no counter slot —
354
- `PARTICLE_COUNTER.SORTED` belongs to the render order, `SIM_PADDED` to the simulation order.
355
-
356
- ### Rendering
357
-
358
- One render pipeline, one blend state = premultiplied: `color: (one, one-minus-src-alpha)`,
359
- `alpha: (one, one-minus-src-alpha)`. Additive vs. alpha is chosen by the shader emitting premultiplied
360
- RGB with alpha 0 (additive) vs. alpha a (normal) — GPU-driven, no pipeline switch. A 6-vertex quad is
361
- expanded per instance (billboard/velocity-stretch/axis-lock/world modes). Fragment: sample atlas
362
- (region + flipbook frame), optional forward lighting via `chunk_shade_standard_material_direct` +
363
- IBL, optional soft-depth fade. Depth-tested, no depth write. Draws into the HDR color target.
364
-
365
- ### Drawing under the renderer: an AVBOIT side channel
366
-
367
- Under `Renderer` the particles are not drawn by the billboard pass above. They go through the AVBOIT
368
- transparency pipeline (`rasterize/native/avboit/`) as a **side channel** — the contract in
369
- `AVBOITSideChannel.js` — so they composite order-independently with the transparent meshes and
370
- against the opaque scene, through the same volume and the same accumulators. Quads are not packed
371
- into meshlets; the orchestrator asks the channel for its contribution at the three points where the
372
- meshlet buckets have just made theirs:
373
-
374
- 1. **occupancy** (`shader_particle_avboit_occupancy`): one thread per emitter row, the emitter's
375
- bounds sphere (the culling sphere, through its node's `global`) as a view-depth interval, dilated
376
- and OR-ed into the occupancy bits. Per emitter, not per particle — no contention, and a particle
377
- that strays lands on a run boundary rather than anywhere wrong. Additive emitters mark nothing.
378
- This gives `bounds_radius` teeth on an emitter that never opted into culling: a particle outside
379
- the sphere still draws, but its extinction lands in whatever physical slice the warp put nearest,
380
- so it occludes at slightly the wrong depth. The sphere should hold the effect.
381
- 2. **splat** (`shader_particle_avboit_splat`): the billboard quads rasterized at the volume's
382
- resolution, coverage from the atlas alpha, extinction through `avboit_splat_slice` — the same
383
- write path as a surface. Additive particles write nothing. A soft-depth particle is faded here
384
- against the same scene depth the draw fades it against, one sample per voxel: what the volume
385
- records has to be the coverage the accumulators are weighted by, or a card that fades out where
386
- it meets the floor goes on darkening the floor through itself.
387
- 3. **draw** (`shader_particle_avboit_draw`): the quads at internal resolution, depth-tested read-only
388
- against the scene (which lets the soft fade read the same depth), fog composited, transmittance
389
- from the integral. Which of the two it is, is decided by coverage rather than by the blend enum,
390
- so an additive emitter and a premultiplied sprite with no alpha take the same path — as they do
391
- in the splat. A covering particle writes `color / norm / extinction` as a surface does, in-scatter
392
- included, weighted by the coverage it took. An emitting one writes the fourth accumulator,
393
- **emission**: light attenuated by the fog and by the volume in front of it, and carrying none of
394
- the fog's own in-scatter, because the resolve adds emission outside the normalization and the
395
- scene behind already has it. That accumulator is the one addition to AVBOIT's resolve for this.
396
- The whole decision is `chunk_particle_avboit_contribution`, and it is spec'd on the emulator.
397
-
398
- No sort: the volume makes order irrelevant, so this path reads the alive list and the draw args
399
- straight from `simulate`. The billboard vertex stage is one chunk
400
- (`chunk_particle_billboard_vertex`) shared by the standalone renderer and both AVBOIT passes.
401
- `GPUParticleSystem` implements the three methods; `Renderer.feature_particles_enabled` runs
402
- `simulate` before the transparency pass and hands the system to the orchestrator. MBOIT is legacy
403
- and gets no particle path.
404
-
405
- ### Ownership: `GPUParticleSystem`
406
-
407
- Constructed once against a `GPUSceneContext`, in the mould of `ReSTIRDI`. It owns every persistent
408
- buffer — pool, alive lists, dead list, counters, both indirect-args buffers, the coherence and sort
409
- scratch, the emitter database (both tables), the packed program buffer, the spawn-command ring — and grows the ones that depend on the emitter set inside `execute`. The
410
- default atlas is a 1x1 white texel. `execute({graph, camera, scene_database, color, depth, dt})`
411
- sweeps the scene, pushes what changed (on a command context opened only when something needs
412
- one), and records the whole frame through `graph_particles`, returning the colour handle after
413
- the draw. An integrator does not encode a particle pass by hand.
414
-
415
- ## File layout
416
-
417
- ```
418
- src/shade/renderer/particles/
419
- DESIGN.md
420
- ParticleConstants.js pool stride, reg count, workgroup sizes, limits
421
- isa/ParticleVMISA.js opcode table (single source of truth)
422
- isa/ParticleProgram.js Program value type + (de)serialization to u32
423
- isa/ParticleAssembler.js tiny assembler (build programs in tests/library)
424
- vm/ParticleVMReference.js JS reference executor
425
- vm/chunk_particle_vm.js WGSL interpreter CodeChunk (parity target)
426
- vm/chunk_particle_vm_registers_fast.js register file in var<private> — the interpreter that runs
427
- vm/chunk_particle_vm_registers_wide.js register file in storage — the fallback
428
- shaders/chunk_particle_simulate_step.js one particle's simulate, shared by both simulate shaders
429
- shaders/chunk_particle_emit_step.js one newborn's emit, shared by both emit shaders
430
- graph/ParticleNodeDescription.js node type base (ports + lowering)
431
- graph/ParticleNodeRegistry.js standard node library (a NodeRegistry)
432
- graph/particle_graph_authoring.js authoring sugar over core/model/node-graph
433
- graph/compile_particle_graph.js graph → Program (register alloc, const intern)
434
- layout/ParticleLayout.js attribute schema ↔ record word offsets
435
- data/PARTICLE_EMITTER_STRUCT.js the emitter row + flag/blend/projection encoding
436
- data/particle_emitter_record.js emitter + context -> one row of that struct
437
- data/PARTICLE_EMITTER_STATE.js the GPU-owned per-emitter state row (second table of the database)
438
- data/PARTICLE_DATABASE_SPEC.js the GPUDatabase schema, and the emitter table's descriptor
439
- data/PARTICLE_COUNTERS.js counter slot indices
440
- data/particle_header.js record header encoding (row + generation), JS + WGSL
441
- coherence/particle_program_buckets.js CPU reference for the simulation-order bucketing
442
- coherence/shader_particle_bucket.js reset / histogram / pad / scatter / fill
443
- coherence/graph_particle_bucket.js frame-graph wiring, around graph_prefix_scan_csdldf
444
- runtime/ProgramHeap.js the one packed [header][constants][code] buffer, allocator-backed
445
- runtime/ProgramPlacement.js where one program sits in it — what an emitter record carries
446
- runtime/ParticleEmitter.js one emitter — a Node3D and an effect, with no GPU in it
447
- runtime/GPUParticleEmitterContext.js what a registry assigns one: row, generation, placement, patch
448
- runtime/EmitterRegistry.js GPUDatabase(emitters) + the live emitters + their contexts + the program heap
449
- runtime/create_particle_effect.js layout + INIT/UPDATE graphs → compiled effect
450
- shaders/shader_particle_spawn_commands.js CPU bursts → pending spawns
451
- shaders/shader_particle_emitter_tick.js per-emitter spawn decision
452
- shaders/shader_particle_emit.js per-newborn INIT, sized by the scan
453
- shaders/shader_particle_simulate.js per-particle UPDATE + compaction + orphan kill
454
- shaders/shader_particle_finalize.js reset counters (+ budget) / build indirect args
455
- shaders/shader_particle_render.js billboard pipeline
456
- shaders/chunk_particle_setup_context.js VM builtins from the emitter's node, + the RNG seed
457
- shaders/chunk_particle_emitter_age.js the tick pass's running age, off the emitter_state table
458
- shaders/chunk_particle_curve_disabled.js the no-curves-bound CURVE provider
459
- shaders/chunk_particle_billboard_vertex.js the billboard vertex stage, shared by every draw
460
- shaders/chunk_particle_emitter_world_sphere.js the culling / occupancy sphere
461
- shaders/shader_particle_avboit_*.js occupancy / splat / draw through AVBOIT
462
- sort/* render-order depth sort
463
- graph_particles.js simulate (to the draw args) + the standalone sort and draw
464
- graph_particles_avboit.js the AVBOIT side channel: occupancy, splat, draw
465
- GPUParticleSystem.js the feature: owns the buffers, sweeps the scene, execute()
466
- particle_test_fixtures.js scene + registry + database words, for the emulator specs
467
- *.spec.js co-located tests (JS + emulator + software device)
468
- ```
469
-
470
- ## Testing
471
-
472
- - Pure JS: layout packing, assembler, register allocator, node compiler, heap/emitter/registry
473
- bookkeeping (via `SoftwareGPUDevice`), the header encoding.
474
- - WGSL via emulator: VM opcode parity (JS vs WGSL), curve sampling, sort key/depth, atlas+flipbook
475
- UV math, billboard expansion, soft-depth fade, and every compute pass on small fixtures — the tick
476
- (rate integration, generations, pending, budget, culling, page-group walk), emit (expansion over
477
- the scan, node position, orphans), simulate, spawn commands, finalize, bucket, sort. Fixtures
478
- (`particle_test_fixtures.js`) build a real scene with a `GPUSceneContext` and a real registry, and
479
- `gpu_database_words` lays both databases out the way the GPU holds them, so the shaders read rows
480
- the real writers wrote.
481
- - Multi-frame capstone (`particle_lifecycle.spec`): reset → commands → tick → (JS scan) → emit →
482
- simulate across frames, with the GPU deciding to spawn, a CPU burst, the budget, and death.
483
- - Orchestration on `SoftwareGPUDevice`: `GPUParticleSystem.spec` (scene sweep, growth, uploads only
484
- when needed, the ring), `graph_particles.spec` (pass sequence, indirect sizing, the two orderings,
485
- every persistent buffer imported exactly once), and the playground's `particle_scene.spec` (the
486
- integration: geometry → hierarchy → particles → draw).
487
- - Cross-parity meta-tests: every ISA opcode implemented by both executors.
488
-
489
- End-to-end GPU is out of scope for node (no binary GPU dep); emulator + JS references give the
490
- coverage the brief asks for.
491
-
492
- ### Emulator caveats discovered (all correct on real hardware)
493
-
494
- The in-repo WGSL emulator (unit-test tool, not GPU parity) has limitations the shaders are written
495
- around, documented in the memory note: `switch` case bodies are dropped (use if/else); repeated
496
- read-modify-write of one location across a loop is corrupted (RNG is verified in isolation, not
497
- through the interpreter loop); u32 multiply is imprecise for large operands (PCG can't be checked
498
- through it); `/` is FLOAT division, so integer division only truncates where the result is stored
499
- straight into a typed array — `(n + 63u) / 64u * 64u` reads as 66 there and 64 on a device, and a
500
- binary search's `mid` is a shift, not a divide; matrix column indexing (`m[0]`) is not evaluated,
501
- so an axis scale is read as `m * vec4(1,0,0,0)`. `array<atomic<u32>>` bindings are modeled as
502
- `{value}` objects in tests; a pass that binds the same buffer as plain `array<u32>` gets a
503
- `Uint32Array`. The scene context writes a row's `global` one publish behind (the hierarchy pass
504
- recomposes it on a device), so a node's matrix under the emulator is the one it had when its row
505
- was seeded.
506
-
507
- ### Renderer integration
508
-
509
- Done: `Renderer.feature_particles_enabled` (off by default), `Renderer.particles(scene_ctx)` creating
510
- one system per scene context on first use, `simulate` recorded before the transparency pass and the
511
- system passed to the AVBOIT orchestrator as its side channel. Remaining:
512
- 1. For lit emitters, compile a simulate/render variant that swaps the disabled curve/lighting hooks
513
- for `chunk_particle_curve_animation` (+ bind `GPUAnimationManager.database.buffer`) and a
514
- shading-chunk-backed lighting hook (+ the `SHADING_LIGHT_RESOURCE_GROUP`).
1
+ # GPU-Driven Particle System — Design
2
+
3
+ A specialized, deeply-integrated, GPU-driven particle engine for the v2 renderer.
4
+ Destiny-style: **one** simulation dispatch runs **all** particles of **all** emitters by
5
+ interpreting a compact per-emitter bytecode program (a shader VM). Effects are authored as
6
+ node graphs and lowered to that bytecode by a testable JS compiler. Emitters are scene nodes, and
7
+ the decision to spawn is the GPU's.
8
+
9
+ ## Goals (from the brief)
10
+
11
+ - Single simulation execution shader (data-driven VM, not one-pipeline-per-effect).
12
+ - Projection modes: camera-facing billboard (+ velocity-stretched, axis-locked, world-oriented).
13
+ - Depth sorting for correct transparency, built on the `csdldf` prefix scan.
14
+ - Node-based composition → shader VM (compile graph → bytecode).
15
+ - Frustum culling of spawning (optional, flagged, sleep-when-culled + bounded catch-up).
16
+ - Flipbooks (atlas sub-rect + frame grid).
17
+ - Texture lookup via atlas (region table; mirrors meep `AtlasPatch` uv offset/scale).
18
+ - Soft-depth fade (flagged).
19
+ - `GPUDatabase` for CPU-managed, low-churn tables (emitters).
20
+ - Scene-node / bone attachment: an emitter **is** a `Node3D`; its world matrix is its row of the
21
+ scene's `transforms` table, composed by the GPU hierarchy pass.
22
+ - Distinct per-particle lifecycle: **no** built-in age/position — the schema + program define state.
23
+ - Billboard renderer reuses engine shader chunks incl. lighting (forward-shaded, optional per emitter).
24
+ - Blend modes (alpha + additive) via a **single** premultiplied blend setup.
25
+ - `AnimationCurve` sampling during simulation (reuse `chunk_animation_curve_evaluate`).
26
+ - Spawning decided on the GPU, per emitter, with the CPU able to ask for bursts through the same path.
27
+ - No per-frame allocation on the CPU beyond the frame graph's own records.
28
+
29
+ ## Layered architecture
30
+
31
+ ```
32
+ authoring: NodeGraph (nodes+ports) ──compile──▶ Program (bytecode: INIT + UPDATE)
33
+ standard node library │
34
+ ▼
35
+ runtime CPU: ParticleEmitter (a Node3D) ── Scene ──▶ GPUSceneContext (transforms table)
36
+ EmitterRegistry ── GPUDatabase(emitters + emitter_state) ─┤ ProgramHeap ([header][constants][code])
37
+ GPUParticleSystem (owns every buffer, records the frame)
38
+ ▼
39
+ runtime GPU: particle pool (array<u32>, uniform stride) + alive/dead lists + counters
40
+ emitter state table (GPU-owned: generation, accumulator, age, sleep, pending)
41
+ ┌─ spawn commands (CPU bursts → pending)
42
+ ├─ emitter tick (per row: cull on the published box, integrate rate, take budget → count)
43
+ ├─ scan (csdldf) → emit dispatch
44
+ ├─ emit (VM INIT program) ── per newborn: find emitter in scan, pop dead → init → append alive_new
45
+ ├─ simulate (VM UPDATE program) ── update → (kill|orphan→dead | keep→alive_new)
46
+ ├─ bucket (histogram → csdldf scan → scatter) ── alive → program-coherent sim list
47
+ ├─ bounds (per particle → per emitter box, published to the state row)
48
+ ├─ build indirect (dispatch + draw args)
49
+ ├─ sort (histogram → csdldf scan → scatter) [render order]
50
+ └─ render (instanced billboard quads, forward + optional lighting + soft depth)
51
+ ```
52
+
53
+ ### Particle storage — single shared pool, uniform stride
54
+
55
+ All particles live in one `array<u32>` pool with a fixed `RECORD_WORDS` stride (config, default 32
56
+ words = 128 B). Word 0 is a reserved header holding the owning **emitter row** and that row's
57
+ **generation** (`data/particle_header.js`); user attributes occupy words `1..RECORD_WORDS`. A single
58
+ pool + single alive/dead lists means **one** emit dispatch and **one** simulate dispatch cover every
59
+ emitter — the Destiny single-shader property.
60
+
61
+ Lifecycle uses the canonical GPU dead-list / alive-list ping-pong:
62
+ - `dead_list`: stack of free slot indices (+ atomic count).
63
+ - `alive_list[2]`: ping-pong compact lists of live slot indices (+ counts).
64
+ - `counters`: atomics {alive_count, dead_count, spawn budget, …} + derived indirect dispatch/draw args.
65
+
66
+ Emit: pop dead → run INIT → append `alive_new`. Simulate: read the **coherence list** last frame's
67
+ bucketing pass built from `alive_current` → run UPDATE → kill pushes to `dead`, survival appends to
68
+ `alive_new`. Swap lists each frame.
69
+
70
+ **The loop is strictly GPU-driven.** The build-indirect pass snapshots the padded length of the
71
+ coherence list into `counters[ALIVE_IN]` and writes the indirect dispatch args next frame's simulate
72
+ consumes (`dispatchWorkgroupsIndirect`) plus the indirect draw args this frame's render consumes
73
+ (`drawIndirect`, sized by the *unpadded* alive count); simulate bound-checks against the ALIVE_IN
74
+ snapshot. The sort's per-particle passes take the dispatch the bucket reset writes from the alive
75
+ count. Sizing any of this from a CPU readback is a correctness bug, not a shortcut: the readback is
76
+ inevitably stale, so a growing population orphans the tail of the alive list (slots leak into limbo)
77
+ and a shrinking one re-simulates stale entries — double-freeing slots so two live particles later
78
+ share one record. The emit pass's dead-list pop guards against wrapped counters
79
+ (`old == 0 || old > capacity`), so racing an exhausted pool drops spawns cleanly instead of handing
80
+ out garbage slots.
81
+
82
+ ### Spawning is the GPU's decision
83
+
84
+ The CPU never walks the emitters per frame and never counts particles. Per frame:
85
+
86
+ 1. **spawn commands** (`shaders/shader_particle_spawn_commands.js`, only when the host queued any):
87
+ one thread per `[row, count]` pair the host authored, `atomicAdd` into that row's `PENDING` in
88
+ the GPU-owned emitter state. This is `GPUParticleSystem#spawn` — the CPU's facade over the GPU
89
+ path, not a second path: a burst goes through the same budget, expansion and INIT program as the
90
+ emitter's rate. It is also the channel a VM opcode would use to spawn from the GPU (a sub-emitter
91
+ or trail would `atomicAdd` its target's `PENDING`).
92
+ 2. **emitter tick** (`shaders/shader_particle_emitter_tick.js`): one thread per emitter table row,
93
+ walked with the table's page iterator so an unallocated page costs one word and a freed row one
94
+ bit. Per live row: recognise the row by generation (a mismatch starts the state over — new
95
+ emitter, or a reused row, with nothing for the CPU to clear); take `PENDING`; publish the
96
+ bounding box emit and simulate measured last frame into the state row and reset the accumulator
97
+ (see **Bounds**); if the emitter opts into culling, test that box against the camera frustum and
98
+ sleep while culled, paying back a bounded catch-up on wake; integrate
99
+ `spawn_rate * dt`; turn the whole part of the accumulator into this frame's count, clamped by the
100
+ **spawn budget** the reset pass seeded from the dead count — so an emitter is only ever debited
101
+ for particles the pool can hold; write the count to `spawn_counts.elements[row]`.
102
+ 3. **scan**: `graph_prefix_scan_csdldf` over the counts — each row's batch ends at its inclusive sum,
103
+ and the last sum is the number of particles born.
104
+ 4. **spawn args**: `shader_prefix_sum_to_command` (the rasteriser's own) turns that total into the
105
+ emit dispatch.
106
+ 5. **emit**: one thread per newborn, locating its emitter by `lower_bound_branchless` over the sums
107
+ and its batch fraction from the sum before — the meshes-to-meshlets expansion applied to
108
+ particles. No per-emitter cap, no thread loops over its particles.
109
+
110
+ On a frame that registered emitters with a `prewarm`, and only then, a loop of the warm-up pipeline
111
+ (`warmup/`) runs after the frame's own simulate: the same decision-scan-emit shape per tick over a
112
+ transient population, plus an advance of its own, then a migration into the main pool. See
113
+ **Pre-warm**.
114
+
115
+ The host contributes two numbers to all of that: the emitter table's element capacity (the scan's
116
+ length — static topology) and the length of the command list it wrote itself.
117
+
118
+ ### The emitter state (`data/PARTICLE_EMITTER_STATE.js`)
119
+
120
+ What the GPU integrates per emitter — generation, accumulator, age, sleep, pending — is the
121
+ `emitter_state` table of the same database, one row per emitter at the emitter's row, **not**
122
+ fields of the emitter record. The record is CPU-shadowed and restaged whole whenever the author
123
+ changes a field, and that restage would overwrite an accumulator the GPU has been adding to for a
124
+ hundred frames. The scene's transforms table mixes CPU fields with a GPU-written `global` only
125
+ because `global` is recomputed from scratch every frame; nothing here is. In its own table the CPU
126
+ touches a row exactly twice — a zero record when the emitter's row is claimed, the table's own
127
+ zero-fill when it is released — and the database's resize copies pages on the GPU, so growth keeps
128
+ what the tick has written. Nothing in it is atomic: each row is one tick thread's, and the
129
+ spawn-command pass is handed one command per row (the system merges bursts on the CPU). The
130
+ emitter's `age` here is the VM's `EMITTER_AGE` builtin — and the one field the warm-up pipeline
131
+ writes, as it runs an emitter's clock forward (see **Pre-warm**).
132
+
133
+ ### Eight storage buffers
134
+
135
+ A compute stage gets eight storage buffers by default, and the design does not ask for more. Two
136
+ things keep every pass under it: the emitter record and the GPU-owned state are two tables of one
137
+ database buffer, read through one binding; and the VM's constants and code are one packed buffer —
138
+ `[header][constants][code]`, the header's first word saying where the code starts
139
+ (`ParticleConstants.PARTICLE_PROGRAM_HEADER_WORDS`). Emit and simulate bind exactly eight; their
140
+ wide-register variants bind nine, which is the one place the system spends past the default and is
141
+ still inside the ten the engine demands of an adapter. `particle_shaders_wgsl_validity.spec` pins
142
+ the count for every pass.
143
+
144
+ ### The VM (`ParticleVM`)
145
+
146
+ Register machine, emulator-safe (bounded instruction loop + big `switch`, no subgroup ops, no
147
+ recursion, no data-dependent jumps — conditionals are predicated via CMP/SELECT, destruction via
148
+ `KILL`).
149
+
150
+ - Register file: `array<f32>`, a register being one slot and a value of width *w* occupying *w*
151
+ consecutive ones. The interpreter reaches it through `vm_reg_read` / `vm_reg_write`, which a
152
+ **backing chunk** supplies, and the two backings are the two interpreters:
153
+ `vm/chunk_particle_vm_registers_fast.js` puts the file in `var<private>`
154
+ (`PARTICLE_VM_FAST_REGISTER_SLOTS`, dynamically indexed, so its size is the shader's register
155
+ footprint), and `vm/chunk_particle_vm_registers_wide.js` puts it in a storage buffer
156
+ (`PARTICLE_VM_REGISTER_SLOTS` — every slot an 8-bit operand base can name). See **Two
157
+ interpreters** below.
158
+ - Instruction: fixed **4×u32** words `[op | width << 8 | kinds << 16 | dst << 24, s0, s1, s2]`.
159
+ Word 0 packs the opcode (bits 0..7), the lane count (8..10), each source's `VM_KIND` (two bits per
160
+ slot at 16..21) and the destination slot (24..31). A source word means what its kind says: `REG`
161
+ is a base slot plus a swizzle, `IMM` is an `f32` bit pattern read in every lane, and `NONE` is a
162
+ raw field the opcode reads for itself (attribute word-offset+comp / curve handle / builtin id /
163
+ pool index).
164
+ - Constant pool: `array<vec4f>` per program, referenced by `LOAD_CONST`.
165
+ - Two entry sections per program: **INIT** and **UPDATE** (offsets in the emitter record). INIT
166
+ gives an already-allocated record its state; it has no say in whether the particle exists.
167
+ - Builtins (`LOAD_BUILTIN`): DT, EMITTER_AGE, EMITTER_POSITION, EMITTER_DIRECTION, EMITTER_UP,
168
+ PARTICLE_INDEX, SPAWN_FRACTION, TIME … assembled per-invocation by the host shader. The emitter's
169
+ position and axes are its node's `global` out of the scene database.
170
+ - `CURVE dst, handle, s0`: `dst.x = animation_curve_evaluate(curves[handle], s0.x)` — reuses the
171
+ engine's GPU AnimationCurve chunk + database.
172
+
173
+ #### Two interpreters
174
+
175
+ The register file normally lives in shader registers, which is what makes it fast and what makes it
176
+ small. An effect whose compiled program wants more slots than it holds is not an error and not an
177
+ authoring ceiling: the whole system swaps emit and simulate for variants whose file is a storage
178
+ buffer, and everything runs, slowly.
179
+
180
+ - **The trigger is the set, not the program.** `ProgramHeap` reports `max_register_count` over the
181
+ programs it holds, and `GPUParticleSystem` is where that number is compared against
182
+ `PARTICLE_VM_FAST_REGISTER_SLOTS` — the store has no opinion about what a register budget means,
183
+ and which backing a pass is compiled against is the VM's business. One hungry effect moves every
184
+ emitter onto the wide pair; there is one pipeline per pass and it is chosen once, per frame, in
185
+ `graph_particles`. It moves back when that effect is unregistered: the heap knows what it holds,
186
+ so the number is a maximum over the live set rather than a high-water mark.
187
+ - **The launch is bounded.** The wide passes need a register file per invocation in flight, so they
188
+ do not run one lane per particle — that would size the buffer by pool capacity. They launch
189
+ `PARTICLE_VM_WIDE_LAUNCH_LANES` lanes whatever the workload and walk their input with a grid-stride
190
+ loop, so the file is `LANES * SLOTS * 4` bytes and nothing else. The stride is a whole number of
191
+ workgroups, which is what keeps a wave inside one coherence bucket and the wave-uniform section
192
+ decode sound.
193
+ - **Slot-major, and never cleared.** `vm_registers[slot * LANES + lane]`, so a wave reading one slot
194
+ reads a coalesced run. The buffer is transient and no pass zeroes it: a register is undefined until
195
+ written, and a program that reads one first is corrupt. That is the one place the two backings are
196
+ not interchangeable — the fast file zeroes, so the reference executor and the fast VM agree on a
197
+ value the wide VM has no opinion about.
198
+ - **Held to the same bar.** `vm/chunk_particle_vm.spec.js` runs every parity case, fuzz included,
199
+ against BOTH backings; `shaders/particle_wide_register_passes.spec.js` runs the wide passes against
200
+ the fast ones under the emulator and exercises the grid-stride loop.
201
+
202
+ **Parity strategy.** The opcode numbers + operand semantics live in one source of truth
203
+ (`isa/ParticleVMISA.js`). Two executors implement it: a JS reference VM (`vm/ParticleVMReference.js`)
204
+ and a hand-written WGSL interpreter (`vm/chunk_particle_vm.js`). A spec runs identical programs
205
+ through both (WGSL via the emulator) and asserts bit-for-bit-close equality. A meta-test asserts both
206
+ executors cover every opcode in the ISA table.
207
+
208
+ ### Node graph → bytecode
209
+
210
+ A core `NodeGraph` holds particle nodes (ports + parameters) for INIT and UPDATE. `compile()`:
211
+ 1. topologically orders reachable nodes,
212
+ 2. lowers each node to VM ops (nodes emit via a small builder that allocates result registers),
213
+ 3. linear-scan register allocation with liveness,
214
+ 4. interns constants into the pool,
215
+ 5. resolves curve/attribute references,
216
+ 6. emits `Program { init, update, constants, reg_count, attributes, curves }`.
217
+
218
+ Standard node library: constants, attribute get/set, arithmetic/vector/math, random (uniform/sphere/
219
+ disk/cone), curve-sample, noise (curl), integrate, gravity, emit-shapes, compare/select, kill-if.
220
+
221
+ ### Emitters are scene nodes
222
+
223
+ `ParticleEmitter extends Node3D`. Add one to a `Scene` — on its own or under a joint — and the
224
+ scene context's membership sweep gives it a `transforms` row like any node; the GPU hierarchy pass
225
+ composes its `global` from its own transform and its parents' every frame, animation included.
226
+ `GPUParticleSystem` sweeps the scene's node list when the scene's membership version moves, gives
227
+ each emitter node a row in the emitter table pointed at its `transforms` row, and retires the row
228
+ when the node leaves. Attaching an emitter to a bone is `joint.addChild(emitter)`; there is no
229
+ per-frame transform copy, and an emitter under a GPU-animated joint follows it — which the old
230
+ CPU-side copy of `transform_global` could not do, since the CPU never updates that matrix for a
231
+ node the GPU owns.
232
+
233
+ **The node knows nothing about the GPU.** A `ParticleEmitter` is a transform, an attribute
234
+ `layout`, a compiled `program`, a `texture` URL, and the emission parameters — all of it the
235
+ author's, all of it meaning the same thing with no device present, all of it what a saved effect is
236
+ made of. Everything the system assigns is a `GPUParticleEmitterContext`, held by the
237
+ `EmitterRegistry` against the emitter and reached through `registry.context(emitter)`: the table
238
+ row, the generation, the program placement, the node's `transforms` row, the atlas patch. So an
239
+ emitter cannot be saved with a row in it, cannot be given one by hand, and "is this emitter live?"
240
+ is asked of the registry rather than read off a `-1` that anything could have written. The record
241
+ writer is where the two meet — `write_emitter_record(record, emitter, context)`, next to the struct
242
+ it fills.
243
+
244
+ Restaging is driven by `Node3D#version`, the node's own change counter, which the renderer already
245
+ reads to decide whether a transform needs re-uploading; the context remembers the version its row
246
+ was staged from. An author who edits an emitter says so the way they say it for any node —
247
+ `needsUpdate = true` — and there is no second flag to forget. A change to the context itself (a
248
+ `transforms` row, a repacked atlas patch) moves no version, so the context marks itself instead.
249
+
250
+ ### Emitter table (`GPUDatabase`)
251
+
252
+ One `emitters` table (CPU-managed, low churn). Row = `PARTICLE_EMITTER_STRUCT`: program
253
+ offsets+lengths, constant-pool offset, seed, spawn rate, **node** (transforms row), **generation**,
254
+ flags (lighting/soft/sort/cull/blend/projection), render binding (attr offsets for
255
+ position/size/color/rotation/frame/velocity), atlas region + flipbook grid, and the program id
256
+ the coherence pass buckets on. The atlas region is resolved, not authored: the emitter names an
257
+ image and the system packs it, the way `ShaderManager` packs the old engine's sprites, so a repack
258
+ moves every patch without any emitter changing. The seed is resolved too — an emitter left at 0 gets
259
+ its registration's generation, which no other emitter of the registry shares, rather than being
260
+ written to. Programs+constants live in a packed `ProgramHeap` buffer
261
+ (variable length → not a fixed-stride table).
262
+
263
+ The struct is the ONLY statement of that layout. The JS writer is the database's
264
+ `write_wgsl_type_value`, and the WGSL reader — the `ParticleEmitter` struct declaration plus
265
+ `database_read_..._element` — is generated from the same struct by the table descriptor. There is no
266
+ hand-packed word offset table to drift out of step with the shaders, and no hand-written reader.
267
+ The bucket histogram, which runs once per live particle and wants only `program_id`, reads it
268
+ through `chunk_read_field` rather than decoding a whole row.
269
+
270
+ **Editing the set at runtime is the point of using a table.** `EmitterRegistry.add` stages one
271
+ record and `remove` frees one row; neither touches the others. Rows are stable for an emitter's
272
+ lifetime because a live particle's record header holds the row of the emitter that spawned it —
273
+ renumbering a survivor would re-point its particles at somebody else. A removed row is handed back
274
+ to the next `add`, so the id space does not grow without bound.
275
+
276
+ **Generations make that reuse safe.** Every add takes the next value of one running counter, never
277
+ zero, and the particles of an emitter carry its generation in their header. The simulate pass kills
278
+ any particle whose generation is not its row's current one — the row is zero-filled (generation 0)
279
+ or a later tenant's — so a hard removal frees the emitter's particles within a frame, and a reused
280
+ row cannot adopt a stranger's. Graceful retirement (rate to zero, remove once the tail has died)
281
+ remains an author's choice rather than a correctness rule.
282
+
283
+ **Programs are edited the same way.** `ProgramHeap` is a pair of range allocators
284
+ (`OffsetAllocator`) over the two sections of the packed buffer, keyed by a `HashMap` on the bytecode
285
+ itself — `ParticleProgram` answers `equals`/`hash`, so content deduplication is the map's doing
286
+ rather than a stringified key's. `acquire` takes a reference and `release` gives one back, so the
287
+ last emitter of an effect is what unpacks it, and `EmitterRegistry.set_program` swaps an emitter onto
288
+ different bytecode while it is live. What makes that possible is that a program is **relocatable**:
289
+ instructions carry no absolute addresses and constant references are program-local, so the heap can
290
+ move one by copying its words and rewriting the placement.
291
+
292
+ Two things are deliberately NOT reclaimed. **Program ids** are reused rather than renumbered — they
293
+ are baked into every emitter record, so compacting them would mean rewriting the whole table — and
294
+ `bucketCount()` is sized from the id capacity, the set's peak, so it does not move every time an
295
+ effect comes or goes. An id below the peak with nothing in it is an empty bin, which rounds up to
296
+ zero workgroups and shifts nothing after it. **Section capacity** only rises: a repack (which happens
297
+ when a section has no contiguous run long enough) compacts what is live and grows only if it must,
298
+ so a set that peaked at a hundred effects does not pay for a resize every time it returns to three.
299
+
300
+ A repack is the one change no emitter announces, since it is not an edit to any of them.
301
+ `EmitterRegistry.flush` asks the heap once per frame — `placement_version` — and restages every row
302
+ when it has moved.
303
+
304
+ Programs are immutable after compilation. Both `registry.set_program(emitter, program)` and replacing
305
+ `emitter.program` followed by `needsUpdate = true` are supported. Flush reconciles replacements
306
+ before inspecting heap relocation, since acquiring a replacement may move other programs too.
307
+ The layout must remain compatible with particles already in flight.
308
+
309
+ A live program assignment increments `program_assignment_version`. On the next frame, before reset
310
+ or simulation, the previous compact alive list is bucketed again and its padded count and dispatch
311
+ are rebuilt. This prevents a previously shared wave from executing one emitter's newly edited program
312
+ over another emitter's particles. Ordinary frames keep the existing end-of-frame bucketing only;
313
+ content-identical replacements need no repair. Both bucket runs on an edit frame reuse the latest
314
+ graph handles for their scratch buffers, so each persistent buffer is imported once.
315
+
316
+ The coherence list is persistent GPU state, including its padding. Growing the bucket capacity copies
317
+ it to the replacement buffer before destroying the old one. The old count and indirect dispatch
318
+ remain valid across growth; only the per-frame histogram/cursor/key scratch may be discarded.
319
+
320
+ Every row of a frame is staged through one module-static record (`EmitterRegistry`), so a flush
321
+ allocates nothing however many rows it restages.
322
+
323
+ ### Bounds
324
+
325
+ An emitter has no bounds it could declare. Its particles are simulated in world space by arbitrary
326
+ bytecode: they go where the program sends them, at whatever size it gives them, for as long as it
327
+ keeps them alive. The only bound an author could write down and still be right is an infinite one,
328
+ which culls nothing and reserves every depth slice. So the box is **measured**, on the GPU, every
329
+ frame.
330
+
331
+ Two passes of its own (`bounds/`), at the end of the frame, after emit and simulate have both
332
+ finished appending to the alive list:
333
+
334
+ - **accumulate** — one thread per live particle: read where its quad is and merge it into its
335
+ emitter's accumulator. The centre is the record's `render_position`; the dilation is
336
+ `particle_billboard_radius`, the largest offset `particle_billboard_offset` can produce for that
337
+ size, velocity and projection, over every corner, every roll and any camera. The box has to hold
338
+ the *quads*, not the centres, and it has to be camera-independent, because it is read on a later
339
+ frame. The merge is `aabb_atomic_min_f32` / `aabb_atomic_max_f32`, the same CAS-spin the
340
+ skinned-mesh bounds refresh uses; both load first and return without a CAS once the box holds the
341
+ contribution, so what survives is load traffic, not exchanges. Dispatched indirectly from
342
+ `sim_dispatch_args`, the args the coherence reset wrote from the GPU's own alive count.
343
+ - **publish** — one thread per emitter row: write the accumulation to `bounds_min` / `bounds_max` of
344
+ the state row (`data/PARTICLE_EMITTER_STATE.js` — measured state is GPU-owned state, and on the
345
+ record it would be clobbered by the next restage), and reset the accumulator. The reset is
346
+ unconditional, which is what lets a row that changes hands start clean.
347
+
348
+ **Why a pass and not a tail on emit and simulate.** It was that first. Booking from inside those two
349
+ measures no single population — simulate sees last frame's survivors and emit sees this frame's
350
+ newborns, so the union spans two frames and is looser than either — and it pays six atomics per
351
+ particle inside the pass whose register pressure decides the system's occupancy, to compute
352
+ something that needs none of the VM, none of the scene, and five words of the record. Over the alive
353
+ list instead, the measurement is of exactly the set that is alive and about to be drawn.
354
+
355
+ **Where the accumulator lives.** Six words per emitter row, in the **counters** buffer, after the
356
+ counters themselves (`data/PARTICLE_COUNTERS.js`). It has to be an `array<atomic<u32>>` binding, and
357
+ the emitter database cannot be it — the passes that read a record read it out of an `array<u32>`
358
+ binding, and one buffer under two access types is aliasing. The counters buffer is already bound in
359
+ both bounds passes, already atomic, already the system's cross-pass scratch, so a buffer of its own
360
+ would add a binding and a growth path to say the same thing. It therefore grows with the emitter
361
+ table, by copy — the free-list length in it is GPU-owned and unreconstructable — and the rows the
362
+ growth adds are written to the empty interval, because a zeroed accumulator reads back as a
363
+ degenerate box at the world origin rather than as an absence.
364
+
365
+ **The empty case.** An emitter with no measured particles has infinite, unknown bounds:
366
+ `bounds_min = [-Infinity, -Infinity, -Infinity]`, `bounds_max = [Infinity, Infinity, Infinity]`.
367
+ The CPU stages this on registration and the publish pass restores it whenever the population is
368
+ empty. The tick treats unknown bounds as visible, so a slow emitter can accumulate its first spawn
369
+ without prewarm even when its node is offscreen. A depleted population can restart the same way.
370
+ AVBOIT occupancy skips unknown bounds because there are no live particles to splat. Both readers
371
+ test the sentinel before geometric arithmetic, avoiding NaNs from operations on infinities.
372
+ The sentinel check compares integer bits. The publish pass swaps the empty accumulator's loaded
373
+ endpoints; neither path constructs an infinite floating-point constant, which WGSL rejects at
374
+ shader creation time.
375
+ Measured, nonempty populations continue to use their finite bounds for culling. No bounds are authored.
376
+ Prewarm controls the initial age distribution only; it is not required for culling correctness.
377
+
378
+ ### Pre-warm
379
+
380
+ A continuous effect that starts empty grows into its steady state over one particle lifetime, in
381
+ front of whoever is watching. `ParticleEmitter.prewarm` — seconds, `0` for off — buys that lifetime
382
+ back: an emitter registered with one is run forward by that long before it is first drawn.
383
+
384
+ **A warm-up is the ordinary loop, run over a population of its own.** Spawn, init, advance, per
385
+ tick, until every emitter in the batch has been running for its own `prewarm`. It is *not* a batch
386
+ of particles dumped in one tick and aged: the emitter keeps spawning throughout, what dies is
387
+ replaced as it is in a live scene, and whatever drives the rate — a spawn curve, `EMITTER_AGE`, the
388
+ program itself — drives it here, because it is the same program being run. What the loop leaves is
389
+ the population the emitter would have had.
390
+
391
+ **It is a pipeline apart from the main one** (`warmup/`), because the main one is tuned and must not
392
+ pay for this. Emit and simulate are the passes whose register pressure sets the system's occupancy,
393
+ and their waves are kept program-uniform by the bucketing; a warm-up folded into either would spend
394
+ registers on every frame for work that happens on one, and a loop inside a lane would break the
395
+ uniform flow across a wave. So the warm-up shares code with the main frame and integrates nothing
396
+ into it: the main **emit pass is reused unchanged**, bound to the warm-up's buffers; the tick is a
397
+ pass of the warm-up's own over a host-written queue (`shader_particle_warmup_tick`) rather than over
398
+ the emitter table; the **advance** is a pass of the warm-up's own (`chunk_particle_warmup_step.js`),
399
+ decoding the program section per lane — it reads the raw alive list, so it cannot lean on the
400
+ bucketing's wave-uniformity, and its occupancy is its own business. The main passes' compiled code
401
+ is byte-identical whether `warmup/` exists or not.
402
+
403
+ **The population is transient.** Pool, both alive lists, free list and counters are frame-graph
404
+ transients that exist for the loop and are then moved, record by record, into the main pool by the
405
+ **migrate** pass — a pop from the main free list per particle, and never a push, so it is the one
406
+ point of contact and it is safe by construction. The loop's own free-list traffic never touches the
407
+ main list, and the main passes never see a half-warmed particle.
408
+
409
+ **Bounded.** `warmup/warmup_schedule.js`: at most `PARTICLE_WARMUP_MAX_TICKS` ticks, no shorter than
410
+ `PARTICLE_WARMUP_TICK_SECONDS`, the tick *length* stretching to cover the longest pre-warm in the
411
+ batch rather than the count growing to cover it — the count is what the frame pays for, at six or
412
+ seven dispatches a tick. An emitter takes part in a tick once the loop's clock is within its own
413
+ pre-warm of now, so two emitters queued with different durations each end having run for exactly
414
+ their own; the shorter one sat out the start.
415
+
416
+ **Jittered, so it does not arrive in bands.** Every particle born in one tick would otherwise reach
417
+ the end with exactly that tick's age, and N ticks would read as N cohorts. The advance jitters the
418
+ tick per particle (`PARTICLE_WARMUP_DT_JITTER`, hashed from slot and tick — a hash, not a draw, so
419
+ the program's own random stream is untouched), and the host jitters each emitter's clock and starting
420
+ accumulator (`PARTICLE_WARMUP_EMITTER_JITTER`), so emitters in one batch do not spawn on the same
421
+ beat either. A program does not care what `DELTA_TIME` it is handed; the smear is free.
422
+
423
+ **Where it sits.** After the frame's own emit and simulate, before the bucketing. Migrate appends to
424
+ the same alive list they did, so the bucketing puts the warmed particles in next frame's simulate
425
+ list, the bounds pass measures them into their emitter's box this frame, and the draw draws them
426
+ this frame. The tick also writes the emitter's
427
+ `age` into its state row as it goes, so `EMITTER_AGE` reads as it would have, and the next real tick
428
+ adds its `dt` to a row that says `prewarm`.
429
+
430
+ **Only on the frame that needs it.** The host knows: an emitter is queued when it is registered with
431
+ a `prewarm` (`GPUParticleSystem#sync_membership`), the queue is uploaded and consumed by the frame
432
+ that records the loop, and an ordinary frame records nothing of this.
433
+
434
+ The editor's CPU preview (`editor/particles/effect/ParticleReferenceSimulation.js`) runs the same
435
+ loop over the same schedule before its first frame, and restarts the effect when the duration
436
+ changes: a pre-warm is a statement about how an effect starts, so the only way to show one is to
437
+ start it.
438
+
439
+ ### Culling & attach
440
+
441
+ An emitter that sets `EMITTER_FLAG.CULL` has the box on its state row tested against the camera
442
+ frustum by the tick pass, with the renderer's shared `chunk_aabb3_intersects_frustum`. The tick does
443
+ not measure it — it reads what the bounds passes published at the end of the previous frame — so all
444
+ that is in the spawn path is one frustum test. Because the box is where the particles are rather
445
+ than where the node is, an effect whose particles have drifted on screen keeps spawning even when
446
+ its emitter has not. Culled emitters spawn nothing and accrue sleep; on wake the tick adds a
447
+ catch-up of at most `max_catchup_seconds` of rate. Live particles keep simulating while their
448
+ emitter is culled, so its box keeps moving and keeps being published. HiZ occlusion is not
449
+ written.
450
+
451
+ ### Two orderings, and why they are separate
452
+
453
+ There are two independent permutations of the live particles each frame. They are produced by
454
+ different passes into different buffers and must not be conflated or derived from one another.
455
+
456
+ **Simulation order** (`coherence/`) — camera-independent, changes only when emitters or their
457
+ programs change. The alive list comes out of compaction (`atomicAdd(&counters[ALIVE])`), so its
458
+ order says nothing about which emitter a particle belongs to and a wave routinely holds particles
459
+ running different VM programs. That costs twice: the interpreter's ~40-case opcode `switch`
460
+ diverges, so a wave serially executes the union of the opcodes its lanes want; and `vm_reg[dst]` /
461
+ `vm_reg[s0]` index a VGPR-resident register file dynamically, which AMD reaches with
462
+ `v_movrels`/`v_movreld` behind a waterfall loop whose trip count is the number of *distinct*
463
+ register indices live in the wave. Grouping the wave by program collapses both at once.
464
+
465
+ The grouping is also **load-bearing**: the simulate shader reads its section parameters (code offset,
466
+ instruction count, constant offset) from the wave's first lane through `subgroupBroadcastFirst`
467
+ (`particle_wave_uniform`), which is what lets the driver fetch instruction words through the scalar
468
+ unit, branch the opcode switch on a scalar and index the register file through M0 without a
469
+ waterfall. A wave holding two programs would run one over the other's particles, so the coherence
470
+ pass is part of simulate's correctness. Emit keeps a per-lane decode: its waves span emitters.
471
+
472
+ The bucket key is the **program**, not the emitter: `ProgramHeap` deduplicates by content and hands
473
+ back a `program_id` the emitter record carries, so emitters compiled from the same graph share
474
+ one bucket. Each bucket is padded out to `PARTICLE_SIMULATE_WORKGROUP_SIZE` (64 — a multiple of both
475
+ plausible subgroup widths, so no shader needs to query `subgroup_size`), which is what makes the
476
+ grouping hold at wave granularity: an unpadded boundary lands mid-wave and that wave spans two
477
+ programs again. Padding entries carry `PARTICLE_SIM_PAD_SLOT` and the simulate shader drops them
478
+ *before* the compaction atomics, so they never enter the alive or dead list.
479
+
480
+ It is an **indirection** list. Particle records never move: slots are referenced by the dead list and
481
+ across frames, so relocating a record would leave both pointing at the wrong particle.
482
+
483
+ The three-phase shape is the depth sort's, with the padding added: histogram (atomic per bucket) →
484
+ round each bin up to a whole workgroup → `graph_prefix_scan_csdldf` (reused as-is — the one scan here
485
+ whose only cross-workgroup wait is CSDLDF's bounded spin-then-fallback) → scatter → fill the tails
486
+ with the sentinel. Nothing reads back to the CPU; the only host-supplied number is the bucket count,
487
+ which is static topology.
488
+
489
+ **Render order** (`sort/`) — per-frame, particles sorted back-to-front by view depth so
490
+ premultiplied transparency composites correctly. Counting sort on a quantised view depth
491
+ (`PARTICLE_SORT_BUCKET_COUNT` bins over `GPUParticleSystem#sort_depth_range`): histogram from the
492
+ camera uniform's view matrix, `graph_prefix_scan_csdldf` the bin counts to offsets, scatter (index
493
+ payload). The sorted index list feeds the instanced draw. Both per-particle passes are bounded by
494
+ the alive counter and dispatched by the args the bucket reset wrote from it.
495
+
496
+ Reordering the render list for program coherence would break transparency; depth-sorting the
497
+ simulation list would destroy the coherence. They share no buffer and no counter slot —
498
+ `PARTICLE_COUNTER.SORTED` belongs to the render order, `SIM_PADDED` to the simulation order.
499
+
500
+ ### Rendering
501
+
502
+ One render pipeline, one blend state = premultiplied: `color: (one, one-minus-src-alpha)`,
503
+ `alpha: (one, one-minus-src-alpha)`. Additive vs. alpha is chosen by the shader emitting premultiplied
504
+ RGB with alpha 0 (additive) vs. alpha a (normal) — GPU-driven, no pipeline switch. A 6-vertex quad is
505
+ expanded per instance (billboard/velocity-stretch/axis-lock/world modes). Fragment: sample atlas
506
+ (region + flipbook frame), optional forward lighting via `chunk_shade_standard_material_direct` +
507
+ IBL, optional soft-depth fade. Depth-tested, no depth write. Draws into the HDR color target.
508
+ Soft-depth fading scales both RGB and alpha for premultiplied inputs; straight-alpha and additive
509
+ inputs have alpha scaled before premultiplication. Both draw paths share `particle_fade_color`.
510
+
511
+ ### Drawing under the renderer: an AVBOIT side channel
512
+
513
+ Under `Renderer` the particles are not drawn by the billboard pass above. They go through the AVBOIT
514
+ transparency pipeline (`rasterize/native/avboit/`) as a **side channel** — the contract in
515
+ `AVBOITSideChannel.js` — so they composite order-independently with the transparent meshes and
516
+ against the opaque scene, through the same volume and the same accumulators. Quads are not packed
517
+ into meshlets; the orchestrator asks the channel for its contribution at the three points where the
518
+ meshlet buckets have just made theirs:
519
+
520
+ 1. **occupancy** (`shader_particle_avboit_occupancy`): one thread per emitter row, the emitter's
521
+ measured box (the culling box, off the state row) as a view-depth interval, dilated and OR-ed
522
+ into the occupancy bits. The interval is the box's centre depth plus or minus its extent along
523
+ the view axis — a support projection, not eight transformed corners. Per emitter, not per
524
+ particle — no contention, and a particle that strays lands on a run boundary rather than anywhere
525
+ wrong. Additive emitters mark nothing. The box matters here even for an emitter that never opted
526
+ into culling, which is the second reason it is measured for every emitter rather than only for
527
+ the ones that cull.
528
+ 2. **splat** (`shader_particle_avboit_splat`): the billboard quads rasterized at the volume's
529
+ resolution, coverage from the atlas alpha, extinction through `avboit_splat_slice` — the same
530
+ write path as a surface. Additive particles write nothing. A soft-depth particle is faded here
531
+ against the same scene depth the draw fades it against, one sample per voxel: what the volume
532
+ records has to be the coverage the accumulators are weighted by, or a card that fades out where
533
+ it meets the floor goes on darkening the floor through itself.
534
+ 3. **draw** (`shader_particle_avboit_draw`): the quads at internal resolution, depth-tested read-only
535
+ against the scene (which lets the soft fade read the same depth), fog composited, transmittance
536
+ from the integral. Which of the two it is, is decided by coverage rather than by the blend enum,
537
+ so an additive emitter and a premultiplied sprite with no alpha take the same path — as they do
538
+ in the splat. A covering particle writes `color / norm / extinction` as a surface does, in-scatter
539
+ included, weighted by the coverage it took. An emitting one writes the fourth accumulator,
540
+ **emission**: light attenuated by the fog and by the volume in front of it, and carrying none of
541
+ the fog's own in-scatter, because the resolve adds emission outside the normalization and the
542
+ scene behind already has it. That accumulator is the one addition to AVBOIT's resolve for this.
543
+ The whole decision is `chunk_particle_avboit_contribution`, and it is spec'd on the emulator.
544
+
545
+ No sort: the volume makes order irrelevant, so this path reads the alive list and the draw args
546
+ straight from `simulate`. The billboard vertex stage is one chunk
547
+ (`chunk_particle_billboard_vertex`) shared by the standalone renderer and both AVBOIT passes.
548
+ `GPUParticleSystem` implements the three methods; `Renderer.feature_particles_enabled` runs
549
+ `simulate` before the transparency pass and hands the system to the orchestrator. MBOIT is legacy
550
+ and gets no particle path.
551
+
552
+ ### Texture ownership: ECS
553
+
554
+ `engine/graphics3/GPUParticleEmitterSystem.js` owns the sprite atlas for one Shade scene. Construct it
555
+ with `(graphicsEngine, scene, assetManager)` and add it to the EntityManager. It observes the scene's
556
+ existing `ParticleEmitter` nodes; there is no second component or transform to synchronize. It is
557
+ opt-in and can coexist with the legacy CPU `ParticleEmitterSystem`.
558
+
559
+ ```js
560
+ entityManager.addSystem(new GPUParticleEmitterSystem(graphics, scene, assetManager));
561
+ ```
562
+
563
+ The system loads image URLs through `AssetManager` (`GameAssetType.Image`) and shares one patch per
564
+ URL. It reuses `TextureAtlas`, whose packer is `MaxRectanglesPacker`. When the last emitter releases
565
+ an image, its patch is removed. Loads completing after a URL change, removal or shutdown are ignored.
566
+ Failed images are reported and remain transparent until the URL changes or the emitter is re-added.
567
+ Untextured emitters use a white texel; pending images use a transparent texel.
568
+
569
+ At `FrameStart`, after the scene context has established transform rows, the system calls
570
+ `GPUParticleSystem.sync_membership()`, uploads changed atlas pixels, and refreshes all patch regions.
571
+ This occurs before particle simulation stages emitter rows. Every repack or resize therefore updates
572
+ existing emitters too. The extension enables the renderer's GPU particle feature for its scene;
573
+ rendering uses the existing AVBOIT path. The system owns a `GPUTextureContext` for the atlas, using
574
+ its resize and cached-view handling and advancing its version after each pixel upload. Shutdown
575
+ restores the default atlas and destroys the owned context. A replacement device gets a fresh upload
576
+ from the retained CPU atlas.
577
+
578
+ `GPUParticleSystem` itself does no asset loading or packing. Standalone callers can still supply
579
+ `atlas`/`atlas_sampler` and call `registry.set_atlas_region()` themselves. The `atlas` binding remains
580
+ a `GPUTextureView`, obtained from its owner's context; the default white texture is also owned
581
+ through `GPUTextureContext`.
582
+
583
+ ### Ownership: `GPUParticleSystem`
584
+
585
+ Constructed once against a `GPUSceneContext`, in the mould of `ReSTIRDI`. It owns every persistent
586
+ buffer — pool, alive lists, dead list, counters, both indirect-args buffers, the coherence and sort
587
+ scratch, the emitter database (both tables), the packed program buffer, the spawn-command ring — and grows the ones that depend on the emitter set inside `execute`. The
588
+ default atlas is a 1x1 white texel. `execute({graph, camera, scene_database, color, depth, dt})`
589
+ sweeps the scene, pushes what changed (on a command context opened only when something needs
590
+ one), and records the whole frame through `graph_particles`, returning the colour handle after
591
+ the draw. An integrator does not encode a particle pass by hand.
592
+
593
+ ## File layout
594
+
595
+ ```
596
+ src/shade/renderer/particles/
597
+ DESIGN.md
598
+ ParticleConstants.js pool stride, reg count, workgroup sizes, limits
599
+ isa/ParticleVMISA.js opcode table (single source of truth)
600
+ isa/ParticleProgram.js Program value type + (de)serialization to u32
601
+ isa/ParticleAssembler.js tiny assembler (build programs in tests/library)
602
+ vm/ParticleVMReference.js JS reference executor
603
+ vm/chunk_particle_vm.js WGSL interpreter CodeChunk (parity target)
604
+ vm/chunk_particle_vm_registers_fast.js register file in var<private> — the interpreter that runs
605
+ vm/chunk_particle_vm_registers_wide.js register file in storage — the fallback
606
+ shaders/chunk_particle_simulate_step.js one particle's simulate, shared by both simulate shaders
607
+ shaders/chunk_particle_emit_step.js one newborn's emit, shared by both emit shaders
608
+ graph/ParticleNodeDescription.js node type base (ports + lowering)
609
+ graph/ParticleNodeRegistry.js standard node library (a NodeRegistry)
610
+ graph/particle_graph_authoring.js authoring sugar over core/model/node-graph
611
+ graph/compile_particle_graph.js graph → Program (register alloc, const intern)
612
+ layout/ParticleLayout.js attribute schema ↔ record word offsets
613
+ data/PARTICLE_EMITTER_STRUCT.js the emitter row + flag/blend/projection encoding
614
+ data/particle_emitter_record.js emitter + context -> one row of that struct
615
+ data/PARTICLE_EMITTER_STATE.js the GPU-owned per-emitter state row (second table of the database)
616
+ data/PARTICLE_DATABASE_SPEC.js the GPUDatabase schema, and the emitter table's descriptor
617
+ data/PARTICLE_COUNTERS.js counter slot indices
618
+ data/particle_header.js record header encoding (row + generation), JS + WGSL
619
+ coherence/particle_program_buckets.js CPU reference for the simulation-order bucketing
620
+ coherence/shader_particle_bucket.js reset / histogram / pad / scatter / fill
621
+ coherence/graph_particle_bucket.js frame-graph wiring, around graph_prefix_scan_csdldf
622
+ runtime/ProgramHeap.js the one packed [header][constants][code] buffer, allocator-backed
623
+ runtime/ProgramPlacement.js where one program sits in it — what an emitter record carries
624
+ runtime/ParticleEmitter.js one emitter — a Node3D and an effect, with no GPU in it
625
+ runtime/GPUParticleEmitterContext.js what a registry assigns one: row, generation, placement, patch
626
+ runtime/EmitterRegistry.js GPUDatabase(emitters) + the live emitters + their contexts + the program heap
627
+ runtime/create_particle_effect.js layout + INIT/UPDATE graphs → compiled effect
628
+ shaders/shader_particle_spawn_commands.js CPU bursts → pending spawns
629
+ shaders/shader_particle_emitter_tick.js per-emitter spawn decision
630
+ shaders/shader_particle_emit.js per-newborn INIT, sized by the scan
631
+ shaders/shader_particle_simulate.js per-particle UPDATE + compaction + orphan kill
632
+ shaders/shader_particle_finalize.js reset counters (+ budget) / build indirect args
633
+ shaders/shader_particle_render.js billboard pipeline
634
+ shaders/chunk_particle_setup_context.js VM builtins from the emitter's node, + the RNG seed
635
+ shaders/chunk_particle_emitter_age.js the tick pass's running age, off the emitter_state table
636
+ shaders/chunk_particle_spawn_budget.js take from the frame's spawn budget; both ticks spend it
637
+ warmup/warmup_schedule.js ticks and tick length for a batch of pre-warms
638
+ warmup/shader_particle_warmup.js init / reset / tick over the host's queue / migrate
639
+ warmup/chunk_particle_warmup_step.js one particle's advance: per-lane decode, jittered tick
640
+ warmup/shader_particle_warmup_advance.js the advance over the fast file (+ _wide.js)
641
+ warmup/graph_particle_warmup.js the loop, recorded on a frame that registered a pre-warm
642
+ shaders/chunk_particle_curve_disabled.js the no-curves-bound CURVE provider
643
+ shaders/chunk_particle_billboard_vertex.js the billboard vertex stage, shared by every draw
644
+ shaders/shader_particle_avboit_*.js occupancy / splat / draw through AVBOIT
645
+ bounds/* the measured emitter box: accumulate, publish
646
+ sort/* render-order depth sort
647
+ graph_particles.js simulate (to the draw args) + the standalone sort and draw
648
+ graph_particles_avboit.js the AVBOIT side channel: occupancy, splat, draw
649
+ GPUParticleSystem.js the feature: owns the buffers, sweeps the scene, execute()
650
+ *.spec.js co-located tests (JS + emulator + software device)
651
+ ```
652
+
653
+ ## Testing
654
+
655
+ - Pure JS: layout packing, assembler, register allocator, node compiler, heap/emitter/registry
656
+ bookkeeping (via `SoftwareGPUDevice`), the header encoding.
657
+ - WGSL via emulator: VM opcode parity (JS vs WGSL), curve sampling, sort key/depth, atlas+flipbook
658
+ UV math, billboard expansion, soft-depth fade, and every compute pass on small fixtures — the tick
659
+ (rate integration, generations, pending, budget, culling, page-group walk), emit (expansion over
660
+ the scan, node position, orphans), simulate, spawn commands, finalize, bucket, sort. Each such spec
661
+ carries its own copy of the scaffolding: a real scene with a `GPUSceneContext` and a real registry,
662
+ laid out by `gpu_database_words` the way the GPU holds both databases, so the shaders read rows the
663
+ real writers wrote.
664
+ - Multi-frame capstone (`particle_lifecycle.spec`): reset → commands → tick → (JS scan) → emit →
665
+ simulate across frames, with the GPU deciding to spawn, a CPU burst, the budget, and death.
666
+ - Orchestration on `SoftwareGPUDevice`: `GPUParticleSystem.spec` (scene sweep, growth, uploads only
667
+ when needed, the ring), `graph_particles.spec` (pass sequence, indirect sizing, the two orderings,
668
+ every persistent buffer imported exactly once), and the playground's `particle_scene.spec` (the
669
+ integration: geometry → hierarchy → particles → draw).
670
+ - Cross-parity meta-tests: every ISA opcode implemented by both executors.
671
+
672
+ Regression checks use the existing tools: seeded software-device buffer copies for list growth,
673
+ recorded pass ordering for program edits, emulator math for blend/fade and unknown bounds, and real
674
+ AssetManager/TextureAtlas bookkeeping with a stub image loader for asynchronous atlas lifetimes.
675
+ No full software WebGPU implementation or cross-frame shader simulator is added. The emulator's
676
+ `subgroupBroadcastFirst` is identity, so grouping correctness is checked at its input contract.
677
+
678
+ End-to-end GPU is out of scope for node (no binary GPU dep); emulator + JS references give the
679
+ coverage the brief asks for.
680
+
681
+ ### Emulator caveats discovered (all correct on real hardware)
682
+
683
+ The in-repo WGSL emulator (unit-test tool, not GPU parity) has limitations the shaders are written
684
+ around, documented in the memory note: `switch` case bodies are dropped (use if/else); repeated
685
+ read-modify-write of one location across a loop is corrupted (RNG is verified in isolation, not
686
+ through the interpreter loop); u32 multiply is imprecise for large operands (PCG can't be checked
687
+ through it); `/` is FLOAT division, so integer division only truncates where the result is stored
688
+ straight into a typed array — `(n + 63u) / 64u * 64u` reads as 66 there and 64 on a device, and a
689
+ binary search's `mid` is a shift, not a divide; matrix column indexing (`m[0]`) is not evaluated,
690
+ so an axis scale is read as `m * vec4(1,0,0,0)`. `array<atomic<u32>>` bindings are modeled as
691
+ `{value}` objects in tests; a pass that binds the same buffer as plain `array<u32>` gets a
692
+ `Uint32Array`. The scene context writes a row's `global` one publish behind (the hierarchy pass
693
+ recomposes it on a device), so a node's matrix under the emulator is the one it had when its row
694
+ was seeded.
695
+
696
+ ### Renderer integration
697
+
698
+ Done: `Renderer.feature_particles_enabled` (off by default), `Renderer.particles(scene_ctx)` creating
699
+ one system per scene context on first use, `simulate` recorded before the transparency pass and the
700
+ system passed to the AVBOIT orchestrator as its side channel. Remaining:
701
+ 1. For lit emitters, compile a simulate/render variant that swaps the disabled curve/lighting hooks
702
+ for `chunk_particle_curve_animation` (+ bind `GPUAnimationManager.database.buffer`) and a
703
+ shading-chunk-backed lighting hook (+ the `SHADING_LIGHT_RESOURCE_GROUP`).