oidn-web 0.3.5 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (65) hide show
  1. package/CHANGELOG.md +94 -0
  2. package/README.md +208 -8
  3. package/dist/oidn.js +4699 -22516
  4. package/dist/oidn.umd.cjs +989 -5796
  5. package/lib/UNet.d.ts +111 -26
  6. package/lib/UNet.js +310 -329
  7. package/lib/UNet.js.map +1 -1
  8. package/lib/WGPUComputePass.d.ts +1 -1
  9. package/lib/WGPUComputePass.js +6 -4
  10. package/lib/WGPUComputePass.js.map +1 -1
  11. package/lib/backend.d.ts +1 -4
  12. package/lib/backend.js +28 -44
  13. package/lib/backend.js.map +1 -1
  14. package/lib/finalRgbShader.d.ts +13 -0
  15. package/lib/finalRgbShader.js +160 -0
  16. package/lib/finalRgbShader.js.map +1 -0
  17. package/lib/graphOptimizer.d.ts +54 -0
  18. package/lib/graphOptimizer.js +215 -0
  19. package/lib/graphOptimizer.js.map +1 -0
  20. package/lib/hdrTransfer.d.ts +14 -0
  21. package/lib/hdrTransfer.js +61 -0
  22. package/lib/hdrTransfer.js.map +1 -0
  23. package/lib/main.d.ts +43 -11
  24. package/lib/main.js +9 -5
  25. package/lib/main.js.map +1 -1
  26. package/lib/modelSpec.d.ts +80 -0
  27. package/lib/modelSpec.js +270 -0
  28. package/lib/modelSpec.js.map +1 -0
  29. package/lib/nativeUNet.d.ts +103 -0
  30. package/lib/nativeUNet.js +2064 -0
  31. package/lib/nativeUNet.js.map +1 -0
  32. package/lib/process.d.ts +5 -11
  33. package/lib/process.js +38 -49
  34. package/lib/process.js.map +1 -1
  35. package/lib/resourceTracker.d.ts +26 -0
  36. package/lib/resourceTracker.js +65 -0
  37. package/lib/resourceTracker.js.map +1 -0
  38. package/lib/tileScheduler.d.ts +61 -0
  39. package/lib/tileScheduler.js +199 -0
  40. package/lib/tileScheduler.js.map +1 -0
  41. package/lib/webnnUNet.d.ts +52 -0
  42. package/lib/webnnUNet.js +535 -0
  43. package/lib/webnnUNet.js.map +1 -0
  44. package/package.json +16 -5
  45. package/src/UNet.ts +463 -437
  46. package/src/WGPUComputePass.ts +6 -4
  47. package/src/backend.ts +33 -59
  48. package/src/finalRgbShader.ts +186 -0
  49. package/src/graphOptimizer.ts +300 -0
  50. package/src/hdrTransfer.ts +88 -0
  51. package/src/main.ts +95 -20
  52. package/src/modelSpec.ts +414 -0
  53. package/src/nativeUNet.ts +2655 -0
  54. package/src/process.ts +46 -71
  55. package/src/resourceTracker.ts +94 -0
  56. package/src/tileScheduler.ts +330 -0
  57. package/src/webnnUNet.ts +812 -0
  58. package/lib/helper.d.ts +0 -4
  59. package/lib/helper.js +0 -33
  60. package/lib/helper.js.map +0 -1
  61. package/lib/kernels.d.ts +0 -1
  62. package/lib/kernels.js +0 -26
  63. package/lib/kernels.js.map +0 -1
  64. package/src/helper.ts +0 -43
  65. package/src/kernels.ts +0 -31
package/CHANGELOG.md ADDED
@@ -0,0 +1,94 @@
1
+ # Changelog
2
+
3
+ All notable changes to this project are documented in this file.
4
+
5
+ ## [0.5.0] - 2026-10-09
6
+
7
+ oidn-web 0.5.0 simplifies device setup, makes tiled execution always settle, and switches FP16 convolutions to implicit GEMM. It contains breaking changes for TypeScript callers that pass `adapterInfo`, for code that constructs `UNet` directly, and for code that relies on square tiles or per-frame tile pacing.
8
+
9
+ ### Breaking changes
10
+
11
+ - `initUNetFromURL` and `initUNetFromBuffer` take `{ device }` as the second argument. `adapterInfo` was only needed by the removed TensorFlow.js backend and is no longer accepted by the type. Extra properties are ignored at runtime, but TypeScript rejects an object literal that still contains `adapterInfo`.
12
+ - `new UNet(tensors, device, options)` takes a `GPUDevice` instead of a `{ device, adapterInfo }` object.
13
+ - `tileExecute` now continues on the event loop between tiles by default instead of waiting for `requestAnimationFrame`. Pass `scheduling: 'animation-frame'` to keep pacing tiles to display frames.
14
+ - Tiles are balanced rectangles with overlap only on edges shared with another tile, instead of fixed-size squares. Tile count, position, and size reported to `progress` differ from 0.4.0. The default `dynamicTile.initialTileSize` is now 432.
15
+ - `UNetExecutionStats` no longer has `tileWidth` and `tileHeight`. Use `tileColumns`, `tileRows`, `tileOverlap`, `inputPixelCount`, and `inputShapeCount` instead.
16
+ - FP16 convolutions use implicit GEMM by default. Output differs slightly from 0.4.0 because accumulation order changed. Set `kernel: 'direct'` to reproduce the 0.4.0 FP16 path.
17
+
18
+ ### Added
19
+
20
+ - `hdrTransfer: 'log'` for RTLightmap HDR models, alongside the default PU transfer.
21
+ - `error` callback on `tileExecute` for asynchronous failures, WebGPU device loss, and exceptions thrown by `progress` or `done`. `progress` and `done` may return promises.
22
+ - `scheduling: 'event-loop' | 'animation-frame'`, `tileOverlap`, and `wholeImage` options on `tileExecute`.
23
+ - `prepareForImage(width, height)` to create per-shape GPU resources before the first denoise.
24
+ - `planTileGrid` and its `TilePlan`, `PlannedTile`, and `TileRect` types.
25
+ - Experimental `gemm` tuning options. These are intended for benchmarking and are not covered by semver.
26
+ - `hdrTransfer` and `activeExecutionCount` in `getRuntimeInfo()`, and the active GEMM configuration under `kernel.gemm`.
27
+
28
+ ### Changed
29
+
30
+ - Implicit GEMM tiles are selected by output alignment and GPU limits, with optimized addressing, weight layout, register tiles, shared-memory layout, and pooling access.
31
+ - The final RGB convolution uses an adaptive shared-memory cache.
32
+ - Adaptive tile sizing uses the smoothed P75 tile GPU time, excludes the cold first tile, ignores cancelled and single-tile work, and buckets input shapes to at most two sizes.
33
+ - `animation-frame` scheduling falls back to a 100 ms timer so execution still completes in hidden tabs.
34
+
35
+ ### Fixed
36
+
37
+ - Tiled execution always settles with `done` or `error`, or stops silently after abort, including when `requestAnimationFrame` never fires or the device is lost.
38
+ - Completed executions release their device-loss listeners.
39
+ - Edge tiles whose size is not a multiple of 16 replicate edge pixels into the padded model input.
40
+ - The npm package no longer includes benchmark results, tests, and scripts.
41
+
42
+ ### Migration notes
43
+
44
+ - Replace `{ device, adapterInfo }` with `{ device }`.
45
+ - Replace `new UNet(tensors, { device, adapterInfo }, options)` with `new UNet(tensors, device, options)`.
46
+ - Interactive renderers that share the GPU with OIDN should pass `scheduling: 'animation-frame'`.
47
+ - Pass an `error` callback to handle failures. Without one, failures are logged to the console.
48
+
49
+ ## [0.4.0] - 2026-08-20
50
+
51
+ oidn-web 0.4.0 replaces the TensorFlow.js inference stack with a purpose-built, model-driven WebGPU runtime. Existing `initUNetFromURL` and `initUNetFromBuffer` integrations remain supported while gaining native FP16, adaptive scheduling, runtime diagnostics, and stronger model validation.
52
+
53
+ ### Highlights
54
+
55
+ - Removed TensorFlow.js and its WebGPU backend from the runtime dependencies.
56
+ - Added a custom WGSL U-Net executor with automatic FP16 selection and a deterministic FP32 fallback.
57
+ - Added built-in topology descriptors for the current OIDN small and large RT models, including clean auxiliary variants.
58
+ - Reduced the reference Vite ESM bundle from approximately 168.8 KB to 31.1 KB gzip compared with 0.3.5, a reduction of about 82%.
59
+ - Reached a 27.6 ms median for a 512 x 512 Clean Aux Large inference on the tested Apple GPU, versus 35.0 ms for the TensorFlow.js baseline (1.27x faster). Performance depends on the browser, GPU, model, and tile size.
60
+
61
+ ### Added
62
+
63
+ - Model-driven graph planning and lifetime-aware activation-buffer reuse.
64
+ - Versioned `UNetModelSpec` descriptors, topology detection, and validation of tensor names, layouts, data types, byte lengths, kernel shapes, bias shapes, and graph channel flow.
65
+ - `modelSpec` initialization option for future or custom OIDN model topologies.
66
+ - `npm run model:inspect` for model hashes, tensor signatures, and descriptor compatibility checks.
67
+ - Dynamic tile sizing with configurable minimum, initial, and maximum tile sizes and a target GPU time.
68
+ - GPU queue backpressure between tiles to keep cancellation and rendering interaction responsive.
69
+ - Runtime diagnostics through `getRuntimeInfo()`, including the selected engine, precision, model, kernel capabilities, tile state, and resource statistics.
70
+ - Per-layer GPU profiling through `profileNextExecution()` and `getLastExecutionProfile()` when `timestamp-query` is enabled.
71
+ - Cross-version browser benchmarks with queue-complete timings, sampled output validation, TFJS baseline comparison, and per-layer GPU profiles.
72
+ - Explicit, idempotent resource disposal and resource lifecycle accounting.
73
+ - Automated model, graph, scheduler, resource leak, allocation rollback, and late-async-cleanup tests.
74
+
75
+ ### Changed
76
+
77
+ - U-Net inference now runs directly in WGSL. The stable `auto` engine selects the native WGSL backend and no longer initializes TensorFlow.js.
78
+ - FP16-capable devices retain TZA half-float weights and use short FP16 FMA accumulation groups folded into FP32 accumulators. The final output remains FP32.
79
+ - Devices without `shader-f16` automatically use native FP32 inference.
80
+ - Convolution tensors and activations use a blocked four-channel layout with channel-specialized shaders.
81
+ - Decoder `upsample + concat + conv` patterns are fused, and all network passes for a tile are encoded into one command buffer.
82
+ - Shape-independent pipelines are compiled asynchronously before initialization resolves to avoid first-denoise shader compilation stalls.
83
+ - `maxTileSize` is now a hard upper bound for the adaptive tile controller.
84
+ - A shared `GPUDevice` uses FP16 only when `shader-f16` was requested during device creation; WebGPU features cannot be enabled afterward.
85
+
86
+ ### Migration notes
87
+
88
+ - No changes are required for basic `initUNetFromURL` or `initUNetFromBuffer` usage.
89
+ - Call `dispose()` when a U-Net instance is no longer needed.
90
+ - When supplying an existing `GPUDevice`, request `shader-f16` before creating the device if FP16 inference is desired.
91
+ - Set `dynamicTile: false` to restore fixed-size tiling.
92
+
93
+ [0.5.0]: https://github.com/pissang/oidn-web/compare/v0.4.0...v0.5.0
94
+ [0.4.0]: https://github.com/pissang/oidn-web/compare/v0.3.5...v0.4.0
package/README.md CHANGED
@@ -9,9 +9,22 @@ It's used in the [Vector to 3D](https://www.figma.com/community/plugin/126460021
9
9
  | :----------------------------------------------------------------------------------------------: | :--------------------------------------------------------------------------------: | :--------------------------------------------------------------------------------------: |
10
10
  | ![](https://github.com/pissang/oidn-web/blob/main/examples/test/ground-truth.png 'Ground Truth') | ![](https://github.com/pissang/oidn-web/blob/main/examples/test/noisy.png 'Noisy') | ![](https://github.com/pissang/oidn-web/blob/main/examples/test/denoised.png 'Denoised') |
11
11
 
12
- ## How it Works.
12
+ ## How it works
13
13
 
14
- It uses [tfjs](https://github.com/tensorflow/tfjs) to build the UNet model used by the OIDN. Then use the model to do prediction with a WebGPU backend from the image data.
14
+ The OIDN U-Net runs directly on WebGPU with model-driven WGSL compute
15
+ pipelines. Convolution activations use a blocked
16
+ four-channel layout, decoder `upsample + concat + conv` patterns are fused,
17
+ and all network dispatches for a tile are submitted in one command buffer.
18
+ FP16 and FP32 convolutions use channel-specialized implicit-GEMM tiles by
19
+ default, with a direct convolution for the final output layer and separate
20
+ max-pool passes.
21
+
22
+ TZA half-float weights stay half-float when the device enables `shader-f16`.
23
+ FP16 products are accumulated in short half-precision groups and periodically
24
+ folded into FP32 accumulators; the final output is FP32. Devices without
25
+ `shader-f16` automatically use the native FP32 path. Shape-independent compute
26
+ pipelines compile asynchronously before initialization resolves, so first-use
27
+ shader compilation does not interrupt an interactive denoise.
15
28
 
16
29
  ## How to Use
17
30
 
@@ -37,12 +50,12 @@ initUNetFromURL('./weights/rt_ldr.tza').then((unet) => {
37
50
  .getContext('2d')
38
51
  .getImageData(0, 0, width, height);
39
52
 
40
- // Tile execute the denoising.
41
- // If the resolution is high. It will split the input into tiles and execute one tile per frame.
53
+ // Tile execute the denoising. High resolutions use balanced rectangular
54
+ // tiles with overlap only at boundaries shared by another tile.
42
55
  const abortDenoising = unet.tileExecute({
43
56
  // The color input for LDR image is 4 channels.
44
57
  // In the format of Uint8ClampedArray or Uint8Array.
45
- color: { data: noisyImageData, width, height },
58
+ color: noisyImageData,
46
59
  done(denoised) {
47
60
  console.log('Finished');
48
61
  },
@@ -76,6 +89,22 @@ initUNetFromURL('./weights/rt_hdr.tza', undefined, {
76
89
  });
77
90
  ```
78
91
 
92
+ HDR transfer defaults to the PU curve used by the regular RT models. Models
93
+ trained for the RTLightmap filter use the logarithmic curve from upstream OIDN;
94
+ select it explicitly when loading such weights:
95
+
96
+ ```ts
97
+ const lightmap = await initUNetFromURL('./weights/rtlightmap_hdr.tza', undefined, {
98
+ hdr: true,
99
+ hdrTransfer: 'log'
100
+ });
101
+ ```
102
+
103
+ The `log` transfer maps `y` to `log(1 + y) / log(65505)` and reverses this
104
+ before writing HDR output. For GPU-buffer inputs, callers may continue to
105
+ pre-scale the complete image once and leave the runtime's `inputScale` at its
106
+ existing default of `1`.
107
+
79
108
  ### Use auxiliary images
80
109
 
81
110
  ```ts
@@ -111,9 +140,8 @@ If you already have a WebGPU path tracer. You can integrate the oidn-web into yo
111
140
  initUNetFromURL(
112
141
  './weights/rt_hdr_alb_nrm.tza',
113
142
  {
114
- // Share GPUDevice and GPUAdapterInfo to the TFJS WebGPU backend
115
- device,
116
- adapterInfo
143
+ // Share the GPUDevice with the native WGSL runtime.
144
+ device
117
145
  },
118
146
  {
119
147
  aux: true,
@@ -153,6 +181,178 @@ initUNetFromURL('./weights/rt_hdr_alb_nrm_small.tza', ...);
153
181
 
154
182
  Other combinations can be found in the [oidn-weights](https://github.com/RenderKit/oidn-weights)
155
183
 
184
+ ### FP16 and runtime information
185
+
186
+ Standalone initialization requests `shader-f16` when the adapter supports it.
187
+ When sharing a device, optional features must be requested when that device is
188
+ created; WebGPU features cannot be enabled afterward.
189
+
190
+ ```ts
191
+ const requiredFeatures = adapter.features.has('shader-f16')
192
+ ? ['shader-f16']
193
+ : [];
194
+ const device = await adapter.requestDevice({ requiredFeatures });
195
+
196
+ const unet = await initUNetFromURL(
197
+ modelUrl,
198
+ { device },
199
+ {
200
+ aux: true,
201
+ hdr: true,
202
+ precision: 'auto' // 'fp16' enforces support; 'fp32' is deterministic fallback
203
+ }
204
+ );
205
+
206
+ console.log(unet.getRuntimeInfo());
207
+ // { gpuEngine: 'wgsl', precision: 'fp16', model: 'oidn-unet-large-v1', ... }
208
+ ```
209
+
210
+ For one-shot native GPU timings, request a profile immediately before an
211
+ execution. This is available when the shared device enabled `timestamp-query`:
212
+
213
+ ```ts
214
+ if (unet.profileNextExecution()) {
215
+ unet.tileExecute({
216
+ color,
217
+ albedo,
218
+ normal,
219
+ done: async () => {
220
+ console.table((await unet.getLastExecutionProfile()).layers);
221
+ }
222
+ });
223
+ }
224
+ ```
225
+
226
+ ### Updating to a new OIDN model
227
+
228
+ TZA stores tensors but not the executable graph. The runtime therefore keeps
229
+ the graph in a versioned `UNetModelSpec`, separate from shader and precision
230
+ code. Built-in descriptors cover the current OIDN small and large RT U-Nets.
231
+ At load time the descriptor is detected from the complete tensor-name set, and
232
+ tensor layout, dtype, byte length, kernel shape, bias shape, and graph channel
233
+ flow are validated before GPU resources are created.
234
+
235
+ If an OIDN update keeps one of these topologies and tensor names, changed
236
+ channel widths are handled automatically. If it adds or renames nodes, add a
237
+ new descriptor (or pass `modelSpec`) and its validation fixture. Existing graph
238
+ fusion rules apply to the new descriptor without changes to WGSL kernels.
239
+
240
+ Use the inspection command to get a stable SHA-256, full tensor signature, and
241
+ descriptor compatibility result for an upstream weight file:
242
+
243
+ ```shell
244
+ npm run model:inspect -- weights/rt_hdr_alb_nrm.tza
245
+ ```
246
+
247
+ ```ts
248
+ const unet = await initUNetFromURL(newModelUrl, undefined, {
249
+ aux: true,
250
+ hdr: true,
251
+ modelSpec: newOidnModelSpec
252
+ });
253
+ ```
254
+
255
+ ### GPU backpressure, scheduling, and dynamic tiles
256
+
257
+ `tileExecute` waits for the submitted GPU work of a tile before scheduling the
258
+ next tile. This keeps at most one OIDN tile in flight, which makes cancellation
259
+ responsive instead of leaving queued denoising work ahead of interactive
260
+ rendering.
261
+
262
+ Between tiles, JavaScript yields according to `scheduling`:
263
+
264
+ - `'event-loop'` (default) continues on the next macrotask. Execution finishes
265
+ as soon as the GPU allows and also completes in background tabs.
266
+ - `'animation-frame'` waits for the next display frame, so at most one tile
267
+ runs per frame. Use it when OIDN shares the GPU with an interactive renderer
268
+ and frame rate matters more than denoise latency. If no frame arrives within
269
+ 100 ms (for example in a hidden tab), the next tile runs anyway.
270
+
271
+ Asynchronous failures, including WebGPU device loss and exceptions thrown by
272
+ `progress` or `done`, are reported to the optional `error` callback. Unless the
273
+ returned abort function is called first, every execution ends by calling `done`
274
+ or `error`; without an `error` callback, failures are logged to the console.
275
+
276
+ Tile sizing is adaptive by default. `maxTileSize` is a hard upper bound. Each
277
+ execution partitions the image into balanced rectangular output regions and
278
+ adds model context only on edges shared with another tile. Input shapes are
279
+ bucketed to at most two sizes so the native execution cache remains stable.
280
+ The smoothed P75 GPU time of completed tiled executions adjusts the maximum tile
281
+ size used by the next execution. The potentially cold first tile is excluded,
282
+ cancelled and single-tile work is ignored, and a layout is held for at least two
283
+ complete executions. The default range starts at 432 pixels, does not go below
284
+ 256, changes in 16-pixel steps, and targets about 16 ms of GPU work per tile.
285
+
286
+ ```ts
287
+ const unet = await initUNetFromURL('./weights/rt_hdr_alb_nrm.tza', undefined, {
288
+ aux: true,
289
+ hdr: true,
290
+ maxTileSize: 512,
291
+ dynamicTile: {
292
+ minTileSize: 256,
293
+ initialTileSize: 432,
294
+ targetTileTimeMs: 16
295
+ }
296
+ });
297
+
298
+ // An interactive renderer can pace tiles to display frames. A smaller halo
299
+ // lowers per-tile cost; the default overlap is half of the model receptive
300
+ // field rounded up to 16 pixels.
301
+ const abortDenoising = unet.tileExecute({
302
+ color,
303
+ albedo,
304
+ normal,
305
+ tileOverlap: 80,
306
+ scheduling: 'animation-frame',
307
+ done(denoised) {
308
+ // ...
309
+ },
310
+ error(reason) {
311
+ console.error(reason);
312
+ }
313
+ });
314
+
315
+ // Restore fixed-size behavior when deterministic tiling is preferred.
316
+ initUNetFromURL('./weights/rt_hdr_alb_nrm.tza', undefined, {
317
+ aux: true,
318
+ hdr: true,
319
+ maxTileSize: 512,
320
+ dynamicTile: false
321
+ });
322
+ ```
323
+
324
+ ### Warm up and whole-image execution
325
+
326
+ The native runtime creates per-shape GPU buffers and pipelines the first time
327
+ it sees a tile input shape. Call `prepareForImage` with the expected image size while a loading
328
+ state is still visible, so the first `tileExecute` does not stall:
329
+
330
+ ```ts
331
+ await unet.prepareForImage(width, height);
332
+ ```
333
+
334
+ Pass `wholeImage: true` to `tileExecute` (and to `prepareForImage`) to run the
335
+ complete image as one tile and skip tiling overhead. It ignores `maxTileSize`,
336
+ so the image must fit the device's buffer and dispatch limits.
337
+
338
+ ### Benchmark the native runtime against TFJS
339
+
340
+ The browser benchmark automatically finds the nearest ancestor whose package
341
+ still depends on TensorFlow.js, builds that commit in a temporary worktree, and
342
+ compares it with the current WGSL FP32 and FP16 runtimes. Each measured run
343
+ waits for the WebGPU queue to finish, so the result includes execution rather
344
+ than only JavaScript command submission. It also samples output against the
345
+ TFJS FP32 result and, when timestamp queries are supported, reports the five
346
+ most expensive native network nodes.
347
+
348
+ ```shell
349
+ npm run benchmark -- --width 512 --height 512 --tile-size 512 --runs 5
350
+ ```
351
+
352
+ Results are printed as a table and written to
353
+ `benchmarks/results/latest.{json,md}`. Use `--baseline <commit>` to pin an
354
+ explicit historical version or `--chrome <path>` to select a browser.
355
+
156
356
  ## Credits
157
357
 
158
358
  Huge thanks to Max Liani for his series: https://maxliani.wordpress.com/2023/03/17/dnnd-1-a-deep-neural-network-dive/. My work is mostly inspired by it.