davinci-resolve-mcp 2.69.3 → 2.70.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,751 @@
1
+ # Running DaVinci Resolve headless
2
+
3
+ Blackmagic's scripting README gives headless mode two sentences: launch with
4
+ `-nogui` and "the various scripting APIs will continue to work as expected."
5
+ That is close enough to true to be misleading, because it says nothing about
6
+ what an agent should *do* differently — and this repository's own working notes
7
+ had drifted to the opposite belief, that Gallery and still-export surfaces were
8
+ headless-only failures.
9
+
10
+ This page replaces both with measurements. Everything under "Verified" was
11
+ observed on this machine against DaVinci Resolve Studio 19.1.3.7 on macOS;
12
+ everything under "Present but not exercised" was read out of the application
13
+ binary and has not been run. The distinction is kept because an unverified flag
14
+ that an agent tries anyway is worse than one it never heard of.
15
+
16
+ The differential itself lives in
17
+ [headless-capability-matrix.md](headless-capability-matrix.md), regenerated by
18
+ `scripts/headless_differential.py`.
19
+
20
+ ## The headline, and its limits
21
+
22
+ **Headless is not a capability-reduced mode.** Across 238 paired observations —
23
+ 139 read-only API probes plus 13 write scenarios covering pages, media pool,
24
+ timeline editing, colour, Fusion comps, Gallery, render-to-disk, interchange
25
+ export, layout presets, playhead and project settings — there were **zero**
26
+ capabilities that worked with a UI and failed without one.
27
+
28
+ The one difference found runs the other way: headless cannot raise a modal
29
+ dialog, and modals are the most common way an agent-driven Resolve session
30
+ wedges.
31
+
32
+ **That is a claim about capability and about modals. It is not a claim about
33
+ stability.** See the next section before pointing a long delivery at it.
34
+
35
+ ## What is and is not measured about stability
36
+
37
+ Measured, and clean: **10 consecutive ProRes 422 HQ renders in one headless
38
+ session**, from a 1045-frame JPEG 2000 / MXF OP1A / 12-bit source — the decode
39
+ path a DCP video track file takes — with no crash, no death, and no leak. See
40
+ the stress table below. That covers repeated jobs in one session, heavy source
41
+ decode, and the encoder teardown at the end of every write, ten times over.
42
+
43
+ Still **not** measured:
44
+
45
+ - **Sustained encode over hours.** Ten 45-second renders is not a feature-length
46
+ delivery. Thermal behaviour and codec-library state over thousands of frames
47
+ remain untested.
48
+ - **The operator's actual footage, codec settings and durations.**
49
+ - **ProRes into MXF** — not because it was skipped, but because this build does
50
+ not offer it: `GetRenderCodecs` returns ProRes only under `mov`.
51
+
52
+ **Field report:** rendering ProRes from DCP sources has crashed *on completion of
53
+ writes* in headless sessions on this machine. The stress leg above exercises that
54
+ same teardown path ten times without a failure, so the write-completion path is
55
+ not trivially broken headless — but ten short renders cannot disprove a fault
56
+ that appears on a long one, and this does not.
57
+
58
+ Note what changed here. An earlier version of this page reported the headless
59
+ stress leg as hanging reproducibly at 0/10, which read as evidence for headless
60
+ instability. It was not: the harness was calling `SaveProject()` on the
61
+ never-saved `Untitled Project`, which blocks the whole application. **A harness
62
+ bug had been sitting in the "headless is unstable" column.** Once guarded, the
63
+ leg ran straight through on the first attempt.
64
+
65
+ Two artefacts worth keeping from the log investigation:
66
+
67
+ - A per-session GUI/headless marker for classifying logs after the fact:
68
+ `Show splash screen` appears in GUI sessions and never in headless ones.
69
+ - A full WARN/ERROR diff between known-headless and known-GUI sessions found
70
+ **nothing headless-specific**. In particular `Encoder audio was not flushed` —
71
+ the most suggestive warning for a write-completion crash — occurs in both
72
+ modes, and more often in the GUI ones. It is not a headless marker.
73
+
74
+ **Practical rule:** headless is qualified for orchestration, edit, conform,
75
+ analysis and renders of this scale. Qualify it on your own footage before a
76
+ long-form or professional-container delivery, and keep a GUI fallback for those.
77
+
78
+ ### Differencing the two modes exhaustively
79
+
80
+ `scripts/mode_matrix.py` is the second-generation differential, and it exists
81
+ because the first one **could not have found the `SaveProject` hang**. That
82
+ harness called everything in one process with no timeouts, so a call that never
83
+ returns simply ended the run; its clean report of "zero mode-dependent failures"
84
+ was partly a description of what it was able to survive.
85
+
86
+ Three changes make exhaustive comparison possible:
87
+
88
+ - **`hang` is a first-class verdict.** A call that returns `False` is
89
+ survivable; one that never returns takes the session with it and degrades the
90
+ instance afterwards. Collapsing them into "failed" hides the worse bug.
91
+ - **Supervisor and worker are separate processes.** The worker streams one
92
+ flushed JSONL result per probe, in catalogue order. When it dies, the
93
+ supervisor knows the first unrecorded probe was in flight, records it as
94
+ `hang`, restarts Resolve, waits for health, and resumes after it. A hang costs
95
+ one probe and one restart, not the run. Nothing in-process can do this: a
96
+ Python signal handler cannot preempt a thread parked in fusionscript's
97
+ `pthread_cond_wait`, which is why the alarm-based approach only ever worked by
98
+ luck.
99
+ - **State is part of each probe.** `SaveProject` is harmless on a named project
100
+ and fatal on the never-saved default, so probes declare which project must be
101
+ current. A catalogue that only ever ran against a well-formed scratch project
102
+ is exactly what reported no differences the first time.
103
+
104
+ Phase breadcrumbs distinguish a hang in a probe's own call from a hang while
105
+ reaching its state — a real distinction, since setting up the "untitled" state
106
+ means closing a project, and closing a *modified* project raises the GUI save
107
+ dialog. Without the breadcrumb that dialog was attributed to the probe that
108
+ never got to run.
109
+
110
+ ### Stressing it without production assets
111
+
112
+ `scripts/render_stress.py` closes most of that gap before real footage is
113
+ available. It renders repeatedly in one session and watches for the things that
114
+ distinguish *causes* rather than merely recording that something died:
115
+
116
+ - `crash_archive.txt` is bracketed per iteration, so a stack is attached to the
117
+ render that produced it. Resolve appends every crash to one growing file with
118
+ no timestamps, so without the bracket "did this render crash" is guesswork.
119
+ - The Resolve PID is watched from outside, so a hard exit is a recorded result
120
+ instead of a hung harness.
121
+ - Resident memory is sampled through each render. A leak and a teardown bug both
122
+ present as "it crashes eventually"; the RSS curve separates them.
123
+ - `--restart-on-crash` continues past the first death, because failing on
124
+ iteration 2 and failing on iteration 40 are different findings.
125
+
126
+ The synthetic source is JPEG 2000 in MXF OP1A at 12-bit 1998×1080, written by
127
+ ffmpeg. Resolve identifies it as `JPEG 2000 / MXF OP1A / 12 bit` — the decode
128
+ path a DCP video track file takes. `--source` swaps in real media later with
129
+ nothing else changing.
130
+
131
+ One thing to get right about the target: **this build offers ProRes in no MXF
132
+ flavour.** `GetRenderCodecs` returns ProRes only under `mov`; neither
133
+ `mxf_op1a` (76 codecs) nor `mxf` (47) lists it. So "MXF to ProRes" means
134
+ J2K-MXF in, ProRes-QuickTime out, and a harness that tried to render ProRes into
135
+ MXF would fail for that reason rather than for any reason worth reporting.
136
+
137
+ ```bash
138
+ python scripts/render_stress.py --mode headless --iterations 20 --duration 60 \
139
+ --restart-on-crash --report headless.json
140
+ python scripts/render_stress.py --mode gui --iterations 20 --duration 60 \
141
+ --restart-on-crash --report gui.json
142
+ ```
143
+
144
+ Confirming a crash on the real job still needs the real job — same source and
145
+ preset, both modes, `ResolveDebug.txt` kept from each. The harness narrows where
146
+ to look; it does not replace that.
147
+
148
+ ### Stress results: both modes, 10 ProRes renders each
149
+
150
+ 10 renders per leg, 1045-frame JPEG 2000 timeline, ProRes 422 HQ QuickTime,
151
+ ~553 MB per render, identical parameters.
152
+
153
+ | | GUI | headless |
154
+ | --- | --- | --- |
155
+ | completed | 10 / 10 | **10 / 10** |
156
+ | process deaths | 0 | **0** |
157
+ | crash blocks captured | 0 | **0** |
158
+ | bytes per render | 553,050,011 | 553,050,011 (identical) |
159
+ | render time mean | 22.6s | **21.5s** |
160
+ | render time first → last | 21.8 → 22.7s | 21.4 → 22.4s |
161
+ | peak RSS first → last | 2096 → 2104 MB | **2002 → 2010 MB** |
162
+
163
+ **Headless is clean, and marginally better than the GUI on both axes** — about
164
+ 1.1s faster per render and ~95 MB lighter throughout. No leak in either mode:
165
+ 8 MB of RSS drift across ten renders, which is noise. Output is byte-identical.
166
+
167
+ An earlier version of this page reported the headless leg as "hung during setup,
168
+ 0/10, reproducibly". That was true at the time and the cause was **not** render
169
+ instability: the harness was hitting `SaveProject()` on the never-saved
170
+ `Untitled Project`, which blocks the whole application (see the modal trap
171
+ above). Once that was guarded, the leg ran straight through on the first
172
+ attempt. A harness bug had been sitting in the "headless is unstable" column.
173
+
174
+ **What this still does not settle:** the field report of crashes on completion of
175
+ writes during DCP→ProRes delivery. This is 10 renders of ~45 seconds each from
176
+ synthetic J2K. It exercises the decode and encode paths and the teardown at the
177
+ end of every write, ten times, without a single failure — but it is not hours of
178
+ sustained encode, and it is not the operator's footage, codec settings or
179
+ durations. Treat it as "the write-completion path is not trivially broken
180
+ headless", not as a clearance for long-form delivery.
181
+
182
+ ### Differencing the two modes exhaustively
183
+
184
+ `scripts/mode_matrix.py` is the second-generation differential, and it exists
185
+ because the first one **could not have found the `SaveProject` hang**. That
186
+ harness called everything in one process with no timeouts, so a call that never
187
+ returns simply ended the run; its clean report of "zero mode-dependent failures"
188
+ was partly a description of what it was able to survive.
189
+
190
+ Three changes make exhaustive comparison possible:
191
+
192
+ - **`hang` is a first-class verdict.** A call that returns `False` is
193
+ survivable; one that never returns takes the session with it and degrades the
194
+ instance afterwards. Collapsing them into "failed" hides the worse bug.
195
+ - **Supervisor and worker are separate processes.** The worker streams one
196
+ flushed JSONL result per probe, in catalogue order. When it dies, the
197
+ supervisor knows the first unrecorded probe was in flight, records it as
198
+ `hang`, restarts Resolve, waits for health, and resumes after it. A hang costs
199
+ one probe and one restart, not the run. Nothing in-process can do this: a
200
+ Python signal handler cannot preempt a thread parked in fusionscript's
201
+ `pthread_cond_wait`, which is why the alarm-based approach only ever worked by
202
+ luck.
203
+ - **State is part of each probe.** `SaveProject` is harmless on a named project
204
+ and fatal on the never-saved default, so probes declare which project must be
205
+ current. A catalogue that only ever ran against a well-formed scratch project
206
+ is exactly what reported no differences the first time.
207
+
208
+ Phase breadcrumbs distinguish a hang in a probe's own call from a hang while
209
+ reaching its state — a real distinction, since setting up the "untitled" state
210
+ means closing a project, and closing a *modified* project raises the GUI save
211
+ dialog. Without the breadcrumb that dialog was attributed to the probe that
212
+ never got to run.
213
+
214
+ ### Stressing it without production assets
215
+
216
+ `scripts/render_stress.py` closes most of that gap before real footage is
217
+ available. It renders repeatedly in one session and watches for the things that
218
+ distinguish *causes* rather than merely recording that something died:
219
+
220
+ - `crash_archive.txt` is bracketed per iteration, so a stack is attached to the
221
+ render that produced it. Resolve appends every crash to one growing file with
222
+ no timestamps, so without the bracket "did this render crash" is guesswork.
223
+ - The Resolve PID is watched from outside, so a hard exit is a recorded result
224
+ instead of a hung harness.
225
+ - Resident memory is sampled through each render. A leak and a teardown bug both
226
+ present as "it crashes eventually"; the RSS curve separates them.
227
+ - `--restart-on-crash` continues past the first death, because failing on
228
+ iteration 2 and failing on iteration 40 are different findings.
229
+
230
+ The synthetic source is JPEG 2000 in MXF OP1A at 12-bit 1998×1080, written by
231
+ ffmpeg. Resolve identifies it as `JPEG 2000 / MXF OP1A / 12 bit` — the decode
232
+ path a DCP video track file takes. `--source` swaps in real media later with
233
+ nothing else changing.
234
+
235
+ One thing to get right about the target: **this build offers ProRes in no MXF
236
+ flavour.** `GetRenderCodecs` returns ProRes only under `mov`; neither
237
+ `mxf_op1a` (76 codecs) nor `mxf` (47) lists it. So "MXF to ProRes" means
238
+ J2K-MXF in, ProRes-QuickTime out, and a harness that tried to render ProRes into
239
+ MXF would fail for that reason rather than for any reason worth reporting.
240
+
241
+ ```bash
242
+ python scripts/render_stress.py --mode headless --iterations 20 --duration 60 \
243
+ --restart-on-crash --report headless.json
244
+ python scripts/render_stress.py --mode gui --iterations 20 --duration 60 \
245
+ --restart-on-crash --report gui.json
246
+ ```
247
+
248
+ Confirming a crash on the real job still needs the real job — same source and
249
+ preset, both modes, `ResolveDebug.txt` kept from each. The harness narrows where
250
+ to look; it does not replace that.
251
+
252
+ ### First stress results (2026-08-01, Studio 19.1.3.7)
253
+
254
+ 10 renders per leg, 1045-frame J2K timeline, ProRes 422 HQ QuickTime, ~553 MB
255
+ per render.
256
+
257
+ | | GUI | headless |
258
+ | --- | --- | --- |
259
+ | iterations completed | 10 / 10 | **0 — hung during setup** |
260
+ | process deaths | 0 | 0 (hung, did not die) |
261
+ | crash blocks captured | 0 | 0 |
262
+ | render time first → last | 21.8s → 22.7s | — |
263
+ | peak RSS first → last | 2096 MB → 2104 MB | — |
264
+
265
+ **The GUI leg is clean.** Ten consecutive ProRes renders, no crash, no leak (8 MB
266
+ of RSS drift across the run), no degradation in render time. Nothing in the
267
+ write-completion path failed with a UI present.
268
+
269
+ **The headless leg never rendered.** Twice, the harness connected, reported
270
+ `headless=True`, and then blocked forever in `Fusion::RemoteApp::WaitPkt` during
271
+ project setup — before generating any media. In both cases a second client
272
+ found `GetCurrentDatabase()` returning `None` on the same instance.
273
+
274
+ What is and is not established:
275
+
276
+ - **Established:** a headless instance can reach a state where scripted project
277
+ work hangs indefinitely, and that state coincides with a null current
278
+ database. Reproduced twice.
279
+ - **Established:** it is not inherent to headless. An earlier headless session
280
+ the same day created a project, rendered, exported five interchange formats,
281
+ and tore down cleanly.
282
+ - **Not established:** the trigger. Both hangs happened when headless was
283
+ started *after* other heavy Resolve activity — once after force-killing a
284
+ wedged instance, once immediately after a 10-render GUI leg and a clean
285
+ `Quit()`. That is a suggestive pattern, not a cause.
286
+ - **Not established:** whether the null database is cause or symptom. The
287
+ reading was taken from a second client while the first was blocked, so it may
288
+ reflect the block rather than explain it.
289
+
290
+ So the honest state of the stability question is: **the GUI path is measured
291
+ clean at this scale, and the headless path has a real, reproducible failure that
292
+ is not yet pinned down.** That is consistent with the field report of headless
293
+ instability, and it is a reason to keep a GUI fallback for deliveries until the
294
+ trigger is understood.
295
+
296
+ ## Verified
297
+
298
+ ### Launch and teardown
299
+
300
+ ```bash
301
+ "/Applications/DaVinci Resolve/DaVinci Resolve.app/Contents/MacOS/Resolve" -nogui
302
+ ```
303
+
304
+ Launch the **binary inside the bundle**, not `open -a`. `open` hands the
305
+ argument list to LaunchServices, which starts the app normally and drops
306
+ `-nogui` on the floor — you get a GUI and no error.
307
+
308
+ | Observation | Measured |
309
+ | --- | --- |
310
+ | Launch → `scriptapp("Resolve")` answers | under 3s (warm cache) |
311
+ | Launch → `ProjectManager` answers `GetProjectListInCurrentFolder()` | same instant, no extra wait |
312
+ | `resolve.Quit()` → process gone | 2.1s |
313
+ | Resident memory, idle, no project | ~1.9 GB |
314
+ | Scripting listener | TCP 15000, bound to all interfaces, same as GUI |
315
+
316
+ `Quit()` is the correct teardown. It discards the open project without
317
+ prompting, which is exactly what you want from a batch process and exactly what
318
+ `CloseProject` on an unsaved project must never be asked to do.
319
+
320
+ `scriptapp()` answering is *not* proof the ProjectManager is ready. It was ready
321
+ immediately here, but earlier sessions on this codebase recorded it lagging, so
322
+ `scripts/headless_differential.py` polls
323
+ `GetProjectListInCurrentFolder()` rather than trusting the handle. Keep that
324
+ pattern; the failure it prevents looks like "headless cannot see any projects".
325
+
326
+ ### Everything that works headless
327
+
328
+ Verified working, identically to the GUI:
329
+
330
+ - **All seven pages.** `OpenPage()` returns True for media, cut, edit, fusion,
331
+ color, fairlight and deliver, and `GetCurrentPage()` reads each one back.
332
+ There is no window, but the page concept is real and page-gated calls behave.
333
+ - **Render to disk.** Format and codec enumeration, `SetRenderSettings`,
334
+ `AddRenderJob`, `StartRendering(isInteractiveMode=False)`, status polling to
335
+ `Complete`, and a non-empty file on disk.
336
+ - **Interchange export.** AAF, EDL, FCP7 XML, DRT and OTIO all written non-empty.
337
+ - **`Project.ExportCurrentFrameAsStill()`.** Returns True and writes a real PNG.
338
+ This contradicts a 2026-07-03 working note that recorded it as a headless-only
339
+ failure — see "Corrections" below.
340
+ - **`Timeline.GrabStill()`.** Returns a live still object.
341
+ - **Fusion.** `resolve.Fusion()`, `Fusion.GetCurrentComp()`, and per-clip
342
+ `AddFusionComp` / `GetFusionCompNameList` / comp counting.
343
+ - **Colour.** Node graph access, `GetNumNodes`, `GetNodeLabel`, `GetLUT`,
344
+ `SetNodeEnabled`, and colour group create / assign / delete.
345
+ - **Layout presets.** `SaveLayoutPreset`, `ExportLayoutPreset` (170 KB written),
346
+ `LoadLayoutPreset`, `DeleteLayoutPreset` — all True, with no UI to lay out.
347
+ Included in the probe precisely because it was expected to fail; it does not.
348
+ - **Ordinary editorial.** Markers, tracks, track names, clip properties, media
349
+ pool folders, project settings, playhead timecode.
350
+
351
+ ### The modal trap — and why headless is WORSE, not immune
352
+
353
+ **An earlier version of this page claimed headless was immune to modal dialogs
354
+ and therefore the safer mode. That was wrong, and it was the most consequential
355
+ error in the study.** Headless is not immune. It cannot *display* a dialog, but
356
+ it still tries to raise one — and the call then never returns.
357
+
358
+ `ProjectManager.SaveProject()` on the default, never-saved project named
359
+ `Untitled Project`:
360
+
361
+ | | GUI | headless |
362
+ | --- | --- | --- |
363
+ | result | returns `False` | **blocks forever** |
364
+ | observed | immediate | no return after 45s; client parked in `Fusion::RemoteApp::WaitPkt` |
365
+ | recoverable by | a human clicking the dialog | nothing — the client must be killed |
366
+
367
+ Measured on a cold headless boot with the database verified attached
368
+ immediately before the call, so nothing else was wrong with the instance.
369
+
370
+ The project has no location to save to and the API offers no `SaveProjectAs`, so
371
+ Resolve wants a Save-As dialog. With a UI it can ask and move on. Without one it
372
+ waits for an answer that can never arrive.
373
+
374
+ **This inverts the usual advice.** The standard defence against losing work on a
375
+ project switch — "call `SaveProject()` first" — is the exact call that hangs a
376
+ headless session, and it hangs on precisely the project that made you want to
377
+ call it. In the GUI the same situation costs a human one click; headless it
378
+ costs the run.
379
+
380
+ It also degrades: after a hung `SaveProject` was interrupted, the same instance
381
+ began returning `None` from `SaveProject` instantly, and in other runs reached
382
+ the no-database wedge described below. So the first hang is not the end of it.
383
+
384
+ **The rule:** never call `SaveProject()` in a headless session without first
385
+ checking that the current project is not the never-saved `Untitled Project`.
386
+ Check `GetCurrentProject().GetName()` and skip the save — there is nothing to
387
+ save, and the call cannot succeed.
388
+
389
+ ```python
390
+ project = pm.GetCurrentProject()
391
+ if project is not None and project.GetName() != "Untitled Project":
392
+ pm.SaveProject() # safe: it has a location
393
+ # else: nothing to save, and calling it headless blocks forever
394
+ ```
395
+
396
+ Name-matching is a blunt guard — a real project deliberately named
397
+ "Untitled Project" would be skipped — but skipping a save on a named project
398
+ costs nothing here, while calling it on the default one costs the session.
399
+
400
+ ### Detecting headless
401
+
402
+ **There is no API tell.** `GetCurrentPage()` answers with a real page,
403
+ `GetProductName()` and `GetVersionString()` are unchanged, and every UI-shaped
404
+ call the probe tried returns the same value as the GUI. Anything that inspects
405
+ the `resolve` handle to decide whether it has a UI is guessing.
406
+
407
+ The only reliable signal is the process's own argument vector — `-nogui` in the
408
+ command line of the running `…/Contents/MacOS/Resolve`. That is what
409
+ `src/utils/resolve_runtime.py` reads.
410
+
411
+ ### The wedge: a Resolve with no database attached
412
+
413
+ The most dangerous failure state found so far, because every cheap health check
414
+ passes through it.
415
+
416
+ After an unclean shutdown, a headless instance came back up **attached to no
417
+ project database**. It accepted scripting connections. `GetProductName`,
418
+ `GetVersionString`, `GetCurrentPage` and `GetCurrentProject` all answered
419
+ normally. And yet:
420
+
421
+ - `CreateProject` and `LoadProject` returned `False` — indefinitely, not flakily.
422
+ - `SaveProject` returned `None`.
423
+ - Some calls never returned at all: the client sat in
424
+ `Fusion::RemoteApp::WaitPkt` forever while a *different* client got instant
425
+ answers from the same instance.
426
+ - It did not recover. Only a restart cleared it.
427
+
428
+ **The tell is `ProjectManager.GetCurrentDatabase()` returning `None`.** Healthy
429
+ looks like `{"DbType": "Disk", "DbName": "Local Database"}`.
430
+ `resolve_control(action="runtime_mode")` reports this as `database_attached`,
431
+ and flips its guidance to `WEDGED:` when it is false. Check it before doing work
432
+ in an unattended session — a liveness check that only proves the API answers
433
+ will sail straight past this.
434
+
435
+ Two related observations from the same episode:
436
+
437
+ - **Slow startup is not a hang.** That instance took **2 minutes 5 seconds** to
438
+ start its script server, against under 3 seconds for a clean boot, sitting at
439
+ 0% CPU in an idle Qt run loop the whole time. A 90-second readiness timeout
440
+ gave up 15 seconds early. Allow at least 5 minutes before declaring a launch
441
+ failed, and never start a second instance on a timeout — that converts a slow
442
+ start into singleton contention.
443
+ - **A stale database entry costs startup time every boot.** This machine's
444
+ `~/Library/Preferences/Blackmagic Design/DaVinci Resolve/dblist.conf` lists a
445
+ disk database whose path no longer exists, and each start logs
446
+ `Cannot connect to <name> database: path not found`. Worth pruning on any
447
+ machine that runs Resolve unattended.
448
+
449
+ - **Killing a scripting client mid-render can wedge the instance it was driving.**
450
+ Prefer letting a render finish, or `Quit()`, over killing the client.
451
+
452
+ ### Singleton rules
453
+
454
+ One Resolve per machine. The application, a render node, and any scripted launch
455
+ all contend for the same singleton, and losing that contention has produced
456
+ crash loops in this codebase's history rather than a clean error — every
457
+ project open/close collision died in Fairlight page teardown.
458
+
459
+ Before starting a headless instance, confirm nothing else has one:
460
+
461
+ ```bash
462
+ ps -Ao pid=,command= | grep "MacOS/Resolve" | grep -v grep
463
+ ```
464
+
465
+ An empty result is the only safe state. If a render node manages Resolve on this
466
+ machine, turn it off first.
467
+
468
+ ## The expanded sweep: 105 probes, still no real mode differences
469
+
470
+ The catalogue grew from 92 to 105 probes, adding the surfaces most likely to
471
+ behave differently without a UI: long-running and Studio-AI operations
472
+ (`TranscribeAudio`, `AutoSyncAudio`, `CreateSubtitlesFromAudio`, proxy
473
+ link/unlink), editorial constructs (compound clips, Fusion clips, take
474
+ selectors, nested timelines, clip linking) and project-level delivery
475
+ (render presets, burn-in presets, `ArchiveProject`, multi-job render queues).
476
+
477
+ Result across 2 GUI × 2 headless runs: **86 parity, 14 both-failed, 0 flaky, and
478
+ 5 apparent mode differences that all evaporated under isolation.**
479
+
480
+ ### Not API-reachable at all
481
+
482
+ Worth stating because their absence is easy to mistake for a headless problem:
483
+ **retimes/speed changes, transitions, and multicam have no scripting API** in
484
+ this reference. Nothing can round-trip what cannot be created. `CreateCompoundClip`
485
+ and `CreateFusionClip` are documented but raise `TypeError` on 19.1.3 — present
486
+ in the docs, absent from this build.
487
+
488
+ ### Five false findings, and the rule they produced
489
+
490
+ A full sweep reported `editorial.nested_timeline` and `editorial.set_clips_linked`
491
+ as `headless_degraded`, and take selectors, multi-job queues and render presets
492
+ as `divergent`. Re-run with `--only` against a fresh fixture, **all five are
493
+ identical in both modes.**
494
+
495
+ The cause is ordering. Probes run in catalogue order against one shared fixture,
496
+ so a probe near the end carries ~90 destructive probes' worth of accumulated
497
+ state, and the two modes drift apart for reasons unrelated to the mode. The
498
+ report now carries that warning inline, and the rule is simple:
499
+
500
+ **An actionable finding from a full sweep is a hypothesis. Re-run it with
501
+ `--only` in both modes before believing it.**
502
+
503
+ That is the third distinct class of false positive this study has produced —
504
+ after single-run flake (fixed by requiring repetition) and mixing catalogue
505
+ versions (fixed by fingerprinting). Every one of them looked like a real
506
+ headless bug first.
507
+
508
+ ## Rendered pixels are identical
509
+
510
+ The strongest result in this study, and the one everything else was missing.
511
+ Every earlier comparison was of structure and return values; none of them could
512
+ have caught a headless render that completes, produces a file of the right
513
+ length in the right codec, and contains the wrong image.
514
+
515
+ Five timelines rendered in each mode to ProRes 422 HQ, compared **after
516
+ decoding** with ffmpeg `framemd5` so container metadata cannot influence the
517
+ result:
518
+
519
+ | effect | changed the picture? | pixels across modes |
520
+ | --- | --- | --- |
521
+ | plain | — | **identical** |
522
+ | transform (zoom/pan/rotate) | yes | **identical** |
523
+ | CDL grade | yes | **identical** |
524
+ | 3D LUT | yes | **identical** |
525
+ | Fusion comp (MediaIn→Blur→MediaOut) | yes | **identical** |
526
+
527
+ The middle column is the control that makes the result mean anything. If a grade
528
+ did not change the picture relative to plain, then "the modes agree" would only
529
+ mean both rendered the same ungraded image. Every treatment is proven to have
530
+ changed the output first — which took two corrections to achieve:
531
+
532
+ - **`SetLUT` silently refused an absolute path.** It returns False and reads
533
+ back an empty string for a LUT Resolve has not "discovered". The first run
534
+ recorded `applied: False` and rendered pixels identical to plain, which would
535
+ have been written up as "LUT survives headless" — a pass produced by a grade
536
+ that never happened. Staging the LUT into the master LUT directory
537
+ (`src/utils/lut_paths.py:ensure_lut_in_master`) and referencing it relatively
538
+ fixes it.
539
+ - **An empty Fusion comp cannot change anything**, so adding one and finding the
540
+ pixels unchanged proves nothing about whether Fusion reaches the renderer.
541
+
542
+ ## Fusion comps DO render from the API — if the graph is rooted
543
+
544
+ Correcting this repository's own note, which said a comp created on a media clip
545
+ through the API is "never applied at render".
546
+
547
+ A comp wired `MediaIn → Blur → MediaOut`, created entirely through the API on an
548
+ ordinary media clip, **renders in both modes**: PSNR between the plain and
549
+ Fusion renders of the same timeline is 22.7 dB (identical would be infinite) and
550
+ the file shrinks 22.5 MB → 14.8 MB, exactly as a blur should.
551
+
552
+ The distinguishing factor appears to be whether MediaOut descends from MediaIn.
553
+ The original observation — a MediaOut fed only by a Text+ with no MediaIn path,
554
+ rendering the untouched clip — was not re-measured and stands for that
555
+ configuration.
556
+
557
+ A related failure worth knowing: a first attempt wired only `MediaOut → Blur`,
558
+ leaving the Blur with no source. That did not silently bypass the comp — the
559
+ **render job came back `Failed` with an 887-byte file.** An unrooted graph can
560
+ take the render down.
561
+
562
+ ## Network scripting reaches a headless instance
563
+
564
+ Confirmed rather than assumed: `scriptapp("Resolve", "127.0.0.1")` and the repo's
565
+ `connect_resolve` with `RESOLVE_SCRIPT_HOST` both connect to a `-nogui` instance
566
+ and see its database and all 242 projects. Headless and network scripting
567
+ compose.
568
+
569
+ `fuscript -p Resolve` also **discovers** a headless instance over the network,
570
+ reporting hostname, IP, version and platform with no Python involved — useful
571
+ for a supervisor checking what is running. Its script-execution forms (`-x` and a
572
+ script file, Lua or py3) printed only the interpreter banner and produced no
573
+ output in this build; not pursued further.
574
+
575
+ **The in-app bridge cannot work headless.** The bridge listener only exists once
576
+ someone runs Workspace > Scripts > resolve_bridge *inside* Resolve, and that menu
577
+ requires a UI. Since external scripting is Studio-only, this means **the free
578
+ edition cannot be driven headless at all** — the bridge is its only transport,
579
+ and headless is the one mode where the bridge cannot be started.
580
+
581
+ ## Isolated instances: yes. Parallel instances: no.
582
+
583
+ The `BMD_RESOLVE_*_DIR` variables found in the binary **do work**, and they
584
+ isolate more thoroughly than expected — but they do not defeat the singleton.
585
+
586
+ **Measured, launching a second `-nogui` instance with isolated directories while
587
+ one was already running:** the second process started, created a complete
588
+ private tree — `config/config.dat`, `support/Resolve Project Library`,
589
+ `support/Fairlight`, `support/DolbyVision`, `support/easyDCP` — and then **exited
590
+ silently**, logging nothing beyond two log4cxx lines. The first instance was
591
+ unharmed and kept port 15000. So **parallel headless workers on one machine are
592
+ not possible this way.** That closes the biggest open architectural question on
593
+ this list, in the negative.
594
+
595
+ **Running alone, the same isolated instance works completely** — its own
596
+ database with **zero projects**, entirely separate from the user's 242. That is
597
+ still valuable: a CI or render worker that cannot see, lock, or damage the
598
+ operator's projects.
599
+
600
+ One non-obvious step is required. A fresh config does **not** enable external
601
+ scripting, so an isolated instance starts unreachable — it runs, opens no
602
+ listener, and looks like a hang. The user's own config has
603
+ `System.Scripting.Mode = 1`; a fresh one has no such key. `config.dat` is plain
604
+ ASCII, so seeding it is a one-liner:
605
+
606
+ ```bash
607
+ ISO=/path/to/isolated
608
+ mkdir -p "$ISO"/{config,support,logs,lut}
609
+
610
+ # One boot to generate the default config, then enable external scripting.
611
+ BMD_RESOLVE_CONFIG_DIR="$ISO/config" BMD_RESOLVE_SUPPORT_DIR="$ISO/support" \
612
+ BMD_RESOLVE_LOGS_DIR="$ISO/logs" BMD_RESOLVE_LUT_DIR="$ISO/lut" \
613
+ "/Applications/DaVinci Resolve/DaVinci Resolve.app/Contents/MacOS/Resolve" -nogui &
614
+ sleep 30; pkill -f "MacOS/Resolve"; sleep 8
615
+ echo "System.Scripting.Mode = 1" >> "$ISO/config/config.dat"
616
+
617
+ # Now it is scriptable, with its own empty project library.
618
+ BMD_RESOLVE_CONFIG_DIR="$ISO/config" BMD_RESOLVE_SUPPORT_DIR="$ISO/support" \
619
+ BMD_RESOLVE_LOGS_DIR="$ISO/logs" BMD_RESOLVE_LUT_DIR="$ISO/lut" \
620
+ "/Applications/DaVinci Resolve/DaVinci Resolve.app/Contents/MacOS/Resolve" -nogui &
621
+ ```
622
+
623
+ `BMD_RESOLVE_LOGS_DIR` is the one that did not fully take: the isolated tree got
624
+ `LogArchive/` and `gpudetect.bin` but no main log.
625
+
626
+ ## `-fastmode`: starts, and is not scriptable
627
+
628
+ Exercised in the isolated sandbox. `Resolve -nogui -fastmode` starts a process
629
+ that stays alive at 0% CPU and **never opens the scripting listener** — port
630
+ 15000 is not bound at all, and it was still unreachable after two minutes.
631
+ Whatever it is for, it is not usable for automation. Do not use it.
632
+
633
+ ## Deliberately NOT exercised
634
+
635
+ `-activate` and `-deactivate` are left untested **on purpose**, not by omission.
636
+ The strings beside them ("License activated successfully.", "An activation
637
+ already exists on this machine.", "License deactivated successfully.") say
638
+ plainly that they mutate this machine's Studio licence activation. Running
639
+ `-deactivate` to find out what it does could cost the operator their activation,
640
+ and no finding here is worth that. If licence automation is ever needed, test it
641
+ on a machine whose activation is expendable.
642
+
643
+ ## Present but not exercised
644
+
645
+ Read out of the application binary's argument table, adjacent to `-nogui` in the
646
+ same string block. Their existence is certain; their behaviour is not tested,
647
+ and none should be used in automation without qualifying it first.
648
+
649
+ | Flag | Adjacent strings suggest |
650
+ | --- | --- |
651
+ | `-fastmode` | Unknown. No accompanying diagnostic strings. |
652
+ | `-activate` / `-deactivate` | Licence activation from the command line — "License activated successfully.", "An activation already exists on this machine.", "License deactivated successfully.", and an interactive `Please enter q or Q to exit:` prompt. Plausibly the supported way to licence a render node without a GUI. |
653
+ | `-versionUpdate` | Update check. |
654
+ | `-reportCrash` | Crash reporter entry point. |
655
+ | `-test_suite` / `-test_suite_gui` | Blackmagic-internal test harness. |
656
+ | `-psn` | macOS process serial number, passed by LaunchServices. |
657
+
658
+ Environment variables in the same binary, likewise unexercised, that appear to
659
+ relocate Resolve's per-user directories — the mechanism an isolated CI instance
660
+ would need:
661
+
662
+ `BMD_RESOLVE_CONFIG_DIR`, `BMD_RESOLVE_LICENSE_DIR`, `BMD_RESOLVE_LOGS_DIR`,
663
+ `BMD_RESOLVE_LUT_DIR`, `BMD_RESOLVE_SUPPORT_DIR`.
664
+
665
+ Whether setting these actually lets two Resolve instances coexist is **the open
666
+ question worth answering next**. It would decide whether parallel headless
667
+ render workers are possible at all.
668
+
669
+ ## Environment variables that are verified
670
+
671
+ The scripting ones, required for any external Python to find the API:
672
+
673
+ ```bash
674
+ export RESOLVE_SCRIPT_API="/Library/Application Support/Blackmagic Design/DaVinci Resolve/Developer/Scripting"
675
+ export RESOLVE_SCRIPT_LIB="/Applications/DaVinci Resolve/DaVinci Resolve.app/Contents/Libraries/Fusion/fusionscript.so"
676
+ export PYTHONPATH="$PYTHONPATH:$RESOLVE_SCRIPT_API/Modules/"
677
+ ```
678
+
679
+ On Linux the paths are `/opt/resolve/Developer/Scripting` and
680
+ `/opt/resolve/libs/Fusion/fusionscript.so`, with the binary at
681
+ `/opt/resolve/bin/resolve`; on Windows, `%PROGRAMDATA%\Blackmagic Design\DaVinci
682
+ Resolve\Support\Developer\Scripting` and `C:\Program Files\Blackmagic
683
+ Design\DaVinci Resolve\fusionscript.dll`, with `Resolve.exe` taking `-nogui`
684
+ directly.
685
+
686
+ `RESOLVE_SCRIPT_HOST` and `RESOLVE_SCRIPT_TIMEOUT` select this server's network
687
+ scripting mode; see the network-mode notes in `docs/SKILL.md`. The listener a
688
+ headless instance opens on port 15000 is the same one that mode connects to, so
689
+ headless and network scripting compose.
690
+
691
+ ## Corrections to earlier notes
692
+
693
+ A 2026-07-03 working note recorded Gallery `ExportStills` and
694
+ `export_frame_as_still` as **headless-only failures**. Settling this took three
695
+ attempts, and the sequence is worth keeping because each wrong answer was
696
+ confidently held:
697
+
698
+ 1. The original note: headless-only failure.
699
+ 2. A hand-written test found it failing in the GUI too, and "corrected" the note.
700
+ That test was bad — its `GrabStill` had returned `False`, so `ExportStills`
701
+ was handed nothing to export and its `False` said nothing about anything.
702
+ 3. Two proper sweeps then showed GUI `True/2 files` versus headless `False/0`,
703
+ which looked like a clean confirmation of the original note.
704
+ 4. Four controlled sweeps (2 per mode, same catalogue version) showed it
705
+ returning `False` in **all four**, both modes — with `GrabStill` succeeding
706
+ in all four.
707
+
708
+ **The answer is that it is panel-dependent, not mode-dependent.** The variable
709
+ that changed between the one GUI success and the four failures is whether the
710
+ Gallery panel was visible on the Color page — restored workspace state, which
711
+ the harness does not control. Headless can never satisfy it; a GUI session
712
+ satisfies it only sometimes. So *in practice* it never works headless, but
713
+ "headless" is not the cause and switching to a GUI session is not a fix.
714
+
715
+ `ExportCurrentFrameAsStill`, by contrast, worked in all four runs in both modes.
716
+ Use it.
717
+
718
+ Two methodological lessons, both learned the hard way here:
719
+
720
+ - **A negative result needs its preconditions verified as carefully as a
721
+ positive one.** Step 2 above was a confident correction built on an
722
+ unverified precondition.
723
+ - **One run per mode reports noise as findings.** Before repetition was added,
724
+ `app.keyframe_mode` was reported as `headless_degraded` in one comparison and
725
+ `gui_degraded` in another — it returns `0` and `None` in *both* modes
726
+ depending on the run. The comparison now requires a probe to be internally
727
+ consistent within a mode before it may be called a difference between modes,
728
+ and reports the rest as `flaky`.
729
+
730
+ This page is regenerable — rerun the sweeps rather than trusting it after a
731
+ Resolve upgrade.
732
+
733
+ ## Reproducing
734
+
735
+ ```bash
736
+ # 1. GUI baseline
737
+ python scripts/headless_differential.py record --label gui --out /tmp/gui.json
738
+
739
+ # 2. quit Resolve, then
740
+ "/Applications/DaVinci Resolve/DaVinci Resolve.app/Contents/MacOS/Resolve" -nogui &
741
+ python scripts/headless_differential.py record --label headless --out /tmp/headless.json
742
+
743
+ # 3. difference them
744
+ python scripts/headless_differential.py compare /tmp/gui.json /tmp/headless.json \
745
+ --out docs/reference/headless-capability-matrix.md
746
+ ```
747
+
748
+ `record` refuses to run if the `--label` contradicts the running process's argv,
749
+ which is the mistake that would otherwise silently invert every verdict. It
750
+ creates and deletes its own scratch project and saves the incumbent first, but
751
+ it does switch projects — do not point it at a session with unsaved work.