@koda-sl/baker-cli 0.123.0 → 0.124.0-dev.2ddde71d7

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -397,6 +397,262 @@ baker ads google keywords metrics --customer-id 1234567890 --keywords "running s
397
397
 
398
398
  ---
399
399
 
400
+ ### Google Ads Library (`baker ads google library`)
401
+
402
+ Manage and search the Google Ads Transparency Center. Track competitor advertisers, browse their ad creatives, and discover who's bidding on keywords.
403
+
404
+ **Typical workflow:** `search-advertiser` → `track` → `search-ads`
405
+
406
+ ---
407
+
408
+ ### `baker ads google library search-advertiser "query"`
409
+
410
+ Search for an advertiser on the Google Ads Transparency Center.
411
+
412
+ > **Recommended:** use the domain running the ads (e.g. `example.com`) for more accurate results.
413
+
414
+ ```bash
415
+ baker ads google library search-advertiser "example.com"
416
+ baker ads google library search-advertiser "Nike"
417
+ ```
418
+
419
+ **Response:**
420
+
421
+ ```json
422
+ {
423
+ "ok": true,
424
+ "data": {
425
+ "results": [
426
+ { "advertiserId": "AR12345678901234567", "name": "Nike, Inc.", "region": "US", "format": "TEXT_IMAGE_VIDEO" }
427
+ ]
428
+ }
429
+ }
430
+ ```
431
+
432
+ **Flags:**
433
+
434
+ | Flag | Description |
435
+ |------------|--------------------------------|
436
+ | `--output` | Format: `json` \| `csv` \| `md` |
437
+
438
+ ---
439
+
440
+ ### `baker ads google library track <id> <name>`
441
+
442
+ Track a new Google advertiser and wait for the initial ad sync to complete. Polls every 5 seconds with a 10-minute timeout. Progress is written to stderr.
443
+
444
+ ```bash
445
+ baker ads google library track AR12345678901234567 "Nike, Inc."
446
+ baker ads google library track AR12345678901234567 "Nike, Inc." --json
447
+ ```
448
+
449
+ **Response (with `--json`):**
450
+
451
+ ```json
452
+ {
453
+ "ok": true,
454
+ "data": {
455
+ "advertiserId": "ar_abc123",
456
+ "accountId": "acc_def456",
457
+ "totalAdCount": 342,
458
+ "activeAdCount": 89
459
+ }
460
+ }
461
+ ```
462
+
463
+ **Flags:**
464
+
465
+ | Flag | Description |
466
+ |----------|----------------------|
467
+ | `--json` | Output in JSON format |
468
+
469
+ ---
470
+
471
+ ### `baker ads google library list-advertisers`
472
+
473
+ List all tracked Google advertisers and their accounts.
474
+
475
+ ```bash
476
+ baker ads google library list-advertisers
477
+ baker ads google library list-advertisers --output md
478
+ ```
479
+
480
+ **Flags:**
481
+
482
+ | Flag | Description |
483
+ |------------|--------------------------------|
484
+ | `--output` | Format: `json` \| `csv` \| `md` |
485
+
486
+ ---
487
+
488
+ ### `baker ads google library sync-status <accountId>`
489
+
490
+ Check the sync status and ad counts of a tracked account.
491
+
492
+ ```bash
493
+ baker ads google library sync-status acc_def456
494
+ ```
495
+
496
+ **Response:**
497
+
498
+ ```json
499
+ {
500
+ "ok": true,
501
+ "data": {
502
+ "syncStatus": null,
503
+ "totalAdCount": 342,
504
+ "activeAdCount": 89
505
+ }
506
+ }
507
+ ```
508
+
509
+ `syncStatus` is `null` when idle, `"syncing"` during a sync, or `"error"` if the last sync failed.
510
+
511
+ ---
512
+
513
+ ### `baker ads google library search-ads <accountId>`
514
+
515
+ Search and filter ads for a tracked account. Supports pagination.
516
+
517
+ ```bash
518
+ baker ads google library search-ads acc_def456
519
+ baker ads google library search-ads acc_def456 --search "summer sale" --isActive --mediaType image
520
+ baker ads google library search-ads acc_def456 --sort newest --limit 50
521
+ baker ads google library search-ads acc_def456 --cursor "eyJwYWdl..."
522
+ ```
523
+
524
+ **Response:**
525
+
526
+ ```json
527
+ {
528
+ "ok": true,
529
+ "data": {
530
+ "page": [
531
+ {
532
+ "_id": "abc123",
533
+ "platform": "google",
534
+ "externalId": "CR_1234567890",
535
+ "isActive": true,
536
+ "mediaType": "image",
537
+ "headline": "Summer Sale — 50% Off Everything",
538
+ "description": "Shop our biggest sale of the year. Free shipping on all orders.",
539
+ "destinationUrl": "https://example.com/summer-sale",
540
+ "bodyText": "Summer Sale — 50% Off Everything",
541
+ "pageName": "Example Store",
542
+ "impressionsMin": 100000,
543
+ "impressionsMax": 200000,
544
+ "startDate": "2025-06-01",
545
+ "endDate": "2025-06-30",
546
+ "firstSeenAt": 1717200000000,
547
+ "lastSeenAt": 1719792000000,
548
+ "publisherPlatforms": ["GOOGLE_ADS"],
549
+ "regionCodes": ["US", "GB"],
550
+ "variations": [
551
+ {
552
+ "headline": "Summer Sale — 50% Off",
553
+ "description": "Shop our biggest sale of the year.",
554
+ "destinationUrl": "https://example.com/summer-sale",
555
+ "imageUrl": "https://...",
556
+ "visibleUrl": "example.com"
557
+ }
558
+ ],
559
+ "regions": [
560
+ { "code": "US", "name": "United States" }
561
+ ],
562
+ "analysisStatus": "completed",
563
+ "aiAnalysis": {
564
+ "aiSummary": "Promotional display ad for a seasonal sale with urgency-driven CTA",
565
+ "hookAngle": "Discount/Price",
566
+ "offerType": "Percentage Discount",
567
+ "ctaStrategy": "Shop Now",
568
+ "funnelStage": "Bottom",
569
+ "targetAudience": "Price-sensitive shoppers",
570
+ "adFormat": "responsive_display",
571
+ "tags": ["sale", "discount", "ecommerce"],
572
+ "trustSignals": ["Free shipping"],
573
+ "keyMessages": ["50% off", "Free shipping"],
574
+ "competitiveAngle": "Price leadership",
575
+ "dominantColors": ["#FF5733", "#FFFFFF"],
576
+ "analyzedAt": 1719792000000
577
+ }
578
+ }
579
+ ],
580
+ "continueCursor": "eyJwYWdl...",
581
+ "isDone": false
582
+ }
583
+ }
584
+ ```
585
+
586
+ **Key response fields:**
587
+
588
+ | Field | Description |
589
+ |-------|-------------|
590
+ | `headline`, `description` | Top-level ad copy (first variation) |
591
+ | `variations[]` | All ad variations with copy, images, videos, and URLs |
592
+ | `regions[]` | Geographic targeting regions |
593
+ | `impressionsMin/Max` | Estimated impression range (Google Ads Transparency data) |
594
+ | `publisherPlatforms` | Where the ad ran (GOOGLE_ADS, YOUTUBE, etc.) |
595
+ | `analysisStatus` | AI analysis state: `pending`, `processing`, `completed`, `failed` |
596
+ | `aiAnalysis` | AI-generated creative analysis (only present when `analysisStatus` is `completed`) |
597
+ | `aiAnalysis.aiSummary` | One-line AI summary of the ad |
598
+ | `aiAnalysis.hookAngle` | Creative hook (Discount, Fear, Social Proof, etc.) |
599
+ | `aiAnalysis.funnelStage` | Funnel position: Top, Middle, Bottom |
600
+ | `aiAnalysis.tags` | AI-generated tags for filtering |
601
+
602
+ **Flags:**
603
+
604
+ | Flag | Description |
605
+ |---------------|------------------------------------------------|
606
+ | `--search` | Search term for ad text |
607
+ | `--isActive` | Filter by active ads only |
608
+ | `--mediaType` | Filter by media type: `image`, `video`, `text` |
609
+ | `--sort` | Sort: `newest` or `oldest` |
610
+ | `--limit` | Max results per page (default 20, max 100) |
611
+ | `--cursor` | Pagination cursor from previous response |
612
+ | `--output` | Format: `json` \| `csv` \| `md` |
613
+
614
+ ---
615
+
616
+ ### `baker ads google library sync <accountId>`
617
+
618
+ Trigger an immediate re-sync for a tracked account. Polls every 5 seconds until complete (10-minute timeout). Progress is written to stderr.
619
+
620
+ ```bash
621
+ baker ads google library sync acc_def456
622
+ ```
623
+
624
+ **Response:**
625
+
626
+ ```json
627
+ {
628
+ "ok": true,
629
+ "data": {
630
+ "totalAdCount": 350,
631
+ "activeAdCount": 92
632
+ }
633
+ }
634
+ ```
635
+
636
+ ---
637
+
638
+ ### `baker ads google library search-competitors "keyword"`
639
+
640
+ Search for competitors running Google ads for a keyword. Uses DataForSEO (same data as `baker research advertisers`).
641
+
642
+ ```bash
643
+ baker ads google library search-competitors "running shoes"
644
+ baker ads google library search-competitors "crm software" --location uk
645
+ ```
646
+
647
+ **Flags:**
648
+
649
+ | Flag | Description |
650
+ |--------------|----------------------------|
651
+ | `--location` | Location name or code |
652
+ | `--json` | Output in JSON format |
653
+
654
+ ---
655
+
400
656
  ### Staged writes (`baker ads google budgets|campaigns|...`)
401
657
 
402
658
  Write commands **never touch the Google Ads API at stage time**. Each command stages a create/update/pause/resume/remove op against the current chat's draft (`BAKER_CHAT_ID`); the dashboard shows it as a pending "Google Ads" change, and the whole draft applies as one atomic `GoogleAdsService.Mutate` when the chat is published. Feature-flagged per company (`companies.googleAdsWriteEnabled`) — off by default = a fully simulated publish with zero real API calls.
@@ -1876,6 +2132,8 @@ Auto-ingests with prefetched bytes (no double-fetch). The screenshot bytes thems
1876
2132
 
1877
2133
  ScreenshotOne caches captures for 30 days, so re-shotting the same URL within that window returns the cached capture.
1878
2134
 
2135
+ If the target page is unreachable or returns a non-2xx status (e.g. a 404 path or a login wall), the command fails with an actionable `VALIDATION_ERROR` naming the page and the status it returned, plus a `fix` hint — verify the URL and retry, or continue without the screenshot. It no longer surfaces a generic "Internal server error".
2136
+
1879
2137
  **Flags:**
1880
2138
 
1881
2139
  | Flag | Description |
@@ -2161,17 +2419,13 @@ baker testimonials tags
2161
2419
 
2162
2420
  ### Winning Ads (`baker winning-ads`)
2163
2421
 
2164
- Search the **ad-dna** corpus of scored "winning" competitor ads for reference creatives to reproduce (e.g. with `baker canvas`), and manage the brands your library tracks (`follow` / `following` / `unfollow`). Each result carries a presigned media URL (~1h TTL), the ad's DNA summary, and scores. The CLI authenticates with the normal `BAKER_API_KEY`; the Baker backend proxies the request to the ad-dna service with a server-held token — no extra credential in the sandbox.
2165
-
2166
- > The corpus has **Meta + LinkedIn** connectors, so `--platform` inputs are limited to `meta,linkedin`. (Older result rows may still carry a legacy platform string.)
2422
+ Search the **ad-dna** corpus of scored "winning" competitor ads for reference creatives to reproduce (e.g. with `baker canvas`). Each result carries a presigned media URL (~1h TTL), the ad's DNA summary, and scores. The CLI authenticates with the normal `BAKER_API_KEY`; the Baker backend proxies the request to the ad-dna service with a server-held token — no extra credential in the sandbox.
2167
2423
 
2168
2424
  > Backend env: the Convex deployment must have `AD_DNA_API_TOKEN` set (`npx convex env set AD_DNA_API_TOKEN …`). `AD_DNA_API_URL` is optional and defaults to `https://ads.withbaker.com`.
2169
2425
 
2170
- > Replaces the old `baker ads google library` tree, which has been removed. Competitor-by-keyword discovery still lives at `baker research advertisers`.
2171
-
2172
2426
  ### `baker winning-ads search <query>`
2173
2427
 
2174
- Semantic search (dense recall + BM25 + rerank) → `POST /api/ad-library/winners/search`. The CLI projects each result to a **lean, decision-focused** shape so the agent's context stays small — default fields: `advertiser`, `advertiser_id`, `platform`, `format`, `relevance`, `winner_score`, `summary` (what the ad is about), `media_url`; plus top-level `pool_size`, `match_confidence` (`high|medium|low`), and `below_floor_count` (matches dropped under the relevance floor). `--full` adds DNA detail (`angle`, `target_persona`, `hook_archetype`, `awareness_stage`, `industry`) + longevity (`days_active`, `reach`, `active`, `winner_category`, `media_kind`). `--output json` (default) returns the lean objects; `--output md` prints a table.
2428
+ Semantic search (dense recall + BM25 + rerank). The CLI projects each result to a **lean, decision-focused** shape so the agent's context stays small — default fields: `advertiser`, `advertiser_id`, `platform`, `format`, `relevance`, `winner_score`, `summary` (what the ad is about), `media_url`; plus top-level `pool_size` and `match_confidence`. `--full` adds DNA detail (`angle`, `target_persona`, `hook_archetype`, `awareness_stage`, `industry`) + longevity (`days_active`, `reach`, `active`, `winner_category`, `media_kind`). `--output json` (default) returns the lean objects; `--output md` prints a table.
2175
2429
 
2176
2430
  > `media_url` is the creative itself: for `static` it's the image, for `video` it's the video file. ad-dna stores **no separate poster** for videos, so a video result has only the video URL.
2177
2431
 
@@ -2193,7 +2447,7 @@ baker winning-ads search --ref-ad-id a_12345 --first-seen-after 2026-01-01T00:00
2193
2447
  | `--limit <n>` | Max results 1–100 (**default 10** — shortlist size) |
2194
2448
  | `--max-per-advertiser <n>` | Cap results per advertiser 1–50 (default 3) |
2195
2449
  | `--min-relevance <0-1>` | Relevance floor; trims weak matches |
2196
- | `--platform <list>` | `meta,linkedin` — pass a single value to search **only** that platform |
2450
+ | `--platform <list>` | One or many of `meta,tiktok,linkedin,google_search,google_display,youtube,reddit,x,pinterest,snapchat` — pass a single value to search **only** that platform |
2197
2451
  | `--format <list>` | `video,static,carousel` |
2198
2452
  | `--winner-category <list>` | `winner,scaled_winner,evergreen,rising,untested,dud,…` (default: all) |
2199
2453
  | `--awareness <list>` | `unaware,problem_aware,solution_aware,product_aware,most_aware` |
@@ -2208,66 +2462,13 @@ Reading the scores: **`relevance`** (0–1) = match of the creative to your quer
2208
2462
 
2209
2463
  ### `baker winning-ads advertisers <brand>`
2210
2464
 
2211
- List corpus brands by name or domain → `GET /api/ad-library/advertisers`. Use it to find **your own** advertiser (to `--exclude-advertiser`) or a **competitor** (to `--advertiser-id` / `winners`). Lean default fields: `advertiser_id`, `label`, `platform_count`, `family_count`, `active_ad_count`, `total_ad_count`; `--full` adds `scraped_name`, `image_url`, `platforms`, `total_reach`, `last_synced_at`. Supports `--platform meta|linkedin`, `--limit`, `--offset`.
2465
+ Resolve a brand name → `advertiser_id`(s) in the corpus. Use it to find **your own** advertiser (to `--exclude-advertiser`) or a **competitor** (to `--advertiser-id`). Returns `advertiser_id`, `label`, `active_ads`, `total_ads`.
2212
2466
 
2213
2467
  ```bash
2214
2468
  baker winning-ads advertisers "Acme" --output md # find our own advertiser id
2215
2469
  baker winning-ads advertisers "Deel" --platform meta --output md
2216
2470
  ```
2217
2471
 
2218
- ### `baker winning-ads follow "<domain | profile URL | brand>" --platform meta|linkedin`
2219
-
2220
- Add **all** of a brand's ads to your library — every platform (Meta + LinkedIn) and every country → `POST /api/ad-library/follow`. `--platform` is **required**, but it only tells us how to read your `input` (Facebook vs LinkedIn URL); it does **not** limit what we track. A bare domain is best: we resolve both the Meta page and the LinkedIn company from it and track both. Discovery is never region-scoped. The result `status` is one of:
2221
-
2222
- - `following` — the brand is already in the corpus; you're now subscribed (no wait).
2223
- - `added` — a new brand was queued for ingestion; its ads appear as discovery completes (a `hints[]` note flags this).
2224
- - `ambiguous` — the input mapped to multiple brands; pick one from `candidates` and re-run with a more specific domain/URL.
2225
-
2226
- ```bash
2227
- baker winning-ads follow "deel.com" --platform meta
2228
- baker winning-ads follow "https://www.linkedin.com/company/acme" --platform linkedin --label "Acme (competitor)"
2229
- ```
2230
-
2231
- ### `baker winning-ads following`
2232
-
2233
- List the brands you follow → `GET /api/ad-library/following`. Each row shows `status` (`ready` vs `adding…`) plus cached counts (`active_ad_count`, `total_ad_count`, `family_count`) and discovery progress (`adding_discovered`, `adding_enqueued`). `--full` adds `image_url` + `platforms`.
2234
-
2235
- ```bash
2236
- baker winning-ads following --output md
2237
- ```
2238
-
2239
- ### `baker winning-ads winners <advertiser>`
2240
-
2241
- Top winning ads for one advertiser id → `GET /api/ad-library/advertiser-winners`. Same lean winner cards as `search` (add `--full` for DNA + longevity). Supports `--top N` and `--platform meta|linkedin`.
2242
-
2243
- ```bash
2244
- baker winning-ads winners adv_123 --top 15 --output md
2245
- ```
2246
-
2247
- ### `baker winning-ads unfollow <advertiser>`
2248
-
2249
- Stop following a brand by advertiser id → `POST /api/ad-library/unfollow`. Returns `{ removed }`.
2250
-
2251
- ```bash
2252
- baker winning-ads unfollow adv_123
2253
- ```
2254
-
2255
- ### `baker winning-ads brief`
2256
-
2257
- Generate a creative brief grounded in strategically-similar winners → `POST /api/ad-library/brief`. Optionally describe the target creative with `--dna` (a JSON object), steer with `--notes`, and cap references with `--k`. Returns `brief_markdown` + `reference_ad_ids`.
2258
-
2259
- ```bash
2260
- baker winning-ads brief --dna '{"angle":"cost savings","awareness_stage":"solution_aware"}' --notes "B2B, LinkedIn video" --k 8
2261
- ```
2262
-
2263
- ### `baker winning-ads patterns --winners <adIds> --duds <adIds>`
2264
-
2265
- Mine what separates two cohorts of ads → `POST /api/ad-library/patterns`. Pass a comma-list of winning ad ids (`--winners`, cohort A) and weaker/dud ad ids (`--duds`, cohort B), plus optional `--top-n`. Returns each discriminating DNA field with the cohort it `favors` (`winners`/`duds`), a `score`, and the top values on each side.
2266
-
2267
- ```bash
2268
- baker winning-ads patterns --winners a_1,a_2,a_3 --duds a_9,a_8 --top-n 10 --output md
2269
- ```
2270
-
2271
2472
  ---
2272
2473
 
2273
2474
  ### Scheduled Actions (`baker scheduled-actions`)
@@ -2502,6 +2703,29 @@ baker canvas run my-canvas.json
2502
2703
  # single node — that old serial workaround is obsolete.
2503
2704
  baker canvas run my-canvas.json --parallel 8
2504
2705
 
2706
+ # 2c. Runs persist across sandboxes/sessions by default: node results sync to a
2707
+ # company-scoped remote cache (small JSON pointers; bytes stay in R2), so a FRESH
2708
+ # sandbox re-runs an already-computed canvas at zero credits — assets rehydrate
2709
+ # from R2, sha-verified. Every run also posts a durable history record (per-node
2710
+ # outputs, credits, cached/fresh) that powers the dashboard's Creatives
2711
+ # generations timeline. Opt out with --remote-cache off (env
2712
+ # BAKER_CANVAS_REMOTE_CACHE=off) and --no-record. With --remote-cache off,
2713
+ # assets are not uploaded, so a recorded run keeps its stats but has no
2714
+ # browsable outputs — pass --no-record too if you want nothing persisted.
2715
+ baker canvas run my-canvas.json --remote-cache off --no-record
2716
+
2717
+ # 2d. Regenerate a node whose prompt is fine (force a fresh roll). The engine is
2718
+ # content-addressed: re-running an UNCHANGED node returns the identical cached
2719
+ # render, never a new draw. To re-roll a node without editing its prompt (a
2720
+ # color drifted, a face came out wrong), force it fresh two ways — NEVER
2721
+ # restructure the canvas (repointing output / deleting nodes) to trick the cache:
2722
+ # • One-shot flag — forces the named nodes + everything downstream fresh this
2723
+ # run, leaving every other node cached (unknown ids fail loudly before billing):
2724
+ baker canvas run my-canvas.json --regenerate gen_4x5,gen_9x16
2725
+ # • Persistent — add/bump a node's `regenerate` field in the canvas JSON
2726
+ # (e.g. "regenerate": 2) and re-run; the fresh render is reproducible in any
2727
+ # later session. Bump it again (3, 4, …) for each additional draw.
2728
+
2505
2729
  # 3. Inspect a finished run (per-node timing, file list, optional video thumbs)
2506
2730
  baker canvas inspect <run_id>
2507
2731
 
@@ -2601,6 +2825,20 @@ A literal string value. Use for prompts, descriptions, copy.
2601
2825
 
2602
2826
  ---
2603
2827
 
2828
+ ##### `collect`
2829
+
2830
+ Gather images from multiple upstream nodes into one ordered array — the standard terminal for **multi-variant canvases** whose final output is several images (e.g. one artwork composited into N scene photos, one `image_generate` branch per scene). Point the canvas `output` at this node and every collected image becomes a final (`final#0`…`final#n-1`, capped at 10 in run records).
2831
+
2832
+ Pure ref passthrough: zero credits, no byte downloads, and each final carries a **label** into the run record — its producer node id (`$ref:gen_billboard_03.images#0` → `gen_billboard_03`) or an explicit `params.labels[i]` — so variants stay identifiable in the dashboard grid and per-output selection. Name branches after their scene/variant to get meaningful labels for free.
2833
+
2834
+ **Inputs:** `images` (`ImageRef[]`, min 1) — wire a literal array of refs, one per branch: `["$ref:gen_a.images#0", "$ref:gen_b.images#0", …]`.
2835
+
2836
+ **Params:** `labels` (string[], optional) — one unique label per wired image; overrides the producer-id default.
2837
+
2838
+ **Outputs:** `images` → `image[]`, same order as wired.
2839
+
2840
+ ---
2841
+
2604
2842
  ##### `ffmpeg`
2605
2843
 
2606
2844
  Local ffmpeg passthrough. Write the argv you'd type, declare outputs, the engine stages inputs and ingests results. See [Local CLI nodes](#local-cli-nodes) for the placeholder safety contract.
@@ -2746,7 +2984,7 @@ Pick a `source` discriminator and declare the kind you expect. See [Ingestion](#
2746
2984
 
2747
2985
  **Outputs:** `asset` → `<params.expect>` / content-determined (URL strategy table) or extension-inferred (path).
2748
2986
 
2749
- **Path-source notes:** the canvas is **not portable** to another machine without the file. Cache key folds the file's `mtime:size`, so editing the file invalidates the cache automatically. Supported extensions: `png`, `jpg`/`jpeg`, `webp`, `gif`, `avif`, `svg`, `mp4`, `webm`, `mov`, `m4v`, `mp3`, `wav`, `m4a`, `ogg`, `flac`, `json`, `txt`, `md`, `markdown`, `html`/`htm`, `csv`, `ttf`, `otf`, `woff`, `woff2`. Unknown extensions fall back to magic-byte sniffing for common image formats (and an SVG content sniff), else `kind_mismatch`. **SVG (`expect: "image"`) is rasterized to a transparent PNG on ingest** — brand logos are usually SVG, and image-generation models can't read SVG markup, so it's upscaled (longest edge near 2048px) with transparency preserved and the resulting asset carries `metadata.rasterized_from: "svg"`. **Video (`expect: "video"`) duration is probed from the file's ISO-BMFF (`mp4`/`mov`/`m4v`) header** and stamped as the canonical `duration_ms` (and `metadata.duration_ms`); other containers (e.g. `webm`) leave it unset. Downstream `video_deconstruct` uses this declared duration to size its ingest-poll timeout and preflight — without it those fall back to worst-case budgets and a single deconstruct step can hit the action time limit.
2987
+ **Path-source notes:** the canvas is **not portable** to another machine without the file. Cache key folds the file's `mtime:size`, so editing the file invalidates the cache automatically. Supported extensions: `png`, `jpg`/`jpeg`, `webp`, `gif`, `avif`, `svg`, `mp4`, `webm`, `mov`, `m4v`, `mp3`, `wav`, `m4a`, `ogg`, `flac`, `json`, `txt`, `md`, `markdown`, `html`/`htm`, `csv`, `ttf`, `otf`, `woff`, `woff2`. Unknown extensions fall back to magic-byte sniffing for common image formats (and an SVG content sniff), else `kind_mismatch`. **Any `expect: "image"` in a format image-generation models can't read (SVG, AVIF, HEIC, TIFF, BMP) is normalized to PNG on ingest** — model-safe rasters (`jpeg`/`png`/`gif`/`webp`) pass through untouched, everything else is transcoded so a reference can never 400 a generation. This applies to **both `source: "path"` and `source: "url"`** (URL images are fetched and normalized locally, since the backend can't run the rasterizer). SVG gets density-aware upscaling (longest edge near 2048px, transparency preserved). The normalized asset carries `metadata.rasterized_from` set to the source format (e.g. `"svg"`, `"avif"`). **Video (`expect: "video"`) duration is probed from the file's ISO-BMFF (`mp4`/`mov`/`m4v`) header** and stamped as the canonical `duration_ms` (and `metadata.duration_ms`); other containers (e.g. `webm`) leave it unset. Downstream `video_deconstruct` uses this declared duration to size its ingest-poll timeout and preflight — without it those fall back to worst-case budgets and a single deconstruct step can hit the action time limit.
2750
2988
 
2751
2989
  **Cost:** 0 engine credits for direct fetch + yt-dlp + local file. Handinger charges per scrape.
2752
2990
 
@@ -3196,9 +3434,9 @@ Accepted ref-image MIMEs vary by model — see per-model sections below.
3196
3434
 
3197
3435
  ###### Model: `bytedance/seedance-2.0`
3198
3436
 
3199
- Production-quality ad-creative model. Routed via **fal.ai** (not OpenRouter) because OpenRouter's Seedance passthrough rejects photorealistic human reference frames via ByteDance's "real person" safety filter.
3437
+ Production-quality ad-creative model. Routed via **Replicate** (`bytedance/seedance-2.0`). NOTE: ByteDance's upstream "real person" likeness filter still blocks photorealistic human reference frames on **any** reseller — the escape is a synthetic/AI presenter face or routing real faces to Veo, not the provider.
3200
3438
 
3201
- Ref-image MIMEs: `image/png`, `image/jpeg`, `image/webp` (via fal.ai).
3439
+ Ref-image MIMEs: `image/png`, `image/jpeg`, `image/webp`.
3202
3440
 
3203
3441
  | Name | Type | Required | Notes |
3204
3442
  |---|---|---|---|
@@ -3252,7 +3490,7 @@ Ref-image MIMEs: `image/png`, `image/jpeg`, `image/webp`, `image/gif` (via OpenR
3252
3490
  >
3253
3491
  > A scaffolded canvas carries this table inline at `metadata.todo.model_constraints`.
3254
3492
 
3255
- > **Content-policy blocks are deterministic, not flaky.** fal.ai/Seedance rejects any first/last frame that reads as a real-person likeness (even an AI-generated face) — surfaced as `content_policy_blocked` (HTTP 422, **non-retryable**), even when fal's proxy chain masks it as a 5xx. Retrying **never** succeeds and wastes credits. Fix the cause: switch the clip to the other curated model (Veo routes around fal's filter), or make the source frame less photorealistic.
3493
+ > **Content-policy blocks are deterministic, not flaky.** ByteDance's Seedance rejects any first/last frame that reads as a real-person likeness (even an AI-generated face) — surfaced as `content_policy_blocked` (HTTP 422, **non-retryable**). This is ByteDance's upstream filter, so it fires on **any** reseller (Replicate or otherwise) — a provider swap does **not** route around it. Retrying **never** succeeds and wastes credits. Fix the cause: use a **synthetic/AI-generated** (non-identifiable) presenter face, or route a real face to **Veo** (`person_generation: allow_adult`), or make the source frame less photorealistic.
3256
3494
 
3257
3495
  ---
3258
3496
 
@@ -3303,7 +3541,7 @@ None.
3303
3541
 
3304
3542
  ##### `video_lipsync`
3305
3543
 
3306
- Lip-sync a video to an audio track via VEED (fal.ai).
3544
+ Lip-sync a video to an audio track via Sync Labs `sync/lipsync-2` (Replicate).
3307
3545
 
3308
3546
  **Inputs**
3309
3547
 
@@ -3330,7 +3568,7 @@ Lip-sync a video to an audio track via VEED (fal.ai).
3330
3568
 
3331
3569
  ##### `video_background_remove`
3332
3570
 
3333
- Strip a video's background → alpha WebM/H264. Powered by fal.ai VEED.
3571
+ Strip a video's background → transparent alpha WebM (VP9) or MOV (ProRes 4444). Powered by `sprited/birefnet-video` (Replicate).
3334
3572
 
3335
3573
  **Inputs**
3336
3574
 
@@ -3506,7 +3744,7 @@ Place and mix several audio clips onto one timeline — a music bed plus timed v
3506
3744
 
3507
3745
  ##### `image_background_remove`
3508
3746
 
3509
- Strip background → transparent PNG (or mask). Powered by fal.ai BiRefNet v2.
3747
+ Strip background → transparent PNG. Powered by `men1scus/birefnet` (Replicate). (Note: mask-only output is no longer produced.)
3510
3748
 
3511
3749
  **Inputs**
3512
3750
 
@@ -3748,25 +3986,30 @@ Validate, then execute the graph. Blocks until done. Logs one line per node. Ret
3748
3986
  | `--run-id <id>` | auto ULID | Override the generated run id. |
3749
3987
  | `--cache-policy <policy>` | `read_write` | `read_write`, `bypass`, or `read_only`. |
3750
3988
  | `--concurrency <n>` | `5` (or `BAKER_CANVAS_CONCURRENCY`) | Max nodes executing at once within a layer. |
3989
+ | `--remote-cache <on\|off>` | `on` (or `BAKER_CANVAS_REMOTE_CACHE`) | Company-scoped remote cache + durable asset persistence. |
3990
+ | `--max-credits <n>` | uncapped (or `BAKER_CANVAS_MAX_CREDITS`) | Credit ceiling: aborts before billing when the estimate exceeds it, and at layer boundaries once actual spend does. Completed nodes stay cached, so retrying with a higher cap loses nothing. |
3991
+ | `--no-record` | records | Skip posting the durable run-history record (and its live progress). |
3992
+
3993
+ **Run history streams live.** The run posts its plan (every node + its dependency edges) the moment validation passes, then re-posts a progress snapshot as each node starts and settles — the dashboard's creative workflow graph shows nodes flipping pending → running → done in real time, with each node's outputs attached as they land. A failed run keeps its per-node trail (what completed, what died). All best-effort: an unreachable backend never changes the run's outcome.
3751
3994
 
3752
3995
  **Failures don't abandon sibling work.** Nodes in a layer run under the concurrency cap and every one **settles** — a failed clip no longer kills its in-flight siblings, whose finished results still land in the content-addressed cache. One failure re-throws as-is; several are reported together (each failed node named). Re-running `baker canvas run` resumes from the cache and re-executes **only** the failed nodes and their descendants — never hand-orchestrate per-node renders. Long `video_generate` clips execute as **backend jobs** (the CLI polls; a CDN/proxy timeout can no longer kill a generation mid-flight).
3753
3996
 
3754
3997
  #### `baker canvas scaffold-video <video> [flags]`
3755
3998
 
3756
- Turn a reference video into a **runnable, self-validated reproduction canvas** in one command — the video counterpart of `scaffold-static-ad`. It runs **billed passes** up front:
3999
+ Turn a reference video into a **runnable, self-validated reproduction canvas** in one command — the video counterpart of `scaffold-static-ad`. The `<video>` positional is a **local path OR an http(s) URL** (a `baker winning-ads` `media_url`, a library URL, any reel link) — a URL is downloaded for you (no manual `curl` first), so pass `--slug`/`--out` with it to give the canvas a home. It runs **billed passes** up front:
3757
4000
 
3758
- 1. **`video_deconstruct`** (`~google/gemini-pro-latest`, full mode) — reverse-engineers the video into a scene-by-scene blueprint + word-level transcript, written next to the canvas as **`prompt.json`**. Each scene's `start_frame_prompt`/`end_frame_prompt` are inlined into the frame nodes (see below); `prompt.json` then rides along as the shared **global style reference** (palette, cast cohesion) and as provenance.
4001
+ 1. **`video_deconstruct`** (`~google/gemini-pro-latest`, full mode) — reverse-engineers the video into a scene-by-scene blueprint + word-level transcript, written next to the canvas as **`prompt.json`** (the human-editable source of truth). Each scene's `start_frame_prompt`/`end_frame_prompt` are inlined into the frame nodes (see below); the shared **global style reference** every frame reads via `target_blueprint` is a **slim projection** written alongside as **`prompt.style.json`** (`global` cast/palette/brand + `reference_elements` only, no per-scene array). A 33-scene blueprint is ~200 KB — inlining it into every one of a dozen frame prompts was pure waste and let a frame blend in another scene's content; the slim is ~5 KB. Edit `prompt.json` and re-scaffold to refresh the style projection.
3759
4002
  2. **recurring-element selection** (`~google/gemini-flash-latest`) — picks only the **recurring, identity-critical** elements (each `global.cast` person, a recurring animal, a showcased product, the brand logo) and the scene indices each appears in. One real reference image grounds each element across **every** frame it appears in, so the same actor stays consistent the whole video. This selection runs as a **second pass over a slimmed blueprint** (cast/branding + each scene's frame prompts only) — a long ad's full blueprint can exceed the engine's inline-prompt limit, so the heavy per-scene detail (dialogue, overlays, transcript) the selector never reads is dropped before the prompt.
3760
4003
 
3761
4004
  Before the deconstruct it runs a **local shot-cut pass** on the source file with **[PySceneDetect](https://www.scenedetect.com)** (`scenedetect` CLI, `detect-content` — the battle-tested HSV content detector, installed in the canvas sandbox) and passes the cut timestamps as `video_deconstruct`'s `shot_cuts`. The deconstruct snaps its scene boundaries onto those real cuts and **splits any scene that spans one**, so a scene's frames can never straddle a hard cut (the failure where a scene's start frame was the couch and its end frame the b-roll). Two knobs tuned for fast social ads: the content **threshold defaults to 18** (PySceneDetect's own default of 27 misses soft reframes) and the **minimum scene length is dropped to 0.25s** (its default ~0.6s merges away rapid montage flashes) — so super-fast cuts survive and become cheap still-holds downstream. The threshold is **adaptive**: if the first pass looks like a continuous shot shredded into many close micro-cuts (a talking-head selfie's natural motion), it re-runs at PySceneDetect's own default of 27 and **merges the two passes** — the high-threshold set is the base, and the low pass's *isolated* extras (real soft blur-morph transitions that vanish at 27) are added back while clustered extras (motion shred) stay dropped. Pinning **`--shot-threshold N`** disables the re-check (lower = more cuts). The backend snap window is likewise **adaptive** (up to 1s onto an unambiguous nearest cut, shrinking around dense cut pairs so a boundary never jumps past the wrong cut), any scene spanning an interior cut is split, and the residual-sliver coalesce is **cut-aware**: a drift sliver folds backward across its non-cut edge and never re-merges across a real cut. If `scenedetect` is unavailable it warns loudly and degrades to LLM-only boundaries.
3762
4005
 
3763
4006
  A shot longer than the video model's per-clip ceiling (Seedance's 15s, passed as `video_deconstruct`'s `max_clip_s`) is split into equal **continuation sub-scenes** that share their splice boundary exactly — so a long shot is reproduced in **full** (no truncation) and joins seamlessly. Each sub-scene carries `continues_previous`.
3764
4007
 
3765
- It then scaffolds the full pipeline like an **editing timeline**: each clip gets a **static-ad-grade start AND end keyframe** (`image_generate`, each with its **own self-contained `params.prompt`** — edit a frame node to change only that frame; `prompt.json` wired as the **authoritative shared `target_blueprint`**, plus a per-element reference legend). Each keyframe is **fully recast** to the dropped `el_*` reference images. The original extracted frame is kept LAST as a **pure composition anchor** (framing / camera angle / shot size / pose) whenever identity is safely locked — i.e. a frame with no person/animal, OR every cast member present is **sheet-backed** (a multi-view turnaround owns identity, so the anchor can reproduce the source's framing without dictating the face). Since every base element is now sheet-backed by default, cast frames keep their framing anchor too — this is what reproduces the source's composition (a side-profile stays a side-profile, the camera angle holds scene to scene) instead of drifting to a fresh guess. The anchor's legend forbids taking identity/text/palette from it. It is dropped only when a cast member rests on a weak lone-snapshot reference (e.g. a `same_as` second-look slot), where the original frame could re-leak the source actor. Both keyframes feed `video_generate` (`first_frame`+`last_frame`, so Seedance interpolates real in-shot motion; ultra-detailed motion brief; duration snapped to the nearest allowed clip length). Every keyframe grounds **only on its own extracted frame + `el_*` slots** — no reference to any other generated frame — so all images render **in parallel** (no cascade). Source-frame URLs are **deduped** (each ingested once). `--frames reuse` wires the real source frame straight in.
4008
+ It then scaffolds the full pipeline like an **editing timeline**: each clip gets a **static-ad-grade start AND end keyframe** (`image_generate`, each with its **own self-contained `params.prompt`** — edit a frame node to change only that frame; the slim `prompt.style.json` wired as the **shared `target_blueprint`** style reference, plus a per-element reference legend). Each keyframe is **fully recast** to the dropped `el_*` reference images. The original extracted frame is kept LAST as a **pure composition anchor** (framing / camera angle / shot size / pose) whenever identity is safely locked — i.e. a frame with no person/animal, OR every cast member present is **sheet-backed** (a multi-view turnaround owns identity, so the anchor can reproduce the source's framing without dictating the face). Since every base element is now sheet-backed by default, cast frames keep their framing anchor too — this is what reproduces the source's composition (a side-profile stays a side-profile, the camera angle holds scene to scene) instead of drifting to a fresh guess. The anchor's legend forbids taking identity/text/palette from it. It is dropped only when a cast member rests on a weak lone-snapshot reference (e.g. a `same_as` second-look slot), where the original frame could re-leak the source actor. Both keyframes feed `video_generate` (`first_frame`+`last_frame`, so Seedance interpolates real in-shot motion; ultra-detailed motion brief; duration snapped to the nearest allowed clip length). Every keyframe grounds **only on its own extracted frame + `el_*` slots** — no reference to any other generated frame — so all images render **in parallel** (no cascade). Source-frame URLs are **deduped** (each ingested once). `--frames reuse` wires the real source frame straight in.
3766
4009
 
3767
4010
  **Composited scenes (split-screen / picture-in-picture / keyed presenter).** Real ads aren't always one full-frame shot — a frame can be **persistently divided** (b-roll on top, a presenter talking on the bottom) or **layer a presenter** over background footage (boxed in a corner, or green-screen keyed). The deconstruct now reports this per scene as `scene.composition` (`layout: split_screen | pip | keyed_overlay`, with one `region` per stream — each its own clean-plate frame + motion brief, the talking-head region flagged `is_presenter`). The scaffold reproduces a composited scene by building **one clip per region** (`s<i>_r0_*`, `s<i>_r1_*`, …) and compositing them with ffmpeg: a split-screen `vstack`/`hstack` (stack direction read from the region **panels**, so a top/bottom split always stacks vertically), or a picture-in-picture `overlay` of the presenter inset at its corner. A **keyed** presenter is first cut to transparency by `video_background_remove` (`s<i>_key`), then overlaid. The presenter region carries the native lip-synced voice; b-roll/render panels stay silent. To change a layout, edit `composition` in `prompt.json` and re-scaffold, or hand-edit the `s<i>_composite` ffmpeg args. Plain full-frame scenes (the default) are unaffected.
3768
4011
 
3769
- **Typed region kinds & real screen surfaces.** Each composition region now carries a `kind` — `camera` (filmed footage, re-generated), `screen_capture` (app/site/document screen recording), `static_graphic` (designed text/graphic panel), or `generated` (3D/motion graphics) — plus an optional `nested` list for video-in-video (a Loom-style camera bubble inside a screen share). `kind` is authoritative for routing (prose keywords remain the fallback for older blueprints): `screen_capture`/`static_graphic` regions are **never generated by the video model** — the scene renders as a clean background plate (its clip prompt is scrubbed of all screen narration and forbids rendering UI) and the real surface is composited on the overlay layer. The route is decided **once per persistent layout run** (consecutive scenes sharing one composition signature), so a layout that runs unbroken across many scenes can't flip between pipelines on wording differences. A persistent surface seeds **ONE grouped stub** in `video-overlay-composition/index.html` spanning its whole window, with a per-scene **state timeline** — build one continuous screen recording/mockup, not one screenshot per scene. A `screen_capture` region also carries `surface_id`: a source video routinely **splices two unrelated screen recordings** under one persistent layout (a live app-processing capture, then an unrelated pre-made demo note) — the deconstruct assigns a stable id while the SAME recording continues and a new one when the on-screen content genuinely changes, so the run splits into **separate stubs** at the splice instead of asking for one screenshot that can't cover both. `baker canvas validate` additionally warns (`VIDEO_UI_IN_PROMPT`) if any clip prompt still narrates a screen surface, and (`VIDEO_BRANDMARK_IN_PROMPT`) if a generate prompt asks the model to paint a brand logo/wordmark (generation garbles marks; source the real one with `baker images logo` and composite it on the overlay layer).
4012
+ **Typed region kinds & real screen surfaces.** Each composition region now carries a `kind` — `camera` (filmed footage, re-generated), `screen_capture` (app/site/document screen recording), `static_graphic` (designed text/graphic panel), or `generated` (3D/motion graphics) — plus an optional `nested` list for video-in-video (a Loom-style camera bubble inside a screen share). `kind` is authoritative for routing (prose keywords remain the fallback for older blueprints): `screen_capture`/`static_graphic` regions are **never generated by the video model** — the scene renders as a clean background plate (its clip prompt is scrubbed of all screen narration and forbids rendering UI) and the real surface is composited on the overlay layer. The route is decided **once per persistent layout run** (consecutive scenes sharing one composition signature), so a layout that runs unbroken across many scenes can't flip between pipelines on wording differences. A persistent surface seeds **ONE grouped stub** in `video-overlay-composition/index.html` spanning its whole window, with a per-scene **state timeline** — build one continuous screen recording/mockup, not one screenshot per scene. A `screen_capture` region also carries `surface_id`: a source video routinely **splices two unrelated screen recordings** under one persistent layout (a live app-processing capture, then an unrelated pre-made demo note) — the deconstruct assigns a stable id while the SAME recording continues and a new one when the on-screen content genuinely changes, so the run splits into **separate stubs** at the splice instead of asking for one screenshot that can't cover both. Full-frame screen scenes reuse the same `surface_id`: consecutive full-frame UI beats of one screen (e.g. a 3-scene import flow) share **ONE `s<i>_screen_ref` ingest** — the operator supplies that screenshot once instead of dropping the same capture into a dozen identical `[TODO]`s (distinct surfaces stay distinct). `baker canvas validate` additionally warns (`VIDEO_UI_IN_PROMPT`) if any clip prompt still narrates a screen surface, and (`VIDEO_BRANDMARK_IN_PROMPT`) if a generate prompt asks the model to paint a brand logo/wordmark (generation garbles marks; source the real one with `baker images logo` and composite it on the overlay layer).
3770
4013
 
3771
4014
  **Designed graphics are rebuilt, not generated.** A `static_graphic` surface (a newspaper-collage panel, a meme card, a marketing composition) seeds a **GRAPHIC PANEL** stub — rebuild it as brand HTML or drop the design asset; it never gets the "screenshot the live page" instruction (there is no live page). A **full-frame** designed-graphic scene (the deconstruct emits one full-frame `static_graphic` region for meme/collage/motion-graphic beats) routes to a real design plate the same way screens do — no `image_generate`/`video_generate` — and dialogue over an all-graphic scene is voiceover by definition (nobody is on screen to lip-sync). A region typed `generated` whose own prose reads like a UI/designed panel is treated as a surface candidate too (the frame-grounded continuity checker delivers the verdict and corrects the kind), so one mistyped kind can't re-open the Seedance-paints-UI hole. Floating FX elements (hearts, sparkles, badges) ride the overlay layer: their narration is **scrubbed from clip briefs** and a categorical no-decorations directive is added, so the model can't bake a second, uneditable copy under the real composited one.
3772
4015
 
@@ -3778,13 +4021,13 @@ It then scaffolds the full pipeline like an **editing timeline**: each clip gets
3778
4021
 
3779
4022
  **Montage flashes held as stills — unless the picture really moves.** A rapid-cut beat shorter than ~2s with no spoken line is a **flash** — Seedance's shortest clip is 4s, so generating one (then trimming away most of it) burns credits for motion no viewer perceives. The scaffold instead **holds one keyframe as a still** for the scene length (a cheap ffmpeg loop, no billed `video_generate`), same look at a fraction of the cost. The deconstruct now stamps each scene's **`motion_level`** (`static` / `subtle` / `dynamic`): a **dynamic** flash (pouring chocolate, hands working, walking) keeps a **real trimmed clip** — freezing a moving montage turns it into a slideshow — while genuinely static beats (a logo card, a pinned photo, a product still) keep the cheap hold. Talking/ambient beats always keep a real clip (they need motion + native audio). The deconstruct also stamps each dialogue line's **`on_camera`** flag — a voice playing over b-roll, a graphic, or a mere *photo* of the speaker stays voiceover, so the scaffold never lip-syncs a scene with no speaking face (the polaroid close-up failure).
3780
4023
 
3781
- **The phrase model (voice cut at pauses, not at visual cuts).** The voice is grouped into **phrases** runs of continuous speech with no real pause, which may span several visual scenes. A phrase is voiced ONCE (so a sentence the deconstruct split at a visual cut never breaks mid-word): if the speaker is **shown** anywhere in the phrase it's a single Seedance clip (`s<anchor>_clip`, native lip-sync + audio) re-voiced to the brand voice; if the speaker is **never shown** it's one ElevenLabs `tts` read. The picture is then assembled **scene by scene**: a scene that shows the speaker **slices its window** out of the phrase clip (`s<i>_seg`, an ffmpeg `-ss`/`-t` cut — video and audio come from the *same* clip, so lip-sync holds), and a **b-roll cutaway** gets its own silent clip while the phrase's voice plays underneath. "Shown" is decided by the **presenter element's per-scene presence**, not just who's speaking — a scene where a cast member narrates over b-roll (their element absent) is treated as a cutaway, so the talking head never appears where the original cut away. A presenter run longer than the **gateway-safe ~10s clip ceiling splits at a scene boundary** into contiguous takes (each its own clip + convert), so a sliced window never reads past its clip. (Seedance's *API* max is 15s, but the generation gateway frequently times out — **HTTP 524** — before it can deliver a clip longer than ~10s, so the scaffold never asks for one that long; 10s is a Seedance-allowed duration, so the split clip still snaps cleanly.) A b-roll cutaway *inside* a phrase lands at an **approximate** time (Seedance exposes no word timing) — nudge the scene boundary if it's off its beat.
4024
+ **One clip per shot separated at complete breaks.** A video is a sequence of clear **shots** with **complete breaks** (hard cuts) between them, and that is what the scaffold separates by: **two adjacent presenter shots at a hard cut become TWO clips**, never glued into one invented take just because the speech runs continuously across the cut. Each presenter shot is one Seedance clip (`s<anchor>_clip`, native lip-sync + audio) re-voiced to the brand voice. What is **NOT** split: a **voiceover** narration stays ONE ElevenLabs `tts` read across the b-roll it plays over, and a **b-roll cutaway** between two on-camera moments leaves the presenter shot continuous the shown scenes aren't adjacent (the insert sits between them), so the clip covers both on-camera windows (sliced as `s<i>_seg`, an ffmpeg `-ss`/`-t` cut — video+audio from the *same* clip so lip-sync holds) while the cutaway plays its own silent clip over the continuing voice. "Shown" is decided by the **presenter element's per-scene presence**, not just who's speaking — a scene where a cast member narrates over b-roll (their element absent) is a cutaway, so the talking head never appears where the original cut away. A single shot longer than the **gateway-safe ~10s clip ceiling** (Seedance's *API* max is 15s, but the gateway often times out — **HTTP 524** — past ~10s) **splits into contiguous takes joined by a shared boundary frame**; the spine then **seam-dedups** that duplicated frame so the concat doesn't freeze on it (`--seam-dedup head|tail|off`, default `head` = drop the second clip's first frame). Timbre stays consistent across all the separate shot clips because every clip's native audio is re-voiced in **one merged per-speaker pass** (not per clip). A b-roll cutaway *inside* a phrase lands at an **approximate** time (Seedance exposes no word timing) — nudge the scene boundary if it's off its beat.
3782
4025
 
3783
4026
  **A starting point, not a locked render.** The canvas mirrors the reference's structure to give you a faithful scaffold, but `metadata.todo.full_flexibility` makes explicit that the agent has **full editing freedom**: add / delete / reorder / split / merge scenes, re-prompt any frame or motion brief, change a scene's layout (full-frame ↔ composite), or rewrite any line — the content-addressed cache re-bills only what changes, and `baker canvas validate` re-checks timing/lip-sync after any edit.
3784
4027
 
3785
4028
  **Sequenced audio.** Dialogue is a back-and-forth on one absolute timeline, so each **contiguous same-speaker turn** becomes its own `tts` placed at its real `start_s` — turns alternate and never stack (the earlier design concatenated each speaker's whole monologue at their earliest timestamp, so two voices played in parallel for the entire video). Each speaker is locked to one shared `voice_select` voice; a `sound_effect` per SFX and a `music` bed (conditioned on the **ad's own script + emotional arc** so the bed supports the message, styled after the AudD-identified track when available, ducked under the voices, and started at the reference's `music.starts_at_s` rather than always at 0) round out the mix (`audio_timeline`). The final mux normalizes the soundtrack to **−14 LUFS (stereo)** so the output plays loud in every player — the raw mix is quiet mono, which reads as "no sound."
3786
4029
 
3787
- **Native talking heads + one voice per person (no post-hoc lip-sync).** Seedance 2.0 generates lip-synced speech **natively** — a presenter phrase puts the full phrase in the clip's prompt with `generate_audio`, so lips and voice are generated together (no `video_lipsync`/veed). Each presenter phrase's audio is extracted and re-voiced through a **per-phrase** `audio_voice_convert` (ElevenLabs Voice Changer; one per phrase keeps each ≤15s clip under the converter's length cap) to the brand voice timing preserved so the lips stay matched. There is **ONE voice per person**: a single `voice_select` is reused for all that person's phrases, and the deconstruct's `voiceover` label folds into the sole on-camera presenter (so on-camera and off-camera narration are the same voice, not two). A scene with **two speakers both on screen** can't be one clip — both lines become `tts` over a plain scene clip. But a scene with **one on-camera speaker trading lines with an OFF-camera voice** (an interviewer, a heard-but-not-shown assistant) keeps the on-camera speaker **native** (lip-synced) and reads the off-camera line as `tts` — "on screen" is decided by the speaker's element presence, so a heard-but-unshown voice no longer drops the whole scene to a silent clip. Every `tts` node is stamped with the spoken track's **`language_code`** when the blueprint states a language (cast localization note / voiceover persona / voice description), so numbers and units are read in the target tongue instead of ElevenLabs' English default (the "6900 read in English" bug). For **NATIVE (Seedance) lines** — which carry no language tag — the scaffold additionally **spells numerals into target-language words** across every part of the clip prompt Seedance can vocalize (the spoken line, the scene summary/action/motion, the transcript), so a French "6930 ?" becomes "six mille neuf cent trente ?" and is never read as English digits. Spelling covers **every language the blueprint can resolve** (fr, es, en, de, it, pt, nl, pl, ar, ja, ko, hi — via `n2words`); a language outside that set leaves digits (the `tts` path still localizes them via `language_code`).
4030
+ **Native talking heads + one voice per person (no post-hoc lip-sync).** Seedance 2.0 generates lip-synced speech **natively** — a presenter phrase puts the full phrase in the clip's prompt with `generate_audio`, so lips and voice are generated together (no `video_lipsync`/veed). Each presenter phrase's audio is extracted (the spoken window only) and the extracts are merged **per speaker** onto one timeline, then re-voiced through a **single** `audio_voice_convert` (`<voice>_conv`, ElevenLabs Voice Changer) to the brand voice — one STS pass over the whole track instead of a convert node per clip, so it's fewer nodes, fewer calls, and a more consistent brand timbre (composite scenes already share this path); timing is preserved so the lips stay matched. There is **ONE voice per person**: a single `voice_select` is reused for all that person's phrases, and the deconstruct's `voiceover` label folds into the sole on-camera presenter (so on-camera and off-camera narration are the same voice, not two). A scene with **two speakers both on screen** can't be one clip — both lines become `tts` over a plain scene clip. But a scene with **one on-camera speaker trading lines with an OFF-camera voice** (an interviewer, a heard-but-not-shown assistant) keeps the on-camera speaker **native** (lip-synced) and reads the off-camera line as `tts` — "on screen" is decided by the speaker's element presence, so a heard-but-unshown voice no longer drops the whole scene to a silent clip. Every `tts` node is stamped with the spoken track's **`language_code`** when the blueprint states a language (cast localization note / voiceover persona / voice description), so numbers and units are read in the target tongue instead of ElevenLabs' English default (the "6900 read in English" bug). For **NATIVE (Seedance) lines** — which carry no language tag — the scaffold additionally **spells numerals into target-language words** across every part of the clip prompt Seedance can vocalize (the spoken line, the scene summary/action/motion, the transcript), so a French "6930 ?" becomes "six mille neuf cent trente ?" and is never read as English digits. Spelling covers **every language the blueprint can resolve** (fr, es, en, de, it, pt, nl, pl, ar, ja, ko, hi — via `n2words`); a language outside that set leaves digits (the `tts` path still localizes them via `language_code`).
3788
4031
 
3789
4032
  **Same-shot lip-sync caution.** A single held shot can carry only ONE lip-synced clip (voiceover turns must not overlap, and Seedance generates one clip per shot), so when the on-camera speaker has further turns in that shot (a rapid "3000? … 4000?" with an off-camera "Plus" between), the first turn is native and the rest play as `tts` over the same clip — where the mouth no longer matches those words. This is inherent to reproducing sparse same-shot dialogue, not a wiring fault; the scaffold lists the affected scenes/lines in **`metadata.video.lip_sync_caution`** (advisory, never gated) so you can cut away to b-roll over those lines or rely on the burned-in captions that already show them.
3790
4033
 
@@ -3800,7 +4043,9 @@ It then scaffolds the full pipeline like an **editing timeline**: each clip gets
3800
4043
 
3801
4044
  **Re-craft the script — the hook is the #1 decision.** A reproduction is *inspiration* from a proven ad, not a clone: its structure (hook → body → CTA) carries the persuasion, and the hook is *targeting*, so a competitor's hook often does **not** transfer. `metadata.todo.script_recraft` tags each scene with its `narrative_role` (from the deconstruct, else inferred) and carries the original line **flagged** so it is never shipped as-is — and the per-scene `recraft` instruction is **role-aware**: the **hook** scene's entry carries the diagnose → decide (keep/adapt/rebuild) → criteria (statement not question, benefit by ~2s, first frame legible **sound-off** in ~1s, no bait-and-switch) inline and routes to the skill's `references/hook-craft.md`. A dedicated top-level **`metadata.todo.hook`** key foregrounds it as the highest-leverage beat, mapped onto scene-0's artifacts (`s0_start` first frame, scene-0 overlay text, `s0_clip` line, micro-hook, hook-ramp).
3802
4045
 
3803
- The emitted canvas is validated (`validateCanvasDeep`) before it's written, so it always runs. It also carries a **`metadata.video`** timing plan that `baker canvas validate` proves **statically, before any billed render**: no two voiceover turns overlap, the audio length the video length, every single-on-camera-speaker scene is a native talking head (its clip carries `generate_audio` and is wired to an `audio_voice_convert` node), **no re-crafted line physically overruns its clip** (`VIDEO_SPEECH_OVERRUN` est. speech > ~1.6× the clip duration fails validate, since Seedance crams or dies on it), and **every clip agrees on one aspect ratio** (`VIDEO_ASPECT_MISMATCH`). The full editable checklist is embedded as **`metadata.todo`** (with a step-by-step guide in `metadata.description`). stdout returns `{ ok, canvas_path, prompt_path, models, stats, checklist }`.
4046
+ **The inspiration video is preserved.** Like `scaffold-static-ad` keeps its reference image, the video command now auto-writes a **`_definition.md`** (so the creative joins the `creatives` collection) recording the source it was built from: `sourceKind: video`, `sourceAdvertiser` (the brand the deconstruct identified, or `--advertiser`), `platform` (`--platform`, default `meta`), and **`sourceReferenceUrl`** the **durable, content-addressed R2 URL** the deconstruct already uploaded the source to (`prompt.json`'s `source.url`), which the dashboard's Inspiration card plays inline. Unlike the static flow it does **not** commit the video into `references/`: a reference clip can be up to 2 GiB and the video canvas never re-ingests the source at run time (it uses the extracted frame URLs), so a git copy would be pure bloat the durable R2 URL is the reference. The `_definition.md` is preserved on re-scaffold, and the same `sourceReferenceUrl` is synced to the backend so the creative shows "built from this ad."
4047
+
4048
+ The emitted canvas is validated (`validateCanvasDeep`) before it's written, so it always runs. It also carries a **`metadata.video`** timing plan that `baker canvas validate` proves **statically, before any billed render**: no two voiceover turns overlap, the audio length ≈ the video length, every single-on-camera-speaker scene is a native talking head (its clip carries `generate_audio` and is wired to an `audio_voice_convert` node), **no re-crafted line physically overruns its clip** (`VIDEO_SPEECH_OVERRUN` — est. speech > ~1.6× the clip duration fails validate, since Seedance crams or dies on it), and **every clip agrees on one aspect ratio** (`VIDEO_ASPECT_MISMATCH`). When a **photoreal on-camera cast** generates on **Seedance**, the checklist carries a **`content_policy_risk`** note: ByteDance's real-person-likeness filter can reject a photoreal AI face with a **non-retryable 422** (`content_policy_blocked`) that **no prompt reframe clears** — the escapes are regenerating on Veo (`--video-model google/veo-3.1-fast`) or a less-photoreal frame. Surfaced before the billed run so a face-heavy ad isn't discovered broken mid-render. The full editable checklist is embedded as **`metadata.todo`** (with a step-by-step guide in `metadata.description`). stdout returns `{ ok, canvas_path, prompt_path, models, stats, checklist }`.
3804
4049
 
3805
4050
  ```bash
3806
4051
  baker canvas scaffold-video ./reference-ad.mp4 --focus "competitor UGC ad for <brand>"
@@ -3813,8 +4058,10 @@ baker canvas run ./reference-ad.video.canvas.json
3813
4058
  | Flag | Default | Effect |
3814
4059
  |---|---|---|
3815
4060
  | `--out <path>` | `<video-dir>/<name>.video.canvas.json` | Where to write the canvas (composition is copied alongside). |
4061
+ | `--slug <slug>` | — | Creative slug (lowercase kebab): writes the canvas to `src/creatives/<slug>/<slug>.canvas.json` — the repo convention that attaches every run to the creative's dashboard generation history. `--out` wins over `--slug`. |
3816
4062
  | `--frames <mode>` | `generate` | `generate` emits ONE recast keyframe per scene (the original frame is dropped so the dropped `el_*` assets drive identity); `reuse` wires the real extracted first+last frames straight into the clips (faithful, cheaper, no recast). |
3817
4063
  | `--ambient` | off | Give silent **b-roll** scenes native diegetic ambient (Seedance `generate_audio`), mixed deep under the music bed. Talking scenes already carry voice; check levels don't muddy the mix before keeping it. |
4064
+ | `--seam-dedup <mode>` | `head` | How to dedup the boundary frame two clips SHARE when a long shot is split for length (the second clip's first frame IS the first clip's last frame, so a plain concat freezes on it for a frame). `head` drops the second clip's first frame, `tail` drops the first clip's last frame, `off` keeps both. Only touches shared-frame continuation joins — a hard cut between two shots shares no frame. |
3818
4065
  | `--max-scenes <n>` | all source scenes | **Cost lever that reduces fidelity** — caps the deconstruct, MERGING away every scene beyond the cap (fewer cuts, lost beats). Prints a warning when set; omit it to reproduce every scene. |
3819
4066
  | `--language <code>` | auto | Transcript/dialogue language hint (e.g. `fr`, `en`). |
3820
4067
  | `--focus <text>` | — | Known provenance/emphasis to ground the deconstruct. |
@@ -3837,10 +4084,10 @@ The two scaffold passes are billed (the full `video_deconstruct` is the heavy on
3837
4084
  Turn a source/inspiration image into a **runnable, self-validated static-ad canvas** — the static counterpart of `scaffold-video`. Like the video scaffold, this runs **billed Gemini passes** up front:
3838
4085
 
3839
4086
  1. **`image_describe`** (`~google/gemini-pro-latest`) — reverse-engineers the image into a blueprint JSON, written next to the canvas as **`prompt.json`**. This is the editable "prompt": you rewrite it by hand into the ad you want (palette, copy, claims, subjects). It feeds the generator directly — there is **no automatic brand-transform step**. The blueprint also names the **`winning_mechanisms`** — the special sauce that makes the ad a candidate winner, each tagged `kind` (verbal: rhyme/pun/rhythm; visual: unexpected crop, visual gag, juxtaposition, pattern interrupt, before/after; structural: hook order/reveal) with a `device` and `why_it_works` — so your rewrite rebuilds the mechanism that makes the ad win instead of adapting only the surface and losing it.
3840
- 2. **element selection** (`~google/gemini-flash-latest`) — picks the **main, identity-critical** elements (the brand logo, a showcased product, a trust badge) **plus any foreground/hero person or animal** — the emotional focal point — even a generic one, because a free-generated face/muzzle reads as AI and grows artifacts; the emotional hero always gets a real-reference slot. Background extras are dropped. Each is stamped back onto its blueprint entry as a `reference_image` label so the JSON self-documents which slot grounds which subject.
4087
+ 2. **element selection** (`~google/gemini-flash-latest`) — picks the **main, identity-critical** elements (the brand logo, a showcased product, a trust badge) **plus any foreground/hero person or animal** — the emotional focal point — even a generic one, because a free-generated face/muzzle reads as AI and grows artifacts; the emotional hero always gets a real-reference slot. When the advertiser's logo appears in **more than one lockup** (a square/icon **mark** and a horizontal **wordmark**), each is emitted as its **own** element (e.g. `LOGO_MARK`, `LOGO_WORDMARK`) so you drop the right file in each slot instead of stretching one logo to cover both. The describe pass also records the ad's **typography** under a `fonts` block (each typeface's classification, a best-guess family, and its weight/case) so you know exactly what to drop at the brand-font slot. Background extras are dropped. Each element is stamped back onto its blueprint entry as a `reference_image` label so the JSON self-documents which slot grounds which subject.
3841
4088
  3. **global layout** (`~google/gemini-flash-latest`) — produces a structured `layout` block in `prompt.json`: the column/row grid, each region's `x_pct`/`y_pct` bounds, panel splits, background/shape, and every text block's relative size/weight/case/alignment. This is what gives the generator a precise composition to rebuild.
3842
4089
 
3843
- It then scaffolds a canvas that ingests `prompt.json`, wires **one `[TODO]` ingest slot per detected element** (plus an optional brand-font → type-specimen) into `image_generate`, and wires the original image in for composition only. The canvas is validated before it's written. stdout returns `{ ok, canvas_path, prompt_path, models, layout_regions, stats, checklist }` — the **checklist** lists every real asset to drop in.
4090
+ It then scaffolds a canvas that ingests `prompt.json`, wires **one `[TODO]` ingest slot per detected element** (plus an optional brand-font → type-specimen) into `image_generate`, and wires the original image in for composition only. Each **person/animal hero** is additionally fused into a generated **multi-view reference sheet** (`image_reference_sheet`, a turnaround built from the one dropped photo) that the render grounds on instead of the lone flat snapshot — the same identity lock the video scaffold uses, so the face/muzzle stays consistent and artifact-free from a single reference. Pass `--skip-actor-sheets` to ground straight on the dropped photo. The canvas is validated before it's written. stdout returns `{ ok, canvas_path, prompt_path, models, layout_regions, stats, checklist }` — the **checklist** lists every real asset to drop in (and which heroes get a sheet).
3844
4091
 
3845
4092
  ```bash
3846
4093
  baker canvas scaffold-static-ad ./reference-ad.png --context "competitor ad for <brand>, <category>, <market>"
@@ -3854,15 +4101,37 @@ baker canvas run ./static-ad.canvas.json
3854
4101
  |---|---|---|
3855
4102
  | `--context <text>` | — | Known provenance (advertiser, category, market) to ground the describe. |
3856
4103
  | `--out <path>` | `<image-dir>/static-ad.canvas.json` (cwd when `<image>` is a URL) | Where to write the canvas (`prompt.json` is written alongside). |
4104
+ | `--slug <slug>` | — | Creative slug (lowercase kebab): writes the canvas to `src/creatives/<slug>/<slug>.canvas.json` — the repo convention that attaches every run to the creative's dashboard generation history. `--out` wins over `--slug`. With a slug, the reference image is **downloaded into `src/creatives/<slug>/references/` and normalized to a model-safe format** (SVG/AVIF/HEIC/… → PNG), named from the actual bytes (not the URL string) so a presigned/extensionless URL never lands as PNG-bytes-in-`.jpg` — the canvas ingests that committed, portable path instead of the expiring URL. |
3857
4105
  | `--describe-model <id>` | registry default (`~google/gemini-pro-latest`) | Override the `image_describe` model. |
3858
4106
  | `--select-model <id>` | registry default (`~google/gemini-flash-latest`) | Override the element-selection `text_generate` model. |
3859
4107
  | `--layout-model <id>` | registry default (`~google/gemini-flash-latest`) | Override the global-layout `text_generate` model. |
3860
4108
  | `--gen-model <id>` | registry default (`openai/gpt-5.4-image-2`) | Override the `image_generate` model. |
3861
4109
  | `--aspect <ratio>` | inferred from the image, else `9:16` | Force the output aspect ratio. |
3862
4110
  | `--skip-font` | off | Skip the brand-font → type-specimen slot. |
4111
+ | `--skip-actor-sheets` | off | Ground each person/animal on its lone dropped photo instead of a generated multi-view reference sheet. |
3863
4112
 
3864
4113
  Scaffolding runs (and bills) the two vision passes; **running** the result generates a billed image. `baker canvas validate` does not check that the `[TODO]` paths exist — supply the real files before `run`.
3865
4114
 
4115
+ **Resuming an interrupted run.** A long `baker canvas run` (multi-clip video) that is killed mid-render — session end, sandbox pause — leaves a marker under the outputs dir. The next `baker canvas run` of the same canvas automatically **resumes** that run: it reuses the run id so still-running billed jobs re-attach instead of being abandoned and re-billed, and completed nodes come from the cache. Resume also works from a **different workspace or a fresh sandbox**: when no local marker survives, the run history is consulted and an interrupted (or stale) run of the exact same canvas is adopted automatically. Ctrl-C / SIGTERM aborts gracefully — no new nodes dispatch, a resumable snapshot is flushed, and the marker survives. A clean completion (or a handled failure) clears the marker, so a normal re-run starts a fresh generation. Force a new run with `--fresh`, or pin a specific run with `--run-id <id>` (also the escape hatch to adopt a run that is reported as concurrently live). Independent same-layer nodes (e.g. video clips) fan out in parallel up to `--parallel`/`--concurrency` (default 8; env `BAKER_CANVAS_CONCURRENCY`).
4116
+
4117
+ #### `baker canvas rerun <slug> [flags]`
4118
+
4119
+ Re-run a creative's latest recorded canvas **from run history** — no local files needed. Every `canvas run` of a creative uploads a portable snapshot (the canvas JSON plus the local files it ingests: prompt blueprints, composition dirs, reference images) to durable storage and records it on the run. `rerun` restores those files into `src/creatives/<slug>/` (sha-verified; files already matching are left untouched) and then executes the normal run flow — so an interrupted run **resumes** (in-flight jobs re-attach) and a completed one re-renders from the cache at zero credits.
4120
+
4121
+ | Flag | Default | Meaning |
4122
+ | --- | --- | --- |
4123
+ | `--force-remote` | off | Overwrite local files whose content differs from the snapshot (otherwise a conflict aborts with the differing paths). |
4124
+ | `--fresh` | off | Start a new run id instead of resuming an interrupted one. |
4125
+ | `--regenerate <ids>` | — | Same as `canvas run --regenerate`. |
4126
+ | `--concurrency <n>` | — | Same as `canvas run --concurrency`. |
4127
+
4128
+ ```bash
4129
+ baker canvas rerun spring-offer-4x5
4130
+ baker canvas rerun spring-offer-4x5 --force-remote --regenerate gen_4x5
4131
+ ```
4132
+
4133
+ Use it when a creative was built in another conversation (or its sandbox is gone) and you need to continue or re-render it here. Files the snapshot could not include (missing at run time, or oversized) are listed as warnings — supply those locally only if the run actually needs to regenerate the nodes that read them.
4134
+
3866
4135
  #### `baker canvas inspect <run_id> [--thumbnails]`
3867
4136
 
3868
4137
  One-page summary of a completed run: per-node duration + cache status, list of files in the run dir, optional video thumbnails (start/middle/end frames extracted via ffmpeg).
@@ -4325,6 +4594,26 @@ import {
4325
4594
  } from "@koda-sl/baker-cli/engine";
4326
4595
  ```
4327
4596
 
4597
+ ## Creatives
4598
+
4599
+ Publish an approved canvas render as a first-class Baker creative. The image uploads to the Baker image library (tagged `creative`), a creative record is created/updated, and the command prints the creative reference JSON the dashboard renders in chat.
4600
+
4601
+ ```bash
4602
+ baker creatives publish ./canvas/<run_id>/<final>.png --title "Spring Offer 4x5" \
4603
+ --slug spring-offer-4x5 --run-id r_01JXYZ... \
4604
+ --source-reference-url "https://www.facebook.com/ads/library/?id=..."
4605
+ ```
4606
+
4607
+ | Flag | Effect |
4608
+ |---|---|
4609
+ | `--title <text>` | Required. Human title for the creative. |
4610
+ | `--slug <slug>` | Creative slug (`src/creatives/<slug>/`) — attaches the image to that creative's row, marks it `published`. |
4611
+ | `--run-id <r_…>` | Pins the approved generation from the creative's run history as the published one. |
4612
+ | `--source-reference-url <url>` | Original reference ad URL, recorded on the creative. |
4613
+ | `--context <text>` | Optional describe context for the uploaded image asset. |
4614
+
4615
+ Without `--slug` the command behaves as before (one creative record per published image). With `--slug` it upserts the repo-convention row — the same one the dashboard's Creatives tab and the `src/creatives/{slug}/` folder describe — so publish, repo sync, and run history all land on a single record regardless of order.
4616
+
4328
4617
  ## Help & Discovery
4329
4618
 
4330
4619
  Every command supports `--help` for usage info: