blamcode 0.4.0__py3-none-any.whl

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (58) hide show
  1. blamcode/__init__.py +9 -0
  2. blamcode/cli.py +122 -0
  3. blamcode/layer/README.md +43 -0
  4. blamcode/layer/config/agent/build.md +118 -0
  5. blamcode/layer/config/agent/debugger.md +32 -0
  6. blamcode/layer/config/agent/designer.md +31 -0
  7. blamcode/layer/config/agent/vision.md +26 -0
  8. blamcode/layer/config/agent/writer.md +28 -0
  9. blamcode/layer/config/command/ask.md +41 -0
  10. blamcode/layer/config/command/blam.md +24 -0
  11. blamcode/layer/config/command/explore.md +14 -0
  12. blamcode/layer/config/command/fix.md +14 -0
  13. blamcode/layer/config/command/open.md +18 -0
  14. blamcode/layer/config/command/review.md +14 -0
  15. blamcode/layer/config/opencode.json +53 -0
  16. blamcode/layer/config/themes/bangladeshi.json +67 -0
  17. blamcode/layer/config/themes/blamcode.json +217 -0
  18. blamcode/layer/config/tui.json +4 -0
  19. blamcode/layer/install.sh +752 -0
  20. blamcode/layer/scripts/__pycache__/patch-brand.cpython-312.pyc +0 -0
  21. blamcode/layer/scripts/blamcode +518 -0
  22. blamcode/layer/scripts/blamcode-browser +315 -0
  23. blamcode/layer/scripts/blamcode-menu +140 -0
  24. blamcode/layer/scripts/blamcode-uninstall +71 -0
  25. blamcode/layer/scripts/blamcode-vision +542 -0
  26. blamcode/layer/scripts/oc-settings.sh +138 -0
  27. blamcode/layer/scripts/patch-brand.py +267 -0
  28. blamcode/layer/skills/android-app/SKILL.md +53 -0
  29. blamcode/layer/skills/api-integration/SKILL.md +49 -0
  30. blamcode/layer/skills/bash-cli-expert/SKILL.md +48 -0
  31. blamcode/layer/skills/bot-development/SKILL.md +53 -0
  32. blamcode/layer/skills/clean-code-performance/SKILL.md +45 -0
  33. blamcode/layer/skills/database/SKILL.md +56 -0
  34. blamcode/layer/skills/debugging-fixes/SKILL.md +45 -0
  35. blamcode/layer/skills/deploy-hosting/SKILL.md +38 -0
  36. blamcode/layer/skills/docker/SKILL.md +74 -0
  37. blamcode/layer/skills/firebase-supabase/SKILL.md +61 -0
  38. blamcode/layer/skills/git-workflow/SKILL.md +63 -0
  39. blamcode/layer/skills/lets-scroll/SKILL.md +877 -0
  40. blamcode/layer/skills/lets-scroll/references/index-template.html +73 -0
  41. blamcode/layer/skills/lets-scroll/references/knockout.py +89 -0
  42. blamcode/layer/skills/lets-scroll/references/pipeline.md +312 -0
  43. blamcode/layer/skills/lets-scroll/references/prompts.md +194 -0
  44. blamcode/layer/skills/lets-scroll/references/scrub-engine.js +448 -0
  45. blamcode/layer/skills/project-structure/SKILL.md +74 -0
  46. blamcode/layer/skills/python-automation/SKILL.md +52 -0
  47. blamcode/layer/skills/react-next-best-practices/SKILL.md +54 -0
  48. blamcode/layer/skills/security-review/SKILL.md +48 -0
  49. blamcode/layer/skills/seo-basics/SKILL.md +44 -0
  50. blamcode/layer/skills/testing/SKILL.md +58 -0
  51. blamcode/layer/skills/ui-ux-responsive/SKILL.md +53 -0
  52. blamcode/layer/skills/website-builder/SKILL.md +47 -0
  53. blamcode-0.4.0.dist-info/METADATA +62 -0
  54. blamcode-0.4.0.dist-info/RECORD +58 -0
  55. blamcode-0.4.0.dist-info/WHEEL +5 -0
  56. blamcode-0.4.0.dist-info/entry_points.txt +2 -0
  57. blamcode-0.4.0.dist-info/licenses/LICENSE +21 -0
  58. blamcode-0.4.0.dist-info/top_level.txt +1 -0
@@ -0,0 +1,877 @@
1
+ ---
2
+ name: lets-scroll
3
+ description: >
4
+ Build an immersive scroll-scrubbed "fly through the world" landing page for any
5
+ industry or brand using Higgsfield. As the visitor scrolls, a pre-rendered camera
6
+ flies from outside each scene into its interior, then flows on to the next scene
7
+ with NO cuts — one continuous connected flight (Emons-style isometric diorama world,
8
+ or any art direction you pick). The skill interviews the user for the topic, the
9
+ story beats/sections, and brand kit, then generates cohesive scenes + seamless camera
10
+ clips with Higgsfield (or the user's own tools, via a prompt + conditioning-frame
11
+ handoff) and wires a portable, framework-agnostic scroll-scrub engine.
12
+ The video chain renders through Monid by default (Seedance 2.0, pay-per-clip
13
+ USD — capability re-checked each build, see Step 4) with Higgsfield credits as
14
+ the fallback biller. Use when the user wants a "3D world" /
15
+ "browse-through-the-industry" hero, a scroll cinematic, a diorama landing, or to
16
+ turn a business into a scrollable world.
17
+ allowed-tools: Bash, Read, Write, Edit, AskUserQuestion, Skill
18
+ ---
19
+
20
+ # lets-scroll
21
+
22
+ Produces a landing page where **scroll drives a camera**: it dives from outside a scene
23
+ into its interior, then flies out and into the next scene, continuously, with no visible
24
+ cuts. The visuals are AI-generated — stills via Higgsfield (or Codex), the video chain
25
+ via **Monid by default** (pay-per-clip Seedance 2.0; Higgsfield credits as fallback) —
26
+ and the page just scrubs pre-rendered video by scroll position. A **manual asset path**
27
+ (Step 1.7) swaps the render calls for a prompt + file handoff: the user generates every
28
+ still and clip in tools of their choice and drops the files into the work folder;
29
+ everything downstream (frame extraction, encode, engine, QA) is identical. This is the same technique behind Apple's scroll-through product
30
+ pages — the camera genuinely moves, scroll only drives time.
31
+
32
+ **What you generate:** N scene stills → N "dive-in" camera clips → N-1 "connector" clips
33
+ that join consecutive scenes seamlessly → a portable scrub engine that plays the whole
34
+ chain as one flight.
35
+
36
+ **The one rule that makes or breaks it:** seams must be *frame-identical*. Read
37
+ [The seamless chain](#step-5--the-seamless-chain-the-critical-part) before generating any
38
+ connector. Getting this wrong is the single most common failure and produces a visible
39
+ "pop" between scenes.
40
+
41
+ Do not assume a frontend framework. The scrub engine in `references/scrub-engine.js` is
42
+ self-contained vanilla JS (it builds its own DOM + injects its own CSS into a container
43
+ you give it), so it drops into plain HTML, Next.js, Vue, a Python-served page, anything.
44
+ The value of this skill is the Higgsfield pipeline, the prompts, and the seam method —
45
+ not the framework.
46
+
47
+ ---
48
+
49
+ ## Step 0 — Bootstrap
50
+
51
+ 1. **Monid CLI — the default video-chain backend.** Check `monid --version`,
52
+ `monid keys list` (active key) and `monid balance` — the chain is billed per
53
+ clip in USD (Step 1.8 has the numbers; a 1080p N=6 chain ≈ $27). If the CLI is
54
+ missing or the balance can't cover the chain, say so and fall back to
55
+ rendering the chain on Higgsfield credits instead — same model, same
56
+ pipeline, different biller (Step 4 → Monid backend).
57
+ 2. **Higgsfield CLI — still required even on the Monid path**: it renders the
58
+ scene stills (`gpt_image_2`) and is the home of the `kling3_0` NSFW fallback
59
+ and the fallback chain. If `higgsfield` is not on `$PATH`, install per the
60
+ `higgsfield-generate` skill. If `higgsfield workspace list` fails auth, ask the user
61
+ to run `higgsfield auth login` (interactive OAuth — you cannot run it) and, if needed,
62
+ `higgsfield workspace set <id>`. Confirm credits cover the stills (~N image
63
+ gens) — plus `(2N-1)` video gens if the chain falls back here.
64
+ 3. **ffmpeg / ffprobe** on `$PATH` (frame extraction + encoding).
65
+ 4. **An image tool** for background knockout if you want floating scenes: PIL
66
+ (`python3 -c "import PIL"`), or `cwebp`/`sips`. Optional — see Step 3.
67
+ 5. **(Optional) Codex CLI** — if `codex` is on `$PATH` (≥ 0.125) and
68
+ `codex login status` reports a ChatGPT login, the scene stills can be generated
69
+ through Codex's built-in `image_gen` (the same gpt-image-2 model) billed to the
70
+ user's ChatGPT subscription instead of Higgsfield credits — offer it at
71
+ Step 1.8, command in Step 2. Absence just removes the option.
72
+ 6. **Manual asset path** (Step 1.7) — if the user will render the assets themselves,
73
+ skip items 1, 2 and 5 entirely (no Monid/Higgsfield/Codex needed); only
74
+ ffmpeg/ffprobe (item 3) and, for the optional knockout, PIL (item 4) matter.
75
+ 7. Caveats: macOS ships **bash 3.2** (no `declare -A`); don't use associative arrays in
76
+ scripts. Higgsfield generations take **3–8 min each** — always run them detached
77
+ (background) and poll, never a foreground blocking call. Reference-by-job-UUID is
78
+ rejected by media flags — pass **local file paths** to `--image/--start-image/--end-image`.
79
+ Video models differ in accepted params (e.g. Kling has no `--resolution`) and in whether
80
+ they support start/end-image conditioning at all — before batching, confirm the chosen
81
+ model's schema with `higgsfield model get <job_type>` and see the Step 4 model table.
82
+
83
+ ---
84
+
85
+ ## Step 1 — Interview the user
86
+
87
+ The **subject is the user's to state — ask it as an open question in plain prose**, never a
88
+ fabricated multiple-choice. A made-up list of industries biases them and reads as you
89
+ deciding their business for them; let them answer in their own words (their real business,
90
+ a client's, or any idea). Reserve structured multiple-choice (`AskUserQuestion` in Claude
91
+ Code; a plain either/or question elsewhere) for the genuinely
92
+ enumerable, lower-stakes choices below — art direction, camera style, and brand-kit
93
+ approach — and even
94
+ there, signal they can go their own way ("Other"). Ask only what you can't sensibly
95
+ default. Cover:
96
+
97
+ 1. **Subject** (ask openly, not multiple-choice) — "What should this world be about? Your
98
+ business, a client's, or any idea — a word or a sentence is fine." Capture the
99
+ industry/product + a one-line pitch (e.g. "a bubble tea company, from leaf to last
100
+ sip"), and a brand name if they have one; otherwise you'll propose one below.
101
+ 2. **Brand kit** — offer three paths, pick one:
102
+ - Import from a URL: `higgsfield marketing-studio brand-kits fetch --url <site> --wait`
103
+ (pulls name, colours, tone). Then read it back with `brand-kits list --json`.
104
+ - The user hands you palette + name + tone directly.
105
+ - You propose a palette + name and let them approve.
106
+ Capture **4–6 named hex values**, a display name, and a tone word or two.
107
+ 3. **Art direction** — default is "soft matte low-poly **clay diorama**, isometric,
108
+ tilt-shift miniature, warm light." Offer alternatives (flat papercraft, glossy toy,
109
+ claymation, neon night). Whatever is chosen becomes the shared **style preamble**
110
+ reused verbatim in every scene prompt (this is what makes the world cohesive).
111
+ 4. **Camera style — ALWAYS ask; it's the film's personality, not a technical
112
+ detail.** Ask by feel (`AskUserQuestion` in Claude Code; a plain question
113
+ elsewhere) and record the answer as `CAMERA`. The options map to the Step 4
114
+ architectures — Step 4 then *implements* the choice, it never re-decides it:
115
+ - **"Fly through the world"** — the camera dives into each scene, pulls up and
116
+ out, and hops across the miniature world to the next; angles change
117
+ constantly, big expressive aerial moves (this is the flagship-demo look).
118
+ → Architecture B. Recommend as the default for diorama/miniature art
119
+ directions.
120
+ - **"One continuous walkthrough"** — a single forward flight that glides
121
+ through each scene straight into the next, never pulling back; expressive
122
+ but always-forward moves per scene (camera grammar table). → Architecture A.
123
+ Recommend as the default for grounded/photoreal art directions.
124
+ - **"Locked isometric glide"** — the camera keeps one fixed angle for the whole
125
+ film, Emons-style; the world slides past/toward it, no rotation, no reveals.
126
+ → Architecture A + the locked-iso clause in every leg prompt (prompts.md).
127
+ State the trade-off in one line each (B reverses direction at seams — charming
128
+ in miniature, jarring in realism; locked-iso is the calmest and cheapest to
129
+ re-roll; walkthrough sits between).
130
+ 5. **The journey (sections) — start with the LEVEL OF ANIMATION; always ask it** as a
131
+ structured choice (`AskUserQuestion` in Claude Code; a plain question elsewhere).
132
+ Three sizes, **6 is the maximum** — record the count as `N`:
133
+ - **2 scenes** — a teaser: opener + hero/CTA. 2 stills, 3 clips (2 dives + 1
134
+ connector). Fastest and cheapest; the right first run for testing the pipeline
135
+ end-to-end.
136
+ - **4 scenes** — a short journey. 4 stills, 7 clips.
137
+ - **6 scenes** — the full film. 6 stills, 11 clips. The most the scroll pacing
138
+ stays comfortable with.
139
+ Then propose that many scenes derived from the subject's own value chain and let
140
+ the user edit. Boba example at N=6: farms → pearl kitchen → flagship shop →
141
+ delivery → community plaza → the hero product. Each section needs: a short subject
142
+ description (what's IN the diorama), an eyebrow, a headline, one line of body, and
143
+ 0–3 tag pills. The last section is usually the hero product + the CTA.
144
+ 6. **Mobile version — ALWAYS ask this; never silently generate both.** Ask as a
145
+ two-option choice (`AskUserQuestion` in Claude Code; a plain question elsewhere):
146
+ *"Want a mobile-optimized version too? The mobile version is a second camera chain
147
+ rendered natively in **9:16 portrait** — composed for phones, not a crop of the
148
+ landscape film — which roughly doubles the Higgsfield credit spend (state the
149
+ estimated number)."*
150
+ Options: "Desktop only" / "Desktop + mobile (native 9:16 — ~2× credits)". The
151
+ credit cost must be stated to the user, not just implied.
152
+ What the answer gates:
153
+ - **Yes** → render the parallel 9:16 portrait chain and ship it as the mobile variants
154
+ (Step 6 / pipeline.md §6b): portrait start canvases → 9:16 dives + connectors
155
+ frame-locked against their own renders → 720-wide `-m.mp4` encodes → `stillMobile`
156
+ portrait posters. Wire `clipMobile`/`connectorsMobile`/`stillMobile` (Step 7); run
157
+ the full mobile QA (Step 8). Budget ~2N-1 extra video gens + NSFW re-rolls.
158
+ **Never ship the centre-crop as the mobile version by default** — if credits can't
159
+ cover the portrait chain, say so and offer the crop encodes (pipeline.md §6) as an
160
+ explicitly-labelled stopgap the user must approve.
161
+ - **No** → skip the mobile encodes and wiring entirely. The engine's phone hardening
162
+ (seek-coalescing, iOS priming, safe-area CSS) is always on regardless — that's not
163
+ a "mobile version," it's just the page not breaking when a phone visits — so a
164
+ desktop-only build still degrades gracefully.
165
+
166
+ 7. **Asset source — automatic or manual. ALWAYS ask; it decides who renders.**
167
+ Two options (`AskUserQuestion` in Claude Code; a plain either/or elsewhere):
168
+ *"How do you want to produce the stills and clips — should I generate them, or
169
+ do you want the prompts to render in tools of your choice?"* Record as
170
+ `ASSET_SOURCE`.
171
+ - **Automatic (default)** — the skill renders everything itself: stills via
172
+ Higgsfield `gpt_image_2` (or Codex `image_gen`), the video chain via Monid
173
+ with Higgsfield as fallback biller. Item 8 prices this path.
174
+ - **Manual** — the skill writes every prompt to a file (`still_<name>.txt` at
175
+ Step 2, `dive_<name>.txt` at Step 4, `conn_<i>.txt` at Step 5) plus the exact
176
+ conditioning frames each clip must start/end on; the user copy-pastes the
177
+ prompts into tools of their choice (any image tool for the stills; a
178
+ start/end-frame-capable video tool for the clips) and drops the finished
179
+ files into `$WORK`. The skill validates what comes back (count, dimensions,
180
+ aspect, duration, seam frames) and continues the same downstream pipeline —
181
+ only the render calls change. State the one hard requirement up front: the
182
+ seamless rule still applies, so their **video** tool must accept a
183
+ first/start frame (both architectures) and, for architecture B connectors, a
184
+ last/end frame — if it can't, steer `CAMERA` to architecture A or have them
185
+ pick a tool that can (Steps 4–5).
186
+
187
+ 8. **Budget — engines shown by cost, decided before anything renders.** On the
188
+ **manual** path (item 7) there is nothing to bill here — the spend happens in
189
+ the user's own tools — so skip the tier/backend picks below; just state the
190
+ workload they're signing up to produce (`N stills + (2N−1) clips [×2 if
191
+ mobile]`, plus re-roll headroom) and get a go. On the **automatic** path:
192
+ present the render tiers (`AskUserQuestion`), then compute and state the
193
+ estimated total for the user's N scenes — `N stills + (2N−1) videos [videos ×2
194
+ if mobile] + ~15% re-roll headroom` — and get a go before generating.
195
+ - **Video tier** (roster only — every option frame-locks seams, Step 4):
196
+
197
+ | Tier | Model | Rough cost |
198
+ |---|---|---|
199
+ | Draft / previz | `seedance_2_0_mini` (720p) | ~¼ of Standard |
200
+ | Standard (default) | `seedance_2_0` (1080p) | baseline |
201
+ | Alternate | `kling3_0` (720p native) | ≈ Standard; different look + content filter |
202
+
203
+ Draft doubles as the previz path: run the whole chain cheap, approve the
204
+ journey, re-render final legs on Standard (pipeline.md Notes) — suggest it
205
+ unprompted when the balance reads tight.
206
+ - **Backend — Monid is the DEFAULT biller for the chain** (Step 0.1; wiring
207
+ in pipeline.md → Monid backend). Same Seedance 2.0, per-clip USD instead of
208
+ credits. Token-priced `width × height × 24 × seconds / 1024` at $7/1M
209
+ (480p/720p) or $7.7/1M (1080p) — measured: 1080p 8s dive ≈ $2.99, 5s
210
+ connector ≈ $1.87; 720p ≈ $1.21 / $0.76; 480p ≈ $0.28 / $0.35. An N=6
211
+ desktop chain ≈ $27 at 1080p / ~$11 at 720p vs Higgsfield Plus-monthly ≈
212
+ $32 / $16 — ~15% cheaper per clip, parity with Plus-annual; structurally
213
+ better for one-off builds (pay-per-use, no monthly expiry). On Monid the
214
+ Draft/previz tier is simply the **same endpoint at 480p** — no model swap,
215
+ so previz→final stays one-model by construction. State `monid balance`
216
+ against the estimate; **fall back to Higgsfield credits** (per-model tiers
217
+ above) when the user prefers their subscription, the balance is short, or
218
+ the model must be `kling3_0` (Higgsfield-only). It's the same underlying
219
+ model (`seedance_2_0` ≙ Monid's `seedance-2.0`), so finishing a stranded
220
+ chain on the other biller is a reasonable rescue — but the serving stacks
221
+ differ and cross-provider seam character is **untested**: eyeball the first
222
+ rescued seam before rendering the rest, same as any model swap.
223
+ - **Stills source** (only offer if the Codex CLI is present, Step 0.5):
224
+ Higgsfield `gpt_image_2` (spends credits) vs **Codex `image_gen`** — the same
225
+ gpt-image-2 model billed to the ChatGPT subscription (zero credits; counts
226
+ toward Codex usage limits; 1536×1024 output — exactly 3:2, slightly under
227
+ Higgsfield's 2k). Stills are plain PNGs handed to `--start-image`, so the
228
+ video chain is indifferent to their source. Command in Step 2. **One source
229
+ for all N stills of a build** — the two render with slightly different
230
+ character (verified: Codex runs warmer/lighter), and mixing sources across
231
+ scenes reads as style drift, same reason the video chain uses one model.
232
+ - **Calibrate costs, don't guess.** The CLI exposes no pricing and plans differ.
233
+ Run ONE still and ONE video first, diff `higgsfield workspace list` before/
234
+ after, extrapolate to the full run, and warn the user whenever the estimate
235
+ exceeds ~70% of the balance. (Observed on a plus plan, 2026-07: Standard
236
+ video ≈ 40–55 credits, still ≈ 15.) A real `not_enough_credits` mid-run is
237
+ recoverable (finished clips survive; resume after top-up) but ugly — the
238
+ whole point of this step is that the user decides *before* the spend.
239
+
240
+ If the user names a video model outside the roster, honor it **only if it can
241
+ frame-lock seams** (Step 4). This skill only ships seamless output, so a model that
242
+ can't frame-lock is declined with a one-line why, not substituted in — use a roster
243
+ model instead.
244
+
245
+ Keep the scroll mechanic fixed (continuous fly-through) — that's the point of the skill.
246
+ See `references/prompts.md` for the intake checklist and copy structure.
247
+
248
+ ---
249
+
250
+ ## Step 2 — Generate the scene stills
251
+
252
+ One image per section, **all sharing the same style preamble** for cohesion. Default
253
+ model **`gpt_image_2`** (crisp, great at isometric illustration; returns a solid/white
254
+ background which is perfect for floating diorama "islands"). Use `nano_banana_2` only if
255
+ the brief is character/cartoon-heavy (note: `nano_banana_2` is a CLI alias — it resolves
256
+ to `nano_banana_pro`; it won't appear under that name in `higgsfield model list`).
257
+
258
+ Prompt shape (full templates in `references/prompts.md`):
259
+
260
+ ```
261
+ <STYLE PREAMBLE, identical every time — already carries the solid <bg> + palette>.
262
+ Render a wide 3:2 landscape image (≥1536 px). The background stays plain solid <bg>
263
+ across the whole frame — empty backdrop: no sky, no clouds, no horizon, no gradient.
264
+ Centered composition, nothing essential at the far edges. No text, no letters, no logos.
265
+ Subject: <what is in THIS diorama>.
266
+ ```
267
+
268
+ Self-contained on purpose: aspect, background lock and palette live in the prompt text,
269
+ because on the manual path each still may be rendered in a different tool or session
270
+ (full template + rationale in `references/prompts.md`).
271
+
272
+ - **Manual stills path** (`ASSET_SOURCE` = manual, Step 1.7) — skip the generate
273
+ commands below. Write the same prompt files (`$WORK/still_<name>.txt`, identical
274
+ preamble, from the prompts.md templates) and hand them over with the spec: one
275
+ image per file, **3:2 landscape, ≥1536 px wide, solid <bg> background, no
276
+ text** — any image tool the user likes. They save the results as
277
+ `$WORK/still_<name>.png` and tell you when the set is in. Validate before
278
+ continuing: all N present, each opens, aspect ≈ 3:2 (a few % tolerance), width
279
+ ≥ ~1200 px (`ffprobe` or PIL). Then run the same cohesion review as below;
280
+ off-style stills go back to the user for a re-roll (suggest re-running the same
281
+ prompt with an approved sibling as style reference). The one-source rule applies
282
+ here too — don't mix manual stills with CLI-rendered ones in one build.
283
+ - Run all N concurrently, detached. Command per scene:
284
+ `higgsfield generate create gpt_image_2 --prompt "$(cat scene_i.txt)" --aspect_ratio 3:2 --resolution 2k --quality high --wait --wait-timeout 15m --json > scene_i.json 2>scene_i.err`
285
+ - Result URL is `.[]0.result_url` in the `--wait --json` output. `curl` it down.
286
+ - **Codex stills variant** (if chosen at Step 1.8 — subscription-billed, zero
287
+ credits): same prompt files, same byte-identical preamble, generated by Codex's
288
+ built-in `image_gen`:
289
+
290
+ ```bash
291
+ codex exec -C "$WORK" -s workspace-write --skip-git-repo-check \
292
+ 'Use the image generation tool ($imagegen) to generate: '"$(cat "$WORK/still_i.txt")"' Wide 3:2 landscape, high resolution. Save it as ./still_i.png. Do not do anything else.' \
293
+ < /dev/null
294
+ ```
295
+
296
+ Single-quote the `$imagegen` segment (the shell must not expand it); if editing
297
+ with reference images, the prompt goes BEFORE any `-i` flag (it's variadic).
298
+ ~1–3 min per image; run a few in parallel, not all N at once — and keep the
299
+ `< /dev/null`: parallel `codex exec` calls sharing a script's stdin hang
300
+ waiting for input (Gotchas). Output lands at
301
+ 1536×1024 (3:2) — fine for `--start-image` and posters. Everything downstream
302
+ (cohesion review, knockout, dives) is unchanged.
303
+ - A generation may fail transiently (HTTP 503) — re-roll that one individually; don't
304
+ restart the batch.
305
+ - **Review the stills before continuing.** They must read as one cohesive world (same
306
+ angle, palette, light). If one is off-style, regenerate it, optionally passing an
307
+ approved scene as `--image` to lock style.
308
+
309
+ See `references/pipeline.md` for the exact batch script.
310
+
311
+ ---
312
+
313
+ ## Step 3 — (Optional) Float the scenes
314
+
315
+ If you want the dioramas to float over an atmospheric background instead of sitting in a
316
+ solid box, knock out the flat background to transparency with
317
+ `references/knockout.py` (border-connected flood fill — preserves interior colour that
318
+ matches the bg, e.g. cream walls). Then encode to webp. If you'd rather keep it simple,
319
+ just make the page background the same colour as the scene background and skip this.
320
+
321
+ These stills double as **video posters and lazy-load fallbacks**, so keep them.
322
+
323
+ ---
324
+
325
+ ## Step 4 — Camera architecture (implements the Step 1.4 choice)
326
+
327
+ How the camera moves *between* scenes is the single biggest quality lever. The user
328
+ already chose the style at the interview (`CAMERA`, Step 1.4): fly-through → **B**,
329
+ walkthrough → **A**, locked-iso → **A + the locked-iso leg clause** (prompts.md). If
330
+ the interview somehow skipped it, ask now — never silently pick for them. The two
331
+ shapes, and the grammar that colors them:
332
+
333
+ ### Video model — pick ONE for the whole chain
334
+
335
+ **This skill only ships seamless output**, so the only usable models are ones that can
336
+ frame-lock a seam: every chained clip must accept `--start-image`, and connectors also
337
+ need `--end-image`. That capability — not preference — is the selection rule. Check any
338
+ model with `higgsfield model get <job_type>` and **skip anything whose media inputs are
339
+ reference-only** (no start/end image): it can only *condition* a generation, not
340
+ *continue* a shot, so it physically can't hold a seam. Schemas below were confirmed
341
+ against the CLI:
342
+
343
+ | Model | start/end image | Notes |
344
+ |---|---|---|
345
+ | `seedance_2_0` (default) | ✓ / ✓ | Full chain (legs + connectors). `--mode std --resolution 1080p`. Its NSFW filter is the touchy one (see Gotchas). |
346
+ | `kling3_0` | ✓ / ✓ | Full chain — tested: `--mode std --sound off --duration 5` with start+end images accepted, seams frame-lock cleanly. **No `--resolution` param** (don't pass one; `--mode std` returns **720p native** — encode what ffprobe reports, never upscale). Sound defaults **on** → `--sound off`. `--duration` default 5, try 10 for legs. Different content filter than Seedance — the sanctioned NSFW fallback. |
347
+ | `seedance_2_0_mini` | ✓ / ✓ | Cheap draft tier that keeps frame-locking (720p). The previz tier: run the whole chain here first, then re-render final legs on the full model — still seamless, so it translates directly. |
348
+
349
+ Those three are the roster — all do both architectures. (`kling3_0_turbo` also frame-locks
350
+ via `--start-image`, but has no `--end-image`, so it's architecture-A-only and can't make
351
+ connectors; it also takes a different flag set — no `--mode`, has `--resolution` — so it
352
+ doesn't drop into the pipeline as-is. It's not in the default roster; only reach for it, and
353
+ wire it by hand, if architecture A's sequential render time is a proven bottleneck and you've
354
+ benchmarked it as actually faster.)
355
+
356
+ One more architecture-A-only candidate, worth knowing because it is by far the cheapest
357
+ probe: **`minimax_hailuo`** (Hailuo-2.3, ~6 credits per 768p/6s clip vs 22–72 for the
358
+ roster). Verified 2026-07: `--start-image` + prompt frame-locks (output frame 0 ≡ input,
359
+ PSNR 33 dB) and a forward-glide prompt was obeyed, gently. Constraints: the 2.3 variant
360
+ rejects `end_image` (no connectors → arch A only), output aspect follows the input image
361
+ (hand it a 16:9 canvas, not a bare 3:2 still), motion runs subtler than seedance, and
362
+ don't pass `--resolution` (the CLI mis-types the enum; the 768 default works — 1080
363
+ supports 6s only). One clip ≠ a chain: qualify a leg-to-leg handoff before betting a
364
+ full build on it.
365
+
366
+ Rules:
367
+ - **One model for all chained clips.** Each renderer has its own motion/color/grain
368
+ character; mixing models mid-chain keeps *position* continuity (frames still hand off)
369
+ but the render-character shift reads as a subtle pop. The one sanctioned exception is
370
+ the NSFW fallback for a single stubborn clip (Gotchas) — a slight character shift on
371
+ one 5s connector beats a missing connector.
372
+ - Default to `seedance_2_0`, rendered through **Monid by default** (per-clip USD —
373
+ next section) with Higgsfield credits as the fallback biller (Step 0.1/1.8); honor
374
+ a user's stated preference **only if the model qualifies** (frame-locking). If it
375
+ doesn't, say so and use a supported model — never ship a non-seamless build to
376
+ satisfy a model request. `kling3_0` and `seedance_2_0_mini` exist only on the
377
+ Higgsfield side.
378
+ - The pipeline scripts take the model as `$VMODEL` with per-model flags already cased
379
+ out (`references/pipeline.md`).
380
+
381
+ ### Monid backend — the DEFAULT chain biller (qualified 2026-07-25)
382
+
383
+ Monid's **`bytedance /v1/video/seedance-2.0`** passed both paid probes on
384
+ 2026-07-25 and is the **default** way this skill renders the chain, for **both
385
+ architectures** — it is the roster's `seedance_2_0` served pay-per-USD (wiring in
386
+ pipeline.md → "Monid backend"; Higgsfield renders the chain only as the fallback
387
+ biller or for Higgsfield-only models):
388
+
389
+ - **Leg probe** (prompt + `first_frame` image): output frame 0 ≡ input still
390
+ (PSNR 31.6 dB), forward-glide prompt obeyed, billed the advertised cell
391
+ ($0.279 / 480p 4s).
392
+ - **Connector probe** (prompt + `first_frame` + `last_frame`): start locked
393
+ (31.6 dB); the end **lands close but not pixel-perfect** (27.5 dB, same
394
+ composition, prop-level drift) — the exact end-image behavior Seedance shows
395
+ on Higgsfield, covered by the engine's seam crossfade and by using the next
396
+ dive's ACTUAL first frame as the end-image (Step 5 law, unchanged).
397
+
398
+ The I/O contract differs from the Higgsfield CLI — three rules:
399
+
400
+ 1. **Images go by URL, never inline.** `content` items are
401
+ `{"type":"image_url","image_url":{"url":…},"role":"first_frame"|"last_frame"}`;
402
+ base64 data URLs are **rejected** ("Must be a public https:// URL or an
403
+ asset://<id> reference"). Local frames travel through Monid's free workspace
404
+ file system: `sfs /put` → `curl -T` the bytes → `sfs /cat` returns a signed
405
+ public URL to paste into the body ($0, explicitly built for this).
406
+ 2. **Pass `ratio` explicitly** (`16:9`, or `9:16` for the mobile chain) — the
407
+ adaptive default follows the input image's aspect instead.
408
+ 3. **Bill-check every clip**: cost is token-priced
409
+ (`w × h × 24 × sec / 1024` at $7–7.7/1M); read `cost.value` off each run.
410
+
411
+ History that shaped these rules (still true as of 2026-07-25): the seedance
412
+ endpoints were text-to-video-only until late July 2026 — **re-`inspect` before
413
+ each build; the catalog moves in both directions.** `minimax
414
+ /v1/video_generation` (Hailuo-2.3) remains disqualified: sending `prompt` +
415
+ `first_frame_image` together silently drops the image (unrelated t2v output,
416
+ wrong price cell); image-only frame-locks (33 dB) but has no camera control.
417
+
418
+ **Qualification protocol for any new/changed Monid endpoint** (each probe is one
419
+ cheap 480p clip): (1) prompt + first-frame from a real still — frame 0 must
420
+ match the input to codec noise (PSNR ≳ 30 dB) and `cost.value` must match the
421
+ advertised cell; (2) for connector duty, add a `last_frame` from a different
422
+ still — the end must land on that composition (Seedance-style near-miss is fine,
423
+ the crossfade covers it). Pass → pay-per-clip tier (arch A if start-only; full
424
+ roster if start+end).
425
+
426
+ ### Manual rendering path (`ASSET_SOURCE` = manual — Step 1.7)
427
+
428
+ Skip the model roster and the Monid backend above; the architecture choice (A vs
429
+ B) still stands — it decides which clips exist and which frames chain them. The
430
+ frame-lock rule is unchanged, so confirm the user's video tool **accepts a
431
+ first/start frame** (both architectures) and, for B's connectors, a **last/end
432
+ frame**. If it can't take an end frame, either re-confirm `CAMERA` as
433
+ architecture A with the user (no connectors) or have them use a tool that can
434
+ (e.g. Kling's start/end-frame mode) — never ship unseamed connectors. Per clip:
435
+
436
+ 1. Write the same prompt file as the automatic path (`$WORK/dive_<name>.txt` /
437
+ `conn_<i>.txt`) and hand it over **with the conditioning image(s) it assumes**:
438
+ arch A leg i → the previous leg's actual last frame; B dive → the scene's
439
+ solid-bg still; B connector → the extracted `last_`/`first_` pair (Step 5).
440
+ 2. Spec for the user: 16:9 landscape (9:16 on the mobile chain), ~8 s dives/legs,
441
+ ~5 s connectors, highest quality their tool offers, no audio. Results saved as
442
+ `$WORK/dive_<name>.mp4` / `$WORK/conn_<i>.mp4`.
443
+ 3. When a batch lands, validate each file **before** chaining: plays, right
444
+ aspect, duration ≈ spec — and frame 0 must match the handed-over start frame
445
+ (extract it and compare). A clip whose tool ignored the start image can't hold
446
+ its seam; send it back for re-generation, don't crossfade over it.
447
+
448
+ Arch A stays logically sequential even though rendering is manual — each leg's
449
+ start frame is extracted from the previous leg's returned file. Everything
450
+ downstream (boundary-frame extraction, encode, engine, QA) is identical to the
451
+ automatic path.
452
+
453
+ **Every manual handoff — stills, dives, connectors — is ALWAYS presented as a spec
454
+ table**, never a prose list: prompt file, conditioning frame(s), the exact output
455
+ filename, and a status column you keep current (pending / rendered / accepted) as
456
+ files arrive and pass validation. The user works through it offline, in any order,
457
+ and the table is the contract for what you're still waiting on. The connector table
458
+ in Step 5 is the canonical example.
459
+
460
+ ### A) Continuous forward take — RECOMMENDED for grounded / realistic / walkthrough
461
+ One camera that only ever glides **forward**, first scene through last, as a single take.
462
+ Generate the legs **sequentially**: leg 0 from scene-0's still (glide forward into it);
463
+ then each leg's `--start-image` = the **previous leg's ACTUAL last frame** (extract with
464
+ ffmpeg), prompt *"continue gliding smoothly FORWARD into [scene i], never pulling back"*
465
+ (or an expressive mid-leg move under the motion-handoff contract — see **Camera grammar**
466
+ below), and **no `--end-image`** — an end-image of a wide establishing shot forces the
467
+ camera to pull back, which is the #1 cause of stutter. Extract each leg's last frame to feed the
468
+ next. Result: every seam is frame-identical **and** the camera never reverses. There are
469
+ **no connectors** (skip Step 5) — the legs ARE the journey. Wire each leg as a section
470
+ clip with `connectors: []` and a small `crossfade` (~0.08). Even without an `--end-image`
471
+ the legs still arrive at distinct rooms (the prompt steers the content). Cost: strictly
472
+ **sequential** (can't parallelize) and slower; interiors trip the NSFW filter, so build in
473
+ re-rolls (3 attempts/leg).
474
+
475
+ ### B) Dive-in + aerial connector — only for diorama / miniature / god's-eye worlds
476
+ A "dive into each scene" clip + a connector that pulls **up and out** and flies over to the
477
+ next scene (Step 5). The pull-out **reverses camera direction at every seam** (forward dive
478
+ → backward pull-out). In a miniature/diorama world that reads as an intentional "zoom out
479
+ to the map, fly to the next island"; in a grounded first-person walkthrough it reads as a
480
+ jarring **rewind/stutter**. Use B only for the map-like aesthetic — which is exactly
481
+ what the "fly through the world" interview answer opts into; the reversal reads as
482
+ intentional there. If the user picked B against a grounded/photoreal direction, say
483
+ why it will read as a stutter and confirm before rendering.
484
+
485
+ ### Camera grammar — the move should fit the concept (A is NOT "forward only")
486
+
487
+ "Forward only" is the *seam* rule, not the *leg* rule. The physics of the chain:
488
+
489
+ - **Position continuity** at a seam comes from the frame handoff (next leg starts from the
490
+ previous leg's actual last frame).
491
+ - **Velocity continuity** at a seam means the camera must never *reverse across a seam* —
492
+ that's the rewind stutter.
493
+ - **Inside a single leg the camera is free.** One leg is one continuous render — there is
494
+ no seam to break mid-leg, so orbits, crane-ups, lateral tracking, even a push-in that
495
+ eases back out are all safe *within* the clip. Reversals are only fatal *across* seams.
496
+
497
+ So give each leg an expressive move chosen from the scene's own logic, under a **motion
498
+ handoff contract**: every leg **ends by settling into a slow, steady forward drift** toward
499
+ the next destination (final ~1 s), and every leg **begins by continuing that same drift**.
500
+ Keep both clauses in the prompts verbatim (templates in `references/prompts.md`).
501
+
502
+ Pick the grammar from the concept:
503
+
504
+ | Concept / tone | Mid-leg move |
505
+ |---|---|
506
+ | Product / luxury retail | slow half-orbit around the hero object, then continue past it |
507
+ | Real estate / hospitality | steadicam glide through doorways; gentle crane-up in atria |
508
+ | Industrial / process / logistics | low lateral track alongside the line, foreground parallax |
509
+ | Travel / outdoors / campus | drone-style rise-and-reveal, then a descending swoop |
510
+ | Food / craft / detail-driven | push in close to the craft moment, ease back, carry on |
511
+ | Playful miniature (arch. B) | dives + aerial hops — the connector IS the grammar |
512
+
513
+ Honest costs: expressive mid-leg moves raise re-roll odds — the model can end a fancy move
514
+ in a state that isn't a clean forward drift. Mitigations: keep the final-second settle
515
+ clause verbatim; **eyeball each leg's last frame before chaining the next** (it should look
516
+ like a frame from a gentle forward glide — if not, re-roll before wasting the next leg);
517
+ budget ~1 extra re-roll per expressive leg. A plain forward glide stays the zero-risk
518
+ default — use it for legs where the scene itself is the show.
519
+
520
+ **Locked-iso variant** (`CAMERA` = locked isometric glide): architecture A where every
521
+ leg pins the view instead of taking a mid-leg move — "the camera keeps exactly the same
522
+ high isometric angle throughout, no rotation, no orbit, no tilt; it only travels
523
+ straight and level, the world sliding past beneath the same view" (verbatim clause in
524
+ prompts.md). The handoff contract is unchanged. Seedance drifts the angle slightly on
525
+ long legs — the existing eyeball-each-last-frame rule is the catch; re-roll a leg whose
526
+ view has rotated. Calmest look, cheapest re-rolls, and the closest to the Emons
527
+ reference.
528
+
529
+ Two related pacing knobs live in the engine (Step 7): per-section `scroll` (more scroll
530
+ distance = longer dwell in that scene) and `linger` (the camera settles mid-scene exactly
531
+ while the copy peaks, then picks up speed toward the seam). Prefer expressive motion in the
532
+ *clip* and restraint in the *scrub mapping* — they compound.
533
+
534
+ And remember scroll is a scrubber: visitors can scroll **up**, so every move also plays in
535
+ reverse. That's free and expected — no extra work — but it's another reason seam velocity
536
+ must be consistent in both directions (a seam that reads fine forward reads as a stutter
537
+ backward too if velocity flips).
538
+
539
+ **For B**, one camera flight per scene: starts high/outside, descends into the interior,
540
+ structure opens. Model: the chain model you picked above (default **`seedance_2_0`**),
541
+ `--start-image = the scene still`.
542
+
543
+ - Use the **solid-background still** (not the knocked-out transparent one) as the
544
+ start image, so the video has a full frame.
545
+ - Prompt: "Single continuous cinematic camera move, no cuts. Begin high and far looking
546
+ at the whole <scene> from outside … descend and fly inside toward <focal point> … the
547
+ roof/walls gently open to reveal the interior. <style>, smooth graceful slow motion.
548
+ No text." (Template in `references/prompts.md`.)
549
+ - Params (seedance): `--mode std --resolution 1080p --aspect_ratio 16:9 --duration 8`.
550
+ For Kling: drop `--resolution` (no such param), add `--sound off`, `--duration 10`.
551
+ Do **not** pass `--generate-audio` (it errors on seedance; audio is wasted anyway —
552
+ you'll mute).
553
+ - Run concurrently, detached, then download each `.result_url`. Re-roll individual
554
+ failures. Keep the raw 1080p sources — you need their frames next.
555
+
556
+ ---
557
+
558
+ ## Step 5 — Connectors (architecture B only)
559
+
560
+ Skip this whole step for architecture **A** — the forward take has no connectors; its legs
561
+ already chain seamlessly. This step applies to **B** (diorama/miniature), and note the
562
+ reversal caveat from Step 4.
563
+
564
+ The connector clips are what make the world feel *connected* instead of cut. A connector
565
+ flies from the end of scene i out and into the start of scene i+1. **Both of its
566
+ endpoints must be the ACTUAL RENDERED FRAMES of the neighbouring clips — never the
567
+ original diorama still.**
568
+
569
+ Why: every Higgsfield generation renders slightly differently. If a connector *ends* on
570
+ a fresh render of "the kitchen diorama," but the next dive clip *starts* on its own
571
+ different render of that same diorama, the two won't match and you get a pop at the seam.
572
+ The fix is to hand off the exact pixels:
573
+
574
+ ```
575
+ For each connector between dive_i and dive_{i+1}:
576
+ start-image = the LAST frame extracted from dive_i's rendered video
577
+ end-image = the FIRST frame extracted from dive_{i+1}'s rendered video
578
+ ```
579
+
580
+ Now every seam is frame-identical on *both* sides:
581
+ `dive_i.end == connector.start` and `connector.end == dive_{i+1}.start`.
582
+
583
+ Extract the boundary frames from the rendered dives (not the stills):
584
+
585
+ ```bash
586
+ ffmpeg -sseof -0.15 -i dive_i.mp4 -frames:v 1 -q:v 2 dive_i_last.png # interior of i
587
+ ffmpeg -ss 0 -i dive_{i+1}.mp4 -frames:v 1 -q:v 2 dive_next_first.png # establishing of i+1
588
+ ```
589
+
590
+ Generate the connector (`--duration 5` is plenty). Connectors need `--end-image`, so the
591
+ model must accept it — any roster model does (`seedance_2_0`, `seedance_2_0_mini`,
592
+ `kling3_0`):
593
+
594
+ ```bash
595
+ higgsfield generate create "$VMODEL" \
596
+ --prompt "$(cat connector_i.txt)" \
597
+ --start-image dive_i_last.png --end-image dive_next_first.png \
598
+ $VOPTS --aspect_ratio 16:9 --duration 5 --wait --json
599
+ # seedance: VOPTS="--mode std --resolution 1080p"; kling3_0: VOPTS="--mode std --sound off"
600
+ ```
601
+
602
+ Connector prompt: "Single continuous camera move, no cuts. Pull up and back out of
603
+ <scene i>, rise into the sky, glide across the connected miniature world, and arrive
604
+ above <scene i+1>, beginning to descend toward it. Seamless flowing aerial transition.
605
+ <style>. No text." (Template in `references/prompts.md`.)
606
+
607
+ Insurance: Seedance lands *close* to the end-image but not always pixel-perfect, so the
608
+ engine still applies a **short crossfade** (a few frames) at each seam. Frame-matched
609
+ endpoints + a small crossfade = no visible cut. Never skip the actual-frame handoff and
610
+ rely on the crossfade alone; a big content jump can't be hidden by a crossfade.
611
+
612
+ **Manual path (Step 1.7):** the frame extraction above is yours to run either way —
613
+ hand the user `conn_<i>.txt` plus the two extracted PNGs and wait for
614
+ `$WORK/conn_<i>.mp4`. Accept it only if frame 0 matches `last_<prev>.png` and the
615
+ final frame lands on the `first_<next>.png` composition (a Seedance-style near-miss
616
+ is fine — the crossfade covers it).
617
+
618
+ The handoff is ALWAYS a spec table, one row per connector, filled with this build's
619
+ real extracted frames:
620
+
621
+ | Prompt file | Start frame | End frame | Save as | Status |
622
+ |---|---|---|---|---|
623
+ | `conn_1.txt` | `last_<scene1>.png` | `first_<scene2>.png` | `$WORK/conn_1.mp4` | pending |
624
+ | `conn_2.txt` | `last_<scene2>.png` | `first_<scene3>.png` | `$WORK/conn_2.mp4` | pending |
625
+ | … | … | … | … | … |
626
+
627
+ State the acceptance rule with it: the start frame must be obeyed exactly; the end
628
+ frame only needs to land on the same composition.
629
+
630
+ ---
631
+
632
+ ## Step 6 — Encode for smooth scrubbing
633
+
634
+ Scrubbing = setting `video.currentTime` from scroll. Two things matter, and they are
635
+ often gotten wrong:
636
+
637
+ 1. **Seekability, not keyframe density, is what makes scrubbing work.** Many static
638
+ hosts (and `python -m http.server`) don't serve HTTP byte-range requests, which pins
639
+ `video.seekable` to `[0,0]` and clamps *every* seek to frame 0 — the video looks
640
+ frozen. The robust fix is to **fetch each clip as a `Blob` and play it from an
641
+ in-memory object URL** (blobs are always fully seekable). The engine does this.
642
+ Because of it, you do **not** need all-intra video.
643
+ 2. **Don't shrink quality to get smooth seeks.** Encode at the **native resolution**
644
+ (1080p from Seedance — don't downscale), `crf ~20`, a **small GOP** (`-g 8`) rather
645
+ than all-intra (all-intra bloats an 8s clip to ~25 MB; GOP 8 is ~8 MB and scrubs
646
+ fine via blob). Strip audio, add faststart, and a light `unsharp` counters video
647
+ softness:
648
+
649
+ ```bash
650
+ ffmpeg -i src.mp4 -an -vf "unsharp=5:5:0.8:5:5:0.0" \
651
+ -c:v libx264 -preset slow -crf 20 -pix_fmt yuv420p \
652
+ -g 8 -keyint_min 8 -sc_threshold 0 -movflags +faststart out.mp4
653
+ ```
654
+
655
+ Encode all 2N-1 clips (dives + connectors) with the same settings for uniform quality.
656
+
657
+ **Mobile encodes (only if the user opted in at Step 1.6).** The mobile version is
658
+ the **native 9:16 portrait chain** (pipeline.md §6b): portrait renders of every dive and
659
+ connector, encoded **720 wide (`scale=720:-2`), `-g 4`** (more keyframes = cheaper seeks —
660
+ phone decoders' seek cost scales with GOP length), crf 23 — wired as `clipMobile` /
661
+ `connectorsMobile`, with each portrait dive's first frame extracted as the section's
662
+ `stillMobile` poster (Step 7). The engine serves them automatically on phones and falls
663
+ back to the desktop clip when absent. The 16:9 centre-crop `encm()` encodes
664
+ (pipeline.md §6) are a **fallback only** — for when credits can't cover the portrait
665
+ chain — and shipping them must be called out to the user, never silent. If the user chose
666
+ desktop-only, skip this — the engine still hardens phone scrubbing regardless
667
+ (seek-coalescing, iOS priming), so the page degrades gracefully rather than breaking.
668
+
669
+ ---
670
+
671
+ ## Step 7 — Assemble the page
672
+
673
+ Copy `references/scrub-engine.js` (and, if you want a fully standalone page, the tiny
674
+ `references/index-template.html`) into the user's project — or adapt into their
675
+ framework. It's config-driven and self-contained:
676
+
677
+ ```js
678
+ mountLetsScroll(document.getElementById('world'), {
679
+ brand: { name: 'Pearl & Co.' },
680
+ diveScroll: 1.3, connScroll: 0.9, // viewport-heights of scroll per clip
681
+ sections: [
682
+ { id:'farm', label:'The Farms', still:'assets/farm.webp',
683
+ clip:'assets/vid/farm.mp4',
684
+ clipMobile:'assets/vid/farm-m.mp4', // mobile opt-in only: native 9:16 render
685
+ stillMobile:'assets/farm-m.webp', // its first frame as the portrait poster
686
+ scroll: 1.6, linger: 0.45, // optional pacing: longer dwell + camera settles mid-scene
687
+ accent:'#8FB98A', eyebrow:'From leaf to last sip', title:'It starts in the hills.',
688
+ body:'…', tags:['Single-origin','Hand-picked'] },
689
+ // …one per section; last may carry a `cta`
690
+ ],
691
+ connectors: ['assets/vid/conn1.mp4','assets/vid/conn2.mp4', /* … length = sections-1 */],
692
+ connectorsMobile: ['assets/vid/conn1-m.mp4','assets/vid/conn2-m.mp4' /* … same length; mobile opt-in only */],
693
+ });
694
+ ```
695
+
696
+ The engine handles: the ordered dive/connector chain, scroll→currentTime with rAF
697
+ smoothing, blob loading, lazy prefetch of nearby clips, frame-matched crossfades, pinned
698
+ per-section copy (first section greets on landing, last holds its CTA), a route rail,
699
+ `prefers-reduced-motion`, and mobile. **Pacing per section:** `scroll` overrides
700
+ `diveScroll` for that scene (more scroll = longer dwell) and `linger` (0–1, keep ≤ 0.6)
701
+ remaps time so the camera settles mid-scene — exactly while the copy peaks — then speeds
702
+ up toward the seam; seam frames are untouched (f(0)=0, f(1)=1). Give the hero and finale
703
+ scenes a higher `scroll` + some `linger`; keep transit scenes brisk. Theme it with CSS variables (`--accent`,
704
+ `--sw-bg`, `--sw-ink`, …) — the visual identity comes from the generated clips, so the
705
+ chrome stays quiet. See the header of `scrub-engine.js` for the full config + CSS vars.
706
+
707
+ **On phones the engine adapts automatically** (coarse pointer or ≤860px): it serves
708
+ `clipMobile` / `connectorsMobile` when present, **coalesces seeks** (never queues a new
709
+ `currentTime` while the decoder is still seeking — this is what stops a fast flick from
710
+ freezing the clip), **keeps the still as a poster until the clip paints its first frame**
711
+ and **primes each video on first touch** (fixes iOS's blank-until-played video), drops the
712
+ drifting particles, ignores URL-bar-only resizes (no scroll jump), and uses safe-area
713
+ insets so copy clears the notch/home indicator. All of this hardening is on by default —
714
+ no config needed. The `clipMobile`/`connectorsMobile` encodes are the opt-in part
715
+ (Step 1.6): only wire them when the user asked for the mobile version.
716
+
717
+ For non-JS backends (Python/Rails/etc.): serve the assets and drop the engine `<script>`
718
+ into the rendered HTML; nothing about it is framework-specific.
719
+
720
+ ---
721
+
722
+ ## Step 8 — QA the seams (don't skip)
723
+
724
+ Drive the page in a headless browser and **verify frame continuity at the seams**, which
725
+ is the thing most likely to be wrong:
726
+
727
+ - Screenshot at scroll positions just before and just after each seam. The two frames
728
+ must be near-identical (the dive's last frame == the connector's first frame). If they
729
+ pop, you used the diorama still instead of the actual rendered frame (redo Step 5), or
730
+ the crossfade band is too short. Calibration: judge seams by *composition*, not raw
731
+ PSNR — at 720p/1080p a correctly frame-locked seam can read ~18–25 dB from detail
732
+ shimmer alone (observed on a verified-good build); a real mismatch shows as different
733
+ composition/props, not just softness.
734
+ - Check the console for errors, confirm `video.seekable.end(0) > 0` (blob working), and
735
+ that `currentTime` tracks scroll across each clip's band.
736
+ - **Mobile — full checklist only if the user opted into the mobile version (Step 1.6).**
737
+ For a desktop-only build, just sanity-check a phone viewport once: page loads, still
738
+ posters show, nothing overlaps — the engine's hardening covers graceful degradation.
739
+ For the mobile build (do this on a real phone or an emulated one, portrait + landscape):
740
+ - Emulate a phone viewport **with CPU throttled 4–6×** and scroll fast — the clip should
741
+ track without freezing (the seek-coalescing + `-m.mp4` encodes are what make this hold).
742
+ - Confirm the first scene shows immediately (its still is the poster) and the video takes
743
+ over the instant you scroll — no blank/black scene (the iOS priming fix). Test iOS Safari
744
+ specifically; it's the one that goes blank if this regresses.
745
+ - Verify the `-m.mp4` variant is actually served on mobile (Network panel), and the
746
+ heavy 1080p master on desktop. The mobile clips must be **natively portrait**
747
+ (`videoWidth < videoHeight` — not a downscaled 16:9 file), and the `stillMobile`
748
+ posters must be served and match each portrait clip's first frame (no
749
+ landscape→portrait flash when the video paints).
750
+ - Slowly scroll so the URL bar collapses — the page must **not jump** (height-only resizes
751
+ are ignored on touch). Rotate the device — layout should recompose cleanly.
752
+ - Only if the crop **fallback** shipped (no credits for the portrait chain): portrait
753
+ crops a 16:9 clip to its centre — confirm the focal subject still reads, and remind
754
+ the user this is the stopgap, not the mobile version.
755
+ - Check reduced-motion (should fall back to the stills, no video, no particles).
756
+
757
+ ---
758
+
759
+ ## Gotchas (hard-won)
760
+
761
+ - **Seam pop** → connector endpoints were the diorama stills, not the neighbouring
762
+ clips' actual frames. Always extract real frames (Step 5).
763
+ - **Seam stutter / camera "jumps backward"** → even with frame-matched seams, if the
764
+ camera *velocity reverses* (forward dive, then a connector that pulls back out) it
765
+ reads as a rewind. This is inherent to architecture B. For any grounded walkthrough use
766
+ architecture A (one continuous forward take — legs chained from actual last frames, no
767
+ pull-back, no `--end-image`); see Step 4.
768
+ - **Frozen video / stuck at frame 0** → `seekable=[0,0]`; the host isn't serving byte
769
+ ranges. Use blob URLs (engine does).
770
+ - **Huge files** → you used all-intra. Use `-g 8` + blob instead.
771
+ - **Soft / low quality** → you downscaled or over-compressed. Encode native 1080p,
772
+ crf ≤ 20, add `unsharp`. Video is inherently softer than the stills — keep the stills
773
+ as the lite fallback for max fidelity.
774
+ - **Concurrent gens 503 / "not_enough_credits" race** → transient when many launch at
775
+ once; re-roll the individual failure, it's not really out of credits (verify with
776
+ `higgsfield workspace list`).
777
+ - **NSFW false-positives (Seedance `status "nsfw"`)** → the video content filter flags
778
+ perfectly innocuous clips, especially **bedroom, pool, spa/wellness** contexts and
779
+ trigger words like "bed", "pool", "waterfall", "wine", "swim". It's partly the prompt
780
+ wording and partly the reference frames. Fixes, in order: (1) re-roll — it's often
781
+ non-deterministic and passes on the 2nd–3rd try; (2) strip trigger words and add
782
+ "empty, unoccupied, no people, no figures, architectural, tasteful"; (3) regenerate
783
+ just that clip on **`kling3_0`** with the same start/end frames — a different
784
+ provider's filter often passes what Seedance blocks. Expect a slight render-character
785
+ shift on that one clip (each model has its own grain/motion feel); for a 5s connector
786
+ behind a crossfade that usually beats option (4): set the connector slot to `null` —
787
+ the engine crossfades that seam directly (optional connectors), so the page still
788
+ completes. Budget extra credits/time for these re-rolls on interiors/real-estate content.
789
+ - **Dark / custom theme** → the engine wraps its default tokens in `@layer sw`, so a
790
+ page-level `:root` / `.sw-root { --sw-bg; --sw-ink; --sw-accent; --sw-font-* }` block
791
+ wins cleanly (no specificity hacks). `--sw-ink` is your primary **text/heading** colour;
792
+ the **accent** fills the primary button and active nav. For a dark theme, set `--sw-bg`
793
+ dark and `--sw-ink` light — the copy scrim and title shadow follow `--sw-bg` automatically.
794
+ - **Phone scrub stutters / freezes on a fast flick** → the 1080p master is too heavy for a
795
+ phone decoder and seeks pile up. Ship the `-m.mp4` mobile encodes (720p, `-g 4`) and wire
796
+ `clipMobile`/`connectorsMobile` (Step 6/7). The engine already coalesces seeks; the lighter
797
+ encode is the other half. Still choppy on a low-end device? Tighten GOP (`-g 2` / all-intra).
798
+ - **Blank / black scene on iOS (desktop was fine)** → an iOS Safari quirk: a muted video that
799
+ was never played won't paint a seeked frame. The engine fixes this by keeping the still as a
800
+ poster until the clip paints and priming each video on first touch — so **don't** hide the
801
+ still on `loadedmetadata` or strip the `playsinline`/`muted` attributes if you adapt the
802
+ engine into a framework.
803
+ - **Page jumps while scrolling on mobile** → something is re-running layout on the URL-bar
804
+ show/hide `resize`. The engine ignores height-only resizes on touch; if you ported it, gate
805
+ your resize handler on a width change (keep the `orientationchange` path for rotation).
806
+ - **Copy hidden behind the URL bar / notch on mobile** → use the engine's safe-area-aware
807
+ bottom offset (`env(safe-area-inset-bottom)` + `dvh`); make sure the page's
808
+ `<meta viewport>` includes `viewport-fit=cover` (the template does).
809
+ - **Portrait crops the scene** → a 16:9 clip on a tall phone shows only its centre — which
810
+ is why the mobile version is the native 9:16 chain (§6b), never the crop. If you're seeing
811
+ this on a mobile build, either the crop fallback shipped (call it out to the user) or the
812
+ 9:16 encodes aren't actually being served (check `videoWidth < videoHeight`). Keeping each
813
+ scene's focal subject centred (prompts.md) still matters for the desktop film itself.
814
+ - **`--generate-audio` errors on seedance** → omit it; mute in HTML and `-an` on encode.
815
+ - **Kling rejects your flags** → `kling3_0` has **no `--resolution` param** (don't pass
816
+ one; encode at whatever native res ffprobe reports) and **sound defaults on** — pass
817
+ `--sound off`. Duration default is 5; legs/dives want 10.
818
+ - **Seam pop only where you "saved credits"** → you swapped models mid-chain, or used a
819
+ start-image-only model where a connector needs an `--end-image`. One model for the whole
820
+ chain; the only cheap tier is `seedance_2_0_mini`, which keeps frame-locking so it stays
821
+ seamless. (Any model with reference-only inputs can't hold a seam at all — Step 4.)
822
+ - **Manual-path clip pops at its seam** → the user's tool ignored the start frame,
823
+ or they generated from the still instead of the extracted handoff frame. Diff
824
+ frame 0 of every manual clip against the handed-over PNG before accepting it —
825
+ drifted means it goes back for re-generation; no crossfade fixes a wrong start.
826
+ Off-world manual stills have the same fix: one tool + the byte-identical style
827
+ preamble for all N, no mixing with CLI renders.
828
+ - **Monid seedance rejects inline images** → "Must be a public https:// URL or an
829
+ asset://<id> reference": frames go through the free `sfs` file system
830
+ (put → `curl -T` → cat → signed URL; pipeline.md → Monid backend), never base64.
831
+ Quirk: `/put` echoes back `home/<path>`, but `/cat` and `/ls` want the **original
832
+ relative path** you gave `/put` — using the echoed path 404s.
833
+ - **Monid clip wrong aspect** → the `ratio` default is adaptive and follows the input
834
+ image (a 3:2 still → a 4:3-ish video). Pass `--ratio` — `16:9` desktop, `9:16`
835
+ mobile chain — explicitly on every chained clip.
836
+ - **Monid CLI "Polling timed out after 120s"** → only the local wait died; the run
837
+ continues server-side. Re-poll with `monid runs get -r <runId> -w 120` (find the id
838
+ in `monid runs list`). Result URLs expire (~24–48 h) — download immediately.
839
+ - **Monid minimax drops the image when a prompt is present** → `prompt` +
840
+ `first_frame_image` together returns an unrelated t2v clip AND bills the wrong matrix
841
+ cell ($0.56 vs $0.28 observed). Image-only frame-locks but has no camera control.
842
+ Until the wrapper is fixed, that endpoint can't chain — use Monid's seedance-2.0.
843
+ (The model itself is fine — the same prompt+image via Higgsfield `minimax_hailuo`
844
+ frame-locks.)
845
+ - **Monid billing surprises** → matrix/token-priced endpoints bill by selector match:
846
+ pass every selector field explicitly (`model`, `resolution`, `duration`) and read
847
+ `cost.value` off the run result after each clip. A big base64 field in any body can
848
+ also bounce as an HTML error page ("Unexpected token '<'") — another reason frames
849
+ travel by sfs URL.
850
+ - **Monid schema changed since last build** → it happens (seedance was t2v-only until
851
+ late July 2026, then gained first/last-frame support). `monid inspect` before each
852
+ build; re-run the Step 4 qualification probes when the Input schema differs from
853
+ what pipeline.md documents.
854
+ - **Codex stills hang at "Reading additional input from stdin..."** → parallel
855
+ `codex exec` calls launched from one script share the parent's stdin; one wins it,
856
+ the rest block forever (observed: 1 of 3 completed, 2 hung, the second batch never
857
+ started). Always append `< /dev/null` to every backgrounded `codex exec` — the
858
+ pipeline's `gen_still_codex` has it; keep it if you adapt the command.
859
+ - **White-box scenes** → `gpt_image_2` returns a solid bg; either match the page bg to it
860
+ or knock it out (Step 3).
861
+ - **bash 3.2** on macOS → no associative arrays in scripts.
862
+ - **Connector grabs the wrong scene's frames** (or errors on a frame that doesn't exist
863
+ yet) → the array loop ran in **zsh** (macOS default interactive shell), where arrays are
864
+ 1-indexed, not bash's 0-indexed. Keep every array-driven chain step in a `#!/bin/bash`
865
+ script run via `bash script.sh` — never inline array loops in the interactive shell.
866
+
867
+ ## References
868
+
869
+ - `references/prompts.md` — the intake checklist, style-preamble pattern, and every
870
+ prompt template (scene still, dive, connector) with fill-in slots.
871
+ - `references/pipeline.md` — copy-paste batch scripts for the whole run (generate →
872
+ extract frames → connectors → encode → mobile encode), bash-3.2-safe.
873
+ - `references/scrub-engine.js` — the portable, config-driven scrub engine (builds DOM +
874
+ injects CSS; blob-seek, lazy load, seam crossfade, copy, route rail, reduced-motion, and
875
+ phone hardening: mobile encodes, seek-coalescing, iOS priming, safe-area, no-jump resize).
876
+ - `references/index-template.html` — a minimal standalone page that mounts the engine.
877
+ - `references/knockout.py` — border-connected background knockout for floating scenes.