@amaster.ai/pi-video-gen 0.1.7 → 0.1.8

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@amaster.ai/pi-video-gen",
3
- "version": "0.1.7",
3
+ "version": "0.1.8",
4
4
  "description": "Pi extension for AI video generation plus local video composition: lossless clip concat and mixed image/video timelines with overlays, TTS, soft or burned subtitles, source audio, BGM, and bundled LGPL/GPL FFmpeg runtimes.",
5
5
  "keywords": [
6
6
  "pi-package",
@@ -53,14 +53,14 @@
53
53
  "dependencies": {
54
54
  "sharp": "^0.34.0",
55
55
  "msedge-tts": "^2.0.7",
56
- "@amaster.ai/pi-shared": "0.1.7"
56
+ "@amaster.ai/pi-shared": "0.1.8"
57
57
  },
58
58
  "optionalDependencies": {
59
- "@amaster.ai/pi-video-gen-ffmpeg-darwin-arm64": "0.1.7",
60
- "@amaster.ai/pi-video-gen-ffmpeg-darwin-x64": "0.1.7",
61
- "@amaster.ai/pi-video-gen-ffmpeg-linux-x64": "0.1.7",
62
- "@amaster.ai/pi-video-gen-ffmpeg-linux-arm64": "0.1.7",
63
- "@amaster.ai/pi-video-gen-ffmpeg-win32-x64": "0.1.7"
59
+ "@amaster.ai/pi-video-gen-ffmpeg-linux-arm64": "0.1.8",
60
+ "@amaster.ai/pi-video-gen-ffmpeg-darwin-arm64": "0.1.8",
61
+ "@amaster.ai/pi-video-gen-ffmpeg-darwin-x64": "0.1.8",
62
+ "@amaster.ai/pi-video-gen-ffmpeg-linux-x64": "0.1.8",
63
+ "@amaster.ai/pi-video-gen-ffmpeg-win32-x64": "0.1.8"
64
64
  },
65
65
  "peerDependencies": {
66
66
  "@earendil-works/pi-ai": ">=0.80.10",
@@ -8,9 +8,10 @@ description: "Video creation and local composition: join existing clips or rende
8
8
  This skill orchestrates four published flows:
9
9
 
10
10
  - **C0 local concat** — `video_compose` joins compatible existing MP4 clips.
11
- - **Timeline local render** — `video_compose` turns images/screenshots and
12
- existing video clips into a video with overlays, TTS, motion, transitions,
13
- soft or burned subtitles, source audio, and optional BGM.
11
+ - **Timeline render** — `video_compose` locally turns images/screenshots and
12
+ existing video clips into a video with overlays, motion, transitions, soft or
13
+ burned subtitles, source audio, and optional BGM. Narration is an optional
14
+ network feature.
14
15
  - **Single AI clip** — one paid `video_generate` call.
15
16
  - **Shot-book AI film** — author a shot book, generate frames with
16
17
  `image_generate`, then ONE paid `video_render` call renders and stitches.
@@ -27,7 +28,7 @@ pi-image-gen's active model.
27
28
  | Existing local mp4 clips to join | `video_compose` (C0 — lossless, local, no paid models) |
28
29
  | One AI-generated moving shot | `video_generate` |
29
30
  | Multi-shot film with keyframes | `video_render` |
30
- | Promo/explainer from images, screenshots & clips | `video_compose` (TimelineSpec — mixed media + overlays + TTS + transitions, local render, near-zero cost) |
31
+ | Promo/explainer from images, screenshots & clips | `video_compose` (TimelineSpec — local mixed-media render; optional network TTS) |
31
32
 
32
33
  1. **Local flows stop here.** For C0 follow §A0; for Timeline follow §A1. Do not
33
34
  run the AI preflight, shot-book steps, or paid confirmation gates below.
@@ -94,8 +95,10 @@ pi-image-gen's active model.
94
95
 
95
96
  ## A1. Timeline compose (`video_compose` with `timeline-input.json`)
96
97
 
97
- Use for promos/explainers from still images and existing clips. Costs ~0
98
- (Edge TTS is free, render is local) — prefer it over AI video for this job type.
98
+ Use for promos/explainers from still images and existing clips. Media rendering
99
+ is local and uses no paid video model. If narration is requested, disclose that
100
+ the explicit `edge-tts:<voice>` option sends narration text to Microsoft before
101
+ using it.
99
102
 
100
103
  1. **Collect existing images/screenshots/clips first**, and use `image_generate`
101
104
  only for missing visual material. Keep all source media outside the job
@@ -107,7 +110,7 @@ Use for promos/explainers from still images and existing clips. Costs ~0
107
110
  {
108
111
  "title": "产品宣传片",
109
112
  "output": { "resolution": "1920x1080", "fps": 25, "codec": "h264" },
110
- "voice": "edge-tts:zh-CN-YunyangNeural", // default; free, no key
113
+ "voice": "edge-tts:zh-CN-YunyangNeural", // explicit opt-in: narration text is sent to Microsoft
111
114
  "ttsFailureMode": "fail", // or "silent-subtitles" only after the user accepts that degradation
112
115
  "subtitles": { "mode": "burn", "fontSize": 36,
113
116
  "textColor": "#ffffff", "backgroundColor": "#000000", "backgroundOpacity": 0.55 },
@@ -134,15 +137,18 @@ Use for promos/explainers from still images and existing clips. Costs ~0
134
137
  ```
135
138
  2. **Chinese text NEVER comes from an image model** — titles/subtitles go in
136
139
  `overlay` and are rendered locally via SVG (no garbled CJK).
137
- 3. **Cost confirmation is unnecessary** (local compute), but still show the
138
- segment count and total planned duration before calling `video_compose`.
140
+ 3. **Paid-model confirmation is unnecessary** for the local media render, but
141
+ still show the segment count and total planned duration before calling
142
+ `video_compose`. Obtain explicit agreement before sending narration text to
143
+ Edge TTS.
139
144
  4. Every segment contains exactly one of `image` or `video`. Video segments
140
145
  are normalized to the output resolution/fps, may be trimmed/scaled, and
141
146
  mix their source audio with narration before optional BGM. Video source
142
147
  audio without a stream degrades to silence; `sourceAudio.muted: true` or
143
148
  `volume: 0` disables it. A video's numeric `durationSec` is its fixed trim
144
149
  window; narration that does not fit is rejected instead of extending it.
145
- 5. Narration uses Edge TTS (free). Measured audio duration drives image
150
+ 5. Narration uses Edge TTS only with an explicit `edge-tts:<voice>` selection;
151
+ its text is sent to Microsoft. Measured audio duration drives image
146
152
  `durationSec: "auto"`; subtitles use each segment's actual video timing.
147
153
  `subtitles.mode` defaults to `"soft"` (`mov_text`); `"burn"` renders the
148
154
  configured font/color/background directly into each narrated segment.
@@ -270,13 +276,17 @@ Write `<outputDir>/<jobId>/render-input.json` (jobId: letters/digits/dash/unders
270
276
  "shots": [{
271
277
  "id": "s1",
272
278
  "videoPrompt": "<motion> + <audio cues>",
273
- "firstFramePath": "/abs/path/from/assets.json.png",
274
- "lastFramePath": "/abs/optional.png",
279
+ "firstFramePath": "<project>/path/from/assets.json.png",
280
+ "lastFramePath": "<project>/optional.png",
275
281
  "durationSec": 5
276
282
  }]
277
283
  }
278
284
  ```
279
285
 
286
+ Every reference-frame path must resolve to a regular png/jpg/webp file inside
287
+ the session cwd. Absolute paths are accepted only when they remain inside that
288
+ approved project directory; symlinks and outside paths are rejected.
289
+
280
290
  Then call `video_render` with that path. Interrupted? Call it again with the
281
291
  same path — it resumes. If an ambiguous submit is reported, do not delete a
282
292
  shot or call render again blindly: run `/video-gen recover <jobId>`, check the