sonilo-cli 0.12.0__tar.gz → 0.13.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,3 +1,19 @@
1
+ Metadata-Version: 2.5
2
+ Name: sonilo-cli
3
+ Version: 0.13.0
4
+ Summary: Command-line interface for the Sonilo API: generate music and sound effects from text or video
5
+ Project-URL: Repository, https://github.com/sonilo-ai/sonilo-python
6
+ Author: Sonilo AI
7
+ License-Expression: MIT
8
+ License-File: LICENSE
9
+ Keywords: ai,cli,music,sfx,sonilo,text-to-music,video-to-music
10
+ Requires-Python: >=3.9
11
+ Requires-Dist: sonilo<0.15,>=0.14.0
12
+ Provides-Extra: dev
13
+ Requires-Dist: pytest>=8; extra == 'dev'
14
+ Requires-Dist: respx>=0.21; extra == 'dev'
15
+ Description-Content-Type: text/markdown
16
+
1
17
  # sonilo-cli
2
18
 
3
19
  Command-line interface for the [Sonilo API](https://github.com/sonilo-ai/sonilo-python) — generate music and sound effects from text or video.
@@ -82,6 +98,8 @@ production sign-in coexist without overwriting each other.
82
98
  sonilo audio-ducking --voice interview.mp4 --music-url https://example.com/bed.wav
83
99
  # ducks the existing music bed under the voice; a video voice comes back
84
100
  # as a new .mp4 with the ducked mix muxed in
101
+ sonilo video-analysis --video clip.mp4 --variants 2
102
+ # prints a creative brief as JSON; generates nothing
85
103
  sonilo dubbing --video-url https://example.com/clip.mp4 --languages es,fr --output dubbed.mp4
86
104
  # writes dubbed.es.mp4 and dubbed.fr.mp4
87
105
  sonilo tasks get <task-id>
@@ -225,6 +243,23 @@ call instead.)
225
243
  there, so the CLI rejects a local video file up front rather than let it be mishandled silently.
226
244
  - Each input is capped at 360 seconds server-side.
227
245
 
246
+ ### Video analysis
247
+
248
+ `video-analysis` analyzes a video and prints a **creative brief** for scoring it. It is the one
249
+ command that produces no media file — nothing is generated:
250
+
251
+ sonilo video-analysis --video clip.mp4 --prompt "focus on the chase" --variants 2
252
+
253
+ - The brief goes to **stdout as JSON** so it can be piped into another tool: `segments` (a
254
+ time-aligned section plan) and `variations` (one ready-to-use generation prompt each). Pass
255
+ `--output brief.json` to write it to a file instead.
256
+ - `--variants` is 1-5 (default 1) and is **billed per brief**.
257
+ - Source videos may be at most 600 seconds long, and billing has a 10-second floor.
258
+ - Feed a variation's prompt straight into the next command:
259
+
260
+ sonilo video-analysis --video clip.mp4 --output brief.json
261
+ sonilo video-to-music --video clip.mp4 --prompt "$(jq -r '.variations[0].prompt' brief.json)"
262
+
228
263
  ### Dubbing
229
264
 
230
265
  `dubbing` dubs a video into one or more target languages in a single async call:
@@ -251,7 +286,7 @@ required:
251
286
 
252
287
  | Free runs | Endpoints |
253
288
  | --- | --- |
254
- | 2 each | text-to-music, text-to-sfx, audio-ducking |
289
+ | 2 each | text-to-music, text-to-sfx, audio-ducking, video-analysis |
255
290
  | 1 each | video-to-music, video-to-sfx, video-to-video-music, video-to-video-sfx, video-to-sound, video-to-video-sound |
256
291
  | 0 | dubbing |
257
292
 
@@ -1,19 +1,3 @@
1
- Metadata-Version: 2.5
2
- Name: sonilo-cli
3
- Version: 0.12.0
4
- Summary: Command-line interface for the Sonilo API: generate music and sound effects from text or video
5
- Project-URL: Repository, https://github.com/sonilo-ai/sonilo-python
6
- Author: Sonilo AI
7
- License-Expression: MIT
8
- License-File: LICENSE
9
- Keywords: ai,cli,music,sfx,sonilo,text-to-music,video-to-music
10
- Requires-Python: >=3.9
11
- Requires-Dist: sonilo<0.14,>=0.13.0
12
- Provides-Extra: dev
13
- Requires-Dist: pytest>=8; extra == 'dev'
14
- Requires-Dist: respx>=0.21; extra == 'dev'
15
- Description-Content-Type: text/markdown
16
-
17
1
  # sonilo-cli
18
2
 
19
3
  Command-line interface for the [Sonilo API](https://github.com/sonilo-ai/sonilo-python) — generate music and sound effects from text or video.
@@ -98,6 +82,8 @@ production sign-in coexist without overwriting each other.
98
82
  sonilo audio-ducking --voice interview.mp4 --music-url https://example.com/bed.wav
99
83
  # ducks the existing music bed under the voice; a video voice comes back
100
84
  # as a new .mp4 with the ducked mix muxed in
85
+ sonilo video-analysis --video clip.mp4 --variants 2
86
+ # prints a creative brief as JSON; generates nothing
101
87
  sonilo dubbing --video-url https://example.com/clip.mp4 --languages es,fr --output dubbed.mp4
102
88
  # writes dubbed.es.mp4 and dubbed.fr.mp4
103
89
  sonilo tasks get <task-id>
@@ -241,6 +227,23 @@ call instead.)
241
227
  there, so the CLI rejects a local video file up front rather than let it be mishandled silently.
242
228
  - Each input is capped at 360 seconds server-side.
243
229
 
230
+ ### Video analysis
231
+
232
+ `video-analysis` analyzes a video and prints a **creative brief** for scoring it. It is the one
233
+ command that produces no media file — nothing is generated:
234
+
235
+ sonilo video-analysis --video clip.mp4 --prompt "focus on the chase" --variants 2
236
+
237
+ - The brief goes to **stdout as JSON** so it can be piped into another tool: `segments` (a
238
+ time-aligned section plan) and `variations` (one ready-to-use generation prompt each). Pass
239
+ `--output brief.json` to write it to a file instead.
240
+ - `--variants` is 1-5 (default 1) and is **billed per brief**.
241
+ - Source videos may be at most 600 seconds long, and billing has a 10-second floor.
242
+ - Feed a variation's prompt straight into the next command:
243
+
244
+ sonilo video-analysis --video clip.mp4 --output brief.json
245
+ sonilo video-to-music --video clip.mp4 --prompt "$(jq -r '.variations[0].prompt' brief.json)"
246
+
244
247
  ### Dubbing
245
248
 
246
249
  `dubbing` dubs a video into one or more target languages in a single async call:
@@ -267,7 +270,7 @@ required:
267
270
 
268
271
  | Free runs | Endpoints |
269
272
  | --- | --- |
270
- | 2 each | text-to-music, text-to-sfx, audio-ducking |
273
+ | 2 each | text-to-music, text-to-sfx, audio-ducking, video-analysis |
271
274
  | 1 each | video-to-music, video-to-sfx, video-to-video-music, video-to-video-sfx, video-to-sound, video-to-video-sound |
272
275
  | 0 | dubbing |
273
276
 
@@ -4,13 +4,13 @@ build-backend = "hatchling.build"
4
4
 
5
5
  [project]
6
6
  name = "sonilo-cli"
7
- version = "0.12.0"
7
+ version = "0.13.0"
8
8
  description = "Command-line interface for the Sonilo API: generate music and sound effects from text or video"
9
9
  readme = "README.md"
10
10
  license = "MIT"
11
11
  requires-python = ">=3.9"
12
12
  authors = [{ name = "Sonilo AI" }]
13
- dependencies = ["sonilo>=0.13.0,<0.14"]
13
+ dependencies = ["sonilo>=0.14.0,<0.15"]
14
14
  keywords = ["sonilo", "cli", "music", "sfx", "text-to-music", "video-to-music", "ai"]
15
15
 
16
16
  [project.urls]
@@ -1,3 +1,3 @@
1
- __version__ = "0.12.0"
1
+ __version__ = "0.13.0"
2
2
 
3
3
  __all__ = ["__version__"]
@@ -510,6 +510,51 @@ def cmd_audio_ducking(client: Sonilo, args: argparse.Namespace) -> None:
510
510
  _wrote(path, path.stat().st_size)
511
511
 
512
512
 
513
+ def _analysis_payload(result: Any) -> dict:
514
+ """Flatten a VideoAnalysisResult back into the API's own envelope shape.
515
+
516
+ Deliberately re-emits the wire format rather than dumping the dataclass:
517
+ this output is meant to be piped into another tool (or read by an agent),
518
+ and it should look identical to what GET /v1/tasks returned. None-valued
519
+ accounting fields are dropped so a brief stays readable.
520
+ """
521
+ payload: dict = {
522
+ "task_id": result.task_id,
523
+ "status": result.status,
524
+ "segments": [
525
+ {"start": s.start, "end": s.end, "label": s.label, "prompt": s.prompt}
526
+ for s in result.segments
527
+ ],
528
+ "variations": [{"prompt": v.prompt} for v in result.variations],
529
+ }
530
+ for key in ("variants_num", "duration_seconds", "cost"):
531
+ value = getattr(result, key, None)
532
+ if value is not None:
533
+ payload[key] = value
534
+ return payload
535
+
536
+
537
+ def cmd_video_analysis(client: Sonilo, args: argparse.Namespace) -> None:
538
+ """video-analysis is the one command that produces no media file. The
539
+ brief goes to stdout so it can be piped straight into the next tool;
540
+ --output is the opt-in for keeping a copy on disk."""
541
+ result = client.video_analysis.analyze(
542
+ video=args.video, video_url=args.video_url,
543
+ prompt=args.prompt, variants_num=args.variants,
544
+ timeout=args.timeout,
545
+ )
546
+ if not result.variations:
547
+ _fail("task succeeded but returned no creative brief")
548
+ payload = _analysis_payload(result)
549
+ if args.output is None:
550
+ _print_json(payload)
551
+ return
552
+ path = Path(args.output)
553
+ path.parent.mkdir(parents=True, exist_ok=True)
554
+ path.write_text(json.dumps(payload, indent=2) + "\n")
555
+ _wrote(path, path.stat().st_size)
556
+
557
+
513
558
  # Matched to the dubbing backend's own ceiling: it polls its pipeline for up
514
559
  # to 7200s (2 hours), so anything shorter abandons a job the user has already
515
560
  # been charged for. The SDK's generic DEFAULT_WAIT_TIMEOUT of 600s is far too
@@ -858,6 +903,32 @@ def build_parser() -> argparse.ArgumentParser:
858
903
  )
859
904
  p_duck.set_defaults(func=cmd_audio_ducking)
860
905
 
906
+ p_va = sub.add_parser(
907
+ "video-analysis",
908
+ help="Analyze a video and print a creative brief for scoring it",
909
+ )
910
+ _add_global(p_va)
911
+ _add_video_source(p_va)
912
+ p_va.add_argument(
913
+ "--prompt", default=None,
914
+ help="Optional guidance for the analysis, e.g. 'focus on the chase'.",
915
+ )
916
+ p_va.add_argument(
917
+ "--variants", type=int, default=None,
918
+ help="How many independent briefs to author for the same video (1-5). "
919
+ "Billed per brief. Default: 1",
920
+ )
921
+ p_va.add_argument(
922
+ "--output", default=None,
923
+ help="Write the brief to this .json file instead of printing it to stdout.",
924
+ )
925
+ p_va.add_argument(
926
+ "--timeout", type=float, default=600.0,
927
+ help="Give up waiting after this many seconds. Default: 600. A timed-out "
928
+ "task may still finish — resume it with `sonilo tasks get <task-id>`.",
929
+ )
930
+ p_va.set_defaults(func=cmd_video_analysis)
931
+
861
932
  p_dub = sub.add_parser("dubbing", help="Dub a video into other languages")
862
933
  _add_global(p_dub)
863
934
  _add_video_source(p_dub)
@@ -1346,3 +1346,92 @@ def test_empty_env_var_falls_through_to_the_credential(monkeypatch):
1346
1346
  )
1347
1347
  main(["account"])
1348
1348
  assert route.calls.last.request.headers["authorization"] == "Bearer sk-stored"
1349
+
1350
+
1351
+ ANALYSIS_ACK = {"task_id": "va1", "status": "processing"}
1352
+ ANALYSIS_BODY = {
1353
+ "task_id": "va1",
1354
+ "type": "video_analysis",
1355
+ "status": "succeeded",
1356
+ "variants_num": 2,
1357
+ "segments": [
1358
+ {"start": 0, "end": 12, "label": "intro", "prompt": "sparse piano"},
1359
+ ],
1360
+ "variations": [
1361
+ {"prompt": "cinematic strings, 90bpm"},
1362
+ {"prompt": "lo-fi hip hop, warm keys"},
1363
+ ],
1364
+ "duration_seconds": 30.0,
1365
+ "cost": 0.24,
1366
+ }
1367
+
1368
+
1369
+ def _mock_analysis():
1370
+ respx.post(f"{BASE}/v1/video-analysis").mock(
1371
+ return_value=httpx.Response(202, json=ANALYSIS_ACK)
1372
+ )
1373
+ respx.get(f"{BASE}/v1/tasks/va1").mock(
1374
+ return_value=httpx.Response(200, json=ANALYSIS_BODY)
1375
+ )
1376
+
1377
+
1378
+ @respx.mock
1379
+ def test_video_analysis_prints_the_brief_as_json(capsys):
1380
+ """The result is a brief, not a file: it goes to stdout so it can be
1381
+ piped, and nothing is written to disk unless --output says so."""
1382
+ _mock_analysis()
1383
+ run(["video-analysis", "--video-url", "https://x/v.mp4"])
1384
+ out = json.loads(capsys.readouterr().out)
1385
+ assert out["task_id"] == "va1"
1386
+ assert out["segments"] == [
1387
+ {"start": 0, "end": 12, "label": "intro", "prompt": "sparse piano"}
1388
+ ]
1389
+ assert [v["prompt"] for v in out["variations"]] == [
1390
+ "cinematic strings, 90bpm",
1391
+ "lo-fi hip hop, warm keys",
1392
+ ]
1393
+
1394
+
1395
+ @respx.mock
1396
+ def test_video_analysis_sends_prompt_and_variants():
1397
+ _mock_analysis()
1398
+ route = respx.routes[0]
1399
+ run([
1400
+ "video-analysis",
1401
+ "--video-url", "https://x/v.mp4",
1402
+ "--prompt", "focus on the chase",
1403
+ "--variants", "2",
1404
+ ])
1405
+ body = unquote_plus(route.calls.last.request.content.decode())
1406
+ assert "prompt=focus on the chase" in body
1407
+ assert "variants_num=2" in body
1408
+
1409
+
1410
+ @respx.mock
1411
+ def test_video_analysis_omits_unset_optionals():
1412
+ _mock_analysis()
1413
+ route = respx.routes[0]
1414
+ run(["video-analysis", "--video-url", "https://x/v.mp4"])
1415
+ body = unquote_plus(route.calls.last.request.content.decode())
1416
+ assert "prompt" not in body
1417
+ assert "variants_num" not in body
1418
+
1419
+
1420
+ @respx.mock
1421
+ def test_video_analysis_output_writes_the_brief_to_a_file(tmp_path, capsys):
1422
+ _mock_analysis()
1423
+ out = tmp_path / "brief.json"
1424
+ run(["video-analysis", "--video-url", "https://x/v.mp4", "--output", str(out)])
1425
+ written = json.loads(out.read_text())
1426
+ assert written["variations"][0]["prompt"] == "cinematic strings, 90bpm"
1427
+ # With --output the brief goes to the file, not to stdout.
1428
+ assert "Wrote" in capsys.readouterr().out
1429
+
1430
+
1431
+ def test_video_analysis_requires_a_video_source(capsys):
1432
+ with pytest.raises(SystemExit) as exc:
1433
+ main(["--api-key", "sk-test", "video-analysis"])
1434
+ assert exc.value.code == 1
1435
+ # Asserted on the message, not just the exit code: an unknown command
1436
+ # also exits 1, so the code alone would pass before the command exists.
1437
+ assert "--video" in capsys.readouterr().err
File without changes
File without changes