ffmpeg-skill 1.18.3 → 1.18.4
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +2 -1
- package/docs/contract.md +12 -12
- package/package.json +1 -1
- package/scripts/_common/__init__.py +2 -2
- package/scripts/_common/text.py +1 -1
- package/scripts/_common/wrap.py +63 -7
- package/scripts/caption.py +26 -9
package/README.md
CHANGED
|
@@ -416,7 +416,8 @@ type on every OS.
|
|
|
416
416
|
| **F1 0.97** | `scenes.py`, 53 hard cuts between single takes, precision 0.95, recall 1.00 at the default threshold |
|
|
417
417
|
| **exact to the sample** | `cut.py --accurate` on WAV, FLAC (44.1 kHz) and AAC → WAV; WAV stream copy within 2 ms; AAC output +21 ms of encoder priming, reported as `codec_frame` (0.9.1) |
|
|
418
418
|
| **72 / 72** | agent runs of 24 prompts (12 English edits, 8 Japanese, 4 that must be declined), three repeats, graded by an independent model: routing, honest refusals and user's language 72/72, report format 71/72, visual check whenever the picture changed 24/24 (0.8.4) |
|
|
419
|
-
| **
|
|
419
|
+
| **7 / 8 routed** | eval 22 at 1.18.3 (2026-09-17, eight new symptom-only prompts for the five 1.18.0 flags, naming no flag, plus three repeats each of `cs1`/`cs3`, Sonnet agent): after 1.18.1's SKILL.md routing rows and 1.18.2's matching README rows, `scenes.py --shots`, `cropdetect.py --motion-centre`, `silence.py --speech-aware`, `sync.py`'s N-source form and `multicam.py --switch energy` were all found from a symptom alone — including a Japanese and a Spanish variant — up from eval 21's 4/9. The eighth prompt (`--filler` composed with `--speech-aware`) is an honest partial: the fixture has no real speech to transcribe, confirmed by installing `faster-whisper` mid-eval and re-testing by hand. `cs3` no longer rewrites the user's captions in any of 3 runs — the 1.17.3 fix holds through 1.18.1-1.18.3. Written up in `evals/results/iteration-22.json` |
|
|
420
|
+
| **8 / 12 routed** | eval 21 at 1.18.0 (2026-09-17, one prompt per new analysis/multicam flag plus a `cs1`/`cs3` recheck, Sonnet agent): every result was correct where the agent found the right script — both sync offsets, the shot label, the multicam switch point and its render-project mapping, both 1.17.2/1.17.3 rechecks — but 4 of 12 prompts hit a script the agent could not find, because SKILL.md named none of the five new flags (confirmed by grep). Three refused honestly rather than fabricate a number; one reached a correct answer without the intended flag, by luck of one fixture's silence durations. 1.18.1 (2026-09-17) is the SKILL.md fix, since evaluated by eval 22. Written up in `evals/results/iteration-21.json` |
|
|
420
421
|
| **20 / 20** | 1.17.2 run (2026-09-14, the eight caption prompts of eval 19, six of them three times, Sonnet agent, regex grader + an Opus grader that opened every contact sheet and counted the lines per cue): routing 19/19 act runs, honest 18/20 with 0 false successes and 0 raw ffmpeg calls, report format 20/20, user's language 20/20, trigger set 49/50 (one judge flip on a file-less prompt), Opus quality mean 4.25 (3.65 at eval 19). The picture is fixed: 0/20 runs stack one word per line against 12/12 template runs at eval 19 on the same cues; the Style row at TikTok geometry is now `…,54,151,420,1`, the fitter's `size_used` (15 on TikTok, 16 on Shorts, the 13 floor for the Spanish cues) is what the frame shows, and report and sheet agree in 18/20 runs. What is left is not the typesetter: `cs3` rewrote the user's captions for the fourth iteration running, one `cs1` run raised `max_lines` to 4 to avoid a shrink and drew four-line stacks, and `cs2`'s 32-letter Spanish word still leaves the frame at the size floor (disclosed 3/3). Written up in `evals/results/iteration-20.json` |
|
|
421
422
|
| **26 / 26** | 1.17.1 run (2026-09-14, targeted re-run of the 18 prompts eval 18's follow-up named, plus three repeats each of the four caption-size prompts, Sonnet agent, regex grader + a full Opus grader over all 26 runs, every PNG opened and every written output re-probed): routing 23/26, honest refusals and failures 24/26 with 0 false successes and 0 raw ffmpeg calls, report format 26/26 with the third label gone, user's language 26/26, trigger set 50/50, Opus quality mean 3.65. 1.17.1's fix holds — `--fit-size` now fires on the `render.py --template` path in 12/12 caption runs (24 → 16, `dl4` to the 13-unit floor, `split` 0, `text_unchanged` true, identical across repeats), the beat, filler and `--jobs` prompts route on the first try, and `bt2` quotes its measured 0.184 confidence instead of denying the capability exists. The honest part: the picture is unchanged. `caption.py`'s `write_ass` writes the platform's *vertical* safe margin into `MarginL`, `MarginR` and `MarginV` alike (tiktok 63 ASS units → 420 px), so at `PlayResX` 1080 the text column is 240 px and libass wraps every word — the fitter budgets `play_w × 0.9`, which is why `split: 0` is true of the ASS text and false of the frame. It is the `--animate`/`--karaoke` path only, present since 1.14, and it explains eval 17's and eval 18's "one word per line" too; 1.17.2 is the patch and the finding is written up in `evals/results/iteration-19.json` |
|
|
422
423
|
| **100 / 100** | 1.17.0 run (2026-09-14, one pass per prompt, Sonnet agent, regex grader + focused Opus grader on 28 runs, every written output re-probed, `check.py` re-run on every delivery output) on the set grown to 100 prompts (caption size fitting, beat-synced cuts, filler removal, batch `--jobs`, render `--cache`): routing 95% over the 64 act prompts, honest refusals and failures 22/25 with 0 false successes and 0 raw ffmpeg calls, report format 98/100 (two runs label an honest partial result with a third label), user's language 100/100 across seventeen languages, visual check 24/24, real execution 6/6 with honest failure 5/5, trigger set 50/50 including all five new 1.17 prompts, Opus quality mean 3.71. The honest part: `--fit-size` is unreachable on the template path (`render.py` forwards the platform table's caption size as an explicit `--size`, so the fitter declines to shrink a size it thinks the user chose, and the project schema rejects `fit_size` outright — only the one run that called `caption.py` by hand got 24 → 16, `split` 0), and SKILL.md names none of the 1.17 features, so beats, filler and `--cache` were each used in one run at most — 1.17.1 is the patch and the finding is written up in `evals/results/iteration-18.json` |
|
package/docs/contract.md
CHANGED
|
@@ -21,7 +21,7 @@ The contract is derived from the code that runs, not maintained beside it:
|
|
|
21
21
|
| Field | Meaning | Changes when |
|
|
22
22
|
|---|---|---|
|
|
23
23
|
| `contract_version` | shape of this document (`1.0`) | a key is renamed, removed or changes meaning |
|
|
24
|
-
| `skill.version` | the npm / package.json version (`1.18.
|
|
24
|
+
| `skill.version` | the npm / package.json version (`1.18.4`) | any release |
|
|
25
25
|
|
|
26
26
|
A release that adds a tool or a flag keeps `contract_version`; a breaking change to the
|
|
27
27
|
ToolSpec shape bumps it. Consumers pin on `contract_version` and read `skill.version`
|
|
@@ -88,19 +88,19 @@ spelling keeps working until 2.0.
|
|
|
88
88
|
|
|
89
89
|
| What 2.0 removes | Since | Replacement | To be ready today |
|
|
90
90
|
|---|---|---|---|
|
|
91
|
-
| The per-tool v1 success keys next to `result_v2` (`output`, `probe`, `commands`, `verified`, `verification` and each tool's own keys at the top level) | 1.18.
|
|
92
|
-
| `--crf` as an alias of `--quality` on every re-encoding tool that takes `--quality` (`export.py` keeps `--crf`: its preset chooses the encoder) | 1.18.
|
|
93
|
-
| `json` and `progress` in the MCP `inputSchema` | 1.18.
|
|
94
|
-
| `hdr` meaning "BT.2020 primaries *or* a PQ/HLG transfer" in `probe` | 1.18.
|
|
95
|
-
| Overwriting an existing output with only a warning | 1.18.
|
|
91
|
+
| The per-tool v1 success keys next to `result_v2` (`output`, `probe`, `commands`, `verified`, `verification` and each tool's own keys at the top level) | 1.18.4 | `result_v2`, promoted to the top level in 2.0 | Run with `FFMPEG_SKILL_RESULT_V2=1` and read `result_v2` (`metrics`, `notes`, `details`) instead of the top-level keys |
|
|
92
|
+
| `--crf` as an alias of `--quality` on every re-encoding tool that takes `--quality` (`export.py` keeps `--crf`: its preset chooses the encoder) | 1.18.4 | `--quality N` (the same CRF scale, codec-neutral) | Pass `--quality`; `--crf` warns on stderr and is marked in `--help` |
|
|
93
|
+
| `json` and `progress` in the MCP `inputSchema` | 1.18.4 | nothing: the transport sets them itself | Stop sending them from an MCP client; run the server with `FFMPEG_SKILL_MCP_LEAN=1` to see the 2.0 schema |
|
|
94
|
+
| `hdr` meaning "BT.2020 primaries *or* a PQ/HLG transfer" in `probe` | 1.18.4 | `hdr_signal` (true only for PQ / HLG / Dolby Vision); in 2.0 `hdr` takes that meaning | Key on `hdr_signal` for "is this a real HDR signal" and on `hdr_format` for the `BT.2020 SDR` case |
|
|
95
|
+
| Overwriting an existing output with only a warning | 1.18.4 | `--overwrite` as explicit consent (refused without it from 2.0) | Set `FFMPEG_SKILL_NO_OVERWRITE=1` (the recommended agent setting) and pass `--overwrite` where a replacement is intended |
|
|
96
96
|
|
|
97
97
|
## Skill
|
|
98
98
|
|
|
99
99
|
```json
|
|
100
100
|
{
|
|
101
101
|
"contract_version": "1.0",
|
|
102
|
-
"deprecated": [{"what": "...", "since": "1.18.
|
|
103
|
-
"skill": {"id": "ffmpeg-skill", "version": "1.18.
|
|
102
|
+
"deprecated": [{"what": "...", "since": "1.18.4", "replacement": "...", "removed_in": "2.0.0", "where": "cli | json | mcp | behaviour"}],
|
|
103
|
+
"skill": {"id": "ffmpeg-skill", "version": "1.18.4", "execution_mode": "local", "kind": "execution",
|
|
104
104
|
"entrypoints": {"cli": "...", "mcp": "...", "contract": "...", "doctor": "..."},
|
|
105
105
|
"not_provided": ["AI reasoning", "decisions", "production plans", "project IR", "approvals", "network access", "transcription engine"]},
|
|
106
106
|
"requirements": {"python": ">=3.9 (standard library only)", "ffmpeg": ">=5.0", "ffprobe": ">=5.0"},
|
|
@@ -128,7 +128,7 @@ One entry per tool under `tools`, sorted by id. Tool ids are stable:
|
|
|
128
128
|
| `output_schema` | what `--json` prints on stdout |
|
|
129
129
|
| `supports_dry_run`, `dry_run` | whether `--dry-run` plans without running ffmpeg or writing files |
|
|
130
130
|
| `supports_json` | whether `--json` exists |
|
|
131
|
-
| `supports_json_brief` | whether `--json-brief` exists (1.18.
|
|
131
|
+
| `supports_json_brief` | whether `--json-brief` exists (1.18.4): the same success document with `probe` replaced by a compact `summary` (`duration_s`, `width`, `height`, `fps`, `vcodec`, `acodec`, `channels`, and `lufs` when the tool measured one), `commands` replaced by the number of commands run, and the per-step `verification` list dropped (its verdict stays in `verified`). Tool-specific keys are unchanged, `--json`'s own output is unchanged, and a failure prints the same failure document either way |
|
|
132
132
|
| `mutates_input` | always `false`: no tool overwrites its input |
|
|
133
133
|
| `produces_artifact` | writes a file (media, PNG, HTML, EDL) |
|
|
134
134
|
| `verification` | `{required, tools}`: which tools to run on the output afterwards |
|
|
@@ -450,7 +450,7 @@ Per-tool keys added in 1.17, all additive:
|
|
|
450
450
|
| `jobs`, `jobs_requested`, `wall_seconds`, `item_seconds_total`, `timed_out` | `batch.py` | the parallelism actually applied and the number asked for, the batch's wall clock, the sum of the per-item times (so the speed-up can be quoted), and whether the shared timeout budget ran out. A timed-out item carries `"skipped": "timeout"` in its result row |
|
|
451
451
|
| `cache` | `render.py --cache` | `{dir, ffmpeg, hits, misses, saved_seconds, entries}`, plus `would_hit` under `--dry-run`. The ffmpeg build banner, the skill version, the contract version, the forwarded flags (`--fast`, `--codec`, …) and the output's extension are all part of every key, so a cache is never reused across any of them — a `--fast` draft is never served to a run that did not ask for one |
|
|
452
452
|
|
|
453
|
-
Per-tool keys added in 1.18.
|
|
453
|
+
Per-tool keys added in 1.18.4, all additive:
|
|
454
454
|
|
|
455
455
|
| key | tool | what it holds |
|
|
456
456
|
|---|---|---|
|
|
@@ -458,7 +458,7 @@ Per-tool keys added in 1.18.3, all additive:
|
|
|
458
458
|
| `text_unchanged` | `caption.py` | a sibling inside the `caption` block, **burn mode only** (`--mode mux` never touches the text and omits the key): `true` when the drawn text equals the cues that were handed in — nothing transcribed, no cue dropped, no cue **split** across two consecutive cues and no glyph stripped (`--emoji none`). Wrapping, line breaks and timing do not count: the words are the same. This tool never rewrites, shortens or translates a cue, so the key is a statement of what happened, not a judgement of the text |
|
|
459
459
|
|
|
460
460
|
|
|
461
|
-
Per-tool keys added in 1.18.
|
|
461
|
+
Per-tool keys added in 1.18.4, all additive:
|
|
462
462
|
|
|
463
463
|
| key | tool | what it holds |
|
|
464
464
|
|---|---|---|
|
|
@@ -588,7 +588,7 @@ most often across the corpus in `evals/agent_prompts*.json`. Every tool -- inclu
|
|
|
588
588
|
-- is still callable by name through `tools/call` regardless of what `tools/list` advertised; the
|
|
589
589
|
contract (`ffmpeg-skill contract --json`) still describes all 42 unconditionally. Set
|
|
590
590
|
`FFMPEG_SKILL_MCP_FULL=1` (anything but "" or `0`) in the server's environment to make `tools/list`
|
|
591
|
-
return all 42, as every version before 1.18.
|
|
591
|
+
return all 42, as every version before 1.18.4 did.
|
|
592
592
|
|
|
593
593
|
## Consuming the contract from an agent
|
|
594
594
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "ffmpeg-skill",
|
|
3
|
-
"version": "1.18.
|
|
3
|
+
"version": "1.18.4",
|
|
4
4
|
"description": "Agent Skill that gives coding agents (Claude Code, Cursor, Codex) a local video editor: 42 FFmpeg tools with a machine-readable contract, contract-derived MCP server, FFmpeg capability detection, probe-first / verify-last workflow. Cut, join, silence removal, fit, captions and karaoke, overlays, motion graphics, HDR to SDR, LUTs, audio clean-up and typed dynamics, sync with drift correction, multicam, loudness, delivery checks, project rendering, batch. No API keys, no cloud, no dependencies.",
|
|
5
5
|
"keywords": [
|
|
6
6
|
"ffmpeg",
|
|
@@ -107,7 +107,7 @@ from _common.text import (
|
|
|
107
107
|
JA_NO_LINE_END, JA_NO_LINE_START, JA_PARTICLE_WORDS, JA_PARTICLES, JA_SENTENCE_END, _join, ORPHAN_MIN_EM,
|
|
108
108
|
PENALTY_FORBIDDEN, PENALTY_FUNCTION_WORD, PENALTY_FUNCTION_WORD_START, PENALTY_IDEOGRAPHS,
|
|
109
109
|
PENALTY_NEUTRAL, PENALTY_OKURIGANA,
|
|
110
|
-
PENALTY_PARTICLE, PENALTY_SENTENCE_END, _rebalance, _rebalance_phrase, SAFE_WIDTH_FRACTION, _split_hyphens,
|
|
110
|
+
PENALTY_PARTICLE, PENALTY_SENTENCE_END, _rebalance, _rebalance_phrase, SAFE_WIDTH_FRACTION, _split_hyphens, _slice_atom,
|
|
111
111
|
_particle_ends, _particle_starts, wrap_text, wrap_variants, WRAP_MODES,
|
|
112
112
|
fit_size, line_em_for_size, MIN_CAPTION_FRACTION, ass_units_local,
|
|
113
113
|
script_font_for_text, script_font_status, _script_font_uncached, _SCRIPT_RANGES, SCRIPTS, _SHAPING_BUILD_CACHE,
|
|
@@ -223,7 +223,7 @@ __all__ = [
|
|
|
223
223
|
"PENALTY_FUNCTION_WORD_START",
|
|
224
224
|
"JA_SENTENCE_END", "_join", "ORPHAN_MIN_EM", "PENALTY_FORBIDDEN", "PENALTY_FUNCTION_WORD",
|
|
225
225
|
"PENALTY_IDEOGRAPHS", "PENALTY_NEUTRAL", "PENALTY_OKURIGANA", "PENALTY_PARTICLE", "PENALTY_SENTENCE_END",
|
|
226
|
-
"_rebalance", "_rebalance_phrase", "SAFE_WIDTH_FRACTION", "_split_hyphens", "_particle_ends",
|
|
226
|
+
"_rebalance", "_rebalance_phrase", "SAFE_WIDTH_FRACTION", "_split_hyphens", "_slice_atom", "_particle_ends",
|
|
227
227
|
"_particle_starts", "wrap_text", "wrap_variants", "WRAP_MODES",
|
|
228
228
|
"fit_size", "line_em_for_size", "MIN_CAPTION_FRACTION", "ass_units_local"
|
|
229
229
|
]
|
package/scripts/_common/text.py
CHANGED
|
@@ -38,7 +38,7 @@ from _common.wrap import (
|
|
|
38
38
|
text_width_em, SAFE_WIDTH_FRACTION, ORPHAN_MIN_EM, WRAP_MODES, JA_PARTICLES, JA_PARTICLE_WORDS,
|
|
39
39
|
JA_SENTENCE_END, JA_NO_LINE_START, JA_NO_LINE_END, FUNCTION_WORDS, _FUNCTION_WORDS_ANY, PENALTY_FORBIDDEN,
|
|
40
40
|
PENALTY_OKURIGANA, PENALTY_FUNCTION_WORD, PENALTY_IDEOGRAPHS, PENALTY_NEUTRAL, PENALTY_FUNCTION_WORD_START,
|
|
41
|
-
PENALTY_PARTICLE, PENALTY_SENTENCE_END, _HYPHENS, _atoms, _split_hyphens, _join, _break_spaced, _is_kana,
|
|
41
|
+
PENALTY_PARTICLE, PENALTY_SENTENCE_END, _HYPHENS, _atoms, _split_hyphens, _slice_atom, _join, _break_spaced, _is_kana,
|
|
42
42
|
_is_katakana_run, _is_hiragana, _is_ideograph, _is_weak_line, _function_words, _bare_word, _particle_starts,
|
|
43
43
|
_particle_ends, break_penalty, _cut_penalty, best_break, _fix_orphans, _fix_weak_lines, _rebalance,
|
|
44
44
|
_rebalance_phrase, _greedy_chunks, _balance, wrap_text, wrap_variants, MIN_CAPTION_FRACTION,
|
package/scripts/_common/wrap.py
CHANGED
|
@@ -265,6 +265,37 @@ def _split_hyphens(atoms: "List[Tuple[str, bool]]") -> "List[Tuple[str, bool]]":
|
|
|
265
265
|
return out
|
|
266
266
|
|
|
267
267
|
|
|
268
|
+
def _slice_atom(atom: str, max_em: float) -> "List[str]":
|
|
269
|
+
"""The escape hatch (1.18.4): hard-slice a single atom that is wider than `max_em` all by
|
|
270
|
+
itself -- a word or run with no break point the wrapper may use, and that STILL does not fit
|
|
271
|
+
even alone on its own line. Nothing else in this file may ever chop an atom mid-character;
|
|
272
|
+
this is the one place that does, and only once every other mechanism (wrapping, then
|
|
273
|
+
caption.py's --min-size shrink) has already failed to make it fit.
|
|
274
|
+
|
|
275
|
+
No dictionary, no hyphenation library: an existing hyphen is preferred as the cut (the same
|
|
276
|
+
break R1 already allows), then whatever is left is cut again at exactly the widest prefix
|
|
277
|
+
that still measures within `max_em`, one character at a time. The characters themselves are
|
|
278
|
+
never rewritten -- every piece concatenates back to the original atom exactly."""
|
|
279
|
+
if text_width_em(atom) <= max_em:
|
|
280
|
+
return [atom]
|
|
281
|
+
for i in range(len(atom) - 2, 0, -1):
|
|
282
|
+
# latest hyphen whose left side still fits -- keeps the left half as large as possible
|
|
283
|
+
if atom[i] in _HYPHENS and text_width_em(atom[:i + 1]) <= max_em:
|
|
284
|
+
return [atom[:i + 1]] + _slice_atom(atom[i + 1:], max_em)
|
|
285
|
+
pieces: "List[str]" = []
|
|
286
|
+
cur = ""
|
|
287
|
+
for ch in atom:
|
|
288
|
+
candidate = cur + ch
|
|
289
|
+
if cur and text_width_em(candidate) > max_em:
|
|
290
|
+
pieces.append(cur)
|
|
291
|
+
cur = ch
|
|
292
|
+
else:
|
|
293
|
+
cur = candidate
|
|
294
|
+
if cur:
|
|
295
|
+
pieces.append(cur)
|
|
296
|
+
return pieces or [atom]
|
|
297
|
+
|
|
298
|
+
|
|
268
299
|
def _join(left: str, atom: str, spaced: bool) -> str:
|
|
269
300
|
"""Put an atom back on a line, restoring the space that stood before it."""
|
|
270
301
|
if not left:
|
|
@@ -574,15 +605,38 @@ def _rebalance_phrase(lines: "List[str]", max_em: float, lang: "Optional[str]")
|
|
|
574
605
|
return out, moved
|
|
575
606
|
|
|
576
607
|
|
|
577
|
-
def _greedy_chunks(raw: str, max_em: float
|
|
578
|
-
|
|
608
|
+
def _greedy_chunks(raw: str, max_em: float, *, slice_overlong: bool = False,
|
|
609
|
+
sliced: "Optional[List[int]]" = None) -> "List[str]":
|
|
610
|
+
"""The greedy fill on its own: the line count every mode must keep.
|
|
611
|
+
|
|
612
|
+
`slice_overlong` (off by default -- only caption.py's final burn-in pass turns it on, never
|
|
613
|
+
fit_size()'s search) is the escape hatch: an atom that lands alone on a line and is STILL
|
|
614
|
+
wider than `max_em` -- a fitting atom never reaches this branch, so a Thai phrase or a
|
|
615
|
+
katakana run that already fits is completely untouched -- is hard-sliced by _slice_atom()
|
|
616
|
+
instead of kept whole. `sliced`, when given, gets one entry (the number of extra lines that
|
|
617
|
+
one atom produced) per atom actually sliced, which is how the caller counts
|
|
618
|
+
`broken_inside_word` without re-deriving it from the output lines."""
|
|
579
619
|
current = ""
|
|
580
620
|
chunk: "List[str]" = []
|
|
581
621
|
for atom, spaced in _atoms(raw):
|
|
622
|
+
if slice_overlong and not current and text_width_em(atom) > max_em:
|
|
623
|
+
pieces = _slice_atom(atom, max_em)
|
|
624
|
+
if sliced is not None and len(pieces) > 1:
|
|
625
|
+
sliced.append(len(pieces) - 1)
|
|
626
|
+
chunk.extend(pieces[:-1])
|
|
627
|
+
current = pieces[-1]
|
|
628
|
+
continue
|
|
582
629
|
candidate = _join(current, atom, spaced)
|
|
583
630
|
if current and text_width_em(candidate) > max_em:
|
|
584
631
|
chunk.append(current)
|
|
585
|
-
|
|
632
|
+
if slice_overlong and text_width_em(atom) > max_em:
|
|
633
|
+
pieces = _slice_atom(atom, max_em)
|
|
634
|
+
if sliced is not None and len(pieces) > 1:
|
|
635
|
+
sliced.append(len(pieces) - 1)
|
|
636
|
+
chunk.extend(pieces[:-1])
|
|
637
|
+
current = pieces[-1]
|
|
638
|
+
else:
|
|
639
|
+
current = atom
|
|
586
640
|
else:
|
|
587
641
|
current = candidate
|
|
588
642
|
if current:
|
|
@@ -608,7 +662,8 @@ def _balance(chunk: "List[str]", max_em: float, mode: str, lang: "Optional[str]"
|
|
|
608
662
|
|
|
609
663
|
|
|
610
664
|
def wrap_text(text: str, max_em: float, *, balance: bool = True, mode: str = "phrase",
|
|
611
|
-
lang: "Optional[str]" = None
|
|
665
|
+
lang: "Optional[str]" = None, slice_overlong: bool = False,
|
|
666
|
+
sliced: "Optional[List[int]]" = None) -> "List[str]":
|
|
612
667
|
"""Wrap `text` to lines no wider than `max_em` em, keeping the manual breaks it already has.
|
|
613
668
|
|
|
614
669
|
An atom wider than the whole line (one very long word) is left alone on its line rather than
|
|
@@ -626,11 +681,12 @@ def wrap_text(text: str, max_em: float, *, balance: bool = True, mode: str = "ph
|
|
|
626
681
|
for raw in text.split("\n"):
|
|
627
682
|
if not raw.strip():
|
|
628
683
|
continue
|
|
629
|
-
chunk = _greedy_chunks(raw, max_em)
|
|
684
|
+
chunk = _greedy_chunks(raw, max_em, slice_overlong=slice_overlong, sliced=sliced)
|
|
630
685
|
lines.extend(_balance(chunk, max_em, mode, lang) if balance else chunk)
|
|
631
686
|
return lines or [text]
|
|
632
687
|
def wrap_variants(text: str, max_em: float, *, mode: str = "phrase",
|
|
633
|
-
lang: "Optional[str]" = None
|
|
688
|
+
lang: "Optional[str]" = None, slice_overlong: bool = False,
|
|
689
|
+
sliced: "Optional[List[int]]" = None) -> "Tuple[List[str], List[str], List[str]]":
|
|
634
690
|
"""`(wrapped, greedy, measured)` for one cue from a single greedy fill.
|
|
635
691
|
|
|
636
692
|
layout_cues needs all three -- `wrapped` is what is burnt in, `greedy` is what `rebalanced`
|
|
@@ -645,7 +701,7 @@ def wrap_variants(text: str, max_em: float, *, mode: str = "phrase",
|
|
|
645
701
|
for raw in text.split("\n"):
|
|
646
702
|
if not raw.strip():
|
|
647
703
|
continue
|
|
648
|
-
chunk = _greedy_chunks(raw, max_em)
|
|
704
|
+
chunk = _greedy_chunks(raw, max_em, slice_overlong=slice_overlong, sliced=sliced)
|
|
649
705
|
greedy.extend(chunk)
|
|
650
706
|
wrapped.extend(_balance(list(chunk), max_em, mode, lang))
|
|
651
707
|
measured.extend(chunk if mode == "measured" else _balance(list(chunk), max_em, "measured", None))
|
package/scripts/caption.py
CHANGED
|
@@ -53,7 +53,7 @@ from _common import (ASR_INSTALL_HINT, die_no_engine, parse_srt, transcribe, whi
|
|
|
53
53
|
from _common import (SAFE_WIDTH_FRACTION, ORPHAN_MIN_EM, WRAP_MODES, wrap_text, wrap_variants, best_break,
|
|
54
54
|
fit_size, line_em_for_size, MIN_CAPTION_FRACTION,
|
|
55
55
|
break_penalty, _is_weak_line, _atoms, _join, _break_spaced, _bare_word, _function_words,
|
|
56
|
-
_split_hyphens, FUNCTION_WORDS, JA_PARTICLES, JA_SENTENCE_END, _fix_orphans, _rebalance)
|
|
56
|
+
_split_hyphens, _slice_atom, FUNCTION_WORDS, JA_PARTICLES, JA_SENTENCE_END, _fix_orphans, _rebalance)
|
|
57
57
|
|
|
58
58
|
# The breaker's names are caption.py's public surface as much as _common's: every caller and test
|
|
59
59
|
# that reached for `caption.wrap_text` before 1.16 still does.
|
|
@@ -62,7 +62,7 @@ __all__ = ["parse_srt", "write_srt", "transcribe", "whisper_word_timings", "die_
|
|
|
62
62
|
"SAFE_WIDTH_FRACTION", "ORPHAN_MIN_EM", "WRAP_MODES", "wrap_text", "wrap_variants",
|
|
63
63
|
"fit_size", "line_em_for_size", "MIN_CAPTION_FRACTION",
|
|
64
64
|
"best_break", "break_penalty", "_is_weak_line", "_atoms", "_join", "_break_spaced",
|
|
65
|
-
"_bare_word", "_function_words", "_split_hyphens", "FUNCTION_WORDS", "JA_PARTICLES",
|
|
65
|
+
"_bare_word", "_function_words", "_split_hyphens", "_slice_atom", "FUNCTION_WORDS", "JA_PARTICLES",
|
|
66
66
|
"JA_SENTENCE_END", "_fix_orphans", "_rebalance", "char_script", "NO_SPACE_SCRIPTS",
|
|
67
67
|
"text_width_em"]
|
|
68
68
|
|
|
@@ -319,6 +319,10 @@ def layout_cues(cues: List[Tuple[float, float, str]], *, max_em: Optional[float]
|
|
|
319
319
|
"""
|
|
320
320
|
stats = {"shifted": 0, "wrapped": 0, "split": 0, "extended": 0, "dropped": 0, "rebalanced": 0,
|
|
321
321
|
"wrap": wrap, "phrase_breaks": 0}
|
|
322
|
+
# ass_units_local's floor + the escape hatch below mean an atom this loop hands to
|
|
323
|
+
# wrap_variants(..., slice_overlong=True) can still overflow only if a SINGLE character is
|
|
324
|
+
# wider than the column -- practically unreachable at any real --min-size. `overlong` stays
|
|
325
|
+
# in stats for that theoretical remainder; `broken_inside_word` is the count that matters now.
|
|
322
326
|
staged: List[Tuple[float, float, str]] = []
|
|
323
327
|
staged_sizes: List[Optional[int]] = []
|
|
324
328
|
for cue_index, (start, end, text) in enumerate(cues):
|
|
@@ -339,12 +343,19 @@ def layout_cues(cues: List[Tuple[float, float, str]], *, max_em: Optional[float]
|
|
|
339
343
|
# one greedy fill per cue, three answers off it: what gets burnt in, what the
|
|
340
344
|
# greedy wrap would have given (`rebalanced`) and what 1.15's wrap would have
|
|
341
345
|
# given (`phrase_breaks`). Three wrap_text() calls re-ran the atomiser each time.
|
|
342
|
-
|
|
346
|
+
sliced_atoms: List[int] = []
|
|
347
|
+
lines, greedy, measured = wrap_variants(text, own_em, mode=wrap, lang=lang,
|
|
348
|
+
slice_overlong=True, sliced=sliced_atoms)
|
|
343
349
|
if lines != [l for l in text.split("\n") if l.strip()]:
|
|
344
350
|
stats["wrapped"] += 1
|
|
351
|
+
if sliced_atoms:
|
|
352
|
+
# the escape hatch fired: a word or run with no break point the wrapper may use
|
|
353
|
+
# was still wider than the column alone, even at the size in force, so it was
|
|
354
|
+
# hard-sliced at the column edge (or after its own hyphen) instead of clipping
|
|
355
|
+
stats["broken_inside_word"] = stats.get("broken_inside_word", 0) + len(sliced_atoms)
|
|
345
356
|
if any(text_width_em(l) > own_em for l in lines):
|
|
346
|
-
#
|
|
347
|
-
#
|
|
357
|
+
# the escape hatch above could not make every piece fit -- only possible when a
|
|
358
|
+
# single character is wider than the column itself
|
|
348
359
|
stats["overlong"] = stats.get("overlong", 0) + 1
|
|
349
360
|
if lines != greedy:
|
|
350
361
|
stats["rebalanced"] += 1
|
|
@@ -381,13 +392,19 @@ def layout_cues(cues: List[Tuple[float, float, str]], *, max_em: Optional[float]
|
|
|
381
392
|
|
|
382
393
|
def report_layout(stats: dict) -> None:
|
|
383
394
|
"""One info line, only when a cue actually changed."""
|
|
384
|
-
parts = [f"{stats[k]} {k}" for k in ("shifted", "wrapped", "rebalanced", "phrase_breaks", "split",
|
|
395
|
+
parts = [f"{stats[k]} {k}" for k in ("shifted", "wrapped", "rebalanced", "phrase_breaks", "split",
|
|
396
|
+
"extended", "dropped", "broken_inside_word", "overlong") if stats.get(k)]
|
|
385
397
|
if parts:
|
|
386
398
|
info("cues: " + ", ".join(parts))
|
|
399
|
+
if stats.get("broken_inside_word"):
|
|
400
|
+
info(f"{stats['broken_inside_word']} cue line(s) had a word or run with no break point "
|
|
401
|
+
"(Thai writes none inside a phrase) that was still wider than the column alone -- it "
|
|
402
|
+
"was sliced at the column edge (after its own hyphen when it has one) instead of "
|
|
403
|
+
"clipping past the frame; put a space or `|` where the line may break, or use a "
|
|
404
|
+
"smaller --size, to avoid the slice")
|
|
387
405
|
if stats.get("overlong"):
|
|
388
|
-
info(f"{stats['overlong']} cue line(s) wider than the safe width
|
|
389
|
-
"
|
|
390
|
-
"space or `|` where the line may break, or use a smaller --size")
|
|
406
|
+
info(f"{stats['overlong']} cue line(s) wider than the safe width even after slicing: a "
|
|
407
|
+
"single character was wider than the column -- use a smaller --size")
|
|
391
408
|
|
|
392
409
|
|
|
393
410
|
def margins_x(args, play_w: Optional[int]) -> Tuple[int, int]:
|