ffmpeg-skill 1.18.3 → 1.18.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -416,7 +416,8 @@ type on every OS.
416
416
  | **F1 0.97** | `scenes.py`, 53 hard cuts between single takes, precision 0.95, recall 1.00 at the default threshold |
417
417
  | **exact to the sample** | `cut.py --accurate` on WAV, FLAC (44.1 kHz) and AAC → WAV; WAV stream copy within 2 ms; AAC output +21 ms of encoder priming, reported as `codec_frame` (0.9.1) |
418
418
  | **72 / 72** | agent runs of 24 prompts (12 English edits, 8 Japanese, 4 that must be declined), three repeats, graded by an independent model: routing, honest refusals and user's language 72/72, report format 71/72, visual check whenever the picture changed 24/24 (0.8.4) |
419
- | **8 / 12 routed** | eval 21 at 1.18.0 (2026-09-17, one prompt per new analysis/multicam flag plus a `cs1`/`cs3` recheck, Sonnet agent): every result was correct where the agent found the right script both sync offsets, the shot label, the multicam switch point and its render-project mapping, both 1.17.2/1.17.3 rechecksbut 4 of 12 prompts hit a script the agent could not find, because SKILL.md named none of the five new flags (confirmed by grep). Three refused honestly rather than fabricate a number; one reached a correct answer without the intended flag, by luck of one fixture's silence durations. 1.18.1 (2026-09-17) is the SKILL.md fix, not yet re-evaluated. Written up in `evals/results/iteration-21.json` |
419
+ | **7 / 8 routed** | eval 22 at 1.18.3 (2026-09-17, eight new symptom-only prompts for the five 1.18.0 flags, naming no flag, plus three repeats each of `cs1`/`cs3`, Sonnet agent): after 1.18.1's SKILL.md routing rows and 1.18.2's matching README rows, `scenes.py --shots`, `cropdetect.py --motion-centre`, `silence.py --speech-aware`, `sync.py`'s N-source form and `multicam.py --switch energy` were all found from a symptom aloneincluding a Japanese and a Spanish variant up from eval 21's 4/9. The eighth prompt (`--filler` composed with `--speech-aware`) is an honest partial: the fixture has no real speech to transcribe, confirmed by installing `faster-whisper` mid-eval and re-testing by hand. `cs3` no longer rewrites the user's captions in any of 3 runs — the 1.17.3 fix holds through 1.18.1-1.18.3. Written up in `evals/results/iteration-22.json` |
420
+ | **8 / 12 routed** | eval 21 at 1.18.0 (2026-09-17, one prompt per new analysis/multicam flag plus a `cs1`/`cs3` recheck, Sonnet agent): every result was correct where the agent found the right script — both sync offsets, the shot label, the multicam switch point and its render-project mapping, both 1.17.2/1.17.3 rechecks — but 4 of 12 prompts hit a script the agent could not find, because SKILL.md named none of the five new flags (confirmed by grep). Three refused honestly rather than fabricate a number; one reached a correct answer without the intended flag, by luck of one fixture's silence durations. 1.18.1 (2026-09-17) is the SKILL.md fix, since evaluated by eval 22. Written up in `evals/results/iteration-21.json` |
420
421
  | **20 / 20** | 1.17.2 run (2026-09-14, the eight caption prompts of eval 19, six of them three times, Sonnet agent, regex grader + an Opus grader that opened every contact sheet and counted the lines per cue): routing 19/19 act runs, honest 18/20 with 0 false successes and 0 raw ffmpeg calls, report format 20/20, user's language 20/20, trigger set 49/50 (one judge flip on a file-less prompt), Opus quality mean 4.25 (3.65 at eval 19). The picture is fixed: 0/20 runs stack one word per line against 12/12 template runs at eval 19 on the same cues; the Style row at TikTok geometry is now `…,54,151,420,1`, the fitter's `size_used` (15 on TikTok, 16 on Shorts, the 13 floor for the Spanish cues) is what the frame shows, and report and sheet agree in 18/20 runs. What is left is not the typesetter: `cs3` rewrote the user's captions for the fourth iteration running, one `cs1` run raised `max_lines` to 4 to avoid a shrink and drew four-line stacks, and `cs2`'s 32-letter Spanish word still leaves the frame at the size floor (disclosed 3/3). Written up in `evals/results/iteration-20.json` |
421
422
  | **26 / 26** | 1.17.1 run (2026-09-14, targeted re-run of the 18 prompts eval 18's follow-up named, plus three repeats each of the four caption-size prompts, Sonnet agent, regex grader + a full Opus grader over all 26 runs, every PNG opened and every written output re-probed): routing 23/26, honest refusals and failures 24/26 with 0 false successes and 0 raw ffmpeg calls, report format 26/26 with the third label gone, user's language 26/26, trigger set 50/50, Opus quality mean 3.65. 1.17.1's fix holds — `--fit-size` now fires on the `render.py --template` path in 12/12 caption runs (24 → 16, `dl4` to the 13-unit floor, `split` 0, `text_unchanged` true, identical across repeats), the beat, filler and `--jobs` prompts route on the first try, and `bt2` quotes its measured 0.184 confidence instead of denying the capability exists. The honest part: the picture is unchanged. `caption.py`'s `write_ass` writes the platform's *vertical* safe margin into `MarginL`, `MarginR` and `MarginV` alike (tiktok 63 ASS units → 420 px), so at `PlayResX` 1080 the text column is 240 px and libass wraps every word — the fitter budgets `play_w × 0.9`, which is why `split: 0` is true of the ASS text and false of the frame. It is the `--animate`/`--karaoke` path only, present since 1.14, and it explains eval 17's and eval 18's "one word per line" too; 1.17.2 is the patch and the finding is written up in `evals/results/iteration-19.json` |
422
423
  | **100 / 100** | 1.17.0 run (2026-09-14, one pass per prompt, Sonnet agent, regex grader + focused Opus grader on 28 runs, every written output re-probed, `check.py` re-run on every delivery output) on the set grown to 100 prompts (caption size fitting, beat-synced cuts, filler removal, batch `--jobs`, render `--cache`): routing 95% over the 64 act prompts, honest refusals and failures 22/25 with 0 false successes and 0 raw ffmpeg calls, report format 98/100 (two runs label an honest partial result with a third label), user's language 100/100 across seventeen languages, visual check 24/24, real execution 6/6 with honest failure 5/5, trigger set 50/50 including all five new 1.17 prompts, Opus quality mean 3.71. The honest part: `--fit-size` is unreachable on the template path (`render.py` forwards the platform table's caption size as an explicit `--size`, so the fitter declines to shrink a size it thinks the user chose, and the project schema rejects `fit_size` outright — only the one run that called `caption.py` by hand got 24 → 16, `split` 0), and SKILL.md names none of the 1.17 features, so beats, filler and `--cache` were each used in one run at most — 1.17.1 is the patch and the finding is written up in `evals/results/iteration-18.json` |
package/docs/contract.md CHANGED
@@ -21,7 +21,7 @@ The contract is derived from the code that runs, not maintained beside it:
21
21
  | Field | Meaning | Changes when |
22
22
  |---|---|---|
23
23
  | `contract_version` | shape of this document (`1.0`) | a key is renamed, removed or changes meaning |
24
- | `skill.version` | the npm / package.json version (`1.18.3`) | any release |
24
+ | `skill.version` | the npm / package.json version (`1.18.4`) | any release |
25
25
 
26
26
  A release that adds a tool or a flag keeps `contract_version`; a breaking change to the
27
27
  ToolSpec shape bumps it. Consumers pin on `contract_version` and read `skill.version`
@@ -88,19 +88,19 @@ spelling keeps working until 2.0.
88
88
 
89
89
  | What 2.0 removes | Since | Replacement | To be ready today |
90
90
  |---|---|---|---|
91
- | The per-tool v1 success keys next to `result_v2` (`output`, `probe`, `commands`, `verified`, `verification` and each tool's own keys at the top level) | 1.18.3 | `result_v2`, promoted to the top level in 2.0 | Run with `FFMPEG_SKILL_RESULT_V2=1` and read `result_v2` (`metrics`, `notes`, `details`) instead of the top-level keys |
92
- | `--crf` as an alias of `--quality` on every re-encoding tool that takes `--quality` (`export.py` keeps `--crf`: its preset chooses the encoder) | 1.18.3 | `--quality N` (the same CRF scale, codec-neutral) | Pass `--quality`; `--crf` warns on stderr and is marked in `--help` |
93
- | `json` and `progress` in the MCP `inputSchema` | 1.18.3 | nothing: the transport sets them itself | Stop sending them from an MCP client; run the server with `FFMPEG_SKILL_MCP_LEAN=1` to see the 2.0 schema |
94
- | `hdr` meaning "BT.2020 primaries *or* a PQ/HLG transfer" in `probe` | 1.18.3 | `hdr_signal` (true only for PQ / HLG / Dolby Vision); in 2.0 `hdr` takes that meaning | Key on `hdr_signal` for "is this a real HDR signal" and on `hdr_format` for the `BT.2020 SDR` case |
95
- | Overwriting an existing output with only a warning | 1.18.3 | `--overwrite` as explicit consent (refused without it from 2.0) | Set `FFMPEG_SKILL_NO_OVERWRITE=1` (the recommended agent setting) and pass `--overwrite` where a replacement is intended |
91
+ | The per-tool v1 success keys next to `result_v2` (`output`, `probe`, `commands`, `verified`, `verification` and each tool's own keys at the top level) | 1.18.4 | `result_v2`, promoted to the top level in 2.0 | Run with `FFMPEG_SKILL_RESULT_V2=1` and read `result_v2` (`metrics`, `notes`, `details`) instead of the top-level keys |
92
+ | `--crf` as an alias of `--quality` on every re-encoding tool that takes `--quality` (`export.py` keeps `--crf`: its preset chooses the encoder) | 1.18.4 | `--quality N` (the same CRF scale, codec-neutral) | Pass `--quality`; `--crf` warns on stderr and is marked in `--help` |
93
+ | `json` and `progress` in the MCP `inputSchema` | 1.18.4 | nothing: the transport sets them itself | Stop sending them from an MCP client; run the server with `FFMPEG_SKILL_MCP_LEAN=1` to see the 2.0 schema |
94
+ | `hdr` meaning "BT.2020 primaries *or* a PQ/HLG transfer" in `probe` | 1.18.4 | `hdr_signal` (true only for PQ / HLG / Dolby Vision); in 2.0 `hdr` takes that meaning | Key on `hdr_signal` for "is this a real HDR signal" and on `hdr_format` for the `BT.2020 SDR` case |
95
+ | Overwriting an existing output with only a warning | 1.18.4 | `--overwrite` as explicit consent (refused without it from 2.0) | Set `FFMPEG_SKILL_NO_OVERWRITE=1` (the recommended agent setting) and pass `--overwrite` where a replacement is intended |
96
96
 
97
97
  ## Skill
98
98
 
99
99
  ```json
100
100
  {
101
101
  "contract_version": "1.0",
102
- "deprecated": [{"what": "...", "since": "1.18.3", "replacement": "...", "removed_in": "2.0.0", "where": "cli | json | mcp | behaviour"}],
103
- "skill": {"id": "ffmpeg-skill", "version": "1.18.3", "execution_mode": "local", "kind": "execution",
102
+ "deprecated": [{"what": "...", "since": "1.18.4", "replacement": "...", "removed_in": "2.0.0", "where": "cli | json | mcp | behaviour"}],
103
+ "skill": {"id": "ffmpeg-skill", "version": "1.18.4", "execution_mode": "local", "kind": "execution",
104
104
  "entrypoints": {"cli": "...", "mcp": "...", "contract": "...", "doctor": "..."},
105
105
  "not_provided": ["AI reasoning", "decisions", "production plans", "project IR", "approvals", "network access", "transcription engine"]},
106
106
  "requirements": {"python": ">=3.9 (standard library only)", "ffmpeg": ">=5.0", "ffprobe": ">=5.0"},
@@ -128,7 +128,7 @@ One entry per tool under `tools`, sorted by id. Tool ids are stable:
128
128
  | `output_schema` | what `--json` prints on stdout |
129
129
  | `supports_dry_run`, `dry_run` | whether `--dry-run` plans without running ffmpeg or writing files |
130
130
  | `supports_json` | whether `--json` exists |
131
- | `supports_json_brief` | whether `--json-brief` exists (1.18.3): the same success document with `probe` replaced by a compact `summary` (`duration_s`, `width`, `height`, `fps`, `vcodec`, `acodec`, `channels`, and `lufs` when the tool measured one), `commands` replaced by the number of commands run, and the per-step `verification` list dropped (its verdict stays in `verified`). Tool-specific keys are unchanged, `--json`'s own output is unchanged, and a failure prints the same failure document either way |
131
+ | `supports_json_brief` | whether `--json-brief` exists (1.18.4): the same success document with `probe` replaced by a compact `summary` (`duration_s`, `width`, `height`, `fps`, `vcodec`, `acodec`, `channels`, and `lufs` when the tool measured one), `commands` replaced by the number of commands run, and the per-step `verification` list dropped (its verdict stays in `verified`). Tool-specific keys are unchanged, `--json`'s own output is unchanged, and a failure prints the same failure document either way |
132
132
  | `mutates_input` | always `false`: no tool overwrites its input |
133
133
  | `produces_artifact` | writes a file (media, PNG, HTML, EDL) |
134
134
  | `verification` | `{required, tools}`: which tools to run on the output afterwards |
@@ -450,7 +450,7 @@ Per-tool keys added in 1.17, all additive:
450
450
  | `jobs`, `jobs_requested`, `wall_seconds`, `item_seconds_total`, `timed_out` | `batch.py` | the parallelism actually applied and the number asked for, the batch's wall clock, the sum of the per-item times (so the speed-up can be quoted), and whether the shared timeout budget ran out. A timed-out item carries `"skipped": "timeout"` in its result row |
451
451
  | `cache` | `render.py --cache` | `{dir, ffmpeg, hits, misses, saved_seconds, entries}`, plus `would_hit` under `--dry-run`. The ffmpeg build banner, the skill version, the contract version, the forwarded flags (`--fast`, `--codec`, …) and the output's extension are all part of every key, so a cache is never reused across any of them — a `--fast` draft is never served to a run that did not ask for one |
452
452
 
453
- Per-tool keys added in 1.18.3, all additive:
453
+ Per-tool keys added in 1.18.4, all additive:
454
454
 
455
455
  | key | tool | what it holds |
456
456
  |---|---|---|
@@ -458,7 +458,7 @@ Per-tool keys added in 1.18.3, all additive:
458
458
  | `text_unchanged` | `caption.py` | a sibling inside the `caption` block, **burn mode only** (`--mode mux` never touches the text and omits the key): `true` when the drawn text equals the cues that were handed in — nothing transcribed, no cue dropped, no cue **split** across two consecutive cues and no glyph stripped (`--emoji none`). Wrapping, line breaks and timing do not count: the words are the same. This tool never rewrites, shortens or translates a cue, so the key is a statement of what happened, not a judgement of the text |
459
459
 
460
460
 
461
- Per-tool keys added in 1.18.3, all additive:
461
+ Per-tool keys added in 1.18.4, all additive:
462
462
 
463
463
  | key | tool | what it holds |
464
464
  |---|---|---|
@@ -588,7 +588,7 @@ most often across the corpus in `evals/agent_prompts*.json`. Every tool -- inclu
588
588
  -- is still callable by name through `tools/call` regardless of what `tools/list` advertised; the
589
589
  contract (`ffmpeg-skill contract --json`) still describes all 42 unconditionally. Set
590
590
  `FFMPEG_SKILL_MCP_FULL=1` (anything but "" or `0`) in the server's environment to make `tools/list`
591
- return all 42, as every version before 1.18.3 did.
591
+ return all 42, as every version before 1.18.4 did.
592
592
 
593
593
  ## Consuming the contract from an agent
594
594
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "ffmpeg-skill",
3
- "version": "1.18.3",
3
+ "version": "1.18.4",
4
4
  "description": "Agent Skill that gives coding agents (Claude Code, Cursor, Codex) a local video editor: 42 FFmpeg tools with a machine-readable contract, contract-derived MCP server, FFmpeg capability detection, probe-first / verify-last workflow. Cut, join, silence removal, fit, captions and karaoke, overlays, motion graphics, HDR to SDR, LUTs, audio clean-up and typed dynamics, sync with drift correction, multicam, loudness, delivery checks, project rendering, batch. No API keys, no cloud, no dependencies.",
5
5
  "keywords": [
6
6
  "ffmpeg",
@@ -107,7 +107,7 @@ from _common.text import (
107
107
  JA_NO_LINE_END, JA_NO_LINE_START, JA_PARTICLE_WORDS, JA_PARTICLES, JA_SENTENCE_END, _join, ORPHAN_MIN_EM,
108
108
  PENALTY_FORBIDDEN, PENALTY_FUNCTION_WORD, PENALTY_FUNCTION_WORD_START, PENALTY_IDEOGRAPHS,
109
109
  PENALTY_NEUTRAL, PENALTY_OKURIGANA,
110
- PENALTY_PARTICLE, PENALTY_SENTENCE_END, _rebalance, _rebalance_phrase, SAFE_WIDTH_FRACTION, _split_hyphens,
110
+ PENALTY_PARTICLE, PENALTY_SENTENCE_END, _rebalance, _rebalance_phrase, SAFE_WIDTH_FRACTION, _split_hyphens, _slice_atom,
111
111
  _particle_ends, _particle_starts, wrap_text, wrap_variants, WRAP_MODES,
112
112
  fit_size, line_em_for_size, MIN_CAPTION_FRACTION, ass_units_local,
113
113
  script_font_for_text, script_font_status, _script_font_uncached, _SCRIPT_RANGES, SCRIPTS, _SHAPING_BUILD_CACHE,
@@ -223,7 +223,7 @@ __all__ = [
223
223
  "PENALTY_FUNCTION_WORD_START",
224
224
  "JA_SENTENCE_END", "_join", "ORPHAN_MIN_EM", "PENALTY_FORBIDDEN", "PENALTY_FUNCTION_WORD",
225
225
  "PENALTY_IDEOGRAPHS", "PENALTY_NEUTRAL", "PENALTY_OKURIGANA", "PENALTY_PARTICLE", "PENALTY_SENTENCE_END",
226
- "_rebalance", "_rebalance_phrase", "SAFE_WIDTH_FRACTION", "_split_hyphens", "_particle_ends",
226
+ "_rebalance", "_rebalance_phrase", "SAFE_WIDTH_FRACTION", "_split_hyphens", "_slice_atom", "_particle_ends",
227
227
  "_particle_starts", "wrap_text", "wrap_variants", "WRAP_MODES",
228
228
  "fit_size", "line_em_for_size", "MIN_CAPTION_FRACTION", "ass_units_local"
229
229
  ]
@@ -38,7 +38,7 @@ from _common.wrap import (
38
38
  text_width_em, SAFE_WIDTH_FRACTION, ORPHAN_MIN_EM, WRAP_MODES, JA_PARTICLES, JA_PARTICLE_WORDS,
39
39
  JA_SENTENCE_END, JA_NO_LINE_START, JA_NO_LINE_END, FUNCTION_WORDS, _FUNCTION_WORDS_ANY, PENALTY_FORBIDDEN,
40
40
  PENALTY_OKURIGANA, PENALTY_FUNCTION_WORD, PENALTY_IDEOGRAPHS, PENALTY_NEUTRAL, PENALTY_FUNCTION_WORD_START,
41
- PENALTY_PARTICLE, PENALTY_SENTENCE_END, _HYPHENS, _atoms, _split_hyphens, _join, _break_spaced, _is_kana,
41
+ PENALTY_PARTICLE, PENALTY_SENTENCE_END, _HYPHENS, _atoms, _split_hyphens, _slice_atom, _join, _break_spaced, _is_kana,
42
42
  _is_katakana_run, _is_hiragana, _is_ideograph, _is_weak_line, _function_words, _bare_word, _particle_starts,
43
43
  _particle_ends, break_penalty, _cut_penalty, best_break, _fix_orphans, _fix_weak_lines, _rebalance,
44
44
  _rebalance_phrase, _greedy_chunks, _balance, wrap_text, wrap_variants, MIN_CAPTION_FRACTION,
@@ -265,6 +265,37 @@ def _split_hyphens(atoms: "List[Tuple[str, bool]]") -> "List[Tuple[str, bool]]":
265
265
  return out
266
266
 
267
267
 
268
+ def _slice_atom(atom: str, max_em: float) -> "List[str]":
269
+ """The escape hatch (1.18.4): hard-slice a single atom that is wider than `max_em` all by
270
+ itself -- a word or run with no break point the wrapper may use, and that STILL does not fit
271
+ even alone on its own line. Nothing else in this file may ever chop an atom mid-character;
272
+ this is the one place that does, and only once every other mechanism (wrapping, then
273
+ caption.py's --min-size shrink) has already failed to make it fit.
274
+
275
+ No dictionary, no hyphenation library: an existing hyphen is preferred as the cut (the same
276
+ break R1 already allows), then whatever is left is cut again at exactly the widest prefix
277
+ that still measures within `max_em`, one character at a time. The characters themselves are
278
+ never rewritten -- every piece concatenates back to the original atom exactly."""
279
+ if text_width_em(atom) <= max_em:
280
+ return [atom]
281
+ for i in range(len(atom) - 2, 0, -1):
282
+ # latest hyphen whose left side still fits -- keeps the left half as large as possible
283
+ if atom[i] in _HYPHENS and text_width_em(atom[:i + 1]) <= max_em:
284
+ return [atom[:i + 1]] + _slice_atom(atom[i + 1:], max_em)
285
+ pieces: "List[str]" = []
286
+ cur = ""
287
+ for ch in atom:
288
+ candidate = cur + ch
289
+ if cur and text_width_em(candidate) > max_em:
290
+ pieces.append(cur)
291
+ cur = ch
292
+ else:
293
+ cur = candidate
294
+ if cur:
295
+ pieces.append(cur)
296
+ return pieces or [atom]
297
+
298
+
268
299
  def _join(left: str, atom: str, spaced: bool) -> str:
269
300
  """Put an atom back on a line, restoring the space that stood before it."""
270
301
  if not left:
@@ -574,15 +605,38 @@ def _rebalance_phrase(lines: "List[str]", max_em: float, lang: "Optional[str]")
574
605
  return out, moved
575
606
 
576
607
 
577
- def _greedy_chunks(raw: str, max_em: float) -> "List[str]":
578
- """The greedy fill on its own: the line count every mode must keep."""
608
+ def _greedy_chunks(raw: str, max_em: float, *, slice_overlong: bool = False,
609
+ sliced: "Optional[List[int]]" = None) -> "List[str]":
610
+ """The greedy fill on its own: the line count every mode must keep.
611
+
612
+ `slice_overlong` (off by default -- only caption.py's final burn-in pass turns it on, never
613
+ fit_size()'s search) is the escape hatch: an atom that lands alone on a line and is STILL
614
+ wider than `max_em` -- a fitting atom never reaches this branch, so a Thai phrase or a
615
+ katakana run that already fits is completely untouched -- is hard-sliced by _slice_atom()
616
+ instead of kept whole. `sliced`, when given, gets one entry (the number of extra lines that
617
+ one atom produced) per atom actually sliced, which is how the caller counts
618
+ `broken_inside_word` without re-deriving it from the output lines."""
579
619
  current = ""
580
620
  chunk: "List[str]" = []
581
621
  for atom, spaced in _atoms(raw):
622
+ if slice_overlong and not current and text_width_em(atom) > max_em:
623
+ pieces = _slice_atom(atom, max_em)
624
+ if sliced is not None and len(pieces) > 1:
625
+ sliced.append(len(pieces) - 1)
626
+ chunk.extend(pieces[:-1])
627
+ current = pieces[-1]
628
+ continue
582
629
  candidate = _join(current, atom, spaced)
583
630
  if current and text_width_em(candidate) > max_em:
584
631
  chunk.append(current)
585
- current = atom
632
+ if slice_overlong and text_width_em(atom) > max_em:
633
+ pieces = _slice_atom(atom, max_em)
634
+ if sliced is not None and len(pieces) > 1:
635
+ sliced.append(len(pieces) - 1)
636
+ chunk.extend(pieces[:-1])
637
+ current = pieces[-1]
638
+ else:
639
+ current = atom
586
640
  else:
587
641
  current = candidate
588
642
  if current:
@@ -608,7 +662,8 @@ def _balance(chunk: "List[str]", max_em: float, mode: str, lang: "Optional[str]"
608
662
 
609
663
 
610
664
  def wrap_text(text: str, max_em: float, *, balance: bool = True, mode: str = "phrase",
611
- lang: "Optional[str]" = None) -> "List[str]":
665
+ lang: "Optional[str]" = None, slice_overlong: bool = False,
666
+ sliced: "Optional[List[int]]" = None) -> "List[str]":
612
667
  """Wrap `text` to lines no wider than `max_em` em, keeping the manual breaks it already has.
613
668
 
614
669
  An atom wider than the whole line (one very long word) is left alone on its line rather than
@@ -626,11 +681,12 @@ def wrap_text(text: str, max_em: float, *, balance: bool = True, mode: str = "ph
626
681
  for raw in text.split("\n"):
627
682
  if not raw.strip():
628
683
  continue
629
- chunk = _greedy_chunks(raw, max_em)
684
+ chunk = _greedy_chunks(raw, max_em, slice_overlong=slice_overlong, sliced=sliced)
630
685
  lines.extend(_balance(chunk, max_em, mode, lang) if balance else chunk)
631
686
  return lines or [text]
632
687
  def wrap_variants(text: str, max_em: float, *, mode: str = "phrase",
633
- lang: "Optional[str]" = None) -> "Tuple[List[str], List[str], List[str]]":
688
+ lang: "Optional[str]" = None, slice_overlong: bool = False,
689
+ sliced: "Optional[List[int]]" = None) -> "Tuple[List[str], List[str], List[str]]":
634
690
  """`(wrapped, greedy, measured)` for one cue from a single greedy fill.
635
691
 
636
692
  layout_cues needs all three -- `wrapped` is what is burnt in, `greedy` is what `rebalanced`
@@ -645,7 +701,7 @@ def wrap_variants(text: str, max_em: float, *, mode: str = "phrase",
645
701
  for raw in text.split("\n"):
646
702
  if not raw.strip():
647
703
  continue
648
- chunk = _greedy_chunks(raw, max_em)
704
+ chunk = _greedy_chunks(raw, max_em, slice_overlong=slice_overlong, sliced=sliced)
649
705
  greedy.extend(chunk)
650
706
  wrapped.extend(_balance(list(chunk), max_em, mode, lang))
651
707
  measured.extend(chunk if mode == "measured" else _balance(list(chunk), max_em, "measured", None))
@@ -53,7 +53,7 @@ from _common import (ASR_INSTALL_HINT, die_no_engine, parse_srt, transcribe, whi
53
53
  from _common import (SAFE_WIDTH_FRACTION, ORPHAN_MIN_EM, WRAP_MODES, wrap_text, wrap_variants, best_break,
54
54
  fit_size, line_em_for_size, MIN_CAPTION_FRACTION,
55
55
  break_penalty, _is_weak_line, _atoms, _join, _break_spaced, _bare_word, _function_words,
56
- _split_hyphens, FUNCTION_WORDS, JA_PARTICLES, JA_SENTENCE_END, _fix_orphans, _rebalance)
56
+ _split_hyphens, _slice_atom, FUNCTION_WORDS, JA_PARTICLES, JA_SENTENCE_END, _fix_orphans, _rebalance)
57
57
 
58
58
  # The breaker's names are caption.py's public surface as much as _common's: every caller and test
59
59
  # that reached for `caption.wrap_text` before 1.16 still does.
@@ -62,7 +62,7 @@ __all__ = ["parse_srt", "write_srt", "transcribe", "whisper_word_timings", "die_
62
62
  "SAFE_WIDTH_FRACTION", "ORPHAN_MIN_EM", "WRAP_MODES", "wrap_text", "wrap_variants",
63
63
  "fit_size", "line_em_for_size", "MIN_CAPTION_FRACTION",
64
64
  "best_break", "break_penalty", "_is_weak_line", "_atoms", "_join", "_break_spaced",
65
- "_bare_word", "_function_words", "_split_hyphens", "FUNCTION_WORDS", "JA_PARTICLES",
65
+ "_bare_word", "_function_words", "_split_hyphens", "_slice_atom", "FUNCTION_WORDS", "JA_PARTICLES",
66
66
  "JA_SENTENCE_END", "_fix_orphans", "_rebalance", "char_script", "NO_SPACE_SCRIPTS",
67
67
  "text_width_em"]
68
68
 
@@ -319,6 +319,10 @@ def layout_cues(cues: List[Tuple[float, float, str]], *, max_em: Optional[float]
319
319
  """
320
320
  stats = {"shifted": 0, "wrapped": 0, "split": 0, "extended": 0, "dropped": 0, "rebalanced": 0,
321
321
  "wrap": wrap, "phrase_breaks": 0}
322
+ # ass_units_local's floor + the escape hatch below mean an atom this loop hands to
323
+ # wrap_variants(..., slice_overlong=True) can still overflow only if a SINGLE character is
324
+ # wider than the column -- practically unreachable at any real --min-size. `overlong` stays
325
+ # in stats for that theoretical remainder; `broken_inside_word` is the count that matters now.
322
326
  staged: List[Tuple[float, float, str]] = []
323
327
  staged_sizes: List[Optional[int]] = []
324
328
  for cue_index, (start, end, text) in enumerate(cues):
@@ -339,12 +343,19 @@ def layout_cues(cues: List[Tuple[float, float, str]], *, max_em: Optional[float]
339
343
  # one greedy fill per cue, three answers off it: what gets burnt in, what the
340
344
  # greedy wrap would have given (`rebalanced`) and what 1.15's wrap would have
341
345
  # given (`phrase_breaks`). Three wrap_text() calls re-ran the atomiser each time.
342
- lines, greedy, measured = wrap_variants(text, own_em, mode=wrap, lang=lang)
346
+ sliced_atoms: List[int] = []
347
+ lines, greedy, measured = wrap_variants(text, own_em, mode=wrap, lang=lang,
348
+ slice_overlong=True, sliced=sliced_atoms)
343
349
  if lines != [l for l in text.split("\n") if l.strip()]:
344
350
  stats["wrapped"] += 1
351
+ if sliced_atoms:
352
+ # the escape hatch fired: a word or run with no break point the wrapper may use
353
+ # was still wider than the column alone, even at the size in force, so it was
354
+ # hard-sliced at the column edge (or after its own hyphen) instead of clipping
355
+ stats["broken_inside_word"] = stats.get("broken_inside_word", 0) + len(sliced_atoms)
345
356
  if any(text_width_em(l) > own_em for l in lines):
346
- # a run with no break point the wrapper may use (a long word, a Thai phrase
347
- # without spaces) stays long rather than chopped: say so, and name the fix
357
+ # the escape hatch above could not make every piece fit -- only possible when a
358
+ # single character is wider than the column itself
348
359
  stats["overlong"] = stats.get("overlong", 0) + 1
349
360
  if lines != greedy:
350
361
  stats["rebalanced"] += 1
@@ -381,13 +392,19 @@ def layout_cues(cues: List[Tuple[float, float, str]], *, max_em: Optional[float]
381
392
 
382
393
  def report_layout(stats: dict) -> None:
383
394
  """One info line, only when a cue actually changed."""
384
- parts = [f"{stats[k]} {k}" for k in ("shifted", "wrapped", "rebalanced", "phrase_breaks", "split", "extended", "dropped", "overlong") if stats.get(k)]
395
+ parts = [f"{stats[k]} {k}" for k in ("shifted", "wrapped", "rebalanced", "phrase_breaks", "split",
396
+ "extended", "dropped", "broken_inside_word", "overlong") if stats.get(k)]
385
397
  if parts:
386
398
  info("cues: " + ", ".join(parts))
399
+ if stats.get("broken_inside_word"):
400
+ info(f"{stats['broken_inside_word']} cue line(s) had a word or run with no break point "
401
+ "(Thai writes none inside a phrase) that was still wider than the column alone -- it "
402
+ "was sliced at the column edge (after its own hyphen when it has one) instead of "
403
+ "clipping past the frame; put a space or `|` where the line may break, or use a "
404
+ "smaller --size, to avoid the slice")
387
405
  if stats.get("overlong"):
388
- info(f"{stats['overlong']} cue line(s) wider than the safe width: a word or a run with no break "
389
- "point (Thai writes none inside a phrase) was kept whole rather than chopped -- put a "
390
- "space or `|` where the line may break, or use a smaller --size")
406
+ info(f"{stats['overlong']} cue line(s) wider than the safe width even after slicing: a "
407
+ "single character was wider than the column -- use a smaller --size")
391
408
 
392
409
 
393
410
  def margins_x(args, play_w: Optional[int]) -> Tuple[int, int]: