arxiv-paper-zh 0.1.2 → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "arxiv-paper-zh",
3
- "version": "0.1.2",
3
+ "version": "0.2.0",
4
4
  "description": "Verify and translate arXiv papers into compilable Chinese editions with paired English and Chinese PDFs.",
5
5
  "author": {
6
6
  "name": "AndrewYq",
package/README.md CHANGED
@@ -33,9 +33,11 @@ npm 包地址:[arxiv-paper-zh](https://www.npmjs.com/package/arxiv-paper-zh)
33
33
 
34
34
  - 核对论文简称、标题、作者、摘要、URL 与 arXiv ID,避免翻错同名论文。
35
35
  - 按 `arxiv-paper/<paper-name>/` 统一保存原始源码、中文源码以及中英文 PDF。
36
- - 按可见英文词量均衡分片,支持多个 subagent 并行处理长论文。
36
+ - 将可见英文切成紧凑任务包,公式、引用、URL、代码和注释使用可逆占位符,不再让模型重复读取整份 TeX。
37
+ - 按词量自适应分配 worker;长论文并行、短论文自动减少 worker,单文件论文也能分片。
37
38
  - 保留公式、数值、引用键、标签、人名、模型名、数据集名和常用缩写。
38
39
  - 翻译正文、章节标题、脚注、表头、表注和 caption。
40
+ - 参考文献标题与条目保持原文,并在翻译分片和漏译审计中自动跳过。
39
41
  - 批量检查并安装缺失的 TeX 宏包,复用共享 TinyTeX/TeX Live 缓存。
40
42
  - 自动执行漏译审计、BibTeX/Biber 构建和引用收敛检查。
41
43
  - 使用 XeLaTeX 生成中文 PDF,并要求逐页视觉核验。
@@ -64,13 +66,17 @@ arxiv-paper-zh/
64
66
  │ │ ├── build_and_check.py
65
67
  │ │ ├── inventory_and_shard.py
66
68
  │ │ ├── prepare_output_layout.py
67
- │ │ └── prepare_tex_runtime.py
69
+ │ │ ├── prepare_tex_runtime.py
70
+ │ │ ├── translation_tasks.py
71
+ │ │ └── tex_translation_utils.py
68
72
  │ └── references/
69
73
  │ └── paper-translation-packages.txt
70
74
  ├── install.sh
71
75
  ├── package.json
72
76
  ├── tests/
73
- │ └── installer.test.mjs
77
+ │ ├── installer.test.mjs
78
+ │ ├── test_bibliography_exclusion.py
79
+ │ └── test_translation_tasks.py
74
80
  ├── .gitignore
75
81
  └── README.md
76
82
  ```
@@ -237,7 +243,7 @@ arxiv-paper/EST/
237
243
 
238
244
  `paper-name` 使用用户熟悉的简短名称并保留大小写,例如 `EST`、`Onetrans`,且只能包含英文字母、数字、点、下划线和连字符。
239
245
 
240
- 参考文献元数据默认不翻译;原图内部文字不修改,只翻译必要 caption。公式内的英文说明按“公式不变”原则保留。
246
+ 整个参考文献部分保持原文,包括标题和全部条目;`.bib`、`.bbl`、内嵌 bibliography 环境和单独的参考文献 TeX 文件均不参与翻译。正文中的文献综述仍照常翻译。原图内部文字不修改,只翻译必要 caption。公式内的英文说明按“公式不变”原则保留。
241
247
 
242
248
  ## 内置工具
243
249
 
@@ -245,8 +251,12 @@ arxiv-paper/EST/
245
251
  # 创建并输出固定的论文产物路径
246
252
  python3 skills/arxiv-paper-zh/scripts/prepare_output_layout.py EST --root arxiv-paper
247
253
 
248
- # 统计文件并生成均衡翻译分片
249
- python3 skills/arxiv-paper-zh/scripts/inventory_and_shard.py arxiv-paper/EST/latex/paper-zh --entry main.tex --workers 3 --json
254
+ # 生成紧凑翻译任务包(3 是 worker 上限,小论文会自动减少)
255
+ python3 skills/arxiv-paper-zh/scripts/translation_tasks.py prepare arxiv-paper/EST/latex/paper-zh --entry main.tex --workers 3 --json
256
+
257
+ # worker 填写各自的 worker-*.task 后,检查进度并安全合并
258
+ python3 skills/arxiv-paper-zh/scripts/translation_tasks.py status arxiv-paper/EST/latex/paper-zh
259
+ python3 skills/arxiv-paper-zh/scripts/translation_tasks.py apply arxiv-paper/EST/latex/paper-zh
250
260
 
251
261
  # 检查预装包和论文专用宏包
252
262
  python3 skills/arxiv-paper-zh/scripts/prepare_tex_runtime.py arxiv-paper/EST/latex/paper-zh --preset --kpsewhich /path/to/kpsewhich
@@ -267,7 +277,7 @@ python3 skills/arxiv-paper-zh/scripts/build_and_check.py arxiv-paper/EST/latex/p
267
277
  - 是否允许网络下载与安装宏包。
268
278
  - 是否提供可写文件系统和本地 XeLaTeX。
269
279
 
270
- 客户端不支持 subagent 时,Agent 应顺序处理分片;不影响翻译规则与构建脚本的使用。
280
+ 客户端不支持 subagent 时,Agent 应顺序填写任务包。任务包仍会剔除不需要发送给模型的内容,因此同样能降低上下文开销;只是不具备并行加速。
271
281
 
272
282
  ## 开发与发布检查
273
283
 
package/package.json CHANGED
@@ -1,10 +1,10 @@
1
1
  {
2
2
  "name": "arxiv-paper-zh",
3
- "version": "0.1.2",
3
+ "version": "0.2.0",
4
4
  "description": "Install the arxiv-paper-zh Agent Skill for Codex, Claude Code, and compatible agents.",
5
5
  "type": "module",
6
6
  "bin": {
7
- "arxiv-paper-zh": "./bin/arxiv-paper-zh.mjs"
7
+ "arxiv-paper-zh": "bin/arxiv-paper-zh.mjs"
8
8
  },
9
9
  "files": [
10
10
  ".codex-plugin/",
@@ -17,7 +17,7 @@
17
17
  ],
18
18
  "scripts": {
19
19
  "test": "node --test tests/*.test.mjs",
20
- "check": "node --check bin/arxiv-paper-zh.mjs && python3 -c \"from pathlib import Path; [compile(path.read_bytes(), str(path), 'exec') for path in Path('skills/arxiv-paper-zh/scripts').glob('*.py')]\" && npm test"
20
+ "check": "node --check bin/arxiv-paper-zh.mjs && python3 -c \"from pathlib import Path; [compile(path.read_bytes(), str(path), 'exec') for path in Path('skills/arxiv-paper-zh/scripts').glob('*.py')]\" && python3 -m unittest discover -s tests -p 'test_*.py' && npm test"
21
21
  },
22
22
  "engines": {
23
23
  "node": ">=18"
@@ -1,15 +1,15 @@
1
1
  ---
2
2
  name: arxiv-paper-zh
3
- description: Verify that a requested paper name, acronym, title, authors, abstract, URL, and arXiv ID identify the intended work; then rapidly download and translate its TeX source into Chinese with balanced subagent sharding, deterministic artifact paths under arxiv-paper, paired English/Chinese PDFs, cached dependencies, audits, XeLaTeX compilation, and full-page visual verification. Use when a user asks to locate and translate an arXiv or LaTeX research paper into a compilable Chinese edition, especially for ambiguous acronyms or long papers requiring parallel translation.
3
+ description: Verify an arXiv or LaTeX paper's identity, then produce complete, compilable English and Chinese sources and PDFs. Use for locating and translating research papers while preserving math, citations, data, and layout; the workflow uses compact protected translation packets, adaptive parallel workers, cached TeX dependencies, audits, XeLaTeX compilation, and full-page visual verification.
4
4
  ---
5
5
 
6
6
  # arXiv 论文快速中文化
7
7
 
8
- 交付保留原始数学、数据、引用和作者信息的中英文可编译源码及经过页面核验的中英文 PDF。身份正确性与译文完整性优先,但要并行执行互不依赖的阶段。
8
+ 交付保留数学、数据、引用和作者信息的中英文源码与经过页面核验的 PDF。身份正确性和译文完整性优先;翻译阶段使用紧凑任务包,避免将整份 TeX 和重复会话上下文发送给每个 worker。
9
9
 
10
- ## 输出目录契约
10
+ ## 输出目录
11
11
 
12
- 将全部用户产物固定放到 `arxiv-paper/<paper-name>/`。`paper-name` 优先使用用户给出的论文简称并保留大小写,例如 `EST`、`Onetrans`;若用户只给出标题,则核验论文身份后生成简短名称。名称必须匹配 `[A-Za-z0-9][A-Za-z0-9._-]*`,不得包含空格、斜杠或 arXiv ID,除非 ID 本身就是用户指定名称。
12
+ 全部产物固定放在 `arxiv-paper/<paper-name>/`:
13
13
 
14
14
  ```text
15
15
  arxiv-paper/<paper-name>/
@@ -17,81 +17,85 @@ arxiv-paper/<paper-name>/
17
17
  │ ├── source.tar
18
18
  │ ├── paper-en/
19
19
  │ └── paper-zh/
20
- ├── paper-en/
21
- │ └── <paper-name>-en.pdf
22
- └── paper-zh/
23
- └── <paper-name>-zh.pdf
20
+ ├── paper-en/<paper-name>-en.pdf
21
+ └── paper-zh/<paper-name>-zh.pdf
24
22
  ```
25
23
 
26
- `latex/paper-en/` 保存未经翻译的解压源码,`latex/paper-zh/` 保存其完整中文副本;两个 PDF 目录只保存最终交付 PDF。开始下载前先运行:
24
+ `paper-name` 优先使用用户简称并保留大小写,且必须匹配 `[A-Za-z0-9][A-Za-z0-9._-]*`。开始下载前运行:
27
25
 
28
26
  ```bash
29
27
  python3 scripts/prepare_output_layout.py <paper-name> --root arxiv-paper
30
28
  ```
31
29
 
32
- 以脚本输出的绝对路径为准,后续不得另建 `work/<arxiv-id>`、顶层 `paper_cn/` 或其他平行交付目录。临时渲染图、构建日志和联系表可放在论文目录下的隐藏临时目录,交付前不得混入 `paper-en/` 或 `paper-zh/`。
30
+ 以脚本输出的绝对路径为准,不另建平行交付目录。`.translation-tasks/`、渲染图和日志属于论文目录内的临时文件,不得混入最终 PDF 目录。
33
31
 
34
- ## 共享论文翻译运行时
32
+ ## 工作流
35
33
 
36
- 优先复用固定位置的 TinyTeX/TeX Live,不要将运行时复制进每个论文目录。首次建立或主动刷新共享运行时时,使用内置预装清单一次检查并一次批量安装:
34
+ 1. 从 arXiv 摘要页或 API 核验规范化 ID、完整标题、作者、摘要、版本和日期。简称不是唯一标识;候选不唯一时让用户确认。向用户明确标题、作者和 arXiv ID。
35
+ 2. 创建输出目录,将 `https://export.arxiv.org/e-print/<ID>` 保存为 `latex/source.tar` 并解压到 `latex/paper-en/`。用主 TeX 的标题、作者或 README 二次核验身份;不一致时停止。
36
+ 3. 将英文源码内容完整复制到 `latex/paper-zh/`,不得多套一层目录;此后只修改中文副本。分别定位中英文入口文件。
37
+ 4. 先完成中文入口的 XeLaTeX/ctex 改造和模板可见字符串本地化,再生成翻译任务。通常加入 `\usepackage[UTF8,fontset=fandol]{ctex}`,移除仅适用于 pdfLaTeX 的 `inputenc` 和 T1 `fontenc`。生成任务后、合并任务前不要再编辑中文 TeX 源码,合并器会检查快照。
38
+ 5. 生成紧凑任务包。脚本包含入口文件,因此单文件论文也能按片段并行;它自动省略参考文献,并用可逆占位符保护公式、引用、URL、代码和注释:
37
39
 
38
- ```bash
39
- python3 scripts/prepare_tex_runtime.py --preset \
40
- --kpsewhich <shared-tex-root>/bin/<platform>/kpsewhich \
41
- --tlmgr <shared-tex-root>/bin/<platform>/tlmgr --install
42
- ```
43
-
44
- 预装清单位于 `references/paper-translation-packages.txt`,覆盖 XeLaTeX 中文排版、BibTeX/Biber、常见数学与字体、表格、算法、代码、绘图、caption,以及 ACM、IEEE、Elsevier、Springer 模板。首次安装需要联网;后续任务先离线检查,只有论文特有依赖缺失时才联网批量补装。不要安装 `scheme-full`。
45
-
46
- ## 快速流水线
47
-
48
- 1. 从 arXiv 摘要页或 API 核验规范化 ID、完整标题、作者、摘要、版本和日期。简称不是唯一标识。只有一个候选与用户主题高度一致时才继续,否则先让用户确认。明确告知用户“标题 + 作者 + arXiv ID”。
49
- 2. 确定并校验 `paper-name`,运行 `scripts/prepare_output_layout.py`。从 `https://export.arxiv.org/e-print/<ID>` 下载原始压缩包到 `latex/source.tar`,解压到 `latex/paper-en/`。用主 TeX 的 `\title{}`、作者或 README 做第二次身份核验;不一致时停止。
50
- 3. 将 `latex/paper-en/` 的全部内容完整复制到 `latex/paper-zh/`,不得产生 `latex/paper-zh/paper-en/` 额外嵌套;此后只修改中文副本。分别定位英文和中文入口文件。
51
- 4. 立即并行启动两条路径:
40
+ ```bash
41
+ python3 scripts/translation_tasks.py prepare \
42
+ arxiv-paper/<paper-name>/latex/paper-zh \
43
+ --entry main.tex --workers 3 --json
44
+ ```
52
45
 
53
- - 翻译路径:按可见英文词量生成均衡文件分片;正文较多时占满可用 subagent 槽位。每个代理只能编辑明确分配的文件。
54
- - 构建路径:主代理同时本地化入口/样式文件、准备中文字体并预检 TeX 依赖。不要等待翻译结束后才开始配置编译。
46
+ 默认 `--workers 3` 是上限;小论文会自动减少 worker,避免启动开销。只有需要改变速度/上下文折中才调整 `--chunk-words` 或 `--min-words-per-worker`。
47
+ 6. 每个 `worker-*.task` 只交给一个 worker。支持隔离上下文时使用空/最小历史,而不是复制完整会话;任务提示只需:
55
48
 
56
- ```bash
57
- python3 scripts/inventory_and_shard.py arxiv-paper/<paper-name>/latex/paper-zh --entry main.tex --workers 3 --json
58
- python3 scripts/prepare_tex_runtime.py arxiv-paper/<paper-name>/latex/paper-zh --preset --kpsewhich /path/to/kpsewhich
49
+ ```text
50
+ 翻译 <packet> 的全部 SOURCE 区块,把译文填入对应 TRANSLATION 区块。
51
+ 严格遵守文件头规则,只编辑该任务包,不读取或修改论文源码。
59
52
  ```
60
53
 
61
- 5. 若依赖缺失,使用同一脚本增加 `--tlmgr /path/to/tlmgr --install`,在一次 `tlmgr` 调用中补齐预装包和论文特有包;不得逐个安装。优先复用共享运行时。只有不存在可用运行时时才在可写缓存目录安装便携 TinyTeX;不得使用 `sudo`、`scheme-full` 或完整 TeX Live collection。
62
- 6. 主代理修改入口文件以支持 XeLaTeX 中文排版,通常使用 `\usepackage[UTF8,fontset=fandol]{ctex}`。删除仅适用于 pdfLaTeX 的 `inputenc` 和 T1 `fontenc`。本地化模板硬编码的 `Abstract`、`References`、`Keywords` 等可见字符串。
63
- 7. 翻译所有渲染的英文自然语言:标题、摘要、正文、章节标题、列表、脚注、表头、表注和 caption。保留公式、数学符号、数值、引用键、label/ref、URL、LaTeX 结构、人名、模型名、数据集名、缩写与通行技术标识。图片只翻译必要 caption;算法/代码只翻译自然语言注释、docstring、caption 和说明。参考文献元数据默认不翻译。
64
- 8. 子代理完成后,主代理只运行一次全局审计并人工复核全部命中;不得仅扫描顶层章节,也不得把子代理的自检当作最终证明:
54
+ Worker 不需要读取本 Skill、整篇论文或其他任务包。不同 worker 并行编辑各自任务包;不支持 subagent 时顺序处理。主 agent 同时编译英文版、检查中文依赖,但不修改已快照的中文 TeX。
55
+ 7. Worker 完成后只运行一次合并。合并器先整体校验任务完整性、占位符、LaTeX 结构和源文件哈希,任何错误都会在写文件前停止:
65
56
 
66
57
  ```bash
58
+ python3 scripts/translation_tasks.py status arxiv-paper/<paper-name>/latex/paper-zh
59
+ python3 scripts/translation_tasks.py apply arxiv-paper/<paper-name>/latex/paper-zh
67
60
  python3 scripts/audit_tex_translation.py arxiv-paper/<paper-name>/latex/paper-zh
68
61
  ```
69
62
 
70
- 9. 先使用原论文声明或兼容的构建引擎编译 `latex/paper-en/`,确保英文源码、参考文献和交叉引用收敛。再使用自动构建脚本识别 BibTeX/Biber,并完成中文 XeLaTeX 收敛:
63
+ `status` 未完成时只返工列出的任务;`apply` 报错时只检查对应 segment,不重新读取或重译全部论文。主 agent 必须人工复核全局审计命中。
64
+ 8. 使用原论文声明或兼容引擎编译英文源码。中文使用自动构建脚本识别 BibTeX/Biber 并完成 XeLaTeX 收敛:
71
65
 
72
66
  ```bash
73
- python3 scripts/build_and_check.py arxiv-paper/<paper-name>/latex/paper-zh/main.tex --tex-bin /path/to/tex/bin
67
+ python3 scripts/build_and_check.py \
68
+ arxiv-paper/<paper-name>/latex/paper-zh/main.tex --tex-bin /path/to/tex/bin
74
69
  ```
75
70
 
76
- 若仍报告缺包,一次批量安装全部新缺失项后重试。不得接受 LaTeX 错误、未定义引用/citation 或缺失字符。
77
- 10. 分别低分辨率渲染中英文 PDF 的全部页面生成联系表,检查裁切、重叠、表格溢出、图片和页数;只对可疑页面高分辨率渲染。另用能正确显示 CJK 的系统 PDF 引擎抽查中文字体。Poppler 缺少 CMap 时不得把空白中文误判为正常。
78
- 11. 将英文成品复制为 `paper-en/<paper-name>-en.pdf`,将中文成品复制为 `paper-zh/<paper-name>-zh.pdf`。回复列出完整标题、作者、arXiv ID、`latex/paper-en/`、`latex/paper-zh/` 以及两个 PDF 的绝对路径。
71
+ 9. 分别低分辨率渲染中英文 PDF 的全部页面并检查裁切、重叠、溢出、图片和页数;只对可疑页面高分辨率渲染。另用支持 CJK 的系统 PDF 引擎抽查中文字体。Poppler 缺少 CMap 时不得把空白中文误判为正常。
72
+ 10. 将英文成品复制为 `paper-en/<paper-name>-en.pdf`,中文成品复制为 `paper-zh/<paper-name>-zh.pdf`。回复列出论文身份、两套源码目录和两个 PDF 的绝对路径。
79
73
 
80
- ## 并行规则
74
+ ## 翻译边界
81
75
 
82
- - 分片前先创建 `latex/paper-zh/`,并按脚本的权重而非文件数分配。
83
- - 入口文件、共享宏和样式只由主代理修改;任何文件不得由两个代理同时编辑。
84
- - 主代理在翻译进行时完成依赖预检、字体方案和编译入口改造。
85
- - 共享运行时只建立一次;论文任务不得重复下载或复制相同运行时。
86
- - 对小论文不强制派发代理;代理启动与合并开销高于预计翻译时间时直接处理。
76
+ - 翻译标题、摘要、正文、章节标题、列表、脚注、表头、表注和 caption。
77
+ - 保留公式、数学符号、数值、引用键、label/ref、URL、LaTeX 结构、人名、模型名、数据集名、缩写与通行技术标识。
78
+ - 图片内部文字默认不修改;算法和代码块(含内嵌注释/docstring)默认保持原样,只翻译 caption 与外部说明。用户明确要求时再单独处理代码内文本。
79
+ - 参考文献元数据保持原文;`.bib`、`.bbl`、内嵌 bibliography 环境和独立参考文献 TeX 文件均不进入任务包。
80
+ - 不把完整 TeX 内容粘贴进 agent 提示,不让 worker 直接编辑论文文件,也不让多个 worker 处理同一任务包。
87
81
 
88
- ## 完成标准
82
+ ## 共享 TeX 运行时
83
+
84
+ 优先复用固定位置的 TinyTeX/TeX Live。首次建立或主动刷新时,使用 `references/paper-translation-packages.txt` 一次检查并批量安装,不安装 `scheme-full`:
85
+
86
+ ```bash
87
+ python3 scripts/prepare_tex_runtime.py --preset \
88
+ --kpsewhich <shared-tex-root>/bin/<platform>/kpsewhich \
89
+ --tlmgr <shared-tex-root>/bin/<platform>/tlmgr --install
90
+ ```
89
91
 
90
- - arXiv 元数据与源码身份两次核验通过,原始源码未修改。
91
- - 英文与中文源码分别完整位于 `latex/paper-en/` 和 `latex/paper-zh/`,原始压缩包位于 `latex/source.tar`;渲染的英文自然语言均已翻译或被人工判定为允许保留项。
92
- - `paper-en/<paper-name>-en.pdf` 与 `paper-zh/<paper-name>-zh.pdf` 均构建成功,参考文献和交叉引用收敛;中文版无缺字。
93
- - 中英文 PDF 的全部页面均已检查,中文字体由 CJK 能力正常的渲染器确认。
92
+ 后续先离线检查;论文特有依赖缺失时再用同一脚本和一次 `tlmgr` 调用批量补装。只有没有可用运行时时才在可写缓存目录建立便携 TinyTeX,不使用 `sudo`。
93
+
94
+ ## 完成标准
94
95
 
95
- ## 方案参考
96
+ - arXiv 元数据与源码身份两次核验通过,`latex/paper-en/` 未修改。
97
+ - 中文任务全部合并,全局漏译审计已人工复核;所有可见自然语言均已翻译或明确允许保留。
98
+ - 中英文 PDF 均构建成功,参考文献和交叉引用收敛,中文版无缺字。
99
+ - 中英文 PDF 全部页面均已检查,中文字体经 CJK 能力正常的渲染器确认。
96
100
 
97
- 本 Skill 的论文中文化方案参考了科学空间文章:[《让 AI 翻译一篇完整的论文》](https://spaces.ac.cn/archives/11578)。具体实现结合 Codex 的并行代理、依赖缓存、自动审计和 XeLaTeX 构建流程进行了调整。
101
+ 方案参考科学空间文章[《让 AI 翻译一篇完整的论文》](https://spaces.ac.cn/archives/11578),并结合紧凑任务包、并行 agent、依赖缓存和自动审计实现。
@@ -7,13 +7,19 @@ import argparse
7
7
  import re
8
8
  from pathlib import Path
9
9
 
10
+ from tex_translation_utils import is_bibliography_file, mask_bibliography
11
+
10
12
 
11
13
  PROSE = re.compile(r"[A-Za-z]{4,}(?:[ \t]+[A-Za-z][A-Za-z'’-]{2,}){2,}")
12
14
  COMMAND_ONLY = re.compile(r"^[ \t]*\\(?:usepackage|documentclass|input|include|addbibresource)\b")
13
15
 
14
16
 
15
17
  def tex_files(root: Path) -> list[Path]:
16
- return sorted(path for path in root.rglob("*.tex") if path.is_file())
18
+ return sorted(
19
+ path
20
+ for path in root.rglob("*.tex")
21
+ if path.is_file() and not is_bibliography_file(path)
22
+ )
17
23
 
18
24
 
19
25
  def visible_part(line: str) -> str:
@@ -45,7 +51,8 @@ def main() -> int:
45
51
 
46
52
  hits = 0
47
53
  for path in files:
48
- for number, raw in enumerate(path.read_text(encoding="utf-8", errors="replace").splitlines(), 1):
54
+ text = mask_bibliography(path.read_text(encoding="utf-8", errors="replace"))
55
+ for number, raw in enumerate(text.splitlines(), 1):
49
56
  line = visible_part(raw)
50
57
  if not line.strip() or COMMAND_ONLY.match(line):
51
58
  continue
@@ -1,8 +1,11 @@
1
1
  #!/usr/bin/env python3
2
- """Inventory reachable TeX sources and greedily balance translation shards."""
2
+ """Inventory translatable TeX sources and greedily balance translation shards."""
3
3
  from __future__ import annotations
4
4
  import argparse, json, re
5
5
  from pathlib import Path
6
+
7
+ from tex_translation_utils import is_bibliography_file, mask_bibliography
8
+
6
9
  INCLUDE = re.compile(r"\\(?:input|include)\s*\{([^}]+)\}")
7
10
  COMMENT = re.compile(r"(?<!\\)%.*$")
8
11
  COMMAND = re.compile(r"\\[A-Za-z@]+\*?(?:\[[^]]*\])?")
@@ -13,17 +16,19 @@ def reachable(entry: Path, root: Path) -> list[Path]:
13
16
  path = pending.pop()
14
17
  if path in seen or not path.is_file() or (root not in path.parents and path != root): continue
15
18
  seen.add(path)
16
- for value in INCLUDE.findall(path.read_text(encoding="utf-8", errors="replace")):
19
+ text = mask_bibliography(path.read_text(encoding="utf-8", errors="replace"))
20
+ for value in INCLUDE.findall(text):
17
21
  child = (path.parent / value.strip()).resolve()
18
22
  pending.append(child if child.suffix else child.with_suffix(".tex"))
19
23
  return sorted(seen)
20
24
  def weight(path: Path) -> int:
21
- text = "\n".join(COMMENT.sub("", line) for line in path.read_text(encoding="utf-8", errors="replace").splitlines())
25
+ text = mask_bibliography(path.read_text(encoding="utf-8", errors="replace"))
26
+ text = "\n".join(COMMENT.sub("", line) for line in text.splitlines())
22
27
  return len(re.findall(r"[A-Za-z]{2,}", COMMAND.sub("", MATH.sub("", text))))
23
28
  def main() -> int:
24
29
  parser = argparse.ArgumentParser(); parser.add_argument("root", type=Path); parser.add_argument("--entry", type=Path); parser.add_argument("--workers", type=int, default=3); parser.add_argument("--min-weight", type=int, default=1); parser.add_argument("--json", action="store_true"); args = parser.parse_args()
25
30
  root = args.root.resolve(); entry = (root / args.entry).resolve() if args.entry else None
26
- files = reachable(entry, root) if entry else sorted(root.rglob("*.tex")); shared = {"macro.tex", "macros.tex", "commands.tex"}; records = [(p, weight(p)) for p in files if p.is_file() and p != entry and p.name.lower() not in shared]; records = [(p, score) for p, score in records if score >= args.min_weight]
31
+ files = reachable(entry, root) if entry else sorted(root.rglob("*.tex")); shared = {"macro.tex", "macros.tex", "commands.tex"}; records = [(p, weight(p)) for p in files if p.is_file() and p != entry and p.name.lower() not in shared and not is_bibliography_file(p)]; records = [(p, score) for p, score in records if score >= args.min_weight]
27
32
  groups = [[] for _ in range(max(1, args.workers))]; totals = [0] * len(groups)
28
33
  for path, score in sorted(records, key=lambda item: item[1], reverse=True):
29
34
  index = min(range(len(groups)), key=totals.__getitem__); groups[index].append((path, score)); totals[index] += score
@@ -0,0 +1,84 @@
1
+ #!/usr/bin/env python3
2
+ """Shared helpers for excluding bibliography content from translation work."""
3
+
4
+ from __future__ import annotations
5
+
6
+ import re
7
+ from pathlib import Path
8
+
9
+
10
+ _BIBLIOGRAPHY_ENVIRONMENT = re.compile(
11
+ r"\\begin\s*\{(?P<name>thebibliography|biblist|references)\}"
12
+ r".*?"
13
+ r"\\end\s*\{(?P=name)\}",
14
+ re.IGNORECASE | re.DOTALL,
15
+ )
16
+ _BIBLIOGRAPHY_HEADING = re.compile(
17
+ r"\\(?P<level>chapter|section)\*?\s*"
18
+ r"(?:\[[^]]*\]\s*)?"
19
+ r"\{\s*(?:references|bibliography|literature\s+cited)\s*\}"
20
+ r".*?"
21
+ r"(?="
22
+ r"\\(?:chapter|section)\*?\s*(?:\[[^]]*\]\s*)?\{"
23
+ r"|\\end\s*\{document\}"
24
+ r"|\Z)",
25
+ re.IGNORECASE | re.DOTALL,
26
+ )
27
+ _BIBITEM_BLOCK = re.compile(
28
+ r"^[ \t]*\\bibitem(?:\[[^]]*\])?\s*\{[^}]*\}.*?"
29
+ r"(?="
30
+ r"^[ \t]*\\bibitem(?:\[[^]]*\])?\s*\{"
31
+ r"|^[ \t]*\\end\s*\{thebibliography\}"
32
+ r"|\Z)",
33
+ re.IGNORECASE | re.MULTILINE | re.DOTALL,
34
+ )
35
+ _BIBLIOGRAPHY_CONTROL = re.compile(
36
+ r"^[ \t]*\\(?:"
37
+ r"addbibresource|bibliography|bibliographystyle|printbibliography"
38
+ r")\b"
39
+ r"|^[ \t]*\\renewcommand\s*\{\\(?:bibname|refname)\}",
40
+ re.IGNORECASE,
41
+ )
42
+ _BIBLIOGRAPHY_EXACT_NAMES = {
43
+ "bib",
44
+ "biblio",
45
+ "bibliography",
46
+ "bibliographies",
47
+ "ref",
48
+ "refs",
49
+ "reference",
50
+ "references",
51
+ }
52
+ _BIBLIOGRAPHY_NAME_PARTS = {
53
+ "bib",
54
+ "biblio",
55
+ "bibliography",
56
+ "bibliographies",
57
+ "refs",
58
+ "references",
59
+ }
60
+
61
+
62
+ def _blank(match: re.Match[str]) -> str:
63
+ """Mask content without changing line numbers."""
64
+ return re.sub(r"[^\n]", " ", match.group(0))
65
+
66
+
67
+ def mask_bibliography(text: str) -> str:
68
+ """Blank bibliography regions and controls while preserving newlines."""
69
+ text = _BIBLIOGRAPHY_ENVIRONMENT.sub(_blank, text)
70
+ text = _BIBLIOGRAPHY_HEADING.sub(_blank, text)
71
+ text = _BIBITEM_BLOCK.sub(_blank, text)
72
+ return "\n".join(
73
+ "" if _BIBLIOGRAPHY_CONTROL.match(line) else line
74
+ for line in text.split("\n")
75
+ )
76
+
77
+
78
+ def is_bibliography_file(path: Path) -> bool:
79
+ """Recognize common filenames used only for reference lists."""
80
+ stem = path.stem.lower()
81
+ if stem in _BIBLIOGRAPHY_EXACT_NAMES:
82
+ return True
83
+ parts = {part for part in re.split(r"[-_.]+", stem) if part}
84
+ return bool(parts & _BIBLIOGRAPHY_NAME_PARTS)
@@ -0,0 +1,594 @@
1
+ #!/usr/bin/env python3
2
+ """Prepare compact TeX translation packets and safely merge their results."""
3
+
4
+ from __future__ import annotations
5
+
6
+ import argparse
7
+ from collections import Counter, defaultdict
8
+ import hashlib
9
+ import json
10
+ import math
11
+ import os
12
+ from pathlib import Path
13
+ import re
14
+ import tempfile
15
+ from typing import Iterable, Optional
16
+
17
+ from tex_translation_utils import is_bibliography_file, mask_bibliography
18
+
19
+
20
+ VERSION = 1
21
+ DEFAULT_TASK_DIR = ".translation-tasks"
22
+ INCLUDE = re.compile(r"\\(?:input|include)\s*\{([^}]+)\}")
23
+ COMMENT = re.compile(r"(?<!\\)%[^\n]*")
24
+ COMMAND = re.compile(r"\\(?:[A-Za-z@]+\*?|.)")
25
+ WORD = re.compile(r"[A-Za-z][A-Za-z'’-]*")
26
+ PLACEHOLDER = re.compile(r"⟪T\d{4}⟫")
27
+ TEXT_COMMAND = re.compile(
28
+ r"\\(?:title|subtitle|chapter|section|subsection|subsubsection|paragraph|"
29
+ r"subparagraph|caption|footnote|thanks|keyword|keywords)\*?\b",
30
+ re.IGNORECASE,
31
+ )
32
+ CONTROL_ONLY = re.compile(
33
+ r"^[ \t]*\\(?:documentclass|usepackage|RequirePackage|input|include|"
34
+ r"addbibresource|bibliography|bibliographystyle|includegraphics|label|"
35
+ r"setlength|newcommand|renewcommand|providecommand|def)\b",
36
+ re.IGNORECASE,
37
+ )
38
+ NON_TEXT_COMMAND = re.compile(
39
+ r"\\(?:documentclass|usepackage|RequirePackage|input|include|"
40
+ r"addbibresource|bibliography|bibliographystyle|includegraphics|label|"
41
+ r"ref|eqref|pageref|autoref|cite\w*|url|path)\*?"
42
+ r"(?:\s*\[[^]]*\])*\s*\{[^{}]*\}",
43
+ re.IGNORECASE,
44
+ )
45
+ MATH_ENVIRONMENT = re.compile(
46
+ r"\\begin\s*\{(?P<name>equation\*?|align\*?|alignat\*?|gather\*?|"
47
+ r"multline\*?|displaymath|math|eqnarray\*?)\}.*?"
48
+ r"\\end\s*\{(?P=name)\}",
49
+ re.IGNORECASE | re.DOTALL,
50
+ )
51
+ VERBATIM_ENVIRONMENT = re.compile(
52
+ r"\\begin\s*\{(?P<name>verbatim\*?|lstlisting|minted)\}.*?"
53
+ r"\\end\s*\{(?P=name)\}",
54
+ re.IGNORECASE | re.DOTALL,
55
+ )
56
+ DISPLAY_MATH = re.compile(r"\$\$.*?\$\$|\\\[.*?\\\]", re.DOTALL)
57
+ INLINE_MATH = re.compile(r"(?<!\\)\$(?!\$).*?(?<!\\)\$|\\\(.*?\\\)", re.DOTALL)
58
+ VERB = re.compile(r"\\verb(?P<delimiter>[^A-Za-z0-9\s]).*?(?P=delimiter)")
59
+ REFERENCE_COMMAND = re.compile(
60
+ r"\\(?:cite\w*|ref|eqref|pageref|autoref|label|url|path|"
61
+ r"includegraphics|input|include)\*?(?:\s*\[[^]]*\])*\s*\{[^{}]*\}",
62
+ re.IGNORECASE,
63
+ )
64
+ HREF_TARGET = re.compile(r"\\href\s*\{[^{}]*\}", re.IGNORECASE)
65
+ PACKET_BLOCK = re.compile(
66
+ r"^@@@ SEGMENT (?P<id>\S+)\n"
67
+ r"^@@@ SOURCE\n(?P<source>.*?)"
68
+ r"^@@@ TRANSLATION\n(?P<translation>.*?)"
69
+ r"^@@@ END(?:\n|\Z)",
70
+ re.MULTILINE | re.DOTALL,
71
+ )
72
+
73
+ PACKET_HEADER = """# 只在每个 TRANSLATION 区块填写简体中文译文;不要改 SOURCE 或标记行。
74
+ # 严格保留 LaTeX 结构及每个 ⟪T0000⟫ 形式的占位符;保留人名、模型名、数据集名、缩写、数值和引用键。
75
+ """
76
+
77
+
78
+ def _inside(path: Path, root: Path) -> bool:
79
+ try:
80
+ path.relative_to(root)
81
+ return True
82
+ except ValueError:
83
+ return False
84
+
85
+
86
+ def reachable(entry: Path, root: Path) -> list[Path]:
87
+ """Return TeX files reachable through input/include, including the entry."""
88
+ pending = [entry.resolve()]
89
+ seen: set[Path] = set()
90
+ while pending:
91
+ path = pending.pop()
92
+ if path in seen or not path.is_file() or not _inside(path, root):
93
+ continue
94
+ seen.add(path)
95
+ text = mask_bibliography(path.read_text(encoding="utf-8", errors="replace"))
96
+ for value in INCLUDE.findall(text):
97
+ child = (path.parent / value.strip()).resolve()
98
+ pending.append(child if child.suffix else child.with_suffix(".tex"))
99
+ return sorted(seen)
100
+
101
+
102
+ def _without_protected_text(text: str) -> str:
103
+ for pattern in (
104
+ VERBATIM_ENVIRONMENT,
105
+ MATH_ENVIRONMENT,
106
+ COMMENT,
107
+ DISPLAY_MATH,
108
+ INLINE_MATH,
109
+ VERB,
110
+ REFERENCE_COMMAND,
111
+ HREF_TARGET,
112
+ ):
113
+ text = pattern.sub(" ", text)
114
+ return text
115
+
116
+
117
+ def english_weight(text: str) -> int:
118
+ """Estimate visible English prose words, excluding common TeX controls."""
119
+ text = NON_TEXT_COMMAND.sub(" ", _without_protected_text(text))
120
+ text = COMMAND.sub(" ", text)
121
+ return len([word for word in WORD.findall(text) if len(word) > 1])
122
+
123
+
124
+ def _is_translatable(text: str) -> bool:
125
+ score = english_weight(text)
126
+ return score >= 3 or (score >= 1 and bool(TEXT_COMMAND.search(text) or "&" in text))
127
+
128
+
129
+ def _protect(source: str) -> tuple[str, list[dict[str, str]]]:
130
+ """Replace expensive immutable spans with checked, reversible placeholders."""
131
+ protected: list[dict[str, str]] = []
132
+ masked = source
133
+
134
+ def replace(match: re.Match[str]) -> str:
135
+ placeholder = f"⟪T{len(protected):04d}⟫"
136
+ protected.append({"placeholder": placeholder, "value": match.group(0)})
137
+ return placeholder
138
+
139
+ for pattern in (
140
+ VERBATIM_ENVIRONMENT,
141
+ MATH_ENVIRONMENT,
142
+ COMMENT,
143
+ DISPLAY_MATH,
144
+ INLINE_MATH,
145
+ VERB,
146
+ REFERENCE_COMMAND,
147
+ HREF_TARGET,
148
+ ):
149
+ masked = pattern.sub(replace, masked)
150
+ return masked, protected
151
+
152
+
153
+ def _paragraph_ranges(masked_text: str) -> Iterable[tuple[int, int]]:
154
+ """Yield zero-based, end-exclusive line ranges separated by hard controls."""
155
+ lines = masked_text.splitlines(keepends=True)
156
+ start: Optional[int] = None
157
+ for index, line in enumerate(lines):
158
+ separator = not line.strip() or bool(CONTROL_ONLY.match(line))
159
+ if separator:
160
+ if start is not None:
161
+ yield start, index
162
+ start = None
163
+ elif start is None:
164
+ start = index
165
+ if start is not None:
166
+ yield start, len(lines)
167
+
168
+
169
+ def _chunk_range(
170
+ lines: list[str], start: int, end: int, chunk_words: int
171
+ ) -> Iterable[tuple[int, int]]:
172
+ cursor = start
173
+ score = 0
174
+ for index in range(start, end):
175
+ score += english_weight(lines[index])
176
+ if score >= chunk_words and index + 1 < end:
177
+ yield cursor, index + 1
178
+ cursor = index + 1
179
+ score = 0
180
+ if cursor < end:
181
+ yield cursor, end
182
+
183
+
184
+ def segments_for(path: Path, root: Path, chunk_words: int) -> list[dict[str, object]]:
185
+ original = path.read_text(encoding="utf-8", errors="replace")
186
+ bibliography_masked = mask_bibliography(original)
187
+ original_lines = original.splitlines(keepends=True)
188
+ masked_lines = bibliography_masked.splitlines(keepends=True)
189
+ units: list[tuple[int, int, int]] = []
190
+
191
+ for paragraph_start, paragraph_end in _paragraph_ranges(bibliography_masked):
192
+ for start, end in _chunk_range(
193
+ masked_lines, paragraph_start, paragraph_end, chunk_words
194
+ ):
195
+ visible = "".join(masked_lines[start:end])
196
+ if not _is_translatable(visible):
197
+ continue
198
+ units.append((start, end, english_weight(visible)))
199
+
200
+ merged: list[list[int]] = []
201
+ for start, end, score in units:
202
+ if merged:
203
+ previous = merged[-1]
204
+ gap = "".join(original_lines[previous[1] : start])
205
+ if not gap.strip() and previous[2] + score <= chunk_words:
206
+ previous[1] = end
207
+ previous[2] += score
208
+ continue
209
+ merged.append([start, end, score])
210
+
211
+ segments: list[dict[str, object]] = []
212
+ for start, end, score in merged:
213
+ source = "".join(original_lines[start:end])
214
+ masked_source, protected = _protect(source)
215
+ packet_source = (
216
+ masked_source if masked_source.endswith("\n") else masked_source + "\n"
217
+ )
218
+ relative = str(path.relative_to(root))
219
+ identity = f"{relative}:{start + 1}:{end}:".encode() + source.encode()
220
+ segment_id = "s" + hashlib.sha256(identity).hexdigest()[:12]
221
+ segments.append(
222
+ {
223
+ "id": segment_id,
224
+ "path": relative,
225
+ "start_line": start + 1,
226
+ "end_line": end,
227
+ "weight": score,
228
+ "source": source,
229
+ "source_sha256": hashlib.sha256(source.encode()).hexdigest(),
230
+ "packet_source": packet_source,
231
+ "protected": protected,
232
+ }
233
+ )
234
+ return segments
235
+
236
+
237
+ def _allocate(
238
+ segments: list[dict[str, object]], requested_workers: int, min_words: int
239
+ ) -> list[list[dict[str, object]]]:
240
+ total = sum(int(segment["weight"]) for segment in segments)
241
+ useful_workers = max(1, math.ceil(total / max(1, min_words)))
242
+ count = min(max(1, requested_workers), useful_workers, max(1, len(segments)))
243
+ groups: list[list[dict[str, object]]] = [[] for _ in range(count)]
244
+ totals = [0] * count
245
+ for segment in sorted(segments, key=lambda item: int(item["weight"]), reverse=True):
246
+ index = min(range(count), key=totals.__getitem__)
247
+ groups[index].append(segment)
248
+ totals[index] += int(segment["weight"])
249
+ for group in groups:
250
+ group.sort(key=lambda item: (str(item["path"]), int(item["start_line"])))
251
+ return groups
252
+
253
+
254
+ def _packet_text(segments: list[dict[str, object]]) -> str:
255
+ parts = [PACKET_HEADER]
256
+ for segment in segments:
257
+ parts.extend(
258
+ [
259
+ f"@@@ SEGMENT {segment['id']}\n",
260
+ "@@@ SOURCE\n",
261
+ str(segment["packet_source"]),
262
+ "@@@ TRANSLATION\n",
263
+ "@@@ END\n",
264
+ ]
265
+ )
266
+ return "".join(parts)
267
+
268
+
269
+ def _write_atomic(path: Path, text: str) -> None:
270
+ path.parent.mkdir(parents=True, exist_ok=True)
271
+ handle = tempfile.NamedTemporaryFile(
272
+ "w", encoding="utf-8", dir=path.parent, delete=False
273
+ )
274
+ temporary = Path(handle.name)
275
+ try:
276
+ with handle:
277
+ handle.write(text)
278
+ os.replace(temporary, path)
279
+ finally:
280
+ if temporary.exists():
281
+ temporary.unlink()
282
+
283
+
284
+ def prepare(args: argparse.Namespace) -> int:
285
+ root = args.root.resolve()
286
+ if not root.is_dir():
287
+ raise SystemExit(f"paper root does not exist: {root}")
288
+ task_dir = (root / args.output).resolve()
289
+ if not _inside(task_dir, root):
290
+ raise SystemExit(f"task directory must stay below paper root: {task_dir}")
291
+ manifest_path = task_dir / "manifest.json"
292
+ if manifest_path.exists() and not args.force:
293
+ raise SystemExit(f"task manifest already exists: {manifest_path}; use --force")
294
+ if args.force and task_dir.is_dir():
295
+ for stale_packet in task_dir.glob("worker-*.task"):
296
+ stale_packet.unlink()
297
+
298
+ if args.entry:
299
+ entry = (root / args.entry).resolve()
300
+ if not entry.is_file():
301
+ raise SystemExit(f"entry file does not exist: {entry}")
302
+ files = reachable(entry, root)
303
+ else:
304
+ entry = None
305
+ files = sorted(root.rglob("*.tex"))
306
+ files = [path for path in files if path.is_file() and not is_bibliography_file(path)]
307
+ segments = [
308
+ segment
309
+ for path in files
310
+ for segment in segments_for(path, root, args.chunk_words)
311
+ ]
312
+ if not segments:
313
+ raise SystemExit(f"no translatable TeX prose found below {root}")
314
+
315
+ groups = _allocate(segments, args.workers, args.min_words_per_worker)
316
+ packets = []
317
+ for index, group in enumerate(groups, 1):
318
+ name = f"worker-{index:02d}.task"
319
+ _write_atomic(task_dir / name, _packet_text(group))
320
+ packets.append(
321
+ {
322
+ "worker": index,
323
+ "path": name,
324
+ "weight": sum(int(segment["weight"]) for segment in group),
325
+ "segments": [str(segment["id"]) for segment in group],
326
+ }
327
+ )
328
+
329
+ scanned_bytes = sum(path.stat().st_size for path in files)
330
+ selected_bytes = sum(len(str(segment["source"]).encode()) for segment in segments)
331
+ packet_source_bytes = sum(
332
+ len(str(segment["packet_source"]).encode()) for segment in segments
333
+ )
334
+ reduction = (
335
+ round(100 * (1 - packet_source_bytes / scanned_bytes), 1)
336
+ if scanned_bytes
337
+ else 0.0
338
+ )
339
+ manifest = {
340
+ "version": VERSION,
341
+ "root": str(root),
342
+ "entry": str(entry.relative_to(root)) if entry else None,
343
+ "packets": packets,
344
+ "segments": segments,
345
+ "metrics": {
346
+ "files": len(files),
347
+ "segments": len(segments),
348
+ "workers": len(groups),
349
+ "visible_english_words": sum(int(segment["weight"]) for segment in segments),
350
+ "scanned_bytes": scanned_bytes,
351
+ "selected_bytes": selected_bytes,
352
+ "packet_source_bytes": packet_source_bytes,
353
+ "input_byte_reduction_percent": reduction,
354
+ },
355
+ }
356
+ _write_atomic(manifest_path, json.dumps(manifest, ensure_ascii=False, indent=2) + "\n")
357
+ public_packets = [
358
+ {
359
+ "worker": packet["worker"],
360
+ "path": packet["path"],
361
+ "weight": packet["weight"],
362
+ "segments": len(packet["segments"]),
363
+ }
364
+ for packet in packets
365
+ ]
366
+ payload = {
367
+ "task_dir": str(task_dir),
368
+ **manifest["metrics"],
369
+ "packets": public_packets,
370
+ }
371
+ if args.json:
372
+ print(json.dumps(payload, ensure_ascii=False, indent=2))
373
+ else:
374
+ print(
375
+ f"prepared {len(segments)} segments for {len(groups)} worker(s); "
376
+ f"estimated input reduction {reduction:.1f}%"
377
+ )
378
+ for packet in packets:
379
+ print(f"{packet['path']}: words={packet['weight']} segments={len(packet['segments'])}")
380
+ return 0
381
+
382
+
383
+ def parse_packet(path: Path) -> dict[str, dict[str, str]]:
384
+ text = path.read_text(encoding="utf-8")
385
+ result: dict[str, dict[str, str]] = {}
386
+ for match in PACKET_BLOCK.finditer(text):
387
+ segment_id = match.group("id")
388
+ if segment_id in result:
389
+ raise ValueError(f"duplicate segment {segment_id} in {path}")
390
+ result[segment_id] = {
391
+ "source": match.group("source"),
392
+ "translation": match.group("translation"),
393
+ }
394
+ return result
395
+
396
+
397
+ def _load_manifest(root: Path, output: Path) -> tuple[Path, dict[str, object]]:
398
+ task_dir = (root / output).resolve()
399
+ if not _inside(task_dir, root):
400
+ raise SystemExit(f"task directory must stay below paper root: {task_dir}")
401
+ manifest_path = task_dir / "manifest.json"
402
+ if not manifest_path.is_file():
403
+ raise SystemExit(f"task manifest does not exist: {manifest_path}")
404
+ manifest = json.loads(manifest_path.read_text(encoding="utf-8"))
405
+ if manifest.get("version") != VERSION:
406
+ raise SystemExit(f"unsupported task manifest version: {manifest.get('version')}")
407
+ return task_dir, manifest
408
+
409
+
410
+ def _read_results(
411
+ task_dir: Path, manifest: dict[str, object]
412
+ ) -> tuple[dict[str, dict[str, str]], list[str]]:
413
+ results: dict[str, dict[str, str]] = {}
414
+ errors: list[str] = []
415
+ for packet in manifest["packets"]: # type: ignore[index]
416
+ packet_path = task_dir / packet["path"]
417
+ try:
418
+ parsed = parse_packet(packet_path)
419
+ except (OSError, ValueError) as error:
420
+ errors.append(str(error))
421
+ continue
422
+ for segment_id, value in parsed.items():
423
+ if segment_id in results:
424
+ errors.append(f"duplicate result for {segment_id}")
425
+ results[segment_id] = value
426
+ return results, errors
427
+
428
+
429
+ def status(args: argparse.Namespace) -> int:
430
+ root = args.root.resolve()
431
+ task_dir, manifest = _load_manifest(root, args.output)
432
+ results, errors = _read_results(task_dir, manifest)
433
+ expected = {str(segment["id"]) for segment in manifest["segments"]} # type: ignore[index]
434
+ completed = {
435
+ segment_id
436
+ for segment_id, value in results.items()
437
+ if value["translation"].strip()
438
+ }
439
+ missing = sorted(expected - completed)
440
+ payload = {
441
+ "segments": len(expected),
442
+ "completed": len(expected) - len(missing),
443
+ "missing": missing,
444
+ "errors": errors,
445
+ }
446
+ if args.json:
447
+ print(json.dumps(payload, ensure_ascii=False, indent=2))
448
+ else:
449
+ print(f"completed={payload['completed']}/{payload['segments']}")
450
+ for error in errors:
451
+ print(f"error: {error}")
452
+ for segment_id in missing:
453
+ print(f"missing: {segment_id}")
454
+ return 0 if not missing and not errors else 1
455
+
456
+
457
+ def _restore(translation: str, protected: list[dict[str, str]]) -> str:
458
+ for item in protected:
459
+ translation = translation.replace(item["placeholder"], item["value"])
460
+ return translation
461
+
462
+
463
+ def _structure_signature(text: str) -> tuple[Counter[str], int, int]:
464
+ return Counter(COMMAND.findall(text)), text.count("{"), text.count("}")
465
+
466
+
467
+ def apply(args: argparse.Namespace) -> int:
468
+ root = args.root.resolve()
469
+ task_dir, manifest = _load_manifest(root, args.output)
470
+ results, errors = _read_results(task_dir, manifest)
471
+ replacements: defaultdict[Path, list[tuple[int, int, str]]] = defaultdict(list)
472
+ expected_sources = {
473
+ (str(segment["path"]), int(segment["start_line"]), int(segment["end_line"])): (
474
+ str(segment["source"]),
475
+ str(segment["source_sha256"]),
476
+ )
477
+ for segment in manifest["segments"] # type: ignore[index]
478
+ }
479
+
480
+ for segment in manifest["segments"]: # type: ignore[index]
481
+ segment_id = str(segment["id"])
482
+ result = results.get(segment_id)
483
+ if result is None:
484
+ errors.append(f"missing segment {segment_id}")
485
+ continue
486
+ if result["source"] != segment["packet_source"]:
487
+ errors.append(f"SOURCE was modified for {segment_id}")
488
+ continue
489
+ translation = result["translation"].lstrip("\n")
490
+ if not translation.strip():
491
+ errors.append(f"empty translation for {segment_id}")
492
+ continue
493
+ expected_placeholders = Counter(
494
+ item["placeholder"] for item in segment["protected"]
495
+ )
496
+ actual_placeholders = Counter(PLACEHOLDER.findall(translation))
497
+ if actual_placeholders != expected_placeholders:
498
+ errors.append(f"placeholder mismatch for {segment_id}")
499
+ continue
500
+ if _structure_signature(translation) != _structure_signature(
501
+ str(segment["packet_source"])
502
+ ):
503
+ errors.append(f"LaTeX structure mismatch for {segment_id}")
504
+ continue
505
+
506
+ source = str(segment["source"])
507
+ if source.endswith("\n"):
508
+ if not translation.endswith("\n"):
509
+ translation += "\n"
510
+ else:
511
+ translation = translation.rstrip("\n")
512
+ restored = _restore(translation, segment["protected"])
513
+ path = root / str(segment["path"])
514
+ replacements[path].append(
515
+ (int(segment["start_line"]), int(segment["end_line"]), restored)
516
+ )
517
+
518
+ for path, items in replacements.items():
519
+ text = path.read_text(encoding="utf-8", errors="replace")
520
+ lines = text.splitlines(keepends=True)
521
+ for start, end, _translation in items:
522
+ current = "".join(lines[start - 1 : end])
523
+ expected, expected_hash = expected_sources[
524
+ (str(path.relative_to(root)), start, end)
525
+ ]
526
+ if (
527
+ current != expected
528
+ or hashlib.sha256(current.encode()).hexdigest() != expected_hash
529
+ ):
530
+ errors.append(f"source changed since prepare: {path.relative_to(root)}:{start}-{end}")
531
+
532
+ if errors:
533
+ for error in errors:
534
+ print(f"error: {error}")
535
+ return 1
536
+ if args.check:
537
+ print(f"validated {sum(len(items) for items in replacements.values())} translations")
538
+ return 0
539
+
540
+ for path, items in replacements.items():
541
+ lines = path.read_text(encoding="utf-8", errors="replace").splitlines(keepends=True)
542
+ for start, end, translation in sorted(items, reverse=True):
543
+ lines[start - 1 : end] = [translation]
544
+ _write_atomic(path, "".join(lines))
545
+ print(
546
+ f"applied {sum(len(items) for items in replacements.values())} translations "
547
+ f"to {len(replacements)} file(s)"
548
+ )
549
+ return 0
550
+
551
+
552
+ def build_parser() -> argparse.ArgumentParser:
553
+ parser = argparse.ArgumentParser(description=__doc__)
554
+ subparsers = parser.add_subparsers(dest="command", required=True)
555
+
556
+ prepare_parser = subparsers.add_parser("prepare", help="create compact worker packets")
557
+ prepare_parser.add_argument("root", type=Path)
558
+ prepare_parser.add_argument("--entry", type=Path)
559
+ prepare_parser.add_argument("--workers", type=int, default=3)
560
+ prepare_parser.add_argument("--chunk-words", type=int, default=900)
561
+ prepare_parser.add_argument("--min-words-per-worker", type=int, default=1200)
562
+ prepare_parser.add_argument("--output", type=Path, default=Path(DEFAULT_TASK_DIR))
563
+ prepare_parser.add_argument("--force", action="store_true")
564
+ prepare_parser.add_argument("--json", action="store_true")
565
+ prepare_parser.set_defaults(handler=prepare)
566
+
567
+ status_parser = subparsers.add_parser("status", help="report packet completion")
568
+ status_parser.add_argument("root", type=Path)
569
+ status_parser.add_argument("--output", type=Path, default=Path(DEFAULT_TASK_DIR))
570
+ status_parser.add_argument("--json", action="store_true")
571
+ status_parser.set_defaults(handler=status)
572
+
573
+ apply_parser = subparsers.add_parser("apply", help="validate and merge packet translations")
574
+ apply_parser.add_argument("root", type=Path)
575
+ apply_parser.add_argument("--output", type=Path, default=Path(DEFAULT_TASK_DIR))
576
+ apply_parser.add_argument("--check", action="store_true")
577
+ apply_parser.set_defaults(handler=apply)
578
+ return parser
579
+
580
+
581
+ def main() -> int:
582
+ parser = build_parser()
583
+ args = parser.parse_args()
584
+ if getattr(args, "workers", 1) < 1:
585
+ parser.error("--workers must be positive")
586
+ if getattr(args, "chunk_words", 1) < 1:
587
+ parser.error("--chunk-words must be positive")
588
+ if getattr(args, "min_words_per_worker", 1) < 1:
589
+ parser.error("--min-words-per-worker must be positive")
590
+ return int(args.handler(args))
591
+
592
+
593
+ if __name__ == "__main__":
594
+ raise SystemExit(main())