arxiv-paper-zh 0.1.2 → 0.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.codex-plugin/plugin.json +1 -1
- package/README.md +17 -7
- package/package.json +3 -3
- package/skills/arxiv-paper-zh/SKILL.md +57 -53
- package/skills/arxiv-paper-zh/scripts/audit_tex_translation.py +9 -2
- package/skills/arxiv-paper-zh/scripts/inventory_and_shard.py +9 -4
- package/skills/arxiv-paper-zh/scripts/tex_translation_utils.py +84 -0
- package/skills/arxiv-paper-zh/scripts/translation_tasks.py +594 -0
package/README.md
CHANGED
|
@@ -33,9 +33,11 @@ npm 包地址:[arxiv-paper-zh](https://www.npmjs.com/package/arxiv-paper-zh)
|
|
|
33
33
|
|
|
34
34
|
- 核对论文简称、标题、作者、摘要、URL 与 arXiv ID,避免翻错同名论文。
|
|
35
35
|
- 按 `arxiv-paper/<paper-name>/` 统一保存原始源码、中文源码以及中英文 PDF。
|
|
36
|
-
-
|
|
36
|
+
- 将可见英文切成紧凑任务包,公式、引用、URL、代码和注释使用可逆占位符,不再让模型重复读取整份 TeX。
|
|
37
|
+
- 按词量自适应分配 worker;长论文并行、短论文自动减少 worker,单文件论文也能分片。
|
|
37
38
|
- 保留公式、数值、引用键、标签、人名、模型名、数据集名和常用缩写。
|
|
38
39
|
- 翻译正文、章节标题、脚注、表头、表注和 caption。
|
|
40
|
+
- 参考文献标题与条目保持原文,并在翻译分片和漏译审计中自动跳过。
|
|
39
41
|
- 批量检查并安装缺失的 TeX 宏包,复用共享 TinyTeX/TeX Live 缓存。
|
|
40
42
|
- 自动执行漏译审计、BibTeX/Biber 构建和引用收敛检查。
|
|
41
43
|
- 使用 XeLaTeX 生成中文 PDF,并要求逐页视觉核验。
|
|
@@ -64,13 +66,17 @@ arxiv-paper-zh/
|
|
|
64
66
|
│ │ ├── build_and_check.py
|
|
65
67
|
│ │ ├── inventory_and_shard.py
|
|
66
68
|
│ │ ├── prepare_output_layout.py
|
|
67
|
-
│ │
|
|
69
|
+
│ │ ├── prepare_tex_runtime.py
|
|
70
|
+
│ │ ├── translation_tasks.py
|
|
71
|
+
│ │ └── tex_translation_utils.py
|
|
68
72
|
│ └── references/
|
|
69
73
|
│ └── paper-translation-packages.txt
|
|
70
74
|
├── install.sh
|
|
71
75
|
├── package.json
|
|
72
76
|
├── tests/
|
|
73
|
-
│
|
|
77
|
+
│ ├── installer.test.mjs
|
|
78
|
+
│ ├── test_bibliography_exclusion.py
|
|
79
|
+
│ └── test_translation_tasks.py
|
|
74
80
|
├── .gitignore
|
|
75
81
|
└── README.md
|
|
76
82
|
```
|
|
@@ -237,7 +243,7 @@ arxiv-paper/EST/
|
|
|
237
243
|
|
|
238
244
|
`paper-name` 使用用户熟悉的简短名称并保留大小写,例如 `EST`、`Onetrans`,且只能包含英文字母、数字、点、下划线和连字符。
|
|
239
245
|
|
|
240
|
-
|
|
246
|
+
整个参考文献部分保持原文,包括标题和全部条目;`.bib`、`.bbl`、内嵌 bibliography 环境和单独的参考文献 TeX 文件均不参与翻译。正文中的文献综述仍照常翻译。原图内部文字不修改,只翻译必要 caption。公式内的英文说明按“公式不变”原则保留。
|
|
241
247
|
|
|
242
248
|
## 内置工具
|
|
243
249
|
|
|
@@ -245,8 +251,12 @@ arxiv-paper/EST/
|
|
|
245
251
|
# 创建并输出固定的论文产物路径
|
|
246
252
|
python3 skills/arxiv-paper-zh/scripts/prepare_output_layout.py EST --root arxiv-paper
|
|
247
253
|
|
|
248
|
-
#
|
|
249
|
-
python3 skills/arxiv-paper-zh/scripts/
|
|
254
|
+
# 生成紧凑翻译任务包(3 是 worker 上限,小论文会自动减少)
|
|
255
|
+
python3 skills/arxiv-paper-zh/scripts/translation_tasks.py prepare arxiv-paper/EST/latex/paper-zh --entry main.tex --workers 3 --json
|
|
256
|
+
|
|
257
|
+
# worker 填写各自的 worker-*.task 后,检查进度并安全合并
|
|
258
|
+
python3 skills/arxiv-paper-zh/scripts/translation_tasks.py status arxiv-paper/EST/latex/paper-zh
|
|
259
|
+
python3 skills/arxiv-paper-zh/scripts/translation_tasks.py apply arxiv-paper/EST/latex/paper-zh
|
|
250
260
|
|
|
251
261
|
# 检查预装包和论文专用宏包
|
|
252
262
|
python3 skills/arxiv-paper-zh/scripts/prepare_tex_runtime.py arxiv-paper/EST/latex/paper-zh --preset --kpsewhich /path/to/kpsewhich
|
|
@@ -267,7 +277,7 @@ python3 skills/arxiv-paper-zh/scripts/build_and_check.py arxiv-paper/EST/latex/p
|
|
|
267
277
|
- 是否允许网络下载与安装宏包。
|
|
268
278
|
- 是否提供可写文件系统和本地 XeLaTeX。
|
|
269
279
|
|
|
270
|
-
客户端不支持 subagent 时,Agent
|
|
280
|
+
客户端不支持 subagent 时,Agent 应顺序填写任务包。任务包仍会剔除不需要发送给模型的内容,因此同样能降低上下文开销;只是不具备并行加速。
|
|
271
281
|
|
|
272
282
|
## 开发与发布检查
|
|
273
283
|
|
package/package.json
CHANGED
|
@@ -1,10 +1,10 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "arxiv-paper-zh",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.2.0",
|
|
4
4
|
"description": "Install the arxiv-paper-zh Agent Skill for Codex, Claude Code, and compatible agents.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
|
7
|
-
"arxiv-paper-zh": "
|
|
7
|
+
"arxiv-paper-zh": "bin/arxiv-paper-zh.mjs"
|
|
8
8
|
},
|
|
9
9
|
"files": [
|
|
10
10
|
".codex-plugin/",
|
|
@@ -17,7 +17,7 @@
|
|
|
17
17
|
],
|
|
18
18
|
"scripts": {
|
|
19
19
|
"test": "node --test tests/*.test.mjs",
|
|
20
|
-
"check": "node --check bin/arxiv-paper-zh.mjs && python3 -c \"from pathlib import Path; [compile(path.read_bytes(), str(path), 'exec') for path in Path('skills/arxiv-paper-zh/scripts').glob('*.py')]\" && npm test"
|
|
20
|
+
"check": "node --check bin/arxiv-paper-zh.mjs && python3 -c \"from pathlib import Path; [compile(path.read_bytes(), str(path), 'exec') for path in Path('skills/arxiv-paper-zh/scripts').glob('*.py')]\" && python3 -m unittest discover -s tests -p 'test_*.py' && npm test"
|
|
21
21
|
},
|
|
22
22
|
"engines": {
|
|
23
23
|
"node": ">=18"
|
|
@@ -1,15 +1,15 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: arxiv-paper-zh
|
|
3
|
-
description: Verify
|
|
3
|
+
description: Verify an arXiv or LaTeX paper's identity, then produce complete, compilable English and Chinese sources and PDFs. Use for locating and translating research papers while preserving math, citations, data, and layout; the workflow uses compact protected translation packets, adaptive parallel workers, cached TeX dependencies, audits, XeLaTeX compilation, and full-page visual verification.
|
|
4
4
|
---
|
|
5
5
|
|
|
6
6
|
# arXiv 论文快速中文化
|
|
7
7
|
|
|
8
|
-
|
|
8
|
+
交付保留数学、数据、引用和作者信息的中英文源码与经过页面核验的 PDF。身份正确性和译文完整性优先;翻译阶段使用紧凑任务包,避免将整份 TeX 和重复会话上下文发送给每个 worker。
|
|
9
9
|
|
|
10
|
-
##
|
|
10
|
+
## 输出目录
|
|
11
11
|
|
|
12
|
-
|
|
12
|
+
全部产物固定放在 `arxiv-paper/<paper-name>/`:
|
|
13
13
|
|
|
14
14
|
```text
|
|
15
15
|
arxiv-paper/<paper-name>/
|
|
@@ -17,81 +17,85 @@ arxiv-paper/<paper-name>/
|
|
|
17
17
|
│ ├── source.tar
|
|
18
18
|
│ ├── paper-en/
|
|
19
19
|
│ └── paper-zh/
|
|
20
|
-
├── paper-en
|
|
21
|
-
|
|
22
|
-
└── paper-zh/
|
|
23
|
-
└── <paper-name>-zh.pdf
|
|
20
|
+
├── paper-en/<paper-name>-en.pdf
|
|
21
|
+
└── paper-zh/<paper-name>-zh.pdf
|
|
24
22
|
```
|
|
25
23
|
|
|
26
|
-
`
|
|
24
|
+
`paper-name` 优先使用用户简称并保留大小写,且必须匹配 `[A-Za-z0-9][A-Za-z0-9._-]*`。开始下载前运行:
|
|
27
25
|
|
|
28
26
|
```bash
|
|
29
27
|
python3 scripts/prepare_output_layout.py <paper-name> --root arxiv-paper
|
|
30
28
|
```
|
|
31
29
|
|
|
32
|
-
|
|
30
|
+
以脚本输出的绝对路径为准,不另建平行交付目录。`.translation-tasks/`、渲染图和日志属于论文目录内的临时文件,不得混入最终 PDF 目录。
|
|
33
31
|
|
|
34
|
-
##
|
|
32
|
+
## 工作流
|
|
35
33
|
|
|
36
|
-
|
|
34
|
+
1. 从 arXiv 摘要页或 API 核验规范化 ID、完整标题、作者、摘要、版本和日期。简称不是唯一标识;候选不唯一时让用户确认。向用户明确标题、作者和 arXiv ID。
|
|
35
|
+
2. 创建输出目录,将 `https://export.arxiv.org/e-print/<ID>` 保存为 `latex/source.tar` 并解压到 `latex/paper-en/`。用主 TeX 的标题、作者或 README 二次核验身份;不一致时停止。
|
|
36
|
+
3. 将英文源码内容完整复制到 `latex/paper-zh/`,不得多套一层目录;此后只修改中文副本。分别定位中英文入口文件。
|
|
37
|
+
4. 先完成中文入口的 XeLaTeX/ctex 改造和模板可见字符串本地化,再生成翻译任务。通常加入 `\usepackage[UTF8,fontset=fandol]{ctex}`,移除仅适用于 pdfLaTeX 的 `inputenc` 和 T1 `fontenc`。生成任务后、合并任务前不要再编辑中文 TeX 源码,合并器会检查快照。
|
|
38
|
+
5. 生成紧凑任务包。脚本包含入口文件,因此单文件论文也能按片段并行;它自动省略参考文献,并用可逆占位符保护公式、引用、URL、代码和注释:
|
|
37
39
|
|
|
38
|
-
```bash
|
|
39
|
-
python3 scripts/
|
|
40
|
-
|
|
41
|
-
|
|
42
|
-
```
|
|
43
|
-
|
|
44
|
-
预装清单位于 `references/paper-translation-packages.txt`,覆盖 XeLaTeX 中文排版、BibTeX/Biber、常见数学与字体、表格、算法、代码、绘图、caption,以及 ACM、IEEE、Elsevier、Springer 模板。首次安装需要联网;后续任务先离线检查,只有论文特有依赖缺失时才联网批量补装。不要安装 `scheme-full`。
|
|
45
|
-
|
|
46
|
-
## 快速流水线
|
|
47
|
-
|
|
48
|
-
1. 从 arXiv 摘要页或 API 核验规范化 ID、完整标题、作者、摘要、版本和日期。简称不是唯一标识。只有一个候选与用户主题高度一致时才继续,否则先让用户确认。明确告知用户“标题 + 作者 + arXiv ID”。
|
|
49
|
-
2. 确定并校验 `paper-name`,运行 `scripts/prepare_output_layout.py`。从 `https://export.arxiv.org/e-print/<ID>` 下载原始压缩包到 `latex/source.tar`,解压到 `latex/paper-en/`。用主 TeX 的 `\title{}`、作者或 README 做第二次身份核验;不一致时停止。
|
|
50
|
-
3. 将 `latex/paper-en/` 的全部内容完整复制到 `latex/paper-zh/`,不得产生 `latex/paper-zh/paper-en/` 额外嵌套;此后只修改中文副本。分别定位英文和中文入口文件。
|
|
51
|
-
4. 立即并行启动两条路径:
|
|
40
|
+
```bash
|
|
41
|
+
python3 scripts/translation_tasks.py prepare \
|
|
42
|
+
arxiv-paper/<paper-name>/latex/paper-zh \
|
|
43
|
+
--entry main.tex --workers 3 --json
|
|
44
|
+
```
|
|
52
45
|
|
|
53
|
-
-
|
|
54
|
-
|
|
46
|
+
默认 `--workers 3` 是上限;小论文会自动减少 worker,避免启动开销。只有需要改变速度/上下文折中才调整 `--chunk-words` 或 `--min-words-per-worker`。
|
|
47
|
+
6. 每个 `worker-*.task` 只交给一个 worker。支持隔离上下文时使用空/最小历史,而不是复制完整会话;任务提示只需:
|
|
55
48
|
|
|
56
|
-
```
|
|
57
|
-
|
|
58
|
-
|
|
49
|
+
```text
|
|
50
|
+
翻译 <packet> 的全部 SOURCE 区块,把译文填入对应 TRANSLATION 区块。
|
|
51
|
+
严格遵守文件头规则,只编辑该任务包,不读取或修改论文源码。
|
|
59
52
|
```
|
|
60
53
|
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
7. 翻译所有渲染的英文自然语言:标题、摘要、正文、章节标题、列表、脚注、表头、表注和 caption。保留公式、数学符号、数值、引用键、label/ref、URL、LaTeX 结构、人名、模型名、数据集名、缩写与通行技术标识。图片只翻译必要 caption;算法/代码只翻译自然语言注释、docstring、caption 和说明。参考文献元数据默认不翻译。
|
|
64
|
-
8. 子代理完成后,主代理只运行一次全局审计并人工复核全部命中;不得仅扫描顶层章节,也不得把子代理的自检当作最终证明:
|
|
54
|
+
Worker 不需要读取本 Skill、整篇论文或其他任务包。不同 worker 并行编辑各自任务包;不支持 subagent 时顺序处理。主 agent 同时编译英文版、检查中文依赖,但不修改已快照的中文 TeX。
|
|
55
|
+
7. Worker 完成后只运行一次合并。合并器先整体校验任务完整性、占位符、LaTeX 结构和源文件哈希,任何错误都会在写文件前停止:
|
|
65
56
|
|
|
66
57
|
```bash
|
|
58
|
+
python3 scripts/translation_tasks.py status arxiv-paper/<paper-name>/latex/paper-zh
|
|
59
|
+
python3 scripts/translation_tasks.py apply arxiv-paper/<paper-name>/latex/paper-zh
|
|
67
60
|
python3 scripts/audit_tex_translation.py arxiv-paper/<paper-name>/latex/paper-zh
|
|
68
61
|
```
|
|
69
62
|
|
|
70
|
-
|
|
63
|
+
`status` 未完成时只返工列出的任务;`apply` 报错时只检查对应 segment,不重新读取或重译全部论文。主 agent 必须人工复核全局审计命中。
|
|
64
|
+
8. 使用原论文声明或兼容引擎编译英文源码。中文使用自动构建脚本识别 BibTeX/Biber 并完成 XeLaTeX 收敛:
|
|
71
65
|
|
|
72
66
|
```bash
|
|
73
|
-
python3 scripts/build_and_check.py
|
|
67
|
+
python3 scripts/build_and_check.py \
|
|
68
|
+
arxiv-paper/<paper-name>/latex/paper-zh/main.tex --tex-bin /path/to/tex/bin
|
|
74
69
|
```
|
|
75
70
|
|
|
76
|
-
|
|
77
|
-
10.
|
|
78
|
-
11. 将英文成品复制为 `paper-en/<paper-name>-en.pdf`,将中文成品复制为 `paper-zh/<paper-name>-zh.pdf`。回复列出完整标题、作者、arXiv ID、`latex/paper-en/`、`latex/paper-zh/` 以及两个 PDF 的绝对路径。
|
|
71
|
+
9. 分别低分辨率渲染中英文 PDF 的全部页面并检查裁切、重叠、溢出、图片和页数;只对可疑页面高分辨率渲染。另用支持 CJK 的系统 PDF 引擎抽查中文字体。Poppler 缺少 CMap 时不得把空白中文误判为正常。
|
|
72
|
+
10. 将英文成品复制为 `paper-en/<paper-name>-en.pdf`,中文成品复制为 `paper-zh/<paper-name>-zh.pdf`。回复列出论文身份、两套源码目录和两个 PDF 的绝对路径。
|
|
79
73
|
|
|
80
|
-
##
|
|
74
|
+
## 翻译边界
|
|
81
75
|
|
|
82
|
-
-
|
|
83
|
-
-
|
|
84
|
-
-
|
|
85
|
-
-
|
|
86
|
-
-
|
|
76
|
+
- 翻译标题、摘要、正文、章节标题、列表、脚注、表头、表注和 caption。
|
|
77
|
+
- 保留公式、数学符号、数值、引用键、label/ref、URL、LaTeX 结构、人名、模型名、数据集名、缩写与通行技术标识。
|
|
78
|
+
- 图片内部文字默认不修改;算法和代码块(含内嵌注释/docstring)默认保持原样,只翻译 caption 与外部说明。用户明确要求时再单独处理代码内文本。
|
|
79
|
+
- 参考文献元数据保持原文;`.bib`、`.bbl`、内嵌 bibliography 环境和独立参考文献 TeX 文件均不进入任务包。
|
|
80
|
+
- 不把完整 TeX 内容粘贴进 agent 提示,不让 worker 直接编辑论文文件,也不让多个 worker 处理同一任务包。
|
|
87
81
|
|
|
88
|
-
##
|
|
82
|
+
## 共享 TeX 运行时
|
|
83
|
+
|
|
84
|
+
优先复用固定位置的 TinyTeX/TeX Live。首次建立或主动刷新时,使用 `references/paper-translation-packages.txt` 一次检查并批量安装,不安装 `scheme-full`:
|
|
85
|
+
|
|
86
|
+
```bash
|
|
87
|
+
python3 scripts/prepare_tex_runtime.py --preset \
|
|
88
|
+
--kpsewhich <shared-tex-root>/bin/<platform>/kpsewhich \
|
|
89
|
+
--tlmgr <shared-tex-root>/bin/<platform>/tlmgr --install
|
|
90
|
+
```
|
|
89
91
|
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
- 中英文 PDF 的全部页面均已检查,中文字体由 CJK 能力正常的渲染器确认。
|
|
92
|
+
后续先离线检查;论文特有依赖缺失时再用同一脚本和一次 `tlmgr` 调用批量补装。只有没有可用运行时时才在可写缓存目录建立便携 TinyTeX,不使用 `sudo`。
|
|
93
|
+
|
|
94
|
+
## 完成标准
|
|
94
95
|
|
|
95
|
-
|
|
96
|
+
- arXiv 元数据与源码身份两次核验通过,`latex/paper-en/` 未修改。
|
|
97
|
+
- 中文任务全部合并,全局漏译审计已人工复核;所有可见自然语言均已翻译或明确允许保留。
|
|
98
|
+
- 中英文 PDF 均构建成功,参考文献和交叉引用收敛,中文版无缺字。
|
|
99
|
+
- 中英文 PDF 全部页面均已检查,中文字体经 CJK 能力正常的渲染器确认。
|
|
96
100
|
|
|
97
|
-
|
|
101
|
+
方案参考科学空间文章[《让 AI 翻译一篇完整的论文》](https://spaces.ac.cn/archives/11578),并结合紧凑任务包、并行 agent、依赖缓存和自动审计实现。
|
|
@@ -7,13 +7,19 @@ import argparse
|
|
|
7
7
|
import re
|
|
8
8
|
from pathlib import Path
|
|
9
9
|
|
|
10
|
+
from tex_translation_utils import is_bibliography_file, mask_bibliography
|
|
11
|
+
|
|
10
12
|
|
|
11
13
|
PROSE = re.compile(r"[A-Za-z]{4,}(?:[ \t]+[A-Za-z][A-Za-z'’-]{2,}){2,}")
|
|
12
14
|
COMMAND_ONLY = re.compile(r"^[ \t]*\\(?:usepackage|documentclass|input|include|addbibresource)\b")
|
|
13
15
|
|
|
14
16
|
|
|
15
17
|
def tex_files(root: Path) -> list[Path]:
|
|
16
|
-
return sorted(
|
|
18
|
+
return sorted(
|
|
19
|
+
path
|
|
20
|
+
for path in root.rglob("*.tex")
|
|
21
|
+
if path.is_file() and not is_bibliography_file(path)
|
|
22
|
+
)
|
|
17
23
|
|
|
18
24
|
|
|
19
25
|
def visible_part(line: str) -> str:
|
|
@@ -45,7 +51,8 @@ def main() -> int:
|
|
|
45
51
|
|
|
46
52
|
hits = 0
|
|
47
53
|
for path in files:
|
|
48
|
-
|
|
54
|
+
text = mask_bibliography(path.read_text(encoding="utf-8", errors="replace"))
|
|
55
|
+
for number, raw in enumerate(text.splitlines(), 1):
|
|
49
56
|
line = visible_part(raw)
|
|
50
57
|
if not line.strip() or COMMAND_ONLY.match(line):
|
|
51
58
|
continue
|
|
@@ -1,8 +1,11 @@
|
|
|
1
1
|
#!/usr/bin/env python3
|
|
2
|
-
"""Inventory
|
|
2
|
+
"""Inventory translatable TeX sources and greedily balance translation shards."""
|
|
3
3
|
from __future__ import annotations
|
|
4
4
|
import argparse, json, re
|
|
5
5
|
from pathlib import Path
|
|
6
|
+
|
|
7
|
+
from tex_translation_utils import is_bibliography_file, mask_bibliography
|
|
8
|
+
|
|
6
9
|
INCLUDE = re.compile(r"\\(?:input|include)\s*\{([^}]+)\}")
|
|
7
10
|
COMMENT = re.compile(r"(?<!\\)%.*$")
|
|
8
11
|
COMMAND = re.compile(r"\\[A-Za-z@]+\*?(?:\[[^]]*\])?")
|
|
@@ -13,17 +16,19 @@ def reachable(entry: Path, root: Path) -> list[Path]:
|
|
|
13
16
|
path = pending.pop()
|
|
14
17
|
if path in seen or not path.is_file() or (root not in path.parents and path != root): continue
|
|
15
18
|
seen.add(path)
|
|
16
|
-
|
|
19
|
+
text = mask_bibliography(path.read_text(encoding="utf-8", errors="replace"))
|
|
20
|
+
for value in INCLUDE.findall(text):
|
|
17
21
|
child = (path.parent / value.strip()).resolve()
|
|
18
22
|
pending.append(child if child.suffix else child.with_suffix(".tex"))
|
|
19
23
|
return sorted(seen)
|
|
20
24
|
def weight(path: Path) -> int:
|
|
21
|
-
text =
|
|
25
|
+
text = mask_bibliography(path.read_text(encoding="utf-8", errors="replace"))
|
|
26
|
+
text = "\n".join(COMMENT.sub("", line) for line in text.splitlines())
|
|
22
27
|
return len(re.findall(r"[A-Za-z]{2,}", COMMAND.sub("", MATH.sub("", text))))
|
|
23
28
|
def main() -> int:
|
|
24
29
|
parser = argparse.ArgumentParser(); parser.add_argument("root", type=Path); parser.add_argument("--entry", type=Path); parser.add_argument("--workers", type=int, default=3); parser.add_argument("--min-weight", type=int, default=1); parser.add_argument("--json", action="store_true"); args = parser.parse_args()
|
|
25
30
|
root = args.root.resolve(); entry = (root / args.entry).resolve() if args.entry else None
|
|
26
|
-
files = reachable(entry, root) if entry else sorted(root.rglob("*.tex")); shared = {"macro.tex", "macros.tex", "commands.tex"}; records = [(p, weight(p)) for p in files if p.is_file() and p != entry and p.name.lower() not in shared]; records = [(p, score) for p, score in records if score >= args.min_weight]
|
|
31
|
+
files = reachable(entry, root) if entry else sorted(root.rglob("*.tex")); shared = {"macro.tex", "macros.tex", "commands.tex"}; records = [(p, weight(p)) for p in files if p.is_file() and p != entry and p.name.lower() not in shared and not is_bibliography_file(p)]; records = [(p, score) for p, score in records if score >= args.min_weight]
|
|
27
32
|
groups = [[] for _ in range(max(1, args.workers))]; totals = [0] * len(groups)
|
|
28
33
|
for path, score in sorted(records, key=lambda item: item[1], reverse=True):
|
|
29
34
|
index = min(range(len(groups)), key=totals.__getitem__); groups[index].append((path, score)); totals[index] += score
|
|
@@ -0,0 +1,84 @@
|
|
|
1
|
+
#!/usr/bin/env python3
|
|
2
|
+
"""Shared helpers for excluding bibliography content from translation work."""
|
|
3
|
+
|
|
4
|
+
from __future__ import annotations
|
|
5
|
+
|
|
6
|
+
import re
|
|
7
|
+
from pathlib import Path
|
|
8
|
+
|
|
9
|
+
|
|
10
|
+
_BIBLIOGRAPHY_ENVIRONMENT = re.compile(
|
|
11
|
+
r"\\begin\s*\{(?P<name>thebibliography|biblist|references)\}"
|
|
12
|
+
r".*?"
|
|
13
|
+
r"\\end\s*\{(?P=name)\}",
|
|
14
|
+
re.IGNORECASE | re.DOTALL,
|
|
15
|
+
)
|
|
16
|
+
_BIBLIOGRAPHY_HEADING = re.compile(
|
|
17
|
+
r"\\(?P<level>chapter|section)\*?\s*"
|
|
18
|
+
r"(?:\[[^]]*\]\s*)?"
|
|
19
|
+
r"\{\s*(?:references|bibliography|literature\s+cited)\s*\}"
|
|
20
|
+
r".*?"
|
|
21
|
+
r"(?="
|
|
22
|
+
r"\\(?:chapter|section)\*?\s*(?:\[[^]]*\]\s*)?\{"
|
|
23
|
+
r"|\\end\s*\{document\}"
|
|
24
|
+
r"|\Z)",
|
|
25
|
+
re.IGNORECASE | re.DOTALL,
|
|
26
|
+
)
|
|
27
|
+
_BIBITEM_BLOCK = re.compile(
|
|
28
|
+
r"^[ \t]*\\bibitem(?:\[[^]]*\])?\s*\{[^}]*\}.*?"
|
|
29
|
+
r"(?="
|
|
30
|
+
r"^[ \t]*\\bibitem(?:\[[^]]*\])?\s*\{"
|
|
31
|
+
r"|^[ \t]*\\end\s*\{thebibliography\}"
|
|
32
|
+
r"|\Z)",
|
|
33
|
+
re.IGNORECASE | re.MULTILINE | re.DOTALL,
|
|
34
|
+
)
|
|
35
|
+
_BIBLIOGRAPHY_CONTROL = re.compile(
|
|
36
|
+
r"^[ \t]*\\(?:"
|
|
37
|
+
r"addbibresource|bibliography|bibliographystyle|printbibliography"
|
|
38
|
+
r")\b"
|
|
39
|
+
r"|^[ \t]*\\renewcommand\s*\{\\(?:bibname|refname)\}",
|
|
40
|
+
re.IGNORECASE,
|
|
41
|
+
)
|
|
42
|
+
_BIBLIOGRAPHY_EXACT_NAMES = {
|
|
43
|
+
"bib",
|
|
44
|
+
"biblio",
|
|
45
|
+
"bibliography",
|
|
46
|
+
"bibliographies",
|
|
47
|
+
"ref",
|
|
48
|
+
"refs",
|
|
49
|
+
"reference",
|
|
50
|
+
"references",
|
|
51
|
+
}
|
|
52
|
+
_BIBLIOGRAPHY_NAME_PARTS = {
|
|
53
|
+
"bib",
|
|
54
|
+
"biblio",
|
|
55
|
+
"bibliography",
|
|
56
|
+
"bibliographies",
|
|
57
|
+
"refs",
|
|
58
|
+
"references",
|
|
59
|
+
}
|
|
60
|
+
|
|
61
|
+
|
|
62
|
+
def _blank(match: re.Match[str]) -> str:
|
|
63
|
+
"""Mask content without changing line numbers."""
|
|
64
|
+
return re.sub(r"[^\n]", " ", match.group(0))
|
|
65
|
+
|
|
66
|
+
|
|
67
|
+
def mask_bibliography(text: str) -> str:
|
|
68
|
+
"""Blank bibliography regions and controls while preserving newlines."""
|
|
69
|
+
text = _BIBLIOGRAPHY_ENVIRONMENT.sub(_blank, text)
|
|
70
|
+
text = _BIBLIOGRAPHY_HEADING.sub(_blank, text)
|
|
71
|
+
text = _BIBITEM_BLOCK.sub(_blank, text)
|
|
72
|
+
return "\n".join(
|
|
73
|
+
"" if _BIBLIOGRAPHY_CONTROL.match(line) else line
|
|
74
|
+
for line in text.split("\n")
|
|
75
|
+
)
|
|
76
|
+
|
|
77
|
+
|
|
78
|
+
def is_bibliography_file(path: Path) -> bool:
|
|
79
|
+
"""Recognize common filenames used only for reference lists."""
|
|
80
|
+
stem = path.stem.lower()
|
|
81
|
+
if stem in _BIBLIOGRAPHY_EXACT_NAMES:
|
|
82
|
+
return True
|
|
83
|
+
parts = {part for part in re.split(r"[-_.]+", stem) if part}
|
|
84
|
+
return bool(parts & _BIBLIOGRAPHY_NAME_PARTS)
|
|
@@ -0,0 +1,594 @@
|
|
|
1
|
+
#!/usr/bin/env python3
|
|
2
|
+
"""Prepare compact TeX translation packets and safely merge their results."""
|
|
3
|
+
|
|
4
|
+
from __future__ import annotations
|
|
5
|
+
|
|
6
|
+
import argparse
|
|
7
|
+
from collections import Counter, defaultdict
|
|
8
|
+
import hashlib
|
|
9
|
+
import json
|
|
10
|
+
import math
|
|
11
|
+
import os
|
|
12
|
+
from pathlib import Path
|
|
13
|
+
import re
|
|
14
|
+
import tempfile
|
|
15
|
+
from typing import Iterable, Optional
|
|
16
|
+
|
|
17
|
+
from tex_translation_utils import is_bibliography_file, mask_bibliography
|
|
18
|
+
|
|
19
|
+
|
|
20
|
+
VERSION = 1
|
|
21
|
+
DEFAULT_TASK_DIR = ".translation-tasks"
|
|
22
|
+
INCLUDE = re.compile(r"\\(?:input|include)\s*\{([^}]+)\}")
|
|
23
|
+
COMMENT = re.compile(r"(?<!\\)%[^\n]*")
|
|
24
|
+
COMMAND = re.compile(r"\\(?:[A-Za-z@]+\*?|.)")
|
|
25
|
+
WORD = re.compile(r"[A-Za-z][A-Za-z'’-]*")
|
|
26
|
+
PLACEHOLDER = re.compile(r"⟪T\d{4}⟫")
|
|
27
|
+
TEXT_COMMAND = re.compile(
|
|
28
|
+
r"\\(?:title|subtitle|chapter|section|subsection|subsubsection|paragraph|"
|
|
29
|
+
r"subparagraph|caption|footnote|thanks|keyword|keywords)\*?\b",
|
|
30
|
+
re.IGNORECASE,
|
|
31
|
+
)
|
|
32
|
+
CONTROL_ONLY = re.compile(
|
|
33
|
+
r"^[ \t]*\\(?:documentclass|usepackage|RequirePackage|input|include|"
|
|
34
|
+
r"addbibresource|bibliography|bibliographystyle|includegraphics|label|"
|
|
35
|
+
r"setlength|newcommand|renewcommand|providecommand|def)\b",
|
|
36
|
+
re.IGNORECASE,
|
|
37
|
+
)
|
|
38
|
+
NON_TEXT_COMMAND = re.compile(
|
|
39
|
+
r"\\(?:documentclass|usepackage|RequirePackage|input|include|"
|
|
40
|
+
r"addbibresource|bibliography|bibliographystyle|includegraphics|label|"
|
|
41
|
+
r"ref|eqref|pageref|autoref|cite\w*|url|path)\*?"
|
|
42
|
+
r"(?:\s*\[[^]]*\])*\s*\{[^{}]*\}",
|
|
43
|
+
re.IGNORECASE,
|
|
44
|
+
)
|
|
45
|
+
MATH_ENVIRONMENT = re.compile(
|
|
46
|
+
r"\\begin\s*\{(?P<name>equation\*?|align\*?|alignat\*?|gather\*?|"
|
|
47
|
+
r"multline\*?|displaymath|math|eqnarray\*?)\}.*?"
|
|
48
|
+
r"\\end\s*\{(?P=name)\}",
|
|
49
|
+
re.IGNORECASE | re.DOTALL,
|
|
50
|
+
)
|
|
51
|
+
VERBATIM_ENVIRONMENT = re.compile(
|
|
52
|
+
r"\\begin\s*\{(?P<name>verbatim\*?|lstlisting|minted)\}.*?"
|
|
53
|
+
r"\\end\s*\{(?P=name)\}",
|
|
54
|
+
re.IGNORECASE | re.DOTALL,
|
|
55
|
+
)
|
|
56
|
+
DISPLAY_MATH = re.compile(r"\$\$.*?\$\$|\\\[.*?\\\]", re.DOTALL)
|
|
57
|
+
INLINE_MATH = re.compile(r"(?<!\\)\$(?!\$).*?(?<!\\)\$|\\\(.*?\\\)", re.DOTALL)
|
|
58
|
+
VERB = re.compile(r"\\verb(?P<delimiter>[^A-Za-z0-9\s]).*?(?P=delimiter)")
|
|
59
|
+
REFERENCE_COMMAND = re.compile(
|
|
60
|
+
r"\\(?:cite\w*|ref|eqref|pageref|autoref|label|url|path|"
|
|
61
|
+
r"includegraphics|input|include)\*?(?:\s*\[[^]]*\])*\s*\{[^{}]*\}",
|
|
62
|
+
re.IGNORECASE,
|
|
63
|
+
)
|
|
64
|
+
HREF_TARGET = re.compile(r"\\href\s*\{[^{}]*\}", re.IGNORECASE)
|
|
65
|
+
PACKET_BLOCK = re.compile(
|
|
66
|
+
r"^@@@ SEGMENT (?P<id>\S+)\n"
|
|
67
|
+
r"^@@@ SOURCE\n(?P<source>.*?)"
|
|
68
|
+
r"^@@@ TRANSLATION\n(?P<translation>.*?)"
|
|
69
|
+
r"^@@@ END(?:\n|\Z)",
|
|
70
|
+
re.MULTILINE | re.DOTALL,
|
|
71
|
+
)
|
|
72
|
+
|
|
73
|
+
PACKET_HEADER = """# 只在每个 TRANSLATION 区块填写简体中文译文;不要改 SOURCE 或标记行。
|
|
74
|
+
# 严格保留 LaTeX 结构及每个 ⟪T0000⟫ 形式的占位符;保留人名、模型名、数据集名、缩写、数值和引用键。
|
|
75
|
+
"""
|
|
76
|
+
|
|
77
|
+
|
|
78
|
+
def _inside(path: Path, root: Path) -> bool:
|
|
79
|
+
try:
|
|
80
|
+
path.relative_to(root)
|
|
81
|
+
return True
|
|
82
|
+
except ValueError:
|
|
83
|
+
return False
|
|
84
|
+
|
|
85
|
+
|
|
86
|
+
def reachable(entry: Path, root: Path) -> list[Path]:
|
|
87
|
+
"""Return TeX files reachable through input/include, including the entry."""
|
|
88
|
+
pending = [entry.resolve()]
|
|
89
|
+
seen: set[Path] = set()
|
|
90
|
+
while pending:
|
|
91
|
+
path = pending.pop()
|
|
92
|
+
if path in seen or not path.is_file() or not _inside(path, root):
|
|
93
|
+
continue
|
|
94
|
+
seen.add(path)
|
|
95
|
+
text = mask_bibliography(path.read_text(encoding="utf-8", errors="replace"))
|
|
96
|
+
for value in INCLUDE.findall(text):
|
|
97
|
+
child = (path.parent / value.strip()).resolve()
|
|
98
|
+
pending.append(child if child.suffix else child.with_suffix(".tex"))
|
|
99
|
+
return sorted(seen)
|
|
100
|
+
|
|
101
|
+
|
|
102
|
+
def _without_protected_text(text: str) -> str:
|
|
103
|
+
for pattern in (
|
|
104
|
+
VERBATIM_ENVIRONMENT,
|
|
105
|
+
MATH_ENVIRONMENT,
|
|
106
|
+
COMMENT,
|
|
107
|
+
DISPLAY_MATH,
|
|
108
|
+
INLINE_MATH,
|
|
109
|
+
VERB,
|
|
110
|
+
REFERENCE_COMMAND,
|
|
111
|
+
HREF_TARGET,
|
|
112
|
+
):
|
|
113
|
+
text = pattern.sub(" ", text)
|
|
114
|
+
return text
|
|
115
|
+
|
|
116
|
+
|
|
117
|
+
def english_weight(text: str) -> int:
|
|
118
|
+
"""Estimate visible English prose words, excluding common TeX controls."""
|
|
119
|
+
text = NON_TEXT_COMMAND.sub(" ", _without_protected_text(text))
|
|
120
|
+
text = COMMAND.sub(" ", text)
|
|
121
|
+
return len([word for word in WORD.findall(text) if len(word) > 1])
|
|
122
|
+
|
|
123
|
+
|
|
124
|
+
def _is_translatable(text: str) -> bool:
|
|
125
|
+
score = english_weight(text)
|
|
126
|
+
return score >= 3 or (score >= 1 and bool(TEXT_COMMAND.search(text) or "&" in text))
|
|
127
|
+
|
|
128
|
+
|
|
129
|
+
def _protect(source: str) -> tuple[str, list[dict[str, str]]]:
|
|
130
|
+
"""Replace expensive immutable spans with checked, reversible placeholders."""
|
|
131
|
+
protected: list[dict[str, str]] = []
|
|
132
|
+
masked = source
|
|
133
|
+
|
|
134
|
+
def replace(match: re.Match[str]) -> str:
|
|
135
|
+
placeholder = f"⟪T{len(protected):04d}⟫"
|
|
136
|
+
protected.append({"placeholder": placeholder, "value": match.group(0)})
|
|
137
|
+
return placeholder
|
|
138
|
+
|
|
139
|
+
for pattern in (
|
|
140
|
+
VERBATIM_ENVIRONMENT,
|
|
141
|
+
MATH_ENVIRONMENT,
|
|
142
|
+
COMMENT,
|
|
143
|
+
DISPLAY_MATH,
|
|
144
|
+
INLINE_MATH,
|
|
145
|
+
VERB,
|
|
146
|
+
REFERENCE_COMMAND,
|
|
147
|
+
HREF_TARGET,
|
|
148
|
+
):
|
|
149
|
+
masked = pattern.sub(replace, masked)
|
|
150
|
+
return masked, protected
|
|
151
|
+
|
|
152
|
+
|
|
153
|
+
def _paragraph_ranges(masked_text: str) -> Iterable[tuple[int, int]]:
|
|
154
|
+
"""Yield zero-based, end-exclusive line ranges separated by hard controls."""
|
|
155
|
+
lines = masked_text.splitlines(keepends=True)
|
|
156
|
+
start: Optional[int] = None
|
|
157
|
+
for index, line in enumerate(lines):
|
|
158
|
+
separator = not line.strip() or bool(CONTROL_ONLY.match(line))
|
|
159
|
+
if separator:
|
|
160
|
+
if start is not None:
|
|
161
|
+
yield start, index
|
|
162
|
+
start = None
|
|
163
|
+
elif start is None:
|
|
164
|
+
start = index
|
|
165
|
+
if start is not None:
|
|
166
|
+
yield start, len(lines)
|
|
167
|
+
|
|
168
|
+
|
|
169
|
+
def _chunk_range(
|
|
170
|
+
lines: list[str], start: int, end: int, chunk_words: int
|
|
171
|
+
) -> Iterable[tuple[int, int]]:
|
|
172
|
+
cursor = start
|
|
173
|
+
score = 0
|
|
174
|
+
for index in range(start, end):
|
|
175
|
+
score += english_weight(lines[index])
|
|
176
|
+
if score >= chunk_words and index + 1 < end:
|
|
177
|
+
yield cursor, index + 1
|
|
178
|
+
cursor = index + 1
|
|
179
|
+
score = 0
|
|
180
|
+
if cursor < end:
|
|
181
|
+
yield cursor, end
|
|
182
|
+
|
|
183
|
+
|
|
184
|
+
def segments_for(path: Path, root: Path, chunk_words: int) -> list[dict[str, object]]:
|
|
185
|
+
original = path.read_text(encoding="utf-8", errors="replace")
|
|
186
|
+
bibliography_masked = mask_bibliography(original)
|
|
187
|
+
original_lines = original.splitlines(keepends=True)
|
|
188
|
+
masked_lines = bibliography_masked.splitlines(keepends=True)
|
|
189
|
+
units: list[tuple[int, int, int]] = []
|
|
190
|
+
|
|
191
|
+
for paragraph_start, paragraph_end in _paragraph_ranges(bibliography_masked):
|
|
192
|
+
for start, end in _chunk_range(
|
|
193
|
+
masked_lines, paragraph_start, paragraph_end, chunk_words
|
|
194
|
+
):
|
|
195
|
+
visible = "".join(masked_lines[start:end])
|
|
196
|
+
if not _is_translatable(visible):
|
|
197
|
+
continue
|
|
198
|
+
units.append((start, end, english_weight(visible)))
|
|
199
|
+
|
|
200
|
+
merged: list[list[int]] = []
|
|
201
|
+
for start, end, score in units:
|
|
202
|
+
if merged:
|
|
203
|
+
previous = merged[-1]
|
|
204
|
+
gap = "".join(original_lines[previous[1] : start])
|
|
205
|
+
if not gap.strip() and previous[2] + score <= chunk_words:
|
|
206
|
+
previous[1] = end
|
|
207
|
+
previous[2] += score
|
|
208
|
+
continue
|
|
209
|
+
merged.append([start, end, score])
|
|
210
|
+
|
|
211
|
+
segments: list[dict[str, object]] = []
|
|
212
|
+
for start, end, score in merged:
|
|
213
|
+
source = "".join(original_lines[start:end])
|
|
214
|
+
masked_source, protected = _protect(source)
|
|
215
|
+
packet_source = (
|
|
216
|
+
masked_source if masked_source.endswith("\n") else masked_source + "\n"
|
|
217
|
+
)
|
|
218
|
+
relative = str(path.relative_to(root))
|
|
219
|
+
identity = f"{relative}:{start + 1}:{end}:".encode() + source.encode()
|
|
220
|
+
segment_id = "s" + hashlib.sha256(identity).hexdigest()[:12]
|
|
221
|
+
segments.append(
|
|
222
|
+
{
|
|
223
|
+
"id": segment_id,
|
|
224
|
+
"path": relative,
|
|
225
|
+
"start_line": start + 1,
|
|
226
|
+
"end_line": end,
|
|
227
|
+
"weight": score,
|
|
228
|
+
"source": source,
|
|
229
|
+
"source_sha256": hashlib.sha256(source.encode()).hexdigest(),
|
|
230
|
+
"packet_source": packet_source,
|
|
231
|
+
"protected": protected,
|
|
232
|
+
}
|
|
233
|
+
)
|
|
234
|
+
return segments
|
|
235
|
+
|
|
236
|
+
|
|
237
|
+
def _allocate(
|
|
238
|
+
segments: list[dict[str, object]], requested_workers: int, min_words: int
|
|
239
|
+
) -> list[list[dict[str, object]]]:
|
|
240
|
+
total = sum(int(segment["weight"]) for segment in segments)
|
|
241
|
+
useful_workers = max(1, math.ceil(total / max(1, min_words)))
|
|
242
|
+
count = min(max(1, requested_workers), useful_workers, max(1, len(segments)))
|
|
243
|
+
groups: list[list[dict[str, object]]] = [[] for _ in range(count)]
|
|
244
|
+
totals = [0] * count
|
|
245
|
+
for segment in sorted(segments, key=lambda item: int(item["weight"]), reverse=True):
|
|
246
|
+
index = min(range(count), key=totals.__getitem__)
|
|
247
|
+
groups[index].append(segment)
|
|
248
|
+
totals[index] += int(segment["weight"])
|
|
249
|
+
for group in groups:
|
|
250
|
+
group.sort(key=lambda item: (str(item["path"]), int(item["start_line"])))
|
|
251
|
+
return groups
|
|
252
|
+
|
|
253
|
+
|
|
254
|
+
def _packet_text(segments: list[dict[str, object]]) -> str:
|
|
255
|
+
parts = [PACKET_HEADER]
|
|
256
|
+
for segment in segments:
|
|
257
|
+
parts.extend(
|
|
258
|
+
[
|
|
259
|
+
f"@@@ SEGMENT {segment['id']}\n",
|
|
260
|
+
"@@@ SOURCE\n",
|
|
261
|
+
str(segment["packet_source"]),
|
|
262
|
+
"@@@ TRANSLATION\n",
|
|
263
|
+
"@@@ END\n",
|
|
264
|
+
]
|
|
265
|
+
)
|
|
266
|
+
return "".join(parts)
|
|
267
|
+
|
|
268
|
+
|
|
269
|
+
def _write_atomic(path: Path, text: str) -> None:
|
|
270
|
+
path.parent.mkdir(parents=True, exist_ok=True)
|
|
271
|
+
handle = tempfile.NamedTemporaryFile(
|
|
272
|
+
"w", encoding="utf-8", dir=path.parent, delete=False
|
|
273
|
+
)
|
|
274
|
+
temporary = Path(handle.name)
|
|
275
|
+
try:
|
|
276
|
+
with handle:
|
|
277
|
+
handle.write(text)
|
|
278
|
+
os.replace(temporary, path)
|
|
279
|
+
finally:
|
|
280
|
+
if temporary.exists():
|
|
281
|
+
temporary.unlink()
|
|
282
|
+
|
|
283
|
+
|
|
284
|
+
def prepare(args: argparse.Namespace) -> int:
|
|
285
|
+
root = args.root.resolve()
|
|
286
|
+
if not root.is_dir():
|
|
287
|
+
raise SystemExit(f"paper root does not exist: {root}")
|
|
288
|
+
task_dir = (root / args.output).resolve()
|
|
289
|
+
if not _inside(task_dir, root):
|
|
290
|
+
raise SystemExit(f"task directory must stay below paper root: {task_dir}")
|
|
291
|
+
manifest_path = task_dir / "manifest.json"
|
|
292
|
+
if manifest_path.exists() and not args.force:
|
|
293
|
+
raise SystemExit(f"task manifest already exists: {manifest_path}; use --force")
|
|
294
|
+
if args.force and task_dir.is_dir():
|
|
295
|
+
for stale_packet in task_dir.glob("worker-*.task"):
|
|
296
|
+
stale_packet.unlink()
|
|
297
|
+
|
|
298
|
+
if args.entry:
|
|
299
|
+
entry = (root / args.entry).resolve()
|
|
300
|
+
if not entry.is_file():
|
|
301
|
+
raise SystemExit(f"entry file does not exist: {entry}")
|
|
302
|
+
files = reachable(entry, root)
|
|
303
|
+
else:
|
|
304
|
+
entry = None
|
|
305
|
+
files = sorted(root.rglob("*.tex"))
|
|
306
|
+
files = [path for path in files if path.is_file() and not is_bibliography_file(path)]
|
|
307
|
+
segments = [
|
|
308
|
+
segment
|
|
309
|
+
for path in files
|
|
310
|
+
for segment in segments_for(path, root, args.chunk_words)
|
|
311
|
+
]
|
|
312
|
+
if not segments:
|
|
313
|
+
raise SystemExit(f"no translatable TeX prose found below {root}")
|
|
314
|
+
|
|
315
|
+
groups = _allocate(segments, args.workers, args.min_words_per_worker)
|
|
316
|
+
packets = []
|
|
317
|
+
for index, group in enumerate(groups, 1):
|
|
318
|
+
name = f"worker-{index:02d}.task"
|
|
319
|
+
_write_atomic(task_dir / name, _packet_text(group))
|
|
320
|
+
packets.append(
|
|
321
|
+
{
|
|
322
|
+
"worker": index,
|
|
323
|
+
"path": name,
|
|
324
|
+
"weight": sum(int(segment["weight"]) for segment in group),
|
|
325
|
+
"segments": [str(segment["id"]) for segment in group],
|
|
326
|
+
}
|
|
327
|
+
)
|
|
328
|
+
|
|
329
|
+
scanned_bytes = sum(path.stat().st_size for path in files)
|
|
330
|
+
selected_bytes = sum(len(str(segment["source"]).encode()) for segment in segments)
|
|
331
|
+
packet_source_bytes = sum(
|
|
332
|
+
len(str(segment["packet_source"]).encode()) for segment in segments
|
|
333
|
+
)
|
|
334
|
+
reduction = (
|
|
335
|
+
round(100 * (1 - packet_source_bytes / scanned_bytes), 1)
|
|
336
|
+
if scanned_bytes
|
|
337
|
+
else 0.0
|
|
338
|
+
)
|
|
339
|
+
manifest = {
|
|
340
|
+
"version": VERSION,
|
|
341
|
+
"root": str(root),
|
|
342
|
+
"entry": str(entry.relative_to(root)) if entry else None,
|
|
343
|
+
"packets": packets,
|
|
344
|
+
"segments": segments,
|
|
345
|
+
"metrics": {
|
|
346
|
+
"files": len(files),
|
|
347
|
+
"segments": len(segments),
|
|
348
|
+
"workers": len(groups),
|
|
349
|
+
"visible_english_words": sum(int(segment["weight"]) for segment in segments),
|
|
350
|
+
"scanned_bytes": scanned_bytes,
|
|
351
|
+
"selected_bytes": selected_bytes,
|
|
352
|
+
"packet_source_bytes": packet_source_bytes,
|
|
353
|
+
"input_byte_reduction_percent": reduction,
|
|
354
|
+
},
|
|
355
|
+
}
|
|
356
|
+
_write_atomic(manifest_path, json.dumps(manifest, ensure_ascii=False, indent=2) + "\n")
|
|
357
|
+
public_packets = [
|
|
358
|
+
{
|
|
359
|
+
"worker": packet["worker"],
|
|
360
|
+
"path": packet["path"],
|
|
361
|
+
"weight": packet["weight"],
|
|
362
|
+
"segments": len(packet["segments"]),
|
|
363
|
+
}
|
|
364
|
+
for packet in packets
|
|
365
|
+
]
|
|
366
|
+
payload = {
|
|
367
|
+
"task_dir": str(task_dir),
|
|
368
|
+
**manifest["metrics"],
|
|
369
|
+
"packets": public_packets,
|
|
370
|
+
}
|
|
371
|
+
if args.json:
|
|
372
|
+
print(json.dumps(payload, ensure_ascii=False, indent=2))
|
|
373
|
+
else:
|
|
374
|
+
print(
|
|
375
|
+
f"prepared {len(segments)} segments for {len(groups)} worker(s); "
|
|
376
|
+
f"estimated input reduction {reduction:.1f}%"
|
|
377
|
+
)
|
|
378
|
+
for packet in packets:
|
|
379
|
+
print(f"{packet['path']}: words={packet['weight']} segments={len(packet['segments'])}")
|
|
380
|
+
return 0
|
|
381
|
+
|
|
382
|
+
|
|
383
|
+
def parse_packet(path: Path) -> dict[str, dict[str, str]]:
|
|
384
|
+
text = path.read_text(encoding="utf-8")
|
|
385
|
+
result: dict[str, dict[str, str]] = {}
|
|
386
|
+
for match in PACKET_BLOCK.finditer(text):
|
|
387
|
+
segment_id = match.group("id")
|
|
388
|
+
if segment_id in result:
|
|
389
|
+
raise ValueError(f"duplicate segment {segment_id} in {path}")
|
|
390
|
+
result[segment_id] = {
|
|
391
|
+
"source": match.group("source"),
|
|
392
|
+
"translation": match.group("translation"),
|
|
393
|
+
}
|
|
394
|
+
return result
|
|
395
|
+
|
|
396
|
+
|
|
397
|
+
def _load_manifest(root: Path, output: Path) -> tuple[Path, dict[str, object]]:
|
|
398
|
+
task_dir = (root / output).resolve()
|
|
399
|
+
if not _inside(task_dir, root):
|
|
400
|
+
raise SystemExit(f"task directory must stay below paper root: {task_dir}")
|
|
401
|
+
manifest_path = task_dir / "manifest.json"
|
|
402
|
+
if not manifest_path.is_file():
|
|
403
|
+
raise SystemExit(f"task manifest does not exist: {manifest_path}")
|
|
404
|
+
manifest = json.loads(manifest_path.read_text(encoding="utf-8"))
|
|
405
|
+
if manifest.get("version") != VERSION:
|
|
406
|
+
raise SystemExit(f"unsupported task manifest version: {manifest.get('version')}")
|
|
407
|
+
return task_dir, manifest
|
|
408
|
+
|
|
409
|
+
|
|
410
|
+
def _read_results(
|
|
411
|
+
task_dir: Path, manifest: dict[str, object]
|
|
412
|
+
) -> tuple[dict[str, dict[str, str]], list[str]]:
|
|
413
|
+
results: dict[str, dict[str, str]] = {}
|
|
414
|
+
errors: list[str] = []
|
|
415
|
+
for packet in manifest["packets"]: # type: ignore[index]
|
|
416
|
+
packet_path = task_dir / packet["path"]
|
|
417
|
+
try:
|
|
418
|
+
parsed = parse_packet(packet_path)
|
|
419
|
+
except (OSError, ValueError) as error:
|
|
420
|
+
errors.append(str(error))
|
|
421
|
+
continue
|
|
422
|
+
for segment_id, value in parsed.items():
|
|
423
|
+
if segment_id in results:
|
|
424
|
+
errors.append(f"duplicate result for {segment_id}")
|
|
425
|
+
results[segment_id] = value
|
|
426
|
+
return results, errors
|
|
427
|
+
|
|
428
|
+
|
|
429
|
+
def status(args: argparse.Namespace) -> int:
|
|
430
|
+
root = args.root.resolve()
|
|
431
|
+
task_dir, manifest = _load_manifest(root, args.output)
|
|
432
|
+
results, errors = _read_results(task_dir, manifest)
|
|
433
|
+
expected = {str(segment["id"]) for segment in manifest["segments"]} # type: ignore[index]
|
|
434
|
+
completed = {
|
|
435
|
+
segment_id
|
|
436
|
+
for segment_id, value in results.items()
|
|
437
|
+
if value["translation"].strip()
|
|
438
|
+
}
|
|
439
|
+
missing = sorted(expected - completed)
|
|
440
|
+
payload = {
|
|
441
|
+
"segments": len(expected),
|
|
442
|
+
"completed": len(expected) - len(missing),
|
|
443
|
+
"missing": missing,
|
|
444
|
+
"errors": errors,
|
|
445
|
+
}
|
|
446
|
+
if args.json:
|
|
447
|
+
print(json.dumps(payload, ensure_ascii=False, indent=2))
|
|
448
|
+
else:
|
|
449
|
+
print(f"completed={payload['completed']}/{payload['segments']}")
|
|
450
|
+
for error in errors:
|
|
451
|
+
print(f"error: {error}")
|
|
452
|
+
for segment_id in missing:
|
|
453
|
+
print(f"missing: {segment_id}")
|
|
454
|
+
return 0 if not missing and not errors else 1
|
|
455
|
+
|
|
456
|
+
|
|
457
|
+
def _restore(translation: str, protected: list[dict[str, str]]) -> str:
|
|
458
|
+
for item in protected:
|
|
459
|
+
translation = translation.replace(item["placeholder"], item["value"])
|
|
460
|
+
return translation
|
|
461
|
+
|
|
462
|
+
|
|
463
|
+
def _structure_signature(text: str) -> tuple[Counter[str], int, int]:
|
|
464
|
+
return Counter(COMMAND.findall(text)), text.count("{"), text.count("}")
|
|
465
|
+
|
|
466
|
+
|
|
467
|
+
def apply(args: argparse.Namespace) -> int:
|
|
468
|
+
root = args.root.resolve()
|
|
469
|
+
task_dir, manifest = _load_manifest(root, args.output)
|
|
470
|
+
results, errors = _read_results(task_dir, manifest)
|
|
471
|
+
replacements: defaultdict[Path, list[tuple[int, int, str]]] = defaultdict(list)
|
|
472
|
+
expected_sources = {
|
|
473
|
+
(str(segment["path"]), int(segment["start_line"]), int(segment["end_line"])): (
|
|
474
|
+
str(segment["source"]),
|
|
475
|
+
str(segment["source_sha256"]),
|
|
476
|
+
)
|
|
477
|
+
for segment in manifest["segments"] # type: ignore[index]
|
|
478
|
+
}
|
|
479
|
+
|
|
480
|
+
for segment in manifest["segments"]: # type: ignore[index]
|
|
481
|
+
segment_id = str(segment["id"])
|
|
482
|
+
result = results.get(segment_id)
|
|
483
|
+
if result is None:
|
|
484
|
+
errors.append(f"missing segment {segment_id}")
|
|
485
|
+
continue
|
|
486
|
+
if result["source"] != segment["packet_source"]:
|
|
487
|
+
errors.append(f"SOURCE was modified for {segment_id}")
|
|
488
|
+
continue
|
|
489
|
+
translation = result["translation"].lstrip("\n")
|
|
490
|
+
if not translation.strip():
|
|
491
|
+
errors.append(f"empty translation for {segment_id}")
|
|
492
|
+
continue
|
|
493
|
+
expected_placeholders = Counter(
|
|
494
|
+
item["placeholder"] for item in segment["protected"]
|
|
495
|
+
)
|
|
496
|
+
actual_placeholders = Counter(PLACEHOLDER.findall(translation))
|
|
497
|
+
if actual_placeholders != expected_placeholders:
|
|
498
|
+
errors.append(f"placeholder mismatch for {segment_id}")
|
|
499
|
+
continue
|
|
500
|
+
if _structure_signature(translation) != _structure_signature(
|
|
501
|
+
str(segment["packet_source"])
|
|
502
|
+
):
|
|
503
|
+
errors.append(f"LaTeX structure mismatch for {segment_id}")
|
|
504
|
+
continue
|
|
505
|
+
|
|
506
|
+
source = str(segment["source"])
|
|
507
|
+
if source.endswith("\n"):
|
|
508
|
+
if not translation.endswith("\n"):
|
|
509
|
+
translation += "\n"
|
|
510
|
+
else:
|
|
511
|
+
translation = translation.rstrip("\n")
|
|
512
|
+
restored = _restore(translation, segment["protected"])
|
|
513
|
+
path = root / str(segment["path"])
|
|
514
|
+
replacements[path].append(
|
|
515
|
+
(int(segment["start_line"]), int(segment["end_line"]), restored)
|
|
516
|
+
)
|
|
517
|
+
|
|
518
|
+
for path, items in replacements.items():
|
|
519
|
+
text = path.read_text(encoding="utf-8", errors="replace")
|
|
520
|
+
lines = text.splitlines(keepends=True)
|
|
521
|
+
for start, end, _translation in items:
|
|
522
|
+
current = "".join(lines[start - 1 : end])
|
|
523
|
+
expected, expected_hash = expected_sources[
|
|
524
|
+
(str(path.relative_to(root)), start, end)
|
|
525
|
+
]
|
|
526
|
+
if (
|
|
527
|
+
current != expected
|
|
528
|
+
or hashlib.sha256(current.encode()).hexdigest() != expected_hash
|
|
529
|
+
):
|
|
530
|
+
errors.append(f"source changed since prepare: {path.relative_to(root)}:{start}-{end}")
|
|
531
|
+
|
|
532
|
+
if errors:
|
|
533
|
+
for error in errors:
|
|
534
|
+
print(f"error: {error}")
|
|
535
|
+
return 1
|
|
536
|
+
if args.check:
|
|
537
|
+
print(f"validated {sum(len(items) for items in replacements.values())} translations")
|
|
538
|
+
return 0
|
|
539
|
+
|
|
540
|
+
for path, items in replacements.items():
|
|
541
|
+
lines = path.read_text(encoding="utf-8", errors="replace").splitlines(keepends=True)
|
|
542
|
+
for start, end, translation in sorted(items, reverse=True):
|
|
543
|
+
lines[start - 1 : end] = [translation]
|
|
544
|
+
_write_atomic(path, "".join(lines))
|
|
545
|
+
print(
|
|
546
|
+
f"applied {sum(len(items) for items in replacements.values())} translations "
|
|
547
|
+
f"to {len(replacements)} file(s)"
|
|
548
|
+
)
|
|
549
|
+
return 0
|
|
550
|
+
|
|
551
|
+
|
|
552
|
+
def build_parser() -> argparse.ArgumentParser:
|
|
553
|
+
parser = argparse.ArgumentParser(description=__doc__)
|
|
554
|
+
subparsers = parser.add_subparsers(dest="command", required=True)
|
|
555
|
+
|
|
556
|
+
prepare_parser = subparsers.add_parser("prepare", help="create compact worker packets")
|
|
557
|
+
prepare_parser.add_argument("root", type=Path)
|
|
558
|
+
prepare_parser.add_argument("--entry", type=Path)
|
|
559
|
+
prepare_parser.add_argument("--workers", type=int, default=3)
|
|
560
|
+
prepare_parser.add_argument("--chunk-words", type=int, default=900)
|
|
561
|
+
prepare_parser.add_argument("--min-words-per-worker", type=int, default=1200)
|
|
562
|
+
prepare_parser.add_argument("--output", type=Path, default=Path(DEFAULT_TASK_DIR))
|
|
563
|
+
prepare_parser.add_argument("--force", action="store_true")
|
|
564
|
+
prepare_parser.add_argument("--json", action="store_true")
|
|
565
|
+
prepare_parser.set_defaults(handler=prepare)
|
|
566
|
+
|
|
567
|
+
status_parser = subparsers.add_parser("status", help="report packet completion")
|
|
568
|
+
status_parser.add_argument("root", type=Path)
|
|
569
|
+
status_parser.add_argument("--output", type=Path, default=Path(DEFAULT_TASK_DIR))
|
|
570
|
+
status_parser.add_argument("--json", action="store_true")
|
|
571
|
+
status_parser.set_defaults(handler=status)
|
|
572
|
+
|
|
573
|
+
apply_parser = subparsers.add_parser("apply", help="validate and merge packet translations")
|
|
574
|
+
apply_parser.add_argument("root", type=Path)
|
|
575
|
+
apply_parser.add_argument("--output", type=Path, default=Path(DEFAULT_TASK_DIR))
|
|
576
|
+
apply_parser.add_argument("--check", action="store_true")
|
|
577
|
+
apply_parser.set_defaults(handler=apply)
|
|
578
|
+
return parser
|
|
579
|
+
|
|
580
|
+
|
|
581
|
+
def main() -> int:
|
|
582
|
+
parser = build_parser()
|
|
583
|
+
args = parser.parse_args()
|
|
584
|
+
if getattr(args, "workers", 1) < 1:
|
|
585
|
+
parser.error("--workers must be positive")
|
|
586
|
+
if getattr(args, "chunk_words", 1) < 1:
|
|
587
|
+
parser.error("--chunk-words must be positive")
|
|
588
|
+
if getattr(args, "min_words_per_worker", 1) < 1:
|
|
589
|
+
parser.error("--min-words-per-worker must be positive")
|
|
590
|
+
return int(args.handler(args))
|
|
591
|
+
|
|
592
|
+
|
|
593
|
+
if __name__ == "__main__":
|
|
594
|
+
raise SystemExit(main())
|