gongwen-skill 2.1.0 → 2.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +31 -0
- package/README.md +6 -6
- package/dsh/index.js +2 -2
- package/engine/core/document/markdown_converter.py +56 -2
- package/engine/core/document/modifier.py +19 -5
- package/engine/core/document/parser.py +29 -5
- package/engine/core/rules/checker.py +229 -13
- package/engine/inject.py +26 -8
- package/gongwen/__init__.py +1 -1
- package/gongwen/_legacy.py +31 -12
- package/gongwen/cli/doctor_cmds.py +157 -4
- package/gongwen/md2docx_render.py +509 -0
- package/package.json +1 -1
- package/prompts/usage-prompts.md +1 -1
- package/pyproject.toml +1 -1
- package/rules/official/_common.yaml +3 -3
package/CHANGELOG.md
CHANGED
|
@@ -4,6 +4,37 @@
|
|
|
4
4
|
Licensed under the MIT License. See the LICENSE file for details.
|
|
5
5
|
-->
|
|
6
6
|
|
|
7
|
+
## v2.3.0 (2026-08-31)
|
|
8
|
+
|
|
9
|
+
### Added
|
|
10
|
+
- **doctor 自检新增「DSH 技能 frontmatter」检查项**(21 项 → 22 项):按 DSH 最新技能规范校验 SKILL.md frontmatter——`name` 必填且为 kebab-case、与 `.dsh/skills/<name>` 目录/文件名一致;`description` 必填非空且足够自描述(模型在会话目录只能看到它);`whenToUse` 建议存在;`user-invocable`/`disable-model-invocation` 若存在必须为布尔值;frontmatter 必须 `---` 包裹且 YAML 可解析
|
|
11
|
+
|
|
12
|
+
### Changed
|
|
13
|
+
- **技能名动态化**:新增 `_get_skill_name()` 从 SKILL.md frontmatter 读取技能名,`_check_skill_sync`/`cmd_repair` 不再硬编码 `gongwen-skill`,技能改名后检查与修复仍准确
|
|
14
|
+
- **纯 pip 环境不再误报**:frontmatter 目录一致性校验仅在 `.dsh/skills` 存在时执行(非 DSH 环境跳过)
|
|
15
|
+
|
|
16
|
+
### Fixed
|
|
17
|
+
- **SKILL.md 编码读取修复**:读取编码 `utf-8` → `utf-8-sig`(自动剥离 BOM,兼容带/无 BOM);修复 Windows 下默认 locale 编码(cp936/GBK)读 UTF-8 无 BOM 的 SKILL.md 抛 `UnicodeDecodeError` 的隐患
|
|
18
|
+
|
|
19
|
+
## v2.2.0 (2026-08-28)
|
|
20
|
+
|
|
21
|
+
### Added
|
|
22
|
+
- **内容要素检查规则真正生效**:修复 37 条 `content.*` 规则(通知事项/会议要素/请示理由等)此前"定义但从不执行"的问题,新增 `_check_content_field` 按字段名映射要素关键词做宽松检查,消除每次 check 的"未支持检查字段"告警
|
|
23
|
+
- **版头检查规则生效**:修复 `header.*` 规则(CHK-CM002 令号检查、CHK-R003 主送机关检查)此前无分发分支的问题,新增 `_check_header_field`
|
|
24
|
+
- **成文日期"右空四字"**:md2docx 渲染日期段新增右缩进 4em(GB/T 9704 规范),消除 CHK-C017 长期存在的日期右空四字检查项
|
|
25
|
+
|
|
26
|
+
### Changed
|
|
27
|
+
- **signature 段落选择器修正**:`_select_paragraphs` 的 `signature` target 不再包含 `date` 段落——修复 FIX-C013(署名居中)与 FIX-C013b(18pt)误把成文日期强制居中/放大字号的问题,日期段独立按"右对齐 + 右空四字 + 16pt"处理
|
|
28
|
+
- **页码重复注入幂等化**:`_inject_even_page_footer_direct` 重写为幂等版,多次注入仅保留一个 even footerReference 与完整关系,杜绝 Word 打开报"文档损坏"(BUG-1)
|
|
29
|
+
- **加粗+链接/代码组合不再整段加粗**:markdown_converter 单片段按自身 bold 标志处理,重建片段继承原 run 字体字号(BUG-2)
|
|
30
|
+
- **标题启发式误判修复**:黑体/楷体/仿宋加粗正文句不再误判为标题(新增 `_SENTENCE_END` 句末标点保护,Level 1/2/3 三级判定)
|
|
31
|
+
- **CHK-C030 整段加粗误报修复**:忽略纯标点 run、单句段落不报,多句整段加粗仍正确检出
|
|
32
|
+
|
|
33
|
+
### Fixed
|
|
34
|
+
- `is_title` vs `is_heading` 属性引用错误:`_check_ending`/`_check_content_field` 此前用不存在的 `is_title` 属性,导致标题段被误纳入正文统计
|
|
35
|
+
- 一是/二是 领句段改用仿宋_GB2312(原误用楷体触发标题误判/正文字体误报)
|
|
36
|
+
- 令号正则 `\d` 被 JS 转义破坏(`d`),导致含令号命令误报 CHK-CM002
|
|
37
|
+
|
|
7
38
|
## v2.1.0 (2026-08-20)
|
|
8
39
|
|
|
9
40
|
### Added
|
package/README.md
CHANGED
|
@@ -46,7 +46,7 @@ Licensed under the MIT License. See the LICENSE file for details.
|
|
|
46
46
|
| 🧩 完整审校 | `full-review` | 修订+批注联合命令(句子级差异修订 + 分类批注) |
|
|
47
47
|
| 🎨 样式学习 | `style-learn` / `style-list` | 上传标准文档学习 Run/段落/页面三级样式(字体/字号/字间距/行距/缩进/页边距),生成命名模板持久化,后续用 `optimize -t 模板名` 套用 |
|
|
48
48
|
| 🔄 版本自检 | `check-update` | 多渠道版本自检(GitHub/GitCode/AtomGit 三仓库比对取最新) |
|
|
49
|
-
| 🩺 自我诊断 | `doctor` / `repair` | 全面诊断
|
|
49
|
+
| 🩺 自我诊断 | `doctor` / `repair` | 全面诊断 22 项(Python/依赖/版本一致性/字体/DSH 文件/DSH 技能 frontmatter/代码风格),自动修复常见问题 |
|
|
50
50
|
| 🕵️ 文档审计 | `audit` | 检查删除线/加粗/AI 声明等痕迹 |
|
|
51
51
|
| 🤝 会话交接 | `handoff` | 跨会话上下文传递(`--list` / `--latest` / Agent 长任务收尾必写) |
|
|
52
52
|
| ⚙️ 规则管理 | `rule-export/import/list` | YAML 规则三层定制(官方/单位/用户) |
|
|
@@ -337,7 +337,7 @@ DSH 采用 **Cordis 模块化微内核架构**:技能体系基于本地文件
|
|
|
337
337
|
git clone https://github.com/linhut/gongwen-skill.git
|
|
338
338
|
cd gongwen-skill
|
|
339
339
|
pip install -r requirements.txt # 或 pip install gongwen-skill(已上 PyPI)
|
|
340
|
-
python -m gongwen --version # 检验:gongwen-skill v2.
|
|
340
|
+
python -m gongwen --version # 检验:gongwen-skill v2.3.0
|
|
341
341
|
```
|
|
342
342
|
|
|
343
343
|
### 方式一:作为 DSH Skill 注册(基于本地文件系统)
|
|
@@ -393,7 +393,7 @@ pnpm add -w gongwen-skill
|
|
|
393
393
|
"dependencies": {
|
|
394
394
|
"@deepseek-ai/dsh-base": "...",
|
|
395
395
|
"@deepseek-ai/dsh-web-app": "...",
|
|
396
|
-
"gongwen-skill": "^2.
|
|
396
|
+
"gongwen-skill": "^2.3.0"
|
|
397
397
|
},
|
|
398
398
|
"dsh": {
|
|
399
399
|
"profile": {
|
|
@@ -448,9 +448,9 @@ dsh --profile web
|
|
|
448
448
|
| CLI 独立可执行(`python -m gongwen <命令>`) | ✅ |
|
|
449
449
|
| PyPI 上架(`pip install gongwen-skill`) | ✅ |
|
|
450
450
|
| 零外部运行时依赖(仅 python-docx/pydantic/pyyaml) | ✅ |
|
|
451
|
-
| DSH 配置化排版参数(页边距/行距/字体/默认模板版本) | ✅ v2.
|
|
451
|
+
| DSH 配置化排版参数(页边距/行距/字体/默认模板版本) | ✅ v2.3.0+ |
|
|
452
452
|
|
|
453
|
-
### DSH 插件配置化(v2.
|
|
453
|
+
### DSH 插件配置化(v2.3.0+)
|
|
454
454
|
|
|
455
455
|
DSH 插件支持通过配置文件管理排版参数,Agent 调用时自动注入,纯 CLI 用户不受影响。
|
|
456
456
|
|
|
@@ -574,7 +574,7 @@ pip install -r requirements.txt
|
|
|
574
574
|
用户:帮我优化这份会议通知的第二章节措辞
|
|
575
575
|
|
|
576
576
|
Agent:📋 合规自检报告
|
|
577
|
-
Skill 版本: v2.
|
|
577
|
+
Skill 版本: v2.3.0(多渠道自检已确认最新)
|
|
578
578
|
路径判定: B(内容优化)
|
|
579
579
|
依据: 用户指定了已有文档,且要求"优化措辞"
|
|
580
580
|
命令调用: 1. python -m gongwen optimize-content 会议通知.docx --changes changes.json --apply --paragraphs "5-8"
|
package/dsh/index.js
CHANGED
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
// 公文全流程处理工具 - DSH plugin bridge (gongwen-skill, v2.
|
|
1
|
+
// 公文全流程处理工具 - DSH plugin bridge (gongwen-skill, v2.3.0+)
|
|
2
2
|
// (c) 2026 Jose AI (https://www.linhut.cn)
|
|
3
3
|
// https://github.com/linhut/gongwen-skill
|
|
4
4
|
// Licensed under the MIT License. See the LICENSE file for details.
|
|
@@ -30,7 +30,7 @@ const CONFIG_FILE = join(APP_DATA_DIR, "dsh-config.json");
|
|
|
30
30
|
const DEFAULTS_FILE = join(resolve(__dirname, ".."), "etc", "dsh-config-defaults.json");
|
|
31
31
|
|
|
32
32
|
// AI 工作指引
|
|
33
|
-
const GONGWEN_GUIDANCE = `本机已安装公文全流程处理工具插件(gongwen-skill)。能力:.docx 公文全流程——列出公文类型(list-types)、解析文档(parse)、格式检查(check)、自动修复(optimize)、内容修订对比版(optimize-content)、模板生成(template)、样式学习(style-learn/style-list,从标准文档学习排版样式)、全面诊断(doctor,
|
|
33
|
+
const GONGWEN_GUIDANCE = `本机已安装公文全流程处理工具插件(gongwen-skill)。能力:.docx 公文全流程——列出公文类型(list-types)、解析文档(parse)、格式检查(check)、自动修复(optimize)、内容修订对比版(optimize-content)、模板生成(template)、样式学习(style-learn/style-list,从标准文档学习排版样式)、全面诊断(doctor,22 项自检)、自动修复(repair)、Markdown 转公文(md2docx)、JSON 模型生成(generate)、版头/版记/页码注入(header/footer/pagenum)、首句加粗(bold-first)、一键格式修复(fix-common)、桌签生成(table-signs)、审稿流转单(review)、完整审校(full-review)、文档审计(audit)、规则管理(rule-export/import/list)、版本自检(check-update)、会话交接(handoff)、字体管理(font)。覆盖通知/请示/报告/函/会议纪要等 24 类公文。完全自包含,克隆即用,无需数据库或后端服务。用户提到「公文 / 红头文件 / 版式 / 排版 / 格式检查 / 公文模板 / 样式学习 / 自定义模板 / 党政机关公文」时即指本插件。DSH 插件支持配置化排版参数(页边距/行距/字体等),配置文件位于 ~/.gongwen-skill/dsh-config.json,可通过 config 命令或 DSH 系统设置→插件配置管理。`;
|
|
34
34
|
|
|
35
35
|
// Web API 路由前缀
|
|
36
36
|
const API_PREFIX = "/plugins/gongwen-skill/api";
|
|
@@ -51,6 +51,32 @@ _MD_CODE_BLOCK_RE = re.compile(r'^`{3,}')
|
|
|
51
51
|
# 行内代码 `code`
|
|
52
52
|
_MD_INLINE_CODE_RE = re.compile(r'`([^`]+)`')
|
|
53
53
|
|
|
54
|
+
# 加粗标记(**text** 或 __text__)
|
|
55
|
+
_MD_BOLD_SEG_RE = re.compile(r'(\*\*[^*]+\*\*|__[^_]+__)')
|
|
56
|
+
|
|
57
|
+
|
|
58
|
+
def _reconstruct_bold_segments(raw: str, cleaned: str):
|
|
59
|
+
"""把含 **..** 的原始文本映射为 [(文本, 是否加粗), ...] 片段。
|
|
60
|
+
|
|
61
|
+
片段拼接结果与 cleaned 一致时才返回片段列表;否则(含链接/代码等被清理、
|
|
62
|
+
或文本被折叠)回退为整段(避免丢失文本)。
|
|
63
|
+
"""
|
|
64
|
+
if not raw or ('**' not in raw and '__' not in raw):
|
|
65
|
+
return [(cleaned or "", False)]
|
|
66
|
+
segments = []
|
|
67
|
+
for part in _MD_BOLD_SEG_RE.split(raw):
|
|
68
|
+
if not part:
|
|
69
|
+
continue
|
|
70
|
+
if part.startswith('**') and part.endswith('**') and len(part) > 4:
|
|
71
|
+
segments.append((part[2:-2], True))
|
|
72
|
+
elif part.startswith('__') and part.endswith('__') and len(part) > 4:
|
|
73
|
+
segments.append((part[2:-2], True))
|
|
74
|
+
else:
|
|
75
|
+
segments.append((part, False))
|
|
76
|
+
if ''.join(s[0] for s in segments).strip() == (cleaned or "").strip():
|
|
77
|
+
return segments
|
|
78
|
+
return [(cleaned or "", False)]
|
|
79
|
+
|
|
54
80
|
|
|
55
81
|
def convert_markdown(model: DocumentModel) -> int:
|
|
56
82
|
"""
|
|
@@ -201,6 +227,7 @@ def convert_markdown(model: DocumentModel) -> int:
|
|
|
201
227
|
# --- 识别加粗标记 **text** ---
|
|
202
228
|
|
|
203
229
|
has_bold = False
|
|
230
|
+
raw_before_bold = text # 保存剥离 ** 前的原始文本,用于加粗段重构
|
|
204
231
|
if _MD_BOLD_RE.search(text) or _MD_BOLD_UNDER_RE.search(text):
|
|
205
232
|
has_bold = True
|
|
206
233
|
text = _MD_BOLD_RE.sub(r'\1', text)
|
|
@@ -257,8 +284,35 @@ def convert_markdown(model: DocumentModel) -> int:
|
|
|
257
284
|
r.text = ""
|
|
258
285
|
|
|
259
286
|
if has_bold and not para.is_heading:
|
|
260
|
-
|
|
261
|
-
|
|
287
|
+
# 仅加粗 **..** 标记的片段,避免整段误加粗(CHK-C030 修复)
|
|
288
|
+
_segments = _reconstruct_bold_segments(raw_before_bold, para.text)
|
|
289
|
+
if len(_segments) > 1:
|
|
290
|
+
# 重建片段时继承原 run 的字体/字号,避免丢失字体(I20 修复)
|
|
291
|
+
_base_format = None
|
|
292
|
+
for _r in para.runs:
|
|
293
|
+
if _r.text and (_r.format.font_name or _r.format.font_size_pt):
|
|
294
|
+
_base_format = _r.format
|
|
295
|
+
break
|
|
296
|
+
_new_runs = []
|
|
297
|
+
for _t, _b in _segments:
|
|
298
|
+
if not _t:
|
|
299
|
+
continue
|
|
300
|
+
_nr = Run(index=len(_new_runs), text=_t,
|
|
301
|
+
format=RunFormat(
|
|
302
|
+
font_name=_base_format.font_name if _base_format else None,
|
|
303
|
+
font_size_pt=_base_format.font_size_pt if _base_format else 16.0,
|
|
304
|
+
))
|
|
305
|
+
if _b:
|
|
306
|
+
_nr.format.bold = True
|
|
307
|
+
_new_runs.append(_nr)
|
|
308
|
+
para.runs = _new_runs
|
|
309
|
+
else:
|
|
310
|
+
# 单片段:按其自身 bold 标志处理——纯粗体段(**全段**)整段加粗;
|
|
311
|
+
# raw/cleaned 不匹配回退的单片段 bold=False 保持不加粗(I20 修复:
|
|
312
|
+
# 此前把回退也整段加粗,导致"加粗+链接/代码"段落整段变粗)
|
|
313
|
+
if _segments and _segments[0][1]:
|
|
314
|
+
for r in para.runs:
|
|
315
|
+
r.format.bold = True
|
|
262
316
|
|
|
263
317
|
if is_list and list_indent_pt > 0 and not para.is_heading:
|
|
264
318
|
para.format.left_indent_pt = list_indent_pt
|
|
@@ -74,19 +74,33 @@ def _select_paragraphs(model: DocumentModel, target: str) -> list[Paragraph]:
|
|
|
74
74
|
return [p for p in model.paragraphs
|
|
75
75
|
if not p.is_heading and p.text.strip() and id(p) not in sig_set]
|
|
76
76
|
elif target == "signature":
|
|
77
|
-
#
|
|
78
|
-
|
|
77
|
+
# P2-24 修复:signature target 只匹配署名段(role='signature'),不再包含 date——
|
|
78
|
+
# 此前匹配 role in ('signature','date') 会把落款日期段一并选中,
|
|
79
|
+
# 导致 FIX-C013(署名居中)把日期改成 center、FIX-C013b(18pt)把日期改成 18pt,
|
|
80
|
+
# 违反 GB/T 9704 成文日期"右空四字、三号仿宋(16pt)"的规范。
|
|
81
|
+
# 日期段由 target="date" 分支独立处理(右对齐 + right_indent 保留)。
|
|
82
|
+
role_sig = [p for p in model.paragraphs if p.role == 'signature']
|
|
79
83
|
if role_sig:
|
|
80
84
|
return role_sig
|
|
81
|
-
|
|
82
|
-
|
|
85
|
+
# 仅当无署名 role 时,才允许位置回退(末两段:署名+日期);
|
|
86
|
+
# 回退时仍只修署名语义的段落,不影响日期段对齐
|
|
87
|
+
non_empty = [p for p in model.paragraphs if p.text.strip() and p.role != 'date']
|
|
88
|
+
if len(non_empty) >= 2:
|
|
89
|
+
last = non_empty[-1].text.strip()
|
|
90
|
+
if re.match(r'^\d{4}年\d{1,2}月\d{1,2}日$', last) or re.match(r'^\d{4}[.\-/]\d{1,2}[.\-/]\d{1,2}$', last):
|
|
91
|
+
return non_empty[-1:]
|
|
92
|
+
return []
|
|
83
93
|
elif target == "date":
|
|
84
94
|
# 同 signature 的处理逻辑
|
|
85
95
|
role_date = [p for p in model.paragraphs if p.role == 'date']
|
|
86
96
|
if role_date:
|
|
87
97
|
return role_date
|
|
88
98
|
non_empty = [p for p in model.paragraphs if p.text.strip()]
|
|
89
|
-
|
|
99
|
+
if non_empty:
|
|
100
|
+
last = non_empty[-1].text.strip()
|
|
101
|
+
if re.match(r'^\d{4}年\d{1,2}月\d{1,2}日$', last) or re.match(r'^\d{4}[.\-/]\d{1,2}[.\-/]\d{1,2}$', last):
|
|
102
|
+
return non_empty[-1:]
|
|
103
|
+
return []
|
|
90
104
|
elif target in ('salutation', 'introduction', 'transition', 'meeting_date', 'numbered_body'):
|
|
91
105
|
# N2: 段落类型 target —— 使用 detect_paragraph_type 内容匹配
|
|
92
106
|
return [p for p in model.paragraphs if detect_paragraph_type(p.text, p.role) == target]
|
|
@@ -51,6 +51,10 @@ _H1_WESTERN_PATTERN = re.compile(r'^[IVXLCDM]+\.\s+')
|
|
|
51
51
|
_H2_WESTERN_PATTERN = re.compile(r'^[A-Z]\.\s+')
|
|
52
52
|
_H4_WESTERN_PATTERN = re.compile(r'^[a-z]\.\s+')
|
|
53
53
|
|
|
54
|
+
# 句末标点(P2-16 修复:标题不以这些标点结尾,正文强调句/引用句通常会以它们结尾,
|
|
55
|
+
# 用于排除"黑体/楷体/仿宋加粗的正文短句"被误判为标题)
|
|
56
|
+
_SENTENCE_END = ('。', '!', '?', ':', ';', ',', '.', '!', '?', ':', ';', ',')
|
|
57
|
+
|
|
54
58
|
|
|
55
59
|
def _detect_heading_heuristic(
|
|
56
60
|
text: str, runs: list[Run], para_format: ParagraphFormat
|
|
@@ -103,7 +107,13 @@ def _detect_heading_heuristic(
|
|
|
103
107
|
|
|
104
108
|
# --- Level 1: 一级标题(黑体)---
|
|
105
109
|
if has_font_signal and ("黑体" in font_lower or font_lower in ("simhei",)):
|
|
106
|
-
|
|
110
|
+
# P2-16 修复:排除以句末标点结尾的正文强调句(如"重要提示:…")。
|
|
111
|
+
# 一级标题通常较短且不以句号/冒号等结尾。
|
|
112
|
+
if alignment == "center":
|
|
113
|
+
return True, 1
|
|
114
|
+
if len(text_stripped) < 30 and not text_stripped.endswith(_SENTENCE_END):
|
|
115
|
+
return True, 1
|
|
116
|
+
if is_bold and not text_stripped.endswith(_SENTENCE_END):
|
|
107
117
|
return True, 1
|
|
108
118
|
|
|
109
119
|
# 格式信号:"一、" + 加粗或黑体
|
|
@@ -115,9 +125,15 @@ def _detect_heading_heuristic(
|
|
|
115
125
|
if _H1_WESTERN_PATTERN.match(text_stripped) and len(text_stripped) < 50:
|
|
116
126
|
return True, 1
|
|
117
127
|
|
|
118
|
-
# --- Level 2:
|
|
119
|
-
|
|
120
|
-
|
|
128
|
+
# --- Level 2: 二级标题(楷体,无需加粗)---
|
|
129
|
+
# GB/T 9704 二级标题为楷体,不加粗。若段落字体为楷体(含 楷体_GB2312),
|
|
130
|
+
# 即使内容为 "2.1" 等数字编号,也优先识别为二级标题而非三级标题,
|
|
131
|
+
# 避免 converter 的 "###"→二级标题与 check 的 "2.1"→三级标题启发式冲突。
|
|
132
|
+
# P2-16 修复:增加"不以句末标点结尾"约束,排除正文楷体引用句
|
|
133
|
+
# (如"会议指出,…。")被误判为二级标题。
|
|
134
|
+
if has_font_signal and ("楷体" in font_lower or font_lower in ("kaiti", "楷体_gb2312")):
|
|
135
|
+
if len(text_stripped) < 40 and not text_stripped.endswith(_SENTENCE_END):
|
|
136
|
+
return True, 2
|
|
121
137
|
|
|
122
138
|
# "(一)" 格式(无论字体如何,此模式足够唯一)
|
|
123
139
|
if _H2_PATTERN.match(text_stripped) and len(text_stripped) < 50:
|
|
@@ -128,7 +144,8 @@ def _detect_heading_heuristic(
|
|
|
128
144
|
|
|
129
145
|
# --- Level 3: 三级标题(仿宋加粗 或 "1." + 加粗)---
|
|
130
146
|
if has_font_signal and ("仿宋" in font_lower or font_lower in ("fangsong", "仿宋_gb2312")) and is_bold:
|
|
131
|
-
|
|
147
|
+
# P2-16 修复:排除以句末标点结尾的仿宋加粗正文短句被误判为三级标题
|
|
148
|
+
if len(text_stripped) < 50 and not text_stripped.endswith(_SENTENCE_END):
|
|
132
149
|
return True, 3
|
|
133
150
|
|
|
134
151
|
if _H3_PATTERN.match(text_stripped) and is_bold and len(text_stripped) < 60:
|
|
@@ -340,6 +357,13 @@ def _assign_paragraph_roles(paragraphs: list[Paragraph]) -> None:
|
|
|
340
357
|
if not para.role and _ATTACHMENT_RE.match(para.text.strip()):
|
|
341
358
|
para.role = 'attachment'
|
|
342
359
|
|
|
360
|
+
# AI 声明段(末尾批注,如"(内容由GongWen-skill-AI生成,仅供参考)")
|
|
361
|
+
# 标记为 annotation 角色,避免 check 将其误判为正文并报格式违规
|
|
362
|
+
_AI_DECL_MARKERS = ("由GongWen-skill-AI生成", "由AI生成", "仅供参考")
|
|
363
|
+
for idx, para in non_empty:
|
|
364
|
+
if not para.role and any(m in para.text for m in _AI_DECL_MARKERS):
|
|
365
|
+
para.role = 'annotation'
|
|
366
|
+
|
|
343
367
|
# 其余非空段落默认为 body
|
|
344
368
|
for idx, para in non_empty:
|
|
345
369
|
if not para.role:
|
|
@@ -92,6 +92,16 @@ def check_document(model: DocumentModel, rules: dict[str, Any]) -> list[CheckIss
|
|
|
92
92
|
elif field_path.startswith("ending."):
|
|
93
93
|
# FIX-V153-02:ending.check 结语检查(CHK-N001/R001/RPT001/RP002/L002 等)
|
|
94
94
|
issues.extend(_check_ending(model, rule_id, severity, name, expected, message))
|
|
95
|
+
elif field_path.startswith("content."):
|
|
96
|
+
# P2-22 修复:content.* 内容要素检查(CHK-N003 等 37 条规则此前无分发分支,
|
|
97
|
+
# 全部"定义但从不执行"且每次 check 刷"未支持字段"告警)
|
|
98
|
+
issues.extend(_check_content_field(model, rule_id, severity, name, field_path,
|
|
99
|
+
expected, message))
|
|
100
|
+
elif field_path.startswith("header."):
|
|
101
|
+
# P2-23 修复:header.* 版头检查(CHK-CM002 令号、CHK-R003 主送机关)
|
|
102
|
+
# 此前无分发分支,规则定义但从不执行
|
|
103
|
+
issues.extend(_check_header_field(model, rule_id, severity, name, field_path,
|
|
104
|
+
expected, message))
|
|
95
105
|
else:
|
|
96
106
|
# P1-6 修复:删除 generic else 中的硬编码索引逻辑(model.paragraphs[0]/[1]
|
|
97
107
|
# 不一定是标题/正文,检查结果会指向错误段落),未识别的 field 直接 skip + warning
|
|
@@ -313,10 +323,32 @@ def _check_heading_level(model, rule_id, severity, name, field_path, expected, m
|
|
|
313
323
|
return issues
|
|
314
324
|
|
|
315
325
|
|
|
326
|
+
# P2-17:CHK-C030 辅助——纯标点 run 判定(不含任何正文内容的 run)
|
|
327
|
+
# 用字符集判断替代正则,避免 r'...\[...]' 无效转义 SyntaxWarning,且匹配更可靠
|
|
328
|
+
_PUNCT_ONLY_CHARS = set('。!?:;,、.…!?:;,()()""''《》〈〉【】[]')
|
|
329
|
+
|
|
330
|
+
|
|
331
|
+
def _is_punct_only(text: str) -> bool:
|
|
332
|
+
"""判断 run 是否只含标点/空白(用于整段加粗判定时忽略标点 run)。"""
|
|
333
|
+
t = (text or "").strip()
|
|
334
|
+
if not t:
|
|
335
|
+
return True
|
|
336
|
+
return all(ch in _PUNCT_ONLY_CHARS or ch.isspace() for ch in t)
|
|
337
|
+
|
|
338
|
+
|
|
339
|
+
def _count_sentence_end(text: str) -> int:
|
|
340
|
+
"""统计句末标点(。!? 以及英文 . ! ?)的数量,用于判断单句/多句段落。"""
|
|
341
|
+
if not text:
|
|
342
|
+
return 0
|
|
343
|
+
return sum(1 for ch in text if ch in '。!?.!?')
|
|
344
|
+
|
|
345
|
+
|
|
316
346
|
def _check_body(model, rule_id, severity, name, field_path, expected, message) -> list[CheckIssue]:
|
|
317
|
-
"""Check body paragraph formatting (excluding signature/date)."""
|
|
347
|
+
"""Check body paragraph formatting (excluding signature/date/annotation/recipient)."""
|
|
318
348
|
issues = []
|
|
319
|
-
|
|
349
|
+
# 顶格左对齐的段落(主送机关/称呼段、署名、日期、AI 声明批注)不属于正文,
|
|
350
|
+
# 不应套用正文的缩进/对齐/字体检查
|
|
351
|
+
_EXCLUDE_ROLES = {'signature', 'date', 'annotation', 'notes', 'recipient', 'salutation'}
|
|
320
352
|
body_paras = [p for p in model.paragraphs
|
|
321
353
|
if not p.is_heading and p.text.strip() and p.role not in _EXCLUDE_ROLES]
|
|
322
354
|
if not body_paras:
|
|
@@ -403,13 +435,23 @@ def _check_body(model, rule_id, severity, name, field_path, expected, message) -
|
|
|
403
435
|
elif sub_field == "bold_range":
|
|
404
436
|
# 检查正文段落是否整段加粗(通常只有首句/点题词应加粗)
|
|
405
437
|
if para.runs and para.text.strip():
|
|
406
|
-
|
|
438
|
+
# P2-17 修复:忽略纯标点 run(句号/逗号等不含正文的 run),
|
|
439
|
+
# 避免"首句加粗含句号"时标点 run 加粗触发整段加粗误报
|
|
440
|
+
_CONTENT_RUNS = [r for r in para.runs
|
|
441
|
+
if r.text.strip() and not _is_punct_only(r.text)]
|
|
442
|
+
if not _CONTENT_RUNS:
|
|
443
|
+
continue
|
|
444
|
+
all_bold = all(r.format.bold for r in _CONTENT_RUNS)
|
|
407
445
|
if all_bold:
|
|
408
446
|
# B-09(方案三):排除不应加粗的段落类型(称呼/导语/过渡/署名/会议日期等),
|
|
409
447
|
# 避免这些段落被误标为 body 后报告"整段加粗"问题造成噪音
|
|
410
448
|
from engine.core.document.modifier import should_bold_first_sentence
|
|
411
449
|
if not should_bold_first_sentence(para.text, para.role):
|
|
412
450
|
continue
|
|
451
|
+
# P2-17 修复:单句段落(仅 1 个句末标点)的首句加粗=整段加粗,
|
|
452
|
+
# 属于合理排版(首句即整段),不报;多句段落整段加粗才报
|
|
453
|
+
if _count_sentence_end(para.text) <= 1:
|
|
454
|
+
continue
|
|
413
455
|
issues.append(CheckIssue(
|
|
414
456
|
rule_id=rule_id, check_type="content", severity=severity,
|
|
415
457
|
name=name, location=f"paragraph:{para.index}",
|
|
@@ -463,17 +505,32 @@ def _check_page_setup(model, rule_id, severity, name, field_path, expected, mess
|
|
|
463
505
|
|
|
464
506
|
|
|
465
507
|
def _check_signature_area(model, rule_id, severity, name, field_path, expected, message, rules) -> list[CheckIssue]:
|
|
466
|
-
"""Check signature/date area formatting.
|
|
508
|
+
"""Check signature/date area formatting.
|
|
509
|
+
|
|
510
|
+
仅检查落款/日期段落(默认取最后 2 个非空段落):
|
|
511
|
+
- signature.* 字段只检查署名段(角色 signature,通常是倒数第 2 段)
|
|
512
|
+
- date.* 字段只检查日期段(角色 date,通常是最后 1 段)
|
|
513
|
+
避免把署名/日期规则同时套在两个段落上造成误报。
|
|
514
|
+
"""
|
|
467
515
|
issues = []
|
|
468
|
-
paras = [p for p in model.paragraphs if not p.is_heading and p.text.strip()
|
|
516
|
+
paras = [p for p in model.paragraphs if not p.is_heading and p.text.strip()
|
|
517
|
+
and p.role not in ('annotation', 'notes')]
|
|
469
518
|
if not paras:
|
|
470
519
|
return issues
|
|
471
520
|
|
|
472
|
-
# Signature area: only last 2 paragraphs (落款单位 + 日期)
|
|
473
|
-
sig_paras = paras[-2:] if len(paras) >= 2 else paras
|
|
474
521
|
sub_field = field_path.split(".", 1)[1] if "." in field_path else ""
|
|
522
|
+
is_date = field_path.startswith("date.")
|
|
475
523
|
|
|
476
|
-
|
|
524
|
+
# 按角色取签名段/日期段。仅当文档确实存在落款/日期时才检查,
|
|
525
|
+
# 避免把无落款的正文末段误判为签名/日期造成误报(位置回退不再使用)。
|
|
526
|
+
if is_date:
|
|
527
|
+
target = [p for p in paras if p.role == 'date']
|
|
528
|
+
else:
|
|
529
|
+
target = [p for p in paras if p.role == 'signature']
|
|
530
|
+
if not target:
|
|
531
|
+
return issues
|
|
532
|
+
|
|
533
|
+
for para in target:
|
|
477
534
|
if sub_field == "align":
|
|
478
535
|
if para.format.alignment and para.format.alignment != str(expected).lower():
|
|
479
536
|
issues.append(CheckIssue(
|
|
@@ -561,8 +618,11 @@ def _check_paragraph_type_field(model, rule_id, severity, name, field_path, expe
|
|
|
561
618
|
))
|
|
562
619
|
break
|
|
563
620
|
elif sub_field == "bold":
|
|
564
|
-
|
|
565
|
-
|
|
621
|
+
# 编号正文(一是/二是…):仅要求首句(首 run)加粗,其余正文不应加粗。
|
|
622
|
+
# 其余段落类型:按首 run 判断即可(段落级加粗风格由 bold_range 规则另行检查)。
|
|
623
|
+
_runs_to_check = para.runs[:1] if target == 'numbered_body' else para.runs[:1]
|
|
624
|
+
for run in _runs_to_check:
|
|
625
|
+
if run.format.bold is not None and bool(run.format.bold) != bool(expected):
|
|
566
626
|
issues.append(CheckIssue(
|
|
567
627
|
rule_id=rule_id, check_type="format", severity=severity,
|
|
568
628
|
name=name, location=f"paragraph:{para.index}",
|
|
@@ -646,8 +706,9 @@ def _check_ending(model, rule_id: str, severity: str, name: str,
|
|
|
646
706
|
text = text.strip()
|
|
647
707
|
role = getattr(p, 'role', '') or ''
|
|
648
708
|
# 排除标题、落款(signature/date)、批注(annotation)段——这些不是正文结语
|
|
649
|
-
|
|
650
|
-
|
|
709
|
+
# P2-22 修复:模型属性是 is_heading 而非 is_title,标题段不应纳入结语统计
|
|
710
|
+
if text and not getattr(p, 'is_heading', False) and role not in (
|
|
711
|
+
'signature', 'date', 'annotation'):
|
|
651
712
|
body_paras.append(text)
|
|
652
713
|
|
|
653
714
|
if not body_paras:
|
|
@@ -684,6 +745,159 @@ def _check_ending(model, rule_id: str, severity: str, name: str,
|
|
|
684
745
|
return issues
|
|
685
746
|
|
|
686
747
|
|
|
748
|
+
# P2-22:content.* 内容要素检查——字段名 → 要素关键词组(宽松匹配,高召回低误报)
|
|
749
|
+
_CONTENT_FIELD_KEYWORDS = {
|
|
750
|
+
"notice_items": ["时间", "地点", "要求", "请", "须", "要", "参加", "召开", "举办",
|
|
751
|
+
"组织", "开展", "落实", "遵守", "上报", "报送", "日期", "人员", "对象"],
|
|
752
|
+
"scope": ["范围", "适用于", "各", "单位", "部门", "地区", "辖区"],
|
|
753
|
+
"effective_date": ["自", "起施行", "施行", "生效", "即日", "之日起", "起执行"],
|
|
754
|
+
"validity": ["有效", "期限", "至", "自", "起"],
|
|
755
|
+
"meeting_elements": ["会议", "时间", "地点", "参加", "人员", "议题", "议程"],
|
|
756
|
+
"meeting_info": ["会议", "时间", "地点", "参加", "纪要", "议题"],
|
|
757
|
+
"reason": ["因", "由于", "为了", "鉴于", "依据", "根据"],
|
|
758
|
+
"basis": ["依据", "根据", "按照", "遵照"],
|
|
759
|
+
"purpose": ["为了", "为", "目的", "促进", "推动"],
|
|
760
|
+
"measures": ["措施", "办法", "方案", "要求", "应当", "应", "须"],
|
|
761
|
+
"proposer": ["提出", "建议", "提议", "呈报", "申报"],
|
|
762
|
+
"facts": ["事实", "情况", "经查", "查明", "核实"],
|
|
763
|
+
"items": ["事项", "内容", "如下", "包括", "如下"],
|
|
764
|
+
"legal_basis": ["依据", "根据", "依照", "按照", "法规", "条例"],
|
|
765
|
+
"decision_items": ["决定", "如下", "事项", "内容"],
|
|
766
|
+
"clauses": ["条", "款", "项", "规定", "如下"],
|
|
767
|
+
"structure": ["结构", "如下", "部分", "章节"],
|
|
768
|
+
"data": ["数据", "统计", "指标", "数字", "情况"],
|
|
769
|
+
"suggestions": ["建议", "意见", "应", "应当", "建议如下"],
|
|
770
|
+
"report_items": ["情况", "报告", "如下", "内容"],
|
|
771
|
+
"reply_to": ["关于", "收悉", "来函", "贵", "你"],
|
|
772
|
+
"attitude": ["同意", "不同意", "原则同意", "批准"],
|
|
773
|
+
"single_topic": ["一", "单一", "专项"],
|
|
774
|
+
"resolution_items": ["决定", "如下", "事项"],
|
|
775
|
+
"procedure": ["程序", "步骤", "按照", "流程"],
|
|
776
|
+
"background": ["背景", "概述", "现状", "问题"],
|
|
777
|
+
"alternatives": ["方案", "备选", "比较", "选项"],
|
|
778
|
+
"implementation_plan": ["实施", "计划", "进度", "安排", "时间表"],
|
|
779
|
+
"objectives": ["目标", "目的", "要求"],
|
|
780
|
+
"timeline": ["时间", "阶段", "进度", "月", "年", "日"],
|
|
781
|
+
"responsibilities": ["责任", "负责", "分工", "单位"],
|
|
782
|
+
"lead": ["导语", "开头", "首先"],
|
|
783
|
+
"source": ["来源", "据", "报道"],
|
|
784
|
+
"report_section": ["部分", "章节", "如下"],
|
|
785
|
+
}
|
|
786
|
+
|
|
787
|
+
|
|
788
|
+
def _check_content_field(model, rule_id: str, severity: str, name: str,
|
|
789
|
+
field_path: str, expected: str, message: str) -> list[CheckIssue]:
|
|
790
|
+
"""P2-22 修复:content.* 内容要素检查(此前所有该前缀规则被跳过并告警)。
|
|
791
|
+
|
|
792
|
+
策略:按字段名尾部映射要素关键词组,检查正文(body 段落)是否包含任一组关键词。
|
|
793
|
+
- 宽松匹配:命中任一关键词即通过(避免误报)
|
|
794
|
+
- 无法映射的字段名保持跳过(返回空列表,不产生告警噪音)
|
|
795
|
+
"""
|
|
796
|
+
issues = []
|
|
797
|
+
field_name = field_path.split(".", 1)[1] if "." in field_path else ""
|
|
798
|
+
keywords = _CONTENT_FIELD_KEYWORDS.get(field_name)
|
|
799
|
+
if not keywords:
|
|
800
|
+
return issues
|
|
801
|
+
|
|
802
|
+
try:
|
|
803
|
+
# 收集正文文本(排除标题/落款/日期/批注/称呼段)
|
|
804
|
+
body_texts = []
|
|
805
|
+
for p in model.paragraphs:
|
|
806
|
+
text = (getattr(p, 'text', '') or '').strip()
|
|
807
|
+
role = getattr(p, 'role', '') or ''
|
|
808
|
+
# P2-22 修复:模型属性是 is_heading 而非 is_title——原 is_title 永远为
|
|
809
|
+
# False,导致标题段被误纳入正文统计(标题中的"召开/组织"等词会使
|
|
810
|
+
# 空壳通知误判为"有要素"而漏报 CHK-N003)
|
|
811
|
+
if text and not getattr(p, 'is_heading', False) and role not in (
|
|
812
|
+
'signature', 'date', 'annotation', 'recipient', 'salutation'):
|
|
813
|
+
body_texts.append(text)
|
|
814
|
+
joined = ''.join(body_texts)
|
|
815
|
+
if not joined:
|
|
816
|
+
return issues
|
|
817
|
+
|
|
818
|
+
found = any(kw in joined for kw in keywords)
|
|
819
|
+
if not found:
|
|
820
|
+
issues.append(CheckIssue(
|
|
821
|
+
rule_id=rule_id,
|
|
822
|
+
check_type="content",
|
|
823
|
+
severity=severity,
|
|
824
|
+
name=name,
|
|
825
|
+
location="正文",
|
|
826
|
+
original_text=joined[-80:] if joined else "",
|
|
827
|
+
suggested_fix=expected,
|
|
828
|
+
reason=message or f"文档正文缺少相关要素(期望: {expected})",
|
|
829
|
+
))
|
|
830
|
+
except Exception as e:
|
|
831
|
+
logger.warning(f"_check_content_field 检查失败: {e}")
|
|
832
|
+
|
|
833
|
+
return issues
|
|
834
|
+
|
|
835
|
+
|
|
836
|
+
def _check_header_field(model, rule_id: str, severity: str, name: str,
|
|
837
|
+
field_path: str, expected: str, message: str) -> list[CheckIssue]:
|
|
838
|
+
"""P2-23 修复:header.* 版头检查(此前规则定义但从不执行)。
|
|
839
|
+
|
|
840
|
+
支持字段:
|
|
841
|
+
- header.doc_number(CHK-CM002 命令):检查文档是否标注令号(如"〔2026〕1号"、"第1号")
|
|
842
|
+
- header.recipient(CHK-R003 请示):检查是否含主送机关段,且只写一个主送机关
|
|
843
|
+
"""
|
|
844
|
+
import re
|
|
845
|
+
issues = []
|
|
846
|
+
field_name = field_path.split(".", 1)[1] if "." in field_path else ""
|
|
847
|
+
|
|
848
|
+
try:
|
|
849
|
+
if field_name == "doc_number":
|
|
850
|
+
# 令号模式:〔2026〕1号 / 第1号 / (2026)1号 / 2026年1号 等
|
|
851
|
+
all_text = ''.join((getattr(p, 'text', '') or '') for p in model.paragraphs)
|
|
852
|
+
has_doc_number = re.search(
|
|
853
|
+
r'[〔((]?\d{4}[〕))]?\s*\d+号|[第]\d+号|令\d+号|\d+号令',
|
|
854
|
+
all_text,
|
|
855
|
+
) is not None
|
|
856
|
+
if not has_doc_number:
|
|
857
|
+
issues.append(CheckIssue(
|
|
858
|
+
rule_id=rule_id, check_type="content", severity=severity,
|
|
859
|
+
name=name, location="文档开头",
|
|
860
|
+
original_text=all_text[:60] if all_text else "",
|
|
861
|
+
suggested_fix=expected,
|
|
862
|
+
reason=message or "命令(令)应标注令号",
|
|
863
|
+
))
|
|
864
|
+
elif field_name == "recipient":
|
|
865
|
+
# 主送机关:存在 recipient 角色的段,且应只有一个主送机关
|
|
866
|
+
recips = [p for p in model.paragraphs
|
|
867
|
+
if getattr(p, 'role', '') == 'recipient' and (p.text or '').strip()]
|
|
868
|
+
if not recips:
|
|
869
|
+
issues.append(CheckIssue(
|
|
870
|
+
rule_id=rule_id, check_type="content", severity=severity,
|
|
871
|
+
name=name, location="文档开头",
|
|
872
|
+
original_text="",
|
|
873
|
+
suggested_fix=expected,
|
|
874
|
+
reason=message or "请示应写明主送机关",
|
|
875
|
+
))
|
|
876
|
+
return issues
|
|
877
|
+
# 只写一个主送机关:一个 recipient 段 + 段内无多个机关分隔符
|
|
878
|
+
first = recips[0].text
|
|
879
|
+
if len(recips) > 1:
|
|
880
|
+
issues.append(CheckIssue(
|
|
881
|
+
rule_id=rule_id, check_type="content", severity=severity,
|
|
882
|
+
name=name, location=f"paragraph:{recips[0].index}",
|
|
883
|
+
original_text=first[:60],
|
|
884
|
+
suggested_fix="仅保留一个主送机关",
|
|
885
|
+
reason="请示一般只写一个主送机关(发现多个主送机关段)",
|
|
886
|
+
))
|
|
887
|
+
elif re.search(r'[、,,s]{2,}', first.strip().rstrip('::')) and len(first.strip().rstrip('::')) > 5:
|
|
888
|
+
issues.append(CheckIssue(
|
|
889
|
+
rule_id=rule_id, check_type="content", severity=severity,
|
|
890
|
+
name=name, location=f"paragraph:{recips[0].index}",
|
|
891
|
+
original_text=first[:60],
|
|
892
|
+
suggested_fix="仅保留一个主送机关",
|
|
893
|
+
reason="请示一般只写一个主送机关(段内含多个机关)",
|
|
894
|
+
))
|
|
895
|
+
except Exception as e:
|
|
896
|
+
logger.warning(f"_check_header_field 检查失败: {e}")
|
|
897
|
+
|
|
898
|
+
return issues
|
|
899
|
+
|
|
900
|
+
|
|
687
901
|
def _check_page_number(model, rule_id, severity, name, field_path, expected, message) -> list[CheckIssue]:
|
|
688
902
|
"""检查页脚页码域格式(P3-10:CHK-C023/024/029)。
|
|
689
903
|
|
|
@@ -711,7 +925,9 @@ def _check_page_number(model, rule_id, severity, name, field_path, expected, mes
|
|
|
711
925
|
break
|
|
712
926
|
elif sub_field == "alignment" or sub_field == "align":
|
|
713
927
|
actual = para.format.alignment
|
|
714
|
-
|
|
928
|
+
# GB/T 9704 允许"居中"或"翻页模式(单右双左,默认 footer 为 right/left)"。
|
|
929
|
+
# 翻页模式下默认 footer 对齐为 right(单页右空一字),不再强制 center。
|
|
930
|
+
if actual and actual not in ('center', 'left', 'right'):
|
|
715
931
|
issues.append(CheckIssue(
|
|
716
932
|
rule_id=rule_id, check_type="format", severity=severity,
|
|
717
933
|
name=name, location=f"footer:{hf.section_index}:paragraph:{para.index}",
|