wechat-article-parser 0.0.6__tar.gz → 0.0.7__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,6 +1,6 @@
1
- Metadata-Version: 2.4
1
+ Metadata-Version: 2.5
2
2
  Name: wechat-article-parser
3
- Version: 0.0.6
3
+ Version: 0.0.7
4
4
  Summary: WeChat MP article parser - extract metadata and content from WeChat public account articles
5
5
  License-Expression: MIT
6
6
  License-File: LICENSE
@@ -26,6 +26,10 @@ Description-Content-Type: text/markdown
26
26
  - 小红书风格图片轮播文章
27
27
  - 全屏布局短文本文章
28
28
 
29
+ 解析时优先读取页面全局的 `item_show_type`:`0` 对应图文、`5` 对应视频分享、
30
+ `8` 对应图片轮播、`10` 对应文字消息,再通过 HTML 或脚本数据区分具体模板。
31
+ 类型字段缺失、值未知或对应分支未提取到正文时,会按原有 HTML 特征匹配顺序兜底。
32
+
29
33
  ## 安装
30
34
 
31
35
  ```bash
@@ -108,13 +112,42 @@ result = parse(
108
112
  | `article_title` | `str` | 文章标题 |
109
113
  | `article_cover_image` | `str` | 文章封面图链接 |
110
114
  | `article_description` | `str` | 文章摘要 |
111
- | `article_markdown` | `str` | 文章正文的 Markdown 内容 |
115
+ | `article_markdown` | `str` | 文章正文的 Markdown 内容;视频分享没有文字附言时可为空 |
112
116
  | `article_publish_time` | `int` | 发布时间(Unix 时间戳) |
117
+ | `article_display_type` | `ArticleDisplayType` | 文章展示类型:`article`、`video`、`image`、`text` 或 `unknown` |
113
118
  | `raw_html` | `str` | 原始 HTML 网页源代码(仅在 `include_raw_html=True` 时填充,否则为空字符串) |
114
119
  | `images` | `list[str]` | 文章中提取的所有图片链接 |
115
120
  | `is_valid` | `bool` | 关键字段是否全部解析成功(属性) |
116
121
 
117
- `is_valid` 为 `True` 的条件:`mp_id`、`mp_name`、`article_id`、`article_msg_id`、`article_idx`、`article_sn`、`article_title`、`article_markdown`、`article_publish_time` 均不为空/零。
122
+ `is_valid` 为 `True` 的条件:`mp_id`、`mp_name`、`article_id`、`article_msg_id`、`article_idx`、`article_sn`、`article_title`、`article_publish_time` 均不为空/零,且 `article_markdown` 非空或 `article_display_type == "video"`。视频分享允许没有文字正文,仍须满足其他关键字段要求;有效性不表示已获取视频文件。
123
+
124
+ ### 判断文章展示类型
125
+
126
+ `ArticleDisplayType` 是字符串枚举,既支持枚举比较,也支持语义字符串比较:
127
+
128
+ | 枚举 | 字符串值 | 含义 |
129
+ |---|---|---|
130
+ | `ArticleDisplayType.ARTICLE` | `article` | 图文,包括转载式 |
131
+ | `ArticleDisplayType.VIDEO` | `video` | 视频分享 |
132
+ | `ArticleDisplayType.IMAGE` | `image` | 图片轮播 |
133
+ | `ArticleDisplayType.TEXT` | `text` | 文字消息 |
134
+ | `ArticleDisplayType.UNKNOWN` | `unknown` | 页面类型缺失或未支持 |
135
+
136
+ ```python
137
+ from wechat_article_parser import parse, ArticleDisplayType
138
+
139
+ result = parse("https://mp.weixin.qq.com/s/xxxxx")
140
+ if result.article_display_type == ArticleDisplayType.VIDEO:
141
+ print("这是视频分享")
142
+
143
+ if result.article_display_type == "video":
144
+ print("视频分享可以没有文字正文")
145
+
146
+ print(result.article_display_type) # 例如 video
147
+ ```
148
+
149
+ `article_display_type` 是返回结果中的数据字段,包含在 `dataclasses.asdict(result)` 中,JSON 序列化时输出语义字符串。
150
+ 微信原始的 `item_show_type` 仅在解析器内部使用,不包含在返回结果中。
118
151
 
119
152
  ### 判断账号类型
120
153
 
@@ -168,7 +201,7 @@ except httpx.HTTPStatusError as e:
168
201
 
169
202
  ### 解析不完整
170
203
 
171
- 部分文章类型可能无法提取所有字段(如某些文章没有封面图或摘要)。这种情况不会抛出异常,但 `result.is_valid` 会返回 `False`。建议在业务逻辑中检查:
204
+ 缺少标题、发布时间等关键字段,或非视频分享缺少正文时,解析结果的 `is_valid` 会返回 `False`。封面图、摘要等非必填字段为空不影响有效性。建议在业务逻辑中检查:
172
205
 
173
206
  ```python
174
207
  result = parse("https://mp.weixin.qq.com/s/xxxxx")
@@ -11,6 +11,10 @@
11
11
  - 小红书风格图片轮播文章
12
12
  - 全屏布局短文本文章
13
13
 
14
+ 解析时优先读取页面全局的 `item_show_type`:`0` 对应图文、`5` 对应视频分享、
15
+ `8` 对应图片轮播、`10` 对应文字消息,再通过 HTML 或脚本数据区分具体模板。
16
+ 类型字段缺失、值未知或对应分支未提取到正文时,会按原有 HTML 特征匹配顺序兜底。
17
+
14
18
  ## 安装
15
19
 
16
20
  ```bash
@@ -93,13 +97,42 @@ result = parse(
93
97
  | `article_title` | `str` | 文章标题 |
94
98
  | `article_cover_image` | `str` | 文章封面图链接 |
95
99
  | `article_description` | `str` | 文章摘要 |
96
- | `article_markdown` | `str` | 文章正文的 Markdown 内容 |
100
+ | `article_markdown` | `str` | 文章正文的 Markdown 内容;视频分享没有文字附言时可为空 |
97
101
  | `article_publish_time` | `int` | 发布时间(Unix 时间戳) |
102
+ | `article_display_type` | `ArticleDisplayType` | 文章展示类型:`article`、`video`、`image`、`text` 或 `unknown` |
98
103
  | `raw_html` | `str` | 原始 HTML 网页源代码(仅在 `include_raw_html=True` 时填充,否则为空字符串) |
99
104
  | `images` | `list[str]` | 文章中提取的所有图片链接 |
100
105
  | `is_valid` | `bool` | 关键字段是否全部解析成功(属性) |
101
106
 
102
- `is_valid` 为 `True` 的条件:`mp_id`、`mp_name`、`article_id`、`article_msg_id`、`article_idx`、`article_sn`、`article_title`、`article_markdown`、`article_publish_time` 均不为空/零。
107
+ `is_valid` 为 `True` 的条件:`mp_id`、`mp_name`、`article_id`、`article_msg_id`、`article_idx`、`article_sn`、`article_title`、`article_publish_time` 均不为空/零,且 `article_markdown` 非空或 `article_display_type == "video"`。视频分享允许没有文字正文,仍须满足其他关键字段要求;有效性不表示已获取视频文件。
108
+
109
+ ### 判断文章展示类型
110
+
111
+ `ArticleDisplayType` 是字符串枚举,既支持枚举比较,也支持语义字符串比较:
112
+
113
+ | 枚举 | 字符串值 | 含义 |
114
+ |---|---|---|
115
+ | `ArticleDisplayType.ARTICLE` | `article` | 图文,包括转载式 |
116
+ | `ArticleDisplayType.VIDEO` | `video` | 视频分享 |
117
+ | `ArticleDisplayType.IMAGE` | `image` | 图片轮播 |
118
+ | `ArticleDisplayType.TEXT` | `text` | 文字消息 |
119
+ | `ArticleDisplayType.UNKNOWN` | `unknown` | 页面类型缺失或未支持 |
120
+
121
+ ```python
122
+ from wechat_article_parser import parse, ArticleDisplayType
123
+
124
+ result = parse("https://mp.weixin.qq.com/s/xxxxx")
125
+ if result.article_display_type == ArticleDisplayType.VIDEO:
126
+ print("这是视频分享")
127
+
128
+ if result.article_display_type == "video":
129
+ print("视频分享可以没有文字正文")
130
+
131
+ print(result.article_display_type) # 例如 video
132
+ ```
133
+
134
+ `article_display_type` 是返回结果中的数据字段,包含在 `dataclasses.asdict(result)` 中,JSON 序列化时输出语义字符串。
135
+ 微信原始的 `item_show_type` 仅在解析器内部使用,不包含在返回结果中。
103
136
 
104
137
  ### 判断账号类型
105
138
 
@@ -153,7 +186,7 @@ except httpx.HTTPStatusError as e:
153
186
 
154
187
  ### 解析不完整
155
188
 
156
- 部分文章类型可能无法提取所有字段(如某些文章没有封面图或摘要)。这种情况不会抛出异常,但 `result.is_valid` 会返回 `False`。建议在业务逻辑中检查:
189
+ 缺少标题、发布时间等关键字段,或非视频分享缺少正文时,解析结果的 `is_valid` 会返回 `False`。封面图、摘要等非必填字段为空不影响有效性。建议在业务逻辑中检查:
157
190
 
158
191
  ```python
159
192
  result = parse("https://mp.weixin.qq.com/s/xxxxx")
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
4
4
 
5
5
  [project]
6
6
  name = "wechat-article-parser"
7
- version = "0.0.6"
7
+ version = "0.0.7"
8
8
  description = "WeChat MP article parser - extract metadata and content from WeChat public account articles"
9
9
  readme = "README.md"
10
10
  requires-python = ">=3.10"
@@ -0,0 +1,11 @@
1
+ from .models import AccountType, ArticleDisplayType, ArticleResult, WeChatVerifyError
2
+ from .parser import parse, parse_async
3
+
4
+ __all__ = [
5
+ "parse",
6
+ "parse_async",
7
+ "ArticleResult",
8
+ "AccountType",
9
+ "ArticleDisplayType",
10
+ "WeChatVerifyError",
11
+ ]
@@ -22,6 +22,22 @@ class AccountType(str, Enum):
22
22
  return format(self.value, format_spec)
23
23
 
24
24
 
25
+ class ArticleDisplayType(str, Enum):
26
+ """文章展示类型,支持枚举和语义字符串比较。"""
27
+
28
+ UNKNOWN = "unknown"
29
+ ARTICLE = "article"
30
+ VIDEO = "video"
31
+ IMAGE = "image"
32
+ TEXT = "text"
33
+
34
+ def __str__(self) -> str:
35
+ return self.value
36
+
37
+ def __format__(self, format_spec: str) -> str:
38
+ return format(self.value, format_spec)
39
+
40
+
25
41
  @dataclass
26
42
  class ArticleResult:
27
43
  """微信公众号文章的解析结果。"""
@@ -52,9 +68,12 @@ class ArticleResult:
52
68
  # 文章中提取的图片列表
53
69
  images: list[str] = field(default_factory=list)
54
70
 
71
+ # 对外返回语义展示类型,微信原始数值仅在解析器内部使用。
72
+ article_display_type: ArticleDisplayType = ArticleDisplayType.UNKNOWN
73
+
55
74
  @property
56
75
  def is_valid(self) -> bool:
57
- """检查关键字段是否解析成功。"""
76
+ """检查关键字段是否解析成功;视频分享允许正文为空。"""
58
77
  return bool(
59
78
  self.mp_id
60
79
  and self.mp_name
@@ -63,6 +82,9 @@ class ArticleResult:
63
82
  and self.article_idx
64
83
  and self.article_sn
65
84
  and self.article_title
66
- and self.article_markdown
85
+ and (
86
+ self.article_markdown
87
+ or self.article_display_type == ArticleDisplayType.VIDEO
88
+ )
67
89
  and self.article_publish_time
68
90
  )
@@ -5,13 +5,15 @@ from __future__ import annotations
5
5
  import base64
6
6
  import html as html_module
7
7
  import re
8
+ from collections import deque
9
+ from dataclasses import replace
8
10
  from urllib.parse import unquote
9
11
 
10
12
  import httpx
11
13
  from bs4 import BeautifulSoup, Tag
12
14
  from markdownify import MarkdownConverter
13
15
 
14
- from .models import AccountType, ArticleResult, WeChatVerifyError
16
+ from .models import AccountType, ArticleDisplayType, ArticleResult, WeChatVerifyError
15
17
 
16
18
  _USER_AGENT = (
17
19
  "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) "
@@ -20,6 +22,13 @@ _USER_AGENT = (
20
22
  )
21
23
  _TIMEOUT = 15
22
24
 
25
+ # 将字符串和注释作为整体,避免其中的括号干扰图片数组的层级判断。
26
+ _JS_DATA_TOKEN_RE = re.compile(
27
+ r"(?P<string>'(?:\\[\s\S]|[^'\\])*'|\"(?:\\[\s\S]|[^\"\\])*\"|`(?:\\[\s\S]|[^`\\])*`)"
28
+ r"|(?P<comment>//[^\r\n]*|/\*[\s\S]*?\*/)"
29
+ r"|[a-zA-Z_$][\w$]*|\d+(?:\.\d*)?(?:[eE][+-]?\d+)?|[^\s]"
30
+ )
31
+
23
32
 
24
33
  # ---------------------------------------------------------------------------
25
34
  # Markdown 转换器
@@ -95,18 +104,59 @@ def _service_type_to_account_type(value: str) -> AccountType:
95
104
 
96
105
 
97
106
  def _extract_picture_cdn_urls(script_text: str) -> list[str]:
98
- """从 picture_page_info_list 中只提取正文图片的 cdn_url,排除 watermark_info 和 share_cover 中的。"""
107
+ """读取图片数组条目的直接 cdn_url,排除数组外和嵌套对象中的地址。
108
+
109
+ 支持 window.picture_page_info_list = [...] 和 cgiDataNew 中的同名属性。
110
+ 页面数据是含表达式的 JS 对象字面量,因此只跟踪结构,不执行脚本。
111
+ """
112
+ if "picture_page_info_list" not in script_text:
113
+ return []
114
+
99
115
  urls: list[str] = []
100
116
  seen: set[str] = set()
101
- for m in re.finditer(r"(watermark_info|share_cover)?\s*(?::\s*\{[^}]*?)?\bcdn_url:\s*'([^']*)'", script_text):
102
- prefix = m.group(1)
103
- url = m.group(2)
104
- if prefix or not url:
117
+ stack: list[str] = []
118
+ array_urls: list[str] = []
119
+ previous = before_previous = ""
120
+ for match in _JS_DATA_TOKEN_RE.finditer(script_text):
121
+ if match.lastgroup == "comment":
105
122
  continue
106
- normalized = _normalize_image_url(url)
107
- if normalized not in seen:
108
- seen.add(normalized)
109
- urls.append(normalized)
123
+ token = match.group()
124
+ is_string = match.lastgroup == "string"
125
+
126
+ if not stack:
127
+ if (
128
+ token == "["
129
+ and previous in (":", "=")
130
+ and before_previous == "picture_page_info_list"
131
+ ):
132
+ stack.append(token)
133
+ elif token in ("[", "{"):
134
+ stack.append(token)
135
+ elif token in ("]", "}"):
136
+ if stack[-1] != {"]": "[", "}": "{"}[token]:
137
+ stack.clear()
138
+ array_urls.clear()
139
+ else:
140
+ stack.pop()
141
+ if not stack:
142
+ # 只提交完整数组,保留图片顺序并去重。
143
+ for url in array_urls:
144
+ if url not in seen:
145
+ seen.add(url)
146
+ urls.append(url)
147
+ array_urls.clear()
148
+ elif (
149
+ stack == ["[", "{"]
150
+ and before_previous == "cdn_url"
151
+ and previous == ":"
152
+ and is_string
153
+ and token[0] in ("'", '"')
154
+ ):
155
+ url = _decode_text(token[1:-1])
156
+ if url:
157
+ array_urls.append(_normalize_image_url(url))
158
+
159
+ before_previous, previous = previous, token[1:-1] if is_string else token
110
160
  return urls
111
161
 
112
162
 
@@ -138,6 +188,38 @@ def _extract_meta(soup: BeautifulSoup, result: ArticleResult) -> None:
138
188
  setattr(result, attr, value)
139
189
 
140
190
 
191
+ def _extract_item_show_type(scripts: list[Tag]) -> int | None:
192
+ """只读取当前文章的全局类型赋值;0 是有效类型,缺失时返回 None。"""
193
+ for script in scripts:
194
+ if "item_show_type" not in script.text:
195
+ continue
196
+ history: deque[str] = deque(maxlen=5)
197
+ brace_depth = 0
198
+ for match in _JS_DATA_TOKEN_RE.finditer(script.text):
199
+ if match.lastgroup == "comment":
200
+ continue
201
+ token = match.group()
202
+ previous = tuple(history)
203
+ is_global_var = (
204
+ brace_depth == 0
205
+ and previous[-3:] == ("var", "item_show_type", "=")
206
+ )
207
+ is_window_assignment = (
208
+ previous[-4:] == ("window", ".", "item_show_type", "=")
209
+ and (len(previous) < 5 or previous[-5] != ".")
210
+ )
211
+ if is_global_var or is_window_assignment:
212
+ value = token[1:-1] if token.startswith(("'", '"')) else token
213
+ if re.fullmatch(r"[0-9]+", value):
214
+ return int(value)
215
+ if token == "{":
216
+ brace_depth += 1
217
+ elif token == "}":
218
+ brace_depth -= 1
219
+ history.append(token)
220
+ return None
221
+
222
+
141
223
  def _extract_account_type(script_text: str, result: ArticleResult) -> None:
142
224
  """从 new_service_type 中提取账号类型:0/1 → 订阅号,2 → 服务号。"""
143
225
  if result.mp_account_type:
@@ -345,25 +427,30 @@ def _extract_plain_text_content(soup: BeautifulSoup, result: ArticleResult) -> N
345
427
 
346
428
  def _extract_swiper_content(soup: BeautifulSoup, result: ArticleResult) -> None:
347
429
  """从小红书风格的图片轮播文章中提取内容。"""
430
+ images: list[str] = []
431
+ seen: set[str] = set()
432
+ description = ""
348
433
  for script in soup.find_all("script", attrs={"type": "text/javascript"}):
349
- if "window.picture_page_info_list =" not in script.text:
350
- continue
351
-
352
- result.images = _extract_picture_cdn_urls(script.text)
353
-
354
- html_parts = [f'<img src="{img}" /><br>' for img in result.images]
434
+ # 图片数据和 window.desc 可能位于不同的 script,分别收集。
435
+ for url in _extract_picture_cdn_urls(script.text):
436
+ if url not in seen:
437
+ seen.add(url)
438
+ images.append(url)
355
439
 
356
440
  m = re.search(r'window.desc = "([^"]+)"', script.text)
357
441
  if m:
358
- text = _decode_text(m.group(1), preserve_newlines=True)
359
- html_parts.append(f"<p>{text}</p>")
442
+ description = _decode_text(m.group(1), preserve_newlines=True)
360
443
 
361
- if html_parts:
362
- result.article_markdown = _to_markdown_plain("".join(html_parts))
444
+ result.images = images
445
+ html_parts = [f'<img src="{img}" /><br>' for img in images]
446
+ if description:
447
+ html_parts.append(f"<p>{description}</p>")
448
+ if html_parts:
449
+ result.article_markdown = _to_markdown_plain("".join(html_parts))
363
450
 
364
451
 
365
452
  def _extract_fullscreen_content(soup: BeautifulSoup, result: ArticleResult) -> None:
366
- """从全屏布局文章中提取内容(appmsg_type 10002)。
453
+ """从全屏布局文字消息中提取内容。
367
454
 
368
455
  此类文章的文本存储在 text_page_info.content 中,通过 JsDecode() 编码;
369
456
  图片存储在 picture_page_info_list 的 cdn_url 字段中。
@@ -402,6 +489,67 @@ def _extract_fullscreen_content(soup: BeautifulSoup, result: ArticleResult) -> N
402
489
  # 主解析流程
403
490
  # ---------------------------------------------------------------------------
404
491
 
492
+ # 顺序与原有 HTML 判断流程一致,用于未知类型及提取失败时的兜底。
493
+ _CONTENT_SELECTORS = {
494
+ "rich_text": "div.rich_media_content",
495
+ "repost": "div.original_page",
496
+ "plain_text": "p#js_text_desc",
497
+ "video": "div#js_common_share_desc_wrap",
498
+ "swiper": "div.share_media_swiper_content",
499
+ "fullscreen": "div#js_fullscreen_layout_padding",
500
+ }
501
+ _ITEM_SHOW_TYPE_BRANCHES = {
502
+ 0: ("rich_text", "repost"),
503
+ 5: ("video",),
504
+ 8: ("swiper",),
505
+ 10: ("plain_text", "fullscreen"),
506
+ }
507
+ _ARTICLE_DISPLAY_TYPES = {
508
+ 0: ArticleDisplayType.ARTICLE,
509
+ 5: ArticleDisplayType.VIDEO,
510
+ 8: ArticleDisplayType.IMAGE,
511
+ 10: ArticleDisplayType.TEXT,
512
+ }
513
+
514
+
515
+ def _try_content_branch(
516
+ branch: str,
517
+ soup: BeautifulSoup,
518
+ scripts: list[Tag],
519
+ result: ArticleResult,
520
+ *,
521
+ require_html: bool,
522
+ ) -> ArticleResult | None:
523
+ content = soup.select_one(_CONTENT_SELECTORS[branch])
524
+ rich_text_template = branch in ("rich_text", "repost")
525
+ if content is None and (require_html or rich_text_template):
526
+ return None
527
+
528
+ # 失败分支的图片、元数据和标题处理不能污染后续分支。
529
+ candidate = replace(result, images=result.images.copy())
530
+ extract_meta = _extract_rich_text_meta if rich_text_template else _extract_swiper_meta
531
+ for script in scripts:
532
+ extract_meta(script.text, candidate)
533
+
534
+ if branch == "rich_text":
535
+ _extract_rich_media_content(content, candidate)
536
+ elif branch == "repost":
537
+ _extract_repost_content(content, candidate)
538
+ elif branch in ("plain_text", "video"):
539
+ _extract_plain_text_content(soup, candidate)
540
+ elif branch == "swiper":
541
+ _extract_swiper_content(soup, candidate)
542
+ elif branch == "fullscreen":
543
+ _extract_fullscreen_content(soup, candidate)
544
+
545
+ if not candidate.article_markdown.strip():
546
+ return None
547
+ if branch in ("plain_text", "fullscreen") and len(candidate.article_title) > 50:
548
+ short = candidate.article_title.split("。")[0]
549
+ candidate.article_title = short if len(short) <= 50 else candidate.article_title[:30]
550
+ return candidate
551
+
552
+
405
553
  def _parse_html(url: str, html: str) -> ArticleResult:
406
554
  """将原始 HTML 解析为 ArticleResult。"""
407
555
  # 检测微信验证码/人机验证页面
@@ -420,59 +568,26 @@ def _parse_html(url: str, html: str) -> ArticleResult:
420
568
 
421
569
  scripts = soup.find_all("script", attrs={"type": "text/javascript"})
422
570
 
423
- # 尝试解析标准富文本文章
424
- content = soup.find("div", class_="rich_media_content")
425
- if content:
426
- for s in scripts:
427
- _extract_rich_text_meta(s.text, result)
428
- _extract_rich_media_content(content, result)
429
- return result
430
-
431
- # 尝试解析转载式文章
432
- content = soup.find("div", class_="original_page")
433
- if content:
434
- for s in scripts:
435
- _extract_rich_text_meta(s.text, result)
436
- _extract_repost_content(content, result)
437
- return result
438
-
439
- # 尝试解析纯文本文章
440
- content = soup.find("p", id="js_text_desc")
441
- if content:
442
- for s in scripts:
443
- _extract_swiper_meta(s.text, result)
444
- _extract_plain_text_content(soup, result)
445
- if result.article_title and len(result.article_title) > 50:
446
- short = result.article_title.split("。")[0]
447
- result.article_title = short if len(short) <= 50 else result.article_title[:30]
448
- return result
449
-
450
- # 尝试解析视频分享类文章
451
- content = soup.find("div", id="js_common_share_desc_wrap")
452
- if content:
453
- for s in scripts:
454
- _extract_swiper_meta(s.text, result)
455
- _extract_plain_text_content(soup, result)
456
- return result
457
-
458
- # 尝试解析小红书风格图片轮播文章
459
- content = soup.find("div", class_="share_media_swiper_content")
460
- if content:
461
- for s in scripts:
462
- _extract_swiper_meta(s.text, result)
463
- _extract_swiper_content(soup, result)
464
- return result
465
-
466
- # 尝试解析全屏布局文章(如短文本帖子,appmsg_type 10002)
467
- content = soup.find("div", id="js_fullscreen_layout_padding")
468
- if content:
469
- for s in scripts:
470
- _extract_swiper_meta(s.text, result)
471
- _extract_fullscreen_content(soup, result)
472
- if result.article_title and len(result.article_title) > 50:
473
- short = result.article_title.split("。")[0]
474
- result.article_title = short if len(short) <= 50 else result.article_title[:30]
475
- return result
571
+ item_show_type = _extract_item_show_type(scripts)
572
+ result.article_display_type = _ARTICLE_DISPLAY_TYPES.get(
573
+ item_show_type, ArticleDisplayType.UNKNOWN,
574
+ )
575
+ preferred = _ITEM_SHOW_TYPE_BRANCHES.get(item_show_type, ())
576
+ attempted: set[str] = set()
577
+ for branch in (*preferred, *_CONTENT_SELECTORS):
578
+ if branch in attempted:
579
+ continue
580
+ attempted.add(branch)
581
+ # 已知类型先用对应的 DOM 或脚本数据细分模板;失败后按原有 HTML 顺序兜底。
582
+ candidate = _try_content_branch(
583
+ branch,
584
+ soup,
585
+ scripts,
586
+ result,
587
+ require_html=branch not in preferred,
588
+ )
589
+ if candidate is not None:
590
+ return candidate
476
591
 
477
592
  # 兜底:尽可能提取元数据
478
593
  for s in scripts:
@@ -0,0 +1,89 @@
1
+ """按展示类型检查解析结果的有效性。"""
2
+
3
+ from dataclasses import asdict, replace
4
+ import json
5
+
6
+ import pytest
7
+
8
+ from wechat_article_parser import ArticleDisplayType, ArticleResult
9
+
10
+
11
+ _COMPLETE = ArticleResult(
12
+ mp_id=123,
13
+ mp_name="测试公众号",
14
+ article_id="ncPoZnxKdvDRLM_s4DBgmA",
15
+ article_msg_id=456,
16
+ article_idx=1,
17
+ article_sn="signature",
18
+ article_title="测试文章",
19
+ article_publish_time=1788860293,
20
+ )
21
+
22
+
23
+ @pytest.mark.parametrize("display_type", list(ArticleDisplayType))
24
+ @pytest.mark.parametrize("markdown", ["", "正文"])
25
+ def test_only_video_can_be_valid_without_markdown(
26
+ display_type: ArticleDisplayType, markdown: str,
27
+ ) -> None:
28
+ result = replace(
29
+ _COMPLETE, article_display_type=display_type, article_markdown=markdown,
30
+ )
31
+
32
+ assert result.is_valid is (bool(markdown) or display_type == ArticleDisplayType.VIDEO)
33
+
34
+
35
+ @pytest.mark.parametrize(
36
+ ("field_name", "empty_value"),
37
+ [
38
+ ("mp_id", 0),
39
+ ("mp_name", ""),
40
+ ("article_id", ""),
41
+ ("article_msg_id", 0),
42
+ ("article_idx", 0),
43
+ ("article_sn", ""),
44
+ ("article_title", ""),
45
+ ("article_publish_time", 0),
46
+ ],
47
+ )
48
+ @pytest.mark.parametrize("markdown", ["", "视频附言"])
49
+ def test_video_still_requires_all_key_metadata(
50
+ field_name: str, empty_value: str | int, markdown: str,
51
+ ) -> None:
52
+ result = replace(
53
+ _COMPLETE,
54
+ article_display_type=ArticleDisplayType.VIDEO,
55
+ article_markdown=markdown,
56
+ **{field_name: empty_value},
57
+ )
58
+
59
+ assert not result.is_valid
60
+
61
+
62
+ @pytest.mark.parametrize(
63
+ "expected",
64
+ [
65
+ "article",
66
+ "video",
67
+ "image",
68
+ "text",
69
+ "unknown",
70
+ ],
71
+ )
72
+ def test_display_type_has_semantic_string_values(expected: str) -> None:
73
+ result = ArticleResult(article_display_type=ArticleDisplayType(expected))
74
+
75
+ assert result.article_display_type is ArticleDisplayType(expected)
76
+ assert result.article_display_type == expected
77
+ assert str(result.article_display_type) == expected
78
+ assert f"{result.article_display_type:>10}" == f"{expected:>10}"
79
+ assert json.dumps(result.article_display_type) == json.dumps(expected)
80
+ serialized = json.loads(json.dumps(asdict(result)))
81
+ assert serialized["article_display_type"] == expected
82
+ assert "item_show_type" not in serialized
83
+ assert not hasattr(result, "item_show_type")
84
+
85
+
86
+ def test_default_display_type_is_unknown() -> None:
87
+ result = ArticleResult()
88
+ assert result.article_display_type is ArticleDisplayType.UNKNOWN
89
+ assert not result.is_valid
@@ -15,6 +15,7 @@ TEST_URLS = [
15
15
  "https://mp.weixin.qq.com/s/h8E6riExCaH2Znmnprj-WQ",
16
16
  "https://mp.weixin.qq.com/s/ySQdtsRlRmAl_skdc5HQ-A",
17
17
  "https://mp.weixin.qq.com/s/MnkArbYQNp3tF29gujMUnQ",
18
+ "https://mp.weixin.qq.com/s/ncPoZnxKdvDRLM_s4DBgmA",
18
19
  ]
19
20
 
20
21
 
@@ -62,6 +63,7 @@ def _assert_result(result: ArticleResult, url: str) -> None:
62
63
  print(f"文章idx: {result.article_idx}")
63
64
  print(f"文章签名: {result.article_sn}")
64
65
  print(f"文章标题: {result.article_title}")
66
+ print(f"展示类型: {result.article_display_type}")
65
67
  print(
66
68
  f"封面图: {result.article_cover_image[:80]}..."
67
69
  if result.article_cover_image
@@ -85,7 +87,8 @@ def _assert_result(result: ArticleResult, url: str) -> None:
85
87
  assert result.article_idx > 0, "article_idx should be positive"
86
88
  assert result.article_sn, "article_sn should not be empty"
87
89
  assert result.article_title, "article_title should not be empty"
88
- assert result.article_markdown, "article_markdown should not be empty"
90
+ if result.article_display_type != "video":
91
+ assert result.article_markdown, "article_markdown should not be empty for non-video articles"
89
92
  assert result.article_publish_time > 0, "article_publish_time should be positive"
90
93
  assert result.is_valid
91
94
 
@@ -110,6 +113,7 @@ def test_fetch_all(url: str, proxy: str | None) -> None:
110
113
  print(f"文章idx: {result.article_idx}")
111
114
  print(f"文章签名: {result.article_sn}")
112
115
  print(f"文章标题: {result.article_title}")
116
+ print(f"展示类型: {result.article_display_type}")
113
117
  print(f"封面图: {result.article_cover_image}")
114
118
  print(f"文章摘要: {result.article_description}")
115
119
  print(f"发布时间: {result.article_publish_time}")
@@ -0,0 +1,150 @@
1
+ """使用精简页面离线回归轮播图片的提取和 Markdown 输出。"""
2
+
3
+ import pytest
4
+
5
+ from wechat_article_parser.parser import _parse_html
6
+
7
+
8
+ _URL = "https://mp.weixin.qq.com/s/0Wz3JeMbtWBL5iWJgYPS_Q"
9
+ _EXPECTED_IMAGES = [
10
+ "https://mmbiz.qpic.cn/mmbiz_png/second/640",
11
+ "https://mmbiz.qpic.cn/mmbiz_png/first/640",
12
+ ]
13
+ _PICTURES = r"""[
14
+ {
15
+ note: '含括号 ] } 和转义引号 \' 的字符串',
16
+ watermark_info: {
17
+ position: {x: 1},
18
+ cdn_url: 'https://mmbiz.qpic.cn/mmbiz_png/watermark/0'
19
+ },
20
+ cdn_url: 'https://mmbiz.qpic.cn/mmbiz_png/second/0?wx_fmt=png',
21
+ share_cover: {
22
+ crop_info: {x: 0},
23
+ cdn_url: 'https://mmbiz.qpic.cn/mmbiz_png/share-cover/0'
24
+ },
25
+ live_photo: {format_info: [{cdn_url: 'https://example.com/live-photo'}]}
26
+ },
27
+ {"cdn_url": "https://mmbiz.qpic.cn/mmbiz_png/first/0?wx_fmt=png"},
28
+ {cdn_url: 'https://mmbiz.qpic.cn/mmbiz_png/second/640'},
29
+ {cdn_url: ''}
30
+ ]"""
31
+ _DESCRIPTION = r'window.desc = "第一行\x0a第二行";'
32
+ _REFERENCE = """
33
+ window.picture_page_info_list =
34
+ ((window.cgiDataNew && window.cgiDataNew.picture_page_info_list) || []).slice(0, 20);
35
+ """
36
+
37
+
38
+ def _html(*scripts: str, fullscreen: bool = False) -> str:
39
+ content = (
40
+ '<div id="js_fullscreen_layout_padding"></div>'
41
+ if fullscreen
42
+ else '<div class="share_media_swiper_content"></div>'
43
+ )
44
+ return content + "".join(
45
+ f'<script type="text/javascript">{script}</script>' for script in scripts
46
+ )
47
+
48
+
49
+ @pytest.mark.parametrize("layout", ["inline", "data_first", "description_first"])
50
+ def test_swiper_collects_body_images_across_scripts(layout: str) -> None:
51
+ data = (
52
+ "window.cgiDataNew = {"
53
+ "cdn_url: 'https://mmbiz.qpic.cn/mmbiz_png/outside-before/0',"
54
+ f"picture_page_info_list: {_PICTURES},"
55
+ "other: {cdn_url: 'https://mmbiz.qpic.cn/mmbiz_png/outside-after/0'}"
56
+ "};"
57
+ )
58
+ reference = _REFERENCE + _DESCRIPTION
59
+ if layout == "inline":
60
+ scripts = [f"window.picture_page_info_list = {_PICTURES};" + _DESCRIPTION]
61
+ elif layout == "data_first":
62
+ scripts = [data, reference]
63
+ else:
64
+ scripts = [reference, data]
65
+
66
+ result = _parse_html(_URL, _html(*scripts))
67
+
68
+ assert result.images == _EXPECTED_IMAGES
69
+ assert result.article_markdown.count("![") == 2
70
+ assert result.article_markdown.index(_EXPECTED_IMAGES[0]) < result.article_markdown.index(_EXPECTED_IMAGES[1])
71
+ assert "第一行 \n第二行" in result.article_markdown
72
+ assert "outside" not in result.article_markdown
73
+ assert "watermark" not in result.article_markdown
74
+ assert "share-cover" not in result.article_markdown
75
+ assert "live-photo" not in result.article_markdown
76
+
77
+
78
+ def test_swiper_later_empty_lists_do_not_erase_images() -> None:
79
+ result = _parse_html(
80
+ _URL,
81
+ _html(
82
+ f"window.picture_page_info_list = {_PICTURES};" + _DESCRIPTION,
83
+ "window.picture_page_info_list = [];",
84
+ _REFERENCE,
85
+ ),
86
+ )
87
+
88
+ assert result.images == _EXPECTED_IMAGES
89
+ assert result.article_markdown.count("![") == 2
90
+ assert "第一行 \n第二行" in result.article_markdown
91
+
92
+
93
+ def test_swiper_ignores_list_text_in_strings_and_comments() -> None:
94
+ decoy = "picture_page_info_list: [{cdn_url: 'https://example.com/decoy'}]"
95
+ result = _parse_html(
96
+ _URL,
97
+ _html(
98
+ f'var example = "{decoy}";\n'
99
+ f"// {decoy}\n"
100
+ f"/* {decoy} */\n"
101
+ f"var template = `{decoy}`;\n",
102
+ f"window.cgiDataNew = {{'picture_page_info_list': {_PICTURES}}};",
103
+ _REFERENCE + _DESCRIPTION,
104
+ ),
105
+ )
106
+
107
+ assert result.images == _EXPECTED_IMAGES
108
+ assert "decoy" not in result.article_markdown
109
+
110
+
111
+ def test_swiper_image_only_content() -> None:
112
+ result = _parse_html(
113
+ _URL,
114
+ _html(f"window.cgiDataNew = {{picture_page_info_list: {_PICTURES}}};", _REFERENCE),
115
+ )
116
+
117
+ assert result.images == _EXPECTED_IMAGES
118
+ assert result.article_markdown.count("![") == 2
119
+
120
+
121
+ def test_swiper_text_only_content() -> None:
122
+ result = _parse_html(
123
+ _URL,
124
+ _html(
125
+ "window.cgiDataNew = {picture_page_info_list: [], cdn_url: 'https://example.com/cover'};",
126
+ _REFERENCE,
127
+ _DESCRIPTION,
128
+ ),
129
+ )
130
+
131
+ assert result.images == []
132
+ assert result.article_markdown == "第一行 \n第二行"
133
+
134
+
135
+ def test_fullscreen_shared_image_extractor_keeps_body_images() -> None:
136
+ result = _parse_html(
137
+ _URL,
138
+ _html(
139
+ "window.cgiDataNew = {"
140
+ "cdn_url: 'https://example.com/cover',"
141
+ f"picture_page_info_list: {_PICTURES},"
142
+ "text_page_info: {content: JsDecode('全屏正文')}"
143
+ "};",
144
+ fullscreen=True,
145
+ ),
146
+ )
147
+
148
+ assert result.images == _EXPECTED_IMAGES
149
+ assert result.article_markdown.count("![") == 2
150
+ assert "全屏正文" in result.article_markdown
@@ -0,0 +1,241 @@
1
+ """类型字段优先分发、模板细分及 HTML 兜底的离线回归测试。"""
2
+
3
+ import pytest
4
+ from bs4 import BeautifulSoup
5
+
6
+ from wechat_article_parser import ArticleDisplayType, parser
7
+
8
+
9
+ _URL = "https://mp.weixin.qq.com/s/0Wz3JeMbtWBL5iWJgYPS_Q"
10
+ _RICH_TEXT = '<div class="rich_media_content"><p>富文本正文</p></div>'
11
+ _REPOST = """
12
+ <div class="original_page">
13
+ <p id="js_share_notice"><script>notice.innerHTML = "转载附言";</script></p>
14
+ <span id="js_share_source" data-url="https://example.com/original"></span>
15
+ </div>
16
+ """
17
+ _PLAIN_DATA = """
18
+ var TextContentNoEncode = '';
19
+ var ContentNoEncode = window.a_value_which_never_exists || '文字正文';
20
+ """
21
+ _SWIPER_DATA = """
22
+ window.cgiDataNew = {
23
+ picture_page_info_list: [{cdn_url: 'https://mmbiz.qpic.cn/mmbiz_png/body/0'}]
24
+ };
25
+ window.desc = "轮播正文";
26
+ """
27
+ _FULLSCREEN_DATA = """
28
+ window.cgiDataNew = {
29
+ picture_page_info_list: [],
30
+ text_page_info: {content: JsDecode('全屏正文')}
31
+ };
32
+ """
33
+
34
+
35
+ def _html(body: str = "", script: str = "", title: str = "原标题") -> str:
36
+ return (
37
+ f'<meta property="og:title" content="{title}">'
38
+ f"{body}<script type=\"text/javascript\">{script}</script>"
39
+ )
40
+
41
+
42
+ @pytest.mark.parametrize(
43
+ ("script", "expected"),
44
+ [
45
+ ('var item_show_type = "0";', 0),
46
+ ("window.item_show_type = '8' || '';", 8),
47
+ ("window.item_show_type = 10;", 10),
48
+ ("var /* 注释 */ item_show_type = 5;", 5),
49
+ ("(function() { window.item_show_type = '10'; })();", 10),
50
+ ("window.item_show_type = '999';", 999),
51
+ ("window.item_show_type = '';", None),
52
+ ("window.item_show_type = 'unknown';", None),
53
+ ("window.item_show_type = null;", None),
54
+ ("window.item_show_type = -1;", None),
55
+ ("window.item_show_type = 8.5;", None),
56
+ ("window.item_show_type = 8e2;", None),
57
+ ("window.item_show_type == 8;", None),
58
+ ("var real_item_show_type = '8';", None),
59
+ ("", None),
60
+ ],
61
+ )
62
+ def test_read_current_article_type(script: str, expected: int | None) -> None:
63
+ soup = BeautifulSoup(_html(script=script), "html.parser")
64
+
65
+ assert parser._extract_item_show_type(soup.find_all("script")) == expected
66
+
67
+
68
+ @pytest.mark.parametrize("current_assignment", ["", "window.item_show_type = '8';"])
69
+ def test_ignore_type_in_references_strings_comments_and_local_variables(
70
+ current_assignment: str,
71
+ ) -> None:
72
+ decoys = """
73
+ var reference = {item_show_type: 5};
74
+ var example = "var item_show_type = '5';";
75
+ var url = 'https://example.com/?item_show_type=5';
76
+ var template = `
77
+ window.item_show_type = '5';
78
+ `;
79
+ // window.item_show_type = '5';
80
+ /* var item_show_type = '5'; */
81
+ function report() { var item_show_type = '5'; }
82
+ report.window.item_show_type = '5';
83
+ window.real_item_show_type = '5';
84
+ """
85
+ soup = BeautifulSoup(_html(script=decoys + current_assignment), "html.parser")
86
+
87
+ expected = 8 if current_assignment else None
88
+ assert parser._extract_item_show_type(soup.find_all("script")) == expected
89
+
90
+
91
+ @pytest.mark.parametrize(
92
+ ("body", "expected"),
93
+ [
94
+ (_RICH_TEXT + _REPOST, "富文本正文"),
95
+ (_REPOST, "转载附言"),
96
+ ('<div class="rich_media_content"></div>' + _REPOST, "转载附言"),
97
+ ],
98
+ )
99
+ def test_type_zero_uses_html_to_distinguish_rich_text_and_repost(
100
+ body: str, expected: str,
101
+ ) -> None:
102
+ result = parser._parse_html(_URL, _html(body, 'var item_show_type = "0";'))
103
+
104
+ assert result.article_display_type is ArticleDisplayType.ARTICLE
105
+ assert expected in result.article_markdown
106
+ if expected == "转载附言":
107
+ assert "[查看原文](https://example.com/original)" in result.article_markdown
108
+ else:
109
+ assert "转载附言" not in result.article_markdown
110
+
111
+
112
+ @pytest.mark.parametrize(
113
+ ("item_type", "data", "expected", "display_type"),
114
+ [
115
+ (5, _PLAIN_DATA, "文字正文", ArticleDisplayType.VIDEO),
116
+ (8, _SWIPER_DATA, "轮播正文", ArticleDisplayType.IMAGE),
117
+ (10, _PLAIN_DATA, "文字正文", ArticleDisplayType.TEXT),
118
+ (10, _FULLSCREEN_DATA, "全屏正文", ArticleDisplayType.TEXT),
119
+ ],
120
+ )
121
+ def test_known_type_takes_priority_over_other_html_templates(
122
+ item_type: int, data: str, expected: str, display_type: ArticleDisplayType,
123
+ ) -> None:
124
+ # 即使存在富文本壳,已知类型仍先读取对应脚本数据;无需专用容器。
125
+ result = parser._parse_html(
126
+ _URL,
127
+ _html(_RICH_TEXT, f"window.item_show_type = '{item_type}';" + data),
128
+ )
129
+
130
+ assert expected in result.article_markdown
131
+ assert "富文本正文" not in result.article_markdown
132
+ assert result.article_display_type is display_type
133
+ assert not hasattr(result, "item_show_type")
134
+ if item_type == 8:
135
+ assert result.images == ["https://mmbiz.qpic.cn/mmbiz_png/body/640"]
136
+
137
+
138
+ @pytest.mark.parametrize(
139
+ "declaration",
140
+ ["", "window.item_show_type = '999';", "window.item_show_type = 'invalid';"],
141
+ )
142
+ def test_missing_unknown_or_invalid_type_falls_back_to_html(declaration: str) -> None:
143
+ result = parser._parse_html(_URL, _html(_RICH_TEXT, declaration))
144
+
145
+ assert result.article_markdown == "富文本正文"
146
+ assert result.article_display_type is ArticleDisplayType.UNKNOWN
147
+
148
+
149
+ @pytest.mark.parametrize("item_type", [5, 8, 10])
150
+ def test_known_type_without_matching_data_falls_back_to_html(item_type: int) -> None:
151
+ result = parser._parse_html(
152
+ _URL,
153
+ _html(_RICH_TEXT, f"window.item_show_type = '{item_type}';"),
154
+ )
155
+
156
+ assert result.article_markdown == "富文本正文"
157
+
158
+
159
+ def test_type_zero_without_matching_html_falls_back_to_other_templates() -> None:
160
+ result = parser._parse_html(
161
+ _URL,
162
+ _html(
163
+ '<div class="share_media_swiper_content"></div>',
164
+ "window.item_show_type = '0';" + _SWIPER_DATA,
165
+ ),
166
+ )
167
+
168
+ assert "轮播正文" in result.article_markdown
169
+ assert len(result.images) == 1
170
+
171
+
172
+ def test_failed_branch_does_not_pollute_fallback_result(monkeypatch) -> None:
173
+ def fail_swiper(soup, result):
174
+ result.article_title = "错误标题"
175
+ result.mp_name = "错误账号"
176
+ result.images.append("https://example.com/wrong-image")
177
+ result.article_markdown = " \n"
178
+
179
+ monkeypatch.setattr(parser, "_extract_swiper_content", fail_swiper)
180
+ result = parser._parse_html(
181
+ _URL,
182
+ _html(_RICH_TEXT, "window.item_show_type = '8';"),
183
+ )
184
+
185
+ assert result.article_title == "原标题"
186
+ assert result.mp_name == ""
187
+ assert result.images == []
188
+ assert result.article_markdown == "富文本正文"
189
+
190
+
191
+ def test_unrecognized_content_still_returns_metadata() -> None:
192
+ result = parser._parse_html(
193
+ _URL,
194
+ _html(script='window.item_show_type = "999"; window.alias = "account";'),
195
+ )
196
+
197
+ assert result.article_title == "原标题"
198
+ assert result.mp_alias == "account"
199
+ assert result.article_display_type is ArticleDisplayType.UNKNOWN
200
+ assert result.article_markdown == ""
201
+ assert not result.is_valid
202
+
203
+
204
+ @pytest.mark.parametrize(
205
+ ("declaration", "expected_type"),
206
+ [
207
+ ("", ArticleDisplayType.UNKNOWN),
208
+ ("window.item_show_type = '5';", ArticleDisplayType.VIDEO),
209
+ ],
210
+ )
211
+ def test_video_without_description_keeps_type_and_metadata(
212
+ declaration: str, expected_type: ArticleDisplayType,
213
+ ) -> None:
214
+ metadata = """
215
+ window.__initCgiDataConfig = function(d) {
216
+ return {
217
+ biz: d.biz ? d.biz : 'MQ==',
218
+ nick_name: d.nick_name ? d.nick_name : '测试公众号',
219
+ mid: d.mid ? d.mid : '123',
220
+ idx: d.idx ? d.idx : '1',
221
+ sn: d.sn ? d.sn : 'signature',
222
+ create_time: d.create_time ? d.create_time : '1788860293'
223
+ };
224
+ };
225
+ var TextContentNoEncode = window.a_value_which_never_exists || '';
226
+ var ContentNoEncode = window.a_value_which_never_exists || '';
227
+ """
228
+ result = parser._parse_html(
229
+ _URL,
230
+ _html(
231
+ '<div id="js_common_share_desc_wrap" style="display: none"></div>'
232
+ '<div id="js_fullscreen_layout_padding"></div>',
233
+ declaration + metadata,
234
+ ),
235
+ )
236
+
237
+ assert result.article_display_type is expected_type
238
+ assert result.mp_id == 1
239
+ assert result.mp_name == "测试公众号"
240
+ assert result.article_markdown == ""
241
+ assert result.is_valid is (expected_type == ArticleDisplayType.VIDEO)
@@ -1,4 +0,0 @@
1
- from .models import AccountType, ArticleResult, WeChatVerifyError
2
- from .parser import parse, parse_async
3
-
4
- __all__ = ["parse", "parse_async", "ArticleResult", "AccountType", "WeChatVerifyError"]