wechat-article-parser 0.0.6__tar.gz → 0.0.7__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- {wechat_article_parser-0.0.6 → wechat_article_parser-0.0.7}/PKG-INFO +38 -5
- {wechat_article_parser-0.0.6 → wechat_article_parser-0.0.7}/README.md +36 -3
- {wechat_article_parser-0.0.6 → wechat_article_parser-0.0.7}/pyproject.toml +1 -1
- wechat_article_parser-0.0.7/src/wechat_article_parser/__init__.py +11 -0
- {wechat_article_parser-0.0.6 → wechat_article_parser-0.0.7}/src/wechat_article_parser/models.py +24 -2
- {wechat_article_parser-0.0.6 → wechat_article_parser-0.0.7}/src/wechat_article_parser/parser.py +189 -74
- wechat_article_parser-0.0.7/tests/test_models.py +89 -0
- {wechat_article_parser-0.0.6 → wechat_article_parser-0.0.7}/tests/test_parser.py +5 -1
- wechat_article_parser-0.0.7/tests/test_swiper_content.py +150 -0
- wechat_article_parser-0.0.7/tests/test_type_routing.py +241 -0
- wechat_article_parser-0.0.6/src/wechat_article_parser/__init__.py +0 -4
- {wechat_article_parser-0.0.6 → wechat_article_parser-0.0.7}/.gitignore +0 -0
- {wechat_article_parser-0.0.6 → wechat_article_parser-0.0.7}/LICENSE +0 -0
- {wechat_article_parser-0.0.6 → wechat_article_parser-0.0.7}/tests/conftest.py +0 -0
|
@@ -1,6 +1,6 @@
|
|
|
1
|
-
Metadata-Version: 2.
|
|
1
|
+
Metadata-Version: 2.5
|
|
2
2
|
Name: wechat-article-parser
|
|
3
|
-
Version: 0.0.
|
|
3
|
+
Version: 0.0.7
|
|
4
4
|
Summary: WeChat MP article parser - extract metadata and content from WeChat public account articles
|
|
5
5
|
License-Expression: MIT
|
|
6
6
|
License-File: LICENSE
|
|
@@ -26,6 +26,10 @@ Description-Content-Type: text/markdown
|
|
|
26
26
|
- 小红书风格图片轮播文章
|
|
27
27
|
- 全屏布局短文本文章
|
|
28
28
|
|
|
29
|
+
解析时优先读取页面全局的 `item_show_type`:`0` 对应图文、`5` 对应视频分享、
|
|
30
|
+
`8` 对应图片轮播、`10` 对应文字消息,再通过 HTML 或脚本数据区分具体模板。
|
|
31
|
+
类型字段缺失、值未知或对应分支未提取到正文时,会按原有 HTML 特征匹配顺序兜底。
|
|
32
|
+
|
|
29
33
|
## 安装
|
|
30
34
|
|
|
31
35
|
```bash
|
|
@@ -108,13 +112,42 @@ result = parse(
|
|
|
108
112
|
| `article_title` | `str` | 文章标题 |
|
|
109
113
|
| `article_cover_image` | `str` | 文章封面图链接 |
|
|
110
114
|
| `article_description` | `str` | 文章摘要 |
|
|
111
|
-
| `article_markdown` | `str` | 文章正文的 Markdown
|
|
115
|
+
| `article_markdown` | `str` | 文章正文的 Markdown 内容;视频分享没有文字附言时可为空 |
|
|
112
116
|
| `article_publish_time` | `int` | 发布时间(Unix 时间戳) |
|
|
117
|
+
| `article_display_type` | `ArticleDisplayType` | 文章展示类型:`article`、`video`、`image`、`text` 或 `unknown` |
|
|
113
118
|
| `raw_html` | `str` | 原始 HTML 网页源代码(仅在 `include_raw_html=True` 时填充,否则为空字符串) |
|
|
114
119
|
| `images` | `list[str]` | 文章中提取的所有图片链接 |
|
|
115
120
|
| `is_valid` | `bool` | 关键字段是否全部解析成功(属性) |
|
|
116
121
|
|
|
117
|
-
`is_valid` 为 `True` 的条件:`mp_id`、`mp_name`、`article_id`、`article_msg_id`、`article_idx`、`article_sn`、`article_title`、`
|
|
122
|
+
`is_valid` 为 `True` 的条件:`mp_id`、`mp_name`、`article_id`、`article_msg_id`、`article_idx`、`article_sn`、`article_title`、`article_publish_time` 均不为空/零,且 `article_markdown` 非空或 `article_display_type == "video"`。视频分享允许没有文字正文,仍须满足其他关键字段要求;有效性不表示已获取视频文件。
|
|
123
|
+
|
|
124
|
+
### 判断文章展示类型
|
|
125
|
+
|
|
126
|
+
`ArticleDisplayType` 是字符串枚举,既支持枚举比较,也支持语义字符串比较:
|
|
127
|
+
|
|
128
|
+
| 枚举 | 字符串值 | 含义 |
|
|
129
|
+
|---|---|---|
|
|
130
|
+
| `ArticleDisplayType.ARTICLE` | `article` | 图文,包括转载式 |
|
|
131
|
+
| `ArticleDisplayType.VIDEO` | `video` | 视频分享 |
|
|
132
|
+
| `ArticleDisplayType.IMAGE` | `image` | 图片轮播 |
|
|
133
|
+
| `ArticleDisplayType.TEXT` | `text` | 文字消息 |
|
|
134
|
+
| `ArticleDisplayType.UNKNOWN` | `unknown` | 页面类型缺失或未支持 |
|
|
135
|
+
|
|
136
|
+
```python
|
|
137
|
+
from wechat_article_parser import parse, ArticleDisplayType
|
|
138
|
+
|
|
139
|
+
result = parse("https://mp.weixin.qq.com/s/xxxxx")
|
|
140
|
+
if result.article_display_type == ArticleDisplayType.VIDEO:
|
|
141
|
+
print("这是视频分享")
|
|
142
|
+
|
|
143
|
+
if result.article_display_type == "video":
|
|
144
|
+
print("视频分享可以没有文字正文")
|
|
145
|
+
|
|
146
|
+
print(result.article_display_type) # 例如 video
|
|
147
|
+
```
|
|
148
|
+
|
|
149
|
+
`article_display_type` 是返回结果中的数据字段,包含在 `dataclasses.asdict(result)` 中,JSON 序列化时输出语义字符串。
|
|
150
|
+
微信原始的 `item_show_type` 仅在解析器内部使用,不包含在返回结果中。
|
|
118
151
|
|
|
119
152
|
### 判断账号类型
|
|
120
153
|
|
|
@@ -168,7 +201,7 @@ except httpx.HTTPStatusError as e:
|
|
|
168
201
|
|
|
169
202
|
### 解析不完整
|
|
170
203
|
|
|
171
|
-
|
|
204
|
+
缺少标题、发布时间等关键字段,或非视频分享缺少正文时,解析结果的 `is_valid` 会返回 `False`。封面图、摘要等非必填字段为空不影响有效性。建议在业务逻辑中检查:
|
|
172
205
|
|
|
173
206
|
```python
|
|
174
207
|
result = parse("https://mp.weixin.qq.com/s/xxxxx")
|
|
@@ -11,6 +11,10 @@
|
|
|
11
11
|
- 小红书风格图片轮播文章
|
|
12
12
|
- 全屏布局短文本文章
|
|
13
13
|
|
|
14
|
+
解析时优先读取页面全局的 `item_show_type`:`0` 对应图文、`5` 对应视频分享、
|
|
15
|
+
`8` 对应图片轮播、`10` 对应文字消息,再通过 HTML 或脚本数据区分具体模板。
|
|
16
|
+
类型字段缺失、值未知或对应分支未提取到正文时,会按原有 HTML 特征匹配顺序兜底。
|
|
17
|
+
|
|
14
18
|
## 安装
|
|
15
19
|
|
|
16
20
|
```bash
|
|
@@ -93,13 +97,42 @@ result = parse(
|
|
|
93
97
|
| `article_title` | `str` | 文章标题 |
|
|
94
98
|
| `article_cover_image` | `str` | 文章封面图链接 |
|
|
95
99
|
| `article_description` | `str` | 文章摘要 |
|
|
96
|
-
| `article_markdown` | `str` | 文章正文的 Markdown
|
|
100
|
+
| `article_markdown` | `str` | 文章正文的 Markdown 内容;视频分享没有文字附言时可为空 |
|
|
97
101
|
| `article_publish_time` | `int` | 发布时间(Unix 时间戳) |
|
|
102
|
+
| `article_display_type` | `ArticleDisplayType` | 文章展示类型:`article`、`video`、`image`、`text` 或 `unknown` |
|
|
98
103
|
| `raw_html` | `str` | 原始 HTML 网页源代码(仅在 `include_raw_html=True` 时填充,否则为空字符串) |
|
|
99
104
|
| `images` | `list[str]` | 文章中提取的所有图片链接 |
|
|
100
105
|
| `is_valid` | `bool` | 关键字段是否全部解析成功(属性) |
|
|
101
106
|
|
|
102
|
-
`is_valid` 为 `True` 的条件:`mp_id`、`mp_name`、`article_id`、`article_msg_id`、`article_idx`、`article_sn`、`article_title`、`
|
|
107
|
+
`is_valid` 为 `True` 的条件:`mp_id`、`mp_name`、`article_id`、`article_msg_id`、`article_idx`、`article_sn`、`article_title`、`article_publish_time` 均不为空/零,且 `article_markdown` 非空或 `article_display_type == "video"`。视频分享允许没有文字正文,仍须满足其他关键字段要求;有效性不表示已获取视频文件。
|
|
108
|
+
|
|
109
|
+
### 判断文章展示类型
|
|
110
|
+
|
|
111
|
+
`ArticleDisplayType` 是字符串枚举,既支持枚举比较,也支持语义字符串比较:
|
|
112
|
+
|
|
113
|
+
| 枚举 | 字符串值 | 含义 |
|
|
114
|
+
|---|---|---|
|
|
115
|
+
| `ArticleDisplayType.ARTICLE` | `article` | 图文,包括转载式 |
|
|
116
|
+
| `ArticleDisplayType.VIDEO` | `video` | 视频分享 |
|
|
117
|
+
| `ArticleDisplayType.IMAGE` | `image` | 图片轮播 |
|
|
118
|
+
| `ArticleDisplayType.TEXT` | `text` | 文字消息 |
|
|
119
|
+
| `ArticleDisplayType.UNKNOWN` | `unknown` | 页面类型缺失或未支持 |
|
|
120
|
+
|
|
121
|
+
```python
|
|
122
|
+
from wechat_article_parser import parse, ArticleDisplayType
|
|
123
|
+
|
|
124
|
+
result = parse("https://mp.weixin.qq.com/s/xxxxx")
|
|
125
|
+
if result.article_display_type == ArticleDisplayType.VIDEO:
|
|
126
|
+
print("这是视频分享")
|
|
127
|
+
|
|
128
|
+
if result.article_display_type == "video":
|
|
129
|
+
print("视频分享可以没有文字正文")
|
|
130
|
+
|
|
131
|
+
print(result.article_display_type) # 例如 video
|
|
132
|
+
```
|
|
133
|
+
|
|
134
|
+
`article_display_type` 是返回结果中的数据字段,包含在 `dataclasses.asdict(result)` 中,JSON 序列化时输出语义字符串。
|
|
135
|
+
微信原始的 `item_show_type` 仅在解析器内部使用,不包含在返回结果中。
|
|
103
136
|
|
|
104
137
|
### 判断账号类型
|
|
105
138
|
|
|
@@ -153,7 +186,7 @@ except httpx.HTTPStatusError as e:
|
|
|
153
186
|
|
|
154
187
|
### 解析不完整
|
|
155
188
|
|
|
156
|
-
|
|
189
|
+
缺少标题、发布时间等关键字段,或非视频分享缺少正文时,解析结果的 `is_valid` 会返回 `False`。封面图、摘要等非必填字段为空不影响有效性。建议在业务逻辑中检查:
|
|
157
190
|
|
|
158
191
|
```python
|
|
159
192
|
result = parse("https://mp.weixin.qq.com/s/xxxxx")
|
|
@@ -4,7 +4,7 @@ build-backend = "hatchling.build"
|
|
|
4
4
|
|
|
5
5
|
[project]
|
|
6
6
|
name = "wechat-article-parser"
|
|
7
|
-
version = "0.0.
|
|
7
|
+
version = "0.0.7"
|
|
8
8
|
description = "WeChat MP article parser - extract metadata and content from WeChat public account articles"
|
|
9
9
|
readme = "README.md"
|
|
10
10
|
requires-python = ">=3.10"
|
|
@@ -0,0 +1,11 @@
|
|
|
1
|
+
from .models import AccountType, ArticleDisplayType, ArticleResult, WeChatVerifyError
|
|
2
|
+
from .parser import parse, parse_async
|
|
3
|
+
|
|
4
|
+
__all__ = [
|
|
5
|
+
"parse",
|
|
6
|
+
"parse_async",
|
|
7
|
+
"ArticleResult",
|
|
8
|
+
"AccountType",
|
|
9
|
+
"ArticleDisplayType",
|
|
10
|
+
"WeChatVerifyError",
|
|
11
|
+
]
|
{wechat_article_parser-0.0.6 → wechat_article_parser-0.0.7}/src/wechat_article_parser/models.py
RENAMED
|
@@ -22,6 +22,22 @@ class AccountType(str, Enum):
|
|
|
22
22
|
return format(self.value, format_spec)
|
|
23
23
|
|
|
24
24
|
|
|
25
|
+
class ArticleDisplayType(str, Enum):
|
|
26
|
+
"""文章展示类型,支持枚举和语义字符串比较。"""
|
|
27
|
+
|
|
28
|
+
UNKNOWN = "unknown"
|
|
29
|
+
ARTICLE = "article"
|
|
30
|
+
VIDEO = "video"
|
|
31
|
+
IMAGE = "image"
|
|
32
|
+
TEXT = "text"
|
|
33
|
+
|
|
34
|
+
def __str__(self) -> str:
|
|
35
|
+
return self.value
|
|
36
|
+
|
|
37
|
+
def __format__(self, format_spec: str) -> str:
|
|
38
|
+
return format(self.value, format_spec)
|
|
39
|
+
|
|
40
|
+
|
|
25
41
|
@dataclass
|
|
26
42
|
class ArticleResult:
|
|
27
43
|
"""微信公众号文章的解析结果。"""
|
|
@@ -52,9 +68,12 @@ class ArticleResult:
|
|
|
52
68
|
# 文章中提取的图片列表
|
|
53
69
|
images: list[str] = field(default_factory=list)
|
|
54
70
|
|
|
71
|
+
# 对外返回语义展示类型,微信原始数值仅在解析器内部使用。
|
|
72
|
+
article_display_type: ArticleDisplayType = ArticleDisplayType.UNKNOWN
|
|
73
|
+
|
|
55
74
|
@property
|
|
56
75
|
def is_valid(self) -> bool:
|
|
57
|
-
"""
|
|
76
|
+
"""检查关键字段是否解析成功;视频分享允许正文为空。"""
|
|
58
77
|
return bool(
|
|
59
78
|
self.mp_id
|
|
60
79
|
and self.mp_name
|
|
@@ -63,6 +82,9 @@ class ArticleResult:
|
|
|
63
82
|
and self.article_idx
|
|
64
83
|
and self.article_sn
|
|
65
84
|
and self.article_title
|
|
66
|
-
and
|
|
85
|
+
and (
|
|
86
|
+
self.article_markdown
|
|
87
|
+
or self.article_display_type == ArticleDisplayType.VIDEO
|
|
88
|
+
)
|
|
67
89
|
and self.article_publish_time
|
|
68
90
|
)
|
{wechat_article_parser-0.0.6 → wechat_article_parser-0.0.7}/src/wechat_article_parser/parser.py
RENAMED
|
@@ -5,13 +5,15 @@ from __future__ import annotations
|
|
|
5
5
|
import base64
|
|
6
6
|
import html as html_module
|
|
7
7
|
import re
|
|
8
|
+
from collections import deque
|
|
9
|
+
from dataclasses import replace
|
|
8
10
|
from urllib.parse import unquote
|
|
9
11
|
|
|
10
12
|
import httpx
|
|
11
13
|
from bs4 import BeautifulSoup, Tag
|
|
12
14
|
from markdownify import MarkdownConverter
|
|
13
15
|
|
|
14
|
-
from .models import AccountType, ArticleResult, WeChatVerifyError
|
|
16
|
+
from .models import AccountType, ArticleDisplayType, ArticleResult, WeChatVerifyError
|
|
15
17
|
|
|
16
18
|
_USER_AGENT = (
|
|
17
19
|
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) "
|
|
@@ -20,6 +22,13 @@ _USER_AGENT = (
|
|
|
20
22
|
)
|
|
21
23
|
_TIMEOUT = 15
|
|
22
24
|
|
|
25
|
+
# 将字符串和注释作为整体,避免其中的括号干扰图片数组的层级判断。
|
|
26
|
+
_JS_DATA_TOKEN_RE = re.compile(
|
|
27
|
+
r"(?P<string>'(?:\\[\s\S]|[^'\\])*'|\"(?:\\[\s\S]|[^\"\\])*\"|`(?:\\[\s\S]|[^`\\])*`)"
|
|
28
|
+
r"|(?P<comment>//[^\r\n]*|/\*[\s\S]*?\*/)"
|
|
29
|
+
r"|[a-zA-Z_$][\w$]*|\d+(?:\.\d*)?(?:[eE][+-]?\d+)?|[^\s]"
|
|
30
|
+
)
|
|
31
|
+
|
|
23
32
|
|
|
24
33
|
# ---------------------------------------------------------------------------
|
|
25
34
|
# Markdown 转换器
|
|
@@ -95,18 +104,59 @@ def _service_type_to_account_type(value: str) -> AccountType:
|
|
|
95
104
|
|
|
96
105
|
|
|
97
106
|
def _extract_picture_cdn_urls(script_text: str) -> list[str]:
|
|
98
|
-
"""
|
|
107
|
+
"""读取图片数组条目的直接 cdn_url,排除数组外和嵌套对象中的地址。
|
|
108
|
+
|
|
109
|
+
支持 window.picture_page_info_list = [...] 和 cgiDataNew 中的同名属性。
|
|
110
|
+
页面数据是含表达式的 JS 对象字面量,因此只跟踪结构,不执行脚本。
|
|
111
|
+
"""
|
|
112
|
+
if "picture_page_info_list" not in script_text:
|
|
113
|
+
return []
|
|
114
|
+
|
|
99
115
|
urls: list[str] = []
|
|
100
116
|
seen: set[str] = set()
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
117
|
+
stack: list[str] = []
|
|
118
|
+
array_urls: list[str] = []
|
|
119
|
+
previous = before_previous = ""
|
|
120
|
+
for match in _JS_DATA_TOKEN_RE.finditer(script_text):
|
|
121
|
+
if match.lastgroup == "comment":
|
|
105
122
|
continue
|
|
106
|
-
|
|
107
|
-
|
|
108
|
-
|
|
109
|
-
|
|
123
|
+
token = match.group()
|
|
124
|
+
is_string = match.lastgroup == "string"
|
|
125
|
+
|
|
126
|
+
if not stack:
|
|
127
|
+
if (
|
|
128
|
+
token == "["
|
|
129
|
+
and previous in (":", "=")
|
|
130
|
+
and before_previous == "picture_page_info_list"
|
|
131
|
+
):
|
|
132
|
+
stack.append(token)
|
|
133
|
+
elif token in ("[", "{"):
|
|
134
|
+
stack.append(token)
|
|
135
|
+
elif token in ("]", "}"):
|
|
136
|
+
if stack[-1] != {"]": "[", "}": "{"}[token]:
|
|
137
|
+
stack.clear()
|
|
138
|
+
array_urls.clear()
|
|
139
|
+
else:
|
|
140
|
+
stack.pop()
|
|
141
|
+
if not stack:
|
|
142
|
+
# 只提交完整数组,保留图片顺序并去重。
|
|
143
|
+
for url in array_urls:
|
|
144
|
+
if url not in seen:
|
|
145
|
+
seen.add(url)
|
|
146
|
+
urls.append(url)
|
|
147
|
+
array_urls.clear()
|
|
148
|
+
elif (
|
|
149
|
+
stack == ["[", "{"]
|
|
150
|
+
and before_previous == "cdn_url"
|
|
151
|
+
and previous == ":"
|
|
152
|
+
and is_string
|
|
153
|
+
and token[0] in ("'", '"')
|
|
154
|
+
):
|
|
155
|
+
url = _decode_text(token[1:-1])
|
|
156
|
+
if url:
|
|
157
|
+
array_urls.append(_normalize_image_url(url))
|
|
158
|
+
|
|
159
|
+
before_previous, previous = previous, token[1:-1] if is_string else token
|
|
110
160
|
return urls
|
|
111
161
|
|
|
112
162
|
|
|
@@ -138,6 +188,38 @@ def _extract_meta(soup: BeautifulSoup, result: ArticleResult) -> None:
|
|
|
138
188
|
setattr(result, attr, value)
|
|
139
189
|
|
|
140
190
|
|
|
191
|
+
def _extract_item_show_type(scripts: list[Tag]) -> int | None:
|
|
192
|
+
"""只读取当前文章的全局类型赋值;0 是有效类型,缺失时返回 None。"""
|
|
193
|
+
for script in scripts:
|
|
194
|
+
if "item_show_type" not in script.text:
|
|
195
|
+
continue
|
|
196
|
+
history: deque[str] = deque(maxlen=5)
|
|
197
|
+
brace_depth = 0
|
|
198
|
+
for match in _JS_DATA_TOKEN_RE.finditer(script.text):
|
|
199
|
+
if match.lastgroup == "comment":
|
|
200
|
+
continue
|
|
201
|
+
token = match.group()
|
|
202
|
+
previous = tuple(history)
|
|
203
|
+
is_global_var = (
|
|
204
|
+
brace_depth == 0
|
|
205
|
+
and previous[-3:] == ("var", "item_show_type", "=")
|
|
206
|
+
)
|
|
207
|
+
is_window_assignment = (
|
|
208
|
+
previous[-4:] == ("window", ".", "item_show_type", "=")
|
|
209
|
+
and (len(previous) < 5 or previous[-5] != ".")
|
|
210
|
+
)
|
|
211
|
+
if is_global_var or is_window_assignment:
|
|
212
|
+
value = token[1:-1] if token.startswith(("'", '"')) else token
|
|
213
|
+
if re.fullmatch(r"[0-9]+", value):
|
|
214
|
+
return int(value)
|
|
215
|
+
if token == "{":
|
|
216
|
+
brace_depth += 1
|
|
217
|
+
elif token == "}":
|
|
218
|
+
brace_depth -= 1
|
|
219
|
+
history.append(token)
|
|
220
|
+
return None
|
|
221
|
+
|
|
222
|
+
|
|
141
223
|
def _extract_account_type(script_text: str, result: ArticleResult) -> None:
|
|
142
224
|
"""从 new_service_type 中提取账号类型:0/1 → 订阅号,2 → 服务号。"""
|
|
143
225
|
if result.mp_account_type:
|
|
@@ -345,25 +427,30 @@ def _extract_plain_text_content(soup: BeautifulSoup, result: ArticleResult) -> N
|
|
|
345
427
|
|
|
346
428
|
def _extract_swiper_content(soup: BeautifulSoup, result: ArticleResult) -> None:
|
|
347
429
|
"""从小红书风格的图片轮播文章中提取内容。"""
|
|
430
|
+
images: list[str] = []
|
|
431
|
+
seen: set[str] = set()
|
|
432
|
+
description = ""
|
|
348
433
|
for script in soup.find_all("script", attrs={"type": "text/javascript"}):
|
|
349
|
-
|
|
350
|
-
|
|
351
|
-
|
|
352
|
-
|
|
353
|
-
|
|
354
|
-
html_parts = [f'<img src="{img}" /><br>' for img in result.images]
|
|
434
|
+
# 图片数据和 window.desc 可能位于不同的 script,分别收集。
|
|
435
|
+
for url in _extract_picture_cdn_urls(script.text):
|
|
436
|
+
if url not in seen:
|
|
437
|
+
seen.add(url)
|
|
438
|
+
images.append(url)
|
|
355
439
|
|
|
356
440
|
m = re.search(r'window.desc = "([^"]+)"', script.text)
|
|
357
441
|
if m:
|
|
358
|
-
|
|
359
|
-
html_parts.append(f"<p>{text}</p>")
|
|
442
|
+
description = _decode_text(m.group(1), preserve_newlines=True)
|
|
360
443
|
|
|
361
|
-
|
|
362
|
-
|
|
444
|
+
result.images = images
|
|
445
|
+
html_parts = [f'<img src="{img}" /><br>' for img in images]
|
|
446
|
+
if description:
|
|
447
|
+
html_parts.append(f"<p>{description}</p>")
|
|
448
|
+
if html_parts:
|
|
449
|
+
result.article_markdown = _to_markdown_plain("".join(html_parts))
|
|
363
450
|
|
|
364
451
|
|
|
365
452
|
def _extract_fullscreen_content(soup: BeautifulSoup, result: ArticleResult) -> None:
|
|
366
|
-
"""
|
|
453
|
+
"""从全屏布局文字消息中提取内容。
|
|
367
454
|
|
|
368
455
|
此类文章的文本存储在 text_page_info.content 中,通过 JsDecode() 编码;
|
|
369
456
|
图片存储在 picture_page_info_list 的 cdn_url 字段中。
|
|
@@ -402,6 +489,67 @@ def _extract_fullscreen_content(soup: BeautifulSoup, result: ArticleResult) -> N
|
|
|
402
489
|
# 主解析流程
|
|
403
490
|
# ---------------------------------------------------------------------------
|
|
404
491
|
|
|
492
|
+
# 顺序与原有 HTML 判断流程一致,用于未知类型及提取失败时的兜底。
|
|
493
|
+
_CONTENT_SELECTORS = {
|
|
494
|
+
"rich_text": "div.rich_media_content",
|
|
495
|
+
"repost": "div.original_page",
|
|
496
|
+
"plain_text": "p#js_text_desc",
|
|
497
|
+
"video": "div#js_common_share_desc_wrap",
|
|
498
|
+
"swiper": "div.share_media_swiper_content",
|
|
499
|
+
"fullscreen": "div#js_fullscreen_layout_padding",
|
|
500
|
+
}
|
|
501
|
+
_ITEM_SHOW_TYPE_BRANCHES = {
|
|
502
|
+
0: ("rich_text", "repost"),
|
|
503
|
+
5: ("video",),
|
|
504
|
+
8: ("swiper",),
|
|
505
|
+
10: ("plain_text", "fullscreen"),
|
|
506
|
+
}
|
|
507
|
+
_ARTICLE_DISPLAY_TYPES = {
|
|
508
|
+
0: ArticleDisplayType.ARTICLE,
|
|
509
|
+
5: ArticleDisplayType.VIDEO,
|
|
510
|
+
8: ArticleDisplayType.IMAGE,
|
|
511
|
+
10: ArticleDisplayType.TEXT,
|
|
512
|
+
}
|
|
513
|
+
|
|
514
|
+
|
|
515
|
+
def _try_content_branch(
|
|
516
|
+
branch: str,
|
|
517
|
+
soup: BeautifulSoup,
|
|
518
|
+
scripts: list[Tag],
|
|
519
|
+
result: ArticleResult,
|
|
520
|
+
*,
|
|
521
|
+
require_html: bool,
|
|
522
|
+
) -> ArticleResult | None:
|
|
523
|
+
content = soup.select_one(_CONTENT_SELECTORS[branch])
|
|
524
|
+
rich_text_template = branch in ("rich_text", "repost")
|
|
525
|
+
if content is None and (require_html or rich_text_template):
|
|
526
|
+
return None
|
|
527
|
+
|
|
528
|
+
# 失败分支的图片、元数据和标题处理不能污染后续分支。
|
|
529
|
+
candidate = replace(result, images=result.images.copy())
|
|
530
|
+
extract_meta = _extract_rich_text_meta if rich_text_template else _extract_swiper_meta
|
|
531
|
+
for script in scripts:
|
|
532
|
+
extract_meta(script.text, candidate)
|
|
533
|
+
|
|
534
|
+
if branch == "rich_text":
|
|
535
|
+
_extract_rich_media_content(content, candidate)
|
|
536
|
+
elif branch == "repost":
|
|
537
|
+
_extract_repost_content(content, candidate)
|
|
538
|
+
elif branch in ("plain_text", "video"):
|
|
539
|
+
_extract_plain_text_content(soup, candidate)
|
|
540
|
+
elif branch == "swiper":
|
|
541
|
+
_extract_swiper_content(soup, candidate)
|
|
542
|
+
elif branch == "fullscreen":
|
|
543
|
+
_extract_fullscreen_content(soup, candidate)
|
|
544
|
+
|
|
545
|
+
if not candidate.article_markdown.strip():
|
|
546
|
+
return None
|
|
547
|
+
if branch in ("plain_text", "fullscreen") and len(candidate.article_title) > 50:
|
|
548
|
+
short = candidate.article_title.split("。")[0]
|
|
549
|
+
candidate.article_title = short if len(short) <= 50 else candidate.article_title[:30]
|
|
550
|
+
return candidate
|
|
551
|
+
|
|
552
|
+
|
|
405
553
|
def _parse_html(url: str, html: str) -> ArticleResult:
|
|
406
554
|
"""将原始 HTML 解析为 ArticleResult。"""
|
|
407
555
|
# 检测微信验证码/人机验证页面
|
|
@@ -420,59 +568,26 @@ def _parse_html(url: str, html: str) -> ArticleResult:
|
|
|
420
568
|
|
|
421
569
|
scripts = soup.find_all("script", attrs={"type": "text/javascript"})
|
|
422
570
|
|
|
423
|
-
|
|
424
|
-
|
|
425
|
-
|
|
426
|
-
|
|
427
|
-
|
|
428
|
-
|
|
429
|
-
|
|
430
|
-
|
|
431
|
-
|
|
432
|
-
|
|
433
|
-
|
|
434
|
-
|
|
435
|
-
|
|
436
|
-
|
|
437
|
-
|
|
438
|
-
|
|
439
|
-
|
|
440
|
-
|
|
441
|
-
|
|
442
|
-
|
|
443
|
-
_extract_swiper_meta(s.text, result)
|
|
444
|
-
_extract_plain_text_content(soup, result)
|
|
445
|
-
if result.article_title and len(result.article_title) > 50:
|
|
446
|
-
short = result.article_title.split("。")[0]
|
|
447
|
-
result.article_title = short if len(short) <= 50 else result.article_title[:30]
|
|
448
|
-
return result
|
|
449
|
-
|
|
450
|
-
# 尝试解析视频分享类文章
|
|
451
|
-
content = soup.find("div", id="js_common_share_desc_wrap")
|
|
452
|
-
if content:
|
|
453
|
-
for s in scripts:
|
|
454
|
-
_extract_swiper_meta(s.text, result)
|
|
455
|
-
_extract_plain_text_content(soup, result)
|
|
456
|
-
return result
|
|
457
|
-
|
|
458
|
-
# 尝试解析小红书风格图片轮播文章
|
|
459
|
-
content = soup.find("div", class_="share_media_swiper_content")
|
|
460
|
-
if content:
|
|
461
|
-
for s in scripts:
|
|
462
|
-
_extract_swiper_meta(s.text, result)
|
|
463
|
-
_extract_swiper_content(soup, result)
|
|
464
|
-
return result
|
|
465
|
-
|
|
466
|
-
# 尝试解析全屏布局文章(如短文本帖子,appmsg_type 10002)
|
|
467
|
-
content = soup.find("div", id="js_fullscreen_layout_padding")
|
|
468
|
-
if content:
|
|
469
|
-
for s in scripts:
|
|
470
|
-
_extract_swiper_meta(s.text, result)
|
|
471
|
-
_extract_fullscreen_content(soup, result)
|
|
472
|
-
if result.article_title and len(result.article_title) > 50:
|
|
473
|
-
short = result.article_title.split("。")[0]
|
|
474
|
-
result.article_title = short if len(short) <= 50 else result.article_title[:30]
|
|
475
|
-
return result
|
|
571
|
+
item_show_type = _extract_item_show_type(scripts)
|
|
572
|
+
result.article_display_type = _ARTICLE_DISPLAY_TYPES.get(
|
|
573
|
+
item_show_type, ArticleDisplayType.UNKNOWN,
|
|
574
|
+
)
|
|
575
|
+
preferred = _ITEM_SHOW_TYPE_BRANCHES.get(item_show_type, ())
|
|
576
|
+
attempted: set[str] = set()
|
|
577
|
+
for branch in (*preferred, *_CONTENT_SELECTORS):
|
|
578
|
+
if branch in attempted:
|
|
579
|
+
continue
|
|
580
|
+
attempted.add(branch)
|
|
581
|
+
# 已知类型先用对应的 DOM 或脚本数据细分模板;失败后按原有 HTML 顺序兜底。
|
|
582
|
+
candidate = _try_content_branch(
|
|
583
|
+
branch,
|
|
584
|
+
soup,
|
|
585
|
+
scripts,
|
|
586
|
+
result,
|
|
587
|
+
require_html=branch not in preferred,
|
|
588
|
+
)
|
|
589
|
+
if candidate is not None:
|
|
590
|
+
return candidate
|
|
476
591
|
|
|
477
592
|
# 兜底:尽可能提取元数据
|
|
478
593
|
for s in scripts:
|
|
@@ -0,0 +1,89 @@
|
|
|
1
|
+
"""按展示类型检查解析结果的有效性。"""
|
|
2
|
+
|
|
3
|
+
from dataclasses import asdict, replace
|
|
4
|
+
import json
|
|
5
|
+
|
|
6
|
+
import pytest
|
|
7
|
+
|
|
8
|
+
from wechat_article_parser import ArticleDisplayType, ArticleResult
|
|
9
|
+
|
|
10
|
+
|
|
11
|
+
_COMPLETE = ArticleResult(
|
|
12
|
+
mp_id=123,
|
|
13
|
+
mp_name="测试公众号",
|
|
14
|
+
article_id="ncPoZnxKdvDRLM_s4DBgmA",
|
|
15
|
+
article_msg_id=456,
|
|
16
|
+
article_idx=1,
|
|
17
|
+
article_sn="signature",
|
|
18
|
+
article_title="测试文章",
|
|
19
|
+
article_publish_time=1788860293,
|
|
20
|
+
)
|
|
21
|
+
|
|
22
|
+
|
|
23
|
+
@pytest.mark.parametrize("display_type", list(ArticleDisplayType))
|
|
24
|
+
@pytest.mark.parametrize("markdown", ["", "正文"])
|
|
25
|
+
def test_only_video_can_be_valid_without_markdown(
|
|
26
|
+
display_type: ArticleDisplayType, markdown: str,
|
|
27
|
+
) -> None:
|
|
28
|
+
result = replace(
|
|
29
|
+
_COMPLETE, article_display_type=display_type, article_markdown=markdown,
|
|
30
|
+
)
|
|
31
|
+
|
|
32
|
+
assert result.is_valid is (bool(markdown) or display_type == ArticleDisplayType.VIDEO)
|
|
33
|
+
|
|
34
|
+
|
|
35
|
+
@pytest.mark.parametrize(
|
|
36
|
+
("field_name", "empty_value"),
|
|
37
|
+
[
|
|
38
|
+
("mp_id", 0),
|
|
39
|
+
("mp_name", ""),
|
|
40
|
+
("article_id", ""),
|
|
41
|
+
("article_msg_id", 0),
|
|
42
|
+
("article_idx", 0),
|
|
43
|
+
("article_sn", ""),
|
|
44
|
+
("article_title", ""),
|
|
45
|
+
("article_publish_time", 0),
|
|
46
|
+
],
|
|
47
|
+
)
|
|
48
|
+
@pytest.mark.parametrize("markdown", ["", "视频附言"])
|
|
49
|
+
def test_video_still_requires_all_key_metadata(
|
|
50
|
+
field_name: str, empty_value: str | int, markdown: str,
|
|
51
|
+
) -> None:
|
|
52
|
+
result = replace(
|
|
53
|
+
_COMPLETE,
|
|
54
|
+
article_display_type=ArticleDisplayType.VIDEO,
|
|
55
|
+
article_markdown=markdown,
|
|
56
|
+
**{field_name: empty_value},
|
|
57
|
+
)
|
|
58
|
+
|
|
59
|
+
assert not result.is_valid
|
|
60
|
+
|
|
61
|
+
|
|
62
|
+
@pytest.mark.parametrize(
|
|
63
|
+
"expected",
|
|
64
|
+
[
|
|
65
|
+
"article",
|
|
66
|
+
"video",
|
|
67
|
+
"image",
|
|
68
|
+
"text",
|
|
69
|
+
"unknown",
|
|
70
|
+
],
|
|
71
|
+
)
|
|
72
|
+
def test_display_type_has_semantic_string_values(expected: str) -> None:
|
|
73
|
+
result = ArticleResult(article_display_type=ArticleDisplayType(expected))
|
|
74
|
+
|
|
75
|
+
assert result.article_display_type is ArticleDisplayType(expected)
|
|
76
|
+
assert result.article_display_type == expected
|
|
77
|
+
assert str(result.article_display_type) == expected
|
|
78
|
+
assert f"{result.article_display_type:>10}" == f"{expected:>10}"
|
|
79
|
+
assert json.dumps(result.article_display_type) == json.dumps(expected)
|
|
80
|
+
serialized = json.loads(json.dumps(asdict(result)))
|
|
81
|
+
assert serialized["article_display_type"] == expected
|
|
82
|
+
assert "item_show_type" not in serialized
|
|
83
|
+
assert not hasattr(result, "item_show_type")
|
|
84
|
+
|
|
85
|
+
|
|
86
|
+
def test_default_display_type_is_unknown() -> None:
|
|
87
|
+
result = ArticleResult()
|
|
88
|
+
assert result.article_display_type is ArticleDisplayType.UNKNOWN
|
|
89
|
+
assert not result.is_valid
|
|
@@ -15,6 +15,7 @@ TEST_URLS = [
|
|
|
15
15
|
"https://mp.weixin.qq.com/s/h8E6riExCaH2Znmnprj-WQ",
|
|
16
16
|
"https://mp.weixin.qq.com/s/ySQdtsRlRmAl_skdc5HQ-A",
|
|
17
17
|
"https://mp.weixin.qq.com/s/MnkArbYQNp3tF29gujMUnQ",
|
|
18
|
+
"https://mp.weixin.qq.com/s/ncPoZnxKdvDRLM_s4DBgmA",
|
|
18
19
|
]
|
|
19
20
|
|
|
20
21
|
|
|
@@ -62,6 +63,7 @@ def _assert_result(result: ArticleResult, url: str) -> None:
|
|
|
62
63
|
print(f"文章idx: {result.article_idx}")
|
|
63
64
|
print(f"文章签名: {result.article_sn}")
|
|
64
65
|
print(f"文章标题: {result.article_title}")
|
|
66
|
+
print(f"展示类型: {result.article_display_type}")
|
|
65
67
|
print(
|
|
66
68
|
f"封面图: {result.article_cover_image[:80]}..."
|
|
67
69
|
if result.article_cover_image
|
|
@@ -85,7 +87,8 @@ def _assert_result(result: ArticleResult, url: str) -> None:
|
|
|
85
87
|
assert result.article_idx > 0, "article_idx should be positive"
|
|
86
88
|
assert result.article_sn, "article_sn should not be empty"
|
|
87
89
|
assert result.article_title, "article_title should not be empty"
|
|
88
|
-
|
|
90
|
+
if result.article_display_type != "video":
|
|
91
|
+
assert result.article_markdown, "article_markdown should not be empty for non-video articles"
|
|
89
92
|
assert result.article_publish_time > 0, "article_publish_time should be positive"
|
|
90
93
|
assert result.is_valid
|
|
91
94
|
|
|
@@ -110,6 +113,7 @@ def test_fetch_all(url: str, proxy: str | None) -> None:
|
|
|
110
113
|
print(f"文章idx: {result.article_idx}")
|
|
111
114
|
print(f"文章签名: {result.article_sn}")
|
|
112
115
|
print(f"文章标题: {result.article_title}")
|
|
116
|
+
print(f"展示类型: {result.article_display_type}")
|
|
113
117
|
print(f"封面图: {result.article_cover_image}")
|
|
114
118
|
print(f"文章摘要: {result.article_description}")
|
|
115
119
|
print(f"发布时间: {result.article_publish_time}")
|
|
@@ -0,0 +1,150 @@
|
|
|
1
|
+
"""使用精简页面离线回归轮播图片的提取和 Markdown 输出。"""
|
|
2
|
+
|
|
3
|
+
import pytest
|
|
4
|
+
|
|
5
|
+
from wechat_article_parser.parser import _parse_html
|
|
6
|
+
|
|
7
|
+
|
|
8
|
+
_URL = "https://mp.weixin.qq.com/s/0Wz3JeMbtWBL5iWJgYPS_Q"
|
|
9
|
+
_EXPECTED_IMAGES = [
|
|
10
|
+
"https://mmbiz.qpic.cn/mmbiz_png/second/640",
|
|
11
|
+
"https://mmbiz.qpic.cn/mmbiz_png/first/640",
|
|
12
|
+
]
|
|
13
|
+
_PICTURES = r"""[
|
|
14
|
+
{
|
|
15
|
+
note: '含括号 ] } 和转义引号 \' 的字符串',
|
|
16
|
+
watermark_info: {
|
|
17
|
+
position: {x: 1},
|
|
18
|
+
cdn_url: 'https://mmbiz.qpic.cn/mmbiz_png/watermark/0'
|
|
19
|
+
},
|
|
20
|
+
cdn_url: 'https://mmbiz.qpic.cn/mmbiz_png/second/0?wx_fmt=png',
|
|
21
|
+
share_cover: {
|
|
22
|
+
crop_info: {x: 0},
|
|
23
|
+
cdn_url: 'https://mmbiz.qpic.cn/mmbiz_png/share-cover/0'
|
|
24
|
+
},
|
|
25
|
+
live_photo: {format_info: [{cdn_url: 'https://example.com/live-photo'}]}
|
|
26
|
+
},
|
|
27
|
+
{"cdn_url": "https://mmbiz.qpic.cn/mmbiz_png/first/0?wx_fmt=png"},
|
|
28
|
+
{cdn_url: 'https://mmbiz.qpic.cn/mmbiz_png/second/640'},
|
|
29
|
+
{cdn_url: ''}
|
|
30
|
+
]"""
|
|
31
|
+
_DESCRIPTION = r'window.desc = "第一行\x0a第二行";'
|
|
32
|
+
_REFERENCE = """
|
|
33
|
+
window.picture_page_info_list =
|
|
34
|
+
((window.cgiDataNew && window.cgiDataNew.picture_page_info_list) || []).slice(0, 20);
|
|
35
|
+
"""
|
|
36
|
+
|
|
37
|
+
|
|
38
|
+
def _html(*scripts: str, fullscreen: bool = False) -> str:
|
|
39
|
+
content = (
|
|
40
|
+
'<div id="js_fullscreen_layout_padding"></div>'
|
|
41
|
+
if fullscreen
|
|
42
|
+
else '<div class="share_media_swiper_content"></div>'
|
|
43
|
+
)
|
|
44
|
+
return content + "".join(
|
|
45
|
+
f'<script type="text/javascript">{script}</script>' for script in scripts
|
|
46
|
+
)
|
|
47
|
+
|
|
48
|
+
|
|
49
|
+
@pytest.mark.parametrize("layout", ["inline", "data_first", "description_first"])
|
|
50
|
+
def test_swiper_collects_body_images_across_scripts(layout: str) -> None:
|
|
51
|
+
data = (
|
|
52
|
+
"window.cgiDataNew = {"
|
|
53
|
+
"cdn_url: 'https://mmbiz.qpic.cn/mmbiz_png/outside-before/0',"
|
|
54
|
+
f"picture_page_info_list: {_PICTURES},"
|
|
55
|
+
"other: {cdn_url: 'https://mmbiz.qpic.cn/mmbiz_png/outside-after/0'}"
|
|
56
|
+
"};"
|
|
57
|
+
)
|
|
58
|
+
reference = _REFERENCE + _DESCRIPTION
|
|
59
|
+
if layout == "inline":
|
|
60
|
+
scripts = [f"window.picture_page_info_list = {_PICTURES};" + _DESCRIPTION]
|
|
61
|
+
elif layout == "data_first":
|
|
62
|
+
scripts = [data, reference]
|
|
63
|
+
else:
|
|
64
|
+
scripts = [reference, data]
|
|
65
|
+
|
|
66
|
+
result = _parse_html(_URL, _html(*scripts))
|
|
67
|
+
|
|
68
|
+
assert result.images == _EXPECTED_IMAGES
|
|
69
|
+
assert result.article_markdown.count("![") == 2
|
|
70
|
+
assert result.article_markdown.index(_EXPECTED_IMAGES[0]) < result.article_markdown.index(_EXPECTED_IMAGES[1])
|
|
71
|
+
assert "第一行 \n第二行" in result.article_markdown
|
|
72
|
+
assert "outside" not in result.article_markdown
|
|
73
|
+
assert "watermark" not in result.article_markdown
|
|
74
|
+
assert "share-cover" not in result.article_markdown
|
|
75
|
+
assert "live-photo" not in result.article_markdown
|
|
76
|
+
|
|
77
|
+
|
|
78
|
+
def test_swiper_later_empty_lists_do_not_erase_images() -> None:
|
|
79
|
+
result = _parse_html(
|
|
80
|
+
_URL,
|
|
81
|
+
_html(
|
|
82
|
+
f"window.picture_page_info_list = {_PICTURES};" + _DESCRIPTION,
|
|
83
|
+
"window.picture_page_info_list = [];",
|
|
84
|
+
_REFERENCE,
|
|
85
|
+
),
|
|
86
|
+
)
|
|
87
|
+
|
|
88
|
+
assert result.images == _EXPECTED_IMAGES
|
|
89
|
+
assert result.article_markdown.count("![") == 2
|
|
90
|
+
assert "第一行 \n第二行" in result.article_markdown
|
|
91
|
+
|
|
92
|
+
|
|
93
|
+
def test_swiper_ignores_list_text_in_strings_and_comments() -> None:
|
|
94
|
+
decoy = "picture_page_info_list: [{cdn_url: 'https://example.com/decoy'}]"
|
|
95
|
+
result = _parse_html(
|
|
96
|
+
_URL,
|
|
97
|
+
_html(
|
|
98
|
+
f'var example = "{decoy}";\n'
|
|
99
|
+
f"// {decoy}\n"
|
|
100
|
+
f"/* {decoy} */\n"
|
|
101
|
+
f"var template = `{decoy}`;\n",
|
|
102
|
+
f"window.cgiDataNew = {{'picture_page_info_list': {_PICTURES}}};",
|
|
103
|
+
_REFERENCE + _DESCRIPTION,
|
|
104
|
+
),
|
|
105
|
+
)
|
|
106
|
+
|
|
107
|
+
assert result.images == _EXPECTED_IMAGES
|
|
108
|
+
assert "decoy" not in result.article_markdown
|
|
109
|
+
|
|
110
|
+
|
|
111
|
+
def test_swiper_image_only_content() -> None:
|
|
112
|
+
result = _parse_html(
|
|
113
|
+
_URL,
|
|
114
|
+
_html(f"window.cgiDataNew = {{picture_page_info_list: {_PICTURES}}};", _REFERENCE),
|
|
115
|
+
)
|
|
116
|
+
|
|
117
|
+
assert result.images == _EXPECTED_IMAGES
|
|
118
|
+
assert result.article_markdown.count("![") == 2
|
|
119
|
+
|
|
120
|
+
|
|
121
|
+
def test_swiper_text_only_content() -> None:
|
|
122
|
+
result = _parse_html(
|
|
123
|
+
_URL,
|
|
124
|
+
_html(
|
|
125
|
+
"window.cgiDataNew = {picture_page_info_list: [], cdn_url: 'https://example.com/cover'};",
|
|
126
|
+
_REFERENCE,
|
|
127
|
+
_DESCRIPTION,
|
|
128
|
+
),
|
|
129
|
+
)
|
|
130
|
+
|
|
131
|
+
assert result.images == []
|
|
132
|
+
assert result.article_markdown == "第一行 \n第二行"
|
|
133
|
+
|
|
134
|
+
|
|
135
|
+
def test_fullscreen_shared_image_extractor_keeps_body_images() -> None:
|
|
136
|
+
result = _parse_html(
|
|
137
|
+
_URL,
|
|
138
|
+
_html(
|
|
139
|
+
"window.cgiDataNew = {"
|
|
140
|
+
"cdn_url: 'https://example.com/cover',"
|
|
141
|
+
f"picture_page_info_list: {_PICTURES},"
|
|
142
|
+
"text_page_info: {content: JsDecode('全屏正文')}"
|
|
143
|
+
"};",
|
|
144
|
+
fullscreen=True,
|
|
145
|
+
),
|
|
146
|
+
)
|
|
147
|
+
|
|
148
|
+
assert result.images == _EXPECTED_IMAGES
|
|
149
|
+
assert result.article_markdown.count("![") == 2
|
|
150
|
+
assert "全屏正文" in result.article_markdown
|
|
@@ -0,0 +1,241 @@
|
|
|
1
|
+
"""类型字段优先分发、模板细分及 HTML 兜底的离线回归测试。"""
|
|
2
|
+
|
|
3
|
+
import pytest
|
|
4
|
+
from bs4 import BeautifulSoup
|
|
5
|
+
|
|
6
|
+
from wechat_article_parser import ArticleDisplayType, parser
|
|
7
|
+
|
|
8
|
+
|
|
9
|
+
_URL = "https://mp.weixin.qq.com/s/0Wz3JeMbtWBL5iWJgYPS_Q"
|
|
10
|
+
_RICH_TEXT = '<div class="rich_media_content"><p>富文本正文</p></div>'
|
|
11
|
+
_REPOST = """
|
|
12
|
+
<div class="original_page">
|
|
13
|
+
<p id="js_share_notice"><script>notice.innerHTML = "转载附言";</script></p>
|
|
14
|
+
<span id="js_share_source" data-url="https://example.com/original"></span>
|
|
15
|
+
</div>
|
|
16
|
+
"""
|
|
17
|
+
_PLAIN_DATA = """
|
|
18
|
+
var TextContentNoEncode = '';
|
|
19
|
+
var ContentNoEncode = window.a_value_which_never_exists || '文字正文';
|
|
20
|
+
"""
|
|
21
|
+
_SWIPER_DATA = """
|
|
22
|
+
window.cgiDataNew = {
|
|
23
|
+
picture_page_info_list: [{cdn_url: 'https://mmbiz.qpic.cn/mmbiz_png/body/0'}]
|
|
24
|
+
};
|
|
25
|
+
window.desc = "轮播正文";
|
|
26
|
+
"""
|
|
27
|
+
_FULLSCREEN_DATA = """
|
|
28
|
+
window.cgiDataNew = {
|
|
29
|
+
picture_page_info_list: [],
|
|
30
|
+
text_page_info: {content: JsDecode('全屏正文')}
|
|
31
|
+
};
|
|
32
|
+
"""
|
|
33
|
+
|
|
34
|
+
|
|
35
|
+
def _html(body: str = "", script: str = "", title: str = "原标题") -> str:
|
|
36
|
+
return (
|
|
37
|
+
f'<meta property="og:title" content="{title}">'
|
|
38
|
+
f"{body}<script type=\"text/javascript\">{script}</script>"
|
|
39
|
+
)
|
|
40
|
+
|
|
41
|
+
|
|
42
|
+
@pytest.mark.parametrize(
|
|
43
|
+
("script", "expected"),
|
|
44
|
+
[
|
|
45
|
+
('var item_show_type = "0";', 0),
|
|
46
|
+
("window.item_show_type = '8' || '';", 8),
|
|
47
|
+
("window.item_show_type = 10;", 10),
|
|
48
|
+
("var /* 注释 */ item_show_type = 5;", 5),
|
|
49
|
+
("(function() { window.item_show_type = '10'; })();", 10),
|
|
50
|
+
("window.item_show_type = '999';", 999),
|
|
51
|
+
("window.item_show_type = '';", None),
|
|
52
|
+
("window.item_show_type = 'unknown';", None),
|
|
53
|
+
("window.item_show_type = null;", None),
|
|
54
|
+
("window.item_show_type = -1;", None),
|
|
55
|
+
("window.item_show_type = 8.5;", None),
|
|
56
|
+
("window.item_show_type = 8e2;", None),
|
|
57
|
+
("window.item_show_type == 8;", None),
|
|
58
|
+
("var real_item_show_type = '8';", None),
|
|
59
|
+
("", None),
|
|
60
|
+
],
|
|
61
|
+
)
|
|
62
|
+
def test_read_current_article_type(script: str, expected: int | None) -> None:
|
|
63
|
+
soup = BeautifulSoup(_html(script=script), "html.parser")
|
|
64
|
+
|
|
65
|
+
assert parser._extract_item_show_type(soup.find_all("script")) == expected
|
|
66
|
+
|
|
67
|
+
|
|
68
|
+
@pytest.mark.parametrize("current_assignment", ["", "window.item_show_type = '8';"])
|
|
69
|
+
def test_ignore_type_in_references_strings_comments_and_local_variables(
|
|
70
|
+
current_assignment: str,
|
|
71
|
+
) -> None:
|
|
72
|
+
decoys = """
|
|
73
|
+
var reference = {item_show_type: 5};
|
|
74
|
+
var example = "var item_show_type = '5';";
|
|
75
|
+
var url = 'https://example.com/?item_show_type=5';
|
|
76
|
+
var template = `
|
|
77
|
+
window.item_show_type = '5';
|
|
78
|
+
`;
|
|
79
|
+
// window.item_show_type = '5';
|
|
80
|
+
/* var item_show_type = '5'; */
|
|
81
|
+
function report() { var item_show_type = '5'; }
|
|
82
|
+
report.window.item_show_type = '5';
|
|
83
|
+
window.real_item_show_type = '5';
|
|
84
|
+
"""
|
|
85
|
+
soup = BeautifulSoup(_html(script=decoys + current_assignment), "html.parser")
|
|
86
|
+
|
|
87
|
+
expected = 8 if current_assignment else None
|
|
88
|
+
assert parser._extract_item_show_type(soup.find_all("script")) == expected
|
|
89
|
+
|
|
90
|
+
|
|
91
|
+
@pytest.mark.parametrize(
|
|
92
|
+
("body", "expected"),
|
|
93
|
+
[
|
|
94
|
+
(_RICH_TEXT + _REPOST, "富文本正文"),
|
|
95
|
+
(_REPOST, "转载附言"),
|
|
96
|
+
('<div class="rich_media_content"></div>' + _REPOST, "转载附言"),
|
|
97
|
+
],
|
|
98
|
+
)
|
|
99
|
+
def test_type_zero_uses_html_to_distinguish_rich_text_and_repost(
|
|
100
|
+
body: str, expected: str,
|
|
101
|
+
) -> None:
|
|
102
|
+
result = parser._parse_html(_URL, _html(body, 'var item_show_type = "0";'))
|
|
103
|
+
|
|
104
|
+
assert result.article_display_type is ArticleDisplayType.ARTICLE
|
|
105
|
+
assert expected in result.article_markdown
|
|
106
|
+
if expected == "转载附言":
|
|
107
|
+
assert "[查看原文](https://example.com/original)" in result.article_markdown
|
|
108
|
+
else:
|
|
109
|
+
assert "转载附言" not in result.article_markdown
|
|
110
|
+
|
|
111
|
+
|
|
112
|
+
@pytest.mark.parametrize(
|
|
113
|
+
("item_type", "data", "expected", "display_type"),
|
|
114
|
+
[
|
|
115
|
+
(5, _PLAIN_DATA, "文字正文", ArticleDisplayType.VIDEO),
|
|
116
|
+
(8, _SWIPER_DATA, "轮播正文", ArticleDisplayType.IMAGE),
|
|
117
|
+
(10, _PLAIN_DATA, "文字正文", ArticleDisplayType.TEXT),
|
|
118
|
+
(10, _FULLSCREEN_DATA, "全屏正文", ArticleDisplayType.TEXT),
|
|
119
|
+
],
|
|
120
|
+
)
|
|
121
|
+
def test_known_type_takes_priority_over_other_html_templates(
|
|
122
|
+
item_type: int, data: str, expected: str, display_type: ArticleDisplayType,
|
|
123
|
+
) -> None:
|
|
124
|
+
# 即使存在富文本壳,已知类型仍先读取对应脚本数据;无需专用容器。
|
|
125
|
+
result = parser._parse_html(
|
|
126
|
+
_URL,
|
|
127
|
+
_html(_RICH_TEXT, f"window.item_show_type = '{item_type}';" + data),
|
|
128
|
+
)
|
|
129
|
+
|
|
130
|
+
assert expected in result.article_markdown
|
|
131
|
+
assert "富文本正文" not in result.article_markdown
|
|
132
|
+
assert result.article_display_type is display_type
|
|
133
|
+
assert not hasattr(result, "item_show_type")
|
|
134
|
+
if item_type == 8:
|
|
135
|
+
assert result.images == ["https://mmbiz.qpic.cn/mmbiz_png/body/640"]
|
|
136
|
+
|
|
137
|
+
|
|
138
|
+
@pytest.mark.parametrize(
|
|
139
|
+
"declaration",
|
|
140
|
+
["", "window.item_show_type = '999';", "window.item_show_type = 'invalid';"],
|
|
141
|
+
)
|
|
142
|
+
def test_missing_unknown_or_invalid_type_falls_back_to_html(declaration: str) -> None:
|
|
143
|
+
result = parser._parse_html(_URL, _html(_RICH_TEXT, declaration))
|
|
144
|
+
|
|
145
|
+
assert result.article_markdown == "富文本正文"
|
|
146
|
+
assert result.article_display_type is ArticleDisplayType.UNKNOWN
|
|
147
|
+
|
|
148
|
+
|
|
149
|
+
@pytest.mark.parametrize("item_type", [5, 8, 10])
|
|
150
|
+
def test_known_type_without_matching_data_falls_back_to_html(item_type: int) -> None:
|
|
151
|
+
result = parser._parse_html(
|
|
152
|
+
_URL,
|
|
153
|
+
_html(_RICH_TEXT, f"window.item_show_type = '{item_type}';"),
|
|
154
|
+
)
|
|
155
|
+
|
|
156
|
+
assert result.article_markdown == "富文本正文"
|
|
157
|
+
|
|
158
|
+
|
|
159
|
+
def test_type_zero_without_matching_html_falls_back_to_other_templates() -> None:
|
|
160
|
+
result = parser._parse_html(
|
|
161
|
+
_URL,
|
|
162
|
+
_html(
|
|
163
|
+
'<div class="share_media_swiper_content"></div>',
|
|
164
|
+
"window.item_show_type = '0';" + _SWIPER_DATA,
|
|
165
|
+
),
|
|
166
|
+
)
|
|
167
|
+
|
|
168
|
+
assert "轮播正文" in result.article_markdown
|
|
169
|
+
assert len(result.images) == 1
|
|
170
|
+
|
|
171
|
+
|
|
172
|
+
def test_failed_branch_does_not_pollute_fallback_result(monkeypatch) -> None:
|
|
173
|
+
def fail_swiper(soup, result):
|
|
174
|
+
result.article_title = "错误标题"
|
|
175
|
+
result.mp_name = "错误账号"
|
|
176
|
+
result.images.append("https://example.com/wrong-image")
|
|
177
|
+
result.article_markdown = " \n"
|
|
178
|
+
|
|
179
|
+
monkeypatch.setattr(parser, "_extract_swiper_content", fail_swiper)
|
|
180
|
+
result = parser._parse_html(
|
|
181
|
+
_URL,
|
|
182
|
+
_html(_RICH_TEXT, "window.item_show_type = '8';"),
|
|
183
|
+
)
|
|
184
|
+
|
|
185
|
+
assert result.article_title == "原标题"
|
|
186
|
+
assert result.mp_name == ""
|
|
187
|
+
assert result.images == []
|
|
188
|
+
assert result.article_markdown == "富文本正文"
|
|
189
|
+
|
|
190
|
+
|
|
191
|
+
def test_unrecognized_content_still_returns_metadata() -> None:
|
|
192
|
+
result = parser._parse_html(
|
|
193
|
+
_URL,
|
|
194
|
+
_html(script='window.item_show_type = "999"; window.alias = "account";'),
|
|
195
|
+
)
|
|
196
|
+
|
|
197
|
+
assert result.article_title == "原标题"
|
|
198
|
+
assert result.mp_alias == "account"
|
|
199
|
+
assert result.article_display_type is ArticleDisplayType.UNKNOWN
|
|
200
|
+
assert result.article_markdown == ""
|
|
201
|
+
assert not result.is_valid
|
|
202
|
+
|
|
203
|
+
|
|
204
|
+
@pytest.mark.parametrize(
|
|
205
|
+
("declaration", "expected_type"),
|
|
206
|
+
[
|
|
207
|
+
("", ArticleDisplayType.UNKNOWN),
|
|
208
|
+
("window.item_show_type = '5';", ArticleDisplayType.VIDEO),
|
|
209
|
+
],
|
|
210
|
+
)
|
|
211
|
+
def test_video_without_description_keeps_type_and_metadata(
|
|
212
|
+
declaration: str, expected_type: ArticleDisplayType,
|
|
213
|
+
) -> None:
|
|
214
|
+
metadata = """
|
|
215
|
+
window.__initCgiDataConfig = function(d) {
|
|
216
|
+
return {
|
|
217
|
+
biz: d.biz ? d.biz : 'MQ==',
|
|
218
|
+
nick_name: d.nick_name ? d.nick_name : '测试公众号',
|
|
219
|
+
mid: d.mid ? d.mid : '123',
|
|
220
|
+
idx: d.idx ? d.idx : '1',
|
|
221
|
+
sn: d.sn ? d.sn : 'signature',
|
|
222
|
+
create_time: d.create_time ? d.create_time : '1788860293'
|
|
223
|
+
};
|
|
224
|
+
};
|
|
225
|
+
var TextContentNoEncode = window.a_value_which_never_exists || '';
|
|
226
|
+
var ContentNoEncode = window.a_value_which_never_exists || '';
|
|
227
|
+
"""
|
|
228
|
+
result = parser._parse_html(
|
|
229
|
+
_URL,
|
|
230
|
+
_html(
|
|
231
|
+
'<div id="js_common_share_desc_wrap" style="display: none"></div>'
|
|
232
|
+
'<div id="js_fullscreen_layout_padding"></div>',
|
|
233
|
+
declaration + metadata,
|
|
234
|
+
),
|
|
235
|
+
)
|
|
236
|
+
|
|
237
|
+
assert result.article_display_type is expected_type
|
|
238
|
+
assert result.mp_id == 1
|
|
239
|
+
assert result.mp_name == "测试公众号"
|
|
240
|
+
assert result.article_markdown == ""
|
|
241
|
+
assert result.is_valid is (expected_type == ArticleDisplayType.VIDEO)
|
|
File without changes
|
|
File without changes
|
|
File without changes
|