@geoly-ai/social-hub-cli 0.3.15 → 0.3.16

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -4,6 +4,14 @@ All notable changes to `@geoly-ai/social-hub-cli` are documented in this file.
4
4
 
5
5
  Release workflow: `pnpm release:cli` → `pnpm publish:cli`.
6
6
 
7
+ ## [0.3.16] - 2026-08-10
8
+
9
+ Add Arctic Shift archived comment bodies, reconstructable comment trees, filtered search, and comment-ID lookup to reddit-voc-volume.
10
+
11
+ ### Added
12
+
13
+ - Reddit VOC 支持 Arctic 历史评论正文与评论树
14
+
7
15
  ## [0.3.15] - 2026-08-09
8
16
 
9
17
  评论样本精确 count 读口:GET /v1/system/subreddit-intelligence/comments/count + SDK countRedditCommentSamples + CLI intel comments-count(不封顶、无分页、拒绝 keyword;行级授权语义与 comments 列表一致)
@@ -1,6 +1,6 @@
1
1
  {
2
- "generatedAt": "2026-08-09T09:26:27.640Z",
3
- "cliVersion": "0.3.15",
2
+ "generatedAt": "2026-08-11T01:41:10.558Z",
3
+ "cliVersion": "0.3.16",
4
4
  "commands": [
5
5
  "accounts",
6
6
  "accounts brand-bindings",
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@geoly-ai/social-hub-cli",
3
- "version": "0.3.15",
3
+ "version": "0.3.16",
4
4
  "type": "module",
5
5
  "description": "social-hub CLI for Social Ops Hub",
6
6
  "repository": {
package/skills/README.md CHANGED
@@ -156,7 +156,7 @@ pnpm skills:check
156
156
  | [reddit-content-writing](reddit-content-writing) | 主帖批量写作(9 主帖型 × 7 写作变体) |
157
157
  | [reddit-content-rewrite](reddit-content-rewrite) | 约束保持的帖/评/标题修复改写 |
158
158
  | [reddit-delivery-check](reddit-delivery-check) | 阶段 5 交付 QA:覆盖/完整性/多样性/合规校验 |
159
- | [reddit-voc-volume](reddit-voc-volume) | VOC 声量采集 + Arctic 历史回溯 + Firecrawl 兜底 |
159
+ | [reddit-voc-volume](reddit-voc-volume) | VOC 声量采集 + Arctic 历史帖子/评论树 + Firecrawl 兜底 |
160
160
  | [reddit-brand-risk-response](reddit-brand-risk-response) | 品牌风险 triage/预防/聚合方向分析 |
161
161
  | [reddit-subreddit-compliance](reddit-subreddit-compliance) | subreddit 合规与板块适配检测(版规/发帖前风险) |
162
162
  | [reddit-answers-research](reddit-answers-research) | Reddit 官方 Ask AI 研究 → 提取引用帖取证 |
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "version": 1,
3
3
  "package": "social-ops-hub",
4
- "cliVersion": "0.3.15",
4
+ "cliVersion": "0.3.16",
5
5
  "repository": "social-ops-hub",
6
6
  "publish": {
7
7
  "registry": "https://skill.sh",
@@ -152,7 +152,7 @@
152
152
  {
153
153
  "id": "reddit-voc-volume",
154
154
  "path": "reddit-voc-volume",
155
- "description": "VOC 声量采集:关键词/品牌词搜索 + subreddit 扫描 + Arctic 历史回溯(全文不截断)+ Firecrawl 兜底"
155
+ "description": "VOC 声量采集:关键词/品牌词搜索 + subreddit 扫描 + Arctic 历史帖子/评论正文与评论树 + Firecrawl 兜底"
156
156
  },
157
157
  {
158
158
  "id": "reddit-brand-risk-response",
@@ -28,7 +28,7 @@
28
28
  "contentSha256": "234d7fd1d8f2db68444674c20418c836c1b2488305b20399114dec89a0de32fc"
29
29
  },
30
30
  "reddit-voc-volume": {
31
- "contentSha256": "00e0675d6c77d5c08f3232fa7b6a7ecd8bea6d75b0ebeb67ae01254483695639"
31
+ "contentSha256": "45332556a0835a5c16386cb31550c04ea2447e2326a7a4ff15fd050f1ae91d55"
32
32
  },
33
33
  "reddit-brand-risk-response": {
34
34
  "contentSha256": "dbba80345ebe68818a25e7604dfaa6951c6bf0e4d785d5013be7a06b22ec5b7c"
@@ -1,20 +1,21 @@
1
1
  ---
2
2
  name: reddit-voc-volume
3
3
  description: |
4
- 获取 Reddit VOC 声量、历史帖文与公开 URL 证据:按关键词/品牌词搜索、指定 subreddit 扫描、批量链接汇总帖子量/评论量/score/子版分布,**按时间窗回溯 2005 年至今的历史帖子(正文全文,不截断)**,读取版规/wiki/subreddit 元数据,并通过 Social Hub 或自托管 Firecrawl 读取公开 About 页周访客/周贡献/周活跃用户兼容字段、Installed Apps、flair filters、通用 page_signals,以及**核验归档 score 与版规是否仍然现行**,输出结构化 JSON。
5
- 用户提到 Reddit VOC、品牌词/竞品讨论量、subreddit 声量、帖子评论量统计、舆情采样、批量 Reddit 链接、about/rules/wiki、版规、**历史帖子/往年数据/时间窗回溯/某年某月的帖子/长期趋势**、置顶帖/pinned/sticky/megathread、公开 About 页周指标、Installed Apps、flair filters、页面公开链接信号,或需要**建 subreddit baseline 卡/统计正文长度分布**时必须用本 skill。
4
+ 获取 Reddit VOC 声量、历史帖文、**历史评论正文/完整评论树**与公开 URL 证据:按关键词/品牌词搜索、指定 subreddit 扫描、批量链接汇总帖子量/评论量/score/子版分布,**按时间窗回溯 2005 年至今的历史帖子(正文全文,不截断)**,按帖子 URL/ID 从 Arctic Shift 拉取归档评论树或筛选/ID 批查评论正文,读取版规/wiki/subreddit 元数据,并通过 Social Hub 或自托管 Firecrawl 读取公开 About 页周访客/周贡献/周活跃用户兼容字段、Installed Apps、flair filters、通用 page_signals,以及**核验归档 score 与版规是否仍然现行**,输出结构化 JSON。
5
+ 用户提到 Reddit VOC、品牌词/竞品讨论量、subreddit 声量、帖子评论量统计、**历史评论/评论正文/评论树/thread replies**、舆情采样、批量 Reddit 链接、about/rules/wiki、版规、**历史帖子/往年数据/时间窗回溯/某年某月的帖子/长期趋势**、置顶帖/pinned/sticky/megathread、公开 About 页周指标、Installed Apps、flair filters、页面公开链接信号,或需要**建 subreddit baseline 卡/统计正文长度分布**时必须用本 skill。
6
6
  v2.12.28 起 reddit `.json` 直采已全部移除:搜索/帖子/about/版规/wiki 一律走 Arctic Shift 归档 API(`history` 模式 + Emergency Acquisition),正文不截断;归档 `score` 带 `score_status`,只有 `Backfilled`/`VerifiedLive` 可用于声量,`--verify-scores` / `--verify-rules` 用 Firecrawl 渲染页核验实时值。**置顶是 live 语义,本地已无采集途径,只能走 Social Hub**,取不到标 NotVisible,不得用归档 stickied 代替。
7
7
  **历史数据类意图直连 Arctic,不走 Hub-First**:历史帖子 / 往年数据 / 时间窗回溯 / 某年某月 / 长期趋势 / 建 subreddit baseline 卡需历史样本,一律直接路由到 `history` 模式(Arctic Shift 深度归档,权威源,正文不截断),**不经 Hub-First 四级链路**——Hub 只有近实时缓存、不持有深度归档,历史查 Hub 是浪费且拿不到。仅「当下声量」类(`keyword` / `subreddit` 及证据类模式)才走 Hub-First。
8
- Hub-First 链路(仅当下声量/证据类,不含 history):`keyword` / `subreddit` 及证据类模式先查 Social Hub(voc-query/evidence-status),缺口由 Hub 派发(dispatch-on-gap/evidence-fetch)并轮询 runs(partial 定向重派一次);仅 Hub failed/timeout/unreachable 才进入本地 Emergency Acquisition,本地成功结果自动回灌 Hub(hot-posts-create/evidence-import),coverage 满足但读不到 items 时输出 itemsUnavailable 而非假成功。不用于全站爬取、绕过登录墙或读取 mod 后台。
8
+ Hub-First 链路(仅当下声量/证据类,不含 history/comments):`keyword` / `subreddit` 及证据类模式先查 Social Hub(voc-query/evidence-status),缺口由 Hub 派发(dispatch-on-gap/evidence-fetch)并轮询 runs(partial 定向重派一次);仅 Hub failed/timeout/unreachable 才进入本地 Emergency Acquisition,本地成功结果自动回灌 Hub(hot-posts-create/evidence-import),coverage 满足但读不到 items 时输出 itemsUnavailable 而非假成功。不用于全站爬取、绕过登录墙或读取 mod 后台。
9
+ v2.12.31 新增 `comments`:默认 `/api/comments/tree` 拉取最多 25,000 条归档评论并保留层级;也支持 `/api/comments/search` 筛选及 `/api/comments/ids` 最多 500 ID 批查。该模式固定直连 Arctic,不走 Hub-First;归档通常滞后约 2 小时且无 uptime 保证,评论 `score` 不是实时值。
9
10
  ---
10
11
 
11
12
  # Reddit VOC 声量
12
13
 
13
- 按**关键词**、**subreddit** 或 **Reddit URL 列表** 采样 Reddit 公开讨论声量和公开页面证据,输出统一 JSON。适合品牌词/竞品 VOC、子版讨论量、帖子互动汇总,以及 subreddit about/rules/wiki 的 DataAcquisition。
14
+ 按**关键词**、**subreddit**、**Reddit URL 列表**或**帖子历史评论**采样 Reddit 公开讨论声量和公开页面证据,输出统一 JSON。适合品牌词/竞品 VOC、子版讨论量、帖子互动汇总、历史评论正文分析,以及 subreddit about/rules/wiki 的 DataAcquisition。
14
15
 
15
16
  ## Arctic Shift 归档后端(v2.12.27)
16
17
 
17
- `history` 模式与缺省的 `--backend arctic-shift` 走 Arctic Shift 归档 API,不经
18
+ `history` / `comments` 模式与缺省的 `--backend arctic-shift` 走 Arctic Shift 归档 API,不经
18
19
  reddit.com。**使用前读 [references/arctic-shift.md](references/arctic-shift.md)**,
19
20
  尤其是「能力边界」表:正文 / `num_comments` / `upvote_ratio` / 版规 / wiki 可用;
20
21
  **`score` 不可信**(实测约 1/3 归档行是发帖瞬间的陈旧值,逐条标 `score_status`,
@@ -32,13 +33,13 @@ reddit.com。**使用前读 [references/arctic-shift.md](references/arctic-shift
32
33
 
33
34
  **边界(勿把"近期"误判为历史)**:「最近一天 / 一周 / 一月、本周、本月、当前热度」仍是**当下声量**,走 `keyword` / `subreddit` Hub-First(Hub 侧另有 `--hub-time-window day|week|month` 覆盖近实时窗口);只有上面那种绝对日期 / 跨年 / 早于 Hub 覆盖 / 历史基线的意图才进 `history`。长期趋势应**按月 / 年分桶多次查询**,不要把单次 `--limit` 样本当成完整总体。
34
35
 
35
- `history` 模式在脚本层就固定走 `arctic_search_posts`,**不接受 `--acquisition-policy`**(`history` 子命令未注册该参数,显式传入会直接报错退出),`provenance` 恒为 `arctic_shift`。对历史意图先撞 Hub 是错误引导——Hub 会返回近实时缓存空窗或缺口,既浪费一次派发轮询又拿不到深归档。
36
+ `history` 模式在脚本层就固定走 `arctic_search_posts`;历史评论固定走 `comments`。两者都**不接受 `--acquisition-policy`**(子命令未注册该参数,显式传入会直接报错退出),`provenance` 恒为 `arctic_shift`。对历史意图先撞 Hub 是错误引导——Hub 会返回近实时缓存空窗或缺口,既浪费一次派发轮询又拿不到深归档。
36
37
 
37
38
  当下声量(`keyword` / `subreddit`)与证据类(rules/wiki/about/pinned/flair)仍按下节 Hub-First。Arctic 后端只是 `keyword` / `subreddit` 的 L3 Emergency Acquisition **实现**(证据 live 类的本地兜底是 Firecrawl,stickies 本地无有效采集路径),不改变 Hub 优先级——这些模式 Hub 可用时永远优先走 Hub。
38
39
 
39
40
  ## Hub-First 数据链路(v2.12.25,必读)
40
41
 
41
- > **例外:`history` 模式不适用 Hub-First。** 历史/时间窗/往年数据固定走 Arctic Shift 归档(见上节路由),不查 Hub、不派发、不读 acquisition-policy。本节只约束 `keyword` / `subreddit` 及证据类模式。
42
+ > **例外:`history` 与 `comments` 模式不适用 Hub-First。** 历史/时间窗/往年数据及归档评论正文固定走 Arctic Shift 归档(见上节路由),不查 Hub、不派发、不读 acquisition-policy。本节只约束 `keyword` / `subreddit` 及证据类模式。
42
43
 
43
44
  **执行当下声量 / 证据类模式前,先读 [references/hub-first.md](references/hub-first.md)** 并按 acquisition-policy(缺省 hub-first)走四级链路:**Hub Cache → Hub Dispatch → Emergency Acquisition → Hub Write-back**。要点:
44
45
 
@@ -117,6 +118,16 @@ python3 "$VOC" flair-filters -r espresso \
117
118
  python3 "$VOC" history -r BuyItForLife --after 2024-01-01 --before 2024-03-31 --pretty
118
119
  python3 "$VOC" history "noise cancelling" -r headphones -r audiophile --after 2023-01-01 --pretty
119
120
 
121
+ # 历史评论正文:默认完整树(最多 25,000),保留 parent_id/depth/tree_path/order
122
+ python3 "$VOC" comments "https://www.reddit.com/r/example/comments/abc123/slug/" --pretty
123
+
124
+ # 平面筛选(最多 100):正文关键词、作者、时间、父评论;top 表示顶层评论
125
+ python3 "$VOC" comments t3_abc123 --source search --body "battery life" \
126
+ --author example_user --after 2024-01-01 --parent-id top --sort asc --pretty
127
+
128
+ # 按评论 ID 批查(最多 500;空格或逗号分隔)
129
+ python3 "$VOC" comments --ids t1_def456 ghi789 --pretty
130
+
120
131
  # Emergency Acquisition 后端切换(缺省已是 arctic-shift)
121
132
  python3 "$VOC" keyword "PLAUD recorder" --backend arctic-shift --pretty
122
133
  python3 "$VOC" keyword "PLAUD recorder" --backend reddit-json --pretty
@@ -155,6 +166,7 @@ python3 "$VOC" links \
155
166
  | `keyword` | Hub-First 全站关键词 VOC | positional `query`;`--acquisition-policy hub-first/hub-only/self-first`;`--hub-poll-timeout-sec`;sort/time |
156
167
  | `subreddit` | Hub-First 子版内关键词 VOC | `-r/--subreddit` 可重复;同上 acquisition-policy/sort/time |
157
168
  | `links` | 批量 Reddit URL;支持帖子 permalink、subreddit about、about/rules、wiki 页面;可显式启用 Social Hub / Firecrawl 页面 fallback | URL 列表或 `--file`;`--fallback social-hub/firecrawl/auto` |
169
+ | `comments` | **v2.12.31 新增。** 从 Arctic Shift 读取帖子历史评论正文;默认树模式,另有筛选搜索和评论 ID 批查 | positional 帖子 URL/ID;`--source tree/search`;`--body` / `--author` / `--after` / `--before` / `--parent-id`;`--ids` |
158
170
  | `stickies` | 从 subreddit hot listing 低频读取 `stickied=true` 置顶帖 | `-r/--subreddit` 可重复;`--limit` 控制 hot listing 读取数量 |
159
171
  | `about-page` | 用 Social Hub 或自托管 Firecrawl 渲染公开 subreddit About 页,提取 weekly visitors、weekly contributions、weekly active users 兼容字段、公开 Installed Apps | `-r/--subreddit` 可重复;`--provider social-hub/firecrawl/auto`;`--firecrawl-base-url` |
160
172
  | `page-signals` | 用 Social Hub 或自托管 Firecrawl 渲染任意公开 Reddit 页,统一解析页面 links/Markdown 里的公开信号 | URL 列表或 `--file`;`--provider social-hub/firecrawl/auto` |
@@ -163,6 +175,35 @@ python3 "$VOC" links \
163
175
  | `firecrawl-check` | 检查本地/self-hosted Firecrawl 服务和 Docker 状态 | `--firecrawl-base-url` |
164
176
  | `social-hub-check` | 检查 Social Hub CLI scrape fallback 是否可用 | 无 |
165
177
 
178
+ ### comments 模式支持范围
179
+
180
+ `comments <post-url-or-id>` 固定直连 Arctic Shift 历史归档,不查 Hub、不派发,也不接受
181
+ `--acquisition-policy`。默认 `--source tree` 调用 `/api/comments/tree`,默认和硬上限均为
182
+ 25,000 条;输出把嵌套树扁平为 `item_type=comment`,正文 `body` 不截断,并保留原始
183
+ `link_id` / `parent_id` 以及 `depth` / `sibling_index` / `tree_order` / `tree_path` /
184
+ `child_ids`。因此下游既可直接做 VOC 文本分析,也能确定性还原层级与原顺序。
185
+
186
+ 若 Arctic 因上限或宽深度折叠返回 `kind=more`,这些节点不会伪装成评论;它们原样进入
187
+ `collapsed[]`(含 `parent_id` / `children` / `count` / 树坐标),同时
188
+ `tree_complete=false`。请求失败或响应不是合法的 Arctic `data[]` 评论树时同样输出
189
+ `ok=false`、`tree_complete=false`;只有请求成功、结构合法且 `collapsed_count=0`
190
+ 才能把当次结果视为完整树。
191
+
192
+ - `--source search` 调 `/api/comments/search`,上限 100;支持 `--body`(Arctic FTS)、
193
+ `--author`、`--after`、`--before`、`--parent-id` 与 `--sort asc|desc`。
194
+ - search 模式的 `--parent-id top` / `root` 会按 Arctic 契约显式发送空
195
+ `parent_id=`,只取顶层评论;普通值可传裸 ID 或 `t1_` fullname。tree 模式只接受
196
+ 评论 ID,用于返回该评论下的子树。
197
+ - `comments --ids ...` 调 `/api/comments/ids`,最多 500 个裸/t1_ 评论 ID,按请求顺序
198
+ 输出;缺失 ID 进入 `missing_comment_ids` 并令 `ok=false`,不会静默丢失;该模式不接受
199
+ `--source search`,避免静默忽略来源选择。
200
+ - tree/search 的 limit 越界、ID 类型错误(例如把 `t3_` 当评论 ID)、非 Reddit URL、
201
+ 不兼容参数组合都会在发网络请求前明确报错并退出 2。
202
+
203
+ Arctic 是免费第三方归档,通常滞后约 2 小时且无 uptime 保证。评论 `body` 适合历史
204
+ VOC;`score_status=ArchivedSnapshot`、`score_realtime=false` 是硬边界,评论 score
205
+ 不得写成实时热度。大批量全站处理使用 Arctic 月度 dump,不要循环刷免费 API。
206
+
166
207
  ### links 模式支持范围
167
208
 
168
209
  `links` 不再只按帖子 permalink 解析。它会按 URL 类型分流:
@@ -344,7 +385,8 @@ Social Hub / Firecrawl 只能提供公开 HTML/Markdown 页面证据;不要把
344
385
 
345
386
  脚本内置硬上限,优先避免 IP 风控:
346
387
 
347
- - 默认 `--limit 10`,硬上限 **50**(更大值自动截断)
388
+ - keyword/subreddit/links 默认 `--limit 10`,硬上限 **50**(更大值自动截断);
389
+ `comments` 使用 Arctic 官方端点上限:tree **25,000**、search **100**、IDs **500**,越界直接失败
348
390
  - 默认 `--sleep-sec 6`,每次请求后 **+1.5–4.5s** 随机抖动
349
391
  - 默认 `--max-pages 1`,单次运行最多 **20** 个 HTTP 请求
350
392
  - 遇 **429** → 立即熔断,输出 `rate_limited: true`
@@ -472,6 +514,7 @@ Cookie **仅**用于减少公开页误判,不用于私密版/验证码。输
472
514
  - `score` 是 Reddit 展示分,≠ 纯点赞数;`upvote_ratio` 更有参考价值。
473
515
  - `flair_*` 来自 Reddit post JSON 的 `link_flair_*` 字段;无 flair 时字段保持 `null` 或空数组。帖子 item 还包含公开 JSON 中可见的 `flair_template_id`、`flair_filter_query`、`author_flair_text`、`author_flair_type`、NSFW/spoiler/gallery/video 等辅助字段。
474
516
  - `item_type=post` 是帖子;`item_type=subreddit_about` / `subreddit_rules` / `wiki_page` 是非帖子公开页面证据,`score` / `num_comments` 置 0 仅为 schema 兼容。
517
+ - `item_type=comment` 来自 `comments`,`body` 为不截断的归档评论正文;`parent_id` 加树坐标可还原层级。其 `score_status=ArchivedSnapshot`、`score_realtime=false`,不得当实时热度。
475
518
  - `item_type=pinned_post` 来自 `stickies` 模式,只代表 hot listing 中公开可见的 `stickied=true` 帖子;可作为 Pinned/Sticky/Megathread 证据输入 `reddit-subreddit-compliance`。
476
519
  - `item_type=firecrawl_page` 是 Social Hub / Firecrawl 页面 fallback 结果,含 `markdown_excerpt`、`html_excerpt`、`metadata`、`page_provider` 和 `page_signals`;只在 `--fallback social-hub|firecrawl|auto` 且 JSON 抓取失败或解析不到目标内容时出现。
477
520
  - `item_type=subreddit_about_page` 是 `about-page` 的 Social Hub / Firecrawl 渲染结果,含 `weekly_visitors`、`weekly_contributions`、`weekly_active_users`、`weekly_active_users_source`、`installed_apps`、`installed_apps_status`、`flair_filters`、`page_signals`、`page_provider`,只能作为公开渲染 About 页证据。
@@ -516,8 +559,8 @@ py -3 -m pip install -r $Requirements
516
559
 
517
560
  ## 与相关 skill 的分工
518
561
 
519
- - **`reddit-link-engagement`**:单链 score/评论树
520
- - **`reddit-voc-volume`**(本 skill):关键词/subreddit/批量链接的**声量汇总 JSON**
562
+ - **`reddit-link-engagement`**:单链互动核验
563
+ - **`reddit-voc-volume`**(本 skill):关键词/subreddit/批量链接的**声量汇总 JSON**,以及 Arctic 历史评论正文/评论树
521
564
  - **`reddit-food-packaging-intel`**:餐饮包装 Reddit 批量采集落盘 CSV
522
565
 
523
566
  ## HandoffContract
@@ -544,4 +587,5 @@ py -3 -m pip install -r $Requirements
544
587
  ## 评测
545
588
 
546
589
  - Eval 定义:[`evals/evals.json`](evals/evals.json)
590
+ - 归档评论离线单测:`python3 -m unittest discover -s tests -p 'test_*.py'`
547
591
  - 工作区:`<repo>/.cursor/skills/reddit-voc-volume-workspace/`
@@ -16,6 +16,8 @@
16
16
  - public rendered subreddit About page metrics from Firecrawl `about-page`
17
17
  - visible Installed Apps links from the public rendered About page
18
18
  - post/comment sample metadata
19
+ - archived comment bodies and reconstructable thread hierarchy from Arctic Shift
20
+ (`comments`; `parent_id` plus tree coordinates; archive score is not live)
19
21
  - public subreddit rules/wiki/about/submit-text evidence
20
22
  - public pinned/stickied/megathread evidence from subreddit hot listings
21
23
 
@@ -23,6 +25,13 @@
23
25
 
24
26
  Search/subreddit/post samples are preference signals only. They must not be used as subreddit rules, account-gate proof, Installed Apps proof, or AutoModerator policy evidence.
25
27
 
28
+ `comments` bodies are historical VOC / preference evidence only. Consumers may reconstruct
29
+ the thread from `parent_id`, `depth`, `sibling_index`, `tree_order`, and `tree_path`. They must
30
+ require `tree_complete=true` before describing a result as the complete thread; otherwise
31
+ `collapsed[]` records the unexpanded IDs. Fetch failures and malformed tree payloads explicitly
32
+ set `tree_complete=false` even when `collapsed[]` is empty. `score_status=ArchivedSnapshot` and
33
+ `score_realtime=false` are binding: archived comment scores cannot support current-hot claims.
34
+
26
35
  `links` output for `subreddit_rules`, `wiki_page`, and `subreddit_about` can be used
27
36
  as public source evidence for `reddit-subreddit-compliance`. VOC still does not
28
37
  interpret compliance risk; compliance must do that.
@@ -1,6 +1,6 @@
1
- # Arctic Shift 归档后端(v2.12.27 新增)
1
+ # Arctic Shift 归档后端(v2.12.27;v2.12.31 增加评论正文)
2
2
 
3
- > **何时读本文件:** 使用 `history` 模式,或 Emergency Acquisition 落到
3
+ > **何时读本文件:** 使用 `history` / `comments` 模式,或 Emergency Acquisition 落到
4
4
  > `--backend arctic-shift`(缺省)时。判断某个字段能不能用,以本文件的
5
5
  > 「能力边界」表为准。
6
6
 
@@ -31,6 +31,7 @@ failed 403`,去掉代理直连亦为 `403 Blocked`)。Arctic Shift 不走 re
31
31
  | 能力 | 可用性 | 说明 |
32
32
  |---|---|---|
33
33
  | 正文全文 / 长度分布 / 语料 | ✅ 全程 | baseline 卡长度维度的合法来源 |
34
+ | 评论正文 / 评论树 | ✅ | `/api/comments/tree|search|ids`;保留 `parent_id`,适合历史 VOC |
34
35
  | `num_comments` | ✅ 有回填 | 可用于声量 |
35
36
  | `upvote_ratio` | ✅ 有回填 | 可用 |
36
37
  | 版规 `/api/subreddits/rules` | ✅ | 替代 `/about/rules.json` |
@@ -39,6 +40,28 @@ failed 403`,去掉代理直连亦为 `403 Blocked`)。Arctic Shift 不走 re
39
40
  | 此刻置顶 / 当前 hot | ❌ | 归档无 live 语义,走 Hub |
40
41
  | About 页周指标 / Installed Apps | ❌ | 渲染页专有,走 Social Hub / Firecrawl |
41
42
 
43
+ ### 评论端点(v2.12.31)
44
+
45
+ | 端点 | 用途 | 硬上限 |
46
+ |---|---|---:|
47
+ | `/api/comments/tree?link_id=t3_<post_id>` | 读取帖子评论树;超限/宽深度折叠返回 `kind=more` | 25,000 |
48
+ | `/api/comments/search?link_id=<post_id>` | 按 `body`、`author`、`after/before`、`parent_id` 筛选 | 100 |
49
+ | `/api/comments/ids?ids=...` | 按裸/t1_ 评论 ID 批查 | 500 |
50
+
51
+ Skill 入口统一为 `comments`。默认 tree;扁平 item 保留完整 `body`、原始
52
+ `link_id` / `parent_id`,并补 `depth` / `sibling_index` / `tree_order` /
53
+ `tree_path` / `child_ids`。`kind=more` 单独进入 `collapsed[]`;只要存在折叠节点,
54
+ `tree_complete=false`,不得声称已经拿到完整评论树。请求失败或响应不是合法
55
+ `data[]` 树时也必须 `ok=false`、`tree_complete=false`,不能因为 `collapsed=[]`
56
+ 误报完整。
57
+
58
+ `parent_id` 语义按端点区分:tree 只接受裸/t1_ 评论 ID(取子树);search 的
59
+ `--parent-id top|root` 会显式发送空 `parent_id=` 以筛顶层评论,不能伪造成 t3_ post ID。
60
+
61
+ 评论数据也受本归档约束:最新通常滞后约 2 小时,服务无 uptime 保证;评论
62
+ `score_status=ArchivedSnapshot` 且 `score_realtime=false`,只把正文用于历史 VOC,
63
+ 不要把 score 当当前热度。
64
+
42
65
  ### `score` 为什么不可信(v2.12.28 实测修正)
43
66
 
44
67
  **失真集中在新帖,不是旧帖。** 早期只抽老帖时看到的误差很小(归档 1 → 实时 1~6),
@@ -99,6 +122,13 @@ python3 "$VOC" history -r BuyItForLife --after 2024-01-01 --before 2024-03-31 --
99
122
  # 关键词 + 跨板块历史
100
123
  python3 "$VOC" history "noise cancelling" -r headphones -r audiophile --after 2023-01-01 --pretty
101
124
 
125
+ # 帖子历史评论树(默认最多 25,000)
126
+ python3 "$VOC" comments "https://www.reddit.com/r/example/comments/abc123/slug/" --pretty
127
+
128
+ # 评论正文筛选 / ID 批查
129
+ python3 "$VOC" comments abc123 --source search --body "battery life" --parent-id top --pretty
130
+ python3 "$VOC" comments --ids t1_def456 ghi789 --pretty
131
+
102
132
  # Emergency Acquisition 显式指定后端(缺省已是 arctic-shift)
103
133
  python3 "$VOC" keyword "PLAUD recorder" --backend arctic-shift --pretty
104
134
  python3 "$VOC" keyword "PLAUD recorder" --backend reddit-json --pretty # 旧的 .json + curl
@@ -112,13 +142,18 @@ python3 "$VOC" keyword "PLAUD recorder" --backend reddit-json --pretty # 旧
112
142
  - `score_status_counts`:三态计数
113
143
  - `body_stats`:`with_text` / `max_chars` / `median_chars`
114
144
  - 每条 item:`score_status`、`archived_at_post_time`、`source_fetch="arctic_shift"`
145
+ - `comments`:`retrieval_mode=tree|search|ids`、`tree_complete`、`collapsed[]`、
146
+ `body_stats`;每条评论含完整 `body`、`parent_id` 与树坐标,score 明标非实时
115
147
 
116
148
  ## 与 Hub 的关系
117
149
 
118
- **不是替代,是互补。** Hub 有 live 语义和可信 score,Arctic 有历史深度和全文。
150
+ **不是替代,是互补。** Hub 有 live 语义和可信 score,Arctic 有历史深度、帖子全文和评论正文。
119
151
  Hub-First 的四级链路不变;Arctic 只替换 L3 Emergency Acquisition 的**实现**,
120
152
  不改变它在链路中的位置——Hub 可用时永远优先走 Hub。
121
153
 
154
+ `comments` 是明确的历史归档读取例外:和 `history` 一样固定直连 Arctic,不先查 Hub、
155
+ 不派发、不接受 `--acquisition-policy`。
156
+
122
157
 
123
158
  ## 实时核验(v2.12.28)
124
159
 
@@ -2,13 +2,13 @@
2
2
 
3
3
  矩阵**当下声量 / 证据类** Reddit 公开数据消费遵循四级链路:**Hub Cache → Hub Dispatch → Emergency Acquisition → Hub Write-back**。本地采集(本 skill 的 Python 工具)正式定名 **Emergency Acquisition**——它不是与 Hub 并列的默认链路,只在 Hub 失败/离线且急用时启用。
4
4
 
5
- > **`history` 模式不在本链路内。** 历史/往年数据/时间窗回溯/长期趋势/建 baseline 卡的历史样本固定走 Arctic Shift 深度归档(权威源),**不查 Hub、不派发、不读 acquisition-policy**——Hub 只有近实时缓存、不持有深度归档,对历史意图先撞 Hub 是浪费且拿不到。历史意图请直接执行 `history` 模式,不要套用下述四级链路。
5
+ > **`history` / `comments` 模式不在本链路内。** 历史帖子与归档评论正文固定走 Arctic Shift 深度归档(权威源),**不查 Hub、不派发、不读 acquisition-policy**——Hub 只有近实时缓存、不持有深度归档。历史帖子用 `history`,帖子历史评论正文/评论树用 `comments`,不要套用下述四级链路。
6
6
 
7
7
  前置:Social Hub CLI ≥0.2.10 已安装、`auth login`、`context use`(见 SETUP.md capability probe)。
8
8
 
9
9
  ## acquisition-policy
10
10
 
11
- acquisition-policy 作用于 `keyword` / `subreddit` 以及注册了该参数的证据类模式(`stickies` / `about-page` / `page-signals` / `flair-filters`);`links` 无此参数、按 `no_hub_dispatch_path` 例外处理,而 **`history` 模式根本不注册该参数(固定 Arctic,见上方例外框),显式传入会报错**。执行者按用户指定(缺省 `hub-first`)选择策略:
11
+ acquisition-policy 作用于 `keyword` / `subreddit` 以及注册了该参数的证据类模式(`stickies` / `about-page` / `page-signals` / `flair-filters`);`links` 无此参数、按 `no_hub_dispatch_path` 例外处理,而 **`history` / `comments` 根本不注册该参数(固定 Arctic,见上方例外框),显式传入会报错**。执行者按用户指定(缺省 `hub-first`)选择策略:
12
12
 
13
13
  | 策略 | 行为 |
14
14
  |---|---|
@@ -30,7 +30,7 @@ python3 scripts/reddit_voc_volume.py subreddit "packaging leak" -r foodpackaging
30
30
  | 本 skill 模式 | 先执行的 Hub 查询 |
31
31
  |---|---|
32
32
  | search / subreddit(VOC 综合) | `social-hub intelligence voc-query --brand-keywords ... [--subreddits ...] [--need-comments] [--need-rules-evidence]`(不带 --dispatch-on-gap = 纯评估) |
33
- | 评论采样 | `social-hub intelligence comments --subreddit ... [--keyword ...] [--intent ...] [--sentiment ...]` |
33
+ | 当下/Hub 评论采样 | `social-hub intelligence comments --subreddit ... [--keyword ...] [--intent ...] [--sentiment ...]`;帖子历史归档评论改走本 skill `comments`,不走本表链路 |
34
34
  | about-page / page-signals / flair-filters / rules / wiki / pinned / installed-apps | `social-hub intelligence evidence-status --subreddit ...` → 有 fresh 的用 `evidence-list --subreddit ... --type ... --fresh-only` 取正文/信号 |
35
35
  | 热帖 | `social-hub intelligence hot-posts [--subreddit ...] [--sort ...]` |
36
36
  | links(任意 permalink 详情) | L1 查 `hot-posts`(可能未采集过而未命中);**该模式无专用 L2 派发**——未命中时允许直接 L3(fallbackReason=self_crawl_forced 不需要,记 no_hub_dispatch_path),产出必须回灌 `hot-posts-create`(帖子)/`comments-import`(评论);rules/wiki/about 类 URL 仍按证据模式走 evidence-status/evidence-fetch |
@@ -59,7 +59,7 @@ python3 scripts/reddit_voc_volume.py subreddit "packaging leak" -r foodpackaging
59
59
 
60
60
  ## L4:Hub Write-back(回灌)
61
61
 
62
- Emergency Acquisition 的产出**必须**尝试回灌。`keyword` / `subreddit` 已自动逐帖调用 `hot-posts-create`;证据型本地 item 会调用 `evidence-import`。本地脚本当前不批量采评论正文,因此只有真实存在的评论 item 才可调用 `comments-import`,不得根据帖子 `num_comments` 伪造评论:
62
+ Emergency Acquisition 的产出**必须**尝试回灌。`keyword` / `subreddit` 已自动逐帖调用 `hot-posts-create`;证据型本地 item 会调用 `evidence-import`。只有真实存在的评论 item 才可调用 `comments-import`,不得根据帖子 `num_comments` 伪造评论。注意:`comments` 是历史归档直读模式,不是 L3 Emergency Acquisition,当前输出明确为 `writebackStatus=not_applicable`,不会自动回灌 Hub:
63
63
 
64
64
  ```bash
65
65
  # 帖子 → hot-posts-create(逐条)
@@ -75,7 +75,7 @@ social-hub intelligence evidence-import -j '{"subreddit":"...","evidenceType":"r
75
75
  - `notVisible:true` 必须带 `visibilityEvidence`(页面明示不可见的原文片段),且不得与 content 同传。
76
76
  - 回灌失败:输出 `writebackStatus=failed` + 保留本地产物与重试指引(端点幂等,可安全重试);成功 `done`;Hub 离线暂存 `pending`。
77
77
 
78
- ## 输出信封(所有模式,旧字段全保留)
78
+ ## 输出信封(Hub-First 模式,旧字段全保留)
79
79
 
80
80
  最终 JSON 在原有结构上**新增**(不修改不删除既有字段,旧下游无需改解析器):
81
81
 
@@ -95,6 +95,10 @@ social-hub intelligence evidence-import -j '{"subreddit":"...","evidenceType":"r
95
95
  }
96
96
  ```
97
97
 
98
+ 固定归档模式使用来源专属信封:`history` / `comments` 的
99
+ `provenance=arctic_shift`、`runId/hubRunId=null`、`writebackStatus=not_applicable`;
100
+ `comments` 另带 `archive_caveats`,归档 score 不得解释为实时值。
101
+
98
102
  **runId 语义(v2.12.23 起)**:`runId` 只指向产出当前 items 的 run;`hubRunId` 记录本次任务涉及的派发 run(含失败/超时的),两者不可混用。`itemsUnavailable=true` 时脚本 `ok=false`,执行者应改用 `social-hub intelligence hot-posts/comments/evidence-list` 直接读 Hub 记录,不得据此转本地采集(coverage 已满足,数据在 Hub)。
99
103
 
100
104
  **v2.12.23 起 evidence 模式同样脚本内编排**:`stickies`(pinned)/`about-page`(about)/`flair-filters`(flair)与 subreddit 范围的 `page-signals`(page_signals)默认执行 `evidence-status → fresh 读 evidence-list --fresh-only → 缺口派 evidence-fetch --types → runs 轮询(共享退避/硬超时/定向重派一次)→ 仅 failed/timeout/unreachable 才本地`。非 subreddit 范围的 page-signals URL 与 `links` 一样走 `no_hub_dispatch_path` 例外(允许直接 L3+回灌)。
@@ -1,6 +1,6 @@
1
1
  #!/usr/bin/env python3
2
2
  """
3
- Reddit VOC volume sampler — keyword / subreddit / links / stickies modes.
3
+ Reddit VOC volume sampler — keyword / subreddit / links / comments / stickies modes.
4
4
  Strict rate limits; browser-like JSON requests; no bulk crawling.
5
5
  """
6
6
  from __future__ import annotations
@@ -42,6 +42,8 @@ RETRY_BACKOFF_SEC = (30.0, 120.0)
42
42
  ARCTIC_BASE_URL = "https://arctic-shift.photon-reddit.com"
43
43
  ARCTIC_MAX_LIMIT = 100 # 单次上限;文档另有 "auto",但会随服务器负载浮动,不用
44
44
  ARCTIC_IDS_MAX = 500
45
+ ARCTIC_COMMENT_TREE_MAX = 25_000
46
+ ARCTIC_COMMENT_SEARCH_MAX = 100
45
47
  ARCTIC_DEFAULT_TIMEOUT_SEC = 60.0
46
48
  ARCTIC_MAX_RETRIES = 3
47
49
  ARCTIC_MIN_SLEEP_SEC = 2.0 # 作者要求 "a couple requests per second" 以内,取更保守值
@@ -143,7 +145,7 @@ SOCIAL_HUB_REQUIRED_SKILLS = ("social-hub-cli", "social-hub-shared")
143
145
  HUB_DEFAULT_POLL_TIMEOUT_SEC = 600.0
144
146
  HUB_POLL_BACKOFF_SEC = (5.0, 10.0, 20.0, 30.0)
145
147
  HUB_INTELLIGENCE_TIMEOUT_SEC = 30.0
146
- COLLECTOR_VERSION = "2.12.30"
148
+ COLLECTOR_VERSION = "2.12.31"
147
149
  ACQUISITION_POLICIES = ("hub-first", "hub-only", "self-first")
148
150
  DEFAULT_FLAIR_FILTER_PAGES = ("", "hot", "new", "top")
149
151
 
@@ -3184,6 +3186,7 @@ def arctic_get(
3184
3186
  user_agent: str,
3185
3187
  timeout: float = ARCTIC_DEFAULT_TIMEOUT_SEC,
3186
3188
  tracker: "RequestTracker | None" = None,
3189
+ keep_empty_params: set[str] | None = None,
3187
3190
  ) -> dict[str, Any]:
3188
3191
  """调用 Arctic Shift API,处理 429 限流与重查询超时。
3189
3192
 
@@ -3191,7 +3194,12 @@ def arctic_get(
3191
3194
  `{"error": "Timeout. Maybe slow down a bit"}` —— 那是查询代价过高而非
3192
3195
  限流,同样退避重试(实测跨 15 年的 aggregate 会稳定触发)。
3193
3196
  """
3194
- clean = {k: v for k, v in params.items() if v is not None and v != ""}
3197
+ keep_empty_params = keep_empty_params or set()
3198
+ clean = {
3199
+ k: v
3200
+ for k, v in params.items()
3201
+ if v is not None and (v != "" or k in keep_empty_params)
3202
+ }
3195
3203
  url = f"{ARCTIC_BASE_URL}{path}?{urllib.parse.urlencode(clean)}"
3196
3204
  last_err = ""
3197
3205
  # 预算按**逻辑请求**计一次,不按重试次数计:重试是同一次取数的补偿,
@@ -3277,6 +3285,306 @@ def arctic_search_posts(
3277
3285
  return [arctic_post_item(r) for r in rows if isinstance(r, dict)]
3278
3286
 
3279
3287
 
3288
+ _ARCTIC_BASE36_ID_RE = re.compile(r"^[0-9a-z]+$", re.I)
3289
+
3290
+
3291
+ def normalize_reddit_post_id(target: str) -> str:
3292
+ """Extract a bare base36 post ID from a Reddit URL, ``t3_`` ID, or bare ID."""
3293
+ raw = str(target or "").strip()
3294
+ if not raw:
3295
+ raise ValueError("post URL or post ID is required (unless --ids is used)")
3296
+
3297
+ # Full Reddit permalinks and redd.it short links. Reject lookalike hosts rather
3298
+ # than accepting any URL that happens to contain `/comments/<id>`.
3299
+ if "://" in raw:
3300
+ parsed = urllib.parse.urlparse(raw)
3301
+ host = (parsed.hostname or "").lower().rstrip(".")
3302
+ if host == "redd.it" or host.endswith(".redd.it"):
3303
+ candidate = parsed.path.strip("/").split("/", 1)[0]
3304
+ elif host == "reddit.com" or host.endswith(".reddit.com"):
3305
+ match = re.search(r"/comments/([0-9a-z]+)(?:/|$)", parsed.path, re.I)
3306
+ candidate = match.group(1) if match else ""
3307
+ else:
3308
+ raise ValueError(f"unsupported post URL host: {host or '<missing>'}")
3309
+ elif "/comments/" in raw:
3310
+ match = re.search(r"/comments/([0-9a-z]+)(?:/|$)", raw, re.I)
3311
+ candidate = match.group(1) if match else ""
3312
+ else:
3313
+ candidate = raw[3:] if raw.lower().startswith("t3_") else raw
3314
+ if raw.lower().startswith("t1_"):
3315
+ raise ValueError("expected a post ID (t3_), received a comment ID (t1_)")
3316
+
3317
+ candidate = candidate.strip().lower()
3318
+ if not candidate or not _ARCTIC_BASE36_ID_RE.fullmatch(candidate):
3319
+ raise ValueError(f"malformed Reddit post ID: {raw}")
3320
+ return candidate
3321
+
3322
+
3323
+ def normalize_reddit_comment_id(raw_id: str) -> str:
3324
+ """Normalize one comment ID to its bare base36 representation."""
3325
+ raw = str(raw_id or "").strip()
3326
+ if raw.lower().startswith("t3_"):
3327
+ raise ValueError(f"expected a comment ID (t1_), received a post ID: {raw}")
3328
+ candidate = raw[3:] if raw.lower().startswith("t1_") else raw
3329
+ candidate = candidate.strip().lower()
3330
+ if not candidate or not _ARCTIC_BASE36_ID_RE.fullmatch(candidate):
3331
+ raise ValueError(f"malformed Reddit comment ID: {raw or '<empty>'}")
3332
+ return candidate
3333
+
3334
+
3335
+ def parse_reddit_comment_ids(values: list[str] | None) -> list[str]:
3336
+ """Parse comma- or whitespace-separated comment IDs, preserving first-seen order."""
3337
+ ids: list[str] = []
3338
+ seen: set[str] = set()
3339
+ for value in values or []:
3340
+ for token in str(value).split(","):
3341
+ if not token.strip():
3342
+ raise ValueError("malformed Reddit comment ID: empty value")
3343
+ comment_id = normalize_reddit_comment_id(token)
3344
+ if comment_id not in seen:
3345
+ ids.append(comment_id)
3346
+ seen.add(comment_id)
3347
+ if len(ids) > ARCTIC_IDS_MAX:
3348
+ raise ValueError(f"too many comment IDs: {len(ids)} (maximum {ARCTIC_IDS_MAX})")
3349
+ return ids
3350
+
3351
+
3352
+ def normalize_comment_parent_id(parent_id: str | None, *, allow_top_level: bool = False) -> str | None:
3353
+ """Normalize a parent comment ID; search uses an empty value for top-level comments."""
3354
+ if parent_id is None:
3355
+ return None
3356
+ raw = str(parent_id).strip()
3357
+ if raw.lower() in ("top", "root"):
3358
+ if allow_top_level:
3359
+ return ""
3360
+ raise ValueError("--parent-id top/root is only valid with --source search")
3361
+ return f"t1_{normalize_reddit_comment_id(raw)}"
3362
+
3363
+
3364
+ def arctic_comment_item(
3365
+ row: dict[str, Any],
3366
+ *,
3367
+ depth: int | None = None,
3368
+ tree_order: int | None = None,
3369
+ sibling_index: int | None = None,
3370
+ tree_path: list[int] | None = None,
3371
+ tree_parent_id: str | None = None,
3372
+ ) -> dict[str, Any]:
3373
+ """Map an Arctic comment row without truncating its archived body."""
3374
+ body = row.get("body")
3375
+ body_text = body if isinstance(body, str) else ("" if body is None else str(body))
3376
+ body_status = classify_body(body_text)
3377
+ measurable = body_status not in ("Removed", "Deleted")
3378
+ comment_id = str(row.get("id") or "").removeprefix("t1_").lower()
3379
+ permalink = str(row.get("permalink") or "")
3380
+ if permalink.startswith("/"):
3381
+ permalink = f"https://www.reddit.com{permalink}"
3382
+ item = {
3383
+ "item_type": "comment",
3384
+ "id": comment_id,
3385
+ "name": str(row.get("name") or (f"t1_{comment_id}" if comment_id else "")),
3386
+ "link_id": row.get("link_id"),
3387
+ # Preserve Arctic's original edge exactly. `tree_parent_id` separately records
3388
+ # the edge observed while traversing the nested tree, useful for consistency checks.
3389
+ "parent_id": row.get("parent_id"),
3390
+ "tree_parent_id": tree_parent_id,
3391
+ "author": row.get("author"),
3392
+ "created_utc": row.get("created_utc"),
3393
+ "score": row.get("score"),
3394
+ "score_status": "ArchivedSnapshot",
3395
+ "score_realtime": False,
3396
+ "subreddit": row.get("subreddit"),
3397
+ "body": body_text,
3398
+ "body_status": body_status,
3399
+ "body_chars": len(body_text) if measurable else 0,
3400
+ "body_words": len(re.findall(r"\S+", body_text)) if measurable else 0,
3401
+ "body_truncated": False if measurable else None,
3402
+ "permalink": permalink or None,
3403
+ "retrieved_on": row.get("retrieved_on"),
3404
+ "distinguished": row.get("distinguished"),
3405
+ "is_submitter": row.get("is_submitter"),
3406
+ "edited": row.get("edited"),
3407
+ "depth": depth,
3408
+ "tree_order": tree_order,
3409
+ "sibling_index": sibling_index,
3410
+ "tree_path": tree_path,
3411
+ "source_fetch": "arctic_shift",
3412
+ }
3413
+ return item
3414
+
3415
+
3416
+ def _arctic_tree_children(replies: Any) -> list[dict[str, Any]]:
3417
+ if replies in (None, ""):
3418
+ return []
3419
+ if not isinstance(replies, dict):
3420
+ raise ArcticError("arctic_comments_tree_malformed_payload: replies must be a listing")
3421
+ data = replies.get("data")
3422
+ if not isinstance(data, dict):
3423
+ raise ArcticError("arctic_comments_tree_malformed_payload: replies.data must be an object")
3424
+ children = data.get("children")
3425
+ if not isinstance(children, list) or any(not isinstance(child, dict) for child in children):
3426
+ raise ArcticError("arctic_comments_tree_malformed_payload: replies.data.children must be object[]")
3427
+ return children
3428
+
3429
+
3430
+ def flatten_arctic_comment_tree(payload: Any) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]:
3431
+ """Flatten Arctic's Reddit-shaped comment tree while retaining exact tree coordinates."""
3432
+ if not isinstance(payload, dict) or not isinstance(payload.get("data"), list):
3433
+ raise ArcticError("arctic_comments_tree_malformed_payload: expected data[]")
3434
+ roots = payload.get("data")
3435
+ if any(not isinstance(node, dict) for node in roots):
3436
+ raise ArcticError("arctic_comments_tree_malformed_payload: data must contain object nodes")
3437
+ nodes = roots
3438
+ items: list[dict[str, Any]] = []
3439
+ collapsed: list[dict[str, Any]] = []
3440
+ order = 0
3441
+
3442
+ def walk(children: list[dict[str, Any]], *, depth: int, path: list[int], parent_id: str | None) -> None:
3443
+ nonlocal order
3444
+ for sibling_index, node in enumerate(children):
3445
+ kind = str(node.get("kind") or "")
3446
+ data = node.get("data")
3447
+ if not isinstance(data, dict):
3448
+ raise ArcticError("arctic_comments_tree_malformed_payload: node.data must be an object")
3449
+ node_path = [*path, sibling_index]
3450
+ node_order = order
3451
+ order += 1
3452
+ if kind == "more":
3453
+ child_ids = [
3454
+ normalize_reddit_comment_id(str(value))
3455
+ for value in (data.get("children") or [])
3456
+ if str(value or "").strip()
3457
+ ]
3458
+ collapsed.append(
3459
+ {
3460
+ "kind": "more",
3461
+ "id": data.get("id"),
3462
+ "parent_id": data.get("parent_id") or parent_id,
3463
+ "count": data.get("count"),
3464
+ "children": child_ids,
3465
+ "depth": depth,
3466
+ "tree_order": node_order,
3467
+ "sibling_index": sibling_index,
3468
+ "tree_path": node_path,
3469
+ }
3470
+ )
3471
+ continue
3472
+ if kind != "t1":
3473
+ raise ArcticError(f"arctic_comments_tree_malformed_payload: unexpected kind={kind or '<empty>'}")
3474
+
3475
+ replies = _arctic_tree_children(data.get("replies"))
3476
+ item = arctic_comment_item(
3477
+ data,
3478
+ depth=depth,
3479
+ tree_order=node_order,
3480
+ sibling_index=sibling_index,
3481
+ tree_path=node_path,
3482
+ tree_parent_id=parent_id,
3483
+ )
3484
+ item["child_ids"] = [
3485
+ str(child.get("data", {}).get("id") or "").removeprefix("t1_").lower()
3486
+ for child in replies
3487
+ if child.get("kind") == "t1" and isinstance(child.get("data"), dict)
3488
+ ]
3489
+ items.append(item)
3490
+ current_parent = str(data.get("name") or (f"t1_{item['id']}" if item.get("id") else "")) or None
3491
+ walk(replies, depth=depth + 1, path=node_path, parent_id=current_parent)
3492
+
3493
+ walk(nodes, depth=0, path=[], parent_id=None)
3494
+ return items, collapsed
3495
+
3496
+
3497
+ def arctic_comments_tree(
3498
+ post_id: str,
3499
+ *,
3500
+ parent_id: str | None = None,
3501
+ limit: int = ARCTIC_COMMENT_TREE_MAX,
3502
+ start_breadth: int | None = None,
3503
+ start_depth: int | None = None,
3504
+ user_agent: str,
3505
+ tracker: "RequestTracker | None" = None,
3506
+ ) -> tuple[list[dict[str, Any]], list[dict[str, Any]]]:
3507
+ payload = arctic_get(
3508
+ "/api/comments/tree",
3509
+ {
3510
+ "link_id": f"t3_{post_id}",
3511
+ "parent_id": parent_id,
3512
+ "limit": limit,
3513
+ "start_breadth": start_breadth,
3514
+ "start_depth": start_depth,
3515
+ },
3516
+ user_agent=user_agent,
3517
+ tracker=tracker,
3518
+ )
3519
+ return flatten_arctic_comment_tree(payload)
3520
+
3521
+
3522
+ def arctic_search_comments(
3523
+ post_id: str,
3524
+ *,
3525
+ body: str | None = None,
3526
+ author: str | None = None,
3527
+ after: str | None = None,
3528
+ before: str | None = None,
3529
+ parent_id: str | None = None,
3530
+ limit: int = ARCTIC_COMMENT_SEARCH_MAX,
3531
+ sort: str = "asc",
3532
+ user_agent: str,
3533
+ tracker: "RequestTracker | None" = None,
3534
+ ) -> list[dict[str, Any]]:
3535
+ payload = arctic_get(
3536
+ "/api/comments/search",
3537
+ {
3538
+ "link_id": post_id,
3539
+ "body": body,
3540
+ "author": author,
3541
+ "after": after,
3542
+ "before": before,
3543
+ "parent_id": parent_id,
3544
+ "limit": limit,
3545
+ "sort": sort,
3546
+ },
3547
+ user_agent=user_agent,
3548
+ tracker=tracker,
3549
+ keep_empty_params={"parent_id"} if parent_id == "" else None,
3550
+ )
3551
+ rows = payload.get("data")
3552
+ if not isinstance(rows, list):
3553
+ return []
3554
+ return [
3555
+ arctic_comment_item(row, tree_order=index)
3556
+ for index, row in enumerate(rows)
3557
+ if isinstance(row, dict)
3558
+ ]
3559
+
3560
+
3561
+ def arctic_comments_by_ids(
3562
+ comment_ids: list[str],
3563
+ *,
3564
+ user_agent: str,
3565
+ tracker: "RequestTracker | None" = None,
3566
+ ) -> tuple[list[dict[str, Any]], list[str]]:
3567
+ payload = arctic_get(
3568
+ "/api/comments/ids",
3569
+ {"ids": ",".join(comment_ids)},
3570
+ user_agent=user_agent,
3571
+ tracker=tracker,
3572
+ )
3573
+ rows = payload.get("data")
3574
+ by_id = {
3575
+ str(row.get("id") or "").removeprefix("t1_").lower(): row
3576
+ for row in (rows if isinstance(rows, list) else [])
3577
+ if isinstance(row, dict)
3578
+ }
3579
+ items = [
3580
+ arctic_comment_item(by_id[comment_id], tree_order=index)
3581
+ for index, comment_id in enumerate(comment_ids)
3582
+ if comment_id in by_id
3583
+ ]
3584
+ missing = [comment_id for comment_id in comment_ids if comment_id not in by_id]
3585
+ return items, missing
3586
+
3587
+
3280
3588
  def arctic_subreddit_rules(subreddits: list[str], *, user_agent: str) -> dict[str, list[dict[str, Any]]]:
3281
3589
  """替代 /about/rules.json(本机对 reddit.com 直连被拦时唯一可用途径)。"""
3282
3590
  names = [_clean_subreddit_name(s) for s in subreddits if s]
@@ -4571,6 +4879,43 @@ def main() -> int:
4571
4879
  )
4572
4880
  pl.add_argument("--pretty", action="store_true")
4573
4881
 
4882
+ pc = sub.add_parser(
4883
+ "comments",
4884
+ help="读取帖子历史评论正文:默认 Arctic 评论树,也支持筛选搜索与评论 ID 批查",
4885
+ )
4886
+ pc.add_argument("target", nargs="?", help="Reddit post URL、t3_<post_id> 或裸 post ID;--ids 模式省略")
4887
+ pc.add_argument(
4888
+ "--source",
4889
+ choices=["tree", "search"],
4890
+ default="tree",
4891
+ help="tree=评论树(默认,最多 25,000);search=平面筛选(最多 100)",
4892
+ )
4893
+ pc.add_argument(
4894
+ "--ids",
4895
+ nargs="+",
4896
+ metavar="COMMENT_ID",
4897
+ help=f"按 t1_/裸评论 ID 批查,可用空格或逗号分隔,最多 {ARCTIC_IDS_MAX}",
4898
+ )
4899
+ pc.add_argument("--body", help="search only:评论正文关键词/FTS 查询")
4900
+ pc.add_argument("--author", help="search only:作者")
4901
+ pc.add_argument("--after", help="search only:起始时间(ISO 日期、epoch 或 Arctic offset)")
4902
+ pc.add_argument("--before", help="search only:结束时间(ISO 日期、epoch 或 Arctic offset)")
4903
+ pc.add_argument(
4904
+ "--parent-id",
4905
+ help="tree/search:父评论 ID;search 中 top/root 表示顶层评论",
4906
+ )
4907
+ pc.add_argument(
4908
+ "--limit",
4909
+ type=int,
4910
+ default=None,
4911
+ help="tree 默认/上限 25000;search 默认/上限 100;越界直接报错",
4912
+ )
4913
+ pc.add_argument("--sort", choices=["asc", "desc"], default=None, help="search only:按 created_utc 排序")
4914
+ pc.add_argument("--start-breadth", type=int, default=None, help="tree only:Arctic 折叠宽度,必须 >=0")
4915
+ pc.add_argument("--start-depth", type=int, default=None, help="tree only:Arctic 折叠深度,必须 >=0")
4916
+ pc.add_argument("--user-agent", default=DEFAULT_UA)
4917
+ pc.add_argument("--pretty", action="store_true")
4918
+
4574
4919
  ph = sub.add_parser(
4575
4920
  "history",
4576
4921
  help="历史检索:按时间窗从 Arctic Shift 归档取帖(覆盖 2005-06 起至近实时)",
@@ -4601,6 +4946,174 @@ def main() -> int:
4601
4946
  psh.add_argument("--pretty", action="store_true")
4602
4947
 
4603
4948
  args = p.parse_args()
4949
+ if args.mode == "comments":
4950
+ try:
4951
+ comment_ids = parse_reddit_comment_ids(args.ids) if args.ids else []
4952
+ if comment_ids and args.target:
4953
+ raise ValueError("provide either a post target or --ids, not both")
4954
+ if comment_ids and args.source != "tree":
4955
+ raise ValueError("--ids cannot be combined with --source search")
4956
+ if not comment_ids and not args.target:
4957
+ raise ValueError("post URL or post ID is required (unless --ids is used)")
4958
+
4959
+ filter_values = (args.body, args.author, args.after, args.before, args.sort)
4960
+ if comment_ids and (any(value is not None for value in filter_values) or args.parent_id is not None):
4961
+ raise ValueError("--ids cannot be combined with post/search/tree filters")
4962
+ if comment_ids and (args.limit is not None or args.start_breadth is not None or args.start_depth is not None):
4963
+ raise ValueError("--ids cannot be combined with --limit/--start-breadth/--start-depth")
4964
+ if args.source == "tree" and any(value is not None for value in filter_values):
4965
+ raise ValueError("--body/--author/--after/--before/--sort require --source search")
4966
+ if args.source == "search" and (args.start_breadth is not None or args.start_depth is not None):
4967
+ raise ValueError("--start-breadth/--start-depth are only valid with --source tree")
4968
+ if args.start_breadth is not None and args.start_breadth < 0:
4969
+ raise ValueError("--start-breadth must be >= 0")
4970
+ if args.start_depth is not None and args.start_depth < 0:
4971
+ raise ValueError("--start-depth must be >= 0")
4972
+
4973
+ post_id = normalize_reddit_post_id(args.target) if args.target else None
4974
+ parent_id = (
4975
+ normalize_comment_parent_id(args.parent_id, allow_top_level=args.source == "search")
4976
+ if args.parent_id is not None and post_id is not None
4977
+ else None
4978
+ )
4979
+ retrieval_mode = "ids" if comment_ids else args.source
4980
+ maximum = ARCTIC_COMMENT_TREE_MAX if retrieval_mode == "tree" else ARCTIC_COMMENT_SEARCH_MAX
4981
+ default_limit = ARCTIC_COMMENT_TREE_MAX if retrieval_mode == "tree" else ARCTIC_COMMENT_SEARCH_MAX
4982
+ limit = default_limit if args.limit is None else args.limit
4983
+ if retrieval_mode != "ids" and not 1 <= limit <= maximum:
4984
+ raise ValueError(f"--limit must be between 1 and {maximum} for {retrieval_mode} mode")
4985
+ except ValueError as exc:
4986
+ pc.error(str(exc))
4987
+
4988
+ tracker = RequestTracker()
4989
+ errors: list[str] = []
4990
+ collapsed: list[dict[str, Any]] = []
4991
+ missing_ids: list[str] = []
4992
+ items: list[dict[str, Any]] = []
4993
+ tree_fetch_succeeded = False
4994
+ try:
4995
+ if retrieval_mode == "ids":
4996
+ items, missing_ids = arctic_comments_by_ids(
4997
+ comment_ids,
4998
+ user_agent=args.user_agent,
4999
+ tracker=tracker,
5000
+ )
5001
+ elif retrieval_mode == "search":
5002
+ items = arctic_search_comments(
5003
+ post_id,
5004
+ body=args.body,
5005
+ author=args.author,
5006
+ after=args.after,
5007
+ before=args.before,
5008
+ parent_id=parent_id,
5009
+ limit=limit,
5010
+ sort=args.sort or "asc",
5011
+ user_agent=args.user_agent,
5012
+ tracker=tracker,
5013
+ )
5014
+ else:
5015
+ items, collapsed = arctic_comments_tree(
5016
+ post_id,
5017
+ parent_id=parent_id,
5018
+ limit=limit,
5019
+ start_breadth=args.start_breadth,
5020
+ start_depth=args.start_depth,
5021
+ user_agent=args.user_agent,
5022
+ tracker=tracker,
5023
+ )
5024
+ tree_fetch_succeeded = True
5025
+ except ArcticError as exc:
5026
+ errors.append(f"arctic_comments_failed:{exc}"[:300])
5027
+ if missing_ids:
5028
+ errors.append(f"comment_ids_not_found:{','.join(missing_ids)}"[:300])
5029
+
5030
+ bodies = [item for item in items if item.get("body_status") == "Present"]
5031
+ removed = [item for item in items if item.get("body_status") in ("Removed", "Deleted")]
5032
+ notes = [
5033
+ "Arctic Shift 是第三方历史归档源,通常滞后约 2 小时,且不保证可用性。",
5034
+ "评论 body 适合历史 VOC;score 是归档快照而非实时值,不得用于当前热度判断。",
5035
+ ]
5036
+ if collapsed:
5037
+ notes.append(
5038
+ f"评论树含 {len(collapsed)} 个 kind=more 折叠节点;"
5039
+ "tree_complete=false,collapsed[].children 保留尚未展开的评论 ID。"
5040
+ )
5041
+ fetched_at = utc_now_iso()
5042
+ canonical_source = f"https://www.reddit.com/comments/{post_id}/" if post_id else None
5043
+ fingerprint = hashlib.sha256(
5044
+ json.dumps(
5045
+ {
5046
+ "mode": "comments",
5047
+ "retrieval_mode": retrieval_mode,
5048
+ "post_id": post_id,
5049
+ "comment_ids": comment_ids,
5050
+ "body": args.body,
5051
+ "author": args.author,
5052
+ "after": args.after,
5053
+ "before": args.before,
5054
+ "parent_id": parent_id,
5055
+ "limit": None if retrieval_mode == "ids" else limit,
5056
+ "sort": args.sort or ("asc" if retrieval_mode == "search" else None),
5057
+ "start_breadth": args.start_breadth,
5058
+ "start_depth": args.start_depth,
5059
+ },
5060
+ ensure_ascii=False,
5061
+ sort_keys=True,
5062
+ ).encode("utf-8")
5063
+ ).hexdigest()
5064
+ out = {
5065
+ "ok": not errors,
5066
+ "mode": "comments",
5067
+ "retrieval_mode": retrieval_mode,
5068
+ "provenance": "arctic_shift",
5069
+ "source_endpoint": f"/api/comments/{'ids' if retrieval_mode == 'ids' else retrieval_mode}",
5070
+ "source_urls": [canonical_source] if canonical_source else [],
5071
+ "post_id": post_id,
5072
+ "link_id": f"t3_{post_id}" if post_id else None,
5073
+ "requested_comment_ids": comment_ids,
5074
+ "missing_comment_ids": missing_ids,
5075
+ "filters": {
5076
+ "body": args.body,
5077
+ "author": args.author,
5078
+ "after": args.after,
5079
+ "before": args.before,
5080
+ "parent_id": parent_id,
5081
+ "sort": args.sort or ("asc" if retrieval_mode == "search" else None),
5082
+ },
5083
+ "count": len(items),
5084
+ "tree_complete": tree_fetch_succeeded and not collapsed if retrieval_mode == "tree" else None,
5085
+ "collapsed_count": len(collapsed),
5086
+ "collapsed": collapsed,
5087
+ "body_stats": {
5088
+ "with_text": len(bodies),
5089
+ "removed_or_deleted": len(removed),
5090
+ "max_chars": max((int(item.get("body_chars") or 0) for item in bodies), default=0),
5091
+ },
5092
+ "items": items,
5093
+ "archive_caveats": {
5094
+ "estimated_lag": "about_2_hours",
5095
+ "uptime_guaranteed": False,
5096
+ "score_realtime": False,
5097
+ },
5098
+ "rate_limit": {
5099
+ "requested_limit": len(comment_ids) if retrieval_mode == "ids" else limit,
5100
+ "effective_limit": len(comment_ids) if retrieval_mode == "ids" else limit,
5101
+ "http_requests": tracker.count,
5102
+ "rate_limited": tracker.rate_limited,
5103
+ },
5104
+ "fetched_at": fetched_at,
5105
+ "fetchedAt": fetched_at,
5106
+ "fingerprint": fingerprint,
5107
+ "runId": None,
5108
+ "hubRunId": None,
5109
+ "fallbackReason": None,
5110
+ "writebackStatus": "not_applicable",
5111
+ "notes": notes,
5112
+ "errors": errors,
5113
+ }
5114
+ print(json.dumps(out, ensure_ascii=False, indent=2 if args.pretty else None))
5115
+ return 0 if out["ok"] else 1
5116
+
4604
5117
  if args.mode == "history":
4605
5118
  subs = [_clean_subreddit_name(s) for s in (args.subreddit or [])] or [None]
4606
5119
  items: list[dict[str, Any]] = []
@@ -0,0 +1,189 @@
1
+ from __future__ import annotations
2
+
3
+ import contextlib
4
+ import importlib.util
5
+ import io
6
+ import json
7
+ from pathlib import Path
8
+ import sys
9
+ import unittest
10
+ from unittest import mock
11
+
12
+
13
+ SCRIPT = Path(__file__).resolve().parents[1] / "scripts" / "reddit_voc_volume.py"
14
+ SPEC = importlib.util.spec_from_file_location("reddit_voc_volume_under_test", SCRIPT)
15
+ assert SPEC is not None and SPEC.loader is not None
16
+ voc = importlib.util.module_from_spec(SPEC)
17
+ SPEC.loader.exec_module(voc)
18
+
19
+
20
+ class ArchivedCommentTests(unittest.TestCase):
21
+ def run_main(self, *args: str) -> tuple[int, dict]:
22
+ stdout = io.StringIO()
23
+ with mock.patch.object(sys, "argv", [str(SCRIPT), *args]), contextlib.redirect_stdout(stdout):
24
+ code = voc.main()
25
+ return code, json.loads(stdout.getvalue())
26
+
27
+ def test_post_and_comment_id_normalization(self) -> None:
28
+ self.assertEqual(
29
+ voc.normalize_reddit_post_id("https://www.reddit.com/r/example/comments/AbC123/a_slug/"),
30
+ "abc123",
31
+ )
32
+ self.assertEqual(voc.normalize_reddit_post_id("t3_ABC123"), "abc123")
33
+ self.assertEqual(voc.normalize_reddit_post_id("https://redd.it/AbC123"), "abc123")
34
+ self.assertEqual(voc.normalize_reddit_comment_id("t1_DeF456"), "def456")
35
+ with self.assertRaisesRegex(ValueError, "comment ID"):
36
+ voc.normalize_reddit_post_id("t1_def456")
37
+ with self.assertRaisesRegex(ValueError, "unsupported post URL host"):
38
+ voc.normalize_reddit_post_id("https://example.com/comments/abc123")
39
+ with self.assertRaisesRegex(ValueError, "maximum 500"):
40
+ voc.parse_reddit_comment_ids([f"c{index:x}" for index in range(501)])
41
+
42
+ def test_tree_mode_flattens_body_and_preserves_hierarchy(self) -> None:
43
+ payload = {
44
+ "data": [
45
+ {
46
+ "kind": "t1",
47
+ "data": {
48
+ "id": "root1",
49
+ "name": "t1_root1",
50
+ "link_id": "t3_post1",
51
+ "parent_id": "t3_post1",
52
+ "author": "alice",
53
+ "body": "Root comment",
54
+ "created_utc": 100,
55
+ "score": 9,
56
+ "subreddit": "example",
57
+ "replies": {
58
+ "data": {
59
+ "children": [
60
+ {
61
+ "kind": "t1",
62
+ "data": {
63
+ "id": "child1",
64
+ "name": "t1_child1",
65
+ "link_id": "t3_post1",
66
+ "parent_id": "t1_root1",
67
+ "author": "bob",
68
+ "body": "Child comment",
69
+ "created_utc": 101,
70
+ "score": 3,
71
+ "subreddit": "example",
72
+ },
73
+ },
74
+ {
75
+ "kind": "more",
76
+ "data": {
77
+ "id": "more1",
78
+ "parent_id": "t1_root1",
79
+ "count": 2,
80
+ "children": ["hidden1", "hidden2"],
81
+ },
82
+ },
83
+ ]
84
+ }
85
+ },
86
+ },
87
+ }
88
+ ]
89
+ }
90
+ with mock.patch.object(voc, "arctic_get", return_value=payload) as get:
91
+ code, out = self.run_main("comments", "https://reddit.com/r/example/comments/post1/slug")
92
+
93
+ self.assertEqual(code, 0)
94
+ self.assertEqual(get.call_args.args[0], "/api/comments/tree")
95
+ self.assertEqual(get.call_args.args[1]["link_id"], "t3_post1")
96
+ self.assertEqual(get.call_args.args[1]["limit"], 25_000)
97
+ self.assertEqual([item["id"] for item in out["items"]], ["root1", "child1"])
98
+ self.assertEqual(out["items"][0]["child_ids"], ["child1"])
99
+ self.assertEqual(out["items"][1]["parent_id"], "t1_root1")
100
+ self.assertEqual(out["items"][1]["depth"], 1)
101
+ self.assertEqual(out["items"][1]["tree_path"], [0, 0])
102
+ self.assertEqual(out["items"][1]["body"], "Child comment")
103
+ self.assertFalse(out["items"][1]["body_truncated"])
104
+ self.assertFalse(out["tree_complete"])
105
+ self.assertEqual(out["collapsed"][0]["children"], ["hidden1", "hidden2"])
106
+ self.assertFalse(out["archive_caveats"]["score_realtime"])
107
+
108
+ def test_search_mode_forwards_supported_filters(self) -> None:
109
+ with mock.patch.object(voc, "arctic_get", return_value={"data": []}) as get:
110
+ code, out = self.run_main(
111
+ "comments",
112
+ "t3_post1",
113
+ "--source",
114
+ "search",
115
+ "--body",
116
+ "battery life",
117
+ "--author",
118
+ "alice",
119
+ "--after",
120
+ "2024-01-01",
121
+ "--before",
122
+ "2024-02-01",
123
+ "--parent-id",
124
+ "top",
125
+ "--limit",
126
+ "10",
127
+ "--sort",
128
+ "desc",
129
+ )
130
+
131
+ self.assertEqual(code, 0)
132
+ self.assertEqual(out["retrieval_mode"], "search")
133
+ self.assertEqual(get.call_args.args[0], "/api/comments/search")
134
+ params = get.call_args.args[1]
135
+ self.assertEqual(params["link_id"], "post1")
136
+ self.assertEqual(params["body"], "battery life")
137
+ self.assertEqual(params["author"], "alice")
138
+ self.assertEqual(params["parent_id"], "")
139
+ self.assertEqual(get.call_args.kwargs["keep_empty_params"], {"parent_id"})
140
+ self.assertEqual(params["limit"], 10)
141
+ self.assertEqual(params["sort"], "desc")
142
+
143
+ def test_ids_mode_restores_requested_order(self) -> None:
144
+ payload = {
145
+ "data": [
146
+ {"id": "second2", "body": "second", "link_id": "t3_post1", "parent_id": "t3_post1"},
147
+ {"id": "first1", "body": "first", "link_id": "t3_post1", "parent_id": "t3_post1"},
148
+ ]
149
+ }
150
+ with mock.patch.object(voc, "arctic_get", return_value=payload) as get:
151
+ code, out = self.run_main("comments", "--ids", "t1_first1,second2")
152
+
153
+ self.assertEqual(code, 0)
154
+ self.assertEqual(get.call_args.args[0], "/api/comments/ids")
155
+ self.assertEqual(get.call_args.args[1]["ids"], "first1,second2")
156
+ self.assertEqual([item["id"] for item in out["items"]], ["first1", "second2"])
157
+
158
+ def test_tree_failure_and_malformed_payload_are_never_reported_complete(self) -> None:
159
+ with mock.patch.object(voc, "arctic_get", side_effect=voc.ArcticError("upstream unavailable")):
160
+ code, out = self.run_main("comments", "post1")
161
+ self.assertEqual(code, 1)
162
+ self.assertFalse(out["tree_complete"])
163
+ self.assertIn("upstream unavailable", out["errors"][0])
164
+
165
+ with mock.patch.object(voc, "arctic_get", return_value={"data": {}}):
166
+ code, out = self.run_main("comments", "post1")
167
+ self.assertEqual(code, 1)
168
+ self.assertFalse(out["tree_complete"])
169
+ self.assertIn("arctic_comments_tree_malformed_payload", out["errors"][0])
170
+
171
+ def test_malformed_id_and_out_of_range_limits_fail_before_network(self) -> None:
172
+ with mock.patch.object(voc, "arctic_get") as get, contextlib.redirect_stderr(io.StringIO()):
173
+ with mock.patch.object(sys, "argv", [str(SCRIPT), "comments", "--ids", "t3_wrongkind"]):
174
+ with self.assertRaisesRegex(SystemExit, "2"):
175
+ voc.main()
176
+ with mock.patch.object(sys, "argv", [str(SCRIPT), "comments", "post1", "--limit", "25001"]):
177
+ with self.assertRaisesRegex(SystemExit, "2"):
178
+ voc.main()
179
+ with mock.patch.object(sys, "argv", [str(SCRIPT), "comments", "post1", "--parent-id", "top"]):
180
+ with self.assertRaisesRegex(SystemExit, "2"):
181
+ voc.main()
182
+ with mock.patch.object(sys, "argv", [str(SCRIPT), "comments", "--ids", "comment1", "--source", "search"]):
183
+ with self.assertRaisesRegex(SystemExit, "2"):
184
+ voc.main()
185
+ get.assert_not_called()
186
+
187
+
188
+ if __name__ == "__main__":
189
+ unittest.main()