@acedatacloud/skills 2026.726.7 → 2026.726.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@acedatacloud/skills",
3
- "version": "2026.726.7",
3
+ "version": "2026.726.9",
4
4
  "description": "Agent Skills for AceDataCloud AI services — music, image, video generation, LLM chat, web search. Compatible with Claude Code, GitHub Copilot, Gemini CLI, OpenAI Codex, and 30+ AI coding agents.",
5
5
  "keywords": [
6
6
  "agent-skills",
@@ -0,0 +1,89 @@
1
+ ---
2
+ name: cto51
3
+ description: Read the connected 51CTO 博客 (blog.51cto.com) account and create Markdown article drafts with the user's own login cookies (BYOC). Use when the user mentions 51CTO, wants to save a 51CTO draft, or asks who their connected 51CTO account is.
4
+ when_to_use: |
5
+ Trigger for the user's 51CTO 博客 account driven by their own login cookie:
6
+ show the connected account, or turn Markdown into a 51CTO article draft.
7
+ The write API creates a draft, so this skill stops there and hands the user
8
+ the editor URL. Writes are gated behind explicit confirmation.
9
+ connections: [cto51]
10
+ allowed_tools: [Bash]
11
+ license: Apache-2.0
12
+ metadata:
13
+ author: acedatacloud
14
+ version: "1.0"
15
+ ---
16
+
17
+ # cto51 — read & draft on 51CTO 博客 via your own cookies
18
+
19
+ Drives the user's **real** 51CTO account through the same `blog.51cto.com`
20
+ endpoints the site uses, authenticated by the login cookie they captured with
21
+ the ACE extension. No browser, no third-party deps — just `urllib`.
22
+
23
+ The connector injects the cookie jar as a JSON env var `$CTO51_COOKIES`. Never
24
+ print it.
25
+
26
+ ```bash
27
+ python3 "$SKILL_DIR/scripts/cto51.py" whoami
28
+ ```
29
+
30
+ If `$SKILL_DIR` points at a different skill loaded in the same turn, resolve
31
+ this skill's directory explicitly before running the commands below.
32
+
33
+ ## Important: draft only
34
+
35
+ 51CTO's write endpoint creates a **draft**. This skill returns the draft's
36
+ editor URL and does not publish. Tell the user plainly that they must open that
37
+ URL and publish themselves — do not claim the article went live.
38
+
39
+ ## Verify the connection first
40
+
41
+ ```bash
42
+ python3 "$SKILL_DIR/scripts/cto51.py" whoami
43
+ ```
44
+
45
+ If this fails with a redirect or auth error, the cookie has expired. Ask the
46
+ user to reconnect at `https://auth.acedata.cloud/user/connections` rather than
47
+ retrying.
48
+
49
+ ## Create a draft — GATED
50
+
51
+ Prepare the complete Markdown in a file. The first call is always a dry run and
52
+ does not write anything.
53
+
54
+ ```bash
55
+ # Dry run — shows exactly what would be written.
56
+ python3 "$SKILL_DIR/scripts/cto51.py" draft \
57
+ --title "标题" --content-file /tmp/article.md --tags "python,api"
58
+
59
+ # Actually create the draft after the user confirms.
60
+ python3 "$SKILL_DIR/scripts/cto51.py" draft \
61
+ --title "标题" --content-file /tmp/article.md --tags "python,api" --confirm
62
+ ```
63
+
64
+ Options: `--content-file <path.md>` (preferred) or `--content "<markdown>"` for
65
+ short inline text; `--tags "a,b"` comma-separated; `--abstract "…"` sets the
66
+ summary shown in listings. The dry run echoes every field that will be written,
67
+ including the abstract — show that output to the user before confirming.
68
+
69
+ `--confirm` is valid only as the final argument. Show the title, tags and full
70
+ content to the user before writing.
71
+
72
+ ## Gotchas
73
+
74
+ - 51CTO sits behind a WAF that answers a bare request with HTTP 567. The CLI
75
+ always sends a full browser fingerprint, so do not strip its headers or
76
+ re-implement the calls with plain `curl`.
77
+ - Both the identity and the `_csrf` token come from the publish page. A
78
+ redirect there means the session is dead — reconnect, do not retry.
79
+ - Content is sent as Markdown (`is_old=0`). Do not pre-render it to HTML.
80
+ - Images referenced by external URL are not re-hosted. If the source host
81
+ blocks hotlinking they will not render; mention this when the article has
82
+ images.
83
+ - Do not retry a timed-out write automatically — the outcome may be unknown and
84
+ a retry can create a duplicate draft.
85
+
86
+ ## Record the output
87
+
88
+ This skill only produces drafts, so do **not** call `publish_artifact`. Report
89
+ the returned `draft_id` and `edit_url` to the user instead.
@@ -0,0 +1,406 @@
1
+ #!/usr/bin/env python3
2
+ """
3
+ cto51 — read & draft on 51CTO 博客 (blog.51cto.com) with the user's own login
4
+ cookies (BYOC). Standard-library only (urllib), no third-party deps, so it runs
5
+ in the bare sandbox without an image change.
6
+
7
+ The connector injects the user's cookie jar as a JSON env var ``CTO51_COOKIES``
8
+ — a list of ``{name, value, domain, ...}`` dicts captured by the ACE extension.
9
+ (The namespace is `cto51`, not `51cto`, because `51CTO_COOKIES` would not be a
10
+ valid shell identifier.)
11
+
12
+ Read commands run directly. ``draft`` is GATED: without a trailing ``--confirm``
13
+ it only dry-runs. ``--confirm`` is honored ONLY as the last argument.
14
+
15
+ NOTE: this creates a DRAFT and returns its editor URL; the user finishes
16
+ publishing in the 51CTO editor.
17
+
18
+ Examples:
19
+ python3 cto51.py whoami
20
+ python3 cto51.py draft --title T --content-file a.md --confirm
21
+ """
22
+
23
+ from __future__ import annotations
24
+
25
+ import argparse
26
+ import gzip
27
+ import http.client
28
+ import json
29
+ import os
30
+ import re
31
+ import socket
32
+ import sys
33
+ import urllib.error
34
+ import urllib.parse
35
+ import urllib.request
36
+
37
+ UA = (
38
+ "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 "
39
+ "(KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36"
40
+ )
41
+ PLATFORM = "cto51"
42
+ BASE = "https://blog.51cto.com"
43
+ PUBLISH_PAGE = f"{BASE}/blogger/publish"
44
+ # Cap on the FORM-ENCODED body (urlencode expands CJK ~3x).
45
+ MAX_ENCODED_BYTES = 8 * 1024 * 1024
46
+
47
+ # Bounded quantifiers so a hostile page cannot cause catastrophic backtracking.
48
+ _CSRF_RE = re.compile(r'<meta\s{1,10}name="csrf-token"\s{1,10}content="([^"]{1,500})"')
49
+ _USER_RE = re.compile(
50
+ r'<li class="more user">\s{0,50}<a[^>]{0,300}href="([^"]{1,300})"[^>]{0,300}>'
51
+ r'\s{0,50}<img[^>]{0,300}src="([^"]{1,300})"'
52
+ )
53
+
54
+ _RAW = sys.argv[1:]
55
+ CONFIRM = bool(_RAW) and _RAW[-1] == "--confirm"
56
+ ARGV = _RAW[:-1] if CONFIRM else list(_RAW)
57
+
58
+
59
+ class _NoRedirect(urllib.request.HTTPRedirectHandler):
60
+ """Refuse redirects outright.
61
+
62
+ add_unredirected_header only protects the Cookie; urllib copies every other
63
+ header (here the `_csrf` form token's companion headers) onto the redirected
64
+ request, so a 30x to a foreign host could hand them over.
65
+ """
66
+
67
+ def redirect_request(self, req, fp, code, msg, headers, newurl):
68
+ return None
69
+
70
+
71
+ _OPENER = urllib.request.build_opener(_NoRedirect())
72
+
73
+
74
+ def out(obj) -> None:
75
+ print(json.dumps(obj, ensure_ascii=False, indent=2, default=str))
76
+
77
+
78
+ def die(msg: str, code: int = 1) -> None:
79
+ out({"error": msg})
80
+ sys.exit(code)
81
+
82
+
83
+ # 51CTO error/WAF pages carry the site chrome, including the csrf-token meta —
84
+ # never echo a raw response body without stripping it. Covers both the HTML meta
85
+ # form (name="csrf-token" content="…") and any JSON/inline "csrfToken":"…".
86
+ _CSRF_LEAK_RES = (
87
+ # name="csrf-token" … content="TOKEN" (and the reversed attribute order)
88
+ re.compile(r'(name=["\']csrf[-_]?token["\'][^>]{0,400}?content=["\'])([^"\']{0,500})', re.I),
89
+ re.compile(r'(content=["\'])([^"\']{0,500})(["\'][^>]{0,400}?name=["\']csrf[-_]?token)', re.I),
90
+ # "_csrf": "TOKEN" / _csrf=TOKEN / csrfToken: 'TOKEN' — 51CTO runs Yii, whose
91
+ # CSRF parameter is literally `_csrf`, so `token` must be OPTIONAL here.
92
+ re.compile(r'(_?csrf[-_]?(?:token)?["\']?\s{0,5}[:=]\s{0,5}["\']?)([^"\'&,\s>;]{1,500})', re.I),
93
+ )
94
+
95
+
96
+ def _redact(text: str) -> str:
97
+ for rx in _CSRF_LEAK_RES:
98
+ text = rx.sub(
99
+ (lambda mo: mo.group(1) + "<redacted>" + (mo.group(3) if mo.lastindex and mo.lastindex >= 3 else "")),
100
+ text,
101
+ )
102
+ return text
103
+
104
+
105
+
106
+ # ── Cookie jar (shared pattern across the cookie-BYOC skills) ────────
107
+
108
+ def load_cookies() -> list:
109
+ env = f"{PLATFORM.upper()}_COOKIES"
110
+ raw = os.environ.get(env)
111
+ if not raw:
112
+ die(f"{env} is not set — connect 51CTO at "
113
+ f"https://auth.acedata.cloud/user/connections, then retry.")
114
+ try:
115
+ jar = json.loads(raw)
116
+ except json.JSONDecodeError as e:
117
+ die(f"{env} is not valid JSON: {e}")
118
+ if not isinstance(jar, list):
119
+ die(f"{env} must be a JSON list of cookies, got {type(jar).__name__}")
120
+ return jar
121
+
122
+
123
+ def _domain_matches(host: str, domain: str) -> bool:
124
+ d = domain.lstrip(".").lower()
125
+ h = host.lower()
126
+ return not d or h == d or h.endswith("." + d)
127
+
128
+
129
+ def cookie_header(jar: list, url: str) -> str:
130
+ host = urllib.parse.urlsplit(url).hostname or ""
131
+ host_in_scope = any(
132
+ c.get("domain") and _domain_matches(host, str(c["domain"])) for c in jar
133
+ )
134
+ parts = []
135
+ for c in jar:
136
+ name, value = c.get("name"), c.get("value")
137
+ if not name or value is None:
138
+ continue
139
+ domain = c.get("domain")
140
+ if domain:
141
+ if not _domain_matches(host, str(domain)):
142
+ continue
143
+ elif not host_in_scope:
144
+ continue
145
+ # http.client raises ValueError("Invalid header value %r" % value) at
146
+ # send time for CR/LF, and UnicodeEncodeError for non-latin1 — and the
147
+ # exception text carries the WHOLE Cookie header. Reject here, with a
148
+ # message that never echoes a value.
149
+ pair = f"{name}={value}"
150
+ if any(ch in pair for ch in "\r\n\x00"):
151
+ die("a cookie in the jar contains a line break and cannot be sent — "
152
+ "reconnect at https://auth.acedata.cloud/user/connections.")
153
+ try:
154
+ pair.encode("latin-1")
155
+ except UnicodeEncodeError:
156
+ die("a cookie in the jar contains characters that cannot be sent in "
157
+ "an HTTP header — reconnect at "
158
+ "https://auth.acedata.cloud/user/connections.")
159
+ parts.append(pair)
160
+ return "; ".join(parts)
161
+
162
+
163
+ def request(method: str, url: str, jar: list, *, headers=None, form=None,
164
+ write: bool = False):
165
+ # 51CTO's WAF 567s a bare request; it needs a full browser fingerprint
166
+ # (same treatment the csdn skill needs).
167
+ hdrs = {
168
+ "User-Agent": UA,
169
+ "Accept": "text/html,application/xhtml+xml,application/xml;q=0.9,*/*;q=0.8",
170
+ "Accept-Language": "zh-CN,zh;q=0.9,en;q=0.8",
171
+ "sec-ch-ua": '"Chromium";v="131", "Not_A Brand";v="24"',
172
+ "sec-ch-ua-mobile": "?0",
173
+ "sec-ch-ua-platform": '"macOS"',
174
+ "Sec-Fetch-Dest": "document",
175
+ "Sec-Fetch-Mode": "navigate",
176
+ "Sec-Fetch-Site": "same-origin",
177
+ "Origin": BASE,
178
+ "Referer": PUBLISH_PAGE,
179
+ }
180
+ if headers:
181
+ hdrs.update(headers)
182
+ data = None
183
+ if form is not None:
184
+ data = urllib.parse.urlencode(form).encode("utf-8")
185
+ hdrs.setdefault("Content-Type", "application/x-www-form-urlencoded; charset=UTF-8")
186
+ req = urllib.request.Request(url, data=data, headers=hdrs, method=method)
187
+ # Unredirected → the cookie is not re-sent if the API 30x-redirects to a
188
+ # different host (e.g. a login page), so the jar never leaks off-site.
189
+ req.add_unredirected_header("Cookie", cookie_header(jar, url))
190
+ try:
191
+ with _OPENER.open(req, timeout=30) as resp:
192
+ raw = resp.read()
193
+ if resp.headers.get("Content-Encoding") == "gzip":
194
+ raw = gzip.decompress(raw)
195
+ return resp.status, raw.decode("utf-8", "replace")
196
+ except urllib.error.HTTPError as e:
197
+ if e.code in (301, 302, 303, 307, 308):
198
+ die("51CTO redirected the request — not followed, so no credential "
199
+ "left 51cto.com. You are most likely logged out; reconnect at "
200
+ "https://auth.acedata.cloud/user/connections."
201
+ + (" This was a WRITE: its outcome is UNKNOWN — check your "
202
+ "drafts before retrying." if write else ""))
203
+ # Draining the error body can itself raise (IncompleteRead / reset).
204
+ # Degrade to an empty body so a truncated 5xx still flows into the
205
+ # normal non-JSON path, which carries the write-UNKNOWN wording.
206
+ try:
207
+ raw = e.read()
208
+ if e.headers.get("Content-Encoding") == "gzip":
209
+ raw = gzip.decompress(raw)
210
+ except Exception:
211
+ raw = b""
212
+ return e.code, raw.decode("utf-8", "replace")
213
+ # URLError subclasses OSError, so it must come first. The receive phase
214
+ # (getresponse) raises TimeoutError/HTTPException OUTSIDE urllib's own
215
+ # OSError wrapper, so a write whose reply never lands would otherwise
216
+ # escape as a bare traceback with no JSON at all.
217
+ except (urllib.error.URLError, TimeoutError, socket.timeout,
218
+ http.client.HTTPException, gzip.BadGzipFile, OSError) as e:
219
+ if write:
220
+ die(f"51CTO write did not return a result ({type(e).__name__}: {e}); "
221
+ f"the outcome is UNKNOWN. Check your 51CTO drafts before "
222
+ f"retrying so you do not create a duplicate.")
223
+ die(f"network error reaching {url}: {type(e).__name__}: {e}")
224
+
225
+
226
+ def publish_page(jar: list) -> tuple[str, dict]:
227
+ """Fetch the publish page — it carries both the identity and the CSRF token."""
228
+ status, html = request("GET", PUBLISH_PAGE, jar)
229
+ if status in (401, 403):
230
+ die("auth failed — cookie expired or invalid. Reconnect at "
231
+ "https://auth.acedata.cloud/user/connections.")
232
+ if status != 200:
233
+ die(f"unexpected status {status} loading the 51CTO publish page")
234
+ m = _USER_RE.search(html)
235
+ if not m:
236
+ # The csrf-token meta also appears on logged-out pages, so it is not a
237
+ # session signal. Distinguish "definitely logged out" from "markup
238
+ # changed" instead of always blaming the cookie.
239
+ if "/user/login" in html or "home.51cto.com/login" in html:
240
+ die("not logged in to 51CTO — reconnect at "
241
+ "https://auth.acedata.cloud/user/connections.")
242
+ die("could not confirm the 51CTO session from the publish page — either "
243
+ "the cookie expired or 51CTO changed its markup. Try reconnecting "
244
+ "at https://auth.acedata.cloud/user/connections; if that does not "
245
+ "help, this skill needs updating.")
246
+ link, avatar = m.group(1), m.group(2)
247
+ csrf_m = _CSRF_RE.search(html)
248
+ if not csrf_m:
249
+ die("could not read the 51CTO CSRF token from the publish page; "
250
+ "reconnect and retry.")
251
+ uid = link.rstrip("/").split("/")[-1]
252
+ return csrf_m.group(1), {"user_id": uid, "url": link, "avatar": avatar}
253
+
254
+
255
+ # ── commands ────────────────────────────────────────────────────────
256
+
257
+ def cmd_whoami(jar, _args):
258
+ _csrf, who = publish_page(jar)
259
+ out({"platform": PLATFORM, **who})
260
+
261
+
262
+ def read_content(args) -> str:
263
+ content = args.content
264
+ if args.content_file:
265
+ try:
266
+ with open(args.content_file, encoding="utf-8") as f:
267
+ content = f.read()
268
+ except OSError as e:
269
+ die(f"cannot read --content-file: {e}")
270
+ if content is None:
271
+ die("provide --content-file <path.md> or --content <markdown>")
272
+ # The body is form-urlencoded, which expands each CJK byte to %XX (~3x), so
273
+ # cap the ENCODED size — a raw-byte cap would let a 10 MiB Chinese article
274
+ # become a ~30 MB POST that 51CTO rejects opaquely mid-write.
275
+ encoded = len(urllib.parse.quote_plus(content))
276
+ if encoded > MAX_ENCODED_BYTES:
277
+ die(f"content is too large: {encoded} bytes once form-encoded "
278
+ f"(limit {MAX_ENCODED_BYTES}). Split the article or trim it.")
279
+ return content
280
+
281
+
282
+ def cmd_draft(jar, args):
283
+ if not args.title:
284
+ die("--title is required")
285
+ content = read_content(args)
286
+ tags = ",".join(t.strip() for t in (args.tags or "").split(",") if t.strip())
287
+
288
+ if not CONFIRM:
289
+ out({
290
+ "dry_run": True, "command": "draft", "platform": PLATFORM,
291
+ "title": args.title, "tags": tags,
292
+ "abstract": args.abstract or "",
293
+ "content_characters": len(content),
294
+ "note": "51CTO content is Markdown (is_old=0). Re-run with --confirm "
295
+ "as the LAST argument to actually create the draft. This "
296
+ "creates a DRAFT — finish publishing in the 51CTO editor.",
297
+ })
298
+ return
299
+
300
+ csrf, _who = publish_page(jar)
301
+ # We hold the exact token, so scrub it by value too — airtight regardless of
302
+ # how 51CTO frames it in an error page (the patterns are defence in depth).
303
+ def scrub(text: str) -> str:
304
+ return _redact(text.replace(csrf, "<redacted>") if csrf else text)
305
+ form = {
306
+ "title": args.title,
307
+ "content": content,
308
+ "cate_id": "",
309
+ "custom_id": "0",
310
+ "tag": tags,
311
+ "abstract": args.abstract or "",
312
+ "banner_type": "0",
313
+ "blog_type": "1",
314
+ "copy_code": "1",
315
+ "is_hide": "0",
316
+ "top_time": "0",
317
+ "is_comment": "0",
318
+ "is_old": "0", # 0 = Markdown
319
+ "blog_id": "",
320
+ "pid": "",
321
+ "did": "",
322
+ "work_id": "",
323
+ "class_id": "",
324
+ "subjectId": "",
325
+ "import_type": "-1",
326
+ "invite_code": "",
327
+ "raffle": "",
328
+ "orig": "",
329
+ "_csrf": csrf,
330
+ }
331
+ status, text = request("POST", f"{BASE}/blogger/draft", jar, write=True, form=form,
332
+ headers={
333
+ "X-Requested-With": "XMLHttpRequest",
334
+ "Accept": "application/json, text/javascript, */*; q=0.01",
335
+ "Sec-Fetch-Dest": "empty",
336
+ "Sec-Fetch-Mode": "cors",
337
+ })
338
+ try:
339
+ res = json.loads(text)
340
+ except json.JSONDecodeError:
341
+ die(f"51CTO write returned a non-JSON response ({status}); the outcome "
342
+ f"is UNKNOWN. Check your drafts before retrying. "
343
+ f"Body: {scrub(text)[:200]}")
344
+ # The endpoint really does answer `"data": []` (a list) for the logged-out
345
+ # case, so never dereference it without checking the shape first.
346
+ if not isinstance(res, dict):
347
+ die(f"51CTO write returned an unexpected response shape ({status}); the "
348
+ f"outcome is UNKNOWN. Check your drafts before retrying.")
349
+ data = res.get("data")
350
+ if res.get("status") != 1 or not isinstance(data, dict):
351
+ # Normalize to JSON, never repr(): a dict/list `msg` rendered with
352
+ # repr() uses single quotes, which _redact's double-quote-anchored
353
+ # patterns cannot match.
354
+ raw_msg = res.get("msg")
355
+ detail = (raw_msg if isinstance(raw_msg, str)
356
+ else json.dumps(raw_msg if raw_msg is not None else res,
357
+ ensure_ascii=False))
358
+ die(f"draft creation failed: {scrub(detail)[:200]}")
359
+ # Only an ASCII-numeric id is a real draft handle. str.isdigit() alone is
360
+ # Unicode-wide ('١٢٣', '²' all pass) and would build a plausible-but-wrong
361
+ # edit_url while reporting ok=true.
362
+ did = str(data.get("did") or "").strip()
363
+ if not re.fullmatch(r"[0-9]{1,20}", did):
364
+ # Redact BEFORE slicing so a token can never be half-exposed by the cut.
365
+ die(f"draft creation returned an unusable id {scrub(did)[:80]!r}; the "
366
+ f"outcome is UNKNOWN — check your 51CTO drafts before retrying. "
367
+ f"{scrub(json.dumps(res, ensure_ascii=False))[:200]}")
368
+ out({
369
+ "ok": True,
370
+ "draft_only": True,
371
+ "draft_id": str(did),
372
+ "edit_url": f"{BASE}/blogger/draft/{did}",
373
+ "note": "Draft saved. Open edit_url to review and publish.",
374
+ })
375
+
376
+
377
+ COMMANDS = {
378
+ "whoami": cmd_whoami,
379
+ "draft": cmd_draft,
380
+ }
381
+
382
+
383
+ def main() -> None:
384
+ p = argparse.ArgumentParser(prog="cto51.py", description="51CTO 博客 cookie CLI")
385
+ sub = p.add_subparsers(dest="command", required=True)
386
+ sub.add_parser("whoami", help="show the logged-in account")
387
+ sp = sub.add_parser("draft", help="create a draft article (GATED by trailing --confirm)")
388
+ sp.add_argument("--title")
389
+ sp.add_argument("--content", help="Markdown content inline")
390
+ sp.add_argument("--content-file", help="path to a Markdown file")
391
+ sp.add_argument("--tags", help="comma-separated tag names")
392
+ sp.add_argument("--abstract", help="short summary shown in listings")
393
+ args = p.parse_args(ARGV)
394
+ jar = load_cookies()
395
+ COMMANDS[args.command](jar, args)
396
+
397
+
398
+ if __name__ == "__main__":
399
+ # The agent consuming this CLI parses stdout as JSON — never let an
400
+ # unexpected exception escape as a bare traceback with empty stdout.
401
+ try:
402
+ main()
403
+ except SystemExit:
404
+ raise
405
+ except BaseException as exc: # noqa: BLE001
406
+ die(f"unexpected {type(exc).__name__}: {_redact(str(exc))[:200]}")
@@ -0,0 +1,123 @@
1
+ ---
2
+ name: toutiao
3
+ description: Read and publish on 今日头条 / Toutiao (mp.toutiao.com) with the user's own login cookies (BYOC) — list their 头条号 articles with impression/read/comment stats, inspect one article, and publish a new 图文 article or draft. Use when the user mentions 今日头条, 头条号, Toutiao, "我的头条文章", reading their article stats (展现/阅读), or 发头条 / publishing to Toutiao.
4
+ when_to_use: |
5
+ Trigger for anything on the user's 今日头条号 (mp.toutiao.com) account driven by
6
+ their own login cookie: show who they are, list their articles with impression /
7
+ read / comment counts, look at one article's stats, or publish a new article.
8
+ This acts as the user's real account, so writes are gated behind an explicit
9
+ confirmation.
10
+ connections: [toutiao]
11
+ allowed_tools: [Bash]
12
+ license: Apache-2.0
13
+ metadata:
14
+ author: acedatacloud
15
+ version: "1.0"
16
+ ---
17
+
18
+ # toutiao — read & publish on 今日头条 via your own cookies
19
+
20
+ Drives the user's **real** 头条号 through the same `mp.toutiao.com` creator APIs
21
+ the web console uses, authenticated by the login cookie they captured with the
22
+ ACE extension. No browser, no third-party deps — just `urllib`.
23
+
24
+ The connector injects the cookie jar as an env var:
25
+
26
+ - `TOUTIAO_COOKIES` — a JSON array of cookies. **Secret — never echo or print
27
+ it.** The CLI reads it for you.
28
+
29
+ > Writes echo the `csrftoken` cookie back as the `X-CSRFToken` header (the CLI
30
+ > does this). Reads and writes are otherwise cookie-only — no request signing.
31
+
32
+ ## CLI
33
+
34
+ The skill ships [`scripts/toutiao.py`](scripts/toutiao.py) — self-contained, stdlib only.
35
+
36
+ ```sh
37
+ # $SKILL_DIR can point at another skill loaded this turn — anchor on our own
38
+ # script, and re-run this at the top of every Bash block (fresh shell each time).
39
+ TT="$SKILL_DIR/scripts/toutiao.py"; [ -f "$TT" ] || TT=$(find /tmp -maxdepth 8 -path '*/skills/*/scripts/toutiao.py' 2>/dev/null | head -1)
40
+ [ -f "$TT" ] || { echo "toutiao script not found (SKILL_DIR=$SKILL_DIR)" >&2; exit 1; }
41
+ python3 "$TT" whoami # who is logged in (+ total article count)
42
+ python3 "$TT" articles --limit 20 # my articles + stats
43
+ python3 "$TT" articles --status draft # only drafts
44
+ python3 "$TT" article <pgc-id> # one article's stats
45
+ ```
46
+
47
+ Stats come straight from 头条: `impression_count` (展现), `read_count` (阅读),
48
+ `comment_count` (评论), `digg_count` (点赞).
49
+
50
+ `--status` accepts `all` (default) / `draft` / `published` / `reviewing` / `failed`.
51
+
52
+ ## Verify the connection first
53
+
54
+ ```sh
55
+ TT="$SKILL_DIR/scripts/toutiao.py"; [ -f "$TT" ] || TT=$(find /tmp -maxdepth 8 -path '*/skills/*/scripts/toutiao.py' 2>/dev/null | head -1)
56
+ python3 "$TT" whoami
57
+ # → {"user_id": ..., "name": "...", "articles_total": 273}
58
+ ```
59
+
60
+ On an auth error the cookie is expired — tell the user to reconnect at
61
+ <https://auth.acedata.cloud/user/connections>. Do **not** retry in a loop.
62
+
63
+ ## Publishing — GATED (dry-run unless trailing `--confirm`)
64
+
65
+ `publish` writes to the user's real 头条号. Content is **Markdown** (converted to
66
+ HTML for 头条's body field). Without a trailing `--confirm` it dry-runs.
67
+ `--confirm` is honored **only as the last argument**. Always show the dry-run,
68
+ get an explicit "yes", then re-run with `--confirm` last.
69
+
70
+ ```sh
71
+ TT="$SKILL_DIR/scripts/toutiao.py"; [ -f "$TT" ] || TT=$(find /tmp -maxdepth 8 -path '*/skills/*/scripts/toutiao.py' 2>/dev/null | head -1)
72
+ python3 "$TT" publish --title "标题" --content-file a.md # dry-run
73
+ python3 "$TT" publish --title "标题" --content-file a.md --draft-only --confirm # private draft
74
+ python3 "$TT" publish --title "标题" --content-file a.md --confirm # PUBLIC, enters 审核
75
+ ```
76
+
77
+ - `--draft-only` saves a private draft (`save=1`) — safe, nothing public.
78
+ - Without `--draft-only` the article is **submitted publicly** under the user's
79
+ name and enters 头条's 审核 queue. Default to `--draft-only` unless the user
80
+ clearly asked to go live.
81
+ - **Titles must be 2–30 characters** — 头条 rejects anything outside that range
82
+ (the CLI fails early with a clear message).
83
+
84
+ ## Images
85
+
86
+ 头条 rejects the **entire article** (`7115 图片uri非法`) if any `<img>` points at
87
+ a non-头条 URL — so external images cannot simply be left alone. `publish`
88
+ uploads every image in the body to 头条's own CDN first and rewrites the tag with
89
+ the CDN attributes 头条 requires.
90
+
91
+ If an image fails to upload, `publish` **aborts and posts nothing**, listing the
92
+ offending URLs — it will not silently publish the user's article with images
93
+ missing. Pass `--drop-failed-images` to publish without them instead.
94
+ `--no-rehost-images` skips the whole step (头条 will then reject the article
95
+ unless the body already carries 头条-hosted images).
96
+
97
+ ## Gotchas — surface before the user is surprised
98
+
99
+ - **This is the user's real 头条号.** Confirm before any publish.
100
+ - **审核**: a published article is not instantly live — 头条 reviews it. The
101
+ returned URL goes live once it passes; a rejected article shows up under
102
+ `articles --status failed`.
103
+ - **Daily publish cap**: 头条 caps 图文 posts per day. Hitting it fails the
104
+ publish with 头条's own message — relay it, don't retry in a loop.
105
+ - **14-day edit window**: 头条 refuses edits to articles published more than 14
106
+ days ago, so this skill does not offer an edit command.
107
+ - **Cookie expiry**: reconnect at auth.acedata.cloud/user/connections.
108
+ - **Never print `TOUTIAO_COOKIES`** — it is full account access.
109
+ - **ToS**: cookie automation acts only on the user's own account with their own
110
+ captured cookie; the user owns that risk.
111
+
112
+ ## Record the output
113
+
114
+ After you successfully publish and obtain the live result URL, call the built-in
115
+ `publish_artifact` tool ONCE so the user can track this deliverable in **My Outputs**:
116
+
117
+ ```
118
+ publish_artifact(kind="article", channel="toutiao", title="<title>", url="<the REAL returned URL>", status="delivered")
119
+ ```
120
+
121
+ Use the real returned URL — never fabricate one. Call it once per published item,
122
+ only after delivery is confirmed; skip it (or use `status="failed"`) if publishing failed.
123
+ See `_shared/artifacts.md`.
@@ -0,0 +1,737 @@
1
+ #!/usr/bin/env python3
2
+ """
3
+ toutiao — read & publish on 今日头条 (mp.toutiao.com) with the user's own login
4
+ cookies (BYOC). Standard-library only (urllib), no third-party deps, so it runs
5
+ in the bare sandbox without an image change.
6
+
7
+ The connector injects the user's cookie jar as a JSON env var ``TOUTIAO_COOKIES``
8
+ — a list of ``{name, value, domain, ...}`` dicts captured by the ACE extension.
9
+
10
+ 头条's creator APIs are cookie-only (no request signing); writes additionally
11
+ need the ``csrftoken`` cookie echoed back as the ``X-CSRFToken`` header.
12
+
13
+ Read commands run directly. ``publish`` is GATED: without a trailing
14
+ ``--confirm`` it only dry-runs. ``--confirm`` is honored ONLY as the last
15
+ argument, so a title/content that merely contains "--confirm" can never silently
16
+ go live. ``--draft-only`` stops after saving a private draft.
17
+
18
+ Examples:
19
+ python3 toutiao.py whoami
20
+ python3 toutiao.py articles --limit 20
21
+ python3 toutiao.py articles --status draft
22
+ python3 toutiao.py article <pgc-id>
23
+ python3 toutiao.py publish --title T --content-file a.md --draft-only --confirm
24
+ """
25
+
26
+ from __future__ import annotations
27
+
28
+ import argparse
29
+ import gzip
30
+ import html as _html
31
+ import ipaddress
32
+ import json
33
+ import os
34
+ import random
35
+ import re
36
+ import socket
37
+ import sys
38
+ import urllib.error
39
+ import urllib.parse
40
+ import urllib.request
41
+ from html.parser import HTMLParser as _HTMLParser
42
+
43
+ UA = (
44
+ "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 "
45
+ "(KHTML, like Gecko) Chrome/124.0.0.0 Safari/537.36"
46
+ )
47
+ PLATFORM = "toutiao"
48
+ MP = "https://mp.toutiao.com"
49
+
50
+ _RAW = sys.argv[1:]
51
+ CONFIRM = bool(_RAW) and _RAW[-1] == "--confirm"
52
+ ARGV = _RAW[:-1] if CONFIRM else list(_RAW)
53
+
54
+
55
+ def out(obj) -> None:
56
+ print(json.dumps(obj, ensure_ascii=False, indent=2, default=str))
57
+
58
+
59
+ def die(msg: str, code: int = 1) -> None:
60
+ out({"error": msg})
61
+ sys.exit(code)
62
+
63
+
64
+ # ── Cookie jar (shared pattern across the cookie-BYOC skills) ────────
65
+
66
+ def load_cookies() -> list:
67
+ env = f"{PLATFORM.upper()}_COOKIES"
68
+ raw = os.environ.get(env)
69
+ if not raw:
70
+ die(f"{env} is not set — connect 今日头条 at "
71
+ f"https://auth.acedata.cloud/user/connections, then retry.")
72
+ try:
73
+ jar = json.loads(raw)
74
+ except json.JSONDecodeError as e:
75
+ die(f"{env} is not valid JSON: {e}")
76
+ if not isinstance(jar, list):
77
+ die(f"{env} must be a JSON list of cookies, got {type(jar).__name__}")
78
+ return jar
79
+
80
+
81
+ def _domain_matches(host: str, domain: str) -> bool:
82
+ d = domain.lstrip(".").lower()
83
+ h = host.lower()
84
+ return not d or h == d or h.endswith("." + d)
85
+
86
+
87
+ def cookie_header(jar: list, url: str) -> str:
88
+ host = urllib.parse.urlsplit(url).hostname or ""
89
+ host_in_scope = any(
90
+ c.get("domain") and _domain_matches(host, str(c["domain"])) for c in jar
91
+ )
92
+ parts = []
93
+ for c in jar:
94
+ name, value = c.get("name"), c.get("value")
95
+ if not name or value is None:
96
+ continue
97
+ domain = c.get("domain")
98
+ if domain:
99
+ if not _domain_matches(host, str(domain)):
100
+ continue
101
+ elif not host_in_scope:
102
+ continue
103
+ parts.append(f"{name}={value}")
104
+ return "; ".join(parts)
105
+
106
+
107
+ def cookie_value(jar: list, name: str):
108
+ for c in jar:
109
+ if c.get("name") == name:
110
+ return c.get("value")
111
+ return None
112
+
113
+
114
+ # ── HTTP ────────────────────────────────────────────────────────────
115
+
116
+ def _headers(jar: list, referer: str) -> dict:
117
+ return {
118
+ "User-Agent": UA,
119
+ "Accept": "application/json, text/plain, */*",
120
+ "Accept-Language": "zh-CN,zh;q=0.9,en;q=0.8",
121
+ "Referer": referer,
122
+ "Origin": MP,
123
+ "sec-ch-ua": '"Chromium";v="124", "Google Chrome";v="124", "Not-A.Brand";v="99"',
124
+ "sec-ch-ua-mobile": "?0",
125
+ "sec-ch-ua-platform": '"macOS"',
126
+ "Sec-Fetch-Dest": "empty",
127
+ "Sec-Fetch-Mode": "cors",
128
+ "Sec-Fetch-Site": "same-origin",
129
+ }
130
+
131
+
132
+ def request(method: str, url: str, jar: list, *, referer, headers=None, body=None,
133
+ nonfatal=False):
134
+ hdrs = _headers(jar, referer)
135
+ if headers:
136
+ hdrs.update(headers)
137
+ data = body.encode("utf-8") if isinstance(body, str) else body
138
+ req = urllib.request.Request(url, data=data, headers=hdrs, method=method)
139
+ # Unredirected → urllib will NOT re-send these if the API 30x-redirects to a
140
+ # different host (e.g. a login page), so neither the jar nor the CSRF token
141
+ # (which is itself a cookie value) leaks off-site.
142
+ req.add_unredirected_header("Cookie", cookie_header(jar, url))
143
+ # Writes are rejected without the csrftoken cookie echoed as a header.
144
+ req.add_unredirected_header("X-CSRFToken", str(cookie_value(jar, "csrftoken") or ""))
145
+ try:
146
+ with urllib.request.urlopen(req, timeout=60) as resp:
147
+ raw = resp.read()
148
+ if resp.headers.get("Content-Encoding") == "gzip":
149
+ raw = gzip.decompress(raw)
150
+ return resp.status, raw.decode("utf-8", "replace")
151
+ except urllib.error.HTTPError as e:
152
+ raw = e.read()
153
+ try:
154
+ if e.headers.get("Content-Encoding") == "gzip":
155
+ raw = gzip.decompress(raw)
156
+ except Exception:
157
+ pass
158
+ return e.code, raw.decode("utf-8", "replace")
159
+ except urllib.error.URLError as e:
160
+ if nonfatal:
161
+ raise RuntimeError(f"network error reaching {url}: {e.reason}")
162
+ die(f"network error reaching {url}: {e.reason}")
163
+
164
+
165
+ def api_envelope(method: str, path: str, jar: list, *, referer=None, body=None, headers=None,
166
+ nonfatal=False):
167
+ """Call an mp.toutiao.com endpoint and return the raw {code, message, ...}
168
+ envelope, dying on a non-zero code. 头条 answers HTTP 200 even for logical
169
+ failures, so the envelope code is the real status."""
170
+ url = f"{MP}{path}"
171
+ ref = referer or f"{MP}/profile_v4/graphic/articles"
172
+ status, text = request(method, url, jar, referer=ref, body=body, headers=headers,
173
+ nonfatal=nonfatal)
174
+ if status in (401, 403) or "/auth/page/login" in text:
175
+ msg = (f"auth failed ({status}) on {path} — cookie likely expired. "
176
+ f"Reconnect 今日头条 at https://auth.acedata.cloud/user/connections.")
177
+ if nonfatal:
178
+ raise RuntimeError(msg)
179
+ die(msg)
180
+ try:
181
+ env = json.loads(text)
182
+ except json.JSONDecodeError:
183
+ if nonfatal:
184
+ raise RuntimeError(f"non-JSON response ({status}) from {path}: {text[:200]}")
185
+ die(f"non-JSON response ({status}) from {path}: {text[:300]}")
186
+ if not isinstance(env, dict):
187
+ if nonfatal:
188
+ raise RuntimeError(f"unexpected response from {path}: {text[:200]}")
189
+ die(f"unexpected response from {path}: {text[:300]}")
190
+ # `.get("code", …)` would return a stored None instead of falling back to
191
+ # err_no, so test explicitly.
192
+ code = env.get("code")
193
+ if code is None:
194
+ code = env.get("err_no")
195
+ if code not in (0, None):
196
+ msg = env.get("message") or env.get("reason") or ""
197
+ if code in (401, 403) or "登录" in str(msg):
198
+ auth_msg = (f"auth failed (code={code}: {msg}) — cookie likely expired. "
199
+ f"Reconnect at https://auth.acedata.cloud/user/connections.")
200
+ if nonfatal:
201
+ raise RuntimeError(auth_msg)
202
+ die(auth_msg)
203
+ if nonfatal:
204
+ raise RuntimeError(f"头条 API error on {path} (code={code}): {msg}")
205
+ die(f"头条 API error on {path} (code={code}): {msg}")
206
+ return env
207
+
208
+
209
+ def api(method: str, path: str, jar: list, **kw):
210
+ """Same as api_envelope but unwraps the `data` payload."""
211
+ return api_envelope(method, path, jar, **kw).get("data")
212
+
213
+
214
+ # ── commands ────────────────────────────────────────────────────────
215
+
216
+ def media_info(jar):
217
+ d = api("GET", "/mp/agw/media/get_media_info/", jar, referer=f"{MP}/profile_v4/index")
218
+ if not isinstance(d, dict):
219
+ die("could not read 头条号 profile (cookie expired?)")
220
+ return d
221
+
222
+
223
+ def cmd_whoami(jar, _args):
224
+ d = media_info(jar)
225
+ media = d.get("media") or {}
226
+ user = d.get("user") or {}
227
+ _, total = _list_page(jar, "all", 1, 1)
228
+ out({
229
+ "user_id": user.get("id"),
230
+ "media_id": media.get("id"),
231
+ "name": user.get("screen_name") or media.get("display_name"),
232
+ "url": f"https://www.toutiao.com/c/user/token/{media.get('id')}/"
233
+ if media.get("id") else None,
234
+ "articles_total": total,
235
+ "media_level": media.get("article_media_level"),
236
+ })
237
+
238
+
239
+ # status → 头条's own filter value. `all` also returns drafts (is_draft=1).
240
+ _STATUS = {"all": "all", "draft": "draft", "published": "published",
241
+ "reviewing": "verifying", "failed": "unpass"}
242
+
243
+
244
+ def _list_page(jar, status, page, size, nonfatal=False):
245
+ q = urllib.parse.urlencode({
246
+ "status": _STATUS.get(status, "all"), "from_time": 0, "start_time": 0,
247
+ "end_time": 0, "search_word": "", "page": page, "size": size,
248
+ })
249
+ d = api("GET", f"/mp/agw/article/list/?{q}", jar, nonfatal=nonfatal) or {}
250
+ return d.get("content") or [], d.get("total")
251
+
252
+
253
+ def _fmt(a: dict) -> dict:
254
+ pgc = a.get("pgc_id") or a.get("item_id") or a.get("id")
255
+ return {
256
+ "pgc_id": str(pgc) if pgc is not None else None,
257
+ "title": a.get("title"),
258
+ "url": a.get("article_url"),
259
+ "is_draft": bool(a.get("is_draft")),
260
+ "status_desc": a.get("status_desc"),
261
+ "impression_count": a.get("impression_count"),
262
+ # go_detail_count_v2 is the read count; show_go_detail_count is a
263
+ # display-toggle BOOLEAN, never a fallback for it.
264
+ "read_count": a.get("go_detail_count_v2"),
265
+ "comment_count": a.get("comment_count"),
266
+ "digg_count": a.get("digg_count"),
267
+ "create_time": a.get("create_time"),
268
+ }
269
+
270
+
271
+ def _iter_articles(jar, status, hard_cap=2000):
272
+ page, seen = 1, 0
273
+ while seen < hard_cap:
274
+ items, total = _list_page(jar, status, page, 50)
275
+ if not items:
276
+ return
277
+ for it in items:
278
+ yield it
279
+ seen += 1
280
+ if total is not None and seen >= total:
281
+ return
282
+ page += 1
283
+
284
+
285
+ def cmd_articles(jar, args):
286
+ items = []
287
+ for it in _iter_articles(jar, args.status):
288
+ items.append(it)
289
+ if len(items) >= args.limit:
290
+ break
291
+ _, total = _list_page(jar, args.status, 1, 1)
292
+ out({"total": total, "count": len(items), "status": args.status,
293
+ "articles": [_fmt(a) for a in items]})
294
+
295
+
296
+ def cmd_article(jar, args):
297
+ # 头条 has no per-article detail endpoint for the creator (article/edit only
298
+ # works within 14 days of publishing), but the list already carries every
299
+ # stat, so resolve by scanning it.
300
+ for it in _iter_articles(jar, "all"):
301
+ if str(it.get("pgc_id")) == str(args.id) or str(it.get("item_id")) == str(args.id):
302
+ res = _fmt(it)
303
+ res["abstract"] = (it.get("abstract") or "")[:200]
304
+ res["word_count"] = it.get("content_word_cnt")
305
+ out(res)
306
+ return
307
+ die(f"article {args.id} not found among your 头条 articles")
308
+
309
+
310
+ # ── image re-host (头条 rejects the whole article if an <img> is external) ──
311
+
312
+ MAX_IMG_BYTES = 12 * 1024 * 1024
313
+ _IMG_SKIP = ("toutiao.com", "byteimg.com", "pstatp.com", "toutiaoimg.com")
314
+
315
+
316
+ class _ImgFinder(_HTMLParser):
317
+ """Locate <img> tags with the stdlib parser instead of a regex.
318
+
319
+ Four regex attempts each traded one unparseable shape for another; the
320
+ ambiguity is inherent (`title="unclosed>` is indistinguishable from an
321
+ attribute containing `>`). HTMLParser resolves tag boundaries by the same
322
+ rules the browser and 头条 use, so what we rewrite is what they render.
323
+ """
324
+
325
+ def __init__(self):
326
+ # convert_charrefs=False → attribute values stay as authored, matching
327
+ # what we re-emit into the HTML body.
328
+ super().__init__(convert_charrefs=False)
329
+ self.spans = [] # (start_offset, end_offset, attrs)
330
+
331
+ def _record(self, tag, attrs):
332
+ if tag.lower() != "img":
333
+ return
334
+ text = self.get_starttag_text() or ""
335
+ line, col = self.getpos()
336
+ start = self._line_starts[line - 1] + col
337
+ self.spans.append((start, start + len(text), attrs))
338
+
339
+ handle_starttag = _record
340
+ handle_startendtag = _record
341
+
342
+ def find(self, html):
343
+ # getpos() is (line, col); precompute line offsets to map to an index.
344
+ self._line_starts, pos = [0], 0
345
+ for ln in html.splitlines(keepends=True):
346
+ pos += len(ln)
347
+ self._line_starts.append(pos)
348
+ self.feed(html)
349
+ self.close()
350
+ return self.spans
351
+
352
+
353
+ def _first_attr(attrs, name):
354
+ """First occurrence wins, as browsers do with a duplicated attribute."""
355
+ for k, v in attrs:
356
+ if k.lower() == name:
357
+ return v or ""
358
+ return None
359
+
360
+
361
+ class _NoRedirect(urllib.request.HTTPRedirectHandler):
362
+ # Refuse redirects on image fetches — a 30x could otherwise reach an internal
363
+ # host that _assert_public_url() never saw (SSRF).
364
+ def redirect_request(self, req, fp, code, msg, headers, newurl):
365
+ raise RuntimeError(f"image redirect blocked ({code}) -> {newurl[:80]}")
366
+
367
+
368
+ _IMG_OPENER = urllib.request.build_opener(_NoRedirect)
369
+
370
+
371
+ def _assert_public_url(url):
372
+ parts = urllib.parse.urlsplit(url)
373
+ if parts.scheme not in ("http", "https") or not parts.hostname:
374
+ raise RuntimeError(f"unsupported image URL: {url[:80]}")
375
+ try:
376
+ addrs = socket.getaddrinfo(parts.hostname, None)
377
+ except OSError as e:
378
+ raise RuntimeError(f"cannot resolve {parts.hostname}: {e}")
379
+ for info in addrs:
380
+ ip = ipaddress.ip_address(info[4][0])
381
+ if (ip.is_private or ip.is_loopback or ip.is_link_local
382
+ or ip.is_reserved or ip.is_multicast or ip.is_unspecified):
383
+ raise RuntimeError(f"blocked non-public image host: {parts.hostname}")
384
+
385
+
386
+ def _download_image(url):
387
+ _assert_public_url(url)
388
+ req = urllib.request.Request(url, headers={"User-Agent": UA})
389
+ with _IMG_OPENER.open(req, timeout=30) as r:
390
+ data = r.read(MAX_IMG_BYTES + 1)
391
+ if len(data) > MAX_IMG_BYTES:
392
+ raise RuntimeError(f"image exceeds {MAX_IMG_BYTES} bytes")
393
+ return data
394
+
395
+
396
+ def _same_or_sub(host, suffix):
397
+ return host == suffix or host.endswith("." + suffix)
398
+
399
+
400
+ def _ext_of(url, default="png"):
401
+ tail = url.rsplit("/", 1)[-1].split("?")[0]
402
+ if "." in tail:
403
+ e = tail.rsplit(".", 1)[-1].lower()
404
+ if e in ("jpg", "jpeg", "png", "gif", "webp"):
405
+ return e
406
+ return default
407
+
408
+
409
+ def _multipart(field, filename, blob, ctype):
410
+ boundary = "----acedata" + "".join(random.choice("0123456789abcdef") for _ in range(20))
411
+ body = b"".join([
412
+ (f'--{boundary}\r\nContent-Disposition: form-data; name="{field}";'
413
+ f' filename="{filename}"\r\nContent-Type: {ctype}\r\n\r\n').encode(),
414
+ blob,
415
+ f"\r\n--{boundary}--\r\n".encode(),
416
+ ])
417
+ return body, boundary
418
+
419
+
420
+ def upload_image(jar, src) -> dict:
421
+ """Upload one image to 头条's CDN, returning its {url, web_uri, width, ...}."""
422
+ img = _download_image(src)
423
+ ext = _ext_of(src)
424
+ mime = "image/jpeg" if ext in ("jpg", "jpeg") else f"image/{ext}"
425
+ # The upload form field MUST be `upfile` — `file`/`image` return 1053.
426
+ body, boundary = _multipart("upfile", f"image.{ext}", img, mime)
427
+ # This endpoint answers with a FLAT envelope (no `data` wrapper), unlike the
428
+ # rest of the creator API, so read the fields off the top level.
429
+ # nonfatal → a per-image failure raises instead of die()ing, so the caller's
430
+ # failure collector (and --drop-failed-images) actually gets to run.
431
+ env = api_envelope("POST", "/mp/agw/article_material/photo/upload_picture/?type=json", jar,
432
+ referer=f"{MP}/profile_v4/graphic/publish", body=body,
433
+ headers={"Content-Type": f"multipart/form-data; boundary={boundary}"},
434
+ nonfatal=True)
435
+ if not env.get("web_url") or not env.get("web_uri"):
436
+ raise RuntimeError(f"upload returned no image URL: {str(env)[:200]}")
437
+ return env
438
+
439
+
440
+ def _img_tag(info: dict, alt: str) -> str:
441
+ """头条 needs the CDN attributes it returned, not a bare src — a plain
442
+ <img src> (or any external URL) makes publish fail with 7115 图片uri非法."""
443
+ return (
444
+ f'<img src="{_attr_url(info["web_url"])}"'
445
+ f' img_width="{int(info.get("width") or 0)}"'
446
+ f' img_height="{int(info.get("height") or 0)}"'
447
+ f' image_type="{int(info.get("image_type") or 1)}"'
448
+ f' mime_type="{_alt(str(info.get("mime_type") or "image/jpeg"))}"'
449
+ f' web_uri="{_alt(str(info["web_uri"]))}"'
450
+ f' alt="{_alt(alt)}">'
451
+ )
452
+
453
+
454
+ def rehost_images(jar, html, drop_failed=False):
455
+ """Rewrite every <img> in the rendered body to a 头条-hosted one.
456
+
457
+ An external <img> is not merely blocked — it makes 头条 reject the WHOLE
458
+ article (7115), so keeping the original is never an option. Default is to
459
+ abort the publish so the user's article is never silently posted without
460
+ its images; ``drop_failed`` opts into dropping them instead.
461
+ """
462
+ failures = []
463
+ try:
464
+ spans = _ImgFinder().find(html)
465
+ except Exception as e: # noqa: BLE001 — malformed beyond parsing
466
+ die(f"could not parse the article HTML to find its images: {e}")
467
+
468
+ out, cursor = [], 0
469
+ for start, end, attrs in spans:
470
+ out.append(html[cursor:start])
471
+ cursor = end
472
+ tag = html[start:end]
473
+ src = _first_attr(attrs, "src")
474
+ if not src:
475
+ # `data-src`-only or valueless src — we cannot fetch it, and 头条
476
+ # would reject the article, so surface it instead of dropping it.
477
+ failures.append(f"{tag[:80]} (no usable src attribute)")
478
+ continue
479
+ # Attribute values arrive unescaped; re-escape alt on the way back out.
480
+ alt = _first_attr(attrs, "alt") or ""
481
+ host = (urllib.parse.urlsplit(src).hostname or "").lower()
482
+ if any(_same_or_sub(host, s) for s in _IMG_SKIP) and _first_attr(attrs, "web_uri"):
483
+ out.append(tag)
484
+ continue
485
+ try:
486
+ info = upload_image(jar, src)
487
+ sys.stderr.write(f"[img] rehosted {src[:60]} -> {info['web_uri']}\n")
488
+ out.append(_img_tag(info, _html.escape(alt, quote=False)))
489
+ except Exception as e: # noqa: BLE001 — collected, then reported below
490
+ failures.append(f"{src[:80]} ({e})")
491
+ out.append(html[cursor:])
492
+ result = "".join(out)
493
+ if failures and not drop_failed:
494
+ die("could not upload these images to 头条, and 头条 rejects any article "
495
+ "with an external image, so nothing was published:\n - "
496
+ + "\n - ".join(failures)
497
+ + "\nFix the image URLs, or re-run with --drop-failed-images to "
498
+ "publish without them.")
499
+ for f in failures:
500
+ sys.stderr.write(f"[img] DROPPED {f}\n")
501
+ return result
502
+
503
+
504
+ # ── Markdown → HTML (头条's `content` field is rendered HTML, not source) ──
505
+
506
+ _IMG_RE = re.compile(r"!\[([^\]]{0,500})\]\(([^)\s]{0,2000})\)")
507
+ _LINK_RE = re.compile(r"(?<!!)\[([^\]]{0,500})\]\(([^)\s]{0,2000})\)")
508
+
509
+
510
+ def _attr_url(u):
511
+ # Strip C0 controls/space FIRST, then allow only http/https/mailto — otherwise
512
+ # `\x01javascript:` would slip past a literal scheme check yet be re-normalised
513
+ # to javascript: by the browser. Finally neutralise attribute-breaking quotes.
514
+ u = re.sub(r"[\x00-\x20\x7f]", "", u or "")
515
+ m = re.match(r"(?i)([a-z][a-z0-9+.\-]*):", u)
516
+ if m and m.group(1).lower() not in ("http", "https", "mailto"):
517
+ return "#"
518
+ return u.replace('"', "%22").replace("'", "%27")
519
+
520
+
521
+ def _alt(s):
522
+ return (s or "").replace('"', "&quot;").replace("'", "&#x27;")
523
+
524
+
525
+ def _inline_md(t):
526
+ t = _html.escape(t, quote=False).replace("\x00", "")
527
+ # Stash code spans FIRST so emphasis/link/image markup inside `…` stays literal.
528
+ spans = []
529
+
530
+ def _stash(m):
531
+ spans.append("<code>" + m.group(1) + "</code>")
532
+ return f"\x00{len(spans) - 1}\x00"
533
+
534
+ t = re.sub(r"`([^`]+)`", _stash, t)
535
+ t = _IMG_RE.sub(lambda m: f'<img src="{_attr_url(m.group(2))}" alt="{_alt(m.group(1))}">', t)
536
+ t = _LINK_RE.sub(lambda m: f'<a href="{_attr_url(m.group(2))}">{m.group(1)}</a>', t)
537
+ t = re.sub(r"\*\*([^*]+)\*\*", r"<strong>\1</strong>", t)
538
+ t = re.sub(r"(?<!\*)\*([^*\n]+)\*(?!\*)", r"<em>\1</em>", t)
539
+ t = re.sub(r"\x00(\d+)\x00", lambda m: spans[int(m.group(1))], t)
540
+ return t
541
+
542
+
543
+ def _md_block_to_html(b):
544
+ b = b.strip("\n")
545
+ if not b.strip():
546
+ return None
547
+ f = b.lstrip()
548
+ m = re.match(r"(#{1,6})\s+(.*)", f)
549
+ if m:
550
+ lvl = len(m.group(1))
551
+ return f"<h{lvl}>{_inline_md(m.group(2).strip())}</h{lvl}>"
552
+ if f.startswith(">"):
553
+ inner = re.sub(r"^>\s?", "", b, flags=re.M).replace("\n", " ")
554
+ return "<blockquote><p>" + _inline_md(inner) + "</p></blockquote>"
555
+ if re.match(r"[-*+]\s+", f):
556
+ items = [_inline_md(re.sub(r"^[-*+]\s+", "", ln)) for ln in b.split("\n") if ln.strip()]
557
+ return "<ul>" + "".join(f"<li>{x}</li>" for x in items) + "</ul>"
558
+ if re.match(r"\d+\.\s+", f):
559
+ items = [_inline_md(re.sub(r"^\d+\.\s+", "", ln)) for ln in b.split("\n") if ln.strip()]
560
+ return "<ol>" + "".join(f"<li>{x}</li>" for x in items) + "</ol>"
561
+ only = re.fullmatch(r"!\[([^\]]{0,500})\]\(([^)\s]{0,2000})\)", f)
562
+ if only:
563
+ return f'<img src="{_attr_url(only.group(2))}" alt="{_alt(_html.escape(only.group(1), quote=False))}">'
564
+ return "<p>" + _inline_md(b.replace("\n", " ").strip()) + "</p>"
565
+
566
+
567
+ def md_to_html(src):
568
+ """Minimal stdlib Markdown→HTML — 头条's body field is HTML, so raw markdown
569
+ would publish as literal `##` / `![]()` on one collapsed line."""
570
+ src = (src or "").strip()
571
+ res = []
572
+ # Pull fenced code out FIRST (it may itself contain blank lines) so the
573
+ # blank-line block splitter below can never tear a ``` … ``` fence apart.
574
+ for i, part in enumerate(re.split(r"(?ms)^(```.*?(?:\n```[ \t]*$|\Z))", src)):
575
+ if i % 2 == 1:
576
+ code = re.sub(r"\A```[^\n]*\n?", "", part)
577
+ code = re.sub(r"\n?```[ \t]*\Z", "", code)
578
+ res.append("<pre><code>" + _html.escape(code) + "</code></pre>")
579
+ continue
580
+ for b in re.split(r"\n[ \t]*\n", part):
581
+ block = _md_block_to_html(b)
582
+ if block:
583
+ res.append(block)
584
+ return "\n".join(res)
585
+
586
+
587
+ def _looks_like_markdown(s):
588
+ # Three-way classification (no full parser): block-level HTML ⇒ already an
589
+ # HTML document → pass through; else markdown markers ⇒ render; else
590
+ # inline-only HTML ⇒ pass through; else plain text ⇒ render.
591
+ s = s or ""
592
+ if re.search(
593
+ r"</?(?:p|div|h[1-6]|ul|ol|li|table|thead|tbody|tr|td|th|blockquote|"
594
+ r"pre|figure|figcaption|section|article|header|footer|nav|aside|hr|"
595
+ r"main|details|summary)\b", s, re.I,
596
+ ):
597
+ return False
598
+ # ^-anchored (re.M) + bounded {0,N} repeats so the scan can't backtrack
599
+ # across newlines or to EOF (no quadratic scanning).
600
+ if re.search(
601
+ r"^#{1,6}\s|!\[[^\]]{0,500}\]\([^)]{0,2000}\)|^[ \t]*[-*+]\s|^[ \t]*\d+\.\s"
602
+ r"|\[[^\]]{0,500}\]\([^)]{0,2000}\)|`[^`]{1,500}`|\*\*[^*]{1,500}\*\*",
603
+ s, re.M,
604
+ ):
605
+ return True
606
+ inline_html = re.search(
607
+ r"</?(?:a|strong|em|b|i|u|s|span|code|br|img|small|mark|sup|sub|"
608
+ r"video|audio|iframe)\b", s, re.I,
609
+ )
610
+ return not inline_html
611
+
612
+
613
+ # ── publish ─────────────────────────────────────────────────────────
614
+
615
+ PUBLISH_PATH = "/mp/agw/article/publish/?source=mp&type=article"
616
+
617
+
618
+ def _resolve_url(jar, pgc_id, draft):
619
+ """Look the freshly-written article up in the list to return its real URL
620
+ (preview URL for a draft, public /item/ URL once published).
621
+
622
+ Runs AFTER the write succeeded, so it must never fail the command — losing
623
+ the pgc_id would read as a failed publish and invite a duplicate post.
624
+ """
625
+ try:
626
+ items, _ = _list_page(jar, "draft" if draft else "all", 1, 20, nonfatal=True)
627
+ for it in items:
628
+ if str(it.get("pgc_id")) == str(pgc_id):
629
+ return it.get("article_url")
630
+ except (Exception, SystemExit): # noqa: BLE001 — URL is a nicety
631
+ pass
632
+ if draft:
633
+ return f"{MP}/preview_article/?pgc_id={pgc_id}"
634
+ return f"https://www.toutiao.com/item/{pgc_id}/"
635
+
636
+
637
+ def cmd_publish(jar, args):
638
+ if not args.title:
639
+ die("--title is required")
640
+ # 头条 rejects out-of-range titles server-side; fail early with a clear message.
641
+ if not 2 <= len(args.title) <= 30:
642
+ die(f"头条 titles must be 2–30 characters; got {len(args.title)}")
643
+ if not args.content_file and args.content is None:
644
+ die("provide --content-file <path.md> or --content <markdown>")
645
+ content = args.content
646
+ if args.content_file:
647
+ try:
648
+ with open(args.content_file, encoding="utf-8") as f:
649
+ content = f.read()
650
+ except OSError as e:
651
+ die(f"cannot read --content-file: {e}")
652
+ content = content or ""
653
+
654
+ if not CONFIRM:
655
+ out({
656
+ "dry_run": True, "command": "publish", "platform": "toutiao",
657
+ "title": args.title, "draft_only": args.draft_only,
658
+ "content_bytes": len(content),
659
+ "note": "头条 content is Markdown (converted to HTML for the body). "
660
+ "Re-run with --confirm as the LAST argument to actually write. "
661
+ "Without --draft-only it publishes a PUBLIC article on the "
662
+ "user's real 头条号 and goes through 审核.",
663
+ })
664
+ return
665
+
666
+ # Writes need the csrftoken cookie echoed as a header; without it 头条 fails
667
+ # deep in its API with an opaque code instead of "reconnect".
668
+ if not cookie_value(jar, "csrftoken"):
669
+ die("the 今日头条 cookie jar has no `csrftoken` — it is incomplete or "
670
+ "expired. Reconnect at https://auth.acedata.cloud/user/connections.")
671
+
672
+ # Render to HTML FIRST, then rewrite <img> tags — 头条 rejects the whole
673
+ # article (7115) if any image isn't hosted on its own CDN with web_uri.
674
+ body_html = md_to_html(content) if _looks_like_markdown(content) else content
675
+ if not args.no_rehost_images:
676
+ body_html = rehost_images(jar, body_html, drop_failed=args.drop_failed_images)
677
+
678
+ form = urllib.parse.urlencode({
679
+ "title": args.title,
680
+ "content": body_html,
681
+ # save=1 → draft; save=0 → submit for 审核 and publish.
682
+ "save": "1" if args.draft_only else "0",
683
+ # 0 = no ad. Other values need 广告/自营 permissions most accounts lack.
684
+ "article_ad_type": "0",
685
+ })
686
+ d = api("POST", PUBLISH_PATH, jar, referer=f"{MP}/profile_v4/graphic/publish",
687
+ body=form, headers={"Content-Type": "application/x-www-form-urlencoded"})
688
+ pgc_id = (d or {}).get("pgc_id")
689
+ if not pgc_id:
690
+ die(f"publish returned no pgc_id: {str(d)[:300]}")
691
+ out({
692
+ "ok": True,
693
+ "draft_only": bool(args.draft_only),
694
+ "published": not args.draft_only,
695
+ "pgc_id": str(pgc_id),
696
+ "url": _resolve_url(jar, pgc_id, args.draft_only),
697
+ "note": None if args.draft_only else "头条 reviews new articles (审核); "
698
+ "the public URL goes live once it passes.",
699
+ })
700
+
701
+
702
+ COMMANDS = {
703
+ "whoami": cmd_whoami,
704
+ "articles": cmd_articles,
705
+ "article": cmd_article,
706
+ "publish": cmd_publish,
707
+ }
708
+
709
+
710
+ def main() -> None:
711
+ p = argparse.ArgumentParser(prog="toutiao.py", description="今日头条 cookie CLI")
712
+ sub = p.add_subparsers(dest="command", required=True)
713
+ sub.add_parser("whoami", help="show the logged-in 头条号")
714
+ sp = sub.add_parser("articles", help="list the user's articles + stats")
715
+ sp.add_argument("--limit", type=int, default=20)
716
+ sp.add_argument("--status", default="all",
717
+ choices=["all", "draft", "published", "reviewing", "failed"])
718
+ sp = sub.add_parser("article", help="one article's stats")
719
+ sp.add_argument("id", help="pgc_id / item_id")
720
+ sp = sub.add_parser("publish", help="create/publish an article (GATED by trailing --confirm)")
721
+ sp.add_argument("--title", help="2–30 characters")
722
+ sp.add_argument("--content", help="Markdown content inline")
723
+ sp.add_argument("--content-file", help="path to a Markdown file")
724
+ sp.add_argument("--draft-only", action="store_true",
725
+ help="save a private draft; do NOT go public")
726
+ sp.add_argument("--no-rehost-images", action="store_true",
727
+ help="keep external image URLs as-is (头条 will reject the article)")
728
+ sp.add_argument("--drop-failed-images", action="store_true",
729
+ help="publish without any image that fails to upload, "
730
+ "instead of aborting")
731
+ args = p.parse_args(ARGV)
732
+ jar = load_cookies()
733
+ COMMANDS[args.command](jar, args)
734
+
735
+
736
+ if __name__ == "__main__":
737
+ main()