master-skill 0.12.10 → 0.12.12

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (35) hide show
  1. package/.claude-plugin/marketplace.json +1 -1
  2. package/.claude-plugin/plugin.json +5 -1
  3. package/.codex/INSTALL.md +23 -16
  4. package/.cursor-plugin/plugin.json +1 -1
  5. package/.opencode/INSTALL.md +29 -16
  6. package/README.md +2 -2
  7. package/README_EN.md +2 -2
  8. package/SKILL.md +2 -2
  9. package/bin/cli.mjs +96 -0
  10. package/gemini-extension.json +1 -1
  11. package/hooks/session_start.py +20 -7
  12. package/package.json +2 -2
  13. package/prebuilt/master-ajahn-chah/SKILL.md +4 -4
  14. package/prebuilt/master-ajahn-chah/sources/sutta-excerpts.md +3 -3
  15. package/prebuilt/master-atisha/SKILL.md +1 -1
  16. package/prebuilt/master-buddhaghosa/SKILL.md +1 -1
  17. package/prebuilt/master-huineng/references/teaching.md +2 -2
  18. package/prebuilt/master-huineng/references/voice.md +3 -3
  19. package/prebuilt/master-mahasi-sayadaw/SKILL.md +2 -2
  20. package/prebuilt/master-mahasi-sayadaw/sources/teachings-excerpts.md +2 -2
  21. package/prebuilt/master-milarepa/SKILL.md +4 -4
  22. package/prebuilt/master-ouyi/SKILL.md +1 -1
  23. package/prebuilt/master-tsongkhapa/SKILL.md +3 -3
  24. package/prebuilt/master-xuyun/SKILL.md +1 -1
  25. package/prebuilt/master-yinguang/sources/INDEX.md +1 -1
  26. package/prebuilt/master-yinguang/sources/yihanbianfu-excerpts.md +6 -5
  27. package/prebuilt/master-zhiyi/SKILL.md +2 -2
  28. package/references/workflow-details.md +27 -9
  29. package/scripts/check-gate-liveness.py +227 -0
  30. package/scripts/validate-quote-attribution.py +126 -0
  31. package/scripts/validate-section-references.py +162 -0
  32. package/tools/compiled-teaching-sources.json +89 -0
  33. package/tools/master_builder.py +88 -4
  34. package/tools/verify_sources.py +245 -2
  35. package/tools/version_manager.py +13 -3
@@ -1,13 +1,14 @@
1
- # 《一函遍復》关键片段
1
+ # 《印光法師文鈔續編》关键片段(一函遍復、复唐能诚居士书)
2
2
 
3
3
  > 印光大师著,出自《印光法師文鈔續編》。CBETA 与 FoJin 均未收录文钞,故无经号、无在线链接。
4
- > 本文件为教学引用用。节选文字尚未与印本逐字核对。
4
+ > 本文件为教学引用用。以下引文已于 2026-09-16 与殆知阁所藏《文鈔續編》全文逐字核对;
5
+ > 此前两段是真语拼接而成的改写,并非原文,已按原书改正。
5
6
 
6
7
  ## 敦伦尽分、闲邪存诚
7
8
 
8
9
  原典(节选):
9
10
 
10
- > 敦伦尽分,闲邪存诚,诸恶莫作,众善奉行。正心诚意,以立人道之本。然后以此回向净土,求生西方。在家人第一须尽各人本分。如父慈子孝,兄友弟恭,夫和妇顺。
11
+ > 又须父慈,子孝,兄友,弟恭,夫和,妇顺,主仁,仆忠,恪尽己分。不计他对我之尽分与否,我总要尽我之分。能于家庭,及与社会,尽谊尽分,是名善人。……念佛最要紧,是敦伦尽分,闲邪存诚,诸恶莫作,众善奉行。存好心,说好话,行好事。
11
12
 
12
13
  **引用格式:**【《印光法師文鈔續編》·一函遍復】
13
14
 
@@ -23,9 +24,9 @@
23
24
 
24
25
  原典(节选):
25
26
 
26
- > 因果者,圣人治天下、佛度众生之大权也。若人人深信因果,则人心自善,风俗自淳,唐虞之治不难复见。
27
+ > 吾常曰,因果者,圣人治天下,佛度众生之大权也。
27
28
 
28
- **引用格式:**【《印光法師文鈔續編》·一函遍復】
29
+ **引用格式:**【《印光法師文鈔續編》·复唐能诚居士书】
29
30
 
30
31
  **教义要点:**
31
32
  - 因果为佛法入门之基础
@@ -39,7 +39,7 @@ verified_at: 2026-04-06
39
39
  - **教义询问**(一念三千 / 三谛 / 五时八教 / 性具善恶)
40
40
  → 读 `references/teaching.md` + 引用 `sources/mohezhiguan-excerpts.md`
41
41
  - **修行方法**(止观 / 一心三观 / 四种三昧 / 六即判位)
42
- → 读 `sources/mohezhiguan-excerpts.md` §止觀法門 + 必要时 `references/teaching.md` §修行方法
42
+ → 读 `sources/mohezhiguan-excerpts.md` §圆顿止观开篇、§二十五方便、§一心三观、§六即佛 + 必要时 `references/teaching.md` §修行方法
43
43
  - **判教体系**(藏通别圆 / 化仪四教 / 五时分判)
44
44
  → 读 `references/teaching.md` §判教 + `sources/fahua-xuanyi-excerpts.md`
45
45
  - **法华经义理**(开权显实 / 会三归一 / 穷子喻)
@@ -138,7 +138,7 @@ GET https://fojin.app/api/search/semantic?q=<URL编码查询>&top_k=5 # 语义
138
138
  | 三谛圆融怎么理解 | `references/teaching.md` §三谛圆融 | 《法華玄義》卷二,T1716 |
139
139
  | 五时八教怎么分 | `references/teaching.md` §判教 | 《法華玄義》卷一,T1716 |
140
140
  | 止观怎么修 | `sources/mohezhiguan-excerpts.md` §二十五方便 | 《摩訶止觀》卷四,T1911 |
141
- | 一心三观 | `sources/mohezhiguan-excerpts.md` §一心三觀 | 《摩訶止觀》卷五,T1911 |
141
+ | 一心三观 | `sources/mohezhiguan-excerpts.md` §一心三观 | 《摩訶止觀》卷五,T1911 |
142
142
  | 四种三昧 | `references/teaching.md` §四种三昧 | 《摩訶止觀》卷二,T1911 |
143
143
  | 性具善恶 | `references/teaching.md` §性具 | 《觀音玄義》卷上,T1726 |
144
144
  | 入门从哪开始 | — | 《修習止觀坐禪法要》(小止觀),T1915 |
@@ -235,7 +235,7 @@ FAIL → 自动修复后重审。
235
235
  ## Step 5:写入文件细则
236
236
 
237
237
  ```bash
238
- python3 ${CLAUDE_SKILL_DIR}/tools/master_builder.py --spec generated-master.json --output masters/
238
+ python3 "${CLAUDE_SKILL_DIR}/tools/master_builder.py" --spec generated-master.json --output "${CLAUDE_SKILL_DIR}/masters/"
239
239
  ```
240
240
 
241
241
  `generated-master.json` 是审查通过后的生成规格,必含 `name`、`tradition`、`school`、`era`、
@@ -245,7 +245,7 @@ python3 ${CLAUDE_SKILL_DIR}/tools/master_builder.py --spec generated-master.json
245
245
  ### 生成后终验
246
246
 
247
247
  ```bash
248
- python3 ${CLAUDE_SKILL_DIR}/tools/verify_sources.py --final-check masters/master-{slug}/
248
+ python3 "${CLAUDE_SKILL_DIR}/tools/verify_sources.py" --final-check "${CLAUDE_SKILL_DIR}/masters/master-{slug}/"
249
249
  ```
250
250
 
251
251
  `--final-check` 离线验证 persona 目录包含 `SKILL.md`、`teaching.md`、`voice.md`、`meta.json`,
@@ -255,8 +255,8 @@ python3 ${CLAUDE_SKILL_DIR}/tools/verify_sources.py --final-check masters/master
255
255
  ### 生成目录结构
256
256
 
257
257
  ```
258
- masters/master-{slug}/
259
- ├── SKILL.md # /master-{slug} 触发(完整角色定义)
258
+ ${CLAUDE_SKILL_DIR}/masters/master-{slug}/
259
+ ├── SKILL.md # 注册后由 /master-{slug} 触发(完整角色定义)
260
260
  ├── teaching.md # 教义体系(可单独使用)
261
261
  ├── voice.md # 说法风格(可单独使用)
262
262
  └── meta.json # 元数据(版本、生成时间、数据来源)
@@ -265,9 +265,27 @@ masters/master-{slug}/
265
265
  ### 角色注册(按运行环境)
266
266
 
267
267
  **Claude Code 用户**
268
- 1. 生成的 SKILL.md 已放置在 `masters/master-{slug}/`
269
- 2. 确保 `masters/` 在 Claude Code skill 搜索路径中(检查 `.claude/settings.json` 的 `skillDirs` 配置)
270
- 3. 完成后自动可通过 `/master-{slug}` 触发
268
+
269
+ 终验通过后注册:
270
+
271
+ ```bash
272
+ python3 "${CLAUDE_SKILL_DIR}/tools/master_builder.py" --register "${CLAUDE_SKILL_DIR}/masters/master-{slug}"
273
+ ```
274
+
275
+ 它在 `~/.claude/skills/master-{slug}` 建一个指向生成目录的链接(Windows 上建 junction),输出 JSON:
276
+
277
+ - `invoke`:调用命令,照此告知用户;
278
+ - `restart_required: true`:`~/.claude/skills/` 是这次才建的,需重启 Claude Code 才会被发现;否则当前会话内即可调用;
279
+ - 退出码 1 且提示 `not replacing`:同名 skill 已存在(例如重新生成了一位预置祖师)。**不要删除或覆盖它**,把情况告诉用户,由用户决定改名或自行移除。
280
+
281
+ 为什么需要这一步:Claude Code 只加载 `<skills 目录>/<名字>/SKILL.md`,不往下扫,也没有 `skillDirs` 之类的设置。
282
+ 2026-09-16 用 Claude Code 2.1.273 在隔离配置中实测:放在 `~/.claude/skills/create-master/masters/` 下的 persona
283
+ 不出现在 `/skills` 里;注册出链接后,同一会话内即出现。生成目录留在 `masters/` 是为了让
284
+ `master-skill update` 更新运行时的时候不丢失用户生成的 persona。
285
+
286
+ 以插件方式安装时,`${CLAUDE_SKILL_DIR}` 是按版本号区分的插件缓存目录:`claude plugin update` 之后
287
+ 它换到新版本目录,新目录下的 `masters/` 是空的(实测 0.12.11 → 0.12.12)。要长期保留自己生成的 persona,
288
+ 请用 `npx master-skill install create-master` 或 git clone 方式安装生成器。
271
289
 
272
290
  **OpenClaw 用户**
273
291
  1. 将 `masters/master-{slug}/` 复制到 OpenClaw 的 skills 目录
@@ -278,8 +296,8 @@ masters/master-{slug}/
278
296
 
279
297
  ```
280
298
  已生成「{master_name}」教学角色
281
- 目录:masters/master-{slug}/
282
- 调用命令:/master-{slug}
299
+ 目录:${CLAUDE_SKILL_DIR}/masters/master-{slug}/
300
+ 调用命令:/master-{slug}(已注册到 ~/.claude/skills/master-{slug})
283
301
  包含文件:SKILL.md, teaching.md, voice.md, meta.json
284
302
  数据来源:{n} 条经文,{m} 个知识图谱实体
285
303
  ```
@@ -25,6 +25,7 @@ Usage:
25
25
  from __future__ import annotations
26
26
 
27
27
  import argparse
28
+ import ast
28
29
  import functools
29
30
  import json
30
31
  import re
@@ -168,6 +169,75 @@ ADVISORY_GATES = {
168
169
  ),
169
170
  }
170
171
 
172
+ # A gate that runs nowhere on a pull request has never guarded a change. This repo
173
+ # has shipped that twice: `validate-curriculum-sources.py` was "wired into no
174
+ # workflow, no npm script and no sub-check — only its own unit tests" (see the
175
+ # sub-check in validate.py that now runs it), and on 2026-09-16
176
+ # `validate-citation-templates.py` and `validate-self-audit-sources.py` were found
177
+ # to live only inside `npm test`, which only npm-publish.yml runs, on a published
178
+ # release. Both passed — the defect was latent, which is exactly why nothing
179
+ # surfaced it.
180
+ #
181
+ # Every entry script under scripts/ must therefore be reachable from a workflow
182
+ # that triggers on `pull_request`, or be declared here with the reason it is not.
183
+ # `check_every_gate_runs_on_a_pr` keeps this true in both directions.
184
+ NOT_A_PR_GATE = {
185
+ "cite.py": (
186
+ "a reader-facing offline lookup tool, documented in README.md and in the "
187
+ "personas' own SKILL.md; not a gate over repository content"
188
+ ),
189
+ "query.py": (
190
+ "a reader-facing offline search tool, documented alongside cite.py; not a "
191
+ "gate over repository content"
192
+ ),
193
+ "check-pe-subsystem.py": (
194
+ "inspects a built Windows executable, which exists only after the desktop "
195
+ "release build — release-desktop.yml is the only place it can run"
196
+ ),
197
+ "reaudit-report.py": (
198
+ "re-audits a committed eval run's stored answers; run by hand after a paid "
199
+ "sweep, against a report that does not exist on a PR"
200
+ ),
201
+ "regrade-report.py": (
202
+ "re-grades a committed eval run against the current judge; same as "
203
+ "reaudit-report.py — it needs a report a PR does not produce"
204
+ ),
205
+ }
206
+
207
+
208
+ # `npm test` is what CONTRIBUTING tells a contributor to run before touching
209
+ # scripts/, in its own words 「避免在 CI 才发现」. A command that exists to pre-empt CI
210
+ # has to cover what CI checks. It has fallen behind twice: pytest was missing from it
211
+ # until 2026-09-03, and on 2026-09-16 four content gates the PR job runs —
212
+ # validate-citation-contract, validate-cross-critique, validate-lore-triggers-content
213
+ # and validate-quote-attribution — were absent, so a contributor could go green
214
+ # locally and still be failed by CI.
215
+ #
216
+ # Anything the per-PR job runs must therefore appear in `npm test` too, or be
217
+ # declared here. `check_npm_test_covers_pr_gates` keeps this true in both directions.
218
+ NOT_IN_NPM_TEST = {
219
+ "check-eval-sdk-surface.py": (
220
+ "asserts the pinned eval SDKs still expose what test-fidelity.py calls — it "
221
+ "needs requirements-eval.txt installed, which a content contributor has no "
222
+ "reason to have"
223
+ ),
224
+ "smoke-eval-sdk.py": (
225
+ "stands up a local server for a keyless end-to-end SDK smoke; same eval-only "
226
+ "dependency, and far slower than the content gates around it"
227
+ ),
228
+ "select-fidelity-smoke.py": (
229
+ "picks which persona the CI smoke grades from job metadata — a CI scheduling "
230
+ "helper, not a check over repository content"
231
+ ),
232
+ "check-audit-ignores.py": (
233
+ "takes the cargo-audit JSON as an argument — security-scan.yml runs "
234
+ "`cargo audit --file desktop/Cargo.lock --json > audit.json` first. Without a "
235
+ "Rust toolchain and the advisory database there is nothing for it to read; "
236
+ "run bare it exits 2 on argparse usage"
237
+ ),
238
+ }
239
+
240
+
171
241
  # The shape of a silent skip: a step that exits 0 because a secret is missing.
172
242
  _SKIP_ON_MISSING_SECRET = re.compile(r'\[\s+-z\s+"\$\{[A-Z_]+:-\}"\s+\]')
173
243
 
@@ -251,6 +321,161 @@ def check_declared_gates_still_exist(workflow_docs: dict[str, dict]) -> list[str
251
321
  # ---------------------------------------------------------------------------
252
322
 
253
323
 
324
+ def _script_references(source: str, scripts: set[str]) -> set[str]:
325
+ """Which other scripts this source actually runs — imports and loaded filenames.
326
+
327
+ Both spellings are in use: `validate.py` loads five siblings through
328
+ `spec_from_file_location(..., "validate-curriculum-sources.py")`, which puts the
329
+ literal filename in the source, while `verify_citations.py` is pulled in as
330
+ `from verify_citations import …`. Counting only workflow text would report both
331
+ as unreachable and invent a defect where there is none.
332
+
333
+ Read through `ast`, not as text. The first version matched filenames anywhere in
334
+ the source and reported five scripts as reachable on the strength of *comments*:
335
+ `verify_citations.py` mentions "scripts/query.py" in a comment about a shared
336
+ guard, and verify_citations is imported by a job the PR runs, so query.py came
337
+ out "reachable". Comments do not survive parsing, and an exact-match on string
338
+ constants keeps a docstring that merely names a path from counting as a call.
339
+ """
340
+ try:
341
+ tree = ast.parse(source)
342
+ except SyntaxError: # pragma: no cover — a syntactically broken script
343
+ return set()
344
+
345
+ wanted = {name: {name, f"scripts/{name}"} for name in scripts}
346
+ modules = {name[:-3]: name for name in scripts if "-" not in name[:-3]}
347
+ hit: set[str] = set()
348
+ for node in ast.walk(tree):
349
+ if isinstance(node, ast.ImportFrom) and node.module in modules:
350
+ hit.add(modules[node.module])
351
+ elif isinstance(node, ast.Import):
352
+ for alias in node.names:
353
+ if alias.name in modules:
354
+ hit.add(modules[alias.name])
355
+ elif isinstance(node, ast.Constant) and isinstance(node.value, str):
356
+ for name, spellings in wanted.items():
357
+ if node.value in spellings:
358
+ hit.add(name)
359
+ return hit
360
+
361
+
362
+ def pr_reachable_scripts(root: Path, workflow_docs: dict[str, dict]) -> set[str]:
363
+ """Scripts a pull request actually executes, following indirect calls."""
364
+ scripts_dir = root / "scripts"
365
+ names = {p.name for p in scripts_dir.glob("*.py")}
366
+ # This file names five scripts in NOT_A_PR_GATE, and this file runs on every PR.
367
+ # Counting its own source as a caller made each declared script "reachable" and
368
+ # then reported the declaration as stale — the checker proving its own entries
369
+ # wrong. Declaring a script is not calling it.
370
+ sources = {
371
+ p.name: p.read_text(encoding="utf-8")
372
+ for p in scripts_dir.glob("*.py")
373
+ if p.name != Path(__file__).name
374
+ }
375
+
376
+ reachable: set[str] = set()
377
+ for path, doc in workflow_docs.items():
378
+ triggers = doc.get("on", doc.get(True))
379
+ keys = set(triggers) if isinstance(triggers, (dict, list)) else set()
380
+ if "pull_request" not in keys:
381
+ continue
382
+ reachable |= {name for name in names if name in (root / path).read_text(encoding="utf-8")}
383
+
384
+ # Fixpoint: a script the PR runs may load or import others.
385
+ while True:
386
+ grown = set(reachable)
387
+ for name in list(reachable):
388
+ grown |= _script_references(sources.get(name, ""), names)
389
+ if grown == reachable:
390
+ return reachable
391
+ reachable = grown
392
+
393
+
394
+ def check_every_gate_runs_on_a_pr(root: Path, workflow_docs: dict[str, dict]) -> list[str]:
395
+ """An entry script must run on a pull request, or say why it does not."""
396
+ scripts_dir = root / "scripts"
397
+ if not scripts_dir.is_dir():
398
+ return []
399
+ entries = {
400
+ p.name
401
+ for p in scripts_dir.glob("*.py")
402
+ if "def main(" in p.read_text(encoding="utf-8")
403
+ }
404
+ reachable = pr_reachable_scripts(root, workflow_docs)
405
+
406
+ problems = [
407
+ f"scripts/{name} runs nowhere on a pull request and is not in NOT_A_PR_GATE "
408
+ "— a gate whose first real execution is the release has guarded nothing"
409
+ for name in sorted(entries - reachable)
410
+ if name not in NOT_A_PR_GATE
411
+ ]
412
+ problems += [
413
+ f"NOT_A_PR_GATE declares {name!r}, but no such script exists — stale entry"
414
+ for name in sorted(NOT_A_PR_GATE)
415
+ if name not in entries
416
+ ]
417
+ problems += [
418
+ f"NOT_A_PR_GATE declares {name!r}, but a pull request does run it now "
419
+ "— drop the entry rather than leave a false caveat standing"
420
+ for name in sorted(NOT_A_PR_GATE)
421
+ if name in reachable
422
+ ]
423
+ return problems
424
+
425
+
426
+ def npm_test_scripts(root: Path) -> set[str]:
427
+ """The scripts the documented pre-push command actually runs."""
428
+ package = root / "package.json"
429
+ if not package.exists():
430
+ return set()
431
+ data = json.loads(package.read_text(encoding="utf-8"))
432
+ command = str((data.get("scripts") or {}).get("test") or "")
433
+ return set(re.findall(r"scripts/([a-z0-9_-]+\.py)", command))
434
+
435
+
436
+ def pr_workflow_scripts(root: Path, workflow_docs: dict[str, dict]) -> set[str]:
437
+ """Scripts named outright by a workflow that triggers on `pull_request`.
438
+
439
+ Direct mentions only, unlike `pr_reachable_scripts`: `npm test` runs commands, so
440
+ what it has to match is the commands CI runs, not everything those import.
441
+ """
442
+ named: set[str] = set()
443
+ for path, doc in workflow_docs.items():
444
+ triggers = doc.get("on", doc.get(True))
445
+ keys = set(triggers) if isinstance(triggers, (dict, list)) else set()
446
+ if "pull_request" not in keys:
447
+ continue
448
+ named |= set(re.findall(r"scripts/([a-z0-9_-]+\.py)", (root / path).read_text(encoding="utf-8")))
449
+ return named
450
+
451
+
452
+ def check_npm_test_covers_pr_gates(root: Path, workflow_docs: dict[str, dict]) -> list[str]:
453
+ """What CI runs on a PR, `npm test` must run too — or say why it does not."""
454
+ if not (root / "package.json").exists():
455
+ return []
456
+ in_ci = pr_workflow_scripts(root, workflow_docs)
457
+ in_npm = npm_test_scripts(root)
458
+
459
+ problems = [
460
+ f"scripts/{name} runs on every PR in CI but is not in `npm test` and not in "
461
+ "NOT_IN_NPM_TEST — the command that exists to pre-empt CI does not cover it"
462
+ for name in sorted(in_ci - in_npm)
463
+ if name not in NOT_IN_NPM_TEST
464
+ ]
465
+ problems += [
466
+ f"NOT_IN_NPM_TEST declares {name!r}, but `npm test` runs it now — drop the "
467
+ "entry rather than leave a false caveat standing"
468
+ for name in sorted(NOT_IN_NPM_TEST)
469
+ if name in in_npm
470
+ ]
471
+ problems += [
472
+ f"NOT_IN_NPM_TEST declares {name!r}, but no PR workflow runs it — stale entry"
473
+ for name in sorted(NOT_IN_NPM_TEST)
474
+ if name not in in_ci
475
+ ]
476
+ return problems
477
+
478
+
254
479
  def discover_test_files(root: Path) -> list[str]:
255
480
  return sorted(
256
481
  str(p.relative_to(root))
@@ -356,6 +581,8 @@ def run_all(root: Path, fidelity_report: Path | None = None) -> list[str]:
356
581
  workflows = read_workflows(root)
357
582
  problems += check_advisory_gates_declared(workflows)
358
583
  problems += check_declared_gates_still_exist(workflows)
584
+ problems += check_every_gate_runs_on_a_pr(root, workflows)
585
+ problems += check_npm_test_covers_pr_gates(root, workflows)
359
586
 
360
587
  # check_graded_suites_graded_something shipped fully written and unit-tested
361
588
  # but unreferenced by run_all — the anti-fake-green script had a check that
@@ -0,0 +1,126 @@
1
+ #!/usr/bin/env python3
2
+ """Gate: every line a persona presents as a quotation must name where it is from.
3
+
4
+ Steps 3h and 3i of the weekly check ask whether a quoted line exists — in CBETA,
5
+ or in the compiled teachings CBETA does not hold. Neither asks whether the reader
6
+ is told *which book* it came from, and that gap is where a misattribution
7
+ survives both: master-nagarjuna's 「宁起我见积若须弥」 is real text, findable in
8
+ CBETA, and not his — it is in 《大宝积经》. "Found in the canon" and "correctly
9
+ attributed" are different questions, and only the first was being asked.
10
+
11
+ This one is offline, so it runs on every PR instead of once a week.
12
+
13
+ Attribution counts when the source is named:
14
+
15
+ - on the quoted line itself — 《书名》, 【…】, `Toh 3861`, `SC: SN 22.59`, 卷N;
16
+ - on a `> 出处:…` line a few lines below (the excerpt files put it there);
17
+ - on a sibling item of the same numbered list (one 出处 covers the block, which
18
+ is how master-fazang's and master-kumarajiva's voice samples are written);
19
+ - in the nearest section heading above (master-ajahn-chah's sutta excerpts name
20
+ the sutta in the `###` heading).
21
+
22
+ It reuses `tools/verify_sources.collect_persona_quotes`, so the set it checks is
23
+ exactly the set 3h and 3i judge — if the collector learns a new quotation shape,
24
+ this gate covers it the same day.
25
+ """
26
+ from __future__ import annotations
27
+
28
+ import os
29
+ import re
30
+ import sys
31
+ from pathlib import Path
32
+
33
+ _TOOLS = os.path.join(os.path.dirname(os.path.abspath(__file__)), os.pardir, "tools")
34
+ if _TOOLS not in sys.path:
35
+ sys.path.insert(0, _TOOLS)
36
+
37
+ try:
38
+ from verify_sources import PREBUILT_DIR, collect_persona_quotes
39
+ except ImportError as exc: # pragma: no cover - the gate must not pass silently
40
+ print(f"FAIL: cannot import the quote collector from tools/ ({exc}).")
41
+ print("A gate that cannot load what it checks has not checked anything.")
42
+ sys.exit(1)
43
+
44
+ # 什么算「写明了出处」。SC:/Toh/PTS 是南传与藏传的写法,卷N 与《…》是汉传的。
45
+ _SOURCE_NAMED = re.compile(
46
+ r"《[^》]{1,40}》|【|Toh[:\s]\d|SC[::]|\b[SMAD]N \d|Vism|PTS|卷[一二三四五六七八九十百千\d]|[Ss]utta"
47
+ )
48
+ _NUMBERED_ITEM = re.compile(r"^\s*\d+\.\s")
49
+ _LOOK = 8
50
+
51
+
52
+ def attribution(lines: list[str], number: int) -> str | None:
53
+ """这条引文的出处写在哪儿;哪儿都没写就返回 None。"""
54
+ i = number - 1
55
+ if i < 0 or i >= len(lines):
56
+ return None
57
+ if _SOURCE_NAMED.search(lines[i]):
58
+ return "on the line"
59
+
60
+ # 摘录文件把「出处」写在引文下方几行。
61
+ for j in range(i + 1, min(len(lines), i + 1 + _LOOK)):
62
+ if "出处" in lines[j]:
63
+ return f"出处 line {j - i} below" if _SOURCE_NAMED.search(lines[j]) else None
64
+
65
+ # 编号示例句:整块共用一个出处,挂在其中一条上。
66
+ if _NUMBERED_ITEM.match(lines[i]):
67
+ low = i
68
+ while low - 1 >= 0 and _NUMBERED_ITEM.match(lines[low - 1]):
69
+ low -= 1
70
+ high = i
71
+ while high + 1 < len(lines) and _NUMBERED_ITEM.match(lines[high + 1]):
72
+ high += 1
73
+ for j in range(low, high + 1):
74
+ if j != i and _SOURCE_NAMED.search(lines[j]):
75
+ return f"sibling list item {j + 1 - low}"
76
+
77
+ # 最近的小节标题。
78
+ for j in range(i - 1, max(-1, i - 1 - _LOOK), -1):
79
+ if lines[j].startswith("#"):
80
+ return f"heading {i - j} above" if _SOURCE_NAMED.search(lines[j]) else None
81
+ return None
82
+
83
+
84
+ def unattributed(quotes: list[tuple[str, str, str]], read) -> list[tuple[str, str]]:
85
+ """[(位置, 引文)]:读者无从知道出自哪部书的引文。"""
86
+ cache: dict[str, list[str]] = {}
87
+ missing: list[tuple[str, str]] = []
88
+ for where, _master, quote in quotes:
89
+ rel, _, number = where.rpartition(":")
90
+ if not number.isdigit():
91
+ continue # meta.json 的 lore_triggers 之类,不是文档行
92
+ if rel not in cache:
93
+ cache[rel] = read(rel)
94
+ if attribution(cache[rel], int(number)) is None:
95
+ missing.append((where, quote))
96
+ return missing
97
+
98
+
99
+ def main() -> int:
100
+ def read(rel: str) -> list[str]:
101
+ return Path(PREBUILT_DIR, rel).read_text(encoding="utf-8").splitlines()
102
+
103
+ quotes = collect_persona_quotes()
104
+ if not quotes:
105
+ print("FAIL: the quote collector returned nothing — this gate examined an empty set.")
106
+ return 1
107
+
108
+ missing = unattributed(quotes, read)
109
+ if missing:
110
+ print(f"FAIL: {len(missing)} quoted line(s) name no source:")
111
+ for where, quote in missing:
112
+ print(f" {where}\n 「{quote[:56]}」")
113
+ print(
114
+ "\nA reader cannot check a quotation whose book is never named, and a line\n"
115
+ "that happens to exist somewhere in the canon passes the weekly checks while\n"
116
+ "still being the wrong master's words. Name the work on the line, in the\n"
117
+ "`> 出处:` line below it, on a sibling numbered item, or in the heading above."
118
+ )
119
+ return 1
120
+
121
+ print(f"OK: all {len(quotes)} quoted lines name the work they come from.")
122
+ return 0
123
+
124
+
125
+ if __name__ == "__main__":
126
+ sys.exit(main())
@@ -0,0 +1,162 @@
1
+ #!/usr/bin/env python3
2
+ """Gate: a "read `file.md` §section" instruction must land on something.
3
+
4
+ Every persona's SKILL.md routes a question to material: a decision tree says
5
+ 读 `references/teaching.md` §参话头, a Quick Reference row says the same. The
6
+ model follows those pointers literally. Nothing checked that the section they
7
+ name exists, and the files they point into have been rewritten many times since
8
+ the routing was written — sections renamed to a book's real chapter titles,
9
+ material moved from references/ to sources/.
10
+
11
+ 2026-09-16, measured across the repo: 18 of 198 section pointers in nine personas
12
+ named a section the target file does not have, and so did one in the generator's
13
+ own SKILL.md. Most were renames (§戒律根本 for 戒律为根本), but not all were
14
+ harmless:
15
+
16
+ - master-milarepa routed 那洛六法 / 拙火 / 气脉明点 to
17
+ `sources/grubum-excerpts.md` §拙火与气脉. That file has no such section on
18
+ purpose — it lists tummo among what it deliberately does not hold, with the
19
+ answer to give instead. The pointer sent the model to look for exactly the
20
+ material the persona is forbidden to supply.
21
+ - master-mahasi-sayadaw routed 十六观智 to §十六观智 in an excerpt file whose
22
+ section is 观智次第 and which says the book numbers seventeen, not sixteen.
23
+ - master-ajahn-chah sent 心的训练 to references/teaching.md; the section is in
24
+ sources/teachings-excerpts.md.
25
+
26
+ A section counts as present when the name appears, ignoring punctuation and
27
+ spacing, in a heading of the target file or in a **bold label** — the persona
28
+ docs use bold labels (**十念法**, **数息观**) as sub-section markers, and a model
29
+ searching for the name finds either. Parenthesised glosses are optional on both
30
+ sides of the match, so §觉受与证悟 finds `6. 觉受 (nyams) 与证悟 (rtogs pa) 的区分`.
31
+ Prose may run on after a name (§版权分级 Tier B/C 流程), so trailing words split
32
+ off by spaces are dropped one at a time; the first word must still be found.
33
+
34
+ A path with a directory in it must name a file that exists even with no section.
35
+ A bare `teaching.md` without one is prose — the generator docs name the files
36
+ they produce that way — and is not a pointer.
37
+ """
38
+ from __future__ import annotations
39
+
40
+ import re
41
+ import sys
42
+ from pathlib import Path
43
+
44
+ ROOT = Path(__file__).resolve().parent.parent
45
+ PREBUILT_DIR = ROOT / "prebuilt"
46
+
47
+ # The docs the model is told to follow: every persona and mode, and the
48
+ # generator skill with its references.
49
+ DOC_GLOBS = ("prebuilt/**/*.md", "SKILL.md", "references/*.md")
50
+
51
+ _PATH = re.compile(r"`((?:[\w-]+/)*[\w.-]+\.md)`")
52
+ # A section name ends at the next separator a routing line uses: `+`, a table
53
+ # pipe, CJK punctuation, a backtick or an opening gloss.
54
+ _SECTIONS = re.compile(r"(?:[ \t、]*§[^§`++\n,、;;|。()()]*)+")
55
+ _SECTION = re.compile(r"§([^§、]*)")
56
+ _BOLD = re.compile(r"\*\*([^*\n]+)\*\*")
57
+ _GLOSS = re.compile(r"[((][^()()]*[))]")
58
+ _NUMBERING = re.compile(r"^\s*(?:\d+\.|[①-⑳])\s*")
59
+ _NOISE = re.compile(r"[\s·、,,。.::—\-–《》「」\"'`*§⚠️]")
60
+
61
+
62
+ def _norm(text: str) -> str:
63
+ return _NOISE.sub("", _NUMBERING.sub("", text)).lower()
64
+
65
+
66
+ def anchor_forms(anchor: str) -> set[str]:
67
+ """一个标题或粗体标签可被匹配的写法:连同括注,或去掉括注。"""
68
+ return {_norm(re.sub(r"[()()]", "", anchor)), _norm(_GLOSS.sub("", anchor))}
69
+
70
+
71
+ def anchors(lines: list[str]) -> list[str]:
72
+ found = [line.lstrip("#").strip() for line in lines if line.startswith("#")]
73
+ for line in lines:
74
+ found.extend(_BOLD.findall(line))
75
+ return found
76
+
77
+
78
+ def section_present(name: str, lines: list[str]) -> bool:
79
+ forms = {form for anchor in anchors(lines) for form in anchor_forms(anchor)}
80
+ words = name.split()
81
+ while words:
82
+ wanted = _norm(" ".join(words))
83
+ if wanted and any(wanted in form for form in forms):
84
+ return True
85
+ words.pop()
86
+ return not _norm(name)
87
+
88
+
89
+ def references(line: str) -> list[tuple[str, list[str]]]:
90
+ """[(路径, [小节名…])]:这一行里的每个文件指针及其所指小节。"""
91
+ out = []
92
+ for m in _PATH.finditer(line):
93
+ tail = _SECTIONS.match(line, m.end())
94
+ names = []
95
+ if tail:
96
+ names = [n.strip() for n in _SECTION.findall(tail.group(0)) if n.strip()]
97
+ out.append((m.group(1), names))
98
+ return out
99
+
100
+
101
+ def _skill_root(doc: Path) -> Path:
102
+ rel = doc.relative_to(ROOT)
103
+ return ROOT / rel.parts[0] / rel.parts[1] if rel.parts[0] == "prebuilt" else ROOT
104
+
105
+
106
+ def resolve(doc: Path, rel: str) -> Path | None:
107
+ for base in (_skill_root(doc), doc.parent, ROOT):
108
+ if (base / rel).is_file():
109
+ return base / rel
110
+ return None
111
+
112
+
113
+ def dangling(docs: list[Path]) -> tuple[int, list[str]]:
114
+ """(检查过的指针数, [问题…])。"""
115
+ examined, problems = 0, []
116
+ cache: dict[Path, list[str]] = {}
117
+ for doc in docs:
118
+ for number, line in enumerate(doc.read_text(encoding="utf-8").splitlines(), 1):
119
+ if "{" in line:
120
+ continue # 模板行,如 `prebuilt/{slug}/…` 下的 `references/teaching.md`
121
+ where = f"{doc.relative_to(ROOT)}:{number}"
122
+ for rel, names in references(line):
123
+ if "/" not in rel and not names:
124
+ continue # 裸文件名是行文,不是指针
125
+ examined += 1
126
+ target = resolve(doc, rel)
127
+ if target is None:
128
+ problems.append(f"{where} `{rel}` does not exist")
129
+ continue
130
+ if target not in cache:
131
+ cache[target] = target.read_text(encoding="utf-8").splitlines()
132
+ for name in names:
133
+ if not section_present(name, cache[target]):
134
+ problems.append(f"{where} `{rel}` has no section §{name}")
135
+ return examined, problems
136
+
137
+
138
+ def model_facing_docs() -> list[Path]:
139
+ return sorted({p for pattern in DOC_GLOBS for p in ROOT.glob(pattern) if p.is_file()})
140
+
141
+
142
+ def main() -> int:
143
+ examined, problems = dangling(model_facing_docs())
144
+ if examined == 0:
145
+ print("FAIL: found no file pointers at all — this gate examined an empty set.")
146
+ return 1
147
+ if problems:
148
+ print(f"FAIL: {len(problems)} of {examined} file pointer(s) do not land:")
149
+ for problem in problems:
150
+ print(f" {problem}")
151
+ print(
152
+ "\nThe model follows these pointers literally. A section that is not there\n"
153
+ "leaves it to improvise the material it was sent to read. Point at a heading\n"
154
+ "or **bold label** the file actually has, or drop the § and name the file."
155
+ )
156
+ return 1
157
+ print(f"OK: all {examined} file pointers land on a file and section that exist.")
158
+ return 0
159
+
160
+
161
+ if __name__ == "__main__":
162
+ sys.exit(main())