eduevidence 6.2.0 → 6.3.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (114) hide show
  1. package/CHANGELOG.md +395 -0
  2. package/README.md +22 -13
  3. package/README.zh-CN.md +15 -8
  4. package/SKILL.md +10 -9
  5. package/benchmarks/evidence-library.json +277 -1
  6. package/docs/architecture.md +6 -3
  7. package/docs/j-ev-experimental.md +250 -0
  8. package/docs/reproducibility.md +138 -0
  9. package/domains/_neutral/copy/few_shots.json +21 -0
  10. package/domains/_neutral/copy/framing_lexicon.json +19 -0
  11. package/domains/_neutral/copy/module_labels.json +5 -0
  12. package/domains/_neutral/copy/module_labels_footer.json +102 -0
  13. package/domains/_neutral/copy/module_labels_modules.json +204 -0
  14. package/domains/_neutral/copy/module_labels_nav.json +126 -0
  15. package/domains/_neutral/copy/module_labels_summary.json +98 -0
  16. package/domains/_neutral/copy/module_labels_tables.json +164 -0
  17. package/domains/_neutral/copy/module_labels_v2.json +90 -0
  18. package/domains/_neutral/copy/risk_constructs.json +20 -0
  19. package/domains/_neutral/copy/section_titles.json +66 -0
  20. package/domains/_neutral/copy/terminology.json +11 -0
  21. package/domains/check_copy_packs.py +103 -0
  22. package/domains/education/copy/few_shots.json +22 -0
  23. package/domains/education/copy/framing_enums.json +167 -0
  24. package/domains/education/copy/framing_lexicon.json +166 -0
  25. package/domains/education/copy/module_labels.json +169 -0
  26. package/domains/education/copy/risk_constructs.json +48 -0
  27. package/domains/education/copy/section_titles.json +186 -0
  28. package/domains/education/copy/terminology.json +70 -0
  29. package/domains/education/manifest.json +1 -1
  30. package/domains/education/outcome_taxonomy.json +2 -2
  31. package/domains/manifest.json +1 -1
  32. package/domains/policy/copy/few_shots.json +22 -0
  33. package/domains/policy/copy/framing_enums.json +94 -0
  34. package/domains/policy/copy/framing_lexicon.json +174 -0
  35. package/domains/policy/copy/module_labels.json +168 -0
  36. package/domains/policy/copy/risk_constructs.json +33 -0
  37. package/domains/policy/copy/section_titles.json +186 -0
  38. package/domains/policy/copy/terminology.json +64 -0
  39. package/engine/capabilities.py +57 -5
  40. package/engine/decision_policy.py +88 -17
  41. package/engine/library_builtin.py +7 -4
  42. package/engine/tribunal.py +17 -23
  43. package/engine/versions.py +1 -1
  44. package/examples/ai-coding-assistant-evidence/EduEvidence_Report.html +4 -4
  45. package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_academic.html +4 -4
  46. package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_claude.html +4 -4
  47. package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_datalab-dark.html +4 -4
  48. package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_datalab.html +4 -4
  49. package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_presentation.html +4 -4
  50. package/examples/spaced-retrieval-practice/EduEvidence_Report.html +2728 -0
  51. package/examples/spaced-retrieval-practice/report.html +2522 -0
  52. package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_academic.html +4 -4
  53. package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_claude.html +4 -4
  54. package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_datalab-dark.html +4 -4
  55. package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_datalab.html +4 -4
  56. package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_presentation.html +4 -4
  57. package/examples/workplace-ai-assistant/EduEvidence_Report.html +2814 -0
  58. package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_academic.html +36 -36
  59. package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_claude.html +36 -36
  60. package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_datalab-dark.html +36 -36
  61. package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_datalab.html +36 -36
  62. package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_presentation.html +36 -36
  63. package/integrations/jev/__init__.py +115 -0
  64. package/integrations/jev/approval.py +212 -0
  65. package/integrations/jev/cli.py +84 -0
  66. package/integrations/jev/config.py +112 -0
  67. package/integrations/jev/gateway.py +128 -0
  68. package/integrations/jev/modes.py +38 -0
  69. package/integrations/jev/tools_classify.py +88 -0
  70. package/integrations/jev/tools_extract.py +111 -0
  71. package/integrations/jev/tools_rerank.py +71 -0
  72. package/integrations/jev/tools_screen.py +87 -0
  73. package/integrations/jev/tools_verify.py +95 -0
  74. package/integrations/jev_mcp.py +22 -0
  75. package/integrations/semantic_decide.py +286 -0
  76. package/integrations/semdecide_cli.py +55 -0
  77. package/package.json +9 -1
  78. package/pyproject.toml +1 -1
  79. package/references/report-copy-style.md +43 -3
  80. package/schemas/v2/decision-snapshot.schema.json +20 -9
  81. package/schemas/v2/intake.schema.json +191 -0
  82. package/scripts/build_evidence_library.py +15 -5
  83. package/scripts/dashboard_server.py +13 -2
  84. package/scripts/intake/__init__.py +31 -0
  85. package/scripts/intake/__main__.py +18 -0
  86. package/scripts/intake/background.py +78 -0
  87. package/scripts/intake/browser.py +79 -0
  88. package/scripts/intake/cli.py +57 -0
  89. package/scripts/intake/constants.py +57 -0
  90. package/scripts/intake/depth.py +53 -0
  91. package/scripts/intake/enhancements.py +106 -0
  92. package/scripts/intake/hooks.py +90 -0
  93. package/scripts/intake/prefs.py +76 -0
  94. package/scripts/intake/prompts.py +85 -0
  95. package/scripts/intake/session.py +152 -0
  96. package/scripts/lint_file_layers.py +126 -0
  97. package/scripts/orchestrator.py +68 -17
  98. package/scripts/pre_verdict_gate.py +21 -7
  99. package/scripts/skill_lint.py +11 -1
  100. package/scripts/skill_payload.py +3 -3
  101. package/scripts/test_adversarial_empirical.py +70 -6
  102. package/skill/agents/evidence-judge.md +49 -7
  103. package/skill/workflows/experimental-jev.md +170 -0
  104. package/skill/workflows/intake.md +120 -0
  105. package/visualization/eduevidence-report/scripts/build_infographics.py +32 -14
  106. package/visualization/eduevidence-report/scripts/build_report.py +75 -662
  107. package/visualization/eduevidence-report/scripts/report_copy_pack.py +296 -0
  108. package/visualization/eduevidence-report/scripts/report_copy_policy_guard.py +47 -0
  109. package/visualization/eduevidence-report/scripts/zh_labels.py +61 -0
  110. package/scripts/build_esl_artifacts.py +0 -1921
  111. package/scripts/build_killer_demo.py +0 -295
  112. package/scripts/enrich_projects_human_and_lieflat.py +0 -315
  113. package/scripts/generate_new_projects.py +0 -686
  114. package/scripts/sync_killer_demo_report.py +0 -270
package/CHANGELOG.md ADDED
@@ -0,0 +1,395 @@
1
+ # Changelog
2
+
3
+ 所有显著变更均记录于此。格式基于 [Keep a Changelog](https://keepachangelog.com/zh-CN/1.0.0/);版本号遵循 [Semantic Versioning](https://semver.org/lang/zh-CN/)。
4
+
5
+ ## [6.3.0] — 2026-09-22
6
+
7
+ > Multi-domain copy packs and the two-round intake, the experimental Jev/SemDecide overlay,
8
+ > one decision-funnel authority, and the post-audit fixes.
9
+ > Fix: the four-state decision could not reach ADOPT through the V1 path, and all three
10
+ > public cases read PILOT regardless of their evidence.
11
+
12
+ ### The ADOPT path is reachable, and provable
13
+ - New `engine/decision_policy.py` is the single authority for the ADOPT gate (primary
14
+ outcome categories per domain, directness threshold, confidence band). `engine/tribunal.py`
15
+ and `scripts/pre_verdict_gate.py` both import it, so the V2 adjudicator and the V1 gate can
16
+ no longer drift apart.
17
+ - `engine/migration.py` no longer hard-codes `directness: 1` on every migrated link. Directness
18
+ and `applicability.scope_match` are derived from the source record's D5 Directness, with an
19
+ explicit downgrade recorded when D5 is absent. A migrated pack used to be structurally
20
+ incapable of ADOPT; `spaced-retrieval-practice` now adjudicates to **ADOPT / High (0.893)**.
21
+ - The Pre-Verdict Gate gained item 12, `decision_action_consistency`: a pack cannot award
22
+ itself ADOPT. The gate re-derives primary-outcome directness from the evidence corpus and
23
+ caps an unsupported adopt to pilot. The gate is now 12 items.
24
+ - `outcome_mapping` blocks High only when a PRIMARY outcome has no evidence; an unmeasured
25
+ secondary or risk outcome is a scope note, since missing evidence is not a zero effect.
26
+ - The gate resolves a pack's domain from `run_manifest.json`, then `frame.extensions.domain`,
27
+ then `result.meta.domain`, instead of assuming education. Policy packs no longer fail their
28
+ own frame check.
29
+
30
+ ### Public cases corrected and re-adjudicated
31
+ - `spaced-retrieval-practice`: claims bound to evidence ids, deterministic confidence and the
32
+ action bound applied, narratives re-written for adoption. Verdict **ADOPT / High / 0.893**.
33
+ - `ai-coding-assistant-evidence`: challenge stage written (`skeptic.json`, nine checks grounded
34
+ in the corpus), `final_verdict.json` added, policy version refreshed to `2026-08-12.v3`.
35
+ Verdict unchanged at **PILOT / Moderate / 0.586** — the evidence still stops at task performance.
36
+ - `workplace-ai-assistant`: per-record quality dimensions added, the pack now declares the
37
+ policy domain's own outcome tokens (`policy_effectiveness` / `implementation_risk`) with the
38
+ teaching-neutral tokens kept in `extensions`, challenge record written, scope calibration
39
+ fields and claim bindings added. Verdict **PILOT / Moderate / 0.578**.
40
+ - `schemas/v2/{study,finding,methodology-audit}.schema.json` accept the `ST-` study-id prefix so
41
+ curated packs migrate without being renamed.
42
+ - All three packs pass the Pre-Verdict Gate (`passed=true`) and record a `gate_report.json`;
43
+ reports, five themes, Studio snapshot, Pages export and the submission package were rebuilt.
44
+ The submission package now ships all three examples, including the ADOPT case.
45
+
46
+ ### Multi-domain copy packs and the two-round intake
47
+ - Domain copy packs land as data: `domains/{education,policy,_neutral}/copy/*.json`
48
+ (framing lexicon, section titles, module labels, risk constructs, terminology,
49
+ few-shots). `report_copy_pack`/`report_copy_policy_guard` in the report layer load
50
+ them, so report chrome and infographics no longer hardcode education vocabulary;
51
+ `domains/check_copy_packs.py` asserts the policy pack stays free of teaching words.
52
+ - `scripts/intake/` is the one-shot two-round intake (research question → execution
53
+ enhancements → depth → Frame boundary confirmation) with `schemas/v2/intake.schema.json`
54
+ as its contract and `skill/workflows/intake.md` as the agent-facing workflow.
55
+ `orchestrator run` gained `--yes` / `--enhancement` / `--depth` and writes `intake.json`
56
+ into the run workspace; non-question preferences persist to `~/.eduevidence/prefs.json`.
57
+
58
+ ### Jev / SemDecide Tier-0 overlay (experimental, opt-in)
59
+ - `integrations/jev/` (detect → approval → gateway, plus the five Tier-0 tools) and
60
+ `integrations/semantic_decide.py` (`semdecide is/filter/choose`) add an optional
61
+ acceleration layer behind `--experimental 0|1|2|3`. It never replaces the Pre-Verdict
62
+ Gate, the nine Skeptic checks or the Tribunal: uncertain/invalid/provider-failed
63
+ results fail closed and escalate to the large model.
64
+ - The five experimental capabilities are registered in
65
+ `engine/capabilities.experimental_capability_registry()` and stay out of the default
66
+ scientific registry, so role-gate checks are unaffected.
67
+ - `docs/j-ev-experimental.md` records the speedup口径, the A/B acceptance table and the
68
+ degradation ladder; `skill/workflows/experimental-jev.md` is the workflow overlay.
69
+
70
+ ### Decision funnel, contracts and packaging
71
+ - The four-state landing is now derived from one authority everywhere: `engine/tribunal.py`
72
+ keeps no private copy of the rule, `scripts/pre_verdict_gate.py` imports it, and
73
+ `schemas/v2/decision-snapshot.schema.json` publishes the same four tokens.
74
+ - `scripts/check_versioned_schemas.py` validates every vNext contract against the payload
75
+ its real producer emits (28 checks), including the autoevolve session report and the v3
76
+ benchmark run manifest.
77
+ - `web/landing.html` states the schema count the repository actually ships, and the npm
78
+ tarball no longer leaks `__pycache__`/`*.pyc` (`**/` patterns in `.npmignore`).
79
+ - The landing page claim is now machine-checked: `studio/verify-schema-count.mjs` derives
80
+ the expected count from `scripts/generate_metrics.py` and runs as
81
+ `npm run verify:schema-count` in the Research Studio CI job.
82
+
83
+ ### Fixes after the 2026-09-22 audit
84
+ - `integrations/semantic_decide.py` was missing `import json` (five call sites in the
85
+ JSON payload and JSONL filter paths) and had no CLI entry; both are fixed, and
86
+ `python3 integrations/semantic_decide.py --experimental 3` now prints the detect JSON.
87
+ - `integrations/jev/approval.is_approval_current()` referenced an undefined `path`
88
+ (NameError). It now takes an explicit `path` argument; the default still re-reads the
89
+ configured approval file, so a missing or unreadable file never counts as current.
90
+ - Intake entry points: `scripts/intake/cli.py` gained its `__main__` guard (the documented
91
+ `python3 -m intake.cli --print-prompts` used to exit 0 with no output) and the package
92
+ now imports its siblings relatively, so `python3 -m intake.cli`, the direct script form
93
+ and `import scripts.intake.*` all work.
94
+ - `scripts/test_adversarial_empirical.py` now drives the class the server actually ships
95
+ (`ThreadingHTTPServer`, 128-slot backlog) instead of a single-threaded `TCPServer` whose
96
+ five-slot accept queue made the kernel reset 24 of 30 concurrent connects. The burst now
97
+ asserts all 30 requests succeed, and the host / `Sec-Fetch-Site` guards are asserted with
98
+ raw requests (403 untrusted host, 403 cross-site, 200 same-origin).
99
+ - `docs/metrics.json` was stale (922/102/49 against 928/103/50); the metrics gate is green
100
+ again, and `docs/reproducibility.md`, `CHANGELOG.md`, `docs/j-ev-experimental.md` plus the
101
+ two missing example `EduEvidence_Report.html` files are now in the npm `files` allowlist.
102
+ - Hygiene: removed the retired `scripts/sync_killer_demo_report.py` (it imported the already
103
+ deleted synthetic generator), dropped `files` excludes that named deleted files, deleted
104
+ the shadowed decision path in `engine/tribunal.py`, corrected the "11-item" Pre-Verdict
105
+ Gate strings to 12, fixed the stale `engine/pilot.py _OUTCOME_CATEGORY` pointers in the
106
+ domain manifests, cleaned unused imports/variables, and restored relative report paths in
107
+ `web/api/projects.json` (18 host-absolute paths) with a CI guard so they cannot return.
108
+ - New offline tests: `tests/test_semdecide_wrappers.py`, `tests/test_jev_approval_gate.py`,
109
+ `tests/test_intake_flow.py`; CI coverage now includes `integrations` and `scripts.intake`.
110
+
111
+ ## [6.2.0] — 2026-09-12
112
+
113
+ > Audit remediation: the domain registry becomes the single authority, the report
114
+ > first screen is written by the adjudicator, and the delivery pipeline proves parity.
115
+
116
+ ### Domain registry as the single authority
117
+ - New `engine/taxonomy.py` reads `domains/<id>/outcome_taxonomy.json`; unknown tokens fail
118
+ closed instead of defaulting to a learning outcome.
119
+ - `engine/pilot.py`, `engine/tribunal.py`, `engine/gaps.py`, `scripts/pre_verdict_gate.py`,
120
+ `scripts/claim_audit.py` and `scripts/build_result.py` no longer keep private copies of
121
+ the token list or the token-to-category map.
122
+ - The frame stage resolves its schema from the run domain; `run` gains `--domain`; a policy
123
+ run validates against `domains/policy/frame.schema.json` instead of the education one.
124
+ - `check_protocol_alignment.py` gained the taxonomy dimension: schema enums, ADOPT gate
125
+ categories and V2 buckets must all agree with the registry.
126
+
127
+ ### Report copy is written, not assembled
128
+ - `verdict.schema.json` declares `strongest_support` / `key_uncertainty` / `main_risk` /
129
+ `next_action` as reader-facing prose; the judge prompt carries the writing contract.
130
+ - The renderer no longer synthesises the first screen from claim fragments.
131
+ - New `references/report-copy-style.md`; the language gate enforces length ceilings and
132
+ structural parallelism between the en and zh evidence sets.
133
+ - Chinese example data corrected: population and claim were rotated across a three-item
134
+ block, and audit notes cited a non-existent evidence id.
135
+
136
+ ### Contract repair
137
+ - Removed the deprecated `direction` from the result evidence contract; aligned
138
+ `report-spec` / `chart-spec` with their real producers; declared the effect-size family.
139
+ - `EvidenceNode.effect_size` no longer defaults to a fabricated zero effect.
140
+ - Added `skeptic.schema.json` / `applicability.schema.json` and wired both into the stage
141
+ gates; the counter-evidence gate now requires all nine checks.
142
+ - Sciverse locators survive the Source contract; `study_audits` gained a producer; V1 to V2
143
+ migration keeps effect magnitudes; the confidence policy version has one authority.
144
+ - The validator implements anyOf / oneOf / allOf, so alternative-shape contracts are real.
145
+
146
+ ### Gates and delivery
147
+ - The red-team suite gained seven real assertions (it previously never failed) and it
148
+ immediately exposed a DID column-mapping defect, now fixed.
149
+ - New `scripts/check_package_parity.py` proves the shipped package is byte-identical to
150
+ the source tree; CI runs it, and also runs the skill linter and covers `visualization/`.
151
+ - Docs, npm ignores and a Python-version guard corrected; audit checklist in
152
+ `docs/audit-2026-09-12.md`.
153
+
154
+ ## [6.1.0] — 2026-09-12
155
+
156
+ > Content and protocol depth: a citation-grade retrieval channel, runnable
157
+ > runbooks, and a mechanical alignment gate.
158
+
159
+ ### Sciverse retrieval channel (key-based academic)
160
+ - New `retrieval/sciverse.py`: `/meta-search`, `/agentic-search`, `/content` and
161
+ `/meta-paper-relations` with typed failure statuses, Unicode code-point
162
+ locators and no credential leakage. Inactive without `SCIVERSE_API_TOKEN`.
163
+ - `SearchHit` gains optional `doc_id` / `chunk_id` / `offset` / `unique_id`;
164
+ `MultiSearchRouter` runs the key-based academic channel ahead of the
165
+ zero-config ones and reports it in `get_provider_status()`.
166
+ - `retrieval/fetch.py::fetch_sciverse_content()` expands a chunk locator into a
167
+ FetchResult-shaped record, so RULE 2 (snippet ≠ evidence) is machine-enforced:
168
+ a chunk must be read through `/content` and pass the validation gate first.
169
+ - Audit exports carry the locator: `chunks.jsonl` plus `doc_id`/`chunk_id`/`offset`
170
+ columns in `source-screening.csv`; DOI-less records are flagged
171
+ `needs_manual_location` instead of receiving a fabricated URL.
172
+ - New `docs/sciverse-api.md`, `references/retrieval-compliance.md` (robots,
173
+ rate limits, paywalls, attribution, credentials) and Sciverse rules in
174
+ `references/retrieval-protocol.md`.
175
+
176
+ ### Runnable content depth
177
+ - The three user workflows became full runbooks (step tables, gates, failure
178
+ handling, human hand-off points, resume semantics, acceptance checklists).
179
+ - All stage briefs were rebuilt on one template (goal / prerequisites / artifacts
180
+ + schema / rules / quality gates / failure modes / language contract / hand-off).
181
+ - All 12 sub-skill recipes were rebuilt on one template (When to Use / Inputs /
182
+ Process / Output Contract / Quality Gates / Anti-Patterns / Worked Example /
183
+ References) and mapped to engine capability IDs via frontmatter.
184
+ - Role prompts declare `role_id` / `capabilities` / `output_contracts` /
185
+ `recommended_reasoning` instead of binding `default_cli` / `default_model`,
186
+ and gain independence + failure-mode sections; independence is graded
187
+ (`different-model-family` for the skeptic, `role-separation` for the reviewer).
188
+ - `skill/roles/registry.yaml` capabilities now use the engine capability IDs.
189
+
190
+ ### Mechanical alignment
191
+ - New `scripts/check_protocol_alignment.py` (+ `tests/test_protocol_alignment.py`,
192
+ wired into CI): stages ↔ briefs ↔ roles ↔ prompts ↔ capabilities ↔ sub-skills ↔
193
+ workflows ↔ packaging must agree, role prompts may not bind model/CLI names,
194
+ and packaging versions must equal `ENGINE_VERSION`.
195
+ - New `CONTRIBUTING.md` (gates, scientific invariants, how to add a capability /
196
+ sub-skill / role / retrieval channel).
197
+ - Fixed real drift: `retrieval/search.py` advertised four providers that do not
198
+ exist, `packaging/scp-manifest.json` was pinned at 4.0.0, and two docs carried
199
+ a stale test count.
200
+
201
+
202
+ ## [6.0.0] — 2026-08-31
203
+
204
+ > Competition kernel convergence: Workflow → Capability → Contract.
205
+
206
+ ### Decision-grade control plane
207
+ - Canonical runtime protocol now includes Applicability; report generation is a Projection outside scientific stages.
208
+ - Added three user-facing workflow entries, durable Project/Run event storage, content-addressed immutable artifacts, and a judge-pack export.
209
+ - Added auditable search plans, bounded provider-attempt records, counter-evidence query coverage, and screening exports.
210
+ - Aligned submission CI with `dist/eduevidence-submission` and reduced report H1 scale across all five baked themes.
211
+ - Root `SKILL.md` replaced with the decision-grade English operating contract (workflow routing, scientific gates, applicability / projection semantics) plus a bilingual tooling contract appendix that keeps lint / version / metrics / release gates stable.
212
+ - Startup approval loop: on `run`, recommend a role to CLI to model table built only from scanned, user-usable models, ask for confirmation, and persist the hash-verified mapping to `~/.eduevidence/agent_mcp_approval.json` for reuse by later runs.
213
+ - Fail-closed model discipline: `safe_spawn` / `cross_model_review` refuse any unapproved CLI/model; benchmark and judge drivers no longer ship a default model (explicit `--model` or `EDUEVIDENCE_LLM_MODEL` required).
214
+
215
+ ## [5.2.0] — 2026-08-24
216
+
217
+ > 可信度修复版:示例 provenance 纠偏 + 版本/口径单一权威(依据 docs/plans/v5.2-v6.0-iteration-plan.md)。
218
+
219
+ ### 版本治理
220
+ - `engine/versions.py` 成为唯一版本权威;`engine/__init__.py` 的死值 2.0.0 改为 re-export。
221
+ - 新增 `scripts/check_version_consistency.py` 并纳入 CI:versions.py ↔ pyproject ↔ CHANGELOG 头条目 ↔ SKILL.md 标题一致性校验。
222
+
223
+ ### 示例 provenance 纠偏(随本版后续提交补充)
224
+ - 全量 DOI 审计(`scripts/audit_dois.py`,Crossref/DataCite 双注册表):88 唯一 DOI 仅 16 真实。
225
+ - 删除伪造旗舰包 `ai-coding-assistant-50`;新建 `ai-coding-assistant-evidence`:
226
+ 8 个注册表核验来源、12 条证据、引擎确定性置信度(Moderate/0.586)、五主题烘焙 ALL CLEAN。
227
+ - esl/math 示例:删除伪造的 agent_mcp_enhanced 运行记录,如实标注 `data_origin=synthetic`,
228
+ 并首次通过 schema 契约(修复 CLM-### claim_id、非法 venue/statement 等历史违规)。
229
+ - 报告头新增 data_origin 徽章;synthetic/hybrid 高亮警示。
230
+ - 口径统一:`docs/metrics.json` 单一权威(752→762 测试函数、37 schema、15 参考、30 金标),
231
+ CI 校验 README/SKILL/install-guide 数字;README 语言命名归位(README.zh-CN.md);
232
+ landing 页清除虚构统计("50 篇顶刊"/+0.38g)与主办方泄漏词;Skill #650 失效外链移除。
233
+ - 可信度功能化:`engine/citation_check.py` + CLI(--write-back 回写 sources.jsonl)、
234
+ schema 增 doi_verified/retracted、来源表 DOI ✓ / RETRACTED 徽章;
235
+ `scripts/retraction_watch.py` 定期撤稿监控。
236
+ - 工程地基:红队压测纳入默认收集并修复至绿(StudioHandler 适配+临时端口)、
237
+ ruff(E9/F) 守门、pytest-cov 覆盖率报告、移除 CI 静默重试、wheel 隔离安装冒烟、
238
+ stdlib logging 贯穿 fetch/search/orchestrator(EDUEVIDENCE_LOG_LEVEL 开关)、
239
+ install.sh 供应链与副作用披露、web archive 快照删除+单一源声明、
240
+ quickstart 30 分钟引导器、Deep Research 对比页。
241
+
242
+
243
+ ## [5.1.1] — 2026-08-23
244
+
245
+ > 排版修复 + 五主题布局约束门(含手机端)。
246
+
247
+ ### Judge(presentation)主题排版问题排查与修复
248
+ - 根因一(移动端/平板横向裁切):`:root[data-theme="X"]`(特异性 0,2,0)压过基座
249
+ `@media (max-width:980px)`(0,1,0)——presentation 的 `full-report-layout` 在 390px 仍为
250
+ `230px 104px` 两列,内容列被挤到 104px;datalab/dark/presentation 的 `outcome-groups`
251
+ `repeat(auto-fit,minmax(240px,1fr))` 内在尺寸爆炸,768px 平板第二列结果消失在视口外
252
+ (shell scrollWidth 1025 > 768)。
253
+ - 修复范式:基座断点改 `minmax(0,1fr)` + 网格 item `min-width:0` 安全网;全部 5 主题
254
+ `auto-fit` 轨道改 `minmax(min(Npx,100%),1fr)`;三主题 `report-page-brief`/`full-report-layout`
255
+ 自带 `@media ≤980` 覆写(同主题文件内,等特异性后声明生效);`scope-grid`/`method-audit-grid`
256
+ 同样处理。
257
+ - 结果:18 份报告(3 示例 × 5 主题 + 3 主报告 + 13 篇示例)× 390/768/1280 × brief/full
258
+ 浏览器实测 114 项全部无溢出、无裁切。
259
+
260
+ ### 展开动画「只有点击才触发」体验修复
261
+ - 根因:首屏卡片在页面加载瞬间即播完动画(用户未看到),滚动中又常被错过,感知为
262
+ "只有点击才动"。`motion.js` 新增入场错峰:页面加载 1.5s 内命中的卡片按
263
+ `140ms + (index%6)*130ms` 逐个播放入场(draw-in 可见);滚动进入与点击重播立即执行;
264
+ `prefers-reduced-motion`/无 JS 仍然静态可见。
265
+
266
+ ### 五主题排版约束机制(skill + 脚本 + 测试)
267
+ - 新建 `scripts/lint_report_layout.py`:静态不变量审计(裸 1fr 轨道、无移动端覆写的双列
268
+ grid、固定 px 最小值 auto-fit 均为违规)+ 可选浏览器级门(调 `check_mobile_layout.js`)。
269
+ - 新建 `visualization/eduevidence-report/scripts/check_mobile_layout.js`:零依赖 Node CDP
270
+ 实测(390/768/1280 × brief/full):页面无横向溢出、可见 shell 无裁切、逃逸元素排除
271
+ 滚动容器/SVG 内部、画廊 reveal 契约(滚入全部 is-live)。
272
+ - 新建 `tests/test_report_layout_mobile.py`:静态门必跑;浏览器级门在有 Chrome+Node 时跑,
273
+ 否则 skip;reveal 契约静态断言。
274
+ - 文档:新增 `visualization/eduevidence-report/references/layout-constraints.md`(守则 + 教训 +
275
+ 验收清单);`component-catalog.md` §10/§11 与 `report-generation` sub-skill §2.4 引用守则。
276
+
277
+ ## [5.1.0] — 2026-08-22
278
+
279
+ > Present 层大改造:**AI 自由组合 Lieflat 图表 + 5 套主题文案排版优化(含展开动画)**。
280
+
281
+ ### 数据驱动 Lieflat 画廊(AI 自由组合)
282
+ - 新增 `visualization/eduevidence-report/scripts/charts_data.py`:16 个提取器(meta.forest / evidence.ranked_effects / year_x_dimension / grouped_distribution / multidim_top / study_type & wwc_composition / outcomes.direction_counts & paired_counts & bipolar_axes / intervention.phase_weeks & activity_weights & phase_groups / decision.confidence_score / methodology.flag_rates / year_x_outcome_counts),数据不足返回 None + reason(镜像 Meaningful Visualization Gate),不编造单位。
283
+ - 重构 `lieflat_engine.py`:REGISTRY(17 个 type ↔ 目录编号 ↔ 提取器 ↔ 渲染器)+ `render_figure` 调度器,未知 type 显式报错;**删除全部硬编码演示数据与 SVG 内嵌 `<style>`**,改为 `lf-pop/lf-fade/lf-draw` 类 + `--motion-delay`(点阵 12ms / 条形 100ms);SVG 最小字号 6.5px、数值 800、面积 sqrt。
284
+ - `build_figures.py`:学术图保持不变;Lieflat 部分只渲染 `resolve_visual_layout` 校验通过的条目,键用条目 chart_id。
285
+ - `build_report.py`:新增 `resolve_visual_layout`(新契约 `{chart_id, type, catalog_ref, title_zh/en, subtitle_zh/en, caption_zh/en, source, params}`;兼容旧 title/subtitle 双语共用并告警;未注册/缺双语/参数非法 → 丢弃 + 原因入 report_spec;全无效 → 确定性安全组合 forest + dot_cascade + bubble_almanac + tick_rows);完整性门新增 `lieflat_data_bound`(渲染值逐一比对提取器 bundle);`render_lieflat_gallery_brief` 重写(四件套卡片 + `data-lieflat`/`data-visual` + 图注 + 抑制清单);`report_spec.json` 增加 `lieflat_gallery`。
286
+ - 新增 `schemas/visual-layout.schema.json`(目录原为空,正放新契约)。
287
+
288
+ ### 展开动画对齐 skill 正本(图表 SVG 为主)
289
+ - `motion/motion.css`:新增 `data-lieflat` 区块——`.js-lf` 门控的 `lf-pop`(scale 0→1,cubic-bezier(.2,.7,.3,1.3),500ms)/ `lf-fade`(900ms ease)/ `lf-draw`(dasharray 1,1s cubic-bezier(.4,0,.2,1)),`--motion-delay` stagger,reduced-motion 与 print 全关,无 JS 静态可见。
290
+ - `motion/motion.js`:`[data-lieflat]` 实现 mono-tokens `obsReveal` 语义——IntersectionObserver(threshold .3)滚入播放一次;点击重播(先清该 id 已登记 timer,防叠加);`CSS.escape` 安全。
291
+ - `references/motion-system.md`:`chart-reveal` 写为固定模板角色,五主题不得另造动画逻辑。
292
+
293
+ ### 5 套主题文案排版优化(设计个性全部保留)
294
+ - 文案:表头 meta 行规范为「模式:… · 生成时间:… · 证据 N 条 · 来源 N 个」(英文对应);润色 `full_report_intro`、12 条 section leads、brief 各块 lead、`benchmark_note`;图表文案规范(标题写结论、副标题说清图例单位、来源行全大写)。
295
+ - 排版:中文优先字体栈(PingFang SC / Noto Sans CJK,academic 保留 Songti 衬线);中文正文不加 letter-spacing;数字 `tabular-nums`;claude 标题 500 字重与 measure 收窄;academic 打印 8.5pt;datalab/datalab-dark 4 列 brief 网格 + 章头双栏 + 控件尺寸统一(暗版对比度 ≥4.5:1 复核);presentation 裁决字号阶梯 + insight 内边距 + 金橙对比度复核。
296
+
297
+ ### 夹具、测试与交付
298
+ - 修复 `examples/ai-coding-assistant` 夹具(决策叙述中的 E-xxx 引用 / null 残留 / overall_risk schema 键 → 语言门禁 0 问题);三个示例项目 `visual_layout` 升级为新双语契约(50 篇示例展示 L20/G15/L14/F11/L15/L16 六种轮廓组合)。
299
+ - 新增 `tests/test_lieflat_composition.py`(提取器溯源、注册表拒绝、抑制、无演示值、schema、数据溯源门);`test_build_report_html.py` 增补画廊卡片 / motion 定义唯一 / reveal 契约 / footer / meta 行;`test_build_figures.py` 按 layout 渲染。
300
+ - 全量 pytest 与 `skill_lint.py` 全绿;`rebake_all_5themes.py` 重烘焙 3 示例 × 5 主题(15 份,integrity 全 PASS)。
301
+
302
+ ### 文档
303
+ - `skill/sub-skills/report-generation/SKILL.md` §2 重写为六步组合工作流;新增 `references/lieflat-composition.md`(注册表 + 契约 + 推荐组合 + 抑制规则);更新 `component-catalog.md` / `full-report-outline.md` / 根 `SKILL.md` Present 行 / `docs/architecture.md` Present 管线。
304
+
305
+ ## [5.0.0] — 2026-08-21
306
+
307
+ > v5 大迭代旗舰版(`ENGINE_VERSION = "5.0.0"`)。主题:**「SSOT 证据图谱 × 杀手级实证闭环 × Claude 本地控制台」**。
308
+
309
+ ### 统一 SSOT 证据图谱引擎 (`engine/evidence_graph.py`)
310
+ - **7 大统一实体节点模型**:`PaperNode` (文献)、`EvidenceNode` (量化效应量 Hedges' g、WWC 5.0 评级)、`OutcomeNode` (5维社科分类)、`ClaimNode` (科学主张)、`RiskNode` (方法学陷阱)、`GapNode` (学术空白)、`DecisionNode` (四态裁决快照)
311
+ - **消除数据孤岛**:所有下游消费端(Tribunal 仲裁、HTML 报告、Web 控制台、GapLens)统一读取此 SSOT 图模型
312
+ - **ECharts 力导向图导出**:支持可视化拓扑导出与交互式路径追踪
313
+
314
+ ### 50 篇真实学术文献杀手级 Demo (`examples/ai-coding-assistant-50/`)
315
+ - **核心命题**:“高校大一程序设计课程引入 AI 编程助手是否真正提升计算思维与独立编程能力?”
316
+ - **揭示社科核心悖论**:任务完成速度提升 $+0.64g$ vs 撤除 AI 后的独立闭卷期末考迁移赤字 $-0.28g$
317
+ - **WWC 5.0 脚手架依赖陷阱检测**:自动触发认知陷阱告警,裁决为 `PILOT` (限制性试点)
318
+ - **12周准实验因果闭环**:自动生成 12 周 DID 准实验设计,支持课堂 CSV 数据回注因果回归
319
+
320
+ ### Claude 风格本地控制台 (`scripts/dashboard_server.py`)
321
+ - **浅色极简美学**:暖白底色 `#FAF9F6`、1px hairline 边框、`Instrument Serif` 衬线字体、呼吸感负空间
322
+ - **实时事件流 (SSE)**:基于 `engine/events.py` 的 EventBus 内存总线实时推送 9 阶段执行日志
323
+ - **Token 消耗与多模型成本矩阵**:实时比对 DeepSeek-V3/R1、Minimax 2.7、Claude 3.5 Sonnet 与 GPT-4o 成本
324
+ - **混合多渠道检索诊断台**:支持 4 大免配置学术渠道与 AIHot 动态趋势渠道实时测试
325
+
326
+ ### 出版级 HTML 报告与图表重构
327
+ - **出版级效应量森林图 (Forest Plot SVG)**:直观展现速度提升 vs 迁移赤字的尖锐分歧
328
+ - **100% 数据一致性门控**:所有动态聚合与图表通过严苛校验,生成单文件 618KB 离线可用报告
329
+ - **三层白话信息架构**:一页白话结论 (30秒) ➔ 可视化概要 (3分钟) ➔ 完整证据档案 (可折叠)
330
+
331
+ ### 外部 Skill 深度集成与安全补强
332
+ - **`skills/aihot-trend-analysis/`**:集成 AIHot 实时 AI/EdTech 技术趋势雷达
333
+ - **`skills/gap-analysis/`**:集成 BioGapLens 空白与矛盾透镜
334
+ - **`skills/ethics-review/`**:引入人类受试者与 IRB 研究伦理审查
335
+ - **`retrieval/corpus_store.py`**:内置 5 领域离线语料库,保障 100% 离线演示容灾
336
+ - **测试矩阵**:补齐 `tests/test_search.py`,全量 703 个 pytest 用例 100% 零错误通过
337
+
338
+ ## [4.0.0] — 2026-08-14
339
+
340
+ > v4 正式版(`ENGINE_VERSION = "4.0.0"`)。主题:**「科艺融合 × 通用智能」——从教育证据引擎到社科证据决策平台**。
341
+
342
+ ### EvidenceCore 领域泛化
343
+ - `domains/` 领域注册表:`eduevidence domain list|select|check`;education 为注册域(零新增逻辑路径,v3 行为不变);policy 为 v4 首个自带契约域(decision_object/intervention/population/stakeholders 政策框架、5 类政策结果、12 项政策证据质量清单、5 篇政策方法学参考)
344
+ - `engine/evidencecore.py`:`DECISION_STATES`/`PROTOCOL_STEPS` 领域无关常量、`load_domain`(JSON Pointer 契约引用校验)、`validate_frame`(按域校验)
345
+
346
+ ### 证据综合层(Evidence Synthesis)
347
+ - `engine/meta_analysis.py`:效应量提取、固定效应(inverse-variance)与随机效应(DerSimonian-Laird,Q/τ²/I²)双口径合并、森林图数据
348
+ - `engine/bias.py`:Egger 回归(stdlib 手写 t 分布 p 值)、Rosenthal fail-safe N
349
+ - `engine/robustness.py`:leave-one-study-out → robust/fragile 标签
350
+
351
+ ### Living Evidence(活证据)
352
+ - `eduevidence living subscribe|refresh|status`:决策快照订阅 → 新证据(人工注入或 retriever 适配器)内容 hash 幂等入图 → 漂移报告(confirmed/changed/needs_review),绝不自动改判
353
+
354
+ ### 内置证据库 + 离线初步裁决
355
+ - `benchmarks/evidence-library.json`(230 条:金标 209 + 示例 21):`preliminary_verdict` 保守裁决(从不 adopt),无检索能力时可用
356
+
357
+ ### LLM Judge 评估
358
+ - `eduevidence benchmark-judge run|report`:5 维 rubric(citation/outcome/scope/contradiction/decision)评审实证响应,与 heuristic 并列对照,披露同族模型独立性限制
359
+
360
+ ### CI 与工程
361
+ - GitHub Actions 三作业并行:测试矩阵(3.10/3.12 + 重试)、schema-smoke(示例全量过契约)、upload 构建(SKILL.md 一致性 + 零泄漏 canary)
362
+ - 测试 612 → 681(69 个 v4 用例),education 域零回归
363
+ ## [3.0.0] — 2026-08-14
364
+
365
+ > v3 正式版(`engine/versions.py` → `ENGINE_VERSION = "3.0.0"`)。v3 三大方向:**可信度收口**(每个结论收紧到可追溯、可反驳、可审计)、**实证 Benchmark**(Layer B 真实模型运行)、**闭环能力**(PILOT → 真实数据 → 再裁决)。
366
+
367
+ ### 可信度收口(Credibility Tightening)
368
+
369
+ - **Evidence Matrix 三列化(P1-1)**:支持 / 反驳 / 中性证据分列呈现,混合列不再被静默统计(`77db8b5`、`303054e`)。
370
+ - **Pre-Verdict Gate 严格化(OPEN-1)**:`uncertain_claims` 必须绑定证据(`evidence_ids`),无绑定即校验失败(`77db8b5`)。
371
+ - **outcome_mapping 报告(OPEN-2)**:决策输出新增 outcome 映射报告,学习效果与任务表现显式区分(`77db8b5`)。
372
+ - **角色 prompt schema 契约(OPEN-4)**:八角色 prompt 内嵌 schema 契约与枚举值表,输出即契约(`352015a`)。
373
+ - **agent MCP 三态(OPEN-5)**:MCP 可用性三态检测,`install` 自动写入 env(`a735cf6`)。
374
+ - **检索层回归测试补强(P0-1)**:检索 / 校验层回归测试补强,并修复 dedupe stale-index 合并(`de741d0`)。
375
+ - **Source fallback 诚实化(P1-2)**:真实 DOI 解析判 tier1,否则 tier5 + incomplete,杜绝编造 canonical URL(`985239e`)。
376
+ - **fetch 基准站点标题无害化**:中文站点标题英文化 + 分类泛化(`d7f068b`)。
377
+
378
+ ### 实证 Benchmark(Empirical Benchmark,Layer B)
379
+
380
+ - **benchmark_v3 harness**:真实模型调用 + run manifest 契约(`schemas/v3/run-manifest.schema.json`),SIMULATED 与 EMPIRICAL 严格区分(`26a1d1b`)。
381
+ - **gold-based evaluator**:对照 `benchmarks/annotations/gold-*.json` 计算六项指标,每条标注 `method: heuristic`,均值按 95% CI 报告(`26a1d1b`)。
382
+ - **30 金标补齐**:Q11–Q30 金标标注 + 一致性校验测试(`fced61c`)。
383
+ - **CliDriver(omp)**:`omp` CLI 执行后端(`--driver cli`),实证运行落地 `benchmarks/empirical/run-empirical-01`(`503abca`)。
384
+
385
+ ### 闭环能力(Phase 2)
386
+
387
+ - **Decision-to-Outcome Loop(pilot)**:PILOT 试点结果回填 → redecide → 再裁决(`715fae7`)。
388
+ - **Multi-project library synthesis**:跨项目证据库综合 + `REPORT_CONTRACT_VERSION 3.0` + pilots 目录(`eb059fb`)。
389
+ - **CLI 新子命令**:`pilot` / `synthesize` / `benchmark`(`400b346`)。
390
+
391
+ ### 交付与文档
392
+
393
+ - **HTML 报告可访问性与诚实性小修**(6.1–6.5,保持视觉风格)(`d54d23a`)。
394
+ - **文档统一**:Canonical Protocol 9 步(6+3)、Schema 13 个、方法论文档 11 篇、V2 contracts 17 个等数字口径修正(`5b0e886`、`907cd11`)。
395
+ - **版本基线**:`3.0.0`(`715fae7` 起 3.0.0-dev,`ee7ae79` 正式化)。
package/README.md CHANGED
@@ -8,7 +8,7 @@
8
8
 
9
9
  ## EduEvidence Research Engine — Evidence Research & Decision Skill
10
10
 
11
- > **From Research Questions to Evidence-Based Decisions.** · Current release **6.2.0**
11
+ > **From Research Questions to Evidence-Based Decisions.** · Current release **6.3.0**
12
12
 
13
13
  > **▶ Live demo:** [Landing](https://37chengshan.github.io/eduevidence/) · [Research Studio](https://37chengshan.github.io/eduevidence/studio/) · [Deep Research comparison](https://37chengshan.github.io/eduevidence/comparison.html)
14
14
 
@@ -296,7 +296,7 @@ Evidence must connect to the real setting — a classroom, a support team, a pol
296
296
 
297
297
  ## Benchmark
298
298
 
299
- 30 education research questions in v1 (`benchmarks/questions.jsonl`), S×10 / M×10 / L×10; 15 in the core domain "AI-assisted university teaching", 10 with human gold annotations (`benchmarks/annotations/`).
299
+ 30 education research questions in v1 (`benchmarks/questions.jsonl`), S×10 / M×10 / L×10; 15 in the core domain "AI-assisted university teaching", all 30 with human gold annotations (`benchmarks/annotations/gold-Q01..Q30`).
300
300
 
301
301
  Baseline design:
302
302
 
@@ -319,13 +319,13 @@ Key metrics: Citation Support Precision / Unsupported Claim Rate / Contradiction
319
319
  `examples/ai-coding-assistant-evidence/` shows the full path from question to decision:
320
320
 
321
321
  - **Evidence** (12 findings from 8 sources): task-performance gains (Kazemitabaar 2023), unguarded access harming independent exam performance by −17% (Bastani 2025, PNAS), guardrails eliminating the negative effect (Bastani 2025), formative-feedback writing evidence (Marzuki 2024).
322
- - **Decision**: **PILOT** — task-performance evidence is strong, but direct learning-effect evidence for university programming courses is missing, and the unguarded-access risk is documented.
322
+ - **Decision**: recorded `PILOT` (Moderate, 0.586); matrix landing (`engine.decision_policy.decision_outcome`) is `INSUFFICIENT_EVIDENCE` (`unresolved_conflict`) — the decisive relations contain both `support_adoption` and `oppose_adoption`, which fails the conflict check before any `ADOPT`/`PILOT` claim; task-performance evidence is strong, but direct learning-effect evidence for university programming courses is missing and the unguarded-access risk is documented.
323
323
  - **Intervention**: 4-phase pilot (Independent Foundation → Explain Don't Solve → Structured Collaboration → Transfer Check).
324
324
  - **Evaluation**: no-AI baseline / post-test / final-exam retention / no-AI transfer task + AI-dependency risk metrics.
325
325
 
326
- A second public example, `examples/workplace-ai-assistant/`, evaluates AI assistance in enterprise customer support using the policy domain: 4 findings from 3 studies, with direct and indirect evidence distinguished. Its proposed supervised pilot has not been executed.
326
+ A second public example, `examples/workplace-ai-assistant/`, evaluates AI assistance in enterprise customer support using the policy domain: 4 findings from 3 studies, with direct and indirect evidence distinguished. Its recorded action is `PILOT` (Moderate, 0.578); the matrix landing is `INSUFFICIENT_EVIDENCE` with `downgrade_reason=None`: all four decisive relations are `conditional`, and `conditional` is not a conflict relation (`engine.decision_policy.CONFLICT_RELATIONS` = `{conflict, mixed}`), so no downgrade reason is recorded and — with no decisive `support_adoption` relation — Moderate confidence cannot reach `PILOT`. Its proposed supervised pilot has not been executed.
327
327
 
328
- The third public example, `examples/spaced-retrieval-practice/`, asks whether spaced repetition and retrieval practice should replace massed review in an introductory programming course. It is the first pack whose sources were located through the **Sciverse** channel (`discovery_provider=sciverse`, `fetch_provider=sciverse_content`) and whose `meta.data_origin` is `real_run_sciverse`: 6 findings from 7 tier-1 DOI sources, decision **ADOPT** (High confidence). It is the worked example that the ADOPT path is reachable: retention and transfer - the two primary outcomes - carry direct, consistent evidence at directness 2, while the coding and workplace cases stay bounded at PILOT because their primary learning evidence is missing.
328
+ The third public example, `examples/spaced-retrieval-practice/`, asks whether spaced repetition and retrieval practice should replace massed review in an introductory programming course. It is the first pack whose sources were located through the **Sciverse** channel (`discovery_provider=sciverse`, `fetch_provider=sciverse_content`) and whose `meta.data_origin` is `real_run_sciverse`: 6 findings from 7 tier-1 DOI sources, decision **ADOPT** (High confidence, 0.893) — the only public case where the recorded action and the matrix landing agree. It is the worked example that the ADOPT path is reachable: retention and transfer - the two primary outcomes - carry direct, consistent evidence at directness 2. The coding and workplace cases are bounded below ADOPT: the coding case's decisive relations mix `support_adoption` with `oppose_adoption` (`unresolved_conflict`), and the workplace case carries no decisive `support_adoption` relation at all (all four are `conditional`), so both land at `INSUFFICIENT_EVIDENCE`.
329
329
 
330
330
  Each pack ships `result.json` + `result.zh.json` (bilingual parallel data), a packaged-`EduEvidence_Report.html` root report, and `reports-5themes/` with the five standalone theme HTML files.
331
331
 
@@ -402,16 +402,17 @@ Read the illustrated single-file walkthrough of the same architecture (nine-step
402
402
  ```text
403
403
  EduEvidence/ (= one Skill package)
404
404
  │
405
- ├─ SKILL.md ← Skill entry: When to Use / Inputs / Workflow / Output Contract
405
+ ├─ SKILL.md ← Skill entry: mission → workflow routing → scientific gates → 9-step protocol → output requirements
406
406
  │
407
407
  ├─ Skill core (required to run)
408
408
  │ ├─ engine/ V2 Research Engine (Project Workspace / immutable Evidence
409
409
  │ │ Graph / Library / synthesis / tribunal / study design /
410
410
  │ │ datasets / analysis / projections / migration)
411
411
  │ ├─ skill/agents/ role protocols (capability execution profiles)
412
- │ ├─ references/ 11 education methodology documents (evidence quality / skeptic /
413
- │ │ tribunal policy / intervention design…)
414
- │ ├─ schemas/ V1 + v2/v3 contracts (13 top-level + 17 v2 + v3 pilot/synthesis/run-manifest)
412
+ │ ├─ references/ 21 methodology documents (evidence quality / skeptic /
413
+ │ │ tribunal policy / intervention design…; counted in docs/metrics.json)
414
+ │ ├─ schemas/ 50 versioned JSON Schema contracts (15 top-level + 18 v2 +
415
+ │ │ 3 v3 + 4 v4 + 10 vNext; counted by `scripts/generate_metrics.py`)
415
416
  │ ├─ domains/ v4 domain registry (manifest.json) + per-domain packages
416
417
  │ │ (education: registration-only; policy: frame schema /
417
418
  │ │ outcome taxonomy / methodology checklist / references)
@@ -431,11 +432,18 @@ EduEvidence/ (= one Skill package)
431
432
  ├─ docs/ architecture / methodology / benchmark / demo / reproducibility
432
433
  ├─ install.sh one-click install (local / multi-agent Skill) + self-check
433
434
  ├─ pyproject.toml packaging metadata (wheel ships CLI, engine and installed runtime resources; stdlib-only core)
434
- └─ README(.en).md bilingual docs
435
+ └─ README.md / README.zh-CN.md bilingual docs
435
436
  ```
436
437
 
437
438
  > Skill-package principle: the **minimal runtime set is `SKILL.md + engine/ + skill/ + references/ + schemas/ + scripts/`**; `retrieval/`, `integrations/`, `visualization/` are the execution/presentation layers that make the Skill actually runnable; `tests/`, `benchmarks/`, `examples/`, `docs/` provide credibility and onboarding — none of them affect the Skill body itself.
438
439
 
440
+ ### Distribution
441
+
442
+ Two release channels, built from one runtime allowlist (`scripts/skill_payload.py`, shared with the host installer):
443
+
444
+ - **Submission folder** — `bash packaging/make_upload.sh` rebuilds `dist/eduevidence-submission/` (flat Skill package + `submission-manifest.json` + `UPLOAD-README.md`); it regenerates the report variants first and aborts without a prebuilt `web/studio/index.html`. The older `upload/` tree is a **frozen deprecated snapshot** that must not be shipped — see [`upload/DEPRECATED.md`](upload/DEPRECATED.md).
445
+ - **npm package** — `npm install -g eduevidence`, then `eduevidence skill --list-hosts` (every supported agent and its install path) or `eduevidence skill --dry-run` (preview; writes nothing). From a checkout the same flags run as `node bin/eduevidence.js skill --list-hosts` / `--dry-run`.
446
+
439
447
  ### v4 Domain Registry (EvidenceCore)
440
448
 
441
449
  The first step of the **EvidenceCore** abstraction: a `domains/` registry
@@ -445,8 +453,9 @@ The first step of the **EvidenceCore** abstraction: a `domains/` registry
445
453
 
446
454
  - **education** is a **registration-only domain**: it points at the existing
447
455
  contracts — `schemas/education-frame.schema.json` (frame),
448
- the 20-token outcome taxonomy from `schemas/evidence.schema.json` (four
449
- categories per `engine/pilot.py` `_OUTCOME_CATEGORY`), the 15-item
456
+ the 20-token education outcome taxonomy from
457
+ `domains/education/outcome_taxonomy.json` (`engine/taxonomy.py` is its sole
458
+ reader; 25 tokens across education + policy), the 15-item
450
459
  methodology checklist from `skill/agents/method-reviewer.md`,
451
460
  `benchmarks/annotations` (golds) and `references/`. It **adds no new logic
452
461
  path** — no new schema, no new validator, no new methodology.
@@ -550,7 +559,7 @@ pytest
550
559
  simulation that proves the evaluation framework runs; it is **not** real model
551
560
  performance (see the ⚠️ note in [Benchmark](#benchmark)).
552
561
  - [x] Skill core & pipeline: 9-step protocol (Research Core 6 + Decision Extension 3),
553
- 13 top-level JSON Schemas, deterministic scripts, 8-role protocols (original Phases 0–6).
562
+ 50 versioned JSON Schemas (`schemas/**`, counted by `scripts/generate_metrics.py`), deterministic scripts, 8-role protocols (original Phases 0–6).
554
563
  - [x] Evidence-to-action: applicability / four-state decision / intervention / evaluation.
555
564
  - [x] Product UI: single-file bilingual HTML report + infographics + academic figures
556
565
  (original Phase 8).
package/README.zh-CN.md CHANGED
@@ -284,13 +284,13 @@ B4 EduEvidence + Agent MCP ← 证明多 Agent 增强价值(B3 vs B4)
284
284
  `examples/ai-coding-assistant-evidence/` 完整展示了从问题到决策的全过程:
285
285
 
286
286
  - **证据**(12 条发现、8 个来源):任务表现提升(Kazemitabaar 2023)、无护栏访问损害独立考试表现 -17%(Bastani 2025, PNAS)、护栏设计消除负效应(Bastani 2025)、形成性反馈写作证据(Marzuki 2024)。
287
- - **决策**:**PILOT** —— 任务表现证据强,但大学编程课程的直接学习效应证据缺失,无护栏风险已被证实。
287
+ - **决策**:落档动作 `PILOT`(Moderate,0.586);矩阵落点(`engine.decision_policy.decision_outcome`)为 `INSUFFICIENT_EVIDENCE`(`unresolved_conflict`)—— decisive relations 中同时存在 `support_adoption` 与 `oppose_adoption`,冲突检查在 `ADOPT`/`PILOT` 判定前即短路;任务表现证据强,但大学编程课程的直接学习效应证据缺失,且无护栏风险已被证实。
288
288
  - **干预**:4 阶段试点(Independent Foundation → Explain Don't Solve → Structured Collaboration → Transfer Check)。
289
289
  - **评价**:无 AI 基线/后测/期末考试保持/无 AI 迁移任务 + AI 依赖风险指标。
290
290
 
291
- 另一公开案例 `examples/workplace-ai-assistant/` 使用组织政策领域,讨论企业客服是否引入 AI 助手:3 项研究、4 条发现,区分直接客服证据与间接写作/咨询证据,判定 **PILOT**(Moderate)。详见 [来源核验与边界](docs/demo-workplace-ai.md)。
291
+ 另一公开案例 `examples/workplace-ai-assistant/` 使用组织政策领域,讨论企业客服是否引入 AI 助手:3 项研究、4 条发现,区分直接客服证据与间接写作/咨询证据,落档动作 `PILOT`(Moderate,0.578),矩阵落点为 `INSUFFICIENT_EVIDENCE` 且 `downgrade_reason=None`:4 条 decisive relation 全为 `conditional`,而 `conditional` 不属于冲突关系(`engine.decision_policy.CONFLICT_RELATIONS` = `{conflict, mixed}`),因此引擎不记录降级理由;又因没有任何 decisive 的 `support_adoption`,Moderate 置信度无法升到 `PILOT`。详见 [来源核验与边界](docs/demo-workplace-ai.md)。
292
292
 
293
- 第三个公开案例 `examples/spaced-retrieval-practice/`(来源经 Sciverse 通道逐条读回原文)讨论间隔重复与检索练习能否替代集中式复习:6 条证据、7 篇 tier-1 DOI 来源,判定 **ADOPT**(High,引擎复算 0.893)。它是“ADOPT 出口真实可达”的实证:延迟保持与迁移两个主要结果上都有 directness=2 的直接且一致证据;另两例因为主要学习结果上缺直接证据而停在 PILOT。
293
+ 第三个公开案例 `examples/spaced-retrieval-practice/`(来源经 Sciverse 通道逐条读回原文)讨论间隔重复与检索练习能否替代集中式复习:6 条证据、7 篇 tier-1 DOI 来源,判定 **ADOPT**(High,引擎复算 0.893)—— 唯一一个落档动作与矩阵落点一致的公开案例。它是“ADOPT 出口真实可达”的实证:延迟保持与迁移两个主要结果上都有 directness=2 的直接且一致证据;另两例低于 ADOPT:编程案中 decisive relations 同时含 `support_adoption` 与 `oppose_adoption`(`unresolved_conflict`),企业客服案中没有任何 decisive 的 `support_adoption`(4 条全为 `conditional`),因此矩阵落点均为 `INSUFFICIENT_EVIDENCE`,不得表述为已落在 PILOT。
294
294
 
295
295
  三个公开案例的数据来源各自如实标注:编程与企业客服两例为人工整理文献(`manual_curated`),间隔重复一例为真实 Sciverse 检索运行记录(`real_run_sciverse`),报告生成不等于九阶段模型研究已运行,也不代表试点已经执行。四个旧教学示例迁入 `tests/fixtures/legacy-examples/`,仅供软件兼容测试,排除于公共目录和分发包;未核验或合成数据不能引用为研究证据。旧 `ai-coding-assistant` 路径保留兼容别名。
296
296
 
@@ -354,14 +354,14 @@ Sciverse 通道把 `/agentic-search` 的 chunk 当作**定位子**:必须经 `
354
354
  ```text
355
355
  EduEvidence/ (= 一个 Skill 包)
356
356
  │
357
- ├─ SKILL.md ← Skill 入口:When to Use / Inputs / 9 步 Workflow / 输出契约
357
+ ├─ SKILL.md ← Skill 入口:使命 → 工作流路由 → 科学门禁 → 9 步协议 → 输出要求
358
358
  │
359
359
  ├─ Skill 本体(运行必需)
360
360
  │ ├─ skill/agents/ 8 个角色协议(Planner / Retriever / Analyst / Skeptic /
361
361
  │ │ Method Reviewer / Judge / Intervention Designer / Evaluation Designer)
362
362
  │ ├─ references/ 方法论文档(证据质量 / 反证协议 / 裁决规则 / 干预设计 / 检索合规 / 文案规范…;数量见 docs/metrics.json)
363
- │ ├─ schemas/ 33 个 JSON Schema 数据契约(13 顶层 + 17 v2 + 3 v3,每步输出的校验门)
364
- │ ├─ scripts/ 17 个确定性逻辑脚本(评分 / 矩阵 / 审计 / 置信度 / Orchestrator / 启动探测)
363
+ │ ├─ schemas/ 50 个 JSON Schema 数据契约(15 顶层 + 18 v2 + 3 v3 + 4 v4 + 10 vNext,每步输出的校验门;递归计数见 `scripts/generate_metrics.py`)
364
+ │ ├─ scripts/ 确定性逻辑脚本(评分 / 矩阵 / 审计 / 置信度 / Orchestrator / 启动探测)
365
365
  │ ├─ retrieval/ 检索与抓取层(fetch / validate / dedupe / failures)
366
366
  │ ├─ integrations/ Agent MCP 增强层 + Smart Web Fetch 集成
367
367
  │ └─ visualization/ 结果呈现层(ECharts / 信息图 / 学术图 / 双语 HTML Composer)
@@ -375,11 +375,18 @@ EduEvidence/ (= 一个 Skill 包)
375
375
  ├─ docs/ 架构 / 方法论 / Benchmark / Demo / 复现指南
376
376
  ├─ install.sh 一键安装(本地 / 多 Agent Skill)+ 自检
377
377
  ├─ pyproject.toml 打包元数据(核心零第三方依赖)
378
- └─ README(.en).md 双语说明
378
+ └─ README.md / README.zh-CN.md 双语说明
379
379
  ```
380
380
 
381
381
  > Skill 包设计原则:**运行所需的最小集是 `SKILL.md + skill/ + references/ + schemas/ + scripts/`**;`retrieval/`、`integrations/`、`visualization/` 是让 Skill 真正"可运行、可呈现"的执行层;`tests/`、`benchmarks/`、`examples/`、`docs/` 是可信度与上手保障,不影响 Skill 本体。
382
382
 
383
+ ### 分发
384
+
385
+ 两条发布通道共用同一份运行时 allowlist(`scripts/skill_payload.py`,宿主安装脚本也用这一份):
386
+
387
+ - **提交包** —— `bash packaging/make_upload.sh` 重建 `dist/eduevidence-submission/`(扁平 Skill 包 + `submission-manifest.json` + `UPLOAD-README.md`);它会先重建五种风格的报告变体,缺少已构建的 `web/studio/index.html` 时直接中止。旧 `upload/` 目录是**冻结的废弃快照**,不得随包分发 —— 见 [`upload/DEPRECATED.md`](upload/DEPRECATED.md)。
388
+ - **npm 包** —— `npm install -g eduevidence`,然后用 `eduevidence skill --list-hosts`(列出全部宿主与落点)或 `eduevidence skill --dry-run`(只预览,不写入)。在仓库目录中等价写法是 `node bin/eduevidence.js skill --list-hosts` / `--dry-run`。
389
+
383
390
  ### SCP / Platform Native Mode
384
391
 
385
392
  EduEvidence 可完全脱离 Agent MCP 独立运行(无需任何外部服务):
@@ -462,7 +469,7 @@ pytest
462
469
  **已完成(仅 harness / 仿真,标注 SIMULATED,非实证):**
463
470
 
464
471
  - [x] Benchmark v2 harness / simulation —— `benchmarks/results/` 是确定性仿真,证明评测框架可运行,**不是**真实模型性能(见上方 [Benchmark](#benchmark) 的 ⚠️ 说明)。
465
- - [x] Skill 核心与管线:9 步协议(Research Core 6 + Decision Extension 3)、13 个顶层 JSON Schema、确定性脚本、8 角色协议(原计划 Phase 0–6)。
472
+ - [x] Skill 核心与管线:9 步协议(Research Core 6 + Decision Extension 3)、50 个版本化 JSON Schema(`schemas/**`,计数见 `scripts/generate_metrics.py`)、确定性脚本、8 角色协议(原计划 Phase 0–6)。
466
473
  - [x] Evidence-to-Action:适用性 / 四态决策 / 干预 / 评价设计。
467
474
  - [x] 产品 UI:单文件双语 HTML 报告 + 信息图 + 学术图(原计划 Phase 8)。
468
475