eduevidence 6.0.0 → 6.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +395 -0
- package/CONTRIBUTING.md +105 -0
- package/README.md +113 -49
- package/README.zh-CN.md +39 -12
- package/SKILL.md +15 -5
- package/assets/readme/landing-tour.gif +0 -0
- package/assets/readme/studio-tour.gif +0 -0
- package/benchmarks/evidence-library.json +277 -1
- package/bin/eduevidence.js +2 -1
- package/docs/architecture.md +325 -46
- package/docs/demo-workplace-ai.md +1 -1
- package/docs/install-guide.md +1 -1
- package/docs/j-ev-experimental.md +250 -0
- package/docs/orchestration-role-model.md +1 -1
- package/docs/release-closeout/README.md +1 -1
- package/docs/reproducibility.md +138 -0
- package/docs/sciverse-api.md +125 -0
- package/domains/_neutral/copy/few_shots.json +21 -0
- package/domains/_neutral/copy/framing_lexicon.json +19 -0
- package/domains/_neutral/copy/module_labels.json +5 -0
- package/domains/_neutral/copy/module_labels_footer.json +102 -0
- package/domains/_neutral/copy/module_labels_modules.json +204 -0
- package/domains/_neutral/copy/module_labels_nav.json +126 -0
- package/domains/_neutral/copy/module_labels_summary.json +98 -0
- package/domains/_neutral/copy/module_labels_tables.json +164 -0
- package/domains/_neutral/copy/module_labels_v2.json +90 -0
- package/domains/_neutral/copy/risk_constructs.json +20 -0
- package/domains/_neutral/copy/section_titles.json +66 -0
- package/domains/_neutral/copy/terminology.json +11 -0
- package/domains/check_copy_packs.py +103 -0
- package/domains/education/copy/few_shots.json +22 -0
- package/domains/education/copy/framing_enums.json +167 -0
- package/domains/education/copy/framing_lexicon.json +166 -0
- package/domains/education/copy/module_labels.json +169 -0
- package/domains/education/copy/risk_constructs.json +48 -0
- package/domains/education/copy/section_titles.json +186 -0
- package/domains/education/copy/terminology.json +70 -0
- package/domains/education/manifest.json +1 -1
- package/domains/education/outcome_taxonomy.json +2 -2
- package/domains/manifest.json +1 -1
- package/domains/policy/copy/few_shots.json +22 -0
- package/domains/policy/copy/framing_enums.json +94 -0
- package/domains/policy/copy/framing_lexicon.json +174 -0
- package/domains/policy/copy/module_labels.json +168 -0
- package/domains/policy/copy/risk_constructs.json +33 -0
- package/domains/policy/copy/section_titles.json +186 -0
- package/domains/policy/copy/terminology.json +64 -0
- package/eduevidence_cli.py +10 -0
- package/engine/capabilities.py +57 -5
- package/engine/decision_policy.py +167 -0
- package/engine/evidence_graph.py +14 -10
- package/engine/gaps.py +42 -22
- package/engine/ids.py +2 -0
- package/engine/library.py +6 -2
- package/engine/library_builtin.py +7 -4
- package/engine/living.py +34 -4
- package/engine/migration.py +88 -3
- package/engine/orchestration.py +5 -5
- package/engine/paths.py +2 -0
- package/engine/pilot.py +34 -32
- package/engine/taxonomy.py +211 -0
- package/engine/tribunal.py +49 -43
- package/engine/versions.py +1 -1
- package/examples/ai-coding-assistant-evidence/EduEvidence_Report.html +1361 -147
- package/examples/ai-coding-assistant-evidence/artifact_manifest.json +3 -3
- package/examples/ai-coding-assistant-evidence/citation_check.json +1 -1
- package/examples/ai-coding-assistant-evidence/final_verdict.json +107 -0
- package/examples/ai-coding-assistant-evidence/gate_report.json +101 -0
- package/examples/ai-coding-assistant-evidence/report_spec.json +23 -12
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_academic.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_claude.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_datalab-dark.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_datalab.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_presentation.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_academic.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_claude.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_datalab-dark.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_datalab.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_presentation.html +1360 -146
- package/examples/ai-coding-assistant-evidence/result.json +13 -9
- package/examples/ai-coding-assistant-evidence/result.zh.json +45 -41
- package/examples/ai-coding-assistant-evidence/skeptic.json +72 -0
- package/examples/ai-coding-assistant-evidence/verdict.json +6 -2
- package/examples/spaced-retrieval-practice/EduEvidence_Report.html +2728 -0
- package/examples/spaced-retrieval-practice/applicability.json +14 -0
- package/examples/spaced-retrieval-practice/artifact_manifest.json +15 -0
- package/examples/spaced-retrieval-practice/claims.jsonl +3 -0
- package/examples/spaced-retrieval-practice/evidence.jsonl +6 -0
- package/examples/spaced-retrieval-practice/final_verdict.json +93 -0
- package/examples/spaced-retrieval-practice/frame.json +58 -0
- package/examples/spaced-retrieval-practice/gate_report.json +101 -0
- package/examples/spaced-retrieval-practice/methodology.json +78 -0
- package/examples/spaced-retrieval-practice/report.html +2522 -0
- package/examples/spaced-retrieval-practice/report_spec.json +212 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_academic.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_claude.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_datalab-dark.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_datalab.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_presentation.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_academic.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_claude.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_datalab-dark.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_datalab.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_presentation.html +2728 -0
- package/examples/spaced-retrieval-practice/result.json +942 -0
- package/examples/spaced-retrieval-practice/result.zh.json +942 -0
- package/examples/spaced-retrieval-practice/skeptic.json +70 -0
- package/examples/spaced-retrieval-practice/sources.jsonl +7 -0
- package/examples/spaced-retrieval-practice/verdict.json +93 -0
- package/examples/workplace-ai-assistant/EduEvidence_Report.html +2814 -0
- package/examples/workplace-ai-assistant/artifact_manifest.json +15 -0
- package/examples/workplace-ai-assistant/claims.jsonl +4 -4
- package/examples/workplace-ai-assistant/evidence.jsonl +4 -4
- package/examples/workplace-ai-assistant/evidence_graph.json +15 -15
- package/examples/workplace-ai-assistant/final_verdict.json +78 -0
- package/examples/workplace-ai-assistant/gate_report.json +101 -0
- package/examples/workplace-ai-assistant/report_spec.json +209 -40
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_academic.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_claude.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_datalab-dark.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_datalab.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_presentation.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/report_academic.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_claude.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_datalab-dark.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_datalab.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_presentation.html +2814 -0
- package/examples/workplace-ai-assistant/result.json +82 -20
- package/examples/workplace-ai-assistant/result.zh.json +82 -20
- package/examples/workplace-ai-assistant/skeptic.json +72 -0
- package/examples/workplace-ai-assistant/verdict.json +36 -10
- package/integrations/agent_mcp.py +2 -2
- package/integrations/jev/__init__.py +115 -0
- package/integrations/jev/approval.py +212 -0
- package/integrations/jev/cli.py +84 -0
- package/integrations/jev/config.py +112 -0
- package/integrations/jev/gateway.py +128 -0
- package/integrations/jev/modes.py +38 -0
- package/integrations/jev/tools_classify.py +88 -0
- package/integrations/jev/tools_extract.py +111 -0
- package/integrations/jev/tools_rerank.py +71 -0
- package/integrations/jev/tools_screen.py +87 -0
- package/integrations/jev/tools_verify.py +95 -0
- package/integrations/jev_mcp.py +22 -0
- package/integrations/semantic_decide.py +286 -0
- package/integrations/semdecide_cli.py +55 -0
- package/package.json +19 -2
- package/pyproject.toml +4 -3
- package/references/report-copy-style.md +107 -0
- package/references/retrieval-compliance.md +75 -0
- package/references/retrieval-protocol.md +20 -0
- package/retrieval/audit.py +27 -3
- package/retrieval/fetch.py +96 -0
- package/retrieval/sciverse.py +398 -0
- package/retrieval/search.py +47 -7
- package/schemas/applicability.schema.json +94 -0
- package/schemas/chart-spec.schema.json +10 -3
- package/schemas/evidence.schema.json +316 -43
- package/schemas/fetch-result.schema.json +2 -1
- package/schemas/report-result.schema.json +3 -3
- package/schemas/report-spec.schema.json +98 -100
- package/schemas/skeptic.schema.json +86 -0
- package/schemas/source.schema.json +21 -2
- package/schemas/v2/decision-snapshot.schema.json +20 -9
- package/schemas/v2/finding.schema.json +5 -1
- package/schemas/v2/intake.schema.json +191 -0
- package/schemas/v2/methodology-audit.schema.json +5 -1
- package/schemas/v2/outcome.schema.json +28 -5
- package/schemas/v2/study.schema.json +5 -1
- package/schemas/vNext/autoevolve-session.schema.json +34 -1
- package/schemas/vNext/eval-snapshot.schema.json +77 -1
- package/schemas/vNext/execution-plan.schema.json +50 -1
- package/schemas/vNext/gap-priority.schema.json +54 -1
- package/schemas/vNext/negative-search-record.schema.json +68 -1
- package/schemas/vNext/research-iteration.schema.json +87 -1
- package/schemas/vNext/research-strategy.schema.json +62 -1
- package/schemas/vNext/skill-experiment.schema.json +90 -1
- package/schemas/vNext/task-spec.schema.json +156 -1
- package/schemas/vNext/worker-result.schema.json +60 -1
- package/schemas/verdict.schema.json +164 -28
- package/scripts/build_evidence_library.py +15 -5
- package/scripts/build_report_variants.py +18 -2
- package/scripts/build_result.py +74 -9
- package/scripts/check_package_parity.py +85 -0
- package/scripts/check_protocol_alignment.py +375 -0
- package/scripts/check_versioned_schemas.py +254 -0
- package/scripts/claim_audit.py +13 -8
- package/scripts/compute_confidence.py +10 -0
- package/scripts/dashboard_server.py +13 -2
- package/scripts/did_regression.py +12 -2
- package/scripts/evidence_score.py +5 -2
- package/scripts/intake/__init__.py +31 -0
- package/scripts/intake/__main__.py +18 -0
- package/scripts/intake/background.py +78 -0
- package/scripts/intake/browser.py +79 -0
- package/scripts/intake/cli.py +57 -0
- package/scripts/intake/constants.py +57 -0
- package/scripts/intake/depth.py +53 -0
- package/scripts/intake/enhancements.py +106 -0
- package/scripts/intake/hooks.py +90 -0
- package/scripts/intake/prefs.py +76 -0
- package/scripts/intake/prompts.py +85 -0
- package/scripts/intake/session.py +152 -0
- package/scripts/lint_file_layers.py +126 -0
- package/scripts/orchestrator.py +187 -40
- package/scripts/pre_verdict_gate.py +241 -29
- package/scripts/quickstart.py +18 -2
- package/scripts/run_workspace.py +7 -1
- package/scripts/skill_lint.py +11 -1
- package/scripts/skill_payload.py +6 -3
- package/scripts/test_adversarial_empirical.py +96 -25
- package/scripts/validate_schema.py +31 -1
- package/skill/agents/evaluation-designer.md +20 -4
- package/skill/agents/evidence-analyst.md +19 -3
- package/skill/agents/evidence-judge.md +98 -8
- package/skill/agents/evidence-retriever.md +20 -3
- package/skill/agents/intervention-designer.md +20 -4
- package/skill/agents/method-reviewer.md +18 -2
- package/skill/agents/{education-planner.md → research-planner.md} +19 -3
- package/skill/agents/skeptic.md +18 -2
- package/skill/roles/registry.yaml +11 -11
- package/skill/sub-skills/aihot-trend-analysis/SKILL.md +28 -9
- package/skill/sub-skills/contradiction-analysis/SKILL.md +31 -11
- package/skill/sub-skills/data-analysis/SKILL.md +34 -15
- package/skill/sub-skills/ethics-review/SKILL.md +33 -10
- package/skill/sub-skills/evidence-extraction/SKILL.md +29 -11
- package/skill/sub-skills/evidence-review/SKILL.md +31 -12
- package/skill/sub-skills/gap-analysis/SKILL.md +31 -9
- package/skill/sub-skills/literature-review/SKILL.md +35 -14
- package/skill/sub-skills/methodology-audit/SKILL.md +29 -12
- package/skill/sub-skills/report-generation/SKILL.md +28 -0
- package/skill/sub-skills/research-planning/SKILL.md +41 -14
- package/skill/sub-skills/study-design/SKILL.md +30 -9
- package/skill/task-briefs/adjudicate.md +32 -7
- package/skill/task-briefs/applicability.md +37 -2
- package/skill/task-briefs/audit.md +32 -7
- package/skill/task-briefs/challenge.md +34 -5
- package/skill/task-briefs/evaluate.md +30 -5
- package/skill/task-briefs/extract.md +31 -8
- package/skill/task-briefs/frame.md +39 -10
- package/skill/task-briefs/intervene.md +32 -6
- package/skill/task-briefs/present.md +32 -8
- package/skill/task-briefs/projection.md +36 -2
- package/skill/task-briefs/retrieve.md +36 -6
- package/skill/workflows/decision-and-pilot.md +76 -1
- package/skill/workflows/evaluate-and-update.md +83 -0
- package/skill/workflows/evidence-review.md +104 -0
- package/skill/workflows/experimental-jev.md +170 -0
- package/skill/workflows/intake.md +120 -0
- package/visualization/eduevidence-report/scripts/build_figures.py +25 -3
- package/visualization/eduevidence-report/scripts/build_infographics.py +37 -15
- package/visualization/eduevidence-report/scripts/build_report.py +435 -575
- package/visualization/eduevidence-report/scripts/charts_data.py +2 -0
- package/visualization/eduevidence-report/scripts/lieflat_engine.py +349 -38
- package/visualization/eduevidence-report/scripts/report_copy_pack.py +296 -0
- package/visualization/eduevidence-report/scripts/report_copy_policy_guard.py +47 -0
- package/visualization/eduevidence-report/scripts/zh_labels.py +141 -1
- package/web/architecture.html +14885 -0
- package/web/studio/assets/index-B8tkF44Q.css +1 -0
- package/web/studio/index.html +2 -2
- package/scripts/build_esl_artifacts.py +0 -1921
- package/scripts/build_killer_demo.py +0 -295
- package/scripts/enrich_projects_human_and_lieflat.py +0 -315
- package/scripts/generate_new_projects.py +0 -686
- package/scripts/sync_killer_demo_report.py +0 -270
- package/web/studio/assets/index-CzXocaGv.css +0 -1
- /package/web/studio/assets/{index-pa7jD7n4.js → index-CQ6Keoyc.js} +0 -0
|
@@ -0,0 +1,250 @@
|
|
|
1
|
+
# Jev 生态实验模式(Experimental Jev / SemDecide Tier-0)
|
|
2
|
+
|
|
3
|
+
> **实验性加速层,不是科学门。** Tier-0(screen / rerank / extract / classify / verify)
|
|
4
|
+
> 只做窄域类型化判断,**不替代** `scripts/pre_verdict_gate.py`、Skeptic 九项检查、
|
|
5
|
+
> 方法学审计与 Tribunal。不确定 / 非法结果 **fail-closed** 升级大模型。
|
|
6
|
+
>
|
|
7
|
+
> 入口开关:`--experimental <0|1|2|3>` 或 `EDU_EXPERIMENTAL_JEV=<mode>`。
|
|
8
|
+
> 密钥仅存 `~/.eduevidence/env`(`AI_GATEWAY_API_KEY` + `JEV_PROVIDER=vercel`,可选
|
|
9
|
+
> `TYPESAFE_API_KEY`;或 FreeJev:`FREEJEV_API_KEY` + `JEV_PROVIDER=freejev`),
|
|
10
|
+
> **禁止写入仓库**。
|
|
11
|
+
|
|
12
|
+
## 1. 定位
|
|
13
|
+
|
|
14
|
+
| 组件 | 实现 | 上游 |
|
|
15
|
+
|---|---|---|
|
|
16
|
+
| `integrations/jev_mcp.py` | detect → approval → urllib 调 Vercel AI Gateway `typesafe-ai/jev` | TypeSafe System One / `@jkudish/jev-mcp` |
|
|
17
|
+
| `integrations/semantic_decide.py` | subprocess `semdecide is / filter / choose` | [sharziki/semdecide](https://github.com/sharziki/semdecide) |
|
|
18
|
+
| `engine/capabilities.py` | 注册 5 个实验能力(`experimental_capability_registry()`;默认科学注册表不含它们,避免协议角色门误伤) | 本仓库能力注册表 |
|
|
19
|
+
| `skill/workflows/experimental-jev.md` | 模式 / 九阶段映射 / fail-closed | Skill 工作流覆盖层 |
|
|
20
|
+
|
|
21
|
+
原则(与 agent-mcp 一致):**检测 → 推荐 → 用户确认 → 调用 → 降级**。不迁移、不复制
|
|
22
|
+
上游实现;上游不可用时整体退回 Platform Native / mode 0,科学协议不变。
|
|
23
|
+
|
|
24
|
+
## 2. 启用方式
|
|
25
|
+
|
|
26
|
+
```bash
|
|
27
|
+
# 0) 密钥(本地,勿提交)
|
|
28
|
+
# ~/.eduevidence/env
|
|
29
|
+
# # Vercel AI Gateway 路径
|
|
30
|
+
# export AI_GATEWAY_API_KEY=...
|
|
31
|
+
# export JEV_PROVIDER=vercel
|
|
32
|
+
# # 或 FreeJev 路径
|
|
33
|
+
# export FREEJEV_API_KEY=...
|
|
34
|
+
# export JEV_PROVIDER=freejev
|
|
35
|
+
|
|
36
|
+
# 1) 检测
|
|
37
|
+
python3 integrations/jev_mcp.py --experimental 3
|
|
38
|
+
python3 integrations/semantic_decide.py --experimental 3
|
|
39
|
+
|
|
40
|
+
# 2) 用户确认门 → ~/.eduevidence/jev_mcp_approval.json
|
|
41
|
+
python3 integrations/jev_mcp.py --experimental 3 --approve
|
|
42
|
+
|
|
43
|
+
# 3) 运行时打开 overlay(宿主 CLI 用 --enhancement 选择;写入 run 的 intake.json)
|
|
44
|
+
eduevidence run --question "..." --run-id <id> --enhancement jev --enhancement semdecide
|
|
45
|
+
# 检测类 CLI 用模式号:--experimental <0|1|2|3>,或 export EDU_EXPERIMENTAL_JEV=3
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
| Mode | 名称 | 后端 |
|
|
49
|
+
|---:|---|---|
|
|
50
|
+
| 0 | standard | 仅大模型(默认,无 `--experimental`) |
|
|
51
|
+
| 1 | jev-mcp | FreeJev / Vercel Gateway urllib **或** 文档化 MCP |
|
|
52
|
+
| 2 | semdecide | `semdecide is / filter / choose` |
|
|
53
|
+
| 3 | hybrid | 1 + 2 |
|
|
54
|
+
|
|
55
|
+
### Provider(Mode 1 / 3)
|
|
56
|
+
|
|
57
|
+
| `JEV_PROVIDER` | 密钥 | Decide 端点 | 备注 |
|
|
58
|
+
|---|---|---|---|
|
|
59
|
+
| `freejev` | `FREEJEV_API_KEY` | `https://freejev.org/api/v1/decide` | 仅 `FREEJEV_API_KEY` 时可自动选中 |
|
|
60
|
+
| `vercel`(默认) | `AI_GATEWAY_API_KEY`(或 `TYPESAFE_API_KEY`) | `https://ai-gateway.vercel.sh/typesafe/v1/systemone` | `model=typesafe-ai/jev` |
|
|
61
|
+
|
|
62
|
+
MCP 备选(不自动 spawn,仅文档化):
|
|
63
|
+
|
|
64
|
+
```bash
|
|
65
|
+
# FreeJev MCP
|
|
66
|
+
# https://freejev.org/mcp + FREEJEV_API_KEY
|
|
67
|
+
# npx 包路径
|
|
68
|
+
claude mcp add jev -- npx -y @jkudish/jev-mcp
|
|
69
|
+
# 通用 JSON:npx -y @jkudish/jev-mcp + TYPESAFE_API_KEY / AI_GATEWAY_API_KEY
|
|
70
|
+
python3 integrations/jev_mcp.py --mcp-guide
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
**FreeJev `request_id` 不重试。** Decide 调用的 `request_id` 是幂等键:同一
|
|
74
|
+
`request_id` **禁止重试**(超时 / 5xx / 未知结果一律不重发)。本调用
|
|
75
|
+
fail-closed 升级大模型;若业务上必须再问一次,使用**新的** `request_id` 发起
|
|
76
|
+
新调用,并单独记录 provenance。
|
|
77
|
+
|
|
78
|
+
## 3. Tier-0 API 列表
|
|
79
|
+
|
|
80
|
+
传输:`POST` 到 provider decide 端点(Vercel System One 或 FreeJev),
|
|
81
|
+
`Authorization: Bearer $AI_GATEWAY_API_KEY` / `$FREEJEV_API_KEY`,
|
|
82
|
+
`model=typesafe-ai/jev`。
|
|
83
|
+
请求体 TypeSafe System One:`{ model, state, questions }` → `{ model, answers, usage }`。
|
|
84
|
+
FreeJev 另带 `request_id`(见上:**不重试**)。
|
|
85
|
+
|
|
86
|
+
| # | Python API | capability_id | MCP 工具 | SemDecide 对应 | 问题形态 | 默认阈值 |
|
|
87
|
+
|---|---|---|---|---|---|---|
|
|
88
|
+
| 1 | `jev_mcp.screen(text, purpose)` | `content_screen` | `jev_screen` | — | 3× Noul:injection / substance / relevance | `block_at=0.75`, `review_at=0.25` |
|
|
89
|
+
| 2 | `jev_mcp.rerank(query, candidates)` | `semantic_rerank` | `jev_rerank` | — | 1× Noul / candidate | `auto_accept=0.70` |
|
|
90
|
+
| 3 | `jev_mcp.extract(document, fields)` | `field_extract` | `jev_extract` | — | 1× Choice / field(regex 候选 + `none_of_them`) | `auto_accept=0.80`, `margin=0.40` |
|
|
91
|
+
| 4 | `jev_mcp.classify(items, classes)` | `classify_check` | `jev_classify` | `choose` 可路由 | 1× Choice / item | `auto_accept=0.85`, `margin=0.50` |
|
|
92
|
+
| 5 | `jev_mcp.verify(claims, evidence)` | `claim_verify` | `jev_verify` | `is` 可预筛 | 1× Choice / claim:supports / contradicts / says_nothing | `auto_accept=0.80` |
|
|
93
|
+
|
|
94
|
+
统一入口:`jev_mcp.safe_call(capability_id, ...)`(先过 approval 门)。
|
|
95
|
+
|
|
96
|
+
SemDecide 封装(exit 2/3/4 → 升级大模型):
|
|
97
|
+
|
|
98
|
+
| API | CLI | 用途 |
|
|
99
|
+
|---|---|---|
|
|
100
|
+
| `semdecide.is_predicate(text, predicate)` | `semdecide is` | 语义谓词 |
|
|
101
|
+
| `semdecide.filter_records(records, predicate)` | `semdecide filter` | JSONL 语义过滤(保序) |
|
|
102
|
+
| `semdecide.choose_route(text, question, options)` | `semdecide choose` | 命名选项路由 |
|
|
103
|
+
|
|
104
|
+
### 状态与 fail-closed(与 `integrations/semantic_decide.py` 对齐)
|
|
105
|
+
|
|
106
|
+
Jev / Tier-0:
|
|
107
|
+
|
|
108
|
+
| 状态 | 含义 | 管道行为 |
|
|
109
|
+
|---|---|---|
|
|
110
|
+
| `ok` + `action/auto` | 可自动采用(仍不是科学门) | 记录 provenance,继续阶段 |
|
|
111
|
+
| `review` / `not_found` / `JEV_INVALID_RESPONSE` | 低置信或非法答案(envelope 缺 `answers` 等) | **升级大模型**;extract 的 `review` 值不得当作已抽取 |
|
|
112
|
+
| `JEV_APPROVAL_REQUIRED` | 无/错 approval | 不发网络请求 |
|
|
113
|
+
| `JEV_UNAVAILABLE` | 无密钥 / 无可用传输 | 本调用升级大模型;hybrid 可切 mode 2 切片 |
|
|
114
|
+
| `JEV_PROVIDER_ERROR` | 上游失败(HTTP 401/403/429/529/5xx) | 本调用升级大模型;FreeJev **`request_id` 不重试** |
|
|
115
|
+
|
|
116
|
+
SemDecide(`EXIT_MEANINGS` / `ESCALATE_EXIT_CODES={2,3,4}`):
|
|
117
|
+
|
|
118
|
+
| exit / status | `EXIT_MEANINGS` | 管道行为 |
|
|
119
|
+
|---|---|---|
|
|
120
|
+
| `0` | `true/selected/match` | 语义结果可用 |
|
|
121
|
+
| `1` | `false/no match` | 语义结果可用;`filter` 的 exit 1 **不**单独升级 |
|
|
122
|
+
| **`2`** | `invalid input` | 修输入重跑,或整段决策**升级大模型**(`escalate_plan`) |
|
|
123
|
+
| **`3`** | `uncertain` | **升级大模型 / 人工**;禁止强行 true/false |
|
|
124
|
+
| **`4`** | `provider failure` | **升级大模型**(仅本调用);记录 attempt |
|
|
125
|
+
| `SEMDECIDE_UNAVAILABLE` | binary 缺失 / spawn 失败 | **升级大模型**(`requires_llm_escalation`) |
|
|
126
|
+
| `SEMDECIDE_TIMEOUT` | 超时 | **升级大模型**(`result.escalate=ESCALATE_TO_LLM`) |
|
|
127
|
+
|
|
128
|
+
`screen` 的 recommendation(pass / review / block / skip)仅为咨询;拦截执行权在调用方。
|
|
129
|
+
|
|
130
|
+
## 4. 阈值
|
|
131
|
+
|
|
132
|
+
| 参数 | 默认 | 调节建议 |
|
|
133
|
+
|---|---:|---|
|
|
134
|
+
| `screen_block_at` | 0.75 | 注入概率 ≥ 则 block(咨询) |
|
|
135
|
+
| `screen_review_at` | 0.25 | 注入概率 ≥ 则 review |
|
|
136
|
+
| `verify_auto_accept` | 0.80 | 关系裁决 confidence ≥ 才 `auto` |
|
|
137
|
+
| `classify_auto_accept` | 0.85 | top 概率门槛 |
|
|
138
|
+
| `classify_minimum_margin` | 0.50 | 冠亚军差门槛(二者同时满足才 auto) |
|
|
139
|
+
| `extract_auto_accept` | 0.80 | 字段挑选 confidence |
|
|
140
|
+
| `extract_minimum_margin` | 0.40 | 候选差门槛;失败 → `review` |
|
|
141
|
+
| `rerank_auto_accept` | 0.70 | 相关 Noul 门槛 |
|
|
142
|
+
| SemDecide `--threshold` | 0.70 | `is` / `filter` |
|
|
143
|
+
| SemDecide `--min-confidence` | 0.70 | `choose`;低于则 exit 3 |
|
|
144
|
+
| SemDecide `--uncertainty-margin` | 0.0 | 边缘区不强行二值化 |
|
|
145
|
+
|
|
146
|
+
阈值是起点(TypeSafe cookbook 基线),**先在自有数据上校准再强制**。
|
|
147
|
+
|
|
148
|
+
## 5. 九阶段映射(加速点)
|
|
149
|
+
|
|
150
|
+
`frame → retrieve → extract → challenge → audit → adjudicate → applicability → intervene → evaluate`
|
|
151
|
+
|
|
152
|
+
| 阶段 | Jev Tier-0 | SemDecide | 不变的科学门 |
|
|
153
|
+
|---|---|---|---|
|
|
154
|
+
| Frame | — | `choose` 可选 | frame schema |
|
|
155
|
+
| Retrieve | screen / rerank / classify | `filter`, `is` | 检索计划 + 反方检索 + Fetch/Validate |
|
|
156
|
+
| Extract | extract / classify | `is` 预检 | Evidence Objects、task ≠ learning |
|
|
157
|
+
| Challenge | **verify 仅辅助** | `is` 辅助 | **Skeptic 九项检查** |
|
|
158
|
+
| Audit | verify 辅助 | `is` 预筛 | 方法学审计 + guard |
|
|
159
|
+
| Adjudicate | —(禁止 Tier-0 定裁决) | — | **Pre-Verdict Gate** + 四态裁决 |
|
|
160
|
+
| Applicability | classify 可选 | `choose` 可选 | 人群/条件/排除/不确定性 |
|
|
161
|
+
| Intervene | — | — | intervention_design |
|
|
162
|
+
| Evaluate | — | — | evaluation_design / data_validation |
|
|
163
|
+
|
|
164
|
+
完整矩阵见 `skill/workflows/experimental-jev.md`。
|
|
165
|
+
|
|
166
|
+
## 6. 加速比口径(speedup)
|
|
167
|
+
|
|
168
|
+
只比较**同一问题、同一 Complexity Gate、同一九阶段协议**下 mode 0 基线 vs 实验模式。
|
|
169
|
+
|
|
170
|
+
| 指标 | 定义 | 统计范围 |
|
|
171
|
+
|---|---|---|
|
|
172
|
+
| `T_stage` | 单阶段 wall-clock(s),从进入阶段到产物过 schema/gate | 仅计 **gate 通过** 的阶段 |
|
|
173
|
+
| `T_run` | Frame→Applicability(或声明的阶段子集)累计 | 同上 |
|
|
174
|
+
| `speedup_stage` | `T_stage(mode0) / T_stage(exp)` | per stage |
|
|
175
|
+
| `speedup_run` | `T_run(mode0) / T_run(exp)` | per run |
|
|
176
|
+
| `cost_tokens` | 主模型 + Tier-0 `usage.input_tokens/output_tokens` | 双方同口径合计 |
|
|
177
|
+
| `escalation_rate` | fail-closed 升级次数 / Tier-0 调用次数 | 实验侧 |
|
|
178
|
+
| `gate_pass_rate` | 过科学门的阶段数 / 执行阶段数 | 双方 |
|
|
179
|
+
|
|
180
|
+
规则:
|
|
181
|
+
|
|
182
|
+
1. **质量门优先**:`gate_pass_rate(exp) < gate_pass_rate(mode0)` 时 speedup **无效**(质量退化不算加速)。
|
|
183
|
+
2. Tier-0 调用耗时与 token 计入实验侧成本,不得只计大模型。
|
|
184
|
+
3. 升级到大模型的调用耗时计入实验侧(含往返),避免“看起来快”。
|
|
185
|
+
4. 未过 gate / 中断 / 人工改写的 run **不进入** speedup 均值,单独列表。
|
|
186
|
+
5. 报告格式:每 run 一行 `{run_id, mode, T_run, speedup_run, cost_tokens, escalation_rate, gate_pass_rate}`。
|
|
187
|
+
|
|
188
|
+
示例(示意,非承诺值):
|
|
189
|
+
|
|
190
|
+
| run | mode | T_run | speedup_run | escalations | gate_pass |
|
|
191
|
+
|---|---:|---:|---:|---:|---:|
|
|
192
|
+
| ai-cs1-b | 0 | 3120s | 1.00× | — | 7/7 |
|
|
193
|
+
| ai-cs1-j | 3 | 1810s | 1.72× | 4/38 | 7/7 |
|
|
194
|
+
|
|
195
|
+
## 7. A/B 验收
|
|
196
|
+
|
|
197
|
+
| 项 | 通过条件 |
|
|
198
|
+
|---|---|
|
|
199
|
+
| 科学完整性 | 挑战阶段九项 Skeptic 全跑;Adjudicate 前 `pre_verdict_gate` 必跑且结论不被 Tier-0 覆盖 |
|
|
200
|
+
| 质量非劣 | `gate_pass_rate` 不低于基线;Pre-Verdict critical 失败数不增加 |
|
|
201
|
+
| 校准 | `verify`/`classify` 的 auto 集抽样人工一致率 ≥ 约定阈值(建议 ≥ 0.9) |
|
|
202
|
+
| Fail-closed | 抽检 100% 的 SemDecide exit 2/3/4、`SEMDECIDE_TIMEOUT`/`SEMDECIDE_UNAVAILABLE` 与 `JEV_INVALID_RESPONSE`/`JEV_PROVIDER_ERROR` 都有升级记录 |
|
|
203
|
+
| 密钥卫生 | 仓库与 run 产物无 `AI_GATEWAY_API_KEY` / `TYPESAFE_API_KEY` / `FREEJEV_API_KEY` 明文 |
|
|
204
|
+
| FreeJev 幂等 | 抽检无同一 `request_id` 重试;失败一律升级大模型 |
|
|
205
|
+
| 可复现 | mode、approval `tools_hash`、阈值、上游 model id 写入 run 元数据 |
|
|
206
|
+
| 加速有效 | `speedup_run ≥ 1.2`(目标,可调)且质量非劣;否则保持 mode 0 |
|
|
207
|
+
|
|
208
|
+
## 8. 降级路径
|
|
209
|
+
|
|
210
|
+
```
|
|
211
|
+
--experimental 3 (hybrid)
|
|
212
|
+
├─ provider key + jev_mcp_approval.json → Tier-0 screen/rerank/extract/classify/verify
|
|
213
|
+
│ (FreeJev: FREEJEV_API_KEY + JEV_PROVIDER=freejev;request_id 不重试)
|
|
214
|
+
├─ semdecide on PATH → is/filter/choose
|
|
215
|
+
├─ exit 2/3/4 / SEMDECIDE_TIMEOUT
|
|
216
|
+
│ / JEV_INVALID_RESPONSE / JEV_PROVIDER_ERROR
|
|
217
|
+
│ → 单次升级大模型(fail-closed)
|
|
218
|
+
├─ Jev 不可用 → 仅 SemDecide(mode 2 切片)
|
|
219
|
+
├─ semdecide 不可用 → 仅 Jev(mode 1 切片)
|
|
220
|
+
└─ 双不可用 / 未开 --experimental → mode 0 standard
|
|
221
|
+
|
|
222
|
+
科学门永不降级:Pre-Verdict Gate / Skeptic 九项 / 方法学审计 / Tribunal 始终在宿主路径执行。
|
|
223
|
+
```
|
|
224
|
+
|
|
225
|
+
## 9. API 速查(代码)
|
|
226
|
+
|
|
227
|
+
```python
|
|
228
|
+
from integrations.jev_mcp import (
|
|
229
|
+
detect_jev_mcp, write_approval, load_approval,
|
|
230
|
+
screen, rerank, extract, classify, verify, safe_call,
|
|
231
|
+
resolve_experimental_mode, npx_mcp_guide,
|
|
232
|
+
)
|
|
233
|
+
from integrations.semantic_decide import (
|
|
234
|
+
detect_semdecide, is_predicate, filter_records, choose_route,
|
|
235
|
+
requires_llm_escalation, escalate_plan,
|
|
236
|
+
)
|
|
237
|
+
|
|
238
|
+
mode = resolve_experimental_mode(3) # 0..3
|
|
239
|
+
approval = load_approval() # ~/.eduevidence/jev_mcp_approval.json
|
|
240
|
+
verdict = safe_call("claim_verify", claims, evidence, approval=approval)
|
|
241
|
+
if verdict.get("escalate"):
|
|
242
|
+
... # fail-closed → large model
|
|
243
|
+
```
|
|
244
|
+
|
|
245
|
+
## 10. 变更记录
|
|
246
|
+
|
|
247
|
+
| 日期 | 内容 |
|
|
248
|
+
|---|---|
|
|
249
|
+
| 2026-09-22 | 首版:Tier-0 五能力、approval 门、mode 0–3、阈值、加速比口径、A/B 验收、降级路径 |
|
|
250
|
+
| 2026-09-22 | 收口:FreeJev(`FREEJEV_API_KEY` / `JEV_PROVIDER=freejev` / MCP `https://freejev.org/mcp` / `request_id` 不重试);fail-closed 表与 `semantic_decide.py` 对齐 |
|
|
@@ -0,0 +1,138 @@
|
|
|
1
|
+
# EduEvidence 复现指南
|
|
2
|
+
|
|
3
|
+
> **复现范围声明(v5.2.0 起,R6 口径收敛)**:可复现承诺分两层,不得混谈。
|
|
4
|
+
>
|
|
5
|
+
> 1. **确定性层(完全可复现)**:schema 校验、置信度计算(规则化公式)、pre-verdict gate、
|
|
6
|
+
> meta-analysis/robustness 数学、报告烘焙与哈希链——同一输入在同一版本上逐字节可复现;
|
|
7
|
+
> 本指南的三步走覆盖这一层。
|
|
8
|
+
> 2. **LLM 层(有边界可复现)**:检索、抽取、挑战、裁决等依赖外部 LLM 的阶段不是逐字节
|
|
9
|
+
> 确定的。其复现条件 = 模型版本 + temperature + prompt 版本 三者同时固定;当前仓库通过
|
|
10
|
+
> run manifest 记录这三项(如 `benchmarks/empirical/run-empirical-01/manifest.json` 的
|
|
11
|
+
> `temperature: 0.0`),但不承诺跨模型供应商的结果一致。
|
|
12
|
+
>
|
|
13
|
+
> 引擎对"证据是否真实"的立场:所有示例包在 `result.json.meta.data_origin` 与报告头徽章中
|
|
14
|
+
> 如实标注数据来源(`real` / `manual_curated` / `synthetic` / `hybrid`);引用的 DOI 经
|
|
15
|
+
> `scripts/audit_dois.py` 对照 Crossref/DataCite 注册表核验(报告:
|
|
16
|
+
> `benchmarks/doi-audit/report.md`)。标注 `synthetic` 的演示包不构成任何实证声明。
|
|
17
|
+
|
|
18
|
+
本指南确保任何人在全新环境中都能复现 EduEvidence 确定性层的校验、测试与 Benchmark 结果。全部命令以仓库根目录 `<repo-root>`(即 pyproject.toml 所在目录)为工作目录执行。
|
|
19
|
+
|
|
20
|
+
## 〇、V2 Research Engine 复现
|
|
21
|
+
|
|
22
|
+
- 引擎模块 `engine/`(stdlib-only)可独立复现:`python3 -m pytest tests/test_v2_*.py`。
|
|
23
|
+
- 图状态复现:`graph/HEAD` + `revisions/rev-N` 快照 + `manifest.json` 的 before/after 哈希链;同一输入 bundle 提交产生确定性哈希。
|
|
24
|
+
- V1 历史产物(`examples/ai-coding-assistant/`、旧 `runs/`)不可变;`eduevidence migrate-v1 --pack <dir>` 生成新 V2 Project 且不改动源包。
|
|
25
|
+
- Shared Research Library 与 Project 图同用不可变 revision 模型;快照导入记录 `library_revision/library_entity_id/content_hash/imported_at`。
|
|
26
|
+
- 测试覆盖:V1 兼容基线、图原子性/孤儿 revision/HEAD 镜像分歧修复、迁移不变量、模式推荐、能力规划、synthesis 独立计数、Tribunal 政策、Gap 推导、数据集隐私门、全周期单 revision 提交、决策 diff、CLI 薄分发。
|
|
27
|
+
|
|
28
|
+
## 一、环境要求
|
|
29
|
+
|
|
30
|
+
- **Python**:≥ 3.10(pyproject.toml 中 `requires-python = ">=3.10"`);
|
|
31
|
+
- **pytest**:≥ 7.0(dev 依赖,安装时自动带入);
|
|
32
|
+
- 无其它强制运行时依赖:EduEvidence 本体是文档+Schema+脚本,Mode A 平台原生模式下不需要 daemon / CLI / Agent MCP;
|
|
33
|
+
- 操作系统:macOS / Linux / Windows 均可;建议使用 venv 隔离环境。
|
|
34
|
+
|
|
35
|
+
### 1.1 创建虚拟环境(推荐)
|
|
36
|
+
|
|
37
|
+
```bash
|
|
38
|
+
cd <repo-root>
|
|
39
|
+
python3.11 -m venv .venv
|
|
40
|
+
source .venv/bin/activate # Windows: .venv\Scripts\activate
|
|
41
|
+
```
|
|
42
|
+
|
|
43
|
+
## 二、安装步骤
|
|
44
|
+
|
|
45
|
+
在虚拟环境内安装本项目(含 dev 依赖):
|
|
46
|
+
|
|
47
|
+
```bash
|
|
48
|
+
pip install -e '.[dev]'
|
|
49
|
+
```
|
|
50
|
+
|
|
51
|
+
- `-e`(editable)安装使脚本与 schema 变更即时生效,无需重装;
|
|
52
|
+
- `.[dev]` 安装 `pytest>=7.0`;
|
|
53
|
+
- 安装成功后可用 `python -c "import sys; print(sys.version)"` 确认版本 ≥ 3.10。
|
|
54
|
+
|
|
55
|
+
## 三、运行测试
|
|
56
|
+
|
|
57
|
+
```bash
|
|
58
|
+
pytest
|
|
59
|
+
```
|
|
60
|
+
|
|
61
|
+
- 测试自动发现 `tests/` 下 `test_*.py` 文件(见 pyproject.toml `[tool.pytest.ini_options]`);
|
|
62
|
+
- 覆盖范围:schema 校验器、confidence 规则化计算、Complexity Gate 判级、benchmark 输出格式等;
|
|
63
|
+
- 预期结果:全部测试通过(`addopts = "-q"` 输出简洁)。
|
|
64
|
+
|
|
65
|
+
## 四、运行 Schema 验证
|
|
66
|
+
|
|
67
|
+
校验某份数据是否符合对应 Schema(mandatory 字段缺失会标记为 UNSUPPORTED):
|
|
68
|
+
|
|
69
|
+
```bash
|
|
70
|
+
python scripts/validate_schema.py \
|
|
71
|
+
--schema schemas/evidence.schema.json \
|
|
72
|
+
--data examples/ai-coding-assistant/evidence.jsonl
|
|
73
|
+
```
|
|
74
|
+
|
|
75
|
+
- `--schema`:指定任一顶层 Schema(schemas/ 下 37 个契约,跨 V1–V4 代际,如 education-frame / evidence / methodology / verdict / intervention / evaluation / report-result 等);
|
|
76
|
+
- `--data`:指定要校验的 JSON 或 JSONL 数据文件;
|
|
77
|
+
- 主 Demo 示例数据位于 `examples/ai-coding-assistant/evidence.jsonl`,验证通过时应输出每条 evidence 的校验状态与 SUPPORTED/UNSUPPORTED 统计;
|
|
78
|
+
- 合格线:**校验通过率 100%**,任一条未通过即非零退出码。
|
|
79
|
+
|
|
80
|
+
## 五、运行 Benchmark
|
|
81
|
+
|
|
82
|
+
逐题运行五条基线(B0–B4)并生成结果:
|
|
83
|
+
|
|
84
|
+
```bash
|
|
85
|
+
python scripts/benchmark.py --questions benchmarks/questions.jsonl
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
- `--questions`:第一版 30 题(S×10 / M×10 / L×10),每题字段为 `id`、`level`(S/M/L)、`domain`、`question`、`expected_outcomes`、`notes`(与 `validate_questions()` 校验一致);人工金标注位于 `benchmarks/annotations/gold-<id>.json`;
|
|
89
|
+
- 运行产物:每题一个 JSON,写入 `benchmarks/results/`;
|
|
90
|
+
- 核心指标(Citation Support Precision、Unsupported Claim Rate、Contradiction Discovery Rate、Outcome Separation Accuracy、Scope Calibration、Intervention Evidence Alignment)由 `benchmarks/evaluator/` 对照 `benchmarks/annotations/` 计算;
|
|
91
|
+
- 可选参数(如 `--ablation` 跑 A1–A7、`--repeat 5` 测稳定性)以 `python scripts/benchmark.py --help` 为准。
|
|
92
|
+
|
|
93
|
+
## 六、输出产物:Research & Decision Pack
|
|
94
|
+
|
|
95
|
+
每次端到端运行(Demo 或 L 级题目)产出一份双层 Research & Decision Pack:
|
|
96
|
+
|
|
97
|
+
- **Visual Brief**:用于快速浏览 Decision、Outcome Separation、Evidence Tribunal、Evidence-to-Action 与关键来源。
|
|
98
|
+
- **Full Report**:由当前研究内容规划为 **5–7 个动态章节**,而不是固定章节模板。
|
|
99
|
+
|
|
100
|
+
完整报告必须覆盖以下语义模块;这些模块可以按研究叙事合并进同一章节,但不得遗漏:
|
|
101
|
+
|
|
102
|
+
| 模块 | 内容 |
|
|
103
|
+
|------|------|
|
|
104
|
+
| decision | 最终决策、Confidence、可说/不可说与主要风险 |
|
|
105
|
+
| scope | Education Research Frame、研究边界与决策标准 |
|
|
106
|
+
| retrieval | 检索策略、来源覆盖与证据选择 |
|
|
107
|
+
| outcomes | Outcome Separation:任务表现、学习、保持、迁移与风险 |
|
|
108
|
+
| evidence | Claim-Level Evidence 与 Evidence Matrix |
|
|
109
|
+
| quality | Methodology / Evidence Quality Audit |
|
|
110
|
+
| conflicts | Skeptic / Counter-Evidence / Conflict Analysis |
|
|
111
|
+
| trace | Evidence Tribunal + Claim → Evidence → Source 追溯 |
|
|
112
|
+
| applicability | 适用性、外推边界与必要条件 |
|
|
113
|
+
| intervention | 最小可验证干预、Guardrails 与 Stop Conditions |
|
|
114
|
+
| evaluation | Baseline / Post-test / Retention / Transfer 评价方案 |
|
|
115
|
+
| sources | Sources、Provenance 与附录 |
|
|
116
|
+
|
|
117
|
+
Pack 的证据链始终可以从 Decision / Claim 反查到 Evidence,再追溯到 Source 与 source location;章节如何合并不改变这一追溯关系。
|
|
118
|
+
|
|
119
|
+
## 七、常见问题(FAQ)
|
|
120
|
+
|
|
121
|
+
| 现象 | 处理 |
|
|
122
|
+
|------|------|
|
|
123
|
+
| `pytest` 无测试收集 | 确认在仓库根目录运行且依赖已安装 |
|
|
124
|
+
| validate_schema 报非零退出 | 逐条查看 UNSUPPORTED 记录,补齐 mandatory 字段 |
|
|
125
|
+
| benchmark 输出缺失 | 检查 `benchmarks/results/` 目录可写;确认 `benchmarks/questions.jsonl` 字段完整 |
|
|
126
|
+
| pip install 找不到 `[dev]` | 确认工作目录是根目录(pyproject.toml 存在) |
|
|
127
|
+
| Python 版本低于 3.10 | 升级解释器或使用 pyenv 安装 ≥3.10 |
|
|
128
|
+
|
|
129
|
+
## 八、端到端复现三步走
|
|
130
|
+
|
|
131
|
+
```bash
|
|
132
|
+
pip install -e '.[dev]' # 1. 安装
|
|
133
|
+
pytest # 2. 验证环境
|
|
134
|
+
python scripts/validate_schema.py --schema schemas/evidence.schema.json --data examples/ai-coding-assistant/evidence.jsonl
|
|
135
|
+
python scripts/benchmark.py --questions benchmarks/questions.jsonl # 3. 复现示例与基准
|
|
136
|
+
```
|
|
137
|
+
|
|
138
|
+
三步全部成功即视为复现完成。
|
|
@@ -0,0 +1,125 @@
|
|
|
1
|
+
# Sciverse 检索通道(API 契约存档)
|
|
2
|
+
|
|
3
|
+
本条记录 EduEvidence 接入 Sciverse 开放平台所用的最小契约,依据官方 `openapi.yaml` **v0.14.2**(`opendatalab/Sciverse-Agent-Tools`)整理。目的是让检索链在**没有网络**时也能被复核:字段名、错误码、限制与语义都在这里,代码变更需同步更新本文件。
|
|
4
|
+
|
|
5
|
+
## 1. 基本信息
|
|
6
|
+
|
|
7
|
+
| 项 | 值 |
|
|
8
|
+
|---|---|
|
|
9
|
+
| Base URL | `https://api.sciverse.space` |
|
|
10
|
+
| 鉴权 | `Authorization: Bearer <SCIVERSE_API_TOKEN>`(HTTP Bearer) |
|
|
11
|
+
| Token 来源 | 控制台 Tokens 页;同账号可通用于 Sciverse / DianShi / SeqStudio 已开通能力 |
|
|
12
|
+
| 实现 | `retrieval/sciverse.py`(stdlib-only) |
|
|
13
|
+
| 通道类型 | key-based 学术通道,在 `MultiSearchRouter.academic_key_providers` 中优先于零配置学术通道 |
|
|
14
|
+
| 无 token 行为 | 通道静默失活(`SCIVERSE_UNAVAILABLE`),零配置通道继续工作 |
|
|
15
|
+
|
|
16
|
+
## 2. 使用的四个端点
|
|
17
|
+
|
|
18
|
+
### 2.1 `POST /meta-search` — 结构化元数据检索
|
|
19
|
+
|
|
20
|
+
用于 Source 级命中:标题、作者、摘要、期刊、年份、DOI。
|
|
21
|
+
|
|
22
|
+
- 请求(所用子集):`collection`(papers/authors/sources)、`query`(BM25)、`filters_advanced[]`(`{field, operator, value}`)、`page`、`page_size`(≤50)。
|
|
23
|
+
- 响应:`results[]`(注意是 `results`,不是 `hits`)、`total_count`、`page`、`page_size`、`total_pages`、`next_cursor`。
|
|
24
|
+
- 关键字段:`unique_id`(元数据全局唯一 ID,**始终存在**)、`doc_id`(全文内容哈希 sha256,**仅当存在全文**)、`is_content_accessible`、`doi`、`author[].name`、`publication_published_year`、`publication_venue_name_unified`。
|
|
25
|
+
- 过滤操作符:`FILTER_OP_EQ/NE/GT/GTE/LT/LTE/IN/NIN/CONTAINS/MATCH/MATCH_PHRASE`;`doi` 用 `EQ`(服务端去 `doi.org` 前缀并转小写后精确匹配)。
|
|
26
|
+
- 排序与加权:`sort_by_year`(`auto` 在带 query 时保留相关性排序)、`freshness_boost` / `impact_boost` / `language_affinity`(`NONE|MILD|STRONG`,仅在 query 非空时生效;加权生效时**不支持深翻页**)。
|
|
27
|
+
|
|
28
|
+
### 2.2 `POST /agentic-search` — 自然语言语义检索(RAG)
|
|
29
|
+
|
|
30
|
+
用于 chunk 级**定位子**:
|
|
31
|
+
|
|
32
|
+
- 请求:`query`(1–200 字最佳)、`top_k`(1–100)、`mode`(`fast` ~200ms / `balanced` ~600ms / `quality` ~2–4s)、可选 `filters`、`source_types`(`web` / `pdf`)。
|
|
33
|
+
- 响应:`hits[]`,每条含 `chunk_id`、`doc_id`、`score`、`title`、`offset`(**Unicode 码点**,可直接作为 `/content` 的 offset)、`page_no`、`source_type`、`chunk`、`abstract`。
|
|
34
|
+
- **限制**:`balanced` 模式服务端约截断至 50 条;同一篇论文最多返回约 3 个 chunk,因此高 `top_k` 需要足够多的不同论文。
|
|
35
|
+
- **软过滤语义**:`filters` 在召回阶段与语义检索同时下推,但 chunk 侧元数据缺失的文档不会被排除(按年份过滤时,缺年份的 chunk 仍可能返回)。结论表述必须写"近似范围";需要严格范围时改用 `meta-search` 的结构化过滤并核对返回记录。
|
|
36
|
+
- 唯一的硬过滤字段是 `doc_id`(命中绝不越出给定集合;去重后上限默认 1000)。
|
|
37
|
+
|
|
38
|
+
### 2.3 `GET /content` — 按码点区间读原文
|
|
39
|
+
|
|
40
|
+
把 chunk 定位子扩展为可抽取的正文:
|
|
41
|
+
|
|
42
|
+
- 参数:`doc_id`(必填)、`offset`(默认 0)、`limit`(默认 4096,服务端上限 524288,超出静默钳制)。
|
|
43
|
+
- **必须显式传 `offset`**:省略时服务端返回整篇全文并忽略 `limit`。
|
|
44
|
+
- 响应:`text`、`bytes_returned`(UTF-8 字节数,仅供参考)、`next_offset`(下一段起始码点,翻页用它而非字节数)、`more`。
|
|
45
|
+
- 单位:`offset` / `limit` 均以 **Unicode 码点**计,与 Python `len` 一致,不是字节。
|
|
46
|
+
|
|
47
|
+
### 2.4 `POST /meta-paper-relations` — 引用 / 被引 / 相关工作
|
|
48
|
+
|
|
49
|
+
用于引文链审计(`SearchQuery.purpose = citation_chain`):
|
|
50
|
+
|
|
51
|
+
- 请求:`unique_id`(**不是 doc_id**)、`relation`(`CITATIONS` 被引 / `REFERENCES` 参考文献 / `RELATED_WORKS`)、`page`、`page_size`(≤200)。
|
|
52
|
+
- 响应:`items[]`(`id` / `id_type` / `title`)、`total_count`、`page`、`page_size`、`total_pages`。
|
|
53
|
+
- 方向语义:`CITATIONS` 是"谁引用了我",`REFERENCES` 是"我引用了谁",两者相反。
|
|
54
|
+
- **上限**:关系数超 10000 返回 `429`;`page × page_size` 超 10000 返回 `400`。两种情况改用 `meta-search` 的 `references_unique_id` 反查(支持深翻页与任意排序)。
|
|
55
|
+
- `total_count` 只统计库内命中,与论文自身 `citation_count` 可能有约 ±1% 差异。
|
|
56
|
+
|
|
57
|
+
## 3. 错误处理与状态映射
|
|
58
|
+
|
|
59
|
+
`retrieval/sciverse.py` 把 HTTP 与网络异常统一映射为**定型状态**,绝不把异常抛进检索管道:
|
|
60
|
+
|
|
61
|
+
| HTTP / 情形 | 状态 | 管道行为 |
|
|
62
|
+
|---|---|---|
|
|
63
|
+
| 200 | `ok` | 正常解析 |
|
|
64
|
+
| 401 / 403 | `SCIVERSE_UNAUTHORIZED` | 记录尝试,切换其他通道 |
|
|
65
|
+
| 400 / 404 / 429 | `SCIVERSE_BAD_REQUEST` | 同上(429 视为配额/上限耗尽) |
|
|
66
|
+
| 502 / 503 / 5xx | `SCIVERSE_UPSTREAM_ERROR` | 同上 |
|
|
67
|
+
| 超时 / DNS / TLS | `SCIVERSE_NETWORK_ERROR` | 同上 |
|
|
68
|
+
| 未配置 token | `SCIVERSE_UNAVAILABLE` | 通道失活,不产生请求 |
|
|
69
|
+
|
|
70
|
+
错误信息只保留服务端 `ApiError.message` 或状态描述,**不携带 Authorization 头或 token 片段**。
|
|
71
|
+
|
|
72
|
+
## 4. 在证据链中的位置(RULE 2 的机器化)
|
|
73
|
+
|
|
74
|
+
```text
|
|
75
|
+
/meta-search → Source 级命中(DOI → https://doi.org/<doi>;无 DOI 则标记 needs_manual_location)
|
|
76
|
+
/agentic-search → chunk 定位子(doc_id + offset)→ 写入 chunks.jsonl,标注 discovery_only_requires_content_fetch
|
|
77
|
+
/content → 按定位读原文 → 经 retrieval/validate.py 校验门 → 才可进入 Extract
|
|
78
|
+
/meta-paper-relations → 引文链记录(写入审计导出,供筛选与饱和判断)
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
要点:**chunk 不是证据**。它是发现线索;只有 `/content` 读到的正文通过校验门(长度、错误页、登录页、验证码、标题匹配等)之后,才允许被抽取为 Evidence Object。这条纪律由 `retrieval/fetch.py::fetch_sciverse_content()` 实现,产出与 `FetchResult` 同形,下游零改动。
|
|
82
|
+
|
|
83
|
+
引用目标永远是论文本身(DOI / `unique_id`),**不是 Sciverse** —— 它是读取路径,不是来源。
|
|
84
|
+
|
|
85
|
+
## 5. 配置
|
|
86
|
+
|
|
87
|
+
```bash
|
|
88
|
+
export SCIVERSE_API_TOKEN=sv-... # 必需;未设置时通道失活
|
|
89
|
+
# 可选:指向测试环境
|
|
90
|
+
export SCIVERSE_BASE_URL=https://api.sciverse.space
|
|
91
|
+
```
|
|
92
|
+
|
|
93
|
+
宿主侧如已安装官方 Skill / MCP,可作为**补充**而非替代(本仓库的实现不依赖它):
|
|
94
|
+
|
|
95
|
+
```bash
|
|
96
|
+
npx skills add https://sciverse.space
|
|
97
|
+
# 或 MCP server
|
|
98
|
+
npm install -g sciverse-mcp-server
|
|
99
|
+
export SCIVERSE_API_TOKEN=sv-...
|
|
100
|
+
```
|
|
101
|
+
|
|
102
|
+
## 6. 验证
|
|
103
|
+
|
|
104
|
+
```bash
|
|
105
|
+
# 契约与降级行为(离线,mock HTTP)
|
|
106
|
+
python -m pytest tests/test_sciverse_channel.py -q
|
|
107
|
+
|
|
108
|
+
# 真实连通冒烟(需 token)
|
|
109
|
+
python - <<'PY'
|
|
110
|
+
from retrieval import sciverse as s
|
|
111
|
+
print('available:', s.available())
|
|
112
|
+
print('meta:', s.meta_search('generative AI coding assistants learning outcomes', limit=3).status)
|
|
113
|
+
r = s.agentic_search('unguarded GPT-4 access harms independent exam performance', top_k=3)
|
|
114
|
+
recs = s.chunk_records(r)
|
|
115
|
+
print('chunks:', len(recs))
|
|
116
|
+
if recs:
|
|
117
|
+
c = s.read_content(recs[0]['doc_id'], offset=recs[0]['offset'], limit=600)
|
|
118
|
+
print('content:', c.status, len(c.data.get('text') or ''))
|
|
119
|
+
PY
|
|
120
|
+
```
|
|
121
|
+
|
|
122
|
+
## 7. 合规
|
|
123
|
+
|
|
124
|
+
配额、限速、缓存与署名规则统一见 `references/retrieval-compliance.md`(§3 Sciverse 附加约定)。规范漂移时以官方 `openapi.yaml` 为准,并同步更新本文件与 `references/retrieval-protocol.md`。
|
|
125
|
+
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
{
|
|
2
|
+
"domain": "_neutral",
|
|
3
|
+
"states": {
|
|
4
|
+
"adopt": {
|
|
5
|
+
"zh": "",
|
|
6
|
+
"en": ""
|
|
7
|
+
},
|
|
8
|
+
"pilot": {
|
|
9
|
+
"zh": "",
|
|
10
|
+
"en": ""
|
|
11
|
+
},
|
|
12
|
+
"reject": {
|
|
13
|
+
"zh": "",
|
|
14
|
+
"en": ""
|
|
15
|
+
},
|
|
16
|
+
"insufficient_evidence": {
|
|
17
|
+
"zh": "",
|
|
18
|
+
"en": ""
|
|
19
|
+
}
|
|
20
|
+
}
|
|
21
|
+
}
|