eduevidence 6.0.0 → 6.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +395 -0
- package/CONTRIBUTING.md +105 -0
- package/README.md +113 -49
- package/README.zh-CN.md +39 -12
- package/SKILL.md +15 -5
- package/assets/readme/landing-tour.gif +0 -0
- package/assets/readme/studio-tour.gif +0 -0
- package/benchmarks/evidence-library.json +277 -1
- package/bin/eduevidence.js +2 -1
- package/docs/architecture.md +325 -46
- package/docs/demo-workplace-ai.md +1 -1
- package/docs/install-guide.md +1 -1
- package/docs/j-ev-experimental.md +250 -0
- package/docs/orchestration-role-model.md +1 -1
- package/docs/release-closeout/README.md +1 -1
- package/docs/reproducibility.md +138 -0
- package/docs/sciverse-api.md +125 -0
- package/domains/_neutral/copy/few_shots.json +21 -0
- package/domains/_neutral/copy/framing_lexicon.json +19 -0
- package/domains/_neutral/copy/module_labels.json +5 -0
- package/domains/_neutral/copy/module_labels_footer.json +102 -0
- package/domains/_neutral/copy/module_labels_modules.json +204 -0
- package/domains/_neutral/copy/module_labels_nav.json +126 -0
- package/domains/_neutral/copy/module_labels_summary.json +98 -0
- package/domains/_neutral/copy/module_labels_tables.json +164 -0
- package/domains/_neutral/copy/module_labels_v2.json +90 -0
- package/domains/_neutral/copy/risk_constructs.json +20 -0
- package/domains/_neutral/copy/section_titles.json +66 -0
- package/domains/_neutral/copy/terminology.json +11 -0
- package/domains/check_copy_packs.py +103 -0
- package/domains/education/copy/few_shots.json +22 -0
- package/domains/education/copy/framing_enums.json +167 -0
- package/domains/education/copy/framing_lexicon.json +166 -0
- package/domains/education/copy/module_labels.json +169 -0
- package/domains/education/copy/risk_constructs.json +48 -0
- package/domains/education/copy/section_titles.json +186 -0
- package/domains/education/copy/terminology.json +70 -0
- package/domains/education/manifest.json +1 -1
- package/domains/education/outcome_taxonomy.json +2 -2
- package/domains/manifest.json +1 -1
- package/domains/policy/copy/few_shots.json +22 -0
- package/domains/policy/copy/framing_enums.json +94 -0
- package/domains/policy/copy/framing_lexicon.json +174 -0
- package/domains/policy/copy/module_labels.json +168 -0
- package/domains/policy/copy/risk_constructs.json +33 -0
- package/domains/policy/copy/section_titles.json +186 -0
- package/domains/policy/copy/terminology.json +64 -0
- package/eduevidence_cli.py +10 -0
- package/engine/capabilities.py +57 -5
- package/engine/decision_policy.py +167 -0
- package/engine/evidence_graph.py +14 -10
- package/engine/gaps.py +42 -22
- package/engine/ids.py +2 -0
- package/engine/library.py +6 -2
- package/engine/library_builtin.py +7 -4
- package/engine/living.py +34 -4
- package/engine/migration.py +88 -3
- package/engine/orchestration.py +5 -5
- package/engine/paths.py +2 -0
- package/engine/pilot.py +34 -32
- package/engine/taxonomy.py +211 -0
- package/engine/tribunal.py +49 -43
- package/engine/versions.py +1 -1
- package/examples/ai-coding-assistant-evidence/EduEvidence_Report.html +1361 -147
- package/examples/ai-coding-assistant-evidence/artifact_manifest.json +3 -3
- package/examples/ai-coding-assistant-evidence/citation_check.json +1 -1
- package/examples/ai-coding-assistant-evidence/final_verdict.json +107 -0
- package/examples/ai-coding-assistant-evidence/gate_report.json +101 -0
- package/examples/ai-coding-assistant-evidence/report_spec.json +23 -12
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_academic.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_claude.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_datalab-dark.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_datalab.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_presentation.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_academic.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_claude.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_datalab-dark.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_datalab.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_presentation.html +1360 -146
- package/examples/ai-coding-assistant-evidence/result.json +13 -9
- package/examples/ai-coding-assistant-evidence/result.zh.json +45 -41
- package/examples/ai-coding-assistant-evidence/skeptic.json +72 -0
- package/examples/ai-coding-assistant-evidence/verdict.json +6 -2
- package/examples/spaced-retrieval-practice/EduEvidence_Report.html +2728 -0
- package/examples/spaced-retrieval-practice/applicability.json +14 -0
- package/examples/spaced-retrieval-practice/artifact_manifest.json +15 -0
- package/examples/spaced-retrieval-practice/claims.jsonl +3 -0
- package/examples/spaced-retrieval-practice/evidence.jsonl +6 -0
- package/examples/spaced-retrieval-practice/final_verdict.json +93 -0
- package/examples/spaced-retrieval-practice/frame.json +58 -0
- package/examples/spaced-retrieval-practice/gate_report.json +101 -0
- package/examples/spaced-retrieval-practice/methodology.json +78 -0
- package/examples/spaced-retrieval-practice/report.html +2522 -0
- package/examples/spaced-retrieval-practice/report_spec.json +212 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_academic.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_claude.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_datalab-dark.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_datalab.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_presentation.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_academic.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_claude.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_datalab-dark.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_datalab.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_presentation.html +2728 -0
- package/examples/spaced-retrieval-practice/result.json +942 -0
- package/examples/spaced-retrieval-practice/result.zh.json +942 -0
- package/examples/spaced-retrieval-practice/skeptic.json +70 -0
- package/examples/spaced-retrieval-practice/sources.jsonl +7 -0
- package/examples/spaced-retrieval-practice/verdict.json +93 -0
- package/examples/workplace-ai-assistant/EduEvidence_Report.html +2814 -0
- package/examples/workplace-ai-assistant/artifact_manifest.json +15 -0
- package/examples/workplace-ai-assistant/claims.jsonl +4 -4
- package/examples/workplace-ai-assistant/evidence.jsonl +4 -4
- package/examples/workplace-ai-assistant/evidence_graph.json +15 -15
- package/examples/workplace-ai-assistant/final_verdict.json +78 -0
- package/examples/workplace-ai-assistant/gate_report.json +101 -0
- package/examples/workplace-ai-assistant/report_spec.json +209 -40
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_academic.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_claude.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_datalab-dark.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_datalab.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_presentation.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/report_academic.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_claude.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_datalab-dark.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_datalab.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_presentation.html +2814 -0
- package/examples/workplace-ai-assistant/result.json +82 -20
- package/examples/workplace-ai-assistant/result.zh.json +82 -20
- package/examples/workplace-ai-assistant/skeptic.json +72 -0
- package/examples/workplace-ai-assistant/verdict.json +36 -10
- package/integrations/agent_mcp.py +2 -2
- package/integrations/jev/__init__.py +115 -0
- package/integrations/jev/approval.py +212 -0
- package/integrations/jev/cli.py +84 -0
- package/integrations/jev/config.py +112 -0
- package/integrations/jev/gateway.py +128 -0
- package/integrations/jev/modes.py +38 -0
- package/integrations/jev/tools_classify.py +88 -0
- package/integrations/jev/tools_extract.py +111 -0
- package/integrations/jev/tools_rerank.py +71 -0
- package/integrations/jev/tools_screen.py +87 -0
- package/integrations/jev/tools_verify.py +95 -0
- package/integrations/jev_mcp.py +22 -0
- package/integrations/semantic_decide.py +286 -0
- package/integrations/semdecide_cli.py +55 -0
- package/package.json +19 -2
- package/pyproject.toml +4 -3
- package/references/report-copy-style.md +107 -0
- package/references/retrieval-compliance.md +75 -0
- package/references/retrieval-protocol.md +20 -0
- package/retrieval/audit.py +27 -3
- package/retrieval/fetch.py +96 -0
- package/retrieval/sciverse.py +398 -0
- package/retrieval/search.py +47 -7
- package/schemas/applicability.schema.json +94 -0
- package/schemas/chart-spec.schema.json +10 -3
- package/schemas/evidence.schema.json +316 -43
- package/schemas/fetch-result.schema.json +2 -1
- package/schemas/report-result.schema.json +3 -3
- package/schemas/report-spec.schema.json +98 -100
- package/schemas/skeptic.schema.json +86 -0
- package/schemas/source.schema.json +21 -2
- package/schemas/v2/decision-snapshot.schema.json +20 -9
- package/schemas/v2/finding.schema.json +5 -1
- package/schemas/v2/intake.schema.json +191 -0
- package/schemas/v2/methodology-audit.schema.json +5 -1
- package/schemas/v2/outcome.schema.json +28 -5
- package/schemas/v2/study.schema.json +5 -1
- package/schemas/vNext/autoevolve-session.schema.json +34 -1
- package/schemas/vNext/eval-snapshot.schema.json +77 -1
- package/schemas/vNext/execution-plan.schema.json +50 -1
- package/schemas/vNext/gap-priority.schema.json +54 -1
- package/schemas/vNext/negative-search-record.schema.json +68 -1
- package/schemas/vNext/research-iteration.schema.json +87 -1
- package/schemas/vNext/research-strategy.schema.json +62 -1
- package/schemas/vNext/skill-experiment.schema.json +90 -1
- package/schemas/vNext/task-spec.schema.json +156 -1
- package/schemas/vNext/worker-result.schema.json +60 -1
- package/schemas/verdict.schema.json +164 -28
- package/scripts/build_evidence_library.py +15 -5
- package/scripts/build_report_variants.py +18 -2
- package/scripts/build_result.py +74 -9
- package/scripts/check_package_parity.py +85 -0
- package/scripts/check_protocol_alignment.py +375 -0
- package/scripts/check_versioned_schemas.py +254 -0
- package/scripts/claim_audit.py +13 -8
- package/scripts/compute_confidence.py +10 -0
- package/scripts/dashboard_server.py +13 -2
- package/scripts/did_regression.py +12 -2
- package/scripts/evidence_score.py +5 -2
- package/scripts/intake/__init__.py +31 -0
- package/scripts/intake/__main__.py +18 -0
- package/scripts/intake/background.py +78 -0
- package/scripts/intake/browser.py +79 -0
- package/scripts/intake/cli.py +57 -0
- package/scripts/intake/constants.py +57 -0
- package/scripts/intake/depth.py +53 -0
- package/scripts/intake/enhancements.py +106 -0
- package/scripts/intake/hooks.py +90 -0
- package/scripts/intake/prefs.py +76 -0
- package/scripts/intake/prompts.py +85 -0
- package/scripts/intake/session.py +152 -0
- package/scripts/lint_file_layers.py +126 -0
- package/scripts/orchestrator.py +187 -40
- package/scripts/pre_verdict_gate.py +241 -29
- package/scripts/quickstart.py +18 -2
- package/scripts/run_workspace.py +7 -1
- package/scripts/skill_lint.py +11 -1
- package/scripts/skill_payload.py +6 -3
- package/scripts/test_adversarial_empirical.py +96 -25
- package/scripts/validate_schema.py +31 -1
- package/skill/agents/evaluation-designer.md +20 -4
- package/skill/agents/evidence-analyst.md +19 -3
- package/skill/agents/evidence-judge.md +98 -8
- package/skill/agents/evidence-retriever.md +20 -3
- package/skill/agents/intervention-designer.md +20 -4
- package/skill/agents/method-reviewer.md +18 -2
- package/skill/agents/{education-planner.md → research-planner.md} +19 -3
- package/skill/agents/skeptic.md +18 -2
- package/skill/roles/registry.yaml +11 -11
- package/skill/sub-skills/aihot-trend-analysis/SKILL.md +28 -9
- package/skill/sub-skills/contradiction-analysis/SKILL.md +31 -11
- package/skill/sub-skills/data-analysis/SKILL.md +34 -15
- package/skill/sub-skills/ethics-review/SKILL.md +33 -10
- package/skill/sub-skills/evidence-extraction/SKILL.md +29 -11
- package/skill/sub-skills/evidence-review/SKILL.md +31 -12
- package/skill/sub-skills/gap-analysis/SKILL.md +31 -9
- package/skill/sub-skills/literature-review/SKILL.md +35 -14
- package/skill/sub-skills/methodology-audit/SKILL.md +29 -12
- package/skill/sub-skills/report-generation/SKILL.md +28 -0
- package/skill/sub-skills/research-planning/SKILL.md +41 -14
- package/skill/sub-skills/study-design/SKILL.md +30 -9
- package/skill/task-briefs/adjudicate.md +32 -7
- package/skill/task-briefs/applicability.md +37 -2
- package/skill/task-briefs/audit.md +32 -7
- package/skill/task-briefs/challenge.md +34 -5
- package/skill/task-briefs/evaluate.md +30 -5
- package/skill/task-briefs/extract.md +31 -8
- package/skill/task-briefs/frame.md +39 -10
- package/skill/task-briefs/intervene.md +32 -6
- package/skill/task-briefs/present.md +32 -8
- package/skill/task-briefs/projection.md +36 -2
- package/skill/task-briefs/retrieve.md +36 -6
- package/skill/workflows/decision-and-pilot.md +76 -1
- package/skill/workflows/evaluate-and-update.md +83 -0
- package/skill/workflows/evidence-review.md +104 -0
- package/skill/workflows/experimental-jev.md +170 -0
- package/skill/workflows/intake.md +120 -0
- package/visualization/eduevidence-report/scripts/build_figures.py +25 -3
- package/visualization/eduevidence-report/scripts/build_infographics.py +37 -15
- package/visualization/eduevidence-report/scripts/build_report.py +435 -575
- package/visualization/eduevidence-report/scripts/charts_data.py +2 -0
- package/visualization/eduevidence-report/scripts/lieflat_engine.py +349 -38
- package/visualization/eduevidence-report/scripts/report_copy_pack.py +296 -0
- package/visualization/eduevidence-report/scripts/report_copy_policy_guard.py +47 -0
- package/visualization/eduevidence-report/scripts/zh_labels.py +141 -1
- package/web/architecture.html +14885 -0
- package/web/studio/assets/index-B8tkF44Q.css +1 -0
- package/web/studio/index.html +2 -2
- package/scripts/build_esl_artifacts.py +0 -1921
- package/scripts/build_killer_demo.py +0 -295
- package/scripts/enrich_projects_human_and_lieflat.py +0 -315
- package/scripts/generate_new_projects.py +0 -686
- package/scripts/sync_killer_demo_report.py +0 -270
- package/web/studio/assets/index-CzXocaGv.css +0 -1
- /package/web/studio/assets/{index-pa7jD7n4.js → index-CQ6Keoyc.js} +0 -0
package/README.md
CHANGED
|
@@ -8,15 +8,28 @@
|
|
|
8
8
|
|
|
9
9
|
## EduEvidence Research Engine — Evidence Research & Decision Skill
|
|
10
10
|
|
|
11
|
-
> **From Research Questions to Evidence-Based Decisions.**
|
|
11
|
+
> **From Research Questions to Evidence-Based Decisions.** · Current release **6.3.0**
|
|
12
|
+
|
|
13
|
+
> **▶ Live demo:** [Landing](https://37chengshan.github.io/eduevidence/) · [Research Studio](https://37chengshan.github.io/eduevidence/studio/) · [Deep Research comparison](https://37chengshan.github.io/eduevidence/comparison.html)
|
|
12
14
|
|
|
13
15
|
EduEvidence is delivered as an **AI Agent Skill**; inside the Skill operates
|
|
14
16
|
the **EduEvidence Research Engine** — a persistent, auditable engine that
|
|
15
|
-
turns
|
|
17
|
+
turns a decision question into an evidence-grounded answer. It is **multi-domain**:
|
|
18
|
+
the domain registry (`domains/manifest.json`) ships **education** and **policy**
|
|
19
|
+
today, each declaring its own frame schema, outcome taxonomy and methodology
|
|
20
|
+
checklist, so one nine-stage protocol serves education and applied social science
|
|
21
|
+
work without forking the engine.
|
|
16
22
|
|
|
17
23
|
- **Three public workflows** — **Evidence Review**, **Decision & Pilot**, and
|
|
18
24
|
**Evaluate & Update**. A full research cycle connects existing evidence,
|
|
19
25
|
grounded knowledge gaps, a study design, new data and a revised decision.
|
|
26
|
+
- **Multi-domain by contract** — `education` and `policy` are registered domains;
|
|
27
|
+
a run validates against its own domain's frame schema and outcome taxonomy
|
|
28
|
+
(`engine/taxonomy.py` is the single authority; unknown tokens fail closed).
|
|
29
|
+
- **Retrieval that stays traceable** — zero-config channels (OpenAlex / Semantic
|
|
30
|
+
Scholar / CrossRef / AIHot / AgentSearch / DuckDuckGo) plus key-based channels:
|
|
31
|
+
**Sciverse** (citation-grade academic retrieval with full-text locators),
|
|
32
|
+
Tavily and Brave. A lookup snippet is a locator, never evidence.
|
|
20
33
|
- **Project Workspace + Evidence Graph** — long-lived Projects with versioned,
|
|
21
34
|
immutable graph revisions; `result.json`/HTML/Markdown are projections, not
|
|
22
35
|
fact stores.
|
|
@@ -32,13 +45,31 @@ turns research questions into evidence-grounded decisions across education and o
|
|
|
32
45
|
server/app; Native Core runs on Python stdlib only and never requires
|
|
33
46
|
Agent MCP or a daemon.
|
|
34
47
|
|
|
35
|
-

|
|
49
|
+
|
|
50
|
+
*Recorded from the actual local Studio — no mockups: overview → report library → five report identities. Below, the introduction page walkthrough:*
|
|
36
51
|
|
|
37
|
-
|
|
52
|
+

|
|
38
53
|
|
|
39
54
|
---
|
|
40
55
|
|
|
41
|
-
## Quick
|
|
56
|
+
## Quick Start
|
|
57
|
+
|
|
58
|
+
**Fastest path — read a finished report (no install):**
|
|
59
|
+
|
|
60
|
+
```bash
|
|
61
|
+
open examples/ai-coding-assistant-evidence/EduEvidence_Report.html
|
|
62
|
+
open examples/spaced-retrieval-practice/EduEvidence_Report.html # real Sciverse run
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
**Look at the console (Python 3.10+, Node not required):**
|
|
66
|
+
|
|
67
|
+
```bash
|
|
68
|
+
python3 scripts/dashboard_server.py --host 127.0.0.1 --port 8765
|
|
69
|
+
# browser: http://127.0.0.1:8765/studio/ (read-only research console)
|
|
70
|
+
```
|
|
71
|
+
|
|
72
|
+
**Install it:**
|
|
42
73
|
|
|
43
74
|
**npm (recommended for Skill install)**
|
|
44
75
|
|
|
@@ -135,7 +166,7 @@ runnable + sample report renderable).
|
|
|
135
166
|
|
|
136
167
|
## What Problem We Solve
|
|
137
168
|
|
|
138
|
-
A typical AI answers
|
|
169
|
+
A typical AI answers a decision question like this:
|
|
139
170
|
|
|
140
171
|
```text
|
|
141
172
|
Question → Search a few sources → Summarize opinions → Give advice
|
|
@@ -144,15 +175,15 @@ Question → Search a few sources → Summarize opinions → Give advice
|
|
|
144
175
|
EduEvidence does this instead:
|
|
145
176
|
|
|
146
177
|
```text
|
|
147
|
-
|
|
148
|
-
→
|
|
178
|
+
Decision question (education or applied social science)
|
|
179
|
+
→ Domain Research Framing (learner or decision object / intervention / comparison / outcomes / context)
|
|
149
180
|
→ Literature & evidence retrieval (supporting evidence + independent counter-evidence)
|
|
150
181
|
→ Claim-Level Evidence Extraction
|
|
151
182
|
→ Skeptic challenge protocol + Method Reviewer audit
|
|
152
183
|
→ Evidence Tribunal
|
|
153
184
|
→ Applicability Analysis
|
|
154
185
|
→ Decision: ADOPT / PILOT / REJECT / INSUFFICIENT EVIDENCE
|
|
155
|
-
→
|
|
186
|
+
→ Intervention (minimum viable pilot)
|
|
156
187
|
→ Evaluation Plan
|
|
157
188
|
```
|
|
158
189
|
|
|
@@ -161,43 +192,44 @@ It answers six questions:
|
|
|
161
192
|
1. What does the current evidence actually support?
|
|
162
193
|
2. What can the current evidence not support?
|
|
163
194
|
3. Why do different studies reach different results?
|
|
164
|
-
4. Which
|
|
165
|
-
5. If an institution adopts it, how
|
|
195
|
+
4. Which population, in which setting, under which conditions does it apply to?
|
|
196
|
+
5. If an institution adopts it, how should it be rolled out with low risk?
|
|
166
197
|
6. How to verify whether it actually works after implementation?
|
|
167
198
|
|
|
168
|
-
## 30-second
|
|
199
|
+
## 30-second tour
|
|
169
200
|
|
|
170
|
-
>
|
|
201
|
+
> Flagship question: **Should first-year C programming students be allowed to use generative AI coding assistants?**
|
|
171
202
|
|
|
172
203
|
| Time | Stage |
|
|
173
204
|
|---|---|
|
|
174
|
-
| 0–20s | Ask the
|
|
175
|
-
| 20–45s |
|
|
205
|
+
| 0–20s | Ask the decision question |
|
|
206
|
+
| 20–45s | Research Frame (domain-specific schema) |
|
|
176
207
|
| 45–75s | Evidence Retrieval |
|
|
177
208
|
| 75–110s | Evidence Matrix |
|
|
178
209
|
| 110–135s | Methodology + Skeptic |
|
|
179
210
|
| 135–155s | Evidence Tribunal |
|
|
180
|
-
| 155–170s |
|
|
211
|
+
| 155–170s | Intervention + Evaluation |
|
|
181
212
|
| 170–180s | Benchmark |
|
|
182
213
|
|
|
183
214
|
Full example pack: [`examples/ai-coding-assistant-evidence/`](examples/ai-coding-assistant-evidence/).
|
|
184
215
|
|
|
185
|
-
## Why
|
|
216
|
+
## Why Evidence Decisions Are Hard
|
|
186
217
|
|
|
187
|
-
|
|
218
|
+
Evidence across education and applied social science shares the same natural pitfalls. EduEvidence's core contribution is standardizing the countermeasures:
|
|
188
219
|
|
|
189
220
|
- **Outcome Separation**: `faster task completion ≠ actually learning to program`; `short-term score gains ≠ long-term retention`; `completing tasks with AI ≠ transferring skills without AI`.
|
|
190
221
|
- **Counter-Evidence Search**: it does not just verify the user's initial assumption — it independently searches for null / negative / contradictory evidence, AI dependency, novelty effects, self-selection bias, and more.
|
|
191
222
|
- **Evidence Tribunal**: instead of listing pros and cons, it judges which studies are more credible, whether conflicts come from samples / measurement / course / tool / design, and what can be concluded so far.
|
|
192
|
-
- **Evidence-to-Action Bridge**: it does not stop at "research shows…" — it connects to applicability, the
|
|
223
|
+
- **Evidence-to-Action Bridge**: it does not stop at "research shows…" — it connects to applicability, the decision, pilot intervention, and evaluation design.
|
|
193
224
|
|
|
194
225
|
## How EduEvidence Works
|
|
195
226
|
|
|
196
227
|
```text
|
|
197
228
|
┌─────────────────────────────────────┐
|
|
198
229
|
│ EduEvidence │
|
|
199
|
-
│
|
|
200
|
-
│
|
|
230
|
+
│ domain contracts (education / │
|
|
231
|
+
│ policy) + decision + intervention │
|
|
232
|
+
│ + evaluation │
|
|
201
233
|
└────────────────┬────────────────────┘
|
|
202
234
|
│
|
|
203
235
|
┌────────────────▼────────────────────┐
|
|
@@ -215,22 +247,24 @@ Education evidence has natural pitfalls. EduEvidence's core contribution is stan
|
|
|
215
247
|
The 9-step workflow:
|
|
216
248
|
|
|
217
249
|
```text
|
|
218
|
-
1. Frame Build the
|
|
250
|
+
1. Frame Build the domain frame (education frame / policy frame)
|
|
219
251
|
2. Retrieve Retrieve literature & evidence (support + independent counter-evidence)
|
|
220
252
|
3. Extract Extract claim-level evidence (bound to outcomes)
|
|
221
253
|
4. Challenge Skeptic protocol (fixed 9 checks)
|
|
222
254
|
5. Audit Method Reviewer audit (15-item checklist)
|
|
223
255
|
6. Adjudicate Evidence Tribunal (Evidence Matrix + Verdict)
|
|
224
256
|
7. Applicability Applicability analysis
|
|
225
|
-
8. Intervene
|
|
257
|
+
8. Intervene Intervention design (minimum viable pilot)
|
|
226
258
|
9. Evaluate Evaluation Plan design
|
|
227
259
|
```
|
|
228
260
|
|
|
229
|
-
Every step is validated against JSON Schemas (`schemas/`), deterministic logic lives in `scripts/`, and the
|
|
261
|
+
Every step is validated against JSON Schemas (`schemas/`), deterministic logic lives in `scripts/`, and the methodology is documented independently in `references/` (21 documents: evidence quality, skeptic protocol, tribunal policy, WWC/GRADE standards, social-science pitfalls, retrieval protocol, report copy style …; counts in `docs/metrics.json`).
|
|
230
262
|
|
|
231
263
|
## Outcome Separation
|
|
232
264
|
|
|
233
|
-
|
|
265
|
+
Outcome tokens are domain-owned. The education taxonomy declares **20 tokens** in four categories (`domains/education/outcome_taxonomy.json`); the policy domain declares its own categories and tokens (`domains/policy/outcome_taxonomy.json`). `engine/taxonomy.py` is the only reader: an unknown token or unregistered domain **fails closed** instead of being silently classified as a learning outcome.
|
|
266
|
+
|
|
267
|
+
The education set (`references/outcome-taxonomy.md`):
|
|
234
268
|
|
|
235
269
|
```text
|
|
236
270
|
Learning: Knowledge Gain / Concept Understanding / Retention / Transfer / Independent Problem Solving
|
|
@@ -239,11 +273,11 @@ Process: Engagement / Motivation / Cognitive Load / Help-Seeking / Metacogni
|
|
|
239
273
|
Risk: AI Dependency / Over-reliance / Reduced Effort / Reduced Transfer / Academic Integrity Risk / False Confidence
|
|
240
274
|
```
|
|
241
275
|
|
|
242
|
-
The demo's highlight: in Kazemitabaar et al. (CHI 2023), the AI code assistant raised task completion by 1.15× and correctness by 1.8×, but the one-week retention test showed no significant difference — **task performance ≠ learning**.
|
|
276
|
+
The flagship demo's highlight: in Kazemitabaar et al. (CHI 2023), the AI code assistant raised task completion by 1.15× and correctness by 1.8×, but the one-week retention test showed no significant difference — **task performance ≠ learning**.
|
|
243
277
|
|
|
244
278
|
## Evidence Tribunal
|
|
245
279
|
|
|
246
|
-
`references/tribunal-policy.md` defines the adjudication rules: input = Frame + Evidence Matrix + Skeptic Findings + Method Reviews; output =
|
|
280
|
+
`references/tribunal-policy.md` defines the adjudication rules: input = Frame + Evidence Matrix + Skeptic Findings + Method Reviews; output = the domain Verdict (`schemas/verdict.schema.json`), including:
|
|
247
281
|
|
|
248
282
|
- supported / uncertain / contradicted claims
|
|
249
283
|
- conflict-source analysis (sample / measurement / course / tool / design)
|
|
@@ -254,15 +288,15 @@ The demo's highlight: in Kazemitabaar et al. (CHI 2023), the AI code assistant r
|
|
|
254
288
|
|
|
255
289
|
## From Evidence to Action
|
|
256
290
|
|
|
257
|
-
Evidence must connect to the real classroom (`references/applicability-policy.md`, `intervention-design.md`, `evaluation-design.md`):
|
|
291
|
+
Evidence must connect to the real setting — a classroom, a support team, a policy roll-out (`references/applicability-policy.md`, `intervention-design.md`, `evaluation-design.md`):
|
|
258
292
|
|
|
259
|
-
- **Applicability**: For whom?
|
|
260
|
-
- **Intervention**: always a "minimum viable pilot", never direct full deployment; includes AI usage rules,
|
|
293
|
+
- **Applicability**: For whom? In which setting? For which outcome? Under what conditions? With what AI usage policy?
|
|
294
|
+
- **Intervention**: always a "minimum viable pilot", never direct full deployment; includes AI usage rules, staff/user roles, reflection requirements, and stop conditions.
|
|
261
295
|
- **Evaluation**: every PILOT/ADOPT recommendation must come with an evaluation plan; distinguishes baseline / post-test / retention / transfer, and task-performance vs learning metrics.
|
|
262
296
|
|
|
263
297
|
## Benchmark
|
|
264
298
|
|
|
265
|
-
30 education research questions in v1 (`benchmarks/questions.jsonl`), S×10 / M×10 / L×10; 15 in the core domain "AI-assisted university teaching",
|
|
299
|
+
30 education research questions in v1 (`benchmarks/questions.jsonl`), S×10 / M×10 / L×10; 15 in the core domain "AI-assisted university teaching", all 30 with human gold annotations (`benchmarks/annotations/gold-Q01..Q30`).
|
|
266
300
|
|
|
267
301
|
Baseline design:
|
|
268
302
|
|
|
@@ -285,20 +319,22 @@ Key metrics: Citation Support Precision / Unsupported Claim Rate / Contradiction
|
|
|
285
319
|
`examples/ai-coding-assistant-evidence/` shows the full path from question to decision:
|
|
286
320
|
|
|
287
321
|
- **Evidence** (12 findings from 8 sources): task-performance gains (Kazemitabaar 2023), unguarded access harming independent exam performance by −17% (Bastani 2025, PNAS), guardrails eliminating the negative effect (Bastani 2025), formative-feedback writing evidence (Marzuki 2024).
|
|
288
|
-
- **Decision**:
|
|
322
|
+
- **Decision**: recorded `PILOT` (Moderate, 0.586); matrix landing (`engine.decision_policy.decision_outcome`) is `INSUFFICIENT_EVIDENCE` (`unresolved_conflict`) — the decisive relations contain both `support_adoption` and `oppose_adoption`, which fails the conflict check before any `ADOPT`/`PILOT` claim; task-performance evidence is strong, but direct learning-effect evidence for university programming courses is missing and the unguarded-access risk is documented.
|
|
289
323
|
- **Intervention**: 4-phase pilot (Independent Foundation → Explain Don't Solve → Structured Collaboration → Transfer Check).
|
|
290
324
|
- **Evaluation**: no-AI baseline / post-test / final-exam retention / no-AI transfer task + AI-dependency risk metrics.
|
|
291
325
|
|
|
292
|
-
A second public example, `examples/workplace-ai-assistant/`, evaluates AI assistance in enterprise customer support using the policy domain: 4 findings from 3 studies, with direct and indirect evidence distinguished. Its proposed supervised pilot has not been executed.
|
|
326
|
+
A second public example, `examples/workplace-ai-assistant/`, evaluates AI assistance in enterprise customer support using the policy domain: 4 findings from 3 studies, with direct and indirect evidence distinguished. Its recorded action is `PILOT` (Moderate, 0.578); the matrix landing is `INSUFFICIENT_EVIDENCE` with `downgrade_reason=None`: all four decisive relations are `conditional`, and `conditional` is not a conflict relation (`engine.decision_policy.CONFLICT_RELATIONS` = `{conflict, mixed}`), so no downgrade reason is recorded and — with no decisive `support_adoption` relation — Moderate confidence cannot reach `PILOT`. Its proposed supervised pilot has not been executed.
|
|
293
327
|
|
|
294
|
-
|
|
328
|
+
The third public example, `examples/spaced-retrieval-practice/`, asks whether spaced repetition and retrieval practice should replace massed review in an introductory programming course. It is the first pack whose sources were located through the **Sciverse** channel (`discovery_provider=sciverse`, `fetch_provider=sciverse_content`) and whose `meta.data_origin` is `real_run_sciverse`: 6 findings from 7 tier-1 DOI sources, decision **ADOPT** (High confidence, 0.893) — the only public case where the recorded action and the matrix landing agree. It is the worked example that the ADOPT path is reachable: retention and transfer - the two primary outcomes - carry direct, consistent evidence at directness 2. The coding and workplace cases are bounded below ADOPT: the coding case's decisive relations mix `support_adoption` with `oppose_adoption` (`unresolved_conflict`), and the workplace case carries no decisive `support_adoption` relation at all (all four are `conditional`), so both land at `INSUFFICIENT_EVIDENCE`.
|
|
295
329
|
|
|
296
|
-
|
|
330
|
+
Each pack ships `result.json` + `result.zh.json` (bilingual parallel data), a packaged-`EduEvidence_Report.html` root report, and `reports-5themes/` with the five standalone theme HTML files.
|
|
331
|
+
|
|
332
|
+
All three public examples are literature demonstrations, and their `data_origin` says exactly what produced them. The coding and workplace cases are **manually curated** (`manual_curated`); the spaced-retrieval case is a recorded **Sciverse-backed run** (`real_run_sciverse`). A rendered report never establishes that an agent completed the whole nine-stage research workflow. See [the workplace evidence notes](docs/demo-workplace-ai.md) for source versions and limitations, and [`docs/reproducibility.md`](docs/reproducibility.md) for how `data_origin` is declared.
|
|
297
333
|
|
|
298
334
|
### Start your own research in ~30 minutes
|
|
299
335
|
|
|
300
336
|
```bash
|
|
301
|
-
python3 scripts/quickstart.py "
|
|
337
|
+
python3 scripts/quickstart.py "你的研究问题" # creates runs/<id> + NEXT_STEPS.md
|
|
302
338
|
# hand the LLM stages to your AI agent per NEXT_STEPS.md, then finish with:
|
|
303
339
|
python3 scripts/orchestrator.py adjudicate --project runs/<id>
|
|
304
340
|
bash scripts/bake_pack.sh <pack_dir> # 5-theme bilingual report
|
|
@@ -321,10 +357,10 @@ After research completes, `result.json` is rendered into three visualization out
|
|
|
321
357
|
|
|
322
358
|
```text
|
|
323
359
|
result.json + result.zh.json (Chinese parallel data)
|
|
324
|
-
├─ build_charts.py → chart_specs.json (ECharts option data; no ECharts runtime bundled)
|
|
325
|
-
├─ build_infographics.py → infographics.json (hand-authored SVGs)
|
|
326
|
-
├─ build_figures.py → figures/ (publication figures: figure_data.json + SVG/PNG/PDF)
|
|
327
|
-
└─ build_report.py → EduEvidence_Report.html (single-file bilingual report + report_spec.json)
|
|
360
|
+
├─ visualization/eduevidence-report/scripts/build_charts.py → chart_specs.json (ECharts option data; no ECharts runtime bundled)
|
|
361
|
+
├─ visualization/eduevidence-report/scripts/build_infographics.py → infographics.json (hand-authored SVGs)
|
|
362
|
+
├─ visualization/eduevidence-report/scripts/build_figures.py → figures/ (publication figures: figure_data.json + SVG/PNG/PDF)
|
|
363
|
+
└─ visualization/eduevidence-report/scripts/build_report.py → EduEvidence_Report.html (single-file bilingual report + report_spec.json)
|
|
328
364
|
```
|
|
329
365
|
|
|
330
366
|
**EduEvidence_Report.html (main deliverable)**:
|
|
@@ -344,23 +380,39 @@ See [Research Studio workflow and delivery guide](docs/research-studio-guide.zh-
|
|
|
344
380
|
|
|
345
381
|
> Open the example directly: `examples/ai-coding-assistant-evidence/EduEvidence_Report.html`
|
|
346
382
|
|
|
383
|
+
|
|
384
|
+
### Optional key-based retrieval channels
|
|
385
|
+
|
|
386
|
+
Zero-config retrieval (OpenAlex / Semantic Scholar / CrossRef / AIHot / AgentSearch) works out of the box. These channels activate once a key is present and stay silently inactive otherwise — the scientific gates never depend on them:
|
|
387
|
+
|
|
388
|
+
```bash
|
|
389
|
+
export SCIVERSE_API_TOKEN=sv-... # citation-grade academic retrieval + full-text location
|
|
390
|
+
export TAVILY_API_KEY=... # general web search
|
|
391
|
+
export BRAVE_API_KEY=... # general web search
|
|
392
|
+
```
|
|
393
|
+
|
|
394
|
+
The Sciverse channel treats an `/agentic-search` chunk as a **locator**: it must be expanded through `/content` and pass the validation gate before it may enter evidence extraction (RULE 2, machine-enforced). Contract: `docs/sciverse-api.md`; compliance: `references/retrieval-compliance.md`.
|
|
395
|
+
|
|
347
396
|
## Architecture
|
|
348
397
|
|
|
349
398
|
The repository is a complete **Skill package**: `SKILL.md` is the entry point; everything else is layered as *skill core → quality assurance → demos*. See [`docs/architecture.md`](docs/architecture.md):
|
|
350
399
|
|
|
400
|
+
Read the illustrated single-file walkthrough of the same architecture (nine-step protocol, roles and independence, artifact/state map, execution and approval loop) at [`web/architecture.html`](web/architecture.html).
|
|
401
|
+
|
|
351
402
|
```text
|
|
352
403
|
EduEvidence/ (= one Skill package)
|
|
353
404
|
│
|
|
354
|
-
├─ SKILL.md ← Skill entry:
|
|
405
|
+
├─ SKILL.md ← Skill entry: mission → workflow routing → scientific gates → 9-step protocol → output requirements
|
|
355
406
|
│
|
|
356
407
|
├─ Skill core (required to run)
|
|
357
408
|
│ ├─ engine/ V2 Research Engine (Project Workspace / immutable Evidence
|
|
358
409
|
│ │ Graph / Library / synthesis / tribunal / study design /
|
|
359
410
|
│ │ datasets / analysis / projections / migration)
|
|
360
411
|
│ ├─ skill/agents/ role protocols (capability execution profiles)
|
|
361
|
-
│ ├─ references/
|
|
362
|
-
│ │ tribunal policy / intervention design
|
|
363
|
-
│ ├─ schemas/
|
|
412
|
+
│ ├─ references/ 21 methodology documents (evidence quality / skeptic /
|
|
413
|
+
│ │ tribunal policy / intervention design…; counted in docs/metrics.json)
|
|
414
|
+
│ ├─ schemas/ 50 versioned JSON Schema contracts (15 top-level + 18 v2 +
|
|
415
|
+
│ │ 3 v3 + 4 v4 + 10 vNext; counted by `scripts/generate_metrics.py`)
|
|
364
416
|
│ ├─ domains/ v4 domain registry (manifest.json) + per-domain packages
|
|
365
417
|
│ │ (education: registration-only; policy: frame schema /
|
|
366
418
|
│ │ outcome taxonomy / methodology checklist / references)
|
|
@@ -380,11 +432,18 @@ EduEvidence/ (= one Skill package)
|
|
|
380
432
|
├─ docs/ architecture / methodology / benchmark / demo / reproducibility
|
|
381
433
|
├─ install.sh one-click install (local / multi-agent Skill) + self-check
|
|
382
434
|
├─ pyproject.toml packaging metadata (wheel ships CLI, engine and installed runtime resources; stdlib-only core)
|
|
383
|
-
└─ README
|
|
435
|
+
└─ README.md / README.zh-CN.md bilingual docs
|
|
384
436
|
```
|
|
385
437
|
|
|
386
438
|
> Skill-package principle: the **minimal runtime set is `SKILL.md + engine/ + skill/ + references/ + schemas/ + scripts/`**; `retrieval/`, `integrations/`, `visualization/` are the execution/presentation layers that make the Skill actually runnable; `tests/`, `benchmarks/`, `examples/`, `docs/` provide credibility and onboarding — none of them affect the Skill body itself.
|
|
387
439
|
|
|
440
|
+
### Distribution
|
|
441
|
+
|
|
442
|
+
Two release channels, built from one runtime allowlist (`scripts/skill_payload.py`, shared with the host installer):
|
|
443
|
+
|
|
444
|
+
- **Submission folder** — `bash packaging/make_upload.sh` rebuilds `dist/eduevidence-submission/` (flat Skill package + `submission-manifest.json` + `UPLOAD-README.md`); it regenerates the report variants first and aborts without a prebuilt `web/studio/index.html`. The older `upload/` tree is a **frozen deprecated snapshot** that must not be shipped — see [`upload/DEPRECATED.md`](upload/DEPRECATED.md).
|
|
445
|
+
- **npm package** — `npm install -g eduevidence`, then `eduevidence skill --list-hosts` (every supported agent and its install path) or `eduevidence skill --dry-run` (preview; writes nothing). From a checkout the same flags run as `node bin/eduevidence.js skill --list-hosts` / `--dry-run`.
|
|
446
|
+
|
|
388
447
|
### v4 Domain Registry (EvidenceCore)
|
|
389
448
|
|
|
390
449
|
The first step of the **EvidenceCore** abstraction: a `domains/` registry
|
|
@@ -394,8 +453,9 @@ The first step of the **EvidenceCore** abstraction: a `domains/` registry
|
|
|
394
453
|
|
|
395
454
|
- **education** is a **registration-only domain**: it points at the existing
|
|
396
455
|
contracts — `schemas/education-frame.schema.json` (frame),
|
|
397
|
-
the 20-token outcome taxonomy from
|
|
398
|
-
|
|
456
|
+
the 20-token education outcome taxonomy from
|
|
457
|
+
`domains/education/outcome_taxonomy.json` (`engine/taxonomy.py` is its sole
|
|
458
|
+
reader; 25 tokens across education + policy), the 15-item
|
|
399
459
|
methodology checklist from `skill/agents/method-reviewer.md`,
|
|
400
460
|
`benchmarks/annotations` (golds) and `references/`. It **adds no new logic
|
|
401
461
|
path** — no new schema, no new validator, no new methodology.
|
|
@@ -448,7 +508,9 @@ python3 scripts/evidence_score.py examples/ai-coding-assistant-evidence/evidence
|
|
|
448
508
|
python3 scripts/evidence_matrix.py examples/ai-coding-assistant-evidence/evidence.jsonl
|
|
449
509
|
|
|
450
510
|
# 4. Run the Citation Audit (claim-evidence traceability)
|
|
451
|
-
python3 scripts/claim_audit.py
|
|
511
|
+
python3 scripts/claim_audit.py \
|
|
512
|
+
--claims examples/ai-coding-assistant-evidence/claims.jsonl \
|
|
513
|
+
--evidence examples/ai-coding-assistant-evidence/evidence.jsonl
|
|
452
514
|
|
|
453
515
|
# 5. Render the Research & Decision Pack (Markdown)
|
|
454
516
|
python3 scripts/render_report.py \
|
|
@@ -466,6 +528,8 @@ python3 visualization/eduevidence-report/scripts/build_report.py \
|
|
|
466
528
|
--out examples/ai-coding-assistant-evidence/EduEvidence_Report.html
|
|
467
529
|
|
|
468
530
|
# 7. Validate the benchmark question set
|
|
531
|
+
# Source checkout only: benchmarks/questions.jsonl is not part of the
|
|
532
|
+
# shipped Skill package (see packaging/upload-layout.md).
|
|
469
533
|
python3 scripts/benchmark.py --questions benchmarks/questions.jsonl
|
|
470
534
|
|
|
471
535
|
# 8. Run the tests
|
|
@@ -495,7 +559,7 @@ pytest
|
|
|
495
559
|
simulation that proves the evaluation framework runs; it is **not** real model
|
|
496
560
|
performance (see the ⚠️ note in [Benchmark](#benchmark)).
|
|
497
561
|
- [x] Skill core & pipeline: 9-step protocol (Research Core 6 + Decision Extension 3),
|
|
498
|
-
|
|
562
|
+
50 versioned JSON Schemas (`schemas/**`, counted by `scripts/generate_metrics.py`), deterministic scripts, 8-role protocols (original Phases 0–6).
|
|
499
563
|
- [x] Evidence-to-action: applicability / four-state decision / intervention / evaluation.
|
|
500
564
|
- [x] Product UI: single-file bilingual HTML report + infographics + academic figures
|
|
501
565
|
(original Phase 8).
|
package/README.zh-CN.md
CHANGED
|
@@ -11,6 +11,8 @@
|
|
|
11
11
|
> **From Research Questions to Evidence-Based Decisions.**
|
|
12
12
|
> **从研究问题,到有证据支撑的决策。**
|
|
13
13
|
|
|
14
|
+
> **▶ 在线演示:** [介绍页](https://37chengshan.github.io/eduevidence/) · [Research Studio](https://37chengshan.github.io/eduevidence/studio/) · [深度调研对比页](https://37chengshan.github.io/eduevidence/comparison.html)
|
|
15
|
+
|
|
14
16
|
EduEvidence 面向研究者与实践决策者,将教育、组织政策和 AI 工具采用等问题转化为**可追溯、可质疑、可验证的证据决策流程**。当前公开案例涵盖编程学习和企业客服,分别使用教育与组织政策领域契约。
|
|
15
17
|
|
|
16
18
|
- **三条公开工作流**:Evidence Review(证据综述)、Decision & Pilot(决策与试点)、Evaluate & Update(评估与更新)。完整研究周期将文献证据、有依据的知识缺口、研究设计、新数据和决策修订连接起来。
|
|
@@ -18,9 +20,11 @@ EduEvidence 面向研究者与实践决策者,将教育、组织政策和 AI
|
|
|
18
20
|
- 🧪 基于真实研究(示例包含 CHI 2023 / PNAS 2025 / ACL 2025 / Springer 2024 的实证证据),不做无来源断言。
|
|
19
21
|
- 🚦 最终输出不是"允许/禁止"的二元结论,而是 **ADOPT / PILOT / REJECT / INSUFFICIENT EVIDENCE** 四态决策 + 可落地的干预与评价方案。
|
|
20
22
|
|
|
21
|
-

|
|
24
|
+
|
|
25
|
+
*本地 Studio 真实录屏(非示意):总览 → 报告阅读室 → 五种报告形态。下方为介绍页滚动实录:*
|
|
22
26
|
|
|
23
|
-
|
|
27
|
+

|
|
24
28
|
|
|
25
29
|
---
|
|
26
30
|
|
|
@@ -280,13 +284,15 @@ B4 EduEvidence + Agent MCP ← 证明多 Agent 增强价值(B3 vs B4)
|
|
|
280
284
|
`examples/ai-coding-assistant-evidence/` 完整展示了从问题到决策的全过程:
|
|
281
285
|
|
|
282
286
|
- **证据**(12 条发现、8 个来源):任务表现提升(Kazemitabaar 2023)、无护栏访问损害独立考试表现 -17%(Bastani 2025, PNAS)、护栏设计消除负效应(Bastani 2025)、形成性反馈写作证据(Marzuki 2024)。
|
|
283
|
-
-
|
|
287
|
+
- **决策**:落档动作 `PILOT`(Moderate,0.586);矩阵落点(`engine.decision_policy.decision_outcome`)为 `INSUFFICIENT_EVIDENCE`(`unresolved_conflict`)—— decisive relations 中同时存在 `support_adoption` 与 `oppose_adoption`,冲突检查在 `ADOPT`/`PILOT` 判定前即短路;任务表现证据强,但大学编程课程的直接学习效应证据缺失,且无护栏风险已被证实。
|
|
284
288
|
- **干预**:4 阶段试点(Independent Foundation → Explain Don't Solve → Structured Collaboration → Transfer Check)。
|
|
285
289
|
- **评价**:无 AI 基线/后测/期末考试保持/无 AI 迁移任务 + AI 依赖风险指标。
|
|
286
290
|
|
|
287
|
-
另一公开案例 `examples/workplace-ai-assistant/` 使用组织政策领域,讨论企业客服是否引入 AI 助手:3 项研究、4
|
|
291
|
+
另一公开案例 `examples/workplace-ai-assistant/` 使用组织政策领域,讨论企业客服是否引入 AI 助手:3 项研究、4 条发现,区分直接客服证据与间接写作/咨询证据,落档动作 `PILOT`(Moderate,0.578),矩阵落点为 `INSUFFICIENT_EVIDENCE` 且 `downgrade_reason=None`:4 条 decisive relation 全为 `conditional`,而 `conditional` 不属于冲突关系(`engine.decision_policy.CONFLICT_RELATIONS` = `{conflict, mixed}`),因此引擎不记录降级理由;又因没有任何 decisive 的 `support_adoption`,Moderate 置信度无法升到 `PILOT`。详见 [来源核验与边界](docs/demo-workplace-ai.md)。
|
|
288
292
|
|
|
289
|
-
|
|
293
|
+
第三个公开案例 `examples/spaced-retrieval-practice/`(来源经 Sciverse 通道逐条读回原文)讨论间隔重复与检索练习能否替代集中式复习:6 条证据、7 篇 tier-1 DOI 来源,判定 **ADOPT**(High,引擎复算 0.893)—— 唯一一个落档动作与矩阵落点一致的公开案例。它是“ADOPT 出口真实可达”的实证:延迟保持与迁移两个主要结果上都有 directness=2 的直接且一致证据;另两例低于 ADOPT:编程案中 decisive relations 同时含 `support_adoption` 与 `oppose_adoption`(`unresolved_conflict`),企业客服案中没有任何 decisive 的 `support_adoption`(4 条全为 `conditional`),因此矩阵落点均为 `INSUFFICIENT_EVIDENCE`,不得表述为已落在 PILOT。
|
|
294
|
+
|
|
295
|
+
三个公开案例的数据来源各自如实标注:编程与企业客服两例为人工整理文献(`manual_curated`),间隔重复一例为真实 Sciverse 检索运行记录(`real_run_sciverse`),报告生成不等于九阶段模型研究已运行,也不代表试点已经执行。四个旧教学示例迁入 `tests/fixtures/legacy-examples/`,仅供软件兼容测试,排除于公共目录和分发包;未核验或合成数据不能引用为研究证据。旧 `ai-coding-assistant` 路径保留兼容别名。
|
|
290
296
|
|
|
291
297
|
## Studio 实际界面
|
|
292
298
|
|
|
@@ -327,27 +333,41 @@ result.json + result.zh.json
|
|
|
327
333
|
|
|
328
334
|
> Open the example directly: `examples/ai-coding-assistant-evidence/EduEvidence_Report.html`
|
|
329
335
|
|
|
336
|
+
### 可选检索通道(key-based)
|
|
337
|
+
|
|
338
|
+
零配置检索(OpenAlex / Semantic Scholar / CrossRef / AIHot / AgentSearch)开箱可用。配置以下 key 后通道自动启用,未配置时静默失活、不影响科学门:
|
|
339
|
+
|
|
340
|
+
```bash
|
|
341
|
+
export SCIVERSE_API_TOKEN=sv-... # 引用级学术检索 + 全文定位(meta-search / agentic-search / content / paper-relations)
|
|
342
|
+
export TAVILY_API_KEY=... # 通用网页检索
|
|
343
|
+
export BRAVE_API_KEY=... # 通用网页检索
|
|
344
|
+
```
|
|
345
|
+
|
|
346
|
+
Sciverse 通道把 `/agentic-search` 的 chunk 当作**定位子**:必须经 `/content` 读原文并通过校验门后,才允许进入证据抽取(RULE 2 的机器化执行)。契约见 `docs/sciverse-api.md`,合规见 `references/retrieval-compliance.md`。
|
|
347
|
+
|
|
330
348
|
## Architecture
|
|
331
349
|
|
|
332
350
|
仓库是一个完整的 **Skill 包**:`SKILL.md` 是入口,其余目录按"Skill 运行必需 → 质量保障 → 演示"分层。详见 [`docs/architecture.md`](docs/architecture.md):
|
|
333
351
|
|
|
352
|
+
同一套架构的图解单页(九步协议 / 角色与独立性 / 产物状态地图 / 执行与审批闭环)见 [`web/architecture.html`](web/architecture.html)。
|
|
353
|
+
|
|
334
354
|
```text
|
|
335
355
|
EduEvidence/ (= 一个 Skill 包)
|
|
336
356
|
│
|
|
337
|
-
├─ SKILL.md ← Skill
|
|
357
|
+
├─ SKILL.md ← Skill 入口:使命 → 工作流路由 → 科学门禁 → 9 步协议 → 输出要求
|
|
338
358
|
│
|
|
339
359
|
├─ Skill 本体(运行必需)
|
|
340
360
|
│ ├─ skill/agents/ 8 个角色协议(Planner / Retriever / Analyst / Skeptic /
|
|
341
361
|
│ │ Method Reviewer / Judge / Intervention Designer / Evaluation Designer)
|
|
342
|
-
│ ├─ references/
|
|
343
|
-
│ ├─ schemas/
|
|
344
|
-
│ ├─ scripts/
|
|
362
|
+
│ ├─ references/ 方法论文档(证据质量 / 反证协议 / 裁决规则 / 干预设计 / 检索合规 / 文案规范…;数量见 docs/metrics.json)
|
|
363
|
+
│ ├─ schemas/ 50 个 JSON Schema 数据契约(15 顶层 + 18 v2 + 3 v3 + 4 v4 + 10 vNext,每步输出的校验门;递归计数见 `scripts/generate_metrics.py`)
|
|
364
|
+
│ ├─ scripts/ 确定性逻辑脚本(评分 / 矩阵 / 审计 / 置信度 / Orchestrator / 启动探测)
|
|
345
365
|
│ ├─ retrieval/ 检索与抓取层(fetch / validate / dedupe / failures)
|
|
346
366
|
│ ├─ integrations/ Agent MCP 增强层 + Smart Web Fetch 集成
|
|
347
367
|
│ └─ visualization/ 结果呈现层(ECharts / 信息图 / 学术图 / 双语 HTML Composer)
|
|
348
368
|
│
|
|
349
369
|
├─ 质量保障
|
|
350
|
-
│ ├─ tests/ pytest
|
|
370
|
+
│ ├─ tests/ pytest 测试矩阵(测试函数与文件数见源码仓库的 docs/metrics.json)
|
|
351
371
|
│ └─ benchmarks/ 30 题 + 30 份金标注 + B0–B4 评测框架
|
|
352
372
|
│
|
|
353
373
|
└─ 演示与分发
|
|
@@ -355,11 +375,18 @@ EduEvidence/ (= 一个 Skill 包)
|
|
|
355
375
|
├─ docs/ 架构 / 方法论 / Benchmark / Demo / 复现指南
|
|
356
376
|
├─ install.sh 一键安装(本地 / 多 Agent Skill)+ 自检
|
|
357
377
|
├─ pyproject.toml 打包元数据(核心零第三方依赖)
|
|
358
|
-
└─ README
|
|
378
|
+
└─ README.md / README.zh-CN.md 双语说明
|
|
359
379
|
```
|
|
360
380
|
|
|
361
381
|
> Skill 包设计原则:**运行所需的最小集是 `SKILL.md + skill/ + references/ + schemas/ + scripts/`**;`retrieval/`、`integrations/`、`visualization/` 是让 Skill 真正"可运行、可呈现"的执行层;`tests/`、`benchmarks/`、`examples/`、`docs/` 是可信度与上手保障,不影响 Skill 本体。
|
|
362
382
|
|
|
383
|
+
### 分发
|
|
384
|
+
|
|
385
|
+
两条发布通道共用同一份运行时 allowlist(`scripts/skill_payload.py`,宿主安装脚本也用这一份):
|
|
386
|
+
|
|
387
|
+
- **提交包** —— `bash packaging/make_upload.sh` 重建 `dist/eduevidence-submission/`(扁平 Skill 包 + `submission-manifest.json` + `UPLOAD-README.md`);它会先重建五种风格的报告变体,缺少已构建的 `web/studio/index.html` 时直接中止。旧 `upload/` 目录是**冻结的废弃快照**,不得随包分发 —— 见 [`upload/DEPRECATED.md`](upload/DEPRECATED.md)。
|
|
388
|
+
- **npm 包** —— `npm install -g eduevidence`,然后用 `eduevidence skill --list-hosts`(列出全部宿主与落点)或 `eduevidence skill --dry-run`(只预览,不写入)。在仓库目录中等价写法是 `node bin/eduevidence.js skill --list-hosts` / `--dry-run`。
|
|
389
|
+
|
|
363
390
|
### SCP / Platform Native Mode
|
|
364
391
|
|
|
365
392
|
EduEvidence 可完全脱离 Agent MCP 独立运行(无需任何外部服务):
|
|
@@ -442,7 +469,7 @@ pytest
|
|
|
442
469
|
**已完成(仅 harness / 仿真,标注 SIMULATED,非实证):**
|
|
443
470
|
|
|
444
471
|
- [x] Benchmark v2 harness / simulation —— `benchmarks/results/` 是确定性仿真,证明评测框架可运行,**不是**真实模型性能(见上方 [Benchmark](#benchmark) 的 ⚠️ 说明)。
|
|
445
|
-
- [x] Skill 核心与管线:9 步协议(Research Core 6 + Decision Extension 3)、
|
|
472
|
+
- [x] Skill 核心与管线:9 步协议(Research Core 6 + Decision Extension 3)、50 个版本化 JSON Schema(`schemas/**`,计数见 `scripts/generate_metrics.py`)、确定性脚本、8 角色协议(原计划 Phase 0–6)。
|
|
446
473
|
- [x] Evidence-to-Action:适用性 / 四态决策 / 干预 / 评价设计。
|
|
447
474
|
- [x] 产品 UI:单文件双语 HTML 报告 + 信息图 + 学术图(原计划 Phase 8)。
|
|
448
475
|
|
package/SKILL.md
CHANGED
|
@@ -2,7 +2,7 @@
|
|
|
2
2
|
name: eduevidence
|
|
3
3
|
description: "Decision-grade evidence synthesis for education and applied social science intervention decisions. Use when a user needs to determine whether, when, for whom, or how to adopt, pilot, evaluate, or revise a teaching method, curriculum change, AI tool, program, or policy intervention. Run an auditable evidence-to-decision workflow spanning systematic retrieval, counter-evidence challenge, methodological quality and evidence-certainty appraisal, provenance-traceable evidence graphs, applicability boundaries, evidence-grounded gap detection, preregistration-ready study or pilot design, empirical evidence re-injection, and decision revision."
|
|
4
4
|
---
|
|
5
|
-
# EduEvidence 6.
|
|
5
|
+
# EduEvidence 6.3 — Decision-Grade Evidence Engine
|
|
6
6
|
> **AI4SS Track | Art–Science Integration · General Intelligence**
|
|
7
7
|
> **From empirical questions to decision-grade evidence and evidence-to-action loops.**
|
|
8
8
|
|
|
@@ -28,6 +28,7 @@ Treat roles, models, Agent MCP, retrievers, scripts, and HTML reports as **execu
|
|
|
28
28
|
---
|
|
29
29
|
|
|
30
30
|
## 2. Route the Request to a Workflow
|
|
31
|
+
Before selecting a primary workflow, run the one-shot two-round intake in `skill/workflows/intake.md` (research question → execution enhancements → depth; then enhancement mappings + Frame boundary confirmation). Persist non-question preferences via `scripts/intake/` to `~/.eduevidence/prefs.json`. After the wrap-up summary is confirmed once, continue unattended.
|
|
31
32
|
Load exactly one primary workflow for the current task:
|
|
32
33
|
- `skill/workflows/evidence-review.md` — use for evidence synthesis and decision appraisal from existing research.
|
|
33
34
|
- `skill/workflows/decision-and-pilot.md` — use when the evidence must be converted into an actionable pilot or intervention plan.
|
|
@@ -109,7 +110,7 @@ Applicability → Intervene → Evaluate
|
|
|
109
110
|
| 3 | **Extract** | `evidence.schema.json` | Extract findings, effect sizes, confidence intervals, sample sizes, outcomes, population characteristics, and study design information when available. |
|
|
110
111
|
| 4 | **Challenge** | `evidence.schema.json` | Search explicitly for null findings, negative findings, contradictory evidence, alternative explanations, and confounders. |
|
|
111
112
|
| 5 | **Audit** | `methodology.schema.json` | Appraise study quality and evidence certainty. Apply WWC 5.0 criteria where relevant to education-study design; use GRADE-informed certainty assessment at the body-of-evidence level where appropriate. |
|
|
112
|
-
| 6 | **Adjudicate** | `verdict.schema.json` | Integrate evidence and emit a bounded decision: `ADOPT`, `PILOT`, `
|
|
113
|
+
| 6 | **Adjudicate** | `verdict.schema.json` | Integrate evidence and emit a bounded decision: `ADOPT`, `PILOT`, `REJECT`, or `INSUFFICIENT_EVIDENCE`. Pass the Pre-Verdict Gate before finalizing. |
|
|
113
114
|
| 7 | **Applicability** | `references/applicability-policy.md` | State who the evidence applies to, in which contexts, for which outcomes, and under what conditions. |
|
|
114
115
|
| 8 | **Intervene** | `intervention.schema.json` | Design the minimum viable intervention or pilot and define explicit success, failure, and stop conditions. |
|
|
115
116
|
| 9 | **Evaluate** | `evaluation.schema.json` | Define baseline, post-intervention, retention/maintenance, transfer, and decision-update logic. |
|
|
@@ -138,7 +139,7 @@ Include, where relevant:
|
|
|
138
139
|
### Gate C — No Direct Learning Evidence, No ADOPT
|
|
139
140
|
For education interventions, do not issue `ADOPT` based only on task speed, task completion, productivity, usability, preference, or subjective experience.
|
|
140
141
|
Require direct evidence on learning, retention, independent transfer, or another explicitly decision-relevant outcome before `ADOPT` can be considered.
|
|
141
|
-
If direct evidence is missing, bound the decision to `PILOT`, `
|
|
142
|
+
If direct evidence is missing, bound the decision to `PILOT`, `INSUFFICIENT_EVIDENCE`, or `REJECT` as justified by the evidence.
|
|
142
143
|
### Gate D — No False Precision
|
|
143
144
|
Never invent missing uncertainty statistics.
|
|
144
145
|
- If a confidence interval is not reported, do not fabricate one.
|
|
@@ -303,6 +304,14 @@ The example demonstrates:
|
|
|
303
304
|
- Empirical evidence re-injection followed by decision revision.
|
|
304
305
|
Do not generalize the flagship verdict to unrelated populations, courses, tools, or policy contexts.
|
|
305
306
|
|
|
307
|
+
The four-state matrix (`engine.decision_policy.decision_outcome`) is the gate-enforced landing; each public case also records the adjudicator's own `recommended_action`. Recomputed from the pack's own `evidence.jsonl` + confidence label:
|
|
308
|
+
|
|
309
|
+
- `examples/ai-coding-assistant-evidence/` - matrix `INSUFFICIENT_EVIDENCE` (`unresolved_conflict`: the decisive relations contain both `support_adoption` and `oppose_adoption`, so the conflict check fires first) / Moderate / 0.586. The pack records `PILOT`; primary evidence stops at task performance (no directness-2 learning outcome), so the decision is bounded — but the matrix does not let a conflicted relation set claim `PILOT`.
|
|
310
|
+
- `examples/spaced-retrieval-practice/` - matrix `ADOPT` / High / 0.893. Retention and transfer, the primary outcomes, carry direct evidence at directness 2. The pack records `ADOPT`; this is the only public case where stated and matrix landings agree.
|
|
311
|
+
- `examples/workplace-ai-assistant/` - matrix `INSUFFICIENT_EVIDENCE` with `downgrade_reason=None` / Moderate / 0.578, using the policy domain contract: all four decisive relations are `conditional`, and `conditional` is not a conflict relation (`CONFLICT_RELATIONS` = `{conflict, mixed}`), so no downgrade reason is recorded and, with no decisive `support_adoption`, Moderate confidence cannot reach `PILOT`. The pack records `PILOT`.
|
|
312
|
+
|
|
313
|
+
A verdict never awards itself an action: the Pre-Verdict Gate re-derives primary-outcome directness from the evidence corpus and caps an unsupported `ADOPT` to `PILOT`. Never present a case as ADOPT without that derivation passing. Where the matrix landing is `INSUFFICIENT_EVIDENCE`, do not describe the case as having landed at `PILOT` — report both the recorded action and the matrix bound, and do not invent evidence to force a stronger landing.
|
|
314
|
+
|
|
306
315
|
---
|
|
307
316
|
|
|
308
317
|
## 16. Keep Presentation as a Projection Layer
|
|
@@ -391,7 +400,8 @@ python3 scripts/dashboard_server.py --port 8765
|
|
|
391
400
|
# Search academic and current evidence
|
|
392
401
|
python3 -m retrieval.search "AI coding assistants learning transfer"
|
|
393
402
|
# Run the DID fixture / empirical analysis path
|
|
394
|
-
python3 scripts/did_regression.py
|
|
403
|
+
python3 scripts/did_regression.py <your.csv> # needs treat / post / outcome columns
|
|
404
|
+
# column names are matched case-insensitively; see scripts/did_regression.py
|
|
395
405
|
# Compute an effect size
|
|
396
406
|
python3 scripts/effect_calculator.py \
|
|
397
407
|
--mean1 78.5 --sd1 10.2 --n1 90 \
|
|
@@ -442,7 +452,7 @@ This appendix keeps the deterministic repository gates (skill lint / version / m
|
|
|
442
452
|
- **EduEvidence Research Engine** — the runnable engine delivered as a Skill package (`engine/`, `retrieval/`, `scripts/`, `schemas/`).
|
|
443
453
|
- **Shared Research Library** — verified external facts (`Source` / `Study` / `Finding` / `Audit`) may be reused across project snapshots; interpretive objects (`Claim` / `EvidenceLink` / `Applicability` / `Decision`) stay project-local.
|
|
444
454
|
- **No new study design without evidence grounding** — every study or pilot must cite an explicit, evidence-supported `KnowledgeGap` identifier.
|
|
445
|
-
- **Schema 版本口径**:schemas/ 顶层
|
|
455
|
+
- **Schema 版本口径**:schemas/ 顶层 15 个 = V1 契约(evidence.schema.json 当前修订 1.1、education-frame / verdict / report-spec 等);schemas/v2/ 18 个 = V2 契约(evidence-link / research-intent / intake / study / graph-revision / project 等);另有 v3 3 个、v4 4 个、vNext 10 个,递归合计 50 个(计数以 `scripts/generate_metrics.py` 为准)。文档与代理配置一律以此口径命名。
|
|
446
456
|
- Five baked report themes (presentation systems, not science):
|
|
447
457
|
|
|
448
458
|
├─ Claude Research [Light]
|
|
Binary file
|
|
Binary file
|