eduevidence 6.0.0 → 6.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CONTRIBUTING.md +105 -0
- package/README.md +93 -38
- package/README.zh-CN.md +26 -6
- package/SKILL.md +11 -2
- package/assets/readme/landing-tour.gif +0 -0
- package/assets/readme/studio-tour.gif +0 -0
- package/bin/eduevidence.js +2 -1
- package/docs/architecture.md +319 -43
- package/docs/demo-workplace-ai.md +1 -1
- package/docs/install-guide.md +1 -1
- package/docs/orchestration-role-model.md +1 -1
- package/docs/release-closeout/README.md +1 -1
- package/docs/sciverse-api.md +125 -0
- package/eduevidence_cli.py +10 -0
- package/engine/decision_policy.py +96 -0
- package/engine/evidence_graph.py +14 -10
- package/engine/gaps.py +42 -22
- package/engine/ids.py +2 -0
- package/engine/library.py +6 -2
- package/engine/living.py +34 -4
- package/engine/migration.py +88 -3
- package/engine/orchestration.py +5 -5
- package/engine/paths.py +2 -0
- package/engine/pilot.py +34 -32
- package/engine/taxonomy.py +211 -0
- package/engine/tribunal.py +43 -31
- package/engine/versions.py +1 -1
- package/examples/ai-coding-assistant-evidence/EduEvidence_Report.html +1360 -146
- package/examples/ai-coding-assistant-evidence/artifact_manifest.json +3 -3
- package/examples/ai-coding-assistant-evidence/citation_check.json +1 -1
- package/examples/ai-coding-assistant-evidence/final_verdict.json +107 -0
- package/examples/ai-coding-assistant-evidence/gate_report.json +101 -0
- package/examples/ai-coding-assistant-evidence/report_spec.json +23 -12
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_academic.html +447 -127
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_claude.html +447 -127
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_datalab-dark.html +447 -127
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_datalab.html +447 -127
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_presentation.html +447 -127
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_academic.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_claude.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_datalab-dark.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_datalab.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_presentation.html +1360 -146
- package/examples/ai-coding-assistant-evidence/result.json +13 -9
- package/examples/ai-coding-assistant-evidence/result.zh.json +45 -41
- package/examples/ai-coding-assistant-evidence/skeptic.json +72 -0
- package/examples/ai-coding-assistant-evidence/verdict.json +6 -2
- package/examples/spaced-retrieval-practice/applicability.json +14 -0
- package/examples/spaced-retrieval-practice/artifact_manifest.json +15 -0
- package/examples/spaced-retrieval-practice/claims.jsonl +3 -0
- package/examples/spaced-retrieval-practice/evidence.jsonl +6 -0
- package/examples/spaced-retrieval-practice/final_verdict.json +93 -0
- package/examples/spaced-retrieval-practice/frame.json +58 -0
- package/examples/spaced-retrieval-practice/gate_report.json +101 -0
- package/examples/spaced-retrieval-practice/methodology.json +78 -0
- package/examples/spaced-retrieval-practice/report_spec.json +212 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_academic.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_claude.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_datalab-dark.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_datalab.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_presentation.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_academic.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_claude.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_datalab-dark.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_datalab.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_presentation.html +2728 -0
- package/examples/spaced-retrieval-practice/result.json +942 -0
- package/examples/spaced-retrieval-practice/result.zh.json +942 -0
- package/examples/spaced-retrieval-practice/skeptic.json +70 -0
- package/examples/spaced-retrieval-practice/sources.jsonl +7 -0
- package/examples/spaced-retrieval-practice/verdict.json +93 -0
- package/examples/workplace-ai-assistant/artifact_manifest.json +15 -0
- package/examples/workplace-ai-assistant/claims.jsonl +4 -4
- package/examples/workplace-ai-assistant/evidence.jsonl +4 -4
- package/examples/workplace-ai-assistant/evidence_graph.json +15 -15
- package/examples/workplace-ai-assistant/final_verdict.json +78 -0
- package/examples/workplace-ai-assistant/gate_report.json +101 -0
- package/examples/workplace-ai-assistant/report_spec.json +209 -40
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_academic.html +435 -105
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_claude.html +435 -105
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_datalab-dark.html +435 -105
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_datalab.html +435 -105
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_presentation.html +435 -105
- package/examples/workplace-ai-assistant/reports-5themes/report_academic.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_claude.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_datalab-dark.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_datalab.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_presentation.html +2814 -0
- package/examples/workplace-ai-assistant/result.json +82 -20
- package/examples/workplace-ai-assistant/result.zh.json +82 -20
- package/examples/workplace-ai-assistant/skeptic.json +72 -0
- package/examples/workplace-ai-assistant/verdict.json +36 -10
- package/integrations/agent_mcp.py +2 -2
- package/package.json +12 -3
- package/pyproject.toml +4 -3
- package/references/report-copy-style.md +67 -0
- package/references/retrieval-compliance.md +75 -0
- package/references/retrieval-protocol.md +20 -0
- package/retrieval/audit.py +27 -3
- package/retrieval/fetch.py +96 -0
- package/retrieval/sciverse.py +398 -0
- package/retrieval/search.py +47 -7
- package/schemas/applicability.schema.json +94 -0
- package/schemas/chart-spec.schema.json +10 -3
- package/schemas/evidence.schema.json +316 -43
- package/schemas/fetch-result.schema.json +2 -1
- package/schemas/report-result.schema.json +3 -3
- package/schemas/report-spec.schema.json +98 -100
- package/schemas/skeptic.schema.json +86 -0
- package/schemas/source.schema.json +21 -2
- package/schemas/v2/finding.schema.json +5 -1
- package/schemas/v2/methodology-audit.schema.json +5 -1
- package/schemas/v2/outcome.schema.json +28 -5
- package/schemas/v2/study.schema.json +5 -1
- package/schemas/vNext/autoevolve-session.schema.json +34 -1
- package/schemas/vNext/eval-snapshot.schema.json +77 -1
- package/schemas/vNext/execution-plan.schema.json +50 -1
- package/schemas/vNext/gap-priority.schema.json +54 -1
- package/schemas/vNext/negative-search-record.schema.json +68 -1
- package/schemas/vNext/research-iteration.schema.json +87 -1
- package/schemas/vNext/research-strategy.schema.json +62 -1
- package/schemas/vNext/skill-experiment.schema.json +90 -1
- package/schemas/vNext/task-spec.schema.json +156 -1
- package/schemas/vNext/worker-result.schema.json +60 -1
- package/schemas/verdict.schema.json +164 -28
- package/scripts/build_esl_artifacts.py +2 -2
- package/scripts/build_report_variants.py +18 -2
- package/scripts/build_result.py +74 -9
- package/scripts/check_package_parity.py +85 -0
- package/scripts/check_protocol_alignment.py +375 -0
- package/scripts/check_versioned_schemas.py +254 -0
- package/scripts/claim_audit.py +13 -8
- package/scripts/compute_confidence.py +10 -0
- package/scripts/did_regression.py +12 -2
- package/scripts/evidence_score.py +5 -2
- package/scripts/generate_new_projects.py +4 -4
- package/scripts/orchestrator.py +120 -24
- package/scripts/pre_verdict_gate.py +224 -26
- package/scripts/quickstart.py +18 -2
- package/scripts/run_workspace.py +7 -1
- package/scripts/skill_payload.py +4 -1
- package/scripts/test_adversarial_empirical.py +26 -19
- package/scripts/validate_schema.py +31 -1
- package/skill/agents/evaluation-designer.md +20 -4
- package/skill/agents/evidence-analyst.md +19 -3
- package/skill/agents/evidence-judge.md +50 -2
- package/skill/agents/evidence-retriever.md +20 -3
- package/skill/agents/intervention-designer.md +20 -4
- package/skill/agents/method-reviewer.md +18 -2
- package/skill/agents/{education-planner.md → research-planner.md} +19 -3
- package/skill/agents/skeptic.md +18 -2
- package/skill/roles/registry.yaml +11 -11
- package/skill/sub-skills/aihot-trend-analysis/SKILL.md +28 -9
- package/skill/sub-skills/contradiction-analysis/SKILL.md +31 -11
- package/skill/sub-skills/data-analysis/SKILL.md +34 -15
- package/skill/sub-skills/ethics-review/SKILL.md +33 -10
- package/skill/sub-skills/evidence-extraction/SKILL.md +29 -11
- package/skill/sub-skills/evidence-review/SKILL.md +31 -12
- package/skill/sub-skills/gap-analysis/SKILL.md +31 -9
- package/skill/sub-skills/literature-review/SKILL.md +35 -14
- package/skill/sub-skills/methodology-audit/SKILL.md +29 -12
- package/skill/sub-skills/report-generation/SKILL.md +28 -0
- package/skill/sub-skills/research-planning/SKILL.md +41 -14
- package/skill/sub-skills/study-design/SKILL.md +30 -9
- package/skill/task-briefs/adjudicate.md +32 -7
- package/skill/task-briefs/applicability.md +37 -2
- package/skill/task-briefs/audit.md +32 -7
- package/skill/task-briefs/challenge.md +34 -5
- package/skill/task-briefs/evaluate.md +30 -5
- package/skill/task-briefs/extract.md +31 -8
- package/skill/task-briefs/frame.md +39 -10
- package/skill/task-briefs/intervene.md +32 -6
- package/skill/task-briefs/present.md +32 -8
- package/skill/task-briefs/projection.md +36 -2
- package/skill/task-briefs/retrieve.md +36 -6
- package/skill/workflows/decision-and-pilot.md +76 -1
- package/skill/workflows/evaluate-and-update.md +83 -0
- package/skill/workflows/evidence-review.md +104 -0
- package/visualization/eduevidence-report/scripts/build_figures.py +25 -3
- package/visualization/eduevidence-report/scripts/build_infographics.py +5 -1
- package/visualization/eduevidence-report/scripts/build_report.py +512 -65
- package/visualization/eduevidence-report/scripts/charts_data.py +2 -0
- package/visualization/eduevidence-report/scripts/lieflat_engine.py +349 -38
- package/visualization/eduevidence-report/scripts/zh_labels.py +80 -1
- package/web/architecture.html +14885 -0
- package/web/studio/assets/index-B8tkF44Q.css +1 -0
- package/web/studio/index.html +2 -2
- package/web/studio/assets/index-CzXocaGv.css +0 -1
- /package/web/studio/assets/{index-pa7jD7n4.js → index-CQ6Keoyc.js} +0 -0
|
@@ -1,8 +1,10 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: evidence-analyst
|
|
3
3
|
description: EduEvidence 证据分析者。把候选 Source 抽取为 Claim-Level Evidence Object(绑定 Outcome、direction、quality_dimensions),执行 Outcome Separation;只结构化,不裁决。
|
|
4
|
-
|
|
5
|
-
|
|
4
|
+
role_id: evidence-analyst
|
|
5
|
+
capabilities: study_extraction, finding_extraction, claim_linking
|
|
6
|
+
output_contracts: evidence.jsonl (schemas/evidence.schema.json)
|
|
7
|
+
recommended_reasoning: medium+ # capability hint only — no model or CLI name is bound here
|
|
6
8
|
default_permission: read
|
|
7
9
|
default_summary_chars: 1200
|
|
8
10
|
default_context_mode: full
|
|
@@ -81,7 +83,7 @@ critical_path: true
|
|
|
81
83
|
|
|
82
84
|
- `relation_to_claim`:该证据支持/反驳某条 claim(Claim Audit 只依据此字段);
|
|
83
85
|
- `effect_direction`:研究观察到的效应方向(Outcome 可视化/聚合只依据此字段);
|
|
84
|
-
- `decision_relation
|
|
86
|
+
- `decision_relation`:对最终决策的意义(Consistency/Tribunal 依据此字段);
|
|
85
87
|
- 旧字段 `direction` 已废弃(deprecated),优先使用 `relation_to_claim`,不要再新写。
|
|
86
88
|
|
|
87
89
|
**类型/格式硬约束(FIX-2 实测违规项,逐条禁止)**:
|
|
@@ -104,3 +106,17 @@ critical_path: true
|
|
|
104
106
|
## 卡住升级
|
|
105
107
|
|
|
106
108
|
原文不可得回传 `NEEDS_CONTEXT: <缺哪篇原文>`;原文声称与抽取冲突回传 BLOCKED 并说明。
|
|
109
|
+
|
|
110
|
+
## 独立性与交叉评审
|
|
111
|
+
|
|
112
|
+
- 抽取只结构化、不裁决;`relation_to_claim` 的最终归属由 Evidence Judge 在裁决阶段复核。
|
|
113
|
+
- 与宿主的模型选择解耦:本文件只声明能力要求(`recommended_reasoning: medium+`、结构化输出强),具体 CLI/模型由用户确认的模型映射决定。
|
|
114
|
+
|
|
115
|
+
## 失败模式与回退
|
|
116
|
+
|
|
117
|
+
| 失败 | 处理 |
|
|
118
|
+
|---|---|
|
|
119
|
+
| 强制字段缺失 | 该对象标 `UNSUPPORTED`,不得带缺陷进入合成。 |
|
|
120
|
+
| 原文不可得 | `NEEDS_CONTEXT: <缺哪篇原文>`;禁止从摘要或记忆补全。 |
|
|
121
|
+
| 原文与结论冲突 | 回传 BLOCKED 并说明冲突点,交审计阶段处理。 |
|
|
122
|
+
| 统计量未报告 | 保持缺失,禁止由显著性反推。 |
|
|
@@ -1,8 +1,10 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: evidence-judge
|
|
3
3
|
description: EduEvidence 证据裁决者。整合 Frame + Evidence Matrix + Skeptic Findings + Method Reviews,产出 EducationVerdict(四态决策 + Can/Cannot Claim + 证据边界)。
|
|
4
|
-
|
|
5
|
-
|
|
4
|
+
role_id: evidence-judge
|
|
5
|
+
capabilities: evidence_synthesis, tribunal, applicability_analysis, knowledge_gap_detection
|
|
6
|
+
output_contracts: final_verdict.json (schemas/verdict.schema.json), applicability.json
|
|
7
|
+
recommended_reasoning: highest # capability hint only — no model or CLI name is bound here
|
|
6
8
|
default_permission: read
|
|
7
9
|
default_summary_chars: 1000
|
|
8
10
|
default_context_mode: full
|
|
@@ -61,10 +63,42 @@ critical_path: true
|
|
|
61
63
|
"missing_evidence": ["..."],
|
|
62
64
|
"recommended_action": "adopt|pilot|reject|insufficient_evidence",
|
|
63
65
|
"decision_rationale": "...",
|
|
66
|
+
"strongest_support": "...",
|
|
67
|
+
"key_uncertainty": "...",
|
|
68
|
+
"main_risk": "...",
|
|
69
|
+
"next_action": "...",
|
|
64
70
|
"exceeds_evidence_boundary": ["..."]
|
|
65
71
|
}
|
|
66
72
|
```
|
|
67
73
|
|
|
74
|
+
## 读者向决策叙事(四件套 · 硬要求)
|
|
75
|
+
|
|
76
|
+
以下四个字段是**成品文案**,不是字段摘录:必须由你一次写成完整句子,
|
|
77
|
+
渲染器只负责呈现,缺字段就显示「未产出」。规范见 `references/report-copy-style.md`。
|
|
78
|
+
|
|
79
|
+
| 字段 | 内容 | 字数上限(中文) |
|
|
80
|
+
|---|---|---|
|
|
81
|
+
| `strongest_support` | 证据支持的最强结论,一句话说清 | ≤60 字 |
|
|
82
|
+
| `key_uncertainty` | 与决策相关的最大不确定性或反证 | ≤70 字 |
|
|
83
|
+
| `main_risk` | 采取行动的主要风险 | ≤60 字 |
|
|
84
|
+
| `next_action` | 建议的下一步 | ≤80 字 |
|
|
85
|
+
|
|
86
|
+
写作要求:
|
|
87
|
+
|
|
88
|
+
- 每条都是可独立阅读的完整句子;读者不需要看别的字段就能理解。
|
|
89
|
+
- 面向非本领域决策者;先结论、后依据;一句话一个意思。
|
|
90
|
+
- 禁止出现内部字段名、存储标识、证据 ID 列表(`E-001、E-006`);引用研究用「作者-年份 + 人话描述」。
|
|
91
|
+
- 缺失信息如实写「尚无直接证据」,不要用模糊措辞掩盖。
|
|
92
|
+
- en / zh 两版各自成篇,语义对齐而非逐字直译。
|
|
93
|
+
|
|
94
|
+
反例(渲染器拼装出来的读感,禁止):
|
|
95
|
+
|
|
96
|
+
```text
|
|
97
|
+
❌ 下一步:当前结果未提供此项信息。
|
|
98
|
+
❌ 最强支持结论:AI 编程助手在训练期提升新手任务表现。(从 what_can_be_claimed[0] 截取)
|
|
99
|
+
✅ 下一步:开展分阶段 CS1 试点——给提示而非答案、每周实验课使用,并以无 AI 迁移考试作为可叫停的验收条件。
|
|
100
|
+
```
|
|
101
|
+
|
|
68
102
|
## 输出契约(必须遵守)
|
|
69
103
|
|
|
70
104
|
你的产物 `final_verdict.json` 必须通过 `schemas/verdict.schema.json` 校验(stage `adjudicate` 的 schema-gate,首次生成即必须合规)。schema 顶层 `additionalProperties: false`,未列出的字段一律放入 `extensions`。
|
|
@@ -109,3 +143,17 @@ critical_path: true
|
|
|
109
143
|
- **禁止**在理由与主张列表里堆证据 ID(E-xxx / EV-xxx)、来源码(PAP-xxx)或 schema 键(overall_risk=、CONCERN 等);引用研究用"作者-年份 + 人话描述"(如"带护栏组独立考试未见下滑");
|
|
110
144
|
- `what_can_be_claimed / what_cannot_be_claimed / missing_evidence / exceeds_evidence_boundary` 同样人话化;统计数字可保留,但用自然表达("效应量 +0.61,差异显著");
|
|
111
145
|
- 无截断残留(null、…)、无中英夹生;en/zh 两个语版分写,语义对齐而非机翻。
|
|
146
|
+
|
|
147
|
+
## 独立性与交叉评审
|
|
148
|
+
|
|
149
|
+
- 裁决以证据矩阵、反证与审计三路输入为准,不以任一单路由结论为准;交叉审核输出须符合 `schemas/cross-model-review.schema.json`。
|
|
150
|
+
- 与宿主的模型选择解耦:本文件只声明能力要求(`recommended_reasoning: highest`、结构化输出强),具体 CLI/模型由用户确认的模型映射决定。
|
|
151
|
+
|
|
152
|
+
## 失败模式与回退
|
|
153
|
+
|
|
154
|
+
| 失败 | 处理 |
|
|
155
|
+
|---|---|
|
|
156
|
+
| `PRE_VERDICT_FAILED` | 修复前置产物后重跑闸门,不得跳过。 |
|
|
157
|
+
| `GATE_CRITICAL_FAILURE` | 封顶置信度,强制降级为 PILOT 或 INSUFFICIENT EVIDENCE。 |
|
|
158
|
+
| `CONFLICT_UNRESOLVED` | 保持不确定,不强行裁决。 |
|
|
159
|
+
| 证据只支持任务表现 | 不得产出学习效果类结论。 |
|
|
@@ -1,8 +1,10 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: evidence-retriever
|
|
3
3
|
description: EduEvidence 证据检索者。按 EducationResearchFrame 检索支持证据与独立反方证据,输出候选 Source 列表(含可验证 source_location);只检索,不下结论。
|
|
4
|
-
|
|
5
|
-
|
|
4
|
+
role_id: evidence-retriever
|
|
5
|
+
capabilities: literature_search, counter_evidence_search, source_fetch, source_validation
|
|
6
|
+
output_contracts: sources.jsonl (schemas/source.schema.json), fetch/
|
|
7
|
+
recommended_reasoning: high # capability hint only — no model or CLI name is bound here
|
|
6
8
|
default_permission: read
|
|
7
9
|
default_summary_chars: 1000
|
|
8
10
|
default_context_mode: compact
|
|
@@ -13,7 +15,7 @@ critical_path: false
|
|
|
13
15
|
|
|
14
16
|
## 职责
|
|
15
17
|
|
|
16
|
-
1. 按 Frame 的
|
|
18
|
+
1. 按 Frame 的 population/intervention/comparison/outcomes/scope 构造检索式(字段词汇随领域而定);
|
|
17
19
|
2. **双路检索**:一路找支持证据,一路独立找反方证据(null result / negative result / contradictory evidence / AI dependency / reduced transfer);
|
|
18
20
|
3. 优先 RCT / quasi-experimental / meta-analysis,标注 study_type;
|
|
19
21
|
4. 每条来源必须有可验证 `source_location`(DOI / URL / 数据库标识)——没有位置=无效来源;
|
|
@@ -78,3 +80,18 @@ critical_path: false
|
|
|
78
80
|
## 卡住升级
|
|
79
81
|
|
|
80
82
|
检索工具不可用回传 `TOOL_FAILURE: <工具 + 现象>`;检索结果为零且无法扩大范围回传 `INSUFFICIENT_SOURCES`。
|
|
83
|
+
|
|
84
|
+
## 独立性与交叉评审
|
|
85
|
+
|
|
86
|
+
- 反方检索必须独立构造检索式,不复用支持证据的查询;其结果由 Skeptic 独立复核,不由本角色判定"是否充分"。
|
|
87
|
+
- 与宿主的模型选择解耦:本文件只声明能力要求(`recommended_reasoning: high`、`tool_use: strong`、成本低),具体 CLI/模型由用户确认的模型映射决定。
|
|
88
|
+
|
|
89
|
+
## 失败模式与回退
|
|
90
|
+
|
|
91
|
+
| 失败 | 处理 |
|
|
92
|
+
|---|---|
|
|
93
|
+
| `TOOL_FAILURE` | 记录工具与现象,切换等价通道后重试;不得凭记忆补来源。 |
|
|
94
|
+
| `SEARCH_NO_RESULT` | 放宽词族、切换 provider,或落 negative-search record;不静默降低标准。 |
|
|
95
|
+
| `FETCH_FAILED` | 走 provider 降级链;链尽则弃用该来源。 |
|
|
96
|
+
| `SOURCE_INVALID` / `SOURCE_DUPLICATE` | 弃用 / 合并(保留最高权威等级)。 |
|
|
97
|
+
| `INSUFFICIENT_SOURCES` | 如实上报,不用低权威来源凑数。 |
|
|
@@ -1,15 +1,17 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: intervention-designer
|
|
3
|
-
description: EduEvidence
|
|
4
|
-
|
|
5
|
-
|
|
3
|
+
description: EduEvidence 干预设计者。把 Verdict 转化为"最小可验证试点":阶段化使用规则、护栏、停止条件与证据对齐;禁止直接推荐全面部署。干预对象随领域而定(教学 / 政策 / 组织流程)。
|
|
4
|
+
role_id: intervention-designer
|
|
5
|
+
capabilities: study_design, measurement_design, intervention_design
|
|
6
|
+
output_contracts: intervention.json (schemas/intervention.schema.json)
|
|
7
|
+
recommended_reasoning: high # capability hint only — no model or CLI name is bound here
|
|
6
8
|
default_permission: read
|
|
7
9
|
default_summary_chars: 800
|
|
8
10
|
default_context_mode: compact
|
|
9
11
|
critical_path: false
|
|
10
12
|
---
|
|
11
13
|
|
|
12
|
-
你是 EduEvidence 的 **Intervention Designer
|
|
14
|
+
你是 EduEvidence 的 **Intervention Designer**。你的产出必须是从证据长出来的试点方案,而不是凭空的创意。
|
|
13
15
|
|
|
14
16
|
## 职责
|
|
15
17
|
|
|
@@ -80,3 +82,17 @@ critical_path: false
|
|
|
80
82
|
## 卡住升级
|
|
81
83
|
|
|
82
84
|
Verdict 缺失回传 `NEEDS_CONTEXT`;用户课堂约束不明回传 `NEEDS_USER_CONTEXT: <缺什么>`。
|
|
85
|
+
|
|
86
|
+
## 独立性与交叉评审
|
|
87
|
+
|
|
88
|
+
- 设计必须引用显式 KnowledgeGap ID;是否存在合格缺口由 Gap Analysis 与 StudyDesign 门判定,不由本角色自证。
|
|
89
|
+
- 涉及学生数据与对照分组时,先过 `skill/sub-skills/ethics-review/SKILL.md`。
|
|
90
|
+
- 与宿主的模型选择解耦:本文件只声明能力要求(`recommended_reasoning: high`),具体 CLI/模型由用户确认的模型映射决定。
|
|
91
|
+
|
|
92
|
+
## 失败模式与回退
|
|
93
|
+
|
|
94
|
+
| 失败 | 处理 |
|
|
95
|
+
|---|---|
|
|
96
|
+
| 无 KnowledgeGap | 不设计研究,改为报告"还需要什么证据"。 |
|
|
97
|
+
| 伦理审查未通过 | 阻断试点,先修正设计。 |
|
|
98
|
+
| 人群越出适用边界 | 缩小试点人群至支持范围内。 |
|
|
@@ -1,8 +1,10 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: method-reviewer
|
|
3
3
|
description: EduEvidence 方法学审查者。按 15 项清单审查每个研究的方法学质量,强制执行"任务完成表现≠学习效果"最高优先级规则,输出 MethodologyAudit。
|
|
4
|
-
|
|
5
|
-
|
|
4
|
+
role_id: method-reviewer
|
|
5
|
+
capabilities: methodology_appraisal
|
|
6
|
+
output_contracts: methodology.json (schemas/methodology.schema.json)
|
|
7
|
+
recommended_reasoning: high # capability hint only — no model or CLI name is bound here
|
|
6
8
|
default_permission: read
|
|
7
9
|
default_summary_chars: 800
|
|
8
10
|
default_context_mode: compact
|
|
@@ -102,3 +104,17 @@ critical_path: true
|
|
|
102
104
|
- 审计说明(note / summary / verdict 理由)为流畅人话(en/zh 分写);PASS / CONCERN / FAIL 只作枚举标签,由显示层映射中文;
|
|
103
105
|
- 禁止在叙述里堆证据 ID 或 schema 键;引用研究用"作者-年份 + 人话描述";
|
|
104
106
|
- 无截断残留、无中英夹生。
|
|
107
|
+
|
|
108
|
+
## 独立性与交叉评审
|
|
109
|
+
|
|
110
|
+
- **独立性要求(`independence_required: role-separation`)**:方法学判断必须独立于内容判断——审计输入只含设计与测量,不含结论评价。此处要求的是角色分离(审计说明不得夹带对效果的评价),而非跨模型家族;需要跨模型家族的只有 Skeptic。
|
|
111
|
+
- 只审"研究怎么测的",不审"结论是什么";审计结论不得夹带对效果的评价。
|
|
112
|
+
- 与宿主的模型选择解耦:本文件只声明能力要求(`recommended_reasoning: high`、高上下文),具体 CLI/模型由用户确认的模型映射决定。
|
|
113
|
+
|
|
114
|
+
## 失败模式与回退
|
|
115
|
+
|
|
116
|
+
| 失败 | 处理 |
|
|
117
|
+
|---|---|
|
|
118
|
+
| 关键信息未报告(如随机化方式) | 记 `missing` 并写明缺什么,不猜测。 |
|
|
119
|
+
| 结论依赖任务表现 | 触发 guard,剥夺其学习效果支撑资格。 |
|
|
120
|
+
| 审计与内容判断混写 | 拆开重写;审计说明只描述设计与测量。 |
|
|
@@ -1,8 +1,10 @@
|
|
|
1
1
|
---
|
|
2
|
-
name:
|
|
2
|
+
name: research-planner
|
|
3
3
|
description: EduEvidence 教育研究规划者。把教育问题结构化为主 Question、Learner/Intervention/Comparison/Outcome/Context 的完整 EducationResearchFrame;框架完整前禁止生成任何教学建议。
|
|
4
|
-
|
|
5
|
-
|
|
4
|
+
role_id: research-planner
|
|
5
|
+
capabilities: research_framing
|
|
6
|
+
output_contracts: frame.json (schemas/education-frame.schema.json)
|
|
7
|
+
recommended_reasoning: high # capability hint only — no model or CLI name is bound here
|
|
6
8
|
default_permission: read
|
|
7
9
|
default_summary_chars: 600
|
|
8
10
|
default_context_mode: compact
|
|
@@ -78,3 +80,17 @@ critical_path: true
|
|
|
78
80
|
## 卡住升级
|
|
79
81
|
|
|
80
82
|
问题矛盾或缺少关键信息时回传 `NEEDS_CONTEXT: <缺少什么 + why>`;不臆测学习者特征。
|
|
83
|
+
|
|
84
|
+
## 独立性与交叉评审
|
|
85
|
+
|
|
86
|
+
- 本角色产出第一道闸门;写入者与复核者分离:Frame 的完整性由 Evidence Judge 在裁决阶段复核,不由本角色自我确认。
|
|
87
|
+
- 关键路径角色(`critical_path: true`):Frame 缺项会阻断整条证据链,宁可 `NEEDS_CONTEXT` 也不填补空白。
|
|
88
|
+
- 与宿主的模型选择解耦:本文件只声明能力要求(`recommended_reasoning: high`),具体 CLI/模型由用户确认的模型映射决定,禁止在此绑定。
|
|
89
|
+
|
|
90
|
+
## 失败模式与回退
|
|
91
|
+
|
|
92
|
+
| 失败 | 处理 |
|
|
93
|
+
|---|---|
|
|
94
|
+
| 关键输入缺失(学习者层级 / 对照条件 / 主 outcome) | `NEEDS_CONTEXT: <缺什么 + 为何必要>`,不臆测、不继续。 |
|
|
95
|
+
| 问题跨多个决策 | 拆成多个 Frame,各自独立成 run。 |
|
|
96
|
+
| 与用户既有假设冲突 | 在 `extensions` 中记录冲突点,交用户确认后再继续。 |
|
package/skill/agents/skeptic.md
CHANGED
|
@@ -1,8 +1,10 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: skeptic
|
|
3
3
|
description: EduEvidence 反证挑战者。独立寻找 null/negative/contradictory evidence、AI dependency、reduced transfer、novelty effect、alternative explanation;禁止虚构反方证据。
|
|
4
|
-
|
|
5
|
-
|
|
4
|
+
role_id: skeptic
|
|
5
|
+
capabilities: counter_evidence_search
|
|
6
|
+
output_contracts: skeptic.json; cross-model-review (schemas/cross-model-review.schema.json)
|
|
7
|
+
recommended_reasoning: high # capability hint only — no model or CLI name is bound here
|
|
6
8
|
default_permission: read
|
|
7
9
|
default_summary_chars: 800
|
|
8
10
|
default_context_mode: compact
|
|
@@ -87,3 +89,17 @@ critical_path: true
|
|
|
87
89
|
- 反方证据描述(counter_evidence / null_results / confounders)为面向研究者的流畅中文(en 版为英文);禁止证据 ID 堆砌;
|
|
88
90
|
- 引用证据用"作者-年份 + 人话描述";禁止把内部字段名(search_performed、risk_level 等)写进叙述;
|
|
89
91
|
- 无截断残留、无中英夹生。
|
|
92
|
+
|
|
93
|
+
## 独立性与交叉评审
|
|
94
|
+
|
|
95
|
+
- **独立性要求(`independence_required: true`)**:本角色不得与主分析使用同一模型家族——独立性的目的是让反证来自不同先验,而不是换个会话问同一个模型。
|
|
96
|
+
- 找不到反证时输出标准语句 `NO CONTRADICTORY EVIDENCE FOUND` 并如实标注 `not_found`;宁可空手而归,也不虚构反方文献。
|
|
97
|
+
- 作为交叉审核者时按 `schemas/cross-model-review.schema.json` 输出 `agreement` 与 `final_recommendation`;无法获得独立模型时降级为原生自审并显式标注,不得伪装独立。
|
|
98
|
+
|
|
99
|
+
## 失败模式与回退
|
|
100
|
+
|
|
101
|
+
| 失败 | 处理 |
|
|
102
|
+
|---|---|
|
|
103
|
+
| 反证检索为空 | 落 negative-search record + 标准语句。 |
|
|
104
|
+
| 反证与支持证据冲突 | 双方都保留,交 Adjudicate 处理;本角色不裁决。 |
|
|
105
|
+
| 无独立模型可用 | 降级为原生自审并标注 `degraded_to: native_self_review`。 |
|
|
@@ -1,42 +1,42 @@
|
|
|
1
1
|
roles:
|
|
2
|
-
|
|
2
|
+
research-planner:
|
|
3
3
|
responsibility: framing completeness and scope
|
|
4
4
|
stages: [frame]
|
|
5
|
-
capabilities: [
|
|
5
|
+
capabilities: [research_framing]
|
|
6
6
|
critical_path: true
|
|
7
7
|
evidence-retriever:
|
|
8
8
|
responsibility: source acquisition and provenance
|
|
9
9
|
stages: [retrieve]
|
|
10
|
-
capabilities: [
|
|
10
|
+
capabilities: [literature_search, counter_evidence_search, source_fetch, source_validation]
|
|
11
11
|
evidence-analyst:
|
|
12
12
|
responsibility: structured finding extraction
|
|
13
13
|
stages: [extract]
|
|
14
|
-
capabilities: [
|
|
14
|
+
capabilities: [study_extraction, finding_extraction, claim_linking]
|
|
15
15
|
skeptic:
|
|
16
16
|
responsibility: independent counter-evidence coverage
|
|
17
17
|
stages: [challenge]
|
|
18
|
-
capabilities: [
|
|
19
|
-
independence_required:
|
|
18
|
+
capabilities: [counter_evidence_search]
|
|
19
|
+
independence_required: different-model-family
|
|
20
20
|
critical_path: true
|
|
21
21
|
method-reviewer:
|
|
22
22
|
responsibility: methodology and construct-validity appraisal
|
|
23
23
|
stages: [audit]
|
|
24
|
-
capabilities: [
|
|
25
|
-
independence_required:
|
|
24
|
+
capabilities: [methodology_appraisal]
|
|
25
|
+
independence_required: role-separation
|
|
26
26
|
critical_path: true
|
|
27
27
|
evidence-judge:
|
|
28
28
|
responsibility: evidence-bounded adjudication and applicability
|
|
29
29
|
stages: [adjudicate, applicability]
|
|
30
|
-
capabilities: [
|
|
30
|
+
capabilities: [evidence_synthesis, tribunal, applicability_analysis, knowledge_gap_detection]
|
|
31
31
|
critical_path: true
|
|
32
32
|
intervention-designer:
|
|
33
33
|
responsibility: grounded intervention or pilot design
|
|
34
34
|
stages: [intervene]
|
|
35
|
-
capabilities: [
|
|
35
|
+
capabilities: [study_design, measurement_design, intervention_design]
|
|
36
36
|
evaluation-designer:
|
|
37
37
|
responsibility: estimable evaluation and update logic
|
|
38
38
|
stages: [evaluate]
|
|
39
|
-
capabilities: [
|
|
39
|
+
capabilities: [evaluation_design, data_validation, data_analysis]
|
|
40
40
|
|
|
41
41
|
execution:
|
|
42
42
|
max_parallel_workers: 6
|
|
@@ -1,31 +1,50 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: aihot-trend-analysis
|
|
3
3
|
description: "Real-time horizon scanning and dynamic trend ingestion for emerging AI educational tools, model benchmarks, and EdTech releases via AIHot."
|
|
4
|
+
capability: literature_search (grey-literature channel)
|
|
4
5
|
---
|
|
5
6
|
|
|
6
|
-
# aihot-trend-analysis — Real-Time AI & EdTech Trend Ingestion
|
|
7
|
+
# aihot-trend-analysis — Real-Time AI & EdTech Trend Ingestion
|
|
7
8
|
|
|
8
9
|
## When to Use
|
|
9
|
-
Triggered when an
|
|
10
|
+
Triggered when an inquiry involves fast-moving generative AI tools (Cursor, Claude, Socratic LLM tutors, Copilot) where peer-reviewed literature may lag 6–18 months.
|
|
10
11
|
|
|
11
|
-
##
|
|
12
|
-
- `keyword`:
|
|
13
|
-
- `time_window
|
|
14
|
-
- `category
|
|
12
|
+
## Inputs
|
|
13
|
+
- `keyword`: target technology or pedagogy topic.
|
|
14
|
+
- `time_window` (optional): `24h` / `7d` / `30d`.
|
|
15
|
+
- `category` (optional): `EdTech` / `Agents` / `Reasoning` / `LLMs`.
|
|
16
|
+
|
|
17
|
+
## Process
|
|
18
|
+
1. Query the AIHot channel through `retrieval/search.py` (`AIHotProvider`).
|
|
19
|
+
2. Record every hit as grey literature with its publication time.
|
|
20
|
+
3. Route any factual claim that would enter the decision back through Retrieve → Fetch → Validate: trend items never bypass RULE 2.
|
|
15
21
|
|
|
16
22
|
## Output Contract
|
|
17
|
-
|
|
23
|
+
`SearchHit` objects tagged `provider: "aihot"` with grey-literature authority (`tier5_general_web`); they inform horizon scanning, not effect estimation.
|
|
18
24
|
|
|
19
25
|
```json
|
|
20
26
|
{
|
|
21
27
|
"trend_items": [
|
|
22
28
|
{
|
|
23
|
-
"title": "
|
|
29
|
+
"title": "Socratic tutoring framework evaluated across 10 universities",
|
|
24
30
|
"url": "https://aihot.virxact.com/api/item/...",
|
|
25
|
-
"summary": "Benchmark evaluation on novice
|
|
31
|
+
"summary": "Benchmark evaluation on novice retention and prompt scaffolding.",
|
|
26
32
|
"category": "EdTech",
|
|
27
33
|
"publish_time": "2026-08-15"
|
|
28
34
|
}
|
|
29
35
|
]
|
|
30
36
|
}
|
|
31
37
|
```
|
|
38
|
+
|
|
39
|
+
## Quality Gates
|
|
40
|
+
- [ ] 每条 trend 项带 URL 与时间戳。
|
|
41
|
+
- [ ] 明确标注为灰来源,不进入效应量合成。
|
|
42
|
+
|
|
43
|
+
## Anti-Patterns
|
|
44
|
+
- 用产品博客宣称的效果当作实证证据;把版本发布日期当研究发表时间。
|
|
45
|
+
|
|
46
|
+
## Worked Example
|
|
47
|
+
关键词 "AI programming assistant" → 返回 30 天内的发布与基准报道,用于判断文献滞后期内是否出现新的风险信号。
|
|
48
|
+
|
|
49
|
+
## References
|
|
50
|
+
- `retrieval/search.py::AIHotProvider`、`references/retrieval-compliance.md`
|
|
@@ -1,17 +1,37 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: contradiction-analysis
|
|
3
3
|
description: "Mines adversarial claims, conflicting effect directions, and boundary condition qualifiers."
|
|
4
|
+
capability: counter_evidence_search (skeptic side)
|
|
4
5
|
---
|
|
6
|
+
|
|
5
7
|
# Contradiction Analysis Skill
|
|
6
8
|
|
|
7
|
-
##
|
|
8
|
-
Trigger during evidence synthesis when studies on the same Claim ID
|
|
9
|
-
|
|
10
|
-
##
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
9
|
+
## When to Use
|
|
10
|
+
Trigger during evidence synthesis when studies on the same Claim ID show conflicting directions (SUPPORTS vs CONTRADICTS) or high heterogeneity.
|
|
11
|
+
|
|
12
|
+
## Inputs
|
|
13
|
+
- `evidence.jsonl`(含方向标签)
|
|
14
|
+
- `frame.json`(判断是否 scope overreach)
|
|
15
|
+
|
|
16
|
+
## Process
|
|
17
|
+
1. **Adversarial Mining (Skeptic)**: identify confounders (teacher training, dosage, novelty); evaluate boundary conditions (does it fail for novices vs experts?).
|
|
18
|
+
2. **Directional Separation**: strictly separate supporting / contradicting / neutral — never blend them into one "mixed" bucket.
|
|
19
|
+
3. **Heterogeneity Attribution**: map conflict to subgroup variation, dosage thresholds, or outcome instrument differences.
|
|
20
|
+
4. **Nine fixed checks** per `skill/agents/skeptic.md`; absence of counter-evidence yields the standard statement, never invented sources.
|
|
21
|
+
|
|
22
|
+
## Output Contract
|
|
23
|
+
`skeptic.json` — `skeptic_findings[]` (`check` / `status` / `detail` / `related_evidence_ids`), `contradictory_evidence_found`, `no_contradictory_evidence_statement`, `threats_to_validity`.
|
|
24
|
+
|
|
25
|
+
## Quality Gates
|
|
26
|
+
- [ ] 九项检查齐全。
|
|
27
|
+
- [ ] 每条 found 绑定证据或来源。
|
|
28
|
+
- [ ] 标准语句与布尔标志语义一致。
|
|
29
|
+
|
|
30
|
+
## Anti-Patterns
|
|
31
|
+
- 把三列合并成"总体看有效";为了显得严谨而虚构反证。
|
|
32
|
+
|
|
33
|
+
## Worked Example
|
|
34
|
+
同一 claim 下出现 g=+0.48(无护栏练习)与 g=−0.17(独立考试),归因到 outcome 测量差异与护栏配置,而非取平均。
|
|
35
|
+
|
|
36
|
+
## References
|
|
37
|
+
- `references/skeptic-protocol.md`、`skill/agents/skeptic.md`、`skill/task-briefs/challenge.md`
|
|
@@ -1,23 +1,42 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: data-analysis
|
|
3
3
|
description: "Runs deterministic statistical regression (DID/OLS) on user-uploaded classroom and field datasets to re-inject local empirical evidence."
|
|
4
|
+
capability: data_validation + data_analysis
|
|
4
5
|
---
|
|
6
|
+
|
|
5
7
|
# Data Analysis Skill
|
|
6
8
|
|
|
7
|
-
##
|
|
9
|
+
## When to Use
|
|
8
10
|
Trigger when the user imports empirical classroom or survey data (CSV/XLSX) from an active field trial or pilot deployment.
|
|
9
11
|
|
|
10
|
-
##
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
-
|
|
22
|
-
|
|
23
|
-
|
|
12
|
+
## Inputs
|
|
13
|
+
- 数据集 + 采集溯源(谁、何时、从哪个人群、何种同意)
|
|
14
|
+
- 预先注册的分析计划(不得在看到数据后修改)
|
|
15
|
+
|
|
16
|
+
## Process
|
|
17
|
+
1. **Data Ingestion & Cleaning**: profile columns, check missingness, identify Treatment and Post indicators; the provenance/hash/missingness gate runs before analysis.
|
|
18
|
+
2. **Deterministic DID Regression** (`scripts/did_regression.py`): Y = β0 + β1·Treat + β2·Post + δ·(Treat×Post) + ε; report δ, SE, t, p and Hedges' g (`scripts/effect_calculator.py`).
|
|
19
|
+
3. **Fail closed**: when the design is not estimable (no baseline, no control, attrition beyond tolerance) return `ANALYSIS_NOT_ESTIMABLE` — never fabricate p-values or silently substitute a weaker estimator.
|
|
20
|
+
4. **Graph Re-adjudication**: add a local Evidence Node (`EVD-LOCAL-*`), commit a new revision, re-run the tribunal.
|
|
21
|
+
|
|
22
|
+
## Output Contract
|
|
23
|
+
`analysis-run` + `dataset-manifest`; the graph revision carries the local node; the decision diff is produced by the update workflow.
|
|
24
|
+
|
|
25
|
+
## Visualization Sync
|
|
26
|
+
- `result.json` → `forest_plot_data`: one entry with `study_label` "Local Field Trial (DID)", outcome dimension from the trial, effect size = Hedges' g, CI bounds from the regression.
|
|
27
|
+
- `result.json` → `evidence`: one evidence object whose `relation_to_claim` follows the sign of δ.
|
|
28
|
+
- `evidence_graph.json`: re-export after adding the `EVD-LOCAL-*` node.
|
|
29
|
+
|
|
30
|
+
## Quality Gates
|
|
31
|
+
- [ ] 溯源、哈希、缺失率在分析前记录。
|
|
32
|
+
- [ ] 分析严格按预注册计划执行。
|
|
33
|
+
- [ ] 本地结果作为"一项研究"并入,不覆盖既有证据。
|
|
34
|
+
|
|
35
|
+
## Anti-Patterns
|
|
36
|
+
- 事后改阈值;把不显著说成无效果;用本地单点结果推翻既有证据体。
|
|
37
|
+
|
|
38
|
+
## Worked Example
|
|
39
|
+
两班前后测数据 → DID δ=+0.21(不显著):报告为"本地未复现",不改变原裁决方向,仅下调适用性置信。
|
|
40
|
+
|
|
41
|
+
## References
|
|
42
|
+
- `scripts/did_regression.py`、`references/evaluation-design.md`、`skill/workflows/evaluate-and-update.md`
|
|
@@ -1,25 +1,48 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: ethics-review
|
|
3
|
-
description: "Evaluates trial designs, intervention protocols, and
|
|
3
|
+
description: "Evaluates trial designs, intervention protocols, and human-subject data collection against IRB and research ethics standards (education and other applied domains)."
|
|
4
|
+
capability: (engine-level gate; no deterministic capability)
|
|
4
5
|
---
|
|
5
6
|
|
|
6
|
-
# ethics-review — Research Ethics & IRB Compliance
|
|
7
|
+
# ethics-review — Research Ethics & IRB Compliance
|
|
7
8
|
|
|
8
9
|
## When to Use
|
|
9
|
-
Triggered
|
|
10
|
+
Triggered before finalising any quasi-experimental / DID field trial design involving human student cohorts, classroom telemetry, or control-group assignment.
|
|
10
11
|
|
|
11
|
-
##
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
12
|
+
## Inputs
|
|
13
|
+
- `intervention.json` / study-design 草案(阶段、人群、对照、数据采集范围)
|
|
14
|
+
- 数据采集清单(分数、日志、提示词、遥测)
|
|
15
|
+
|
|
16
|
+
## Process — Ethical Audit Checklist
|
|
17
|
+
1. **Control Group Harm Prevention**: the control group must not be deprived of essential learning opportunities (use delayed crossover or active alternatives).
|
|
18
|
+
2. **Participant Privacy & Telemetry Protection**: pseudonymise prompts, interaction logs, and outcome records; comply with the applicable regime (FERPA / GDPR or the domain's equivalent).
|
|
19
|
+
3. **Informed Consent & Voluntary Participation**: opt-out without academic penalty.
|
|
20
|
+
4. **Algorithmic Bias & Equity Check**: audit whether the tool introduces grading bias or accessibility barriers for underrepresented groups.
|
|
21
|
+
5. **Data Retention**: state what is stored, where, and for how long; commercial LLM endpoints must not retain prompts.
|
|
22
|
+
|
|
23
|
+
## Output Contract
|
|
24
|
+
Ethics review record attached to the study design; a non-passing review blocks the pilot.
|
|
16
25
|
|
|
17
|
-
## Output Schema
|
|
18
26
|
```json
|
|
19
27
|
{
|
|
20
28
|
"ethics_status": "APPROVED_WITH_CONDITIONS",
|
|
21
29
|
"irb_tier": "Exempt / Expedited Educational Research (Category 1)",
|
|
22
30
|
"privacy_safeguards": ["Anonymized student IDs", "Zero prompt retention on commercial LLM endpoints"],
|
|
23
|
-
"equity_protections": "Provide universal
|
|
31
|
+
"equity_protections": "Provide universal campus lab access to eliminate hardware disparities."
|
|
24
32
|
}
|
|
25
33
|
```
|
|
34
|
+
|
|
35
|
+
## Quality Gates
|
|
36
|
+
- [ ] 对照组的可接受替代方案已给出。
|
|
37
|
+
- [ ] 采集字段清单与去标识方式明确。
|
|
38
|
+
- [ ] 退出机制不产生学业惩罚。
|
|
39
|
+
- [ ] 设备/网络差异导致的公平性问题被处理。
|
|
40
|
+
|
|
41
|
+
## Anti-Patterns
|
|
42
|
+
- 以"教改豁免"跳过伦理审查;保留可回指到个人的提示词日志;用"照常上课"掩盖对照组机会剥夺。
|
|
43
|
+
|
|
44
|
+
## Worked Example
|
|
45
|
+
12 周准实验含对照班 → APPROVED_WITH_CONDITIONS:延迟交叉 + 匿名 ID + 关闭端点日志留存 + 统一机房访问。
|
|
46
|
+
|
|
47
|
+
## References
|
|
48
|
+
- `references/evaluation-design.md`、`references/social_science_pitfalls.md`、`skill/task-briefs/intervene.md`
|
|
@@ -1,19 +1,37 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: evidence-extraction
|
|
3
3
|
description: "Extracts fine-grained claims, effect sizes (Hedges g), sample sizes, and methodology variables from validated full-text sources."
|
|
4
|
+
capability: study_extraction + finding_extraction
|
|
4
5
|
---
|
|
6
|
+
|
|
5
7
|
# Evidence Extraction Skill
|
|
6
8
|
|
|
7
|
-
##
|
|
9
|
+
## When to Use
|
|
8
10
|
Trigger on fetched and validated source texts to perform claim-level feature and statistical extraction.
|
|
9
11
|
|
|
10
|
-
##
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
2. **Methodology Extraction**:
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
12
|
+
## Inputs
|
|
13
|
+
- `sources.jsonl`(全部 FETCH_VALID / 确认的 FETCH_PARTIAL)
|
|
14
|
+
- `fetch/` 清正文;必要时经 Sciverse `/content` 读取定位段落
|
|
15
|
+
|
|
16
|
+
## Process
|
|
17
|
+
1. **Statistical Extraction**: sample sizes (N_treatment, N_control); means/SDs; standardized effect size (Hedges' g, Cohen's d, Odds Ratio via `scripts/effect_calculator.py`); 95% CIs and p-values when reported.
|
|
18
|
+
2. **Methodology Extraction**: design type (RCT, quasi-experimental DID/PSM/RDD, correlational); outcome classification (task performance vs conceptual learning vs delayed retention).
|
|
19
|
+
3. **Locate precisely**: keep `source_location` (page/section/offset) so every claim can be re-opened.
|
|
20
|
+
|
|
21
|
+
## Output Contract
|
|
22
|
+
Evidence Objects per `schemas/evidence.schema.json` (V1 top-level, revision 1.1) into `evidence.jsonl`; graph projections use `schemas/v2/evidence-link.schema.json` (V2).
|
|
23
|
+
|
|
24
|
+
## Quality Gates
|
|
25
|
+
- [ ] 每行通过 evidence schema,枚举合法。
|
|
26
|
+
- [ ] 任务表现与学习效果记录分离。
|
|
27
|
+
- [ ] 缺失统计量保持缺失(禁止由显著性反推效应量)。
|
|
28
|
+
- [ ] 每条记录可定位回原文。
|
|
29
|
+
|
|
30
|
+
## Anti-Patterns
|
|
31
|
+
- 把 `relation_to_claim` 写到研究本体;把即测分数当保持/迁移;从摘要估算效应量。
|
|
32
|
+
|
|
33
|
+
## Worked Example
|
|
34
|
+
PNAS 2025 三臂 RCT → 三条 evidence:练习表现(task_performance, support)、独立考试(learning, contradict)、护栏组(learning, support),各带 N 与效应量。
|
|
35
|
+
|
|
36
|
+
## References
|
|
37
|
+
- `references/evidence-quality.md`、`references/effect_size_formulas.md`、`skill/task-briefs/extract.md`
|