eduevidence 6.0.0 → 6.2.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CONTRIBUTING.md +105 -0
- package/README.md +93 -38
- package/README.zh-CN.md +26 -6
- package/SKILL.md +11 -2
- package/assets/readme/landing-tour.gif +0 -0
- package/assets/readme/studio-tour.gif +0 -0
- package/bin/eduevidence.js +2 -1
- package/docs/architecture.md +319 -43
- package/docs/demo-workplace-ai.md +1 -1
- package/docs/install-guide.md +1 -1
- package/docs/orchestration-role-model.md +1 -1
- package/docs/release-closeout/README.md +1 -1
- package/docs/sciverse-api.md +125 -0
- package/eduevidence_cli.py +10 -0
- package/engine/decision_policy.py +96 -0
- package/engine/evidence_graph.py +14 -10
- package/engine/gaps.py +42 -22
- package/engine/ids.py +2 -0
- package/engine/library.py +6 -2
- package/engine/living.py +34 -4
- package/engine/migration.py +88 -3
- package/engine/orchestration.py +5 -5
- package/engine/paths.py +2 -0
- package/engine/pilot.py +34 -32
- package/engine/taxonomy.py +211 -0
- package/engine/tribunal.py +43 -31
- package/engine/versions.py +1 -1
- package/examples/ai-coding-assistant-evidence/EduEvidence_Report.html +1360 -146
- package/examples/ai-coding-assistant-evidence/artifact_manifest.json +3 -3
- package/examples/ai-coding-assistant-evidence/citation_check.json +1 -1
- package/examples/ai-coding-assistant-evidence/final_verdict.json +107 -0
- package/examples/ai-coding-assistant-evidence/gate_report.json +101 -0
- package/examples/ai-coding-assistant-evidence/report_spec.json +23 -12
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_academic.html +447 -127
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_claude.html +447 -127
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_datalab-dark.html +447 -127
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_datalab.html +447 -127
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_presentation.html +447 -127
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_academic.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_claude.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_datalab-dark.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_datalab.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_presentation.html +1360 -146
- package/examples/ai-coding-assistant-evidence/result.json +13 -9
- package/examples/ai-coding-assistant-evidence/result.zh.json +45 -41
- package/examples/ai-coding-assistant-evidence/skeptic.json +72 -0
- package/examples/ai-coding-assistant-evidence/verdict.json +6 -2
- package/examples/spaced-retrieval-practice/applicability.json +14 -0
- package/examples/spaced-retrieval-practice/artifact_manifest.json +15 -0
- package/examples/spaced-retrieval-practice/claims.jsonl +3 -0
- package/examples/spaced-retrieval-practice/evidence.jsonl +6 -0
- package/examples/spaced-retrieval-practice/final_verdict.json +93 -0
- package/examples/spaced-retrieval-practice/frame.json +58 -0
- package/examples/spaced-retrieval-practice/gate_report.json +101 -0
- package/examples/spaced-retrieval-practice/methodology.json +78 -0
- package/examples/spaced-retrieval-practice/report_spec.json +212 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_academic.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_claude.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_datalab-dark.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_datalab.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_presentation.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_academic.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_claude.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_datalab-dark.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_datalab.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_presentation.html +2728 -0
- package/examples/spaced-retrieval-practice/result.json +942 -0
- package/examples/spaced-retrieval-practice/result.zh.json +942 -0
- package/examples/spaced-retrieval-practice/skeptic.json +70 -0
- package/examples/spaced-retrieval-practice/sources.jsonl +7 -0
- package/examples/spaced-retrieval-practice/verdict.json +93 -0
- package/examples/workplace-ai-assistant/artifact_manifest.json +15 -0
- package/examples/workplace-ai-assistant/claims.jsonl +4 -4
- package/examples/workplace-ai-assistant/evidence.jsonl +4 -4
- package/examples/workplace-ai-assistant/evidence_graph.json +15 -15
- package/examples/workplace-ai-assistant/final_verdict.json +78 -0
- package/examples/workplace-ai-assistant/gate_report.json +101 -0
- package/examples/workplace-ai-assistant/report_spec.json +209 -40
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_academic.html +435 -105
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_claude.html +435 -105
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_datalab-dark.html +435 -105
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_datalab.html +435 -105
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_presentation.html +435 -105
- package/examples/workplace-ai-assistant/reports-5themes/report_academic.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_claude.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_datalab-dark.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_datalab.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_presentation.html +2814 -0
- package/examples/workplace-ai-assistant/result.json +82 -20
- package/examples/workplace-ai-assistant/result.zh.json +82 -20
- package/examples/workplace-ai-assistant/skeptic.json +72 -0
- package/examples/workplace-ai-assistant/verdict.json +36 -10
- package/integrations/agent_mcp.py +2 -2
- package/package.json +12 -3
- package/pyproject.toml +4 -3
- package/references/report-copy-style.md +67 -0
- package/references/retrieval-compliance.md +75 -0
- package/references/retrieval-protocol.md +20 -0
- package/retrieval/audit.py +27 -3
- package/retrieval/fetch.py +96 -0
- package/retrieval/sciverse.py +398 -0
- package/retrieval/search.py +47 -7
- package/schemas/applicability.schema.json +94 -0
- package/schemas/chart-spec.schema.json +10 -3
- package/schemas/evidence.schema.json +316 -43
- package/schemas/fetch-result.schema.json +2 -1
- package/schemas/report-result.schema.json +3 -3
- package/schemas/report-spec.schema.json +98 -100
- package/schemas/skeptic.schema.json +86 -0
- package/schemas/source.schema.json +21 -2
- package/schemas/v2/finding.schema.json +5 -1
- package/schemas/v2/methodology-audit.schema.json +5 -1
- package/schemas/v2/outcome.schema.json +28 -5
- package/schemas/v2/study.schema.json +5 -1
- package/schemas/vNext/autoevolve-session.schema.json +34 -1
- package/schemas/vNext/eval-snapshot.schema.json +77 -1
- package/schemas/vNext/execution-plan.schema.json +50 -1
- package/schemas/vNext/gap-priority.schema.json +54 -1
- package/schemas/vNext/negative-search-record.schema.json +68 -1
- package/schemas/vNext/research-iteration.schema.json +87 -1
- package/schemas/vNext/research-strategy.schema.json +62 -1
- package/schemas/vNext/skill-experiment.schema.json +90 -1
- package/schemas/vNext/task-spec.schema.json +156 -1
- package/schemas/vNext/worker-result.schema.json +60 -1
- package/schemas/verdict.schema.json +164 -28
- package/scripts/build_esl_artifacts.py +2 -2
- package/scripts/build_report_variants.py +18 -2
- package/scripts/build_result.py +74 -9
- package/scripts/check_package_parity.py +85 -0
- package/scripts/check_protocol_alignment.py +375 -0
- package/scripts/check_versioned_schemas.py +254 -0
- package/scripts/claim_audit.py +13 -8
- package/scripts/compute_confidence.py +10 -0
- package/scripts/did_regression.py +12 -2
- package/scripts/evidence_score.py +5 -2
- package/scripts/generate_new_projects.py +4 -4
- package/scripts/orchestrator.py +120 -24
- package/scripts/pre_verdict_gate.py +224 -26
- package/scripts/quickstart.py +18 -2
- package/scripts/run_workspace.py +7 -1
- package/scripts/skill_payload.py +4 -1
- package/scripts/test_adversarial_empirical.py +26 -19
- package/scripts/validate_schema.py +31 -1
- package/skill/agents/evaluation-designer.md +20 -4
- package/skill/agents/evidence-analyst.md +19 -3
- package/skill/agents/evidence-judge.md +50 -2
- package/skill/agents/evidence-retriever.md +20 -3
- package/skill/agents/intervention-designer.md +20 -4
- package/skill/agents/method-reviewer.md +18 -2
- package/skill/agents/{education-planner.md → research-planner.md} +19 -3
- package/skill/agents/skeptic.md +18 -2
- package/skill/roles/registry.yaml +11 -11
- package/skill/sub-skills/aihot-trend-analysis/SKILL.md +28 -9
- package/skill/sub-skills/contradiction-analysis/SKILL.md +31 -11
- package/skill/sub-skills/data-analysis/SKILL.md +34 -15
- package/skill/sub-skills/ethics-review/SKILL.md +33 -10
- package/skill/sub-skills/evidence-extraction/SKILL.md +29 -11
- package/skill/sub-skills/evidence-review/SKILL.md +31 -12
- package/skill/sub-skills/gap-analysis/SKILL.md +31 -9
- package/skill/sub-skills/literature-review/SKILL.md +35 -14
- package/skill/sub-skills/methodology-audit/SKILL.md +29 -12
- package/skill/sub-skills/report-generation/SKILL.md +28 -0
- package/skill/sub-skills/research-planning/SKILL.md +41 -14
- package/skill/sub-skills/study-design/SKILL.md +30 -9
- package/skill/task-briefs/adjudicate.md +32 -7
- package/skill/task-briefs/applicability.md +37 -2
- package/skill/task-briefs/audit.md +32 -7
- package/skill/task-briefs/challenge.md +34 -5
- package/skill/task-briefs/evaluate.md +30 -5
- package/skill/task-briefs/extract.md +31 -8
- package/skill/task-briefs/frame.md +39 -10
- package/skill/task-briefs/intervene.md +32 -6
- package/skill/task-briefs/present.md +32 -8
- package/skill/task-briefs/projection.md +36 -2
- package/skill/task-briefs/retrieve.md +36 -6
- package/skill/workflows/decision-and-pilot.md +76 -1
- package/skill/workflows/evaluate-and-update.md +83 -0
- package/skill/workflows/evidence-review.md +104 -0
- package/visualization/eduevidence-report/scripts/build_figures.py +25 -3
- package/visualization/eduevidence-report/scripts/build_infographics.py +5 -1
- package/visualization/eduevidence-report/scripts/build_report.py +512 -65
- package/visualization/eduevidence-report/scripts/charts_data.py +2 -0
- package/visualization/eduevidence-report/scripts/lieflat_engine.py +349 -38
- package/visualization/eduevidence-report/scripts/zh_labels.py +80 -1
- package/web/architecture.html +14885 -0
- package/web/studio/assets/index-B8tkF44Q.css +1 -0
- package/web/studio/index.html +2 -2
- package/web/studio/assets/index-CzXocaGv.css +0 -1
- /package/web/studio/assets/{index-pa7jD7n4.js → index-CQ6Keoyc.js} +0 -0
|
@@ -1,18 +1,37 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: evidence-review
|
|
3
3
|
description: "Synthesizes extracted claims and evidence nodes into the project Evidence Graph with meta-analysis pooling."
|
|
4
|
+
capability: claim_linking + evidence_synthesis
|
|
4
5
|
---
|
|
6
|
+
|
|
5
7
|
# Evidence Review Skill
|
|
6
8
|
|
|
7
|
-
##
|
|
8
|
-
Trigger to aggregate
|
|
9
|
-
|
|
10
|
-
##
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
9
|
+
## When to Use
|
|
10
|
+
Trigger to aggregate validated evidence nodes into the project's single source of truth (SSOT) Claim Graph and execute quantitative synthesis.
|
|
11
|
+
|
|
12
|
+
## Inputs
|
|
13
|
+
- `evidence.jsonl` + `methodology.json`
|
|
14
|
+
- 项目 Evidence Graph 当前 revision
|
|
15
|
+
|
|
16
|
+
## Process
|
|
17
|
+
1. **Evidence Graph State Machine**: build the directed graph of claims, sources, findings; determine claim status (SUPPORTED / CONTRADICTED / MIXED / UNCERTAIN).
|
|
18
|
+
2. **Meta-Analysis Pooling** (`engine/meta_analysis.py`): fixed-effect inverse-variance pooling and DerSimonian–Laird random-effects pooling; Cochran's Q, df, I².
|
|
19
|
+
3. **Publication Bias & Robustness** (`engine/bias.py` / `engine/robustness.py`): Egger regression, Rosenthal fail-safe N, leave-one-out sensitivity.
|
|
20
|
+
4. **Counting rule**: pool by independent **study**, never by finding count.
|
|
21
|
+
|
|
22
|
+
## Output Contract
|
|
23
|
+
Graph revision (append-only) plus synthesis artifacts; projections (`result.json`, HTML) are derived, never the source of truth.
|
|
24
|
+
|
|
25
|
+
## Quality Gates
|
|
26
|
+
- [ ] 按独立研究计数,非按 finding 计数。
|
|
27
|
+
- [ ] 异质性(I²、Q)与偏倚检验结果一并报告。
|
|
28
|
+
- [ ] 三个方向列未被合并隐藏。
|
|
29
|
+
|
|
30
|
+
## Anti-Patterns
|
|
31
|
+
- 同一研究的多个 outcome 当作多项独立证据;对不可合并的结果强行做标准化合并。
|
|
32
|
+
|
|
33
|
+
## Worked Example
|
|
34
|
+
8 来源 12 条发现 → 仅部分可合并;不可合并者以叙述式证据矩阵呈现并标注原因。
|
|
35
|
+
|
|
36
|
+
## References
|
|
37
|
+
- `references/evidence-quality.md`、`docs/evidence-synthesis.md`、`engine/meta_analysis.py`
|
|
@@ -1,25 +1,47 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: gap-analysis
|
|
3
|
-
description: "Identifies population, measurement, and methodological gaps in the Evidence Graph and diagnoses cross-study empirical contradictions
|
|
3
|
+
description: "Identifies population, measurement, and methodological gaps in the Evidence Graph and diagnoses cross-study empirical contradictions."
|
|
4
|
+
capability: knowledge_gap_detection
|
|
4
5
|
---
|
|
5
6
|
|
|
6
|
-
# gap-analysis — Research Gap Discovery & Contradiction Lens
|
|
7
|
+
# gap-analysis — Research Gap Discovery & Contradiction Lens
|
|
7
8
|
|
|
8
9
|
## When to Use
|
|
9
|
-
Triggered after
|
|
10
|
+
Triggered after evidence extraction and meta-analysis, before study design: trial designs must be grounded on verified empirical gaps, never on generic templates.
|
|
10
11
|
|
|
11
|
-
##
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
12
|
+
## Inputs
|
|
13
|
+
- Evidence Graph(含 claim 状态与 outcome 维度)
|
|
14
|
+
- 矛盾诊断(方向冲突、异质性来源)
|
|
15
|
+
|
|
16
|
+
## Process
|
|
17
|
+
1. **Measurement & Retention Gap Audit**: if evidence only measures immediate task speed, flag the missing delayed unassisted retention measurement.
|
|
18
|
+
2. **Population Heterogeneity Audit**: check whether studies cover only elite CS majors or introductory cohorts; flag advanced-transfer gaps.
|
|
19
|
+
3. **Contradiction Lens**: when directions diverge (g > +0.3 vs g < −0.1), isolate the moderating variable (e.g. scaffolded vs unguided use).
|
|
20
|
+
4. **Priority**: rank gaps by decision-materiality so study design is not spent on immaterial questions.
|
|
21
|
+
|
|
22
|
+
## Output Contract
|
|
23
|
+
`GapNode` list written directly to the SSOT Evidence Graph, each with `gap_id`, `gap_type`, `description`, `target_outcome`, `recommended_trial_design`.
|
|
15
24
|
|
|
16
|
-
## Output Contract (`GapNode` list written directly to SSOT `EvidenceGraph`)
|
|
17
25
|
```json
|
|
18
26
|
{
|
|
19
27
|
"gap_id": "GAP-RETENTION-001",
|
|
20
28
|
"gap_type": "Measurement/Retention Gap",
|
|
21
29
|
"description": "Lack of 12-week longitudinal retention data measuring unassisted transfer in CS1.",
|
|
22
30
|
"target_outcome": "Delayed Unassisted Problem Solving",
|
|
23
|
-
"recommended_trial_design": "12-
|
|
31
|
+
"recommended_trial_design": "12-week cluster randomized trial with 4-week delayed post-test without AI access"
|
|
24
32
|
}
|
|
25
33
|
```
|
|
34
|
+
|
|
35
|
+
## Quality Gates
|
|
36
|
+
- [ ] 每个 gap 绑定具体未满足的测量/人群/方法维度。
|
|
37
|
+
- [ ] gap 与决策材料性关联,而非"文献少"。
|
|
38
|
+
- [ ] 不把"论文数量少"当作研究缺口。
|
|
39
|
+
|
|
40
|
+
## Anti-Patterns
|
|
41
|
+
- 用"研究不足"概括一切;把已有证据能回答的问题标成缺口。
|
|
42
|
+
|
|
43
|
+
## Worked Example
|
|
44
|
+
现有证据多为即测任务表现 → 产出 GAP-RETENTION-001,约束后续 StudyDesign 必须包含无 AI 的延迟后测。
|
|
45
|
+
|
|
46
|
+
## References
|
|
47
|
+
- `engine/gap_lens.py`、`engine/gaps.py`、`references/scientific-invariants.md`
|
|
@@ -1,21 +1,42 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: literature-review
|
|
3
|
-
description: "Executes multi-source academic and web retrieval across OpenAlex, Semantic Scholar, CrossRef, AIHot, AgentSearch, and user-configured providers
|
|
3
|
+
description: "Executes multi-source academic and web retrieval across OpenAlex, Semantic Scholar, CrossRef, AIHot, AgentSearch, Sciverse, and user-configured providers."
|
|
4
|
+
capability: literature_search + counter_evidence_search + source_fetch + source_validation
|
|
4
5
|
---
|
|
6
|
+
|
|
5
7
|
# Literature Review Skill
|
|
6
8
|
|
|
7
|
-
##
|
|
9
|
+
## When to Use
|
|
8
10
|
Trigger after research intent is established to gather candidate empirical studies, peer-reviewed papers, and verified grey literature.
|
|
9
11
|
|
|
10
|
-
##
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
12
|
+
## Inputs
|
|
13
|
+
- `frame.json`(检索边界、纳排标准)
|
|
14
|
+
- 检索预算(S/M/L 决定)与可选 key:`SCIVERSE_API_TOKEN` / `TAVILY_API_KEY` / `BRAVE_API_KEY`
|
|
15
|
+
|
|
16
|
+
## Process
|
|
17
|
+
1. **Plan first** — 写 `SearchPlan`(core / expansion / **counter_evidence**),经 `retrieval/audit.py` 执行并导出审计四件套。
|
|
18
|
+
2. **Zero-Config Academic Providers**: OpenAlex (250M+ works with DOIs), Semantic Scholar (graph citations + abstracts), CrossRef (DOI registry), AIHot (real-time AI/EdTech feed), AgentSearch / ArXiv.
|
|
19
|
+
3. **Key-based academic channel**: Sciverse (`SCIVERSE_API_TOKEN`) — `/meta-search` 产出 Source 级命中;`/agentic-search` 产出 chunk 定位子,必须经 `/content` 读原文并过校验门,定位写入 `chunks.jsonl`。
|
|
20
|
+
4. **Configured web providers**: Tavily, Brave.
|
|
21
|
+
5. **Fetch & Validation**: 候选 URL 走 `retrieval/fetch.py` 降级链并由 `retrieval/validate.py` 校验;严格拒绝 snippet 幻觉。
|
|
22
|
+
6. **Compliance**: 遵守 `references/retrieval-compliance.md`(robots / 限速 / paywall / 署名与缓存)。
|
|
23
|
+
|
|
24
|
+
## Output Contract
|
|
25
|
+
- `sources.jsonl`(`schemas/source.schema.json`)+ `fetch/`(raw + clean + provenance + fallback_chain)
|
|
26
|
+
- 审计导出:`search-provenance.json` / `search-attempts.jsonl` / `source-screening.csv` / `exclusion-log.csv`
|
|
27
|
+
- Sciverse 定位:`chunks.jsonl`(`locator_state=discovery_only_requires_content_fetch`)
|
|
28
|
+
|
|
29
|
+
## Quality Gates
|
|
30
|
+
- [ ] 反方查询独立构造且计数 > 0。
|
|
31
|
+
- [ ] 每条来源为 FETCH_VALID 或经确认的 FETCH_PARTIAL。
|
|
32
|
+
- [ ] 无 DOI/URL 记录标 `needs_manual_location`,未伪造定位。
|
|
33
|
+
- [ ] 排除记录写明原因。
|
|
34
|
+
|
|
35
|
+
## Anti-Patterns
|
|
36
|
+
- 用支持证据的检索式找反证;把 abstract 当证据内容;为凑数填充无关来源。
|
|
37
|
+
|
|
38
|
+
## Worked Example
|
|
39
|
+
`python scripts/search_provenance.py "first-year CS generative AI coding assistant" --out runs/x/provenance --concept "AI coding assistant"` → 审计四件套 + 候选来源表,随后逐条 fetch/validate。
|
|
40
|
+
|
|
41
|
+
## References
|
|
42
|
+
- `references/retrieval-protocol.md`、`references/retrieval-compliance.md`、`references/source-validity.md`、`skill/task-briefs/retrieve.md`
|
|
@@ -1,20 +1,37 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: methodology-audit
|
|
3
3
|
description: "Audits empirical studies against WWC 5.0, GRADE risk-of-bias frameworks, and social science pitfalls."
|
|
4
|
+
capability: methodology_appraisal
|
|
4
5
|
---
|
|
6
|
+
|
|
5
7
|
# Methodology Audit Skill (Methodology Tribunal)
|
|
6
8
|
|
|
7
|
-
##
|
|
9
|
+
## When to Use
|
|
8
10
|
Trigger before claim synthesis to evaluate threats to internal and external validity, applying methodological confidence scoring.
|
|
9
11
|
|
|
10
|
-
##
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
12
|
+
## Inputs
|
|
13
|
+
- `evidence.jsonl` + `fetch/` 原文
|
|
14
|
+
- `references/wwc_standards.md`、`references/grade_framework.md`、`references/social_science_pitfalls.md`
|
|
15
|
+
|
|
16
|
+
## Process
|
|
17
|
+
1. **WWC 5.0 Rating**: Tier 1 (meets standards without reservations — clean RCT, low attrition), Tier 2 (with reservations — QED with baseline equivalence), Tier 3 (correlational / promising).
|
|
18
|
+
2. **Social Science Pitfalls**: task performance ≠ genuine learning; short-term score ≠ retention (4+ weeks); AI-assisted performance ≠ independent transfer; correlation ≠ causation.
|
|
19
|
+
3. **GRADE Certainty**: High / Moderate / Low / Very Low at the body-of-evidence level.
|
|
20
|
+
4. **15-item checklist** in fixed order (control_group … dropout), each `met|partial|missing|not_applicable`.
|
|
21
|
+
|
|
22
|
+
## Output Contract
|
|
23
|
+
`methodology.json` per `schemas/methodology.schema.json`, including `task_vs_learning_guard.equates_task_with_learning`.
|
|
24
|
+
|
|
25
|
+
## Quality Gates
|
|
26
|
+
- [ ] 15 项齐全且仅用四值枚举。
|
|
27
|
+
- [ ] guard 结论明确。
|
|
28
|
+
- [ ] 审计只判"证据是否成立",不判"证据说什么"。
|
|
29
|
+
|
|
30
|
+
## Anti-Patterns
|
|
31
|
+
- 用样本量大掩盖无对照;把相关性研究列入因果结论;缺项直接记 `met`。
|
|
32
|
+
|
|
33
|
+
## Worked Example
|
|
34
|
+
某准实验无前测等价性 → Tier 2 (with reservations) + `pre_test: missing`;其结论只可用于"提示可能",不进强支持列。
|
|
35
|
+
|
|
36
|
+
## References
|
|
37
|
+
- `references/methodology-audit.md`、`references/tribunal-policy.md`、`skill/agents/method-reviewer.md`、`skill/task-briefs/audit.md`
|
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: report-generation
|
|
3
3
|
description: "Renders 5 baked-theme single-file bilingual HTML reports, executive Visual Briefs, and Markdown reports, powered by Lieflat Charts editorial visualization standards and AI-composed, data-driven chart galleries."
|
|
4
|
+
capability: report_projection + report_rendering
|
|
4
5
|
---
|
|
5
6
|
# Report Generation Skill
|
|
6
7
|
|
|
@@ -55,3 +56,30 @@ Dark themes draw on the theme's `card_bg`; text contrast is checked against the
|
|
|
55
56
|
|
|
56
57
|
## 4. Web Studio Sync
|
|
57
58
|
The Local Web Studio (`scripts/dashboard_server.py`) serves the baked HTML reports and Lieflat figures directly.
|
|
59
|
+
|
|
60
|
+
## Inputs
|
|
61
|
+
- `result.json` / `result.zh.json`(叙述字段先过语言门禁 `check_language_parallel`)
|
|
62
|
+
- 当前 Graph Revision 与 decision snapshot 标识
|
|
63
|
+
|
|
64
|
+
## Quality Gates
|
|
65
|
+
- [ ] 双语语义对齐(数字 / ID / 枚举 / URL 不变)。
|
|
66
|
+
- [ ] 渲染完整性门通过:显示数值可回溯到 `result.json`,探针无 `REPORT_INVALID`。
|
|
67
|
+
- [ ] 布局不变量通过 `scripts/lint_report_layout.py`(390 / 768 / 1280 × brief/full)。
|
|
68
|
+
- [ ] 溯源表格保留(来源表 / 证据矩阵 / Claim Trace),图表只作补充。
|
|
69
|
+
- [ ] `artifact_manifest.json` 记录来源 revision 与各产物哈希。
|
|
70
|
+
|
|
71
|
+
## Anti-Patterns
|
|
72
|
+
- 由模型写入图表数值(数值必须来自 `scripts/charts_data.py` 提取器)。
|
|
73
|
+
- 为了"更可视化"而删除可核验表格;用占位或虚构数据补齐缺失图表。
|
|
74
|
+
- 报告与快照不一致时改报告不改结论;在 HTML 内提供运行时换肤。
|
|
75
|
+
|
|
76
|
+
## Failure Handling
|
|
77
|
+
| 失败 | 处理 |
|
|
78
|
+
|---|---|
|
|
79
|
+
| `REPORT_INVALID` | 阻断发布并重跑渲染,不带缺陷投放。 |
|
|
80
|
+
| 双语不对齐 | 修复 `result.zh.json` 后重新烘焙。 |
|
|
81
|
+
| 数据不足以支撑某图 | 抑制该图并记录原因,不补造数据。 |
|
|
82
|
+
| 快照缺失 | 回到上游科学阶段补齐;禁止用默认值生成报告。 |
|
|
83
|
+
|
|
84
|
+
## References
|
|
85
|
+
- `skill/task-briefs/present.md`、`skill/task-briefs/projection.md`、`visualization/eduevidence-report/references/layout-constraints.md`
|
|
@@ -1,21 +1,48 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: research-planning
|
|
3
|
-
description: "Extracts the structured
|
|
3
|
+
description: "Extracts the structured Research Frame (PICO + decision target, in the selected domain's vocabulary) from user natural language and determines execution mode (S/M/L) via the Complexity Gate."
|
|
4
|
+
capability: research_framing
|
|
4
5
|
---
|
|
6
|
+
|
|
5
7
|
# Research Planning Skill
|
|
6
8
|
|
|
7
|
-
##
|
|
8
|
-
Trigger this skill when the user initiates a new empirical
|
|
9
|
+
## When to Use
|
|
10
|
+
Trigger this skill when the user initiates a new empirical inquiry — education, policy, organisational practice, or any other registered domain — or updates the research scope. This is the **Frame** stage of the Canonical Protocol (`docs/architecture.md`).
|
|
11
|
+
|
|
12
|
+
## Inputs
|
|
13
|
+
- 用户原始问题(自然语言,可能含隐含假设)
|
|
14
|
+
- 可选结构化线索:population / intervention / comparison / outcomes / constraints / depth
|
|
15
|
+
- 领域(决定 frame 词汇):education(learner/course)/ policy(decision_object/population/stakeholders),或其他注册在 `domains/` 下的领域
|
|
9
16
|
|
|
10
|
-
##
|
|
11
|
-
1. **Frame Formulation (
|
|
12
|
-
- **Population (P)**:
|
|
13
|
-
- **Intervention (I)**:
|
|
17
|
+
## Process
|
|
18
|
+
1. **Frame Formulation (Research Frame,字段随领域而定)**
|
|
19
|
+
- **Population (P)**: Who or what the decision acts on — learners and courses in education, affected groups and stakeholders in policy, customers or staff in organisations. Use the field vocabulary of the selected domain.
|
|
20
|
+
- **Intervention (I)**: The specific practice under decision — a teaching method, tool, AI system, curriculum change, regulation, process or programme.
|
|
14
21
|
- **Comparison (C)**: Active control, business-as-usual, or non-intervention baseline.
|
|
15
|
-
- **Outcomes (O)**: Primary and secondary outcome metrics
|
|
16
|
-
- **
|
|
17
|
-
2. **Complexity Gating (S/M/L)**
|
|
18
|
-
- S (Quick Fact):
|
|
19
|
-
- M (Standard Review):
|
|
20
|
-
- L (Deep Causal Cycle):
|
|
21
|
-
3. **
|
|
22
|
+
- **Outcomes (O)**: Primary and secondary outcome metrics; task performance and delayed transfer must stay separate.
|
|
23
|
+
- **decision target + scope + inclusion/exclusion criteria**: required by the selected domain's frame schema.
|
|
24
|
+
2. **Complexity Gating (S/M/L)** — `scripts/complexity_gate.py` is the single authority:
|
|
25
|
+
- S (Quick Fact): fact-check a single claim (k=3–5).
|
|
26
|
+
- M (Standard Review): full multi-source evidence review (k=8–15).
|
|
27
|
+
- L (Deep Causal Cycle): synthesis + trial design + empirical DID regression.
|
|
28
|
+
3. **Missing inputs**: ask the user (`NEEDS_USER_CONTEXT`); never invent a default population, comparator, or outcome.
|
|
29
|
+
|
|
30
|
+
## Output Contract
|
|
31
|
+
- Frame body validates against the selected domain's frame schema (`domains/<id>/manifest.json` → `frame_schema`): `schemas/education-frame.schema.json` for education, `domains/policy/frame.schema.json` for policy.
|
|
32
|
+
- Research-intent projection uses `schemas/v2/research-intent.schema.json` (V2 contract).
|
|
33
|
+
- **No intervention advice may be produced before framing completes** — the domain layer's first gate.
|
|
34
|
+
|
|
35
|
+
## Quality Gates
|
|
36
|
+
- [ ] PICO 五要素齐全,comparator 明确,且字段词汇与所选领域一致。
|
|
37
|
+
- [ ] primary outcome 与 secondary/risk outcome 分离。
|
|
38
|
+
- [ ] scope 与纳排标准可操作,非"相关研究"式表述。
|
|
39
|
+
- [ ] 随 frame 输出 S/M/L 等级建议。
|
|
40
|
+
|
|
41
|
+
## Anti-Patterns
|
|
42
|
+
- 把"相关研究"当作 scope;把任务完成速度当作学习效果;为缺失人群编造默认值;frame 未完成即给干预建议;把某个领域的字段强制套到另一个领域。
|
|
43
|
+
|
|
44
|
+
## Worked Example
|
|
45
|
+
输入"大一 C 语言课程要不要允许用 AI 编程助手?" → education 域:population=CS1 学生;intervention=生成式 AI 编程助手(含护栏);comparison=无 AI 常规教学;primary outcome=独立考试成绩(learning);decision_target=teaching_decision。同一协议在 policy 域写成:population=客服团队与受影响客户;intervention=带人工审核的 AI 助手;decision_object=adopt。
|
|
46
|
+
|
|
47
|
+
## References
|
|
48
|
+
- `schemas/education-frame.schema.json`、`references/education-framing.md`、`references/outcome-taxonomy.md`、`skill/task-briefs/frame.md`
|
|
@@ -1,16 +1,37 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: study-design
|
|
3
3
|
description: "Generates pre-registered quasi-experimental DID and RCT trial designs to fill identified knowledge gaps."
|
|
4
|
+
capability: study_design + measurement_design
|
|
4
5
|
---
|
|
6
|
+
|
|
5
7
|
# Study Design Skill
|
|
6
8
|
|
|
7
|
-
##
|
|
8
|
-
Trigger
|
|
9
|
+
## When to Use
|
|
10
|
+
Trigger when the Evidence Graph yields MIXED, UNCERTAIN, or INSUFFICIENT outcomes — converting knowledge gaps into actionable empirical protocols. **A new design requires an explicit evidence-grounded KnowledgeGap ID.**
|
|
11
|
+
|
|
12
|
+
## Inputs
|
|
13
|
+
- `GapNode`(`GAP-*`)+ claim 状态
|
|
14
|
+
- 目标人群、可用课时、伦理审查结论
|
|
15
|
+
|
|
16
|
+
## Process
|
|
17
|
+
1. **Trial Specification**: sampling frame; cluster randomization or matched control classes; 4-phase rollout — baseline pre-test (W1), treatment exposure (W2–W9), immediate post-test (W10), 4-week delayed transfer test (W14).
|
|
18
|
+
2. **Pre-Registration Manifest**: power calculation (e.g. N > 120 for power ≥ 0.80 at g = 0.35); primary/secondary instruments; stopping rules; pre-specified regression model.
|
|
19
|
+
3. **Measurement plan**: hand over to evaluation design so thresholds are fixed before enrolment.
|
|
20
|
+
|
|
21
|
+
## Output Contract
|
|
22
|
+
StudyDesign + measurement plan validated against the grounded-study gate (`scripts/complexity_gate.py` + engine study-design checks). "Few papers found" alone never authorises a pilot.
|
|
23
|
+
|
|
24
|
+
## Quality Gates
|
|
25
|
+
- [ ] 设计引用显式 KnowledgeGap ID。
|
|
26
|
+
- [ ] 功效计算与停止规则齐备。
|
|
27
|
+
- [ ] 包含无 AI 的延迟迁移测量。
|
|
28
|
+
- [ ] 伦理审查已通过。
|
|
29
|
+
|
|
30
|
+
## Anti-Patterns
|
|
31
|
+
- 因"文献少"直接开研究;事后补功效计算;把即测当作迁移。
|
|
32
|
+
|
|
33
|
+
## Worked Example
|
|
34
|
+
GAP-RETENTION-001 → 12 周集群随机设计:W1 前测 / W10 即测 / W14 无 AI 迁移测,N≈160,预先注册分析模型。
|
|
9
35
|
|
|
10
|
-
##
|
|
11
|
-
|
|
12
|
-
- Sampling frame, cluster randomization / matched control classrooms.
|
|
13
|
-
- 4-phase rollout: Baseline pre-test (W1), Treatment exposure (W2-W9), Immediate post-test (W10), 4-week delayed transfer test (W14).
|
|
14
|
-
2. **Pre-Registration Manifest**:
|
|
15
|
-
- Statistical power calculation (N > 120 for power >= 0.80 at effect size g = 0.35).
|
|
16
|
-
- Primary/secondary outcome instruments, stopping rules, and pre-specified regression model.
|
|
36
|
+
## References
|
|
37
|
+
- `engine/study_design.py`、`references/intervention-design.md`、`skill/sub-skills/ethics-review/SKILL.md`
|
|
@@ -4,14 +4,39 @@
|
|
|
4
4
|
整合 Frame + Evidence Matrix + Skeptic Findings + Method Reviews,产出四态裁决
|
|
5
5
|
(ADOPT/PILOT/REJECT/INSUFFICIENT EVIDENCE)+ Can/Cannot Claim + 证据边界。
|
|
6
6
|
|
|
7
|
-
##
|
|
8
|
-
- frame.json
|
|
7
|
+
## 前置输入
|
|
8
|
+
- `frame.json`、`evidence.jsonl`、`skeptic.json`、`methodology.json`
|
|
9
|
+
- 确定性置信度脚本 `scripts/compute_confidence.py` 与裁决前闸门 `scripts/pre_verdict_gate.py`
|
|
9
10
|
|
|
10
11
|
## 产出
|
|
11
|
-
- raw_verdict.json
|
|
12
|
-
→ final_verdict.json
|
|
12
|
+
- `raw_verdict.json`(模型裁决)→ orchestrator 跑 Pre-Verdict Gate + 确定性置信度
|
|
13
|
+
→ `final_verdict.json`(`schemas/verdict.schema.json`)。
|
|
13
14
|
|
|
14
|
-
##
|
|
15
|
-
- decision_rationale 必须为 ≤4 句面向读者的流畅散文(en/zh 分写);
|
|
15
|
+
## 执行规则(人话化硬标准,第一页决策语言)
|
|
16
|
+
- `decision_rationale` 必须为 ≤4 句面向读者的流畅散文(en/zh 分写);
|
|
16
17
|
- 禁证据 ID 列表(E-xxx/EV-xxx)、禁 schema 键(overall_risk= 等)、禁截断残留(null);
|
|
17
|
-
- what_can/cannot_be_claimed 等列表同样人话化;统计数字可保留但以自然表达呈现。
|
|
18
|
+
- what_can/cannot_be_claimed 等列表同样人话化;统计数字可保留但以自然表达呈现。
|
|
19
|
+
- **先过 Pre-Verdict Gate 再定稿**:critical 失败会封顶置信度并禁止高置信 ADOPT。
|
|
20
|
+
- 置信度由规则化公式给出(0.30 质量 + 0.25 一致性 + 0.20 直接性 + 0.25 独立研究数 − 双罚分),不由模型自评。
|
|
21
|
+
- 证据不足时输出 INSUFFICIENT EVIDENCE,并写明"什么证据会改变这个决定"。
|
|
22
|
+
- 语言人话化规则见 `references/scientific-invariants.md` 与角色提示词。
|
|
23
|
+
|
|
24
|
+
## 质量门
|
|
25
|
+
- [ ] `final_verdict.json` 通过 verdict schema,action 取自四态枚举。
|
|
26
|
+
- [ ] Pre-Verdict Gate 已执行并留痕(`gate_report.json`)。
|
|
27
|
+
- [ ] 置信度可复算,分解项可审计。
|
|
28
|
+
- [ ] 反证与 threats_to_validity 均被显式回应(采纳、限定或驳回并说明理由)。
|
|
29
|
+
|
|
30
|
+
## 失败模式与回退
|
|
31
|
+
| 失败 | 处理 |
|
|
32
|
+
|---|---|
|
|
33
|
+
| `PRE_VERDICT_FAILED` | 修复前置产物(schema / 交叉评审 / 方法学)后重跑闸门。 |
|
|
34
|
+
| `GATE_CRITICAL_FAILURE` | 封顶置信度,强制降至 PILOT 或 INSUFFICIENT EVIDENCE。 |
|
|
35
|
+
| `CONFLICT_UNRESOLVED` | 保持不确定,不强行裁决。 |
|
|
36
|
+
| 证据仅支持任务表现 | 不得据以产出学习效果类结论。 |
|
|
37
|
+
|
|
38
|
+
## 语言与呈现契约
|
|
39
|
+
面向决策者的第一页语言;ID、枚举、URL 保留可追溯性,但不出现在叙述句中。
|
|
40
|
+
|
|
41
|
+
## 交接说明
|
|
42
|
+
`final_verdict.json` 是 Applicability / Intervene / Evaluate 与一切投影的唯一依据。
|
|
@@ -1,3 +1,38 @@
|
|
|
1
|
-
#
|
|
1
|
+
# Task Brief — stage: applicability(角色:evidence-judge)
|
|
2
2
|
|
|
3
|
-
|
|
3
|
+
## 目标
|
|
4
|
+
判定证据可迁移到目标人群与场景的程度:明确支持人群、实施条件、被排除人群、结局边界与不确定性。
|
|
5
|
+
|
|
6
|
+
## 前置输入
|
|
7
|
+
- `evidence.jsonl`(含人群、场景、剂量与结局信息)
|
|
8
|
+
- `final_verdict.json`(裁决与边界)
|
|
9
|
+
- `references/applicability-policy.md`
|
|
10
|
+
|
|
11
|
+
## 产出
|
|
12
|
+
- `applicability.json`:supported / unsupported populations、conditions、outcome limits、
|
|
13
|
+
boundary of transfer、uncertainty statement;供 Intervene 判断试点人群是否落在支持范围内。
|
|
14
|
+
|
|
15
|
+
## 执行规则
|
|
16
|
+
- **阳性结果不自动迁移**:不能因为"研究有效"就推定目标人群同样有效。
|
|
17
|
+
- 逐条对齐:目标人群 vs 研究人群、目标场景 vs 研究场景、目标 outcome vs 已测量 outcome。
|
|
18
|
+
- 明确写出被排除人群(例如被排除的补习班学生、特殊需求学习者)与理由。
|
|
19
|
+
- 剂量与实施条件(师资培训、课时、工具版本、护栏)必须显式列出——这些常常是效果成立的前提。
|
|
20
|
+
- 不确定之处如实标注为不确定,不用"一般来说"这类泛化措辞掩盖缺口。
|
|
21
|
+
|
|
22
|
+
## 质量门
|
|
23
|
+
- [ ] 支持人群、条件、结局边界、排除人群四类信息齐全。
|
|
24
|
+
- [ ] 每条边界都追溯到具体证据,而非泛泛陈述。
|
|
25
|
+
- [ ] 不确定性显式声明,且与裁决置信度一致。
|
|
26
|
+
|
|
27
|
+
## 失败模式与回退
|
|
28
|
+
| 失败 | 处理 |
|
|
29
|
+
|---|---|
|
|
30
|
+
| `SCOPE_MISMATCH` | 缩小结论范围而非放宽适用条件。 |
|
|
31
|
+
| 关键人群信息缺失 | 记入不确定项;不得默认"与目标人群相同"。 |
|
|
32
|
+
| 边界与裁决口径冲突 | 以证据为准修正边界,并回退到 Adjudicate 复核。 |
|
|
33
|
+
|
|
34
|
+
## 语言与呈现契约
|
|
35
|
+
面向决策者的人话表述:谁适用、在什么条件下、对哪些结果、到什么程度为止。
|
|
36
|
+
|
|
37
|
+
## 交接说明
|
|
38
|
+
`applicability.json` 约束 Intervene 的试点人群;越界设计必须被拒绝或缩回支持范围。
|
|
@@ -3,13 +3,38 @@
|
|
|
3
3
|
## 目标
|
|
4
4
|
按审计清单审查每个研究的方法学质量,强制执行"任务完成表现 ≠ 学习效果"最高优先级规则。
|
|
5
5
|
|
|
6
|
-
##
|
|
7
|
-
- evidence.jsonl + 来源 fetch 内容
|
|
6
|
+
## 前置输入
|
|
7
|
+
- `evidence.jsonl` + 来源 fetch 内容
|
|
8
|
+
- 审计框架:`references/wwc_standards.md`、`references/grade_framework.md`、`references/methodology-audit.md`
|
|
8
9
|
|
|
9
10
|
## 产出
|
|
10
|
-
- methodology.json
|
|
11
|
-
|
|
11
|
+
- `methodology.json`:15 项 `audit_items`(control_group / randomization / pre_test / post_test /
|
|
12
|
+
retention_test / transfer_test / sample_bias / self_selection / measurement_validity / confounders /
|
|
13
|
+
instructor_effect / novelty_effect / tool_version_effect / ai_usage_policy / dropout,每项 status:
|
|
14
|
+
met|partial|missing|not_applicable)+ 每条 PASS/CONCERN/FAIL verdict + `task_vs_learning_guard`
|
|
15
|
+
(显示层经 zh_labels 映射中文),须通过 `schemas/methodology.schema.json`。
|
|
12
16
|
|
|
13
|
-
##
|
|
14
|
-
- 只审"证据站不站得住",不审"证据说什么"
|
|
15
|
-
|
|
17
|
+
## 执行规则
|
|
18
|
+
- 只审"证据站不站得住",不审"证据说什么"——审计结论不得夹带对效果的判断。
|
|
19
|
+
- `task_vs_learning_guard.equates_task_with_learning` 必须显式给出;为 true 时该研究不得支撑任何学习效果声明。
|
|
20
|
+
- 审计说明(note/summary)为人话叙述,枚举/代号(PASS/CONCERN/FAIL)只作标签。
|
|
21
|
+
- 独立性:Method Reviewer 与内容判断角色分离(`independence_required: true`),且需强上下文能力。
|
|
22
|
+
|
|
23
|
+
## 质量门
|
|
24
|
+
- [ ] 15 项审计项齐全,无缺项、无自造项。
|
|
25
|
+
- [ ] 每项状态取自四值枚举,`not_applicable` 必须说明为何不适用。
|
|
26
|
+
- [ ] `task_vs_learning_guard` 存在且结论明确。
|
|
27
|
+
- [ ] 每条 PASS/CONCERN/FAIL 有可核验依据(对应原文特征)。
|
|
28
|
+
|
|
29
|
+
## 失败模式与回退
|
|
30
|
+
| 失败 | 处理 |
|
|
31
|
+
|---|---|
|
|
32
|
+
| 关键信息缺失(如未报告随机化) | 记 `missing` 并在说明中写清缺什么,不猜测。 |
|
|
33
|
+
| 结论依赖任务表现 | 触发 guard,剥夺其学习效果支撑资格。 |
|
|
34
|
+
| schema 校验失败 | 修复后重跑审计。 |
|
|
35
|
+
|
|
36
|
+
## 语言与呈现契约
|
|
37
|
+
审计叙述为人话;状态枚举只作标签显示。
|
|
38
|
+
|
|
39
|
+
## 交接说明
|
|
40
|
+
`methodology.json` 直接决定每条证据能否支持 claim;Adjudicate 必须读取 `task_vs_learning_guard` 与各条 verdict。
|
|
@@ -5,11 +5,40 @@
|
|
|
5
5
|
novelty effect、alternative explanation;禁止虚构反方证据;没有反方证据时输出
|
|
6
6
|
NO CONTRADICTORY EVIDENCE FOUND。
|
|
7
7
|
|
|
8
|
-
##
|
|
9
|
-
- evidence.jsonl
|
|
8
|
+
## 前置输入
|
|
9
|
+
- `evidence.jsonl`
|
|
10
|
+
- `frame.json`(用于判断结论是否超出研究范围)
|
|
10
11
|
|
|
11
12
|
## 产出
|
|
12
|
-
- skeptic.json
|
|
13
|
+
- `skeptic.json`:`search_performed=true`;`skeptic_findings[]`(九项检查,字段名稳定:
|
|
14
|
+
`check` / `status`(found|not_found)/ `detail` / `related_evidence_ids`);
|
|
15
|
+
`contradictory_evidence_found`;`no_contradictory_evidence_statement`;`threats_to_validity`。
|
|
16
|
+
- 作为独立交叉审核角色时的输出须符合 `schemas/cross-model-review.schema.json`
|
|
17
|
+
(必需字段 `agreement`、`final_recommendation`)。
|
|
13
18
|
|
|
14
|
-
##
|
|
15
|
-
|
|
19
|
+
## 执行规则(固定 9 项检查,缺一不可)
|
|
20
|
+
1. null result 2. negative result 3. contradictory evidence 4. alternative explanation
|
|
21
|
+
5. measurement mismatch(测的是任务完成还是学习) 6. sampling bias
|
|
22
|
+
7. novelty effect 8. AI dependency / over-reliance 9. scope overreach
|
|
23
|
+
|
|
24
|
+
- 反方检索必须独立构造查询,不复用支持证据的检索式。
|
|
25
|
+
- 找不到反证就是 "not_found",并输出标准语句;**不得为了凑数虚构反方文献**。
|
|
26
|
+
- 独立性:Skeptic 不得与主分析使用同一模型家族(`independence_required: true`)。
|
|
27
|
+
|
|
28
|
+
## 质量门
|
|
29
|
+
- [ ] 九项检查全部有 `status`,无遗漏项。
|
|
30
|
+
- [ ] 每条 `found` 的 finding 都绑定 `related_evidence_ids` 或明确来源。
|
|
31
|
+
- [ ] `no_contradictory_evidence_statement` 与 `contradictory_evidence_found` 语义一致。
|
|
32
|
+
|
|
33
|
+
## 失败模式与回退
|
|
34
|
+
| 失败 | 处理 |
|
|
35
|
+
|---|---|
|
|
36
|
+
| 反证检索为空 | 记录 negative-search record;输出标准语句,不填充虚构来源。 |
|
|
37
|
+
| 反证与支持证据冲突 | 保留双方,交 Adjudicate 处理;不在此阶段裁决。 |
|
|
38
|
+
| 交叉审核不可用(无独立模型) | 降级为原生自审并显式标注,不得伪装为独立审核。 |
|
|
39
|
+
|
|
40
|
+
## 语言与呈现契约(人话化硬标准)
|
|
41
|
+
叙述字段为面向研究者的流畅中文;禁证据 ID 堆砌(用"作者-年份 + 人话描述"替代)。
|
|
42
|
+
|
|
43
|
+
## 交接说明
|
|
44
|
+
`skeptic.json` 是 Adjudicate 的必读输入;其 `threats_to_validity` 必须被裁决明确回应或采纳。
|
|
@@ -3,11 +3,36 @@
|
|
|
3
3
|
## 目标
|
|
4
4
|
为 PILOT/ADOPT 配套评价方案:基线/后测/保持/迁移 + 过程/学习/风险指标 + 成功阈值与停止条件。
|
|
5
5
|
|
|
6
|
-
##
|
|
7
|
-
- intervention.json + final_verdict.json
|
|
6
|
+
## 前置输入
|
|
7
|
+
- `intervention.json` + `final_verdict.json`
|
|
8
|
+
- 评价框架:`references/evaluation-design.md`、`references/evaluation-policy.md`、`references/outcome-taxonomy.md`
|
|
8
9
|
|
|
9
10
|
## 产出
|
|
10
|
-
- evaluation.json
|
|
11
|
+
- `evaluation.json`(`schemas/evaluation.schema.json`):baseline / post / retention / transfer 测量点,
|
|
12
|
+
过程与风险指标,成功与失败阈值,分析计划(可被 `scripts/did_regression.py` 直接执行)。
|
|
11
13
|
|
|
12
|
-
##
|
|
13
|
-
- 分离 Task Performance 与 Learning Effect
|
|
14
|
+
## 执行规则
|
|
15
|
+
- 分离 Task Performance 与 Learning Effect;两者指标不得合并成一个"效果"。
|
|
16
|
+
- 保持(retention)与迁移(transfer)必须单独设计测量点,不能只做即测。
|
|
17
|
+
- 阈值在数据到达前固定;事后改阈值属协议偏离,须记录而非吸收。
|
|
18
|
+
- 分析计划要具体到可执行(设计类型、比较组、协变量、缺失处理),使 DID/效应量脚本无需再做决定。
|
|
19
|
+
- 叙述为人话。
|
|
20
|
+
|
|
21
|
+
## 质量门
|
|
22
|
+
- [ ] `evaluation.json` 通过 evaluation schema。
|
|
23
|
+
- [ ] 基线、即测、保持、迁移四类测量点齐全。
|
|
24
|
+
- [ ] 成功、失败、停止阈值三者都在,且先于数据固定。
|
|
25
|
+
- [ ] 分析计划可被确定性脚本执行(无残留人工决策)。
|
|
26
|
+
|
|
27
|
+
## 失败模式与回退
|
|
28
|
+
| 失败 | 处理 |
|
|
29
|
+
|---|---|
|
|
30
|
+
| 无法做对照 | 改为单组前后测并显式标注设计局限,不得仍按因果口径表述。 |
|
|
31
|
+
| 阈值事后调整 | 记录为协议偏离并降级结论强度。 |
|
|
32
|
+
| 样本量不足 | 报告功效局限;不得以不显著当作"无效果"。 |
|
|
33
|
+
|
|
34
|
+
## 语言与呈现契约
|
|
35
|
+
人话评价方案;统计符号与阈值数字保留原样以便复算。
|
|
36
|
+
|
|
37
|
+
## 交接说明
|
|
38
|
+
试用数据回收后由 `evaluate-and-update` 工作流重新注入证据图谱并再裁决。
|