eduevidence 6.0.0 → 6.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +395 -0
- package/CONTRIBUTING.md +105 -0
- package/README.md +113 -49
- package/README.zh-CN.md +39 -12
- package/SKILL.md +15 -5
- package/assets/readme/landing-tour.gif +0 -0
- package/assets/readme/studio-tour.gif +0 -0
- package/benchmarks/evidence-library.json +277 -1
- package/bin/eduevidence.js +2 -1
- package/docs/architecture.md +325 -46
- package/docs/demo-workplace-ai.md +1 -1
- package/docs/install-guide.md +1 -1
- package/docs/j-ev-experimental.md +250 -0
- package/docs/orchestration-role-model.md +1 -1
- package/docs/release-closeout/README.md +1 -1
- package/docs/reproducibility.md +138 -0
- package/docs/sciverse-api.md +125 -0
- package/domains/_neutral/copy/few_shots.json +21 -0
- package/domains/_neutral/copy/framing_lexicon.json +19 -0
- package/domains/_neutral/copy/module_labels.json +5 -0
- package/domains/_neutral/copy/module_labels_footer.json +102 -0
- package/domains/_neutral/copy/module_labels_modules.json +204 -0
- package/domains/_neutral/copy/module_labels_nav.json +126 -0
- package/domains/_neutral/copy/module_labels_summary.json +98 -0
- package/domains/_neutral/copy/module_labels_tables.json +164 -0
- package/domains/_neutral/copy/module_labels_v2.json +90 -0
- package/domains/_neutral/copy/risk_constructs.json +20 -0
- package/domains/_neutral/copy/section_titles.json +66 -0
- package/domains/_neutral/copy/terminology.json +11 -0
- package/domains/check_copy_packs.py +103 -0
- package/domains/education/copy/few_shots.json +22 -0
- package/domains/education/copy/framing_enums.json +167 -0
- package/domains/education/copy/framing_lexicon.json +166 -0
- package/domains/education/copy/module_labels.json +169 -0
- package/domains/education/copy/risk_constructs.json +48 -0
- package/domains/education/copy/section_titles.json +186 -0
- package/domains/education/copy/terminology.json +70 -0
- package/domains/education/manifest.json +1 -1
- package/domains/education/outcome_taxonomy.json +2 -2
- package/domains/manifest.json +1 -1
- package/domains/policy/copy/few_shots.json +22 -0
- package/domains/policy/copy/framing_enums.json +94 -0
- package/domains/policy/copy/framing_lexicon.json +174 -0
- package/domains/policy/copy/module_labels.json +168 -0
- package/domains/policy/copy/risk_constructs.json +33 -0
- package/domains/policy/copy/section_titles.json +186 -0
- package/domains/policy/copy/terminology.json +64 -0
- package/eduevidence_cli.py +10 -0
- package/engine/capabilities.py +57 -5
- package/engine/decision_policy.py +167 -0
- package/engine/evidence_graph.py +14 -10
- package/engine/gaps.py +42 -22
- package/engine/ids.py +2 -0
- package/engine/library.py +6 -2
- package/engine/library_builtin.py +7 -4
- package/engine/living.py +34 -4
- package/engine/migration.py +88 -3
- package/engine/orchestration.py +5 -5
- package/engine/paths.py +2 -0
- package/engine/pilot.py +34 -32
- package/engine/taxonomy.py +211 -0
- package/engine/tribunal.py +49 -43
- package/engine/versions.py +1 -1
- package/examples/ai-coding-assistant-evidence/EduEvidence_Report.html +1361 -147
- package/examples/ai-coding-assistant-evidence/artifact_manifest.json +3 -3
- package/examples/ai-coding-assistant-evidence/citation_check.json +1 -1
- package/examples/ai-coding-assistant-evidence/final_verdict.json +107 -0
- package/examples/ai-coding-assistant-evidence/gate_report.json +101 -0
- package/examples/ai-coding-assistant-evidence/report_spec.json +23 -12
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_academic.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_claude.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_datalab-dark.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_datalab.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_presentation.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_academic.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_claude.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_datalab-dark.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_datalab.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_presentation.html +1360 -146
- package/examples/ai-coding-assistant-evidence/result.json +13 -9
- package/examples/ai-coding-assistant-evidence/result.zh.json +45 -41
- package/examples/ai-coding-assistant-evidence/skeptic.json +72 -0
- package/examples/ai-coding-assistant-evidence/verdict.json +6 -2
- package/examples/spaced-retrieval-practice/EduEvidence_Report.html +2728 -0
- package/examples/spaced-retrieval-practice/applicability.json +14 -0
- package/examples/spaced-retrieval-practice/artifact_manifest.json +15 -0
- package/examples/spaced-retrieval-practice/claims.jsonl +3 -0
- package/examples/spaced-retrieval-practice/evidence.jsonl +6 -0
- package/examples/spaced-retrieval-practice/final_verdict.json +93 -0
- package/examples/spaced-retrieval-practice/frame.json +58 -0
- package/examples/spaced-retrieval-practice/gate_report.json +101 -0
- package/examples/spaced-retrieval-practice/methodology.json +78 -0
- package/examples/spaced-retrieval-practice/report.html +2522 -0
- package/examples/spaced-retrieval-practice/report_spec.json +212 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_academic.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_claude.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_datalab-dark.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_datalab.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_presentation.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_academic.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_claude.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_datalab-dark.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_datalab.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_presentation.html +2728 -0
- package/examples/spaced-retrieval-practice/result.json +942 -0
- package/examples/spaced-retrieval-practice/result.zh.json +942 -0
- package/examples/spaced-retrieval-practice/skeptic.json +70 -0
- package/examples/spaced-retrieval-practice/sources.jsonl +7 -0
- package/examples/spaced-retrieval-practice/verdict.json +93 -0
- package/examples/workplace-ai-assistant/EduEvidence_Report.html +2814 -0
- package/examples/workplace-ai-assistant/artifact_manifest.json +15 -0
- package/examples/workplace-ai-assistant/claims.jsonl +4 -4
- package/examples/workplace-ai-assistant/evidence.jsonl +4 -4
- package/examples/workplace-ai-assistant/evidence_graph.json +15 -15
- package/examples/workplace-ai-assistant/final_verdict.json +78 -0
- package/examples/workplace-ai-assistant/gate_report.json +101 -0
- package/examples/workplace-ai-assistant/report_spec.json +209 -40
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_academic.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_claude.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_datalab-dark.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_datalab.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_presentation.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/report_academic.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_claude.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_datalab-dark.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_datalab.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_presentation.html +2814 -0
- package/examples/workplace-ai-assistant/result.json +82 -20
- package/examples/workplace-ai-assistant/result.zh.json +82 -20
- package/examples/workplace-ai-assistant/skeptic.json +72 -0
- package/examples/workplace-ai-assistant/verdict.json +36 -10
- package/integrations/agent_mcp.py +2 -2
- package/integrations/jev/__init__.py +115 -0
- package/integrations/jev/approval.py +212 -0
- package/integrations/jev/cli.py +84 -0
- package/integrations/jev/config.py +112 -0
- package/integrations/jev/gateway.py +128 -0
- package/integrations/jev/modes.py +38 -0
- package/integrations/jev/tools_classify.py +88 -0
- package/integrations/jev/tools_extract.py +111 -0
- package/integrations/jev/tools_rerank.py +71 -0
- package/integrations/jev/tools_screen.py +87 -0
- package/integrations/jev/tools_verify.py +95 -0
- package/integrations/jev_mcp.py +22 -0
- package/integrations/semantic_decide.py +286 -0
- package/integrations/semdecide_cli.py +55 -0
- package/package.json +19 -2
- package/pyproject.toml +4 -3
- package/references/report-copy-style.md +107 -0
- package/references/retrieval-compliance.md +75 -0
- package/references/retrieval-protocol.md +20 -0
- package/retrieval/audit.py +27 -3
- package/retrieval/fetch.py +96 -0
- package/retrieval/sciverse.py +398 -0
- package/retrieval/search.py +47 -7
- package/schemas/applicability.schema.json +94 -0
- package/schemas/chart-spec.schema.json +10 -3
- package/schemas/evidence.schema.json +316 -43
- package/schemas/fetch-result.schema.json +2 -1
- package/schemas/report-result.schema.json +3 -3
- package/schemas/report-spec.schema.json +98 -100
- package/schemas/skeptic.schema.json +86 -0
- package/schemas/source.schema.json +21 -2
- package/schemas/v2/decision-snapshot.schema.json +20 -9
- package/schemas/v2/finding.schema.json +5 -1
- package/schemas/v2/intake.schema.json +191 -0
- package/schemas/v2/methodology-audit.schema.json +5 -1
- package/schemas/v2/outcome.schema.json +28 -5
- package/schemas/v2/study.schema.json +5 -1
- package/schemas/vNext/autoevolve-session.schema.json +34 -1
- package/schemas/vNext/eval-snapshot.schema.json +77 -1
- package/schemas/vNext/execution-plan.schema.json +50 -1
- package/schemas/vNext/gap-priority.schema.json +54 -1
- package/schemas/vNext/negative-search-record.schema.json +68 -1
- package/schemas/vNext/research-iteration.schema.json +87 -1
- package/schemas/vNext/research-strategy.schema.json +62 -1
- package/schemas/vNext/skill-experiment.schema.json +90 -1
- package/schemas/vNext/task-spec.schema.json +156 -1
- package/schemas/vNext/worker-result.schema.json +60 -1
- package/schemas/verdict.schema.json +164 -28
- package/scripts/build_evidence_library.py +15 -5
- package/scripts/build_report_variants.py +18 -2
- package/scripts/build_result.py +74 -9
- package/scripts/check_package_parity.py +85 -0
- package/scripts/check_protocol_alignment.py +375 -0
- package/scripts/check_versioned_schemas.py +254 -0
- package/scripts/claim_audit.py +13 -8
- package/scripts/compute_confidence.py +10 -0
- package/scripts/dashboard_server.py +13 -2
- package/scripts/did_regression.py +12 -2
- package/scripts/evidence_score.py +5 -2
- package/scripts/intake/__init__.py +31 -0
- package/scripts/intake/__main__.py +18 -0
- package/scripts/intake/background.py +78 -0
- package/scripts/intake/browser.py +79 -0
- package/scripts/intake/cli.py +57 -0
- package/scripts/intake/constants.py +57 -0
- package/scripts/intake/depth.py +53 -0
- package/scripts/intake/enhancements.py +106 -0
- package/scripts/intake/hooks.py +90 -0
- package/scripts/intake/prefs.py +76 -0
- package/scripts/intake/prompts.py +85 -0
- package/scripts/intake/session.py +152 -0
- package/scripts/lint_file_layers.py +126 -0
- package/scripts/orchestrator.py +187 -40
- package/scripts/pre_verdict_gate.py +241 -29
- package/scripts/quickstart.py +18 -2
- package/scripts/run_workspace.py +7 -1
- package/scripts/skill_lint.py +11 -1
- package/scripts/skill_payload.py +6 -3
- package/scripts/test_adversarial_empirical.py +96 -25
- package/scripts/validate_schema.py +31 -1
- package/skill/agents/evaluation-designer.md +20 -4
- package/skill/agents/evidence-analyst.md +19 -3
- package/skill/agents/evidence-judge.md +98 -8
- package/skill/agents/evidence-retriever.md +20 -3
- package/skill/agents/intervention-designer.md +20 -4
- package/skill/agents/method-reviewer.md +18 -2
- package/skill/agents/{education-planner.md → research-planner.md} +19 -3
- package/skill/agents/skeptic.md +18 -2
- package/skill/roles/registry.yaml +11 -11
- package/skill/sub-skills/aihot-trend-analysis/SKILL.md +28 -9
- package/skill/sub-skills/contradiction-analysis/SKILL.md +31 -11
- package/skill/sub-skills/data-analysis/SKILL.md +34 -15
- package/skill/sub-skills/ethics-review/SKILL.md +33 -10
- package/skill/sub-skills/evidence-extraction/SKILL.md +29 -11
- package/skill/sub-skills/evidence-review/SKILL.md +31 -12
- package/skill/sub-skills/gap-analysis/SKILL.md +31 -9
- package/skill/sub-skills/literature-review/SKILL.md +35 -14
- package/skill/sub-skills/methodology-audit/SKILL.md +29 -12
- package/skill/sub-skills/report-generation/SKILL.md +28 -0
- package/skill/sub-skills/research-planning/SKILL.md +41 -14
- package/skill/sub-skills/study-design/SKILL.md +30 -9
- package/skill/task-briefs/adjudicate.md +32 -7
- package/skill/task-briefs/applicability.md +37 -2
- package/skill/task-briefs/audit.md +32 -7
- package/skill/task-briefs/challenge.md +34 -5
- package/skill/task-briefs/evaluate.md +30 -5
- package/skill/task-briefs/extract.md +31 -8
- package/skill/task-briefs/frame.md +39 -10
- package/skill/task-briefs/intervene.md +32 -6
- package/skill/task-briefs/present.md +32 -8
- package/skill/task-briefs/projection.md +36 -2
- package/skill/task-briefs/retrieve.md +36 -6
- package/skill/workflows/decision-and-pilot.md +76 -1
- package/skill/workflows/evaluate-and-update.md +83 -0
- package/skill/workflows/evidence-review.md +104 -0
- package/skill/workflows/experimental-jev.md +170 -0
- package/skill/workflows/intake.md +120 -0
- package/visualization/eduevidence-report/scripts/build_figures.py +25 -3
- package/visualization/eduevidence-report/scripts/build_infographics.py +37 -15
- package/visualization/eduevidence-report/scripts/build_report.py +435 -575
- package/visualization/eduevidence-report/scripts/charts_data.py +2 -0
- package/visualization/eduevidence-report/scripts/lieflat_engine.py +349 -38
- package/visualization/eduevidence-report/scripts/report_copy_pack.py +296 -0
- package/visualization/eduevidence-report/scripts/report_copy_policy_guard.py +47 -0
- package/visualization/eduevidence-report/scripts/zh_labels.py +141 -1
- package/web/architecture.html +14885 -0
- package/web/studio/assets/index-B8tkF44Q.css +1 -0
- package/web/studio/index.html +2 -2
- package/scripts/build_esl_artifacts.py +0 -1921
- package/scripts/build_killer_demo.py +0 -295
- package/scripts/enrich_projects_human_and_lieflat.py +0 -315
- package/scripts/generate_new_projects.py +0 -686
- package/scripts/sync_killer_demo_report.py +0 -270
- package/web/studio/assets/index-CzXocaGv.css +0 -1
- /package/web/studio/assets/{index-pa7jD7n4.js → index-CQ6Keoyc.js} +0 -0
|
@@ -1,42 +1,42 @@
|
|
|
1
1
|
roles:
|
|
2
|
-
|
|
2
|
+
research-planner:
|
|
3
3
|
responsibility: framing completeness and scope
|
|
4
4
|
stages: [frame]
|
|
5
|
-
capabilities: [
|
|
5
|
+
capabilities: [research_framing]
|
|
6
6
|
critical_path: true
|
|
7
7
|
evidence-retriever:
|
|
8
8
|
responsibility: source acquisition and provenance
|
|
9
9
|
stages: [retrieve]
|
|
10
|
-
capabilities: [
|
|
10
|
+
capabilities: [literature_search, counter_evidence_search, source_fetch, source_validation]
|
|
11
11
|
evidence-analyst:
|
|
12
12
|
responsibility: structured finding extraction
|
|
13
13
|
stages: [extract]
|
|
14
|
-
capabilities: [
|
|
14
|
+
capabilities: [study_extraction, finding_extraction, claim_linking]
|
|
15
15
|
skeptic:
|
|
16
16
|
responsibility: independent counter-evidence coverage
|
|
17
17
|
stages: [challenge]
|
|
18
|
-
capabilities: [
|
|
19
|
-
independence_required:
|
|
18
|
+
capabilities: [counter_evidence_search]
|
|
19
|
+
independence_required: different-model-family
|
|
20
20
|
critical_path: true
|
|
21
21
|
method-reviewer:
|
|
22
22
|
responsibility: methodology and construct-validity appraisal
|
|
23
23
|
stages: [audit]
|
|
24
|
-
capabilities: [
|
|
25
|
-
independence_required:
|
|
24
|
+
capabilities: [methodology_appraisal]
|
|
25
|
+
independence_required: role-separation
|
|
26
26
|
critical_path: true
|
|
27
27
|
evidence-judge:
|
|
28
28
|
responsibility: evidence-bounded adjudication and applicability
|
|
29
29
|
stages: [adjudicate, applicability]
|
|
30
|
-
capabilities: [
|
|
30
|
+
capabilities: [evidence_synthesis, tribunal, applicability_analysis, knowledge_gap_detection]
|
|
31
31
|
critical_path: true
|
|
32
32
|
intervention-designer:
|
|
33
33
|
responsibility: grounded intervention or pilot design
|
|
34
34
|
stages: [intervene]
|
|
35
|
-
capabilities: [
|
|
35
|
+
capabilities: [study_design, measurement_design, intervention_design]
|
|
36
36
|
evaluation-designer:
|
|
37
37
|
responsibility: estimable evaluation and update logic
|
|
38
38
|
stages: [evaluate]
|
|
39
|
-
capabilities: [
|
|
39
|
+
capabilities: [evaluation_design, data_validation, data_analysis]
|
|
40
40
|
|
|
41
41
|
execution:
|
|
42
42
|
max_parallel_workers: 6
|
|
@@ -1,31 +1,50 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: aihot-trend-analysis
|
|
3
3
|
description: "Real-time horizon scanning and dynamic trend ingestion for emerging AI educational tools, model benchmarks, and EdTech releases via AIHot."
|
|
4
|
+
capability: literature_search (grey-literature channel)
|
|
4
5
|
---
|
|
5
6
|
|
|
6
|
-
# aihot-trend-analysis — Real-Time AI & EdTech Trend Ingestion
|
|
7
|
+
# aihot-trend-analysis — Real-Time AI & EdTech Trend Ingestion
|
|
7
8
|
|
|
8
9
|
## When to Use
|
|
9
|
-
Triggered when an
|
|
10
|
+
Triggered when an inquiry involves fast-moving generative AI tools (Cursor, Claude, Socratic LLM tutors, Copilot) where peer-reviewed literature may lag 6–18 months.
|
|
10
11
|
|
|
11
|
-
##
|
|
12
|
-
- `keyword`:
|
|
13
|
-
- `time_window
|
|
14
|
-
- `category
|
|
12
|
+
## Inputs
|
|
13
|
+
- `keyword`: target technology or pedagogy topic.
|
|
14
|
+
- `time_window` (optional): `24h` / `7d` / `30d`.
|
|
15
|
+
- `category` (optional): `EdTech` / `Agents` / `Reasoning` / `LLMs`.
|
|
16
|
+
|
|
17
|
+
## Process
|
|
18
|
+
1. Query the AIHot channel through `retrieval/search.py` (`AIHotProvider`).
|
|
19
|
+
2. Record every hit as grey literature with its publication time.
|
|
20
|
+
3. Route any factual claim that would enter the decision back through Retrieve → Fetch → Validate: trend items never bypass RULE 2.
|
|
15
21
|
|
|
16
22
|
## Output Contract
|
|
17
|
-
|
|
23
|
+
`SearchHit` objects tagged `provider: "aihot"` with grey-literature authority (`tier5_general_web`); they inform horizon scanning, not effect estimation.
|
|
18
24
|
|
|
19
25
|
```json
|
|
20
26
|
{
|
|
21
27
|
"trend_items": [
|
|
22
28
|
{
|
|
23
|
-
"title": "
|
|
29
|
+
"title": "Socratic tutoring framework evaluated across 10 universities",
|
|
24
30
|
"url": "https://aihot.virxact.com/api/item/...",
|
|
25
|
-
"summary": "Benchmark evaluation on novice
|
|
31
|
+
"summary": "Benchmark evaluation on novice retention and prompt scaffolding.",
|
|
26
32
|
"category": "EdTech",
|
|
27
33
|
"publish_time": "2026-08-15"
|
|
28
34
|
}
|
|
29
35
|
]
|
|
30
36
|
}
|
|
31
37
|
```
|
|
38
|
+
|
|
39
|
+
## Quality Gates
|
|
40
|
+
- [ ] 每条 trend 项带 URL 与时间戳。
|
|
41
|
+
- [ ] 明确标注为灰来源,不进入效应量合成。
|
|
42
|
+
|
|
43
|
+
## Anti-Patterns
|
|
44
|
+
- 用产品博客宣称的效果当作实证证据;把版本发布日期当研究发表时间。
|
|
45
|
+
|
|
46
|
+
## Worked Example
|
|
47
|
+
关键词 "AI programming assistant" → 返回 30 天内的发布与基准报道,用于判断文献滞后期内是否出现新的风险信号。
|
|
48
|
+
|
|
49
|
+
## References
|
|
50
|
+
- `retrieval/search.py::AIHotProvider`、`references/retrieval-compliance.md`
|
|
@@ -1,17 +1,37 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: contradiction-analysis
|
|
3
3
|
description: "Mines adversarial claims, conflicting effect directions, and boundary condition qualifiers."
|
|
4
|
+
capability: counter_evidence_search (skeptic side)
|
|
4
5
|
---
|
|
6
|
+
|
|
5
7
|
# Contradiction Analysis Skill
|
|
6
8
|
|
|
7
|
-
##
|
|
8
|
-
Trigger during evidence synthesis when studies on the same Claim ID
|
|
9
|
-
|
|
10
|
-
##
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
9
|
+
## When to Use
|
|
10
|
+
Trigger during evidence synthesis when studies on the same Claim ID show conflicting directions (SUPPORTS vs CONTRADICTS) or high heterogeneity.
|
|
11
|
+
|
|
12
|
+
## Inputs
|
|
13
|
+
- `evidence.jsonl`(含方向标签)
|
|
14
|
+
- `frame.json`(判断是否 scope overreach)
|
|
15
|
+
|
|
16
|
+
## Process
|
|
17
|
+
1. **Adversarial Mining (Skeptic)**: identify confounders (teacher training, dosage, novelty); evaluate boundary conditions (does it fail for novices vs experts?).
|
|
18
|
+
2. **Directional Separation**: strictly separate supporting / contradicting / neutral — never blend them into one "mixed" bucket.
|
|
19
|
+
3. **Heterogeneity Attribution**: map conflict to subgroup variation, dosage thresholds, or outcome instrument differences.
|
|
20
|
+
4. **Nine fixed checks** per `skill/agents/skeptic.md`; absence of counter-evidence yields the standard statement, never invented sources.
|
|
21
|
+
|
|
22
|
+
## Output Contract
|
|
23
|
+
`skeptic.json` — `skeptic_findings[]` (`check` / `status` / `detail` / `related_evidence_ids`), `contradictory_evidence_found`, `no_contradictory_evidence_statement`, `threats_to_validity`.
|
|
24
|
+
|
|
25
|
+
## Quality Gates
|
|
26
|
+
- [ ] 九项检查齐全。
|
|
27
|
+
- [ ] 每条 found 绑定证据或来源。
|
|
28
|
+
- [ ] 标准语句与布尔标志语义一致。
|
|
29
|
+
|
|
30
|
+
## Anti-Patterns
|
|
31
|
+
- 把三列合并成"总体看有效";为了显得严谨而虚构反证。
|
|
32
|
+
|
|
33
|
+
## Worked Example
|
|
34
|
+
同一 claim 下出现 g=+0.48(无护栏练习)与 g=−0.17(独立考试),归因到 outcome 测量差异与护栏配置,而非取平均。
|
|
35
|
+
|
|
36
|
+
## References
|
|
37
|
+
- `references/skeptic-protocol.md`、`skill/agents/skeptic.md`、`skill/task-briefs/challenge.md`
|
|
@@ -1,23 +1,42 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: data-analysis
|
|
3
3
|
description: "Runs deterministic statistical regression (DID/OLS) on user-uploaded classroom and field datasets to re-inject local empirical evidence."
|
|
4
|
+
capability: data_validation + data_analysis
|
|
4
5
|
---
|
|
6
|
+
|
|
5
7
|
# Data Analysis Skill
|
|
6
8
|
|
|
7
|
-
##
|
|
9
|
+
## When to Use
|
|
8
10
|
Trigger when the user imports empirical classroom or survey data (CSV/XLSX) from an active field trial or pilot deployment.
|
|
9
11
|
|
|
10
|
-
##
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
-
|
|
22
|
-
|
|
23
|
-
|
|
12
|
+
## Inputs
|
|
13
|
+
- 数据集 + 采集溯源(谁、何时、从哪个人群、何种同意)
|
|
14
|
+
- 预先注册的分析计划(不得在看到数据后修改)
|
|
15
|
+
|
|
16
|
+
## Process
|
|
17
|
+
1. **Data Ingestion & Cleaning**: profile columns, check missingness, identify Treatment and Post indicators; the provenance/hash/missingness gate runs before analysis.
|
|
18
|
+
2. **Deterministic DID Regression** (`scripts/did_regression.py`): Y = β0 + β1·Treat + β2·Post + δ·(Treat×Post) + ε; report δ, SE, t, p and Hedges' g (`scripts/effect_calculator.py`).
|
|
19
|
+
3. **Fail closed**: when the design is not estimable (no baseline, no control, attrition beyond tolerance) return `ANALYSIS_NOT_ESTIMABLE` — never fabricate p-values or silently substitute a weaker estimator.
|
|
20
|
+
4. **Graph Re-adjudication**: add a local Evidence Node (`EVD-LOCAL-*`), commit a new revision, re-run the tribunal.
|
|
21
|
+
|
|
22
|
+
## Output Contract
|
|
23
|
+
`analysis-run` + `dataset-manifest`; the graph revision carries the local node; the decision diff is produced by the update workflow.
|
|
24
|
+
|
|
25
|
+
## Visualization Sync
|
|
26
|
+
- `result.json` → `forest_plot_data`: one entry with `study_label` "Local Field Trial (DID)", outcome dimension from the trial, effect size = Hedges' g, CI bounds from the regression.
|
|
27
|
+
- `result.json` → `evidence`: one evidence object whose `relation_to_claim` follows the sign of δ.
|
|
28
|
+
- `evidence_graph.json`: re-export after adding the `EVD-LOCAL-*` node.
|
|
29
|
+
|
|
30
|
+
## Quality Gates
|
|
31
|
+
- [ ] 溯源、哈希、缺失率在分析前记录。
|
|
32
|
+
- [ ] 分析严格按预注册计划执行。
|
|
33
|
+
- [ ] 本地结果作为"一项研究"并入,不覆盖既有证据。
|
|
34
|
+
|
|
35
|
+
## Anti-Patterns
|
|
36
|
+
- 事后改阈值;把不显著说成无效果;用本地单点结果推翻既有证据体。
|
|
37
|
+
|
|
38
|
+
## Worked Example
|
|
39
|
+
两班前后测数据 → DID δ=+0.21(不显著):报告为"本地未复现",不改变原裁决方向,仅下调适用性置信。
|
|
40
|
+
|
|
41
|
+
## References
|
|
42
|
+
- `scripts/did_regression.py`、`references/evaluation-design.md`、`skill/workflows/evaluate-and-update.md`
|
|
@@ -1,25 +1,48 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: ethics-review
|
|
3
|
-
description: "Evaluates trial designs, intervention protocols, and
|
|
3
|
+
description: "Evaluates trial designs, intervention protocols, and human-subject data collection against IRB and research ethics standards (education and other applied domains)."
|
|
4
|
+
capability: (engine-level gate; no deterministic capability)
|
|
4
5
|
---
|
|
5
6
|
|
|
6
|
-
# ethics-review — Research Ethics & IRB Compliance
|
|
7
|
+
# ethics-review — Research Ethics & IRB Compliance
|
|
7
8
|
|
|
8
9
|
## When to Use
|
|
9
|
-
Triggered
|
|
10
|
+
Triggered before finalising any quasi-experimental / DID field trial design involving human student cohorts, classroom telemetry, or control-group assignment.
|
|
10
11
|
|
|
11
|
-
##
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
12
|
+
## Inputs
|
|
13
|
+
- `intervention.json` / study-design 草案(阶段、人群、对照、数据采集范围)
|
|
14
|
+
- 数据采集清单(分数、日志、提示词、遥测)
|
|
15
|
+
|
|
16
|
+
## Process — Ethical Audit Checklist
|
|
17
|
+
1. **Control Group Harm Prevention**: the control group must not be deprived of essential learning opportunities (use delayed crossover or active alternatives).
|
|
18
|
+
2. **Participant Privacy & Telemetry Protection**: pseudonymise prompts, interaction logs, and outcome records; comply with the applicable regime (FERPA / GDPR or the domain's equivalent).
|
|
19
|
+
3. **Informed Consent & Voluntary Participation**: opt-out without academic penalty.
|
|
20
|
+
4. **Algorithmic Bias & Equity Check**: audit whether the tool introduces grading bias or accessibility barriers for underrepresented groups.
|
|
21
|
+
5. **Data Retention**: state what is stored, where, and for how long; commercial LLM endpoints must not retain prompts.
|
|
22
|
+
|
|
23
|
+
## Output Contract
|
|
24
|
+
Ethics review record attached to the study design; a non-passing review blocks the pilot.
|
|
16
25
|
|
|
17
|
-
## Output Schema
|
|
18
26
|
```json
|
|
19
27
|
{
|
|
20
28
|
"ethics_status": "APPROVED_WITH_CONDITIONS",
|
|
21
29
|
"irb_tier": "Exempt / Expedited Educational Research (Category 1)",
|
|
22
30
|
"privacy_safeguards": ["Anonymized student IDs", "Zero prompt retention on commercial LLM endpoints"],
|
|
23
|
-
"equity_protections": "Provide universal
|
|
31
|
+
"equity_protections": "Provide universal campus lab access to eliminate hardware disparities."
|
|
24
32
|
}
|
|
25
33
|
```
|
|
34
|
+
|
|
35
|
+
## Quality Gates
|
|
36
|
+
- [ ] 对照组的可接受替代方案已给出。
|
|
37
|
+
- [ ] 采集字段清单与去标识方式明确。
|
|
38
|
+
- [ ] 退出机制不产生学业惩罚。
|
|
39
|
+
- [ ] 设备/网络差异导致的公平性问题被处理。
|
|
40
|
+
|
|
41
|
+
## Anti-Patterns
|
|
42
|
+
- 以"教改豁免"跳过伦理审查;保留可回指到个人的提示词日志;用"照常上课"掩盖对照组机会剥夺。
|
|
43
|
+
|
|
44
|
+
## Worked Example
|
|
45
|
+
12 周准实验含对照班 → APPROVED_WITH_CONDITIONS:延迟交叉 + 匿名 ID + 关闭端点日志留存 + 统一机房访问。
|
|
46
|
+
|
|
47
|
+
## References
|
|
48
|
+
- `references/evaluation-design.md`、`references/social_science_pitfalls.md`、`skill/task-briefs/intervene.md`
|
|
@@ -1,19 +1,37 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: evidence-extraction
|
|
3
3
|
description: "Extracts fine-grained claims, effect sizes (Hedges g), sample sizes, and methodology variables from validated full-text sources."
|
|
4
|
+
capability: study_extraction + finding_extraction
|
|
4
5
|
---
|
|
6
|
+
|
|
5
7
|
# Evidence Extraction Skill
|
|
6
8
|
|
|
7
|
-
##
|
|
9
|
+
## When to Use
|
|
8
10
|
Trigger on fetched and validated source texts to perform claim-level feature and statistical extraction.
|
|
9
11
|
|
|
10
|
-
##
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
2. **Methodology Extraction**:
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
12
|
+
## Inputs
|
|
13
|
+
- `sources.jsonl`(全部 FETCH_VALID / 确认的 FETCH_PARTIAL)
|
|
14
|
+
- `fetch/` 清正文;必要时经 Sciverse `/content` 读取定位段落
|
|
15
|
+
|
|
16
|
+
## Process
|
|
17
|
+
1. **Statistical Extraction**: sample sizes (N_treatment, N_control); means/SDs; standardized effect size (Hedges' g, Cohen's d, Odds Ratio via `scripts/effect_calculator.py`); 95% CIs and p-values when reported.
|
|
18
|
+
2. **Methodology Extraction**: design type (RCT, quasi-experimental DID/PSM/RDD, correlational); outcome classification (task performance vs conceptual learning vs delayed retention).
|
|
19
|
+
3. **Locate precisely**: keep `source_location` (page/section/offset) so every claim can be re-opened.
|
|
20
|
+
|
|
21
|
+
## Output Contract
|
|
22
|
+
Evidence Objects per `schemas/evidence.schema.json` (V1 top-level, revision 1.1) into `evidence.jsonl`; graph projections use `schemas/v2/evidence-link.schema.json` (V2).
|
|
23
|
+
|
|
24
|
+
## Quality Gates
|
|
25
|
+
- [ ] 每行通过 evidence schema,枚举合法。
|
|
26
|
+
- [ ] 任务表现与学习效果记录分离。
|
|
27
|
+
- [ ] 缺失统计量保持缺失(禁止由显著性反推效应量)。
|
|
28
|
+
- [ ] 每条记录可定位回原文。
|
|
29
|
+
|
|
30
|
+
## Anti-Patterns
|
|
31
|
+
- 把 `relation_to_claim` 写到研究本体;把即测分数当保持/迁移;从摘要估算效应量。
|
|
32
|
+
|
|
33
|
+
## Worked Example
|
|
34
|
+
PNAS 2025 三臂 RCT → 三条 evidence:练习表现(task_performance, support)、独立考试(learning, contradict)、护栏组(learning, support),各带 N 与效应量。
|
|
35
|
+
|
|
36
|
+
## References
|
|
37
|
+
- `references/evidence-quality.md`、`references/effect_size_formulas.md`、`skill/task-briefs/extract.md`
|
|
@@ -1,18 +1,37 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: evidence-review
|
|
3
3
|
description: "Synthesizes extracted claims and evidence nodes into the project Evidence Graph with meta-analysis pooling."
|
|
4
|
+
capability: claim_linking + evidence_synthesis
|
|
4
5
|
---
|
|
6
|
+
|
|
5
7
|
# Evidence Review Skill
|
|
6
8
|
|
|
7
|
-
##
|
|
8
|
-
Trigger to aggregate
|
|
9
|
-
|
|
10
|
-
##
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
9
|
+
## When to Use
|
|
10
|
+
Trigger to aggregate validated evidence nodes into the project's single source of truth (SSOT) Claim Graph and execute quantitative synthesis.
|
|
11
|
+
|
|
12
|
+
## Inputs
|
|
13
|
+
- `evidence.jsonl` + `methodology.json`
|
|
14
|
+
- 项目 Evidence Graph 当前 revision
|
|
15
|
+
|
|
16
|
+
## Process
|
|
17
|
+
1. **Evidence Graph State Machine**: build the directed graph of claims, sources, findings; determine claim status (SUPPORTED / CONTRADICTED / MIXED / UNCERTAIN).
|
|
18
|
+
2. **Meta-Analysis Pooling** (`engine/meta_analysis.py`): fixed-effect inverse-variance pooling and DerSimonian–Laird random-effects pooling; Cochran's Q, df, I².
|
|
19
|
+
3. **Publication Bias & Robustness** (`engine/bias.py` / `engine/robustness.py`): Egger regression, Rosenthal fail-safe N, leave-one-out sensitivity.
|
|
20
|
+
4. **Counting rule**: pool by independent **study**, never by finding count.
|
|
21
|
+
|
|
22
|
+
## Output Contract
|
|
23
|
+
Graph revision (append-only) plus synthesis artifacts; projections (`result.json`, HTML) are derived, never the source of truth.
|
|
24
|
+
|
|
25
|
+
## Quality Gates
|
|
26
|
+
- [ ] 按独立研究计数,非按 finding 计数。
|
|
27
|
+
- [ ] 异质性(I²、Q)与偏倚检验结果一并报告。
|
|
28
|
+
- [ ] 三个方向列未被合并隐藏。
|
|
29
|
+
|
|
30
|
+
## Anti-Patterns
|
|
31
|
+
- 同一研究的多个 outcome 当作多项独立证据;对不可合并的结果强行做标准化合并。
|
|
32
|
+
|
|
33
|
+
## Worked Example
|
|
34
|
+
8 来源 12 条发现 → 仅部分可合并;不可合并者以叙述式证据矩阵呈现并标注原因。
|
|
35
|
+
|
|
36
|
+
## References
|
|
37
|
+
- `references/evidence-quality.md`、`docs/evidence-synthesis.md`、`engine/meta_analysis.py`
|
|
@@ -1,25 +1,47 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: gap-analysis
|
|
3
|
-
description: "Identifies population, measurement, and methodological gaps in the Evidence Graph and diagnoses cross-study empirical contradictions
|
|
3
|
+
description: "Identifies population, measurement, and methodological gaps in the Evidence Graph and diagnoses cross-study empirical contradictions."
|
|
4
|
+
capability: knowledge_gap_detection
|
|
4
5
|
---
|
|
5
6
|
|
|
6
|
-
# gap-analysis — Research Gap Discovery & Contradiction Lens
|
|
7
|
+
# gap-analysis — Research Gap Discovery & Contradiction Lens
|
|
7
8
|
|
|
8
9
|
## When to Use
|
|
9
|
-
Triggered after
|
|
10
|
+
Triggered after evidence extraction and meta-analysis, before study design: trial designs must be grounded on verified empirical gaps, never on generic templates.
|
|
10
11
|
|
|
11
|
-
##
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
12
|
+
## Inputs
|
|
13
|
+
- Evidence Graph(含 claim 状态与 outcome 维度)
|
|
14
|
+
- 矛盾诊断(方向冲突、异质性来源)
|
|
15
|
+
|
|
16
|
+
## Process
|
|
17
|
+
1. **Measurement & Retention Gap Audit**: if evidence only measures immediate task speed, flag the missing delayed unassisted retention measurement.
|
|
18
|
+
2. **Population Heterogeneity Audit**: check whether studies cover only elite CS majors or introductory cohorts; flag advanced-transfer gaps.
|
|
19
|
+
3. **Contradiction Lens**: when directions diverge (g > +0.3 vs g < −0.1), isolate the moderating variable (e.g. scaffolded vs unguided use).
|
|
20
|
+
4. **Priority**: rank gaps by decision-materiality so study design is not spent on immaterial questions.
|
|
21
|
+
|
|
22
|
+
## Output Contract
|
|
23
|
+
`GapNode` list written directly to the SSOT Evidence Graph, each with `gap_id`, `gap_type`, `description`, `target_outcome`, `recommended_trial_design`.
|
|
15
24
|
|
|
16
|
-
## Output Contract (`GapNode` list written directly to SSOT `EvidenceGraph`)
|
|
17
25
|
```json
|
|
18
26
|
{
|
|
19
27
|
"gap_id": "GAP-RETENTION-001",
|
|
20
28
|
"gap_type": "Measurement/Retention Gap",
|
|
21
29
|
"description": "Lack of 12-week longitudinal retention data measuring unassisted transfer in CS1.",
|
|
22
30
|
"target_outcome": "Delayed Unassisted Problem Solving",
|
|
23
|
-
"recommended_trial_design": "12-
|
|
31
|
+
"recommended_trial_design": "12-week cluster randomized trial with 4-week delayed post-test without AI access"
|
|
24
32
|
}
|
|
25
33
|
```
|
|
34
|
+
|
|
35
|
+
## Quality Gates
|
|
36
|
+
- [ ] 每个 gap 绑定具体未满足的测量/人群/方法维度。
|
|
37
|
+
- [ ] gap 与决策材料性关联,而非"文献少"。
|
|
38
|
+
- [ ] 不把"论文数量少"当作研究缺口。
|
|
39
|
+
|
|
40
|
+
## Anti-Patterns
|
|
41
|
+
- 用"研究不足"概括一切;把已有证据能回答的问题标成缺口。
|
|
42
|
+
|
|
43
|
+
## Worked Example
|
|
44
|
+
现有证据多为即测任务表现 → 产出 GAP-RETENTION-001,约束后续 StudyDesign 必须包含无 AI 的延迟后测。
|
|
45
|
+
|
|
46
|
+
## References
|
|
47
|
+
- `engine/gap_lens.py`、`engine/gaps.py`、`references/scientific-invariants.md`
|
|
@@ -1,21 +1,42 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: literature-review
|
|
3
|
-
description: "Executes multi-source academic and web retrieval across OpenAlex, Semantic Scholar, CrossRef, AIHot, AgentSearch, and user-configured providers
|
|
3
|
+
description: "Executes multi-source academic and web retrieval across OpenAlex, Semantic Scholar, CrossRef, AIHot, AgentSearch, Sciverse, and user-configured providers."
|
|
4
|
+
capability: literature_search + counter_evidence_search + source_fetch + source_validation
|
|
4
5
|
---
|
|
6
|
+
|
|
5
7
|
# Literature Review Skill
|
|
6
8
|
|
|
7
|
-
##
|
|
9
|
+
## When to Use
|
|
8
10
|
Trigger after research intent is established to gather candidate empirical studies, peer-reviewed papers, and verified grey literature.
|
|
9
11
|
|
|
10
|
-
##
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
|
|
12
|
+
## Inputs
|
|
13
|
+
- `frame.json`(检索边界、纳排标准)
|
|
14
|
+
- 检索预算(S/M/L 决定)与可选 key:`SCIVERSE_API_TOKEN` / `TAVILY_API_KEY` / `BRAVE_API_KEY`
|
|
15
|
+
|
|
16
|
+
## Process
|
|
17
|
+
1. **Plan first** — 写 `SearchPlan`(core / expansion / **counter_evidence**),经 `retrieval/audit.py` 执行并导出审计四件套。
|
|
18
|
+
2. **Zero-Config Academic Providers**: OpenAlex (250M+ works with DOIs), Semantic Scholar (graph citations + abstracts), CrossRef (DOI registry), AIHot (real-time AI/EdTech feed), AgentSearch / ArXiv.
|
|
19
|
+
3. **Key-based academic channel**: Sciverse (`SCIVERSE_API_TOKEN`) — `/meta-search` 产出 Source 级命中;`/agentic-search` 产出 chunk 定位子,必须经 `/content` 读原文并过校验门,定位写入 `chunks.jsonl`。
|
|
20
|
+
4. **Configured web providers**: Tavily, Brave.
|
|
21
|
+
5. **Fetch & Validation**: 候选 URL 走 `retrieval/fetch.py` 降级链并由 `retrieval/validate.py` 校验;严格拒绝 snippet 幻觉。
|
|
22
|
+
6. **Compliance**: 遵守 `references/retrieval-compliance.md`(robots / 限速 / paywall / 署名与缓存)。
|
|
23
|
+
|
|
24
|
+
## Output Contract
|
|
25
|
+
- `sources.jsonl`(`schemas/source.schema.json`)+ `fetch/`(raw + clean + provenance + fallback_chain)
|
|
26
|
+
- 审计导出:`search-provenance.json` / `search-attempts.jsonl` / `source-screening.csv` / `exclusion-log.csv`
|
|
27
|
+
- Sciverse 定位:`chunks.jsonl`(`locator_state=discovery_only_requires_content_fetch`)
|
|
28
|
+
|
|
29
|
+
## Quality Gates
|
|
30
|
+
- [ ] 反方查询独立构造且计数 > 0。
|
|
31
|
+
- [ ] 每条来源为 FETCH_VALID 或经确认的 FETCH_PARTIAL。
|
|
32
|
+
- [ ] 无 DOI/URL 记录标 `needs_manual_location`,未伪造定位。
|
|
33
|
+
- [ ] 排除记录写明原因。
|
|
34
|
+
|
|
35
|
+
## Anti-Patterns
|
|
36
|
+
- 用支持证据的检索式找反证;把 abstract 当证据内容;为凑数填充无关来源。
|
|
37
|
+
|
|
38
|
+
## Worked Example
|
|
39
|
+
`python scripts/search_provenance.py "first-year CS generative AI coding assistant" --out runs/x/provenance --concept "AI coding assistant"` → 审计四件套 + 候选来源表,随后逐条 fetch/validate。
|
|
40
|
+
|
|
41
|
+
## References
|
|
42
|
+
- `references/retrieval-protocol.md`、`references/retrieval-compliance.md`、`references/source-validity.md`、`skill/task-briefs/retrieve.md`
|
|
@@ -1,20 +1,37 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: methodology-audit
|
|
3
3
|
description: "Audits empirical studies against WWC 5.0, GRADE risk-of-bias frameworks, and social science pitfalls."
|
|
4
|
+
capability: methodology_appraisal
|
|
4
5
|
---
|
|
6
|
+
|
|
5
7
|
# Methodology Audit Skill (Methodology Tribunal)
|
|
6
8
|
|
|
7
|
-
##
|
|
9
|
+
## When to Use
|
|
8
10
|
Trigger before claim synthesis to evaluate threats to internal and external validity, applying methodological confidence scoring.
|
|
9
11
|
|
|
10
|
-
##
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
|
|
15
|
-
|
|
16
|
-
|
|
17
|
-
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
12
|
+
## Inputs
|
|
13
|
+
- `evidence.jsonl` + `fetch/` 原文
|
|
14
|
+
- `references/wwc_standards.md`、`references/grade_framework.md`、`references/social_science_pitfalls.md`
|
|
15
|
+
|
|
16
|
+
## Process
|
|
17
|
+
1. **WWC 5.0 Rating**: Tier 1 (meets standards without reservations — clean RCT, low attrition), Tier 2 (with reservations — QED with baseline equivalence), Tier 3 (correlational / promising).
|
|
18
|
+
2. **Social Science Pitfalls**: task performance ≠ genuine learning; short-term score ≠ retention (4+ weeks); AI-assisted performance ≠ independent transfer; correlation ≠ causation.
|
|
19
|
+
3. **GRADE Certainty**: High / Moderate / Low / Very Low at the body-of-evidence level.
|
|
20
|
+
4. **15-item checklist** in fixed order (control_group … dropout), each `met|partial|missing|not_applicable`.
|
|
21
|
+
|
|
22
|
+
## Output Contract
|
|
23
|
+
`methodology.json` per `schemas/methodology.schema.json`, including `task_vs_learning_guard.equates_task_with_learning`.
|
|
24
|
+
|
|
25
|
+
## Quality Gates
|
|
26
|
+
- [ ] 15 项齐全且仅用四值枚举。
|
|
27
|
+
- [ ] guard 结论明确。
|
|
28
|
+
- [ ] 审计只判"证据是否成立",不判"证据说什么"。
|
|
29
|
+
|
|
30
|
+
## Anti-Patterns
|
|
31
|
+
- 用样本量大掩盖无对照;把相关性研究列入因果结论;缺项直接记 `met`。
|
|
32
|
+
|
|
33
|
+
## Worked Example
|
|
34
|
+
某准实验无前测等价性 → Tier 2 (with reservations) + `pre_test: missing`;其结论只可用于"提示可能",不进强支持列。
|
|
35
|
+
|
|
36
|
+
## References
|
|
37
|
+
- `references/methodology-audit.md`、`references/tribunal-policy.md`、`skill/agents/method-reviewer.md`、`skill/task-briefs/audit.md`
|
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: report-generation
|
|
3
3
|
description: "Renders 5 baked-theme single-file bilingual HTML reports, executive Visual Briefs, and Markdown reports, powered by Lieflat Charts editorial visualization standards and AI-composed, data-driven chart galleries."
|
|
4
|
+
capability: report_projection + report_rendering
|
|
4
5
|
---
|
|
5
6
|
# Report Generation Skill
|
|
6
7
|
|
|
@@ -55,3 +56,30 @@ Dark themes draw on the theme's `card_bg`; text contrast is checked against the
|
|
|
55
56
|
|
|
56
57
|
## 4. Web Studio Sync
|
|
57
58
|
The Local Web Studio (`scripts/dashboard_server.py`) serves the baked HTML reports and Lieflat figures directly.
|
|
59
|
+
|
|
60
|
+
## Inputs
|
|
61
|
+
- `result.json` / `result.zh.json`(叙述字段先过语言门禁 `check_language_parallel`)
|
|
62
|
+
- 当前 Graph Revision 与 decision snapshot 标识
|
|
63
|
+
|
|
64
|
+
## Quality Gates
|
|
65
|
+
- [ ] 双语语义对齐(数字 / ID / 枚举 / URL 不变)。
|
|
66
|
+
- [ ] 渲染完整性门通过:显示数值可回溯到 `result.json`,探针无 `REPORT_INVALID`。
|
|
67
|
+
- [ ] 布局不变量通过 `scripts/lint_report_layout.py`(390 / 768 / 1280 × brief/full)。
|
|
68
|
+
- [ ] 溯源表格保留(来源表 / 证据矩阵 / Claim Trace),图表只作补充。
|
|
69
|
+
- [ ] `artifact_manifest.json` 记录来源 revision 与各产物哈希。
|
|
70
|
+
|
|
71
|
+
## Anti-Patterns
|
|
72
|
+
- 由模型写入图表数值(数值必须来自 `scripts/charts_data.py` 提取器)。
|
|
73
|
+
- 为了"更可视化"而删除可核验表格;用占位或虚构数据补齐缺失图表。
|
|
74
|
+
- 报告与快照不一致时改报告不改结论;在 HTML 内提供运行时换肤。
|
|
75
|
+
|
|
76
|
+
## Failure Handling
|
|
77
|
+
| 失败 | 处理 |
|
|
78
|
+
|---|---|
|
|
79
|
+
| `REPORT_INVALID` | 阻断发布并重跑渲染,不带缺陷投放。 |
|
|
80
|
+
| 双语不对齐 | 修复 `result.zh.json` 后重新烘焙。 |
|
|
81
|
+
| 数据不足以支撑某图 | 抑制该图并记录原因,不补造数据。 |
|
|
82
|
+
| 快照缺失 | 回到上游科学阶段补齐;禁止用默认值生成报告。 |
|
|
83
|
+
|
|
84
|
+
## References
|
|
85
|
+
- `skill/task-briefs/present.md`、`skill/task-briefs/projection.md`、`visualization/eduevidence-report/references/layout-constraints.md`
|