eduevidence 5.2.0 → 6.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (386) hide show
  1. package/CONTRIBUTING.md +105 -0
  2. package/README.md +142 -75
  3. package/README.zh-CN.md +73 -30
  4. package/SKILL.md +397 -131
  5. package/agents/openai.yaml +4 -0
  6. package/assets/readme/controlled-execution.svg +34 -0
  7. package/assets/readme/landing-tour.gif +0 -0
  8. package/assets/readme/logo.png +0 -0
  9. package/assets/readme/research-workflow.svg +56 -0
  10. package/assets/readme/studio-graph.png +0 -0
  11. package/assets/readme/studio-overview.png +0 -0
  12. package/assets/readme/studio-reports.png +0 -0
  13. package/assets/readme/studio-tour.gif +0 -0
  14. package/autoevolve/config.yaml +17 -0
  15. package/autoevolve/program.md +25 -0
  16. package/autoevolve/protected.manifest.yaml +34 -0
  17. package/benchmarks/adversarial/cases.jsonl +7 -0
  18. package/benchmarks/evidence-library.json +5268 -0
  19. package/benchmarks/partitions.json +8 -0
  20. package/bin/eduevidence.js +2 -1
  21. package/docs/architecture.md +496 -0
  22. package/docs/autoresearch-evolution-plan.md +2903 -0
  23. package/docs/autoresearch-implementation-status.md +101 -0
  24. package/docs/demo-storyboard.md +20 -0
  25. package/docs/demo-workplace-ai.md +92 -0
  26. package/docs/demo.md +32 -0
  27. package/docs/install-guide.md +150 -0
  28. package/docs/orchestration-role-model.md +1254 -0
  29. package/docs/release-closeout/README.md +17 -0
  30. package/docs/release-closeout/frontend-acceptance.md +23 -0
  31. package/docs/release-closeout/issues.md +19 -0
  32. package/docs/release-closeout/verification.md +28 -0
  33. package/docs/release-contract.md +108 -0
  34. package/docs/research-studio-guide.zh-CN.md +166 -0
  35. package/docs/sciverse-api.md +125 -0
  36. package/eduevidence_cli.py +29 -13
  37. package/engine/_resources.py +13 -0
  38. package/engine/autoevolve/__init__.py +3 -0
  39. package/engine/autoevolve/agent_view.py +167 -0
  40. package/engine/autoevolve/core.py +357 -0
  41. package/engine/autoevolve/events.py +11 -0
  42. package/engine/autoevolve/git_workspace.py +77 -0
  43. package/engine/autoevolve/projection.py +23 -0
  44. package/engine/autoevolve/runner.py +413 -0
  45. package/engine/autoevolve/trust.py +146 -0
  46. package/engine/autoresearch/__init__.py +6 -0
  47. package/engine/autoresearch/commit.py +132 -0
  48. package/engine/autoresearch/contracts.py +126 -0
  49. package/engine/autoresearch/controller.py +207 -0
  50. package/engine/autoresearch/events.py +12 -0
  51. package/engine/autoresearch/gap_priority.py +168 -0
  52. package/engine/autoresearch/projection.py +30 -0
  53. package/engine/autoresearch/research_memory.py +59 -0
  54. package/engine/autoresearch/saturation.py +91 -0
  55. package/engine/briefs.py +2 -1
  56. package/engine/capabilities.py +1 -0
  57. package/engine/contracts.py +3 -1
  58. package/engine/decision_policy.py +96 -0
  59. package/engine/evidence_graph.py +14 -10
  60. package/engine/evidencecore.py +7 -5
  61. package/engine/gaps.py +132 -73
  62. package/engine/ids.py +2 -0
  63. package/engine/judge_pack.py +65 -0
  64. package/engine/library.py +6 -2
  65. package/engine/library_builtin.py +3 -1
  66. package/engine/living.py +36 -5
  67. package/engine/meta_synthesis.py +3 -1
  68. package/engine/migration.py +88 -3
  69. package/engine/orchestration.py +460 -0
  70. package/engine/paths.py +2 -0
  71. package/engine/pilot.py +36 -33
  72. package/engine/project.py +2 -2
  73. package/engine/research_service.py +113 -0
  74. package/engine/studio_read_model.py +400 -0
  75. package/engine/taxonomy.py +211 -0
  76. package/engine/tribunal.py +44 -33
  77. package/engine/update.py +1 -0
  78. package/engine/versions.py +1 -1
  79. package/engine/worker_result.py +109 -0
  80. package/engine/workflows.py +70 -0
  81. package/examples/ai-coding-assistant-evidence/EduEvidence_Report.html +2934 -0
  82. package/examples/ai-coding-assistant-evidence/artifact_manifest.json +15 -0
  83. package/examples/ai-coding-assistant-evidence/citation_check.json +79 -0
  84. package/examples/ai-coding-assistant-evidence/claims.jsonl +12 -0
  85. package/examples/ai-coding-assistant-evidence/evaluation.json +35 -0
  86. package/examples/ai-coding-assistant-evidence/evidence.jsonl +12 -0
  87. package/examples/ai-coding-assistant-evidence/final_verdict.json +107 -0
  88. package/examples/ai-coding-assistant-evidence/frame.json +48 -0
  89. package/examples/ai-coding-assistant-evidence/gate_report.json +101 -0
  90. package/examples/ai-coding-assistant-evidence/intervention.json +51 -0
  91. package/examples/ai-coding-assistant-evidence/methodology.json +36 -0
  92. package/examples/ai-coding-assistant-evidence/raw_verdict.json +86 -0
  93. package/examples/ai-coding-assistant-evidence/report_spec.json +230 -0
  94. package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_academic.html +2934 -0
  95. package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_claude.html +2934 -0
  96. package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_datalab-dark.html +2934 -0
  97. package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_datalab.html +2934 -0
  98. package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_presentation.html +2934 -0
  99. package/examples/ai-coding-assistant-evidence/reports-5themes/report_academic.html +2934 -0
  100. package/examples/ai-coding-assistant-evidence/reports-5themes/report_claude.html +2934 -0
  101. package/examples/ai-coding-assistant-evidence/reports-5themes/report_datalab-dark.html +2934 -0
  102. package/examples/ai-coding-assistant-evidence/reports-5themes/report_datalab.html +2934 -0
  103. package/examples/ai-coding-assistant-evidence/reports-5themes/report_presentation.html +2934 -0
  104. package/examples/ai-coding-assistant-evidence/result.json +1457 -0
  105. package/examples/ai-coding-assistant-evidence/result.zh.json +1457 -0
  106. package/examples/ai-coding-assistant-evidence/skeptic.json +72 -0
  107. package/examples/ai-coding-assistant-evidence/sources.jsonl +8 -0
  108. package/examples/ai-coding-assistant-evidence/verdict.json +107 -0
  109. package/examples/spaced-retrieval-practice/applicability.json +14 -0
  110. package/examples/spaced-retrieval-practice/artifact_manifest.json +15 -0
  111. package/examples/spaced-retrieval-practice/claims.jsonl +3 -0
  112. package/examples/spaced-retrieval-practice/evidence.jsonl +6 -0
  113. package/examples/spaced-retrieval-practice/final_verdict.json +93 -0
  114. package/examples/spaced-retrieval-practice/frame.json +58 -0
  115. package/examples/spaced-retrieval-practice/gate_report.json +101 -0
  116. package/examples/spaced-retrieval-practice/methodology.json +78 -0
  117. package/examples/spaced-retrieval-practice/report_spec.json +212 -0
  118. package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_academic.html +2728 -0
  119. package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_claude.html +2728 -0
  120. package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_datalab-dark.html +2728 -0
  121. package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_datalab.html +2728 -0
  122. package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_presentation.html +2728 -0
  123. package/examples/spaced-retrieval-practice/reports-5themes/report_academic.html +2728 -0
  124. package/examples/spaced-retrieval-practice/reports-5themes/report_claude.html +2728 -0
  125. package/examples/spaced-retrieval-practice/reports-5themes/report_datalab-dark.html +2728 -0
  126. package/examples/spaced-retrieval-practice/reports-5themes/report_datalab.html +2728 -0
  127. package/examples/spaced-retrieval-practice/reports-5themes/report_presentation.html +2728 -0
  128. package/examples/spaced-retrieval-practice/result.json +942 -0
  129. package/examples/spaced-retrieval-practice/result.zh.json +942 -0
  130. package/examples/spaced-retrieval-practice/skeptic.json +70 -0
  131. package/examples/spaced-retrieval-practice/sources.jsonl +7 -0
  132. package/examples/spaced-retrieval-practice/verdict.json +93 -0
  133. package/examples/workplace-ai-assistant/artifact_manifest.json +15 -0
  134. package/examples/workplace-ai-assistant/claims.jsonl +4 -0
  135. package/examples/workplace-ai-assistant/evaluation.json +19 -0
  136. package/examples/workplace-ai-assistant/evidence.jsonl +4 -0
  137. package/examples/workplace-ai-assistant/evidence_graph.json +444 -0
  138. package/examples/workplace-ai-assistant/final_verdict.json +78 -0
  139. package/examples/workplace-ai-assistant/frame.json +41 -0
  140. package/examples/workplace-ai-assistant/gate_report.json +101 -0
  141. package/examples/workplace-ai-assistant/intervention.json +27 -0
  142. package/examples/workplace-ai-assistant/legacy-link-check.json +16 -0
  143. package/examples/workplace-ai-assistant/methodology.json +60 -0
  144. package/examples/workplace-ai-assistant/report_spec.json +224 -0
  145. package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_academic.html +2814 -0
  146. package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_claude.html +2814 -0
  147. package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_datalab-dark.html +2814 -0
  148. package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_datalab.html +2814 -0
  149. package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_presentation.html +2814 -0
  150. package/examples/workplace-ai-assistant/reports-5themes/report_academic.html +2814 -0
  151. package/examples/workplace-ai-assistant/reports-5themes/report_claude.html +2814 -0
  152. package/examples/workplace-ai-assistant/reports-5themes/report_datalab-dark.html +2814 -0
  153. package/examples/workplace-ai-assistant/reports-5themes/report_datalab.html +2814 -0
  154. package/examples/workplace-ai-assistant/reports-5themes/report_presentation.html +2814 -0
  155. package/examples/workplace-ai-assistant/result.json +615 -0
  156. package/examples/workplace-ai-assistant/result.zh.json +615 -0
  157. package/examples/workplace-ai-assistant/search_log.json +19 -0
  158. package/examples/workplace-ai-assistant/skeptic.json +72 -0
  159. package/examples/workplace-ai-assistant/sources.jsonl +3 -0
  160. package/examples/workplace-ai-assistant/validation_result.json +9 -0
  161. package/examples/workplace-ai-assistant/verdict.json +78 -0
  162. package/install.sh +7 -7
  163. package/integrations/agent_mcp.py +2 -2
  164. package/integrations/orchestration_dispatch.py +146 -0
  165. package/package.json +46 -3
  166. package/pyproject.toml +14 -22
  167. package/references/autoresearch.md +30 -0
  168. package/references/evaluation-policy.md +24 -0
  169. package/references/orchestration.md +22 -0
  170. package/references/report-copy-style.md +67 -0
  171. package/references/retrieval-compliance.md +75 -0
  172. package/references/retrieval-protocol.md +20 -0
  173. package/references/scientific-invariants.md +19 -0
  174. package/retrieval/audit.py +178 -0
  175. package/retrieval/fetch.py +96 -0
  176. package/retrieval/sciverse.py +398 -0
  177. package/retrieval/search.py +47 -7
  178. package/schemas/applicability.schema.json +94 -0
  179. package/schemas/chart-spec.schema.json +10 -3
  180. package/schemas/evidence.schema.json +316 -43
  181. package/schemas/fetch-result.schema.json +2 -1
  182. package/schemas/intervention.schema.json +106 -21
  183. package/schemas/report-result.schema.json +12 -4
  184. package/schemas/report-spec.schema.json +98 -100
  185. package/schemas/skeptic.schema.json +86 -0
  186. package/schemas/source.schema.json +21 -2
  187. package/schemas/v2/finding.schema.json +5 -1
  188. package/schemas/v2/methodology-audit.schema.json +5 -1
  189. package/schemas/v2/outcome.schema.json +28 -5
  190. package/schemas/v2/project.schema.json +2 -2
  191. package/schemas/v2/run.schema.json +1 -1
  192. package/schemas/v2/study.schema.json +5 -1
  193. package/schemas/vNext/autoevolve-session.schema.json +34 -0
  194. package/schemas/vNext/eval-snapshot.schema.json +77 -0
  195. package/schemas/vNext/execution-plan.schema.json +50 -0
  196. package/schemas/vNext/gap-priority.schema.json +54 -0
  197. package/schemas/vNext/negative-search-record.schema.json +68 -0
  198. package/schemas/vNext/research-iteration.schema.json +87 -0
  199. package/schemas/vNext/research-strategy.schema.json +62 -0
  200. package/schemas/vNext/skill-experiment.schema.json +90 -0
  201. package/schemas/vNext/task-spec.schema.json +156 -0
  202. package/schemas/vNext/worker-result.schema.json +60 -0
  203. package/schemas/verdict.schema.json +164 -28
  204. package/scripts/benchmark_judge.py +2 -2
  205. package/scripts/benchmark_v3.py +26 -43
  206. package/scripts/build_esl_artifacts.py +4 -4
  207. package/scripts/build_evidence_library.py +2 -2
  208. package/scripts/build_gh_pages.py +98 -0
  209. package/scripts/build_readme_diagrams.py +72 -0
  210. package/scripts/build_report_variants.py +101 -0
  211. package/scripts/build_result.py +74 -9
  212. package/scripts/check_autoresearch_invariants.py +95 -0
  213. package/scripts/check_package_parity.py +85 -0
  214. package/scripts/check_protocol_alignment.py +375 -0
  215. package/scripts/check_versioned_schemas.py +254 -0
  216. package/scripts/claim_audit.py +13 -8
  217. package/scripts/compute_confidence.py +10 -0
  218. package/scripts/daily_evolve.py +30 -0
  219. package/scripts/dashboard_server.py +130 -101
  220. package/scripts/did_regression.py +17 -32
  221. package/scripts/enrich_projects_human_and_lieflat.py +1 -1
  222. package/scripts/evidence_score.py +5 -2
  223. package/scripts/generate_metrics.py +4 -3
  224. package/scripts/generate_new_projects.py +5 -5
  225. package/scripts/orchestrator.py +286 -36
  226. package/scripts/pre_verdict_gate.py +224 -26
  227. package/scripts/quickstart.py +18 -2
  228. package/scripts/rebake_all_5themes.py +1 -2
  229. package/scripts/research_auto_cli.py +475 -0
  230. package/scripts/run_workspace.py +24 -8
  231. package/scripts/search_provenance.py +64 -0
  232. package/scripts/serve_web.py +9 -10
  233. package/scripts/skill_lint.py +1 -1
  234. package/scripts/skill_payload.py +81 -0
  235. package/scripts/test_adversarial_empirical.py +26 -19
  236. package/scripts/validate_schema.py +46 -2
  237. package/scripts/vnext_cli.py +133 -0
  238. package/setup.py +12 -0
  239. package/skill/agents/evaluation-designer.md +20 -4
  240. package/skill/agents/evidence-analyst.md +19 -3
  241. package/skill/agents/evidence-judge.md +50 -2
  242. package/skill/agents/evidence-retriever.md +20 -3
  243. package/skill/agents/intervention-designer.md +20 -4
  244. package/skill/agents/method-reviewer.md +18 -2
  245. package/skill/agents/{education-planner.md → research-planner.md} +19 -3
  246. package/skill/agents/skeptic.md +18 -2
  247. package/skill/roles/registry.yaml +45 -0
  248. package/skill/sub-skills/aihot-trend-analysis/SKILL.md +28 -9
  249. package/skill/sub-skills/contradiction-analysis/SKILL.md +31 -11
  250. package/skill/sub-skills/data-analysis/SKILL.md +34 -15
  251. package/skill/sub-skills/ethics-review/SKILL.md +33 -10
  252. package/skill/sub-skills/evidence-extraction/SKILL.md +29 -11
  253. package/skill/sub-skills/evidence-review/SKILL.md +31 -12
  254. package/skill/sub-skills/gap-analysis/SKILL.md +31 -9
  255. package/skill/sub-skills/literature-review/SKILL.md +35 -14
  256. package/skill/sub-skills/methodology-audit/SKILL.md +29 -12
  257. package/skill/sub-skills/report-generation/SKILL.md +40 -6
  258. package/skill/sub-skills/research-planning/SKILL.md +41 -14
  259. package/skill/sub-skills/study-design/SKILL.md +30 -9
  260. package/skill/task-briefs/adjudicate.md +32 -7
  261. package/skill/task-briefs/applicability.md +38 -0
  262. package/skill/task-briefs/audit.md +32 -7
  263. package/skill/task-briefs/challenge.md +34 -5
  264. package/skill/task-briefs/evaluate.md +30 -5
  265. package/skill/task-briefs/extract.md +31 -8
  266. package/skill/task-briefs/frame.md +39 -10
  267. package/skill/task-briefs/intervene.md +32 -6
  268. package/skill/task-briefs/present.md +32 -8
  269. package/skill/task-briefs/projection.md +37 -0
  270. package/skill/task-briefs/retrieve.md +36 -6
  271. package/skill/workflows/decision-and-pilot.md +85 -0
  272. package/skill/workflows/evaluate-and-update.md +93 -0
  273. package/skill/workflows/evidence-review.md +117 -0
  274. package/visualization/eduevidence-report/assets/base.css +2 -2
  275. package/visualization/eduevidence-report/assets/reader.css +752 -0
  276. package/visualization/eduevidence-report/assets/reader.js +132 -0
  277. package/visualization/eduevidence-report/references/chart-selection-catalog.md +109 -0
  278. package/visualization/eduevidence-report/references/lieflat-composition.md +3 -1
  279. package/visualization/eduevidence-report/scripts/build_figures.py +25 -3
  280. package/visualization/eduevidence-report/scripts/build_infographics.py +5 -1
  281. package/visualization/eduevidence-report/scripts/build_report.py +561 -121
  282. package/visualization/eduevidence-report/scripts/charts_data.py +2 -0
  283. package/visualization/eduevidence-report/scripts/lieflat_engine.py +371 -136
  284. package/visualization/eduevidence-report/scripts/zh_labels.py +80 -1
  285. package/visualization/eduevidence-report/themes/academic.css +1 -1
  286. package/visualization/eduevidence-report/themes/claude.css +1 -1
  287. package/visualization/eduevidence-report/themes/datalab-dark.css +2 -2
  288. package/visualization/eduevidence-report/themes/datalab.css +2 -2
  289. package/visualization/eduevidence-report/themes/presentation.css +2 -2
  290. package/web/README.md +18 -0
  291. package/web/architecture.html +14885 -0
  292. package/web/index.html +53 -0
  293. package/web/studio/THIRD_PARTY_LICENSES.txt +146 -0
  294. package/web/studio/assets/index-B8tkF44Q.css +1 -0
  295. package/web/studio/assets/index-CQ6Keoyc.js +230 -0
  296. package/web/studio/config.json +1 -0
  297. package/web/studio/index.html +14 -0
  298. package/engine/__pycache__/__init__.cpython-312.pyc +0 -0
  299. package/engine/__pycache__/analysis.cpython-312.pyc +0 -0
  300. package/engine/__pycache__/bias.cpython-312.pyc +0 -0
  301. package/engine/__pycache__/briefs.cpython-312.pyc +0 -0
  302. package/engine/__pycache__/capabilities.cpython-312.pyc +0 -0
  303. package/engine/__pycache__/citation_check.cpython-312.pyc +0 -0
  304. package/engine/__pycache__/contracts.cpython-312.pyc +0 -0
  305. package/engine/__pycache__/datasets.cpython-312.pyc +0 -0
  306. package/engine/__pycache__/events.cpython-312.pyc +0 -0
  307. package/engine/__pycache__/evidence_graph.cpython-312.pyc +0 -0
  308. package/engine/__pycache__/evidence_review.cpython-312.pyc +0 -0
  309. package/engine/__pycache__/evidencecore.cpython-312.pyc +0 -0
  310. package/engine/__pycache__/gap_lens.cpython-312.pyc +0 -0
  311. package/engine/__pycache__/gaps.cpython-312.pyc +0 -0
  312. package/engine/__pycache__/graph_store.cpython-312.pyc +0 -0
  313. package/engine/__pycache__/graph_validate.cpython-312.pyc +0 -0
  314. package/engine/__pycache__/ids.cpython-312.pyc +0 -0
  315. package/engine/__pycache__/library.cpython-312.pyc +0 -0
  316. package/engine/__pycache__/library_builtin.cpython-312.pyc +0 -0
  317. package/engine/__pycache__/living.cpython-312.pyc +0 -0
  318. package/engine/__pycache__/log.cpython-312.pyc +0 -0
  319. package/engine/__pycache__/meta_analysis.cpython-312.pyc +0 -0
  320. package/engine/__pycache__/meta_synthesis.cpython-312.pyc +0 -0
  321. package/engine/__pycache__/migration.cpython-312.pyc +0 -0
  322. package/engine/__pycache__/mode_router.cpython-312.pyc +0 -0
  323. package/engine/__pycache__/paths.cpython-312.pyc +0 -0
  324. package/engine/__pycache__/pilot.cpython-312.pyc +0 -0
  325. package/engine/__pycache__/planner.cpython-312.pyc +0 -0
  326. package/engine/__pycache__/project.cpython-312.pyc +0 -0
  327. package/engine/__pycache__/projections.cpython-312.pyc +0 -0
  328. package/engine/__pycache__/robustness.cpython-312.pyc +0 -0
  329. package/engine/__pycache__/run.cpython-312.pyc +0 -0
  330. package/engine/__pycache__/semantics.cpython-312.pyc +0 -0
  331. package/engine/__pycache__/study_design.cpython-312.pyc +0 -0
  332. package/engine/__pycache__/synthesis.cpython-312.pyc +0 -0
  333. package/engine/__pycache__/tribunal.cpython-312.pyc +0 -0
  334. package/engine/__pycache__/update.cpython-312.pyc +0 -0
  335. package/engine/__pycache__/versions.cpython-312.pyc +0 -0
  336. package/integrations/__pycache__/__init__.cpython-312.pyc +0 -0
  337. package/integrations/__pycache__/agent_mcp.cpython-312.pyc +0 -0
  338. package/integrations/__pycache__/smart_web_fetch.cpython-312.pyc +0 -0
  339. package/retrieval/__pycache__/__init__.cpython-312.pyc +0 -0
  340. package/retrieval/__pycache__/corpus_store.cpython-312.pyc +0 -0
  341. package/retrieval/__pycache__/dedupe.cpython-312.pyc +0 -0
  342. package/retrieval/__pycache__/failures.cpython-312.pyc +0 -0
  343. package/retrieval/__pycache__/fetch.cpython-312.pyc +0 -0
  344. package/retrieval/__pycache__/search.cpython-312.pyc +0 -0
  345. package/retrieval/__pycache__/source.cpython-312.pyc +0 -0
  346. package/retrieval/__pycache__/validate.cpython-312.pyc +0 -0
  347. package/scripts/__pycache__/__init__.cpython-312.pyc +0 -0
  348. package/scripts/__pycache__/benchmark.cpython-312.pyc +0 -0
  349. package/scripts/__pycache__/benchmark_evaluator.cpython-312.pyc +0 -0
  350. package/scripts/__pycache__/benchmark_judge.cpython-312.pyc +0 -0
  351. package/scripts/__pycache__/benchmark_routing.cpython-312.pyc +0 -0
  352. package/scripts/__pycache__/benchmark_v2.cpython-312.pyc +0 -0
  353. package/scripts/__pycache__/benchmark_v3.cpython-312.pyc +0 -0
  354. package/scripts/__pycache__/build_result.cpython-312.pyc +0 -0
  355. package/scripts/__pycache__/claim_audit.cpython-312.pyc +0 -0
  356. package/scripts/__pycache__/complexity_gate.cpython-312.pyc +0 -0
  357. package/scripts/__pycache__/compute_confidence.cpython-312.pyc +0 -0
  358. package/scripts/__pycache__/dashboard_server.cpython-312.pyc +0 -0
  359. package/scripts/__pycache__/did_regression.cpython-312.pyc +0 -0
  360. package/scripts/__pycache__/effect_calculator.cpython-312.pyc +0 -0
  361. package/scripts/__pycache__/evidence_matrix.cpython-312.pyc +0 -0
  362. package/scripts/__pycache__/evidence_score.cpython-312.pyc +0 -0
  363. package/scripts/__pycache__/evidence_semantics.cpython-312.pyc +0 -0
  364. package/scripts/__pycache__/fetch_benchmark.cpython-312.pyc +0 -0
  365. package/scripts/__pycache__/lint_report_layout.cpython-312.pyc +0 -0
  366. package/scripts/__pycache__/orchestrator.cpython-312.pyc +0 -0
  367. package/scripts/__pycache__/pre_verdict_gate.cpython-312.pyc +0 -0
  368. package/scripts/__pycache__/recompute_demo_quality.cpython-312.pyc +0 -0
  369. package/scripts/__pycache__/render_report.cpython-312.pyc +0 -0
  370. package/scripts/__pycache__/render_report_html.cpython-312.pyc +0 -0
  371. package/scripts/__pycache__/run_workspace.cpython-312.pyc +0 -0
  372. package/scripts/__pycache__/skill_lint.cpython-312.pyc +0 -0
  373. package/scripts/__pycache__/startup_probe.cpython-312.pyc +0 -0
  374. package/scripts/__pycache__/sync_killer_demo_report.cpython-312.pyc +0 -0
  375. package/scripts/__pycache__/test_adversarial_empirical.cpython-312-pytest-9.0.2.pyc +0 -0
  376. package/scripts/__pycache__/test_adversarial_empirical.cpython-312-pytest-9.1.1.pyc +0 -0
  377. package/scripts/__pycache__/validate_schema.cpython-312.pyc +0 -0
  378. package/visualization/eduevidence-report/scripts/__pycache__/adapter_contract.cpython-312.pyc +0 -0
  379. package/visualization/eduevidence-report/scripts/__pycache__/build_artifact_manifest.cpython-312.pyc +0 -0
  380. package/visualization/eduevidence-report/scripts/__pycache__/build_charts.cpython-312.pyc +0 -0
  381. package/visualization/eduevidence-report/scripts/__pycache__/build_figures.cpython-312.pyc +0 -0
  382. package/visualization/eduevidence-report/scripts/__pycache__/build_infographics.cpython-312.pyc +0 -0
  383. package/visualization/eduevidence-report/scripts/__pycache__/build_report.cpython-312.pyc +0 -0
  384. package/visualization/eduevidence-report/scripts/__pycache__/charts_data.cpython-312.pyc +0 -0
  385. package/visualization/eduevidence-report/scripts/__pycache__/lieflat_engine.cpython-312.pyc +0 -0
  386. package/visualization/eduevidence-report/scripts/__pycache__/zh_labels.cpython-312.pyc +0 -0
@@ -1,8 +1,10 @@
1
1
  ---
2
2
  name: method-reviewer
3
3
  description: EduEvidence 方法学审查者。按 15 项清单审查每个研究的方法学质量,强制执行"任务完成表现≠学习效果"最高优先级规则,输出 MethodologyAudit。
4
- default_cli: claude
5
- default_model: claude-opus-4-6
4
+ role_id: method-reviewer
5
+ capabilities: methodology_appraisal
6
+ output_contracts: methodology.json (schemas/methodology.schema.json)
7
+ recommended_reasoning: high # capability hint only — no model or CLI name is bound here
6
8
  default_permission: read
7
9
  default_summary_chars: 800
8
10
  default_context_mode: compact
@@ -102,3 +104,17 @@ critical_path: true
102
104
  - 审计说明(note / summary / verdict 理由)为流畅人话(en/zh 分写);PASS / CONCERN / FAIL 只作枚举标签,由显示层映射中文;
103
105
  - 禁止在叙述里堆证据 ID 或 schema 键;引用研究用"作者-年份 + 人话描述";
104
106
  - 无截断残留、无中英夹生。
107
+
108
+ ## 独立性与交叉评审
109
+
110
+ - **独立性要求(`independence_required: role-separation`)**:方法学判断必须独立于内容判断——审计输入只含设计与测量,不含结论评价。此处要求的是角色分离(审计说明不得夹带对效果的评价),而非跨模型家族;需要跨模型家族的只有 Skeptic。
111
+ - 只审"研究怎么测的",不审"结论是什么";审计结论不得夹带对效果的评价。
112
+ - 与宿主的模型选择解耦:本文件只声明能力要求(`recommended_reasoning: high`、高上下文),具体 CLI/模型由用户确认的模型映射决定。
113
+
114
+ ## 失败模式与回退
115
+
116
+ | 失败 | 处理 |
117
+ |---|---|
118
+ | 关键信息未报告(如随机化方式) | 记 `missing` 并写明缺什么,不猜测。 |
119
+ | 结论依赖任务表现 | 触发 guard,剥夺其学习效果支撑资格。 |
120
+ | 审计与内容判断混写 | 拆开重写;审计说明只描述设计与测量。 |
@@ -1,8 +1,10 @@
1
1
  ---
2
- name: education-planner
2
+ name: research-planner
3
3
  description: EduEvidence 教育研究规划者。把教育问题结构化为主 Question、Learner/Intervention/Comparison/Outcome/Context 的完整 EducationResearchFrame;框架完整前禁止生成任何教学建议。
4
- default_cli: claude
5
- default_model: claude-opus-4-6
4
+ role_id: research-planner
5
+ capabilities: research_framing
6
+ output_contracts: frame.json (schemas/education-frame.schema.json)
7
+ recommended_reasoning: high # capability hint only — no model or CLI name is bound here
6
8
  default_permission: read
7
9
  default_summary_chars: 600
8
10
  default_context_mode: compact
@@ -78,3 +80,17 @@ critical_path: true
78
80
  ## 卡住升级
79
81
 
80
82
  问题矛盾或缺少关键信息时回传 `NEEDS_CONTEXT: <缺少什么 + why>`;不臆测学习者特征。
83
+
84
+ ## 独立性与交叉评审
85
+
86
+ - 本角色产出第一道闸门;写入者与复核者分离:Frame 的完整性由 Evidence Judge 在裁决阶段复核,不由本角色自我确认。
87
+ - 关键路径角色(`critical_path: true`):Frame 缺项会阻断整条证据链,宁可 `NEEDS_CONTEXT` 也不填补空白。
88
+ - 与宿主的模型选择解耦:本文件只声明能力要求(`recommended_reasoning: high`),具体 CLI/模型由用户确认的模型映射决定,禁止在此绑定。
89
+
90
+ ## 失败模式与回退
91
+
92
+ | 失败 | 处理 |
93
+ |---|---|
94
+ | 关键输入缺失(学习者层级 / 对照条件 / 主 outcome) | `NEEDS_CONTEXT: <缺什么 + 为何必要>`,不臆测、不继续。 |
95
+ | 问题跨多个决策 | 拆成多个 Frame,各自独立成 run。 |
96
+ | 与用户既有假设冲突 | 在 `extensions` 中记录冲突点,交用户确认后再继续。 |
@@ -1,8 +1,10 @@
1
1
  ---
2
2
  name: skeptic
3
3
  description: EduEvidence 反证挑战者。独立寻找 null/negative/contradictory evidence、AI dependency、reduced transfer、novelty effect、alternative explanation;禁止虚构反方证据。
4
- default_cli: claude
5
- default_model: claude-opus-4-6
4
+ role_id: skeptic
5
+ capabilities: counter_evidence_search
6
+ output_contracts: skeptic.json; cross-model-review (schemas/cross-model-review.schema.json)
7
+ recommended_reasoning: high # capability hint only — no model or CLI name is bound here
6
8
  default_permission: read
7
9
  default_summary_chars: 800
8
10
  default_context_mode: compact
@@ -87,3 +89,17 @@ critical_path: true
87
89
  - 反方证据描述(counter_evidence / null_results / confounders)为面向研究者的流畅中文(en 版为英文);禁止证据 ID 堆砌;
88
90
  - 引用证据用"作者-年份 + 人话描述";禁止把内部字段名(search_performed、risk_level 等)写进叙述;
89
91
  - 无截断残留、无中英夹生。
92
+
93
+ ## 独立性与交叉评审
94
+
95
+ - **独立性要求(`independence_required: true`)**:本角色不得与主分析使用同一模型家族——独立性的目的是让反证来自不同先验,而不是换个会话问同一个模型。
96
+ - 找不到反证时输出标准语句 `NO CONTRADICTORY EVIDENCE FOUND` 并如实标注 `not_found`;宁可空手而归,也不虚构反方文献。
97
+ - 作为交叉审核者时按 `schemas/cross-model-review.schema.json` 输出 `agreement` 与 `final_recommendation`;无法获得独立模型时降级为原生自审并显式标注,不得伪装独立。
98
+
99
+ ## 失败模式与回退
100
+
101
+ | 失败 | 处理 |
102
+ |---|---|
103
+ | 反证检索为空 | 落 negative-search record + 标准语句。 |
104
+ | 反证与支持证据冲突 | 双方都保留,交 Adjudicate 处理;本角色不裁决。 |
105
+ | 无独立模型可用 | 降级为原生自审并标注 `degraded_to: native_self_review`。 |
@@ -0,0 +1,45 @@
1
+ roles:
2
+ research-planner:
3
+ responsibility: framing completeness and scope
4
+ stages: [frame]
5
+ capabilities: [research_framing]
6
+ critical_path: true
7
+ evidence-retriever:
8
+ responsibility: source acquisition and provenance
9
+ stages: [retrieve]
10
+ capabilities: [literature_search, counter_evidence_search, source_fetch, source_validation]
11
+ evidence-analyst:
12
+ responsibility: structured finding extraction
13
+ stages: [extract]
14
+ capabilities: [study_extraction, finding_extraction, claim_linking]
15
+ skeptic:
16
+ responsibility: independent counter-evidence coverage
17
+ stages: [challenge]
18
+ capabilities: [counter_evidence_search]
19
+ independence_required: different-model-family
20
+ critical_path: true
21
+ method-reviewer:
22
+ responsibility: methodology and construct-validity appraisal
23
+ stages: [audit]
24
+ capabilities: [methodology_appraisal]
25
+ independence_required: role-separation
26
+ critical_path: true
27
+ evidence-judge:
28
+ responsibility: evidence-bounded adjudication and applicability
29
+ stages: [adjudicate, applicability]
30
+ capabilities: [evidence_synthesis, tribunal, applicability_analysis, knowledge_gap_detection]
31
+ critical_path: true
32
+ intervention-designer:
33
+ responsibility: grounded intervention or pilot design
34
+ stages: [intervene]
35
+ capabilities: [study_design, measurement_design, intervention_design]
36
+ evaluation-designer:
37
+ responsibility: estimable evaluation and update logic
38
+ stages: [evaluate]
39
+ capabilities: [evaluation_design, data_validation, data_analysis]
40
+
41
+ execution:
42
+ max_parallel_workers: 6
43
+ split_by: evidence_axis
44
+ single_writer: lead-orchestrator
45
+ recursive_swarm: false
@@ -1,31 +1,50 @@
1
1
  ---
2
2
  name: aihot-trend-analysis
3
3
  description: "Real-time horizon scanning and dynamic trend ingestion for emerging AI educational tools, model benchmarks, and EdTech releases via AIHot."
4
+ capability: literature_search (grey-literature channel)
4
5
  ---
5
6
 
6
- # aihot-trend-analysis — Real-Time AI & EdTech Trend Ingestion Sub-Skill
7
+ # aihot-trend-analysis — Real-Time AI & EdTech Trend Ingestion
7
8
 
8
9
  ## When to Use
9
- Triggered when an educational or social science research inquiry involves fast-moving generative AI tools (e.g. Cursor, Claude 3.5, Socratic LLM tutors, Copilot) where peer-reviewed academic literature may have a 6-18 month publication lag.
10
+ Triggered when an inquiry involves fast-moving generative AI tools (Cursor, Claude, Socratic LLM tutors, Copilot) where peer-reviewed literature may lag 6–18 months.
10
11
 
11
- ## Input Requirements
12
- - `keyword`: Target technology or pedagogy topic (e.g. `"AI programming assistant"`, `"Socratic coding tutor"`).
13
- - `time_window`: Optional lookup horizon (`"24h"`, `"7d"`, `"30d"`).
14
- - `category`: `"EdTech"`, `"Agents"`, `"Reasoning"`, `"LLMs"`.
12
+ ## Inputs
13
+ - `keyword`: target technology or pedagogy topic.
14
+ - `time_window` (optional): `24h` / `7d` / `30d`.
15
+ - `category` (optional): `EdTech` / `Agents` / `Reasoning` / `LLMs`.
16
+
17
+ ## Process
18
+ 1. Query the AIHot channel through `retrieval/search.py` (`AIHotProvider`).
19
+ 2. Record every hit as grey literature with its publication time.
20
+ 3. Route any factual claim that would enter the decision back through Retrieve → Fetch → Validate: trend items never bypass RULE 2.
15
21
 
16
22
  ## Output Contract
17
- Returns structured `SearchHit` objects tagged with `provider: "aihot"` and `tier: 5` (grey literature / technical trend), providing zero-day context before empirical trials are designed.
23
+ `SearchHit` objects tagged `provider: "aihot"` with grey-literature authority (`tier5_general_web`); they inform horizon scanning, not effect estimation.
18
24
 
19
25
  ```json
20
26
  {
21
27
  "trend_items": [
22
28
  {
23
- "title": "OpenAI Socratic Tutoring Framework Evaluated Across 10 Universities",
29
+ "title": "Socratic tutoring framework evaluated across 10 universities",
24
30
  "url": "https://aihot.virxact.com/api/item/...",
25
- "summary": "Benchmark evaluation on novice cognitive retention and prompt scaffolding.",
31
+ "summary": "Benchmark evaluation on novice retention and prompt scaffolding.",
26
32
  "category": "EdTech",
27
33
  "publish_time": "2026-08-15"
28
34
  }
29
35
  ]
30
36
  }
31
37
  ```
38
+
39
+ ## Quality Gates
40
+ - [ ] 每条 trend 项带 URL 与时间戳。
41
+ - [ ] 明确标注为灰来源,不进入效应量合成。
42
+
43
+ ## Anti-Patterns
44
+ - 用产品博客宣称的效果当作实证证据;把版本发布日期当研究发表时间。
45
+
46
+ ## Worked Example
47
+ 关键词 "AI programming assistant" → 返回 30 天内的发布与基准报道,用于判断文献滞后期内是否出现新的风险信号。
48
+
49
+ ## References
50
+ - `retrieval/search.py::AIHotProvider`、`references/retrieval-compliance.md`
@@ -1,17 +1,37 @@
1
1
  ---
2
2
  name: contradiction-analysis
3
3
  description: "Mines adversarial claims, conflicting effect directions, and boundary condition qualifiers."
4
+ capability: counter_evidence_search (skeptic side)
4
5
  ---
6
+
5
7
  # Contradiction Analysis Skill
6
8
 
7
- ## 1. When to Use
8
- Trigger during evidence synthesis when studies on the same Claim ID exhibit conflicting directional tags (SUPPORTS vs CONTRADICTS) or high heterogeneity.
9
-
10
- ## 2. Process
11
- 1. **Adversarial Mining (Skeptic)**:
12
- - Identify confounding variables (e.g., teacher training differences, dosage, novelty effect).
13
- - Evaluate boundary conditions: Does intervention fail for novice vs expert learners?
14
- 2. **Directional Separation**:
15
- - Strictly separate evidence into three columns: Supporting, Contradicting, and Neutral.
16
- 3. **Heterogeneity Attribution**:
17
- - Map conflict to subgroup variations, dosage thresholds, or outcome instrument differences.
9
+ ## When to Use
10
+ Trigger during evidence synthesis when studies on the same Claim ID show conflicting directions (SUPPORTS vs CONTRADICTS) or high heterogeneity.
11
+
12
+ ## Inputs
13
+ - `evidence.jsonl`(含方向标签)
14
+ - `frame.json`(判断是否 scope overreach)
15
+
16
+ ## Process
17
+ 1. **Adversarial Mining (Skeptic)**: identify confounders (teacher training, dosage, novelty); evaluate boundary conditions (does it fail for novices vs experts?).
18
+ 2. **Directional Separation**: strictly separate supporting / contradicting / neutral — never blend them into one "mixed" bucket.
19
+ 3. **Heterogeneity Attribution**: map conflict to subgroup variation, dosage thresholds, or outcome instrument differences.
20
+ 4. **Nine fixed checks** per `skill/agents/skeptic.md`; absence of counter-evidence yields the standard statement, never invented sources.
21
+
22
+ ## Output Contract
23
+ `skeptic.json` — `skeptic_findings[]` (`check` / `status` / `detail` / `related_evidence_ids`), `contradictory_evidence_found`, `no_contradictory_evidence_statement`, `threats_to_validity`.
24
+
25
+ ## Quality Gates
26
+ - [ ] 九项检查齐全。
27
+ - [ ] 每条 found 绑定证据或来源。
28
+ - [ ] 标准语句与布尔标志语义一致。
29
+
30
+ ## Anti-Patterns
31
+ - 把三列合并成"总体看有效";为了显得严谨而虚构反证。
32
+
33
+ ## Worked Example
34
+ 同一 claim 下出现 g=+0.48(无护栏练习)与 g=−0.17(独立考试),归因到 outcome 测量差异与护栏配置,而非取平均。
35
+
36
+ ## References
37
+ - `references/skeptic-protocol.md`、`skill/agents/skeptic.md`、`skill/task-briefs/challenge.md`
@@ -1,23 +1,42 @@
1
1
  ---
2
2
  name: data-analysis
3
3
  description: "Runs deterministic statistical regression (DID/OLS) on user-uploaded classroom and field datasets to re-inject local empirical evidence."
4
+ capability: data_validation + data_analysis
4
5
  ---
6
+
5
7
  # Data Analysis Skill
6
8
 
7
- ## 1. When to Use
9
+ ## When to Use
8
10
  Trigger when the user imports empirical classroom or survey data (CSV/XLSX) from an active field trial or pilot deployment.
9
11
 
10
- ## 2. Process
11
- 1. **Data Ingestion & Cleaning**:
12
- - Profile columns, check missingness, identify Treatment and Post indicators.
13
- 2. **Deterministic DID Regression**:
14
- - Run Difference-in-Differences OLS: Y = beta0 + beta1*Treat + beta2*Post + delta*(Treat*Post) + epsilon.
15
- - Calculate treatment effect delta, standard error, t-statistic, p-value, and Hedges' g.
16
- 3. **Graph Re-adjudication**:
17
- - Generate local Evidence Node (EVD-LOCAL-*) and re-run tribunal to update decision snapshot.
18
-
19
- ## 3. Visualization Sync
20
- Append the local trial result to the project's visualization contract so the Data Visualization page reflects it without hard-coding:
21
- - result.json → forest_plot_data: one entry with study_label "Local Field Trial (DID)", outcome_dimension from the trial outcome, effect_size = Hedges' g, ci_lower/ci_upper from the regression CI.
22
- - result.json → evidence: one evidence object with relation_to_claim derived from the DID delta sign.
23
- - evidence_graph.json: re-export after adding the EVD-LOCAL-* node.
12
+ ## Inputs
13
+ - 数据集 + 采集溯源(谁、何时、从哪个人群、何种同意)
14
+ - 预先注册的分析计划(不得在看到数据后修改)
15
+
16
+ ## Process
17
+ 1. **Data Ingestion & Cleaning**: profile columns, check missingness, identify Treatment and Post indicators; the provenance/hash/missingness gate runs before analysis.
18
+ 2. **Deterministic DID Regression** (`scripts/did_regression.py`): Y = β0 + β1·Treat + β2·Post + δ·(Treat×Post) + ε; report δ, SE, t, p and Hedges' g (`scripts/effect_calculator.py`).
19
+ 3. **Fail closed**: when the design is not estimable (no baseline, no control, attrition beyond tolerance) return `ANALYSIS_NOT_ESTIMABLE` — never fabricate p-values or silently substitute a weaker estimator.
20
+ 4. **Graph Re-adjudication**: add a local Evidence Node (`EVD-LOCAL-*`), commit a new revision, re-run the tribunal.
21
+
22
+ ## Output Contract
23
+ `analysis-run` + `dataset-manifest`; the graph revision carries the local node; the decision diff is produced by the update workflow.
24
+
25
+ ## Visualization Sync
26
+ - `result.json` → `forest_plot_data`: one entry with `study_label` "Local Field Trial (DID)", outcome dimension from the trial, effect size = Hedges' g, CI bounds from the regression.
27
+ - `result.json` → `evidence`: one evidence object whose `relation_to_claim` follows the sign of δ.
28
+ - `evidence_graph.json`: re-export after adding the `EVD-LOCAL-*` node.
29
+
30
+ ## Quality Gates
31
+ - [ ] 溯源、哈希、缺失率在分析前记录。
32
+ - [ ] 分析严格按预注册计划执行。
33
+ - [ ] 本地结果作为"一项研究"并入,不覆盖既有证据。
34
+
35
+ ## Anti-Patterns
36
+ - 事后改阈值;把不显著说成无效果;用本地单点结果推翻既有证据体。
37
+
38
+ ## Worked Example
39
+ 两班前后测数据 → DID δ=+0.21(不显著):报告为"本地未复现",不改变原裁决方向,仅下调适用性置信。
40
+
41
+ ## References
42
+ - `scripts/did_regression.py`、`references/evaluation-design.md`、`skill/workflows/evaluate-and-update.md`
@@ -1,25 +1,48 @@
1
1
  ---
2
2
  name: ethics-review
3
- description: "Evaluates trial designs, intervention protocols, and student data collection against Institutional Review Board (IRB) and educational research ethics standards."
3
+ description: "Evaluates trial designs, intervention protocols, and human-subject data collection against IRB and research ethics standards (education and other applied domains)."
4
+ capability: (engine-level gate; no deterministic capability)
4
5
  ---
5
6
 
6
- # ethics-review — Research Ethics & IRB Compliance Sub-Skill
7
+ # ethics-review — Research Ethics & IRB Compliance
7
8
 
8
9
  ## When to Use
9
- Triggered prior to finalizing any 12-week Quasi-Experimental / DID field trial design involving human student cohorts, classroom telemetry, or control group assignment.
10
+ Triggered before finalising any quasi-experimental / DID field trial design involving human student cohorts, classroom telemetry, or control-group assignment.
10
11
 
11
- ## Ethical Audit Checklist
12
- 1. **Control Group Harm Prevention**: Ensures the control group is not deprived of essential pedagogical learning opportunities (recommends delayed crossover or active alternative pedagogies).
13
- 2. **Student Privacy & Telemetry Protection**: Verifies that LLM interaction prompts, code logs, and exam scores are pseudonymized and compliant with FERPA/GDPR educational data privacy.
14
- 3. **Informed Consent & Voluntary Participation**: Enforces opt-out mechanisms without academic penalty.
15
- 4. **Algorithmic Bias & Equity Check**: Audits whether AI tools introduce unfair grading or accessibility barriers for underrepresented student groups.
12
+ ## Inputs
13
+ - `intervention.json` / study-design 草案(阶段、人群、对照、数据采集范围)
14
+ - 数据采集清单(分数、日志、提示词、遥测)
15
+
16
+ ## Process — Ethical Audit Checklist
17
+ 1. **Control Group Harm Prevention**: the control group must not be deprived of essential learning opportunities (use delayed crossover or active alternatives).
18
+ 2. **Participant Privacy & Telemetry Protection**: pseudonymise prompts, interaction logs, and outcome records; comply with the applicable regime (FERPA / GDPR or the domain's equivalent).
19
+ 3. **Informed Consent & Voluntary Participation**: opt-out without academic penalty.
20
+ 4. **Algorithmic Bias & Equity Check**: audit whether the tool introduces grading bias or accessibility barriers for underrepresented groups.
21
+ 5. **Data Retention**: state what is stored, where, and for how long; commercial LLM endpoints must not retain prompts.
22
+
23
+ ## Output Contract
24
+ Ethics review record attached to the study design; a non-passing review blocks the pilot.
16
25
 
17
- ## Output Schema
18
26
  ```json
19
27
  {
20
28
  "ethics_status": "APPROVED_WITH_CONDITIONS",
21
29
  "irb_tier": "Exempt / Expedited Educational Research (Category 1)",
22
30
  "privacy_safeguards": ["Anonymized student IDs", "Zero prompt retention on commercial LLM endpoints"],
23
- "equity_protections": "Provide universal high-speed campus lab access to eliminate socioeconomic hardware disparities."
31
+ "equity_protections": "Provide universal campus lab access to eliminate hardware disparities."
24
32
  }
25
33
  ```
34
+
35
+ ## Quality Gates
36
+ - [ ] 对照组的可接受替代方案已给出。
37
+ - [ ] 采集字段清单与去标识方式明确。
38
+ - [ ] 退出机制不产生学业惩罚。
39
+ - [ ] 设备/网络差异导致的公平性问题被处理。
40
+
41
+ ## Anti-Patterns
42
+ - 以"教改豁免"跳过伦理审查;保留可回指到个人的提示词日志;用"照常上课"掩盖对照组机会剥夺。
43
+
44
+ ## Worked Example
45
+ 12 周准实验含对照班 → APPROVED_WITH_CONDITIONS:延迟交叉 + 匿名 ID + 关闭端点日志留存 + 统一机房访问。
46
+
47
+ ## References
48
+ - `references/evaluation-design.md`、`references/social_science_pitfalls.md`、`skill/task-briefs/intervene.md`
@@ -1,19 +1,37 @@
1
1
  ---
2
2
  name: evidence-extraction
3
3
  description: "Extracts fine-grained claims, effect sizes (Hedges g), sample sizes, and methodology variables from validated full-text sources."
4
+ capability: study_extraction + finding_extraction
4
5
  ---
6
+
5
7
  # Evidence Extraction Skill
6
8
 
7
- ## 1. When to Use
9
+ ## When to Use
8
10
  Trigger on fetched and validated source texts to perform claim-level feature and statistical extraction.
9
11
 
10
- ## 2. Process
11
- 1. **Statistical Extraction**:
12
- - Sample sizes (N_treatment, N_control).
13
- - Means and standard deviations (M1, SD1, M2, SD2).
14
- - Standardized effect size metric: compute Hedges' g, Cohen's d, or Odds Ratio.
15
- - 95% Confidence Intervals [CI_lower, CI_upper] and p-values.
16
- 2. **Methodology Extraction**:
17
- - Design type: RCT, Quasi-Experimental (DID, PSM, RDD), Correlational.
18
- - Outcome classification: Task Performance vs Conceptual Learning vs Delayed Retention.
19
- 3. **Output Contract**: Emit evidence objects per schemas/evidence.schema.json(V1 顶层契约,修订 1.1)into evidence.jsonl; 图谱层投射为 EvidenceLink 时用 schemas/v2/evidence-link.schema.json(V2 契约)。
12
+ ## Inputs
13
+ - `sources.jsonl`(全部 FETCH_VALID / 确认的 FETCH_PARTIAL)
14
+ - `fetch/` 清正文;必要时经 Sciverse `/content` 读取定位段落
15
+
16
+ ## Process
17
+ 1. **Statistical Extraction**: sample sizes (N_treatment, N_control); means/SDs; standardized effect size (Hedges' g, Cohen's d, Odds Ratio via `scripts/effect_calculator.py`); 95% CIs and p-values when reported.
18
+ 2. **Methodology Extraction**: design type (RCT, quasi-experimental DID/PSM/RDD, correlational); outcome classification (task performance vs conceptual learning vs delayed retention).
19
+ 3. **Locate precisely**: keep `source_location` (page/section/offset) so every claim can be re-opened.
20
+
21
+ ## Output Contract
22
+ Evidence Objects per `schemas/evidence.schema.json` (V1 top-level, revision 1.1) into `evidence.jsonl`; graph projections use `schemas/v2/evidence-link.schema.json` (V2).
23
+
24
+ ## Quality Gates
25
+ - [ ] 每行通过 evidence schema,枚举合法。
26
+ - [ ] 任务表现与学习效果记录分离。
27
+ - [ ] 缺失统计量保持缺失(禁止由显著性反推效应量)。
28
+ - [ ] 每条记录可定位回原文。
29
+
30
+ ## Anti-Patterns
31
+ - 把 `relation_to_claim` 写到研究本体;把即测分数当保持/迁移;从摘要估算效应量。
32
+
33
+ ## Worked Example
34
+ PNAS 2025 三臂 RCT → 三条 evidence:练习表现(task_performance, support)、独立考试(learning, contradict)、护栏组(learning, support),各带 N 与效应量。
35
+
36
+ ## References
37
+ - `references/evidence-quality.md`、`references/effect_size_formulas.md`、`skill/task-briefs/extract.md`
@@ -1,18 +1,37 @@
1
1
  ---
2
2
  name: evidence-review
3
3
  description: "Synthesizes extracted claims and evidence nodes into the project Evidence Graph with meta-analysis pooling."
4
+ capability: claim_linking + evidence_synthesis
4
5
  ---
6
+
5
7
  # Evidence Review Skill
6
8
 
7
- ## 1. When to Use
8
- Trigger to aggregate all validated evidence nodes into the project's single source of truth (SSOT) Claim Graph and execute quantitative synthesis.
9
-
10
- ## 2. Process
11
- 1. **Evidence Graph State Machine**:
12
- - Construct directed graph of Claims, Sources, and Findings.
13
- - Determine claim status: SUPPORTED, CONTRADICTED, MIXED, or UNCERTAIN.
14
- 2. **Meta-Analysis Pooling**:
15
- - Fixed-effect inverse-variance pooling and DerSimonian-Laird random-effects pooling.
16
- - Cochran's Q, degrees of freedom, and I^2 heterogeneity index.
17
- - Egger regression and Rosenthal Fail-Safe N publication bias checks.
18
- - Leave-one-out study sensitivity analysis.
9
+ ## When to Use
10
+ Trigger to aggregate validated evidence nodes into the project's single source of truth (SSOT) Claim Graph and execute quantitative synthesis.
11
+
12
+ ## Inputs
13
+ - `evidence.jsonl` + `methodology.json`
14
+ - 项目 Evidence Graph 当前 revision
15
+
16
+ ## Process
17
+ 1. **Evidence Graph State Machine**: build the directed graph of claims, sources, findings; determine claim status (SUPPORTED / CONTRADICTED / MIXED / UNCERTAIN).
18
+ 2. **Meta-Analysis Pooling** (`engine/meta_analysis.py`): fixed-effect inverse-variance pooling and DerSimonian–Laird random-effects pooling; Cochran's Q, df, I².
19
+ 3. **Publication Bias & Robustness** (`engine/bias.py` / `engine/robustness.py`): Egger regression, Rosenthal fail-safe N, leave-one-out sensitivity.
20
+ 4. **Counting rule**: pool by independent **study**, never by finding count.
21
+
22
+ ## Output Contract
23
+ Graph revision (append-only) plus synthesis artifacts; projections (`result.json`, HTML) are derived, never the source of truth.
24
+
25
+ ## Quality Gates
26
+ - [ ] 按独立研究计数,非按 finding 计数。
27
+ - [ ] 异质性(I²、Q)与偏倚检验结果一并报告。
28
+ - [ ] 三个方向列未被合并隐藏。
29
+
30
+ ## Anti-Patterns
31
+ - 同一研究的多个 outcome 当作多项独立证据;对不可合并的结果强行做标准化合并。
32
+
33
+ ## Worked Example
34
+ 8 来源 12 条发现 → 仅部分可合并;不可合并者以叙述式证据矩阵呈现并标注原因。
35
+
36
+ ## References
37
+ - `references/evidence-quality.md`、`docs/evidence-synthesis.md`、`engine/meta_analysis.py`
@@ -1,25 +1,47 @@
1
1
  ---
2
2
  name: gap-analysis
3
- description: "Identifies population, measurement, and methodological gaps in the Evidence Graph and diagnoses cross-study empirical contradictions using BioGapLens PICO taxonomy."
3
+ description: "Identifies population, measurement, and methodological gaps in the Evidence Graph and diagnoses cross-study empirical contradictions."
4
+ capability: knowledge_gap_detection
4
5
  ---
5
6
 
6
- # gap-analysis — Research Gap Discovery & Contradiction Lens Sub-Skill
7
+ # gap-analysis — Research Gap Discovery & Contradiction Lens
7
8
 
8
9
  ## When to Use
9
- Triggered after Evidence Extraction and Meta-Analysis (between Step 8 and Step 9) to ensure that trial designs are strictly grounded on verified empirical gaps rather than generic templates.
10
+ Triggered after evidence extraction and meta-analysis, before study design: trial designs must be grounded on verified empirical gaps, never on generic templates.
10
11
 
11
- ## Execution Rules
12
- 1. **Measurement & Retention Gap Audit**: If evidence only measures immediate task speed (OutcomeDimension: PROCEDURAL_EFFICIENCY), flag a missing delayed unassisted retention gap.
13
- 2. **Population Heterogeneity Audit**: Check if studies exclusively evaluate elite CS majors or introductory cohorts; flag advanced algorithmic transfer gaps.
14
- 3. **Contradiction Lens**: When studies report divergent effect directions ($g > +0.3$ vs $g < -0.1$), isolate the moderating variable (e.g. Socratic scaffolding vs unguided copy-pasting).
12
+ ## Inputs
13
+ - Evidence Graph(含 claim 状态与 outcome 维度)
14
+ - 矛盾诊断(方向冲突、异质性来源)
15
+
16
+ ## Process
17
+ 1. **Measurement & Retention Gap Audit**: if evidence only measures immediate task speed, flag the missing delayed unassisted retention measurement.
18
+ 2. **Population Heterogeneity Audit**: check whether studies cover only elite CS majors or introductory cohorts; flag advanced-transfer gaps.
19
+ 3. **Contradiction Lens**: when directions diverge (g > +0.3 vs g < −0.1), isolate the moderating variable (e.g. scaffolded vs unguided use).
20
+ 4. **Priority**: rank gaps by decision-materiality so study design is not spent on immaterial questions.
21
+
22
+ ## Output Contract
23
+ `GapNode` list written directly to the SSOT Evidence Graph, each with `gap_id`, `gap_type`, `description`, `target_outcome`, `recommended_trial_design`.
15
24
 
16
- ## Output Contract (`GapNode` list written directly to SSOT `EvidenceGraph`)
17
25
  ```json
18
26
  {
19
27
  "gap_id": "GAP-RETENTION-001",
20
28
  "gap_type": "Measurement/Retention Gap",
21
29
  "description": "Lack of 12-week longitudinal retention data measuring unassisted transfer in CS1.",
22
30
  "target_outcome": "Delayed Unassisted Problem Solving",
23
- "recommended_trial_design": "12-Week Cluster Randomized Trial with 4-week delayed post-test without AI access"
31
+ "recommended_trial_design": "12-week cluster randomized trial with 4-week delayed post-test without AI access"
24
32
  }
25
33
  ```
34
+
35
+ ## Quality Gates
36
+ - [ ] 每个 gap 绑定具体未满足的测量/人群/方法维度。
37
+ - [ ] gap 与决策材料性关联,而非"文献少"。
38
+ - [ ] 不把"论文数量少"当作研究缺口。
39
+
40
+ ## Anti-Patterns
41
+ - 用"研究不足"概括一切;把已有证据能回答的问题标成缺口。
42
+
43
+ ## Worked Example
44
+ 现有证据多为即测任务表现 → 产出 GAP-RETENTION-001,约束后续 StudyDesign 必须包含无 AI 的延迟后测。
45
+
46
+ ## References
47
+ - `engine/gap_lens.py`、`engine/gaps.py`、`references/scientific-invariants.md`
@@ -1,21 +1,42 @@
1
1
  ---
2
2
  name: literature-review
3
- description: "Executes multi-source academic and web retrieval across OpenAlex, Semantic Scholar, CrossRef, AIHot, AgentSearch, and user-configured providers (Tavily, Brave, SerpAPI)."
3
+ description: "Executes multi-source academic and web retrieval across OpenAlex, Semantic Scholar, CrossRef, AIHot, AgentSearch, Sciverse, and user-configured providers."
4
+ capability: literature_search + counter_evidence_search + source_fetch + source_validation
4
5
  ---
6
+
5
7
  # Literature Review Skill
6
8
 
7
- ## 1. When to Use
9
+ ## When to Use
8
10
  Trigger after research intent is established to gather candidate empirical studies, peer-reviewed papers, and verified grey literature.
9
11
 
10
- ## 2. Multi-Channel Search Routing
11
- 1. **Zero-Config Academic Providers**:
12
- - OpenAlex: Search 250M+ scholarly works for peer-reviewed studies with DOIs.
13
- - Semantic Scholar: Graph-based citation and abstract retrieval.
14
- - CrossRef: Official DOI resolution and metadata extraction.
15
- - AIHot: Real-time AI research trends and technical reports.
16
- - AgentSearch: SciPhi open academic and vector search.
17
- 2. **Configured API Key Providers**:
18
- - Tavily, Brave Search, SerpAPI, Exa, Bocha.
19
- 3. **Fetch & Validation**:
20
- - Pass candidate URLs through retrieval/fetch.py and validate via retrieval/validate.py.
21
- - Strictly reject ungrounded snippet hallucinations.
12
+ ## Inputs
13
+ - `frame.json`(检索边界、纳排标准)
14
+ - 检索预算(S/M/L 决定)与可选 key:`SCIVERSE_API_TOKEN` / `TAVILY_API_KEY` / `BRAVE_API_KEY`
15
+
16
+ ## Process
17
+ 1. **Plan first** — 写 `SearchPlan`(core / expansion / **counter_evidence**),经 `retrieval/audit.py` 执行并导出审计四件套。
18
+ 2. **Zero-Config Academic Providers**: OpenAlex (250M+ works with DOIs), Semantic Scholar (graph citations + abstracts), CrossRef (DOI registry), AIHot (real-time AI/EdTech feed), AgentSearch / ArXiv.
19
+ 3. **Key-based academic channel**: Sciverse (`SCIVERSE_API_TOKEN`) — `/meta-search` 产出 Source 级命中;`/agentic-search` 产出 chunk 定位子,必须经 `/content` 读原文并过校验门,定位写入 `chunks.jsonl`。
20
+ 4. **Configured web providers**: Tavily, Brave.
21
+ 5. **Fetch & Validation**: 候选 URL 走 `retrieval/fetch.py` 降级链并由 `retrieval/validate.py` 校验;严格拒绝 snippet 幻觉。
22
+ 6. **Compliance**: 遵守 `references/retrieval-compliance.md`(robots / 限速 / paywall / 署名与缓存)。
23
+
24
+ ## Output Contract
25
+ - `sources.jsonl`(`schemas/source.schema.json`)+ `fetch/`(raw + clean + provenance + fallback_chain)
26
+ - 审计导出:`search-provenance.json` / `search-attempts.jsonl` / `source-screening.csv` / `exclusion-log.csv`
27
+ - Sciverse 定位:`chunks.jsonl`(`locator_state=discovery_only_requires_content_fetch`)
28
+
29
+ ## Quality Gates
30
+ - [ ] 反方查询独立构造且计数 > 0。
31
+ - [ ] 每条来源为 FETCH_VALID 或经确认的 FETCH_PARTIAL。
32
+ - [ ] 无 DOI/URL 记录标 `needs_manual_location`,未伪造定位。
33
+ - [ ] 排除记录写明原因。
34
+
35
+ ## Anti-Patterns
36
+ - 用支持证据的检索式找反证;把 abstract 当证据内容;为凑数填充无关来源。
37
+
38
+ ## Worked Example
39
+ `python scripts/search_provenance.py "first-year CS generative AI coding assistant" --out runs/x/provenance --concept "AI coding assistant"` → 审计四件套 + 候选来源表,随后逐条 fetch/validate。
40
+
41
+ ## References
42
+ - `references/retrieval-protocol.md`、`references/retrieval-compliance.md`、`references/source-validity.md`、`skill/task-briefs/retrieve.md`