eduevidence 6.0.0 → 6.3.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +395 -0
- package/CONTRIBUTING.md +105 -0
- package/README.md +113 -49
- package/README.zh-CN.md +39 -12
- package/SKILL.md +15 -5
- package/assets/readme/landing-tour.gif +0 -0
- package/assets/readme/studio-tour.gif +0 -0
- package/benchmarks/evidence-library.json +277 -1
- package/bin/eduevidence.js +2 -1
- package/docs/architecture.md +325 -46
- package/docs/demo-workplace-ai.md +1 -1
- package/docs/install-guide.md +1 -1
- package/docs/j-ev-experimental.md +250 -0
- package/docs/orchestration-role-model.md +1 -1
- package/docs/release-closeout/README.md +1 -1
- package/docs/reproducibility.md +138 -0
- package/docs/sciverse-api.md +125 -0
- package/domains/_neutral/copy/few_shots.json +21 -0
- package/domains/_neutral/copy/framing_lexicon.json +19 -0
- package/domains/_neutral/copy/module_labels.json +5 -0
- package/domains/_neutral/copy/module_labels_footer.json +102 -0
- package/domains/_neutral/copy/module_labels_modules.json +204 -0
- package/domains/_neutral/copy/module_labels_nav.json +126 -0
- package/domains/_neutral/copy/module_labels_summary.json +98 -0
- package/domains/_neutral/copy/module_labels_tables.json +164 -0
- package/domains/_neutral/copy/module_labels_v2.json +90 -0
- package/domains/_neutral/copy/risk_constructs.json +20 -0
- package/domains/_neutral/copy/section_titles.json +66 -0
- package/domains/_neutral/copy/terminology.json +11 -0
- package/domains/check_copy_packs.py +103 -0
- package/domains/education/copy/few_shots.json +22 -0
- package/domains/education/copy/framing_enums.json +167 -0
- package/domains/education/copy/framing_lexicon.json +166 -0
- package/domains/education/copy/module_labels.json +169 -0
- package/domains/education/copy/risk_constructs.json +48 -0
- package/domains/education/copy/section_titles.json +186 -0
- package/domains/education/copy/terminology.json +70 -0
- package/domains/education/manifest.json +1 -1
- package/domains/education/outcome_taxonomy.json +2 -2
- package/domains/manifest.json +1 -1
- package/domains/policy/copy/few_shots.json +22 -0
- package/domains/policy/copy/framing_enums.json +94 -0
- package/domains/policy/copy/framing_lexicon.json +174 -0
- package/domains/policy/copy/module_labels.json +168 -0
- package/domains/policy/copy/risk_constructs.json +33 -0
- package/domains/policy/copy/section_titles.json +186 -0
- package/domains/policy/copy/terminology.json +64 -0
- package/eduevidence_cli.py +10 -0
- package/engine/capabilities.py +57 -5
- package/engine/decision_policy.py +167 -0
- package/engine/evidence_graph.py +14 -10
- package/engine/gaps.py +42 -22
- package/engine/ids.py +2 -0
- package/engine/library.py +6 -2
- package/engine/library_builtin.py +7 -4
- package/engine/living.py +34 -4
- package/engine/migration.py +88 -3
- package/engine/orchestration.py +5 -5
- package/engine/paths.py +2 -0
- package/engine/pilot.py +34 -32
- package/engine/taxonomy.py +211 -0
- package/engine/tribunal.py +49 -43
- package/engine/versions.py +1 -1
- package/examples/ai-coding-assistant-evidence/EduEvidence_Report.html +1361 -147
- package/examples/ai-coding-assistant-evidence/artifact_manifest.json +3 -3
- package/examples/ai-coding-assistant-evidence/citation_check.json +1 -1
- package/examples/ai-coding-assistant-evidence/final_verdict.json +107 -0
- package/examples/ai-coding-assistant-evidence/gate_report.json +101 -0
- package/examples/ai-coding-assistant-evidence/report_spec.json +23 -12
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_academic.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_claude.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_datalab-dark.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_datalab.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/EduEvidence_Report_presentation.html +448 -128
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_academic.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_claude.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_datalab-dark.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_datalab.html +1360 -146
- package/examples/ai-coding-assistant-evidence/reports-5themes/report_presentation.html +1360 -146
- package/examples/ai-coding-assistant-evidence/result.json +13 -9
- package/examples/ai-coding-assistant-evidence/result.zh.json +45 -41
- package/examples/ai-coding-assistant-evidence/skeptic.json +72 -0
- package/examples/ai-coding-assistant-evidence/verdict.json +6 -2
- package/examples/spaced-retrieval-practice/EduEvidence_Report.html +2728 -0
- package/examples/spaced-retrieval-practice/applicability.json +14 -0
- package/examples/spaced-retrieval-practice/artifact_manifest.json +15 -0
- package/examples/spaced-retrieval-practice/claims.jsonl +3 -0
- package/examples/spaced-retrieval-practice/evidence.jsonl +6 -0
- package/examples/spaced-retrieval-practice/final_verdict.json +93 -0
- package/examples/spaced-retrieval-practice/frame.json +58 -0
- package/examples/spaced-retrieval-practice/gate_report.json +101 -0
- package/examples/spaced-retrieval-practice/methodology.json +78 -0
- package/examples/spaced-retrieval-practice/report.html +2522 -0
- package/examples/spaced-retrieval-practice/report_spec.json +212 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_academic.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_claude.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_datalab-dark.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_datalab.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/EduEvidence_Report_presentation.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_academic.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_claude.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_datalab-dark.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_datalab.html +2728 -0
- package/examples/spaced-retrieval-practice/reports-5themes/report_presentation.html +2728 -0
- package/examples/spaced-retrieval-practice/result.json +942 -0
- package/examples/spaced-retrieval-practice/result.zh.json +942 -0
- package/examples/spaced-retrieval-practice/skeptic.json +70 -0
- package/examples/spaced-retrieval-practice/sources.jsonl +7 -0
- package/examples/spaced-retrieval-practice/verdict.json +93 -0
- package/examples/workplace-ai-assistant/EduEvidence_Report.html +2814 -0
- package/examples/workplace-ai-assistant/artifact_manifest.json +15 -0
- package/examples/workplace-ai-assistant/claims.jsonl +4 -4
- package/examples/workplace-ai-assistant/evidence.jsonl +4 -4
- package/examples/workplace-ai-assistant/evidence_graph.json +15 -15
- package/examples/workplace-ai-assistant/final_verdict.json +78 -0
- package/examples/workplace-ai-assistant/gate_report.json +101 -0
- package/examples/workplace-ai-assistant/report_spec.json +209 -40
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_academic.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_claude.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_datalab-dark.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_datalab.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/EduEvidence_Report_presentation.html +449 -119
- package/examples/workplace-ai-assistant/reports-5themes/report_academic.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_claude.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_datalab-dark.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_datalab.html +2814 -0
- package/examples/workplace-ai-assistant/reports-5themes/report_presentation.html +2814 -0
- package/examples/workplace-ai-assistant/result.json +82 -20
- package/examples/workplace-ai-assistant/result.zh.json +82 -20
- package/examples/workplace-ai-assistant/skeptic.json +72 -0
- package/examples/workplace-ai-assistant/verdict.json +36 -10
- package/integrations/agent_mcp.py +2 -2
- package/integrations/jev/__init__.py +115 -0
- package/integrations/jev/approval.py +212 -0
- package/integrations/jev/cli.py +84 -0
- package/integrations/jev/config.py +112 -0
- package/integrations/jev/gateway.py +128 -0
- package/integrations/jev/modes.py +38 -0
- package/integrations/jev/tools_classify.py +88 -0
- package/integrations/jev/tools_extract.py +111 -0
- package/integrations/jev/tools_rerank.py +71 -0
- package/integrations/jev/tools_screen.py +87 -0
- package/integrations/jev/tools_verify.py +95 -0
- package/integrations/jev_mcp.py +22 -0
- package/integrations/semantic_decide.py +286 -0
- package/integrations/semdecide_cli.py +55 -0
- package/package.json +19 -2
- package/pyproject.toml +4 -3
- package/references/report-copy-style.md +107 -0
- package/references/retrieval-compliance.md +75 -0
- package/references/retrieval-protocol.md +20 -0
- package/retrieval/audit.py +27 -3
- package/retrieval/fetch.py +96 -0
- package/retrieval/sciverse.py +398 -0
- package/retrieval/search.py +47 -7
- package/schemas/applicability.schema.json +94 -0
- package/schemas/chart-spec.schema.json +10 -3
- package/schemas/evidence.schema.json +316 -43
- package/schemas/fetch-result.schema.json +2 -1
- package/schemas/report-result.schema.json +3 -3
- package/schemas/report-spec.schema.json +98 -100
- package/schemas/skeptic.schema.json +86 -0
- package/schemas/source.schema.json +21 -2
- package/schemas/v2/decision-snapshot.schema.json +20 -9
- package/schemas/v2/finding.schema.json +5 -1
- package/schemas/v2/intake.schema.json +191 -0
- package/schemas/v2/methodology-audit.schema.json +5 -1
- package/schemas/v2/outcome.schema.json +28 -5
- package/schemas/v2/study.schema.json +5 -1
- package/schemas/vNext/autoevolve-session.schema.json +34 -1
- package/schemas/vNext/eval-snapshot.schema.json +77 -1
- package/schemas/vNext/execution-plan.schema.json +50 -1
- package/schemas/vNext/gap-priority.schema.json +54 -1
- package/schemas/vNext/negative-search-record.schema.json +68 -1
- package/schemas/vNext/research-iteration.schema.json +87 -1
- package/schemas/vNext/research-strategy.schema.json +62 -1
- package/schemas/vNext/skill-experiment.schema.json +90 -1
- package/schemas/vNext/task-spec.schema.json +156 -1
- package/schemas/vNext/worker-result.schema.json +60 -1
- package/schemas/verdict.schema.json +164 -28
- package/scripts/build_evidence_library.py +15 -5
- package/scripts/build_report_variants.py +18 -2
- package/scripts/build_result.py +74 -9
- package/scripts/check_package_parity.py +85 -0
- package/scripts/check_protocol_alignment.py +375 -0
- package/scripts/check_versioned_schemas.py +254 -0
- package/scripts/claim_audit.py +13 -8
- package/scripts/compute_confidence.py +10 -0
- package/scripts/dashboard_server.py +13 -2
- package/scripts/did_regression.py +12 -2
- package/scripts/evidence_score.py +5 -2
- package/scripts/intake/__init__.py +31 -0
- package/scripts/intake/__main__.py +18 -0
- package/scripts/intake/background.py +78 -0
- package/scripts/intake/browser.py +79 -0
- package/scripts/intake/cli.py +57 -0
- package/scripts/intake/constants.py +57 -0
- package/scripts/intake/depth.py +53 -0
- package/scripts/intake/enhancements.py +106 -0
- package/scripts/intake/hooks.py +90 -0
- package/scripts/intake/prefs.py +76 -0
- package/scripts/intake/prompts.py +85 -0
- package/scripts/intake/session.py +152 -0
- package/scripts/lint_file_layers.py +126 -0
- package/scripts/orchestrator.py +187 -40
- package/scripts/pre_verdict_gate.py +241 -29
- package/scripts/quickstart.py +18 -2
- package/scripts/run_workspace.py +7 -1
- package/scripts/skill_lint.py +11 -1
- package/scripts/skill_payload.py +6 -3
- package/scripts/test_adversarial_empirical.py +96 -25
- package/scripts/validate_schema.py +31 -1
- package/skill/agents/evaluation-designer.md +20 -4
- package/skill/agents/evidence-analyst.md +19 -3
- package/skill/agents/evidence-judge.md +98 -8
- package/skill/agents/evidence-retriever.md +20 -3
- package/skill/agents/intervention-designer.md +20 -4
- package/skill/agents/method-reviewer.md +18 -2
- package/skill/agents/{education-planner.md → research-planner.md} +19 -3
- package/skill/agents/skeptic.md +18 -2
- package/skill/roles/registry.yaml +11 -11
- package/skill/sub-skills/aihot-trend-analysis/SKILL.md +28 -9
- package/skill/sub-skills/contradiction-analysis/SKILL.md +31 -11
- package/skill/sub-skills/data-analysis/SKILL.md +34 -15
- package/skill/sub-skills/ethics-review/SKILL.md +33 -10
- package/skill/sub-skills/evidence-extraction/SKILL.md +29 -11
- package/skill/sub-skills/evidence-review/SKILL.md +31 -12
- package/skill/sub-skills/gap-analysis/SKILL.md +31 -9
- package/skill/sub-skills/literature-review/SKILL.md +35 -14
- package/skill/sub-skills/methodology-audit/SKILL.md +29 -12
- package/skill/sub-skills/report-generation/SKILL.md +28 -0
- package/skill/sub-skills/research-planning/SKILL.md +41 -14
- package/skill/sub-skills/study-design/SKILL.md +30 -9
- package/skill/task-briefs/adjudicate.md +32 -7
- package/skill/task-briefs/applicability.md +37 -2
- package/skill/task-briefs/audit.md +32 -7
- package/skill/task-briefs/challenge.md +34 -5
- package/skill/task-briefs/evaluate.md +30 -5
- package/skill/task-briefs/extract.md +31 -8
- package/skill/task-briefs/frame.md +39 -10
- package/skill/task-briefs/intervene.md +32 -6
- package/skill/task-briefs/present.md +32 -8
- package/skill/task-briefs/projection.md +36 -2
- package/skill/task-briefs/retrieve.md +36 -6
- package/skill/workflows/decision-and-pilot.md +76 -1
- package/skill/workflows/evaluate-and-update.md +83 -0
- package/skill/workflows/evidence-review.md +104 -0
- package/skill/workflows/experimental-jev.md +170 -0
- package/skill/workflows/intake.md +120 -0
- package/visualization/eduevidence-report/scripts/build_figures.py +25 -3
- package/visualization/eduevidence-report/scripts/build_infographics.py +37 -15
- package/visualization/eduevidence-report/scripts/build_report.py +435 -575
- package/visualization/eduevidence-report/scripts/charts_data.py +2 -0
- package/visualization/eduevidence-report/scripts/lieflat_engine.py +349 -38
- package/visualization/eduevidence-report/scripts/report_copy_pack.py +296 -0
- package/visualization/eduevidence-report/scripts/report_copy_policy_guard.py +47 -0
- package/visualization/eduevidence-report/scripts/zh_labels.py +141 -1
- package/web/architecture.html +14885 -0
- package/web/studio/assets/index-B8tkF44Q.css +1 -0
- package/web/studio/index.html +2 -2
- package/scripts/build_esl_artifacts.py +0 -1921
- package/scripts/build_killer_demo.py +0 -295
- package/scripts/enrich_projects_human_and_lieflat.py +0 -315
- package/scripts/generate_new_projects.py +0 -686
- package/scripts/sync_killer_demo_report.py +0 -270
- package/web/studio/assets/index-CzXocaGv.css +0 -1
- /package/web/studio/assets/{index-pa7jD7n4.js → index-CQ6Keoyc.js} +0 -0
|
@@ -44,7 +44,6 @@ from engine.evidence_graph import (
|
|
|
44
44
|
)
|
|
45
45
|
from engine.gap_lens import gap_lens
|
|
46
46
|
from engine.semantics import OutcomeClassifier, OutcomeDimension
|
|
47
|
-
from engine.tribunal import _confidence, _decision_action, _study_implication
|
|
48
47
|
from retrieval.corpus_store import DomainCorpusStore, corpus_store
|
|
49
48
|
from retrieval.search import search_evidence, search_router
|
|
50
49
|
from scripts.did_regression import run_did_analysis
|
|
@@ -87,9 +86,8 @@ def test_eventbus_concurrency():
|
|
|
87
86
|
t.join()
|
|
88
87
|
|
|
89
88
|
print(f"[*] Concurrent subscribe/unsubscribe completed. Race errors caught: {len(race_errors)}")
|
|
90
|
-
|
|
91
|
-
|
|
92
|
-
print(f" - [BUG FOUND] {err}")
|
|
89
|
+
assert not race_errors, (
|
|
90
|
+
f"EventBus raised under concurrent subscribe/unsubscribe: {race_errors[:3]}")
|
|
93
91
|
|
|
94
92
|
# 2. Test Subscribe Non-Atomic Check (Duplicate Subscriber Appending)
|
|
95
93
|
bus._subscribers.clear()
|
|
@@ -105,8 +103,9 @@ def test_eventbus_concurrency():
|
|
|
105
103
|
t.join()
|
|
106
104
|
|
|
107
105
|
print(f"[*] Subscribed same callback across 10 threads. Total registered subscribers: {len(bus._subscribers)} (Expected: 1)")
|
|
108
|
-
|
|
109
|
-
|
|
106
|
+
assert len(bus._subscribers) == 1, (
|
|
107
|
+
f"subscribe() is not atomic: {len(bus._subscribers)} copies of one callback "
|
|
108
|
+
"registered from 10 threads")
|
|
110
109
|
|
|
111
110
|
# 3. Concurrent Publish & Unbounded Memory Leak
|
|
112
111
|
bus.clear()
|
|
@@ -156,21 +155,23 @@ def test_did_regression_adversarial():
|
|
|
156
155
|
results = {}
|
|
157
156
|
|
|
158
157
|
# Case 2.1: Column Name Parsing Collision
|
|
158
|
+
# Enough rows for a DID fit (n - 4 > 0): this case is about COLUMN MAPPING,
|
|
159
|
+
# not about saturation, so a 4-row fixture would test the wrong thing.
|
|
159
160
|
with tempfile.NamedTemporaryFile("w", suffix=".csv", delete=False) as f:
|
|
160
161
|
writer = csv.writer(f)
|
|
161
162
|
writer.writerow(["student_id", "treatment_group", "post_test_score", "time_period"])
|
|
162
|
-
|
|
163
|
-
|
|
164
|
-
|
|
165
|
-
|
|
163
|
+
for i in range(1, 13):
|
|
164
|
+
treat = 1 if i % 2 else 0
|
|
165
|
+
post = 1 if i % 3 else 0
|
|
166
|
+
writer.writerow([i, treat, 60.0 + 4 * treat + 3 * post + 2 * treat * post, post])
|
|
166
167
|
col_test_path = f.name
|
|
167
168
|
|
|
168
169
|
res_col = run_did_analysis(col_test_path)
|
|
169
170
|
os.remove(col_test_path)
|
|
170
171
|
print(f"[*] Case 2.1: Column name collision ('treatment_group', 'post_test_score', 'time_period'):")
|
|
171
172
|
print(f" Result: {res_col}")
|
|
172
|
-
|
|
173
|
-
|
|
173
|
+
assert res_col.get("status") != "error", (
|
|
174
|
+
"DID column mapper failed to parse the outcome column (keyword collision)")
|
|
174
175
|
results["column_mapping_bug"] = res_col
|
|
175
176
|
|
|
176
177
|
# Case 2.2: Perfect Multicollinearity / Singular Design Matrix
|
|
@@ -187,8 +188,10 @@ def test_did_regression_adversarial():
|
|
|
187
188
|
os.remove(collinear_path)
|
|
188
189
|
print(f"\n[*] Case 2.2: Singular Matrix / Perfect Multicollinearity:")
|
|
189
190
|
print(f" Result: {json.dumps(res_coll, indent=2)}")
|
|
190
|
-
|
|
191
|
-
|
|
191
|
+
assert not (res_coll.get("status") == "success"
|
|
192
|
+
and res_coll.get("standard_error") == 1.0), (
|
|
193
|
+
"singular design matrix produced a fabricated standard error of 1.0 "
|
|
194
|
+
"instead of reporting collinearity")
|
|
192
195
|
results["singular_matrix_fallback"] = res_coll
|
|
193
196
|
|
|
194
197
|
# Case 2.3: Zero Variance in Outcome
|
|
@@ -228,8 +231,9 @@ def test_did_regression_adversarial():
|
|
|
228
231
|
# Case 2.5: WWC Baseline Equivalence Rating for QED
|
|
229
232
|
print(f"\n[*] Case 2.5: WWC 5.0 Baseline Rating Check for QED:")
|
|
230
233
|
print(f" When baseline_equivalence_g = {res_sat.get('baseline_equivalence_g')}, WWC rating reported is: '{res_sat.get('wwc_baseline_rating')}'")
|
|
231
|
-
|
|
232
|
-
|
|
234
|
+
assert res_sat.get("wwc_baseline_rating") != "Meets Standards Without Reservations", (
|
|
235
|
+
"methodological violation: a quasi-experimental design was rated as meeting "
|
|
236
|
+
"WWC standards without reservations")
|
|
233
237
|
results["wwc_rating_bug"] = res_sat.get("wwc_baseline_rating")
|
|
234
238
|
|
|
235
239
|
# Case 2.6: Small Sample Normal Z-Test vs Student t-distribution
|
|
@@ -307,8 +311,8 @@ def test_evidence_graph_adversarial():
|
|
|
307
311
|
print(f"[*] Meta-synthesis with NaN effect and negative weight:")
|
|
308
312
|
print(f" - Pooled g: {proc_syn.get('pooled_g')}")
|
|
309
313
|
print(f" - Q statistic: {proc_syn.get('q_statistic')}")
|
|
310
|
-
|
|
311
|
-
|
|
314
|
+
assert not math.isnan(proc_syn.get("pooled_g", 0.0)), (
|
|
315
|
+
"NaN effect size propagated into meta-synthesis without validation")
|
|
312
316
|
|
|
313
317
|
return {
|
|
314
318
|
"json_roundtrip": len(reloaded.edges),
|
|
@@ -439,15 +443,21 @@ def test_offline_corpus_and_network_isolation():
|
|
|
439
443
|
def test_dashboard_server_adversarial():
|
|
440
444
|
print_section("TEST 6: Dashboard Server Concurrency, SSE Disconnection & Source Leakage")
|
|
441
445
|
|
|
442
|
-
import
|
|
446
|
+
import http.server
|
|
443
447
|
from scripts.dashboard_server import StudioHandler
|
|
444
448
|
|
|
445
|
-
class
|
|
446
|
-
#
|
|
449
|
+
class _AdversarialServer(http.server.ThreadingHTTPServer):
|
|
450
|
+
# Exercise the class the shipped server actually runs
|
|
451
|
+
# (scripts/dashboard_server.py run_dashboard_server -> ThreadingHTTPServer).
|
|
452
|
+
# A single-threaded socketserver.TCPServer answers a 30-client burst from a
|
|
453
|
+
# 5-slot accept queue, and the kernel resets the overflow ([Errno 54]
|
|
454
|
+
# on macOS). That harness choice - not the product - caused the resets.
|
|
447
455
|
allow_reuse_address = True
|
|
456
|
+
request_queue_size = 128 # the 30 simultaneous connects below must fit
|
|
457
|
+
daemon_threads = True
|
|
448
458
|
|
|
449
459
|
test_port = 0 # 临时端口:避免与常驻服务/上次残留监听冲突
|
|
450
|
-
server =
|
|
460
|
+
server = _AdversarialServer(("127.0.0.1", test_port), StudioHandler)
|
|
451
461
|
test_port = server.server_address[1]
|
|
452
462
|
|
|
453
463
|
t = threading.Thread(target=server.serve_forever, daemon=True)
|
|
@@ -495,8 +505,10 @@ def test_dashboard_server_adversarial():
|
|
|
495
505
|
except Exception as e:
|
|
496
506
|
leakage_results[tf] = False
|
|
497
507
|
|
|
498
|
-
|
|
499
|
-
|
|
508
|
+
leaked = [path for path, ok in leakage_results.items() if ok]
|
|
509
|
+
assert not leaked, (
|
|
510
|
+
"StudioHandler served local source files over unauthenticated GET: "
|
|
511
|
+
f"{leaked}")
|
|
500
512
|
results["file_leakage"] = leakage_results
|
|
501
513
|
|
|
502
514
|
# 3. Concurrency Stress Test (30 Concurrent HTTP Clients)
|
|
@@ -510,9 +522,68 @@ def test_dashboard_server_adversarial():
|
|
|
510
522
|
futures = [executor.submit(fetch_data, i) for i in range(30)]
|
|
511
523
|
statuses = [f.result() for f in futures]
|
|
512
524
|
elapsed = time.time() - t0
|
|
525
|
+
assert statuses == [200] * 30, (
|
|
526
|
+
"30 concurrent /api/projects requests did not all succeed: "
|
|
527
|
+
f"statuses={sorted(set(statuses))}")
|
|
513
528
|
print(f" - 30 concurrent requests finished in {elapsed:.3f}s (all status: 200).")
|
|
514
|
-
print(f" - Note:
|
|
529
|
+
print(f" - Note: the shipped server is ThreadingHTTPServer; requests are served concurrently.")
|
|
515
530
|
results["concurrency_30_time"] = elapsed
|
|
531
|
+
results["concurrency_30_statuses"] = statuses
|
|
532
|
+
|
|
533
|
+
# 4. Host / CSRF guards (dashboard_server.py do_GET) must fail closed on
|
|
534
|
+
# the projection API, without blocking legitimate same-origin readers.
|
|
535
|
+
def raw_get(request: str) -> tuple[int, str]:
|
|
536
|
+
"""Send one raw HTTP request; return (status_code, body snippet)."""
|
|
537
|
+
sock = socket.create_connection(("127.0.0.1", test_port), timeout=5)
|
|
538
|
+
sock.settimeout(3)
|
|
539
|
+
try:
|
|
540
|
+
sock.sendall(request.encode("latin-1"))
|
|
541
|
+
buf = b""
|
|
542
|
+
while len(buf) < 16384:
|
|
543
|
+
try:
|
|
544
|
+
chunk = sock.recv(4096)
|
|
545
|
+
except socket.timeout:
|
|
546
|
+
break
|
|
547
|
+
if not chunk:
|
|
548
|
+
break
|
|
549
|
+
buf += chunk
|
|
550
|
+
finally:
|
|
551
|
+
sock.close()
|
|
552
|
+
head, _, body = buf.partition(b"\r\n\r\n")
|
|
553
|
+
first_line = head.split(b"\r\n", 1)[0].decode("latin-1", errors="ignore")
|
|
554
|
+
fields = first_line.split(" ")
|
|
555
|
+
status = int(fields[1]) if len(fields) > 1 and fields[1].isdigit() else 0
|
|
556
|
+
return status, body.decode("utf-8", errors="ignore")[:200]
|
|
557
|
+
|
|
558
|
+
untrusted_status, untrusted_body = raw_get(
|
|
559
|
+
"GET /api/projects HTTP/1.1\r\nHost: evil.example.com\r\n\r\n")
|
|
560
|
+
assert untrusted_status == 403, (
|
|
561
|
+
"an untrusted Host header reached the projection API: "
|
|
562
|
+
f"status={untrusted_status}")
|
|
563
|
+
assert "untrusted host" in untrusted_body, untrusted_body
|
|
564
|
+
|
|
565
|
+
csrf_status, csrf_body = raw_get(
|
|
566
|
+
f"GET /api/projects HTTP/1.1\r\nHost: 127.0.0.1:{test_port}\r\n"
|
|
567
|
+
"Sec-Fetch-Site: cross-site\r\n\r\n")
|
|
568
|
+
assert csrf_status == 403, (
|
|
569
|
+
f"a cross-site request reached the projection API: status={csrf_status}")
|
|
570
|
+
assert "cross-origin" in csrf_body, csrf_body
|
|
571
|
+
|
|
572
|
+
same_origin_status, _ = raw_get(
|
|
573
|
+
f"GET /api/projects HTTP/1.1\r\nHost: 127.0.0.1:{test_port}\r\n"
|
|
574
|
+
"Sec-Fetch-Site: same-origin\r\n\r\n")
|
|
575
|
+
assert same_origin_status == 200, (
|
|
576
|
+
f"a legitimate same-origin reader was blocked: status={same_origin_status}")
|
|
577
|
+
|
|
578
|
+
print("\n[*] Host / CSRF guard probes:")
|
|
579
|
+
print(f" - untrusted Host -> {untrusted_status} ({untrusted_body.strip()})")
|
|
580
|
+
print(f" - Sec-Fetch-Site: cross-site -> {csrf_status} ({csrf_body.strip()})")
|
|
581
|
+
print(f" - Sec-Fetch-Site: same-origin -> {same_origin_status}")
|
|
582
|
+
results["host_csrf_guard"] = {
|
|
583
|
+
"untrusted_host": untrusted_status,
|
|
584
|
+
"cross_site": csrf_status,
|
|
585
|
+
"same_origin": same_origin_status,
|
|
586
|
+
}
|
|
516
587
|
|
|
517
588
|
finally:
|
|
518
589
|
server.shutdown()
|
|
@@ -3,7 +3,8 @@
|
|
|
3
3
|
|
|
4
4
|
Zero-dependency JSON Schema (draft-07 subset) validator covering the constructs
|
|
5
5
|
used by schemas/*.schema.json: $id, title, description, type, properties,
|
|
6
|
-
required, enum, minimum, maximum, minLength, additionalProperties, $ref
|
|
6
|
+
required, enum, minimum, maximum, minLength, additionalProperties, $ref,
|
|
7
|
+
anyOf / oneOf / allOf
|
|
7
8
|
(local #/definitions and relative-file references), const, format (uri,
|
|
8
9
|
date-time), pattern.
|
|
9
10
|
|
|
@@ -138,6 +139,35 @@ class Validator:
|
|
|
138
139
|
self.validate(value, self._resolve_ref(ref, path), path)
|
|
139
140
|
return
|
|
140
141
|
|
|
142
|
+
# Combinators (draft-07 subset): anyOf / oneOf / allOf. Without
|
|
143
|
+
# these a schema can express a real alternative shape and be
|
|
144
|
+
# silently ignored - which is how report-spec (two accepted
|
|
145
|
+
# shapes) and visual-layout (oneOf) became dead weight.
|
|
146
|
+
for key in ("anyOf", "oneOf", "allOf"):
|
|
147
|
+
subschemas = schema.get(key)
|
|
148
|
+
if not isinstance(subschemas, list) or not subschemas:
|
|
149
|
+
continue
|
|
150
|
+
matched = 0
|
|
151
|
+
first_error = None
|
|
152
|
+
for sub in subschemas:
|
|
153
|
+
try:
|
|
154
|
+
self.validate(value, sub, path)
|
|
155
|
+
matched += 1
|
|
156
|
+
except SchemaError as exc:
|
|
157
|
+
if first_error is None:
|
|
158
|
+
first_error = exc
|
|
159
|
+
if key == "allOf" and matched != len(subschemas):
|
|
160
|
+
raise first_error or SchemaError(
|
|
161
|
+
f"{path}: allOf not satisfied")
|
|
162
|
+
if key in ("anyOf", "oneOf") and matched == 0:
|
|
163
|
+
detail = f" (first: {first_error})" if first_error else ""
|
|
164
|
+
raise SchemaError(
|
|
165
|
+
f"{path}: matches none of the {key} alternatives" + detail)
|
|
166
|
+
if key == "oneOf" and matched > 1:
|
|
167
|
+
raise SchemaError(
|
|
168
|
+
f"{path}: matches {matched} oneOf alternatives "
|
|
169
|
+
"(exactly one required)")
|
|
170
|
+
|
|
141
171
|
if "type" in schema:
|
|
142
172
|
types = schema["type"]
|
|
143
173
|
if isinstance(types, str):
|
|
@@ -1,8 +1,10 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: evaluation-designer
|
|
3
3
|
description: EduEvidence 效果评价设计者。为任何 PILOT/ADOPT 建议附 EvaluationPlan:基线/后测/保持/迁移 + 过程/学习/风险指标 + 成功阈值与停止条件;区分任务表现与学习效果。
|
|
4
|
-
|
|
5
|
-
|
|
4
|
+
role_id: evaluation-designer
|
|
5
|
+
capabilities: evaluation_design, data_validation, data_analysis
|
|
6
|
+
output_contracts: evaluation.json (schemas/evaluation.schema.json)
|
|
7
|
+
recommended_reasoning: high # capability hint only — no model or CLI name is bound here
|
|
6
8
|
default_permission: read
|
|
7
9
|
default_summary_chars: 800
|
|
8
10
|
default_context_mode: compact
|
|
@@ -54,7 +56,7 @@ critical_path: false
|
|
|
54
56
|
- `groups` 是对象,必须含 `treatment` 与 `comparison` 两个字段;
|
|
55
57
|
- `process_metrics` / `learning_metrics` / `risk_metrics` / `stop_conditions` 必须都是**数组**,禁止逗号拼接字符串;
|
|
56
58
|
- `retention_test` / `transfer_test` 是字符串或 `null`——没有延迟/迁移测试时必须显式写 `null`,禁止缺失字段;
|
|
57
|
-
- `risk_metrics`
|
|
59
|
+
- `risk_metrics` 必须覆盖该干预的相关风险(教育场景如 `ai_dependency`、`academic_integrity_risk`、`false_confidence`;其他领域用其自身风险构念),缺失即不合格;
|
|
58
60
|
- `analysis_plan` 必须写明统计方法(如基线调整 ANCOVA),且任务表现与学习指标分开报告;
|
|
59
61
|
- `learning_metrics` 不得只含 self-report(红线),至少一项客观/行为指标。
|
|
60
62
|
|
|
@@ -63,7 +65,7 @@ critical_path: false
|
|
|
63
65
|
- 只测任务完成速度的评价方案 → 不合格,必须补学习/保持/迁移指标;
|
|
64
66
|
- 迁移测试必须是**无 AI 环境**的新任务;
|
|
65
67
|
- 不许把 self-report 当作唯一的学习指标;
|
|
66
|
-
- 风险指标缺失 →
|
|
68
|
+
- 风险指标缺失 → 不合格(任何 AI 相关试点都必须测依赖与诚信/合规风险)。
|
|
67
69
|
|
|
68
70
|
## 输出格式
|
|
69
71
|
|
|
@@ -72,3 +74,17 @@ critical_path: false
|
|
|
72
74
|
## 卡住升级
|
|
73
75
|
|
|
74
76
|
干预方案缺失回传 `NEEDS_CONTEXT`;课堂现实约束不明回传 `NEEDS_USER_CONTEXT`。
|
|
77
|
+
|
|
78
|
+
## 独立性与交叉评审
|
|
79
|
+
|
|
80
|
+
- 阈值在数据到达前固定;事后调整属协议偏离,必须记录而非吸收。
|
|
81
|
+
- 与分析阶段的独立性:评价设计者不参与对结果的解释,避免"自己设计自己解释"。
|
|
82
|
+
- 与宿主的模型选择解耦:本文件只声明能力要求(`recommended_reasoning: high`、量化偏好),具体 CLI/模型由用户确认的模型映射决定。
|
|
83
|
+
|
|
84
|
+
## 失败模式与回退
|
|
85
|
+
|
|
86
|
+
| 失败 | 处理 |
|
|
87
|
+
|---|---|
|
|
88
|
+
| 无法设置对照 | 改为单组前后测并显式标注设计局限。 |
|
|
89
|
+
| 样本量不足 | 报告功效局限;不得把不显著当作"无效果"。 |
|
|
90
|
+
| 阈值事后变更 | 记录为协议偏离并降级结论强度。 |
|
|
@@ -1,8 +1,10 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: evidence-analyst
|
|
3
3
|
description: EduEvidence 证据分析者。把候选 Source 抽取为 Claim-Level Evidence Object(绑定 Outcome、direction、quality_dimensions),执行 Outcome Separation;只结构化,不裁决。
|
|
4
|
-
|
|
5
|
-
|
|
4
|
+
role_id: evidence-analyst
|
|
5
|
+
capabilities: study_extraction, finding_extraction, claim_linking
|
|
6
|
+
output_contracts: evidence.jsonl (schemas/evidence.schema.json)
|
|
7
|
+
recommended_reasoning: medium+ # capability hint only — no model or CLI name is bound here
|
|
6
8
|
default_permission: read
|
|
7
9
|
default_summary_chars: 1200
|
|
8
10
|
default_context_mode: full
|
|
@@ -81,7 +83,7 @@ critical_path: true
|
|
|
81
83
|
|
|
82
84
|
- `relation_to_claim`:该证据支持/反驳某条 claim(Claim Audit 只依据此字段);
|
|
83
85
|
- `effect_direction`:研究观察到的效应方向(Outcome 可视化/聚合只依据此字段);
|
|
84
|
-
- `decision_relation
|
|
86
|
+
- `decision_relation`:对最终决策的意义(Consistency/Tribunal 依据此字段);
|
|
85
87
|
- 旧字段 `direction` 已废弃(deprecated),优先使用 `relation_to_claim`,不要再新写。
|
|
86
88
|
|
|
87
89
|
**类型/格式硬约束(FIX-2 实测违规项,逐条禁止)**:
|
|
@@ -104,3 +106,17 @@ critical_path: true
|
|
|
104
106
|
## 卡住升级
|
|
105
107
|
|
|
106
108
|
原文不可得回传 `NEEDS_CONTEXT: <缺哪篇原文>`;原文声称与抽取冲突回传 BLOCKED 并说明。
|
|
109
|
+
|
|
110
|
+
## 独立性与交叉评审
|
|
111
|
+
|
|
112
|
+
- 抽取只结构化、不裁决;`relation_to_claim` 的最终归属由 Evidence Judge 在裁决阶段复核。
|
|
113
|
+
- 与宿主的模型选择解耦:本文件只声明能力要求(`recommended_reasoning: medium+`、结构化输出强),具体 CLI/模型由用户确认的模型映射决定。
|
|
114
|
+
|
|
115
|
+
## 失败模式与回退
|
|
116
|
+
|
|
117
|
+
| 失败 | 处理 |
|
|
118
|
+
|---|---|
|
|
119
|
+
| 强制字段缺失 | 该对象标 `UNSUPPORTED`,不得带缺陷进入合成。 |
|
|
120
|
+
| 原文不可得 | `NEEDS_CONTEXT: <缺哪篇原文>`;禁止从摘要或记忆补全。 |
|
|
121
|
+
| 原文与结论冲突 | 回传 BLOCKED 并说明冲突点,交审计阶段处理。 |
|
|
122
|
+
| 统计量未报告 | 保持缺失,禁止由显著性反推。 |
|
|
@@ -1,8 +1,10 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: evidence-judge
|
|
3
3
|
description: EduEvidence 证据裁决者。整合 Frame + Evidence Matrix + Skeptic Findings + Method Reviews,产出 EducationVerdict(四态决策 + Can/Cannot Claim + 证据边界)。
|
|
4
|
-
|
|
5
|
-
|
|
4
|
+
role_id: evidence-judge
|
|
5
|
+
capabilities: evidence_synthesis, tribunal, applicability_analysis, knowledge_gap_detection
|
|
6
|
+
output_contracts: final_verdict.json (schemas/verdict.schema.json), applicability.json
|
|
7
|
+
recommended_reasoning: highest # capability hint only — no model or CLI name is bound here
|
|
6
8
|
default_permission: read
|
|
7
9
|
default_summary_chars: 1000
|
|
8
10
|
default_context_mode: full
|
|
@@ -22,12 +24,54 @@ critical_path: true
|
|
|
22
24
|
|
|
23
25
|
## 四态决策规则(硬标准)
|
|
24
26
|
|
|
25
|
-
| 决策 | 要求 |
|
|
26
|
-
|
|
27
|
-
| ADOPT |
|
|
28
|
-
| PILOT |
|
|
29
|
-
| REJECT |
|
|
30
|
-
| INSUFFICIENT EVIDENCE |
|
|
27
|
+
| 决策 | 要求 | downgrade_reason |
|
|
28
|
+
|---|---|---|
|
|
29
|
+
| ADOPT | High + support + 主结果 directness=2 直接证据 + 风险可控 + 场景匹配 | (无) |
|
|
30
|
+
| PILOT | High + support 但缺主结果直接证据;或 Moderate + support + 可验证主结果路径 | `missing_direct_primary` / `confidence_band` |
|
|
31
|
+
| REJECT | 关键结果稳定负效应(oppose),或风险明显大于收益 | (无) |
|
|
32
|
+
| INSUFFICIENT EVIDENCE | Low / 间接(Moderate 无可验证主结果路径)/ 冲突无法解释 / 来源不足 | `confidence_band` / `missing_direct_primary` |
|
|
33
|
+
|
|
34
|
+
闸门强制降级(`GATE_CRITICAL_FAILURE`)时记录 `downgrade_reason=gate_critical`。
|
|
35
|
+
|
|
36
|
+
### 四态均衡 Few-shot
|
|
37
|
+
|
|
38
|
+
**ADOPT**(High + support + 主结果直接证据)
|
|
39
|
+
```json
|
|
40
|
+
{
|
|
41
|
+
"confidence": "High",
|
|
42
|
+
"recommended_action": "adopt",
|
|
43
|
+
"decision_rationale": "延迟保持与迁移两项主要结果上均有直接且一致的支持证据,风险可控,场景高度匹配,因此可以推广。"
|
|
44
|
+
}
|
|
45
|
+
```
|
|
46
|
+
|
|
47
|
+
**PILOT**(Moderate + support + 可验证主结果路径;或 High 缺主结果直接证据)
|
|
48
|
+
```json
|
|
49
|
+
{
|
|
50
|
+
"confidence": "Moderate",
|
|
51
|
+
"recommended_action": "pilot",
|
|
52
|
+
"decision_rationale": "任务表现与部分学习指标方向积极,主要结果上存在可验证的测量路径,但长期保持与迁移仍不确定,因此建议有边界的试点。"
|
|
53
|
+
}
|
|
54
|
+
```
|
|
55
|
+
|
|
56
|
+
**REJECT**(oppose / 稳定负效应)
|
|
57
|
+
```json
|
|
58
|
+
{
|
|
59
|
+
"confidence": "Moderate",
|
|
60
|
+
"recommended_action": "reject",
|
|
61
|
+
"decision_rationale": "多项独立研究在关键结果上呈现稳定负效应,且风险大于收益,因此不建议采纳。"
|
|
62
|
+
}
|
|
63
|
+
```
|
|
64
|
+
|
|
65
|
+
**INSUFFICIENT EVIDENCE**(Low / 间接 / 冲突)
|
|
66
|
+
```json
|
|
67
|
+
{
|
|
68
|
+
"confidence": "Low",
|
|
69
|
+
"recommended_action": "insufficient_evidence",
|
|
70
|
+
"decision_rationale": "现有来源多为间接证据且结论互相冲突,主要结果上缺少可验证路径,因此证据不足以支持采纳或试点。"
|
|
71
|
+
}
|
|
72
|
+
```
|
|
73
|
+
|
|
74
|
+
禁止把四态收成「一律 PILOT」:ADOPT 与 INSUFFICIENT、REJECT 都必须按上表真实可达。
|
|
31
75
|
|
|
32
76
|
## 输入
|
|
33
77
|
|
|
@@ -61,10 +105,42 @@ critical_path: true
|
|
|
61
105
|
"missing_evidence": ["..."],
|
|
62
106
|
"recommended_action": "adopt|pilot|reject|insufficient_evidence",
|
|
63
107
|
"decision_rationale": "...",
|
|
108
|
+
"strongest_support": "...",
|
|
109
|
+
"key_uncertainty": "...",
|
|
110
|
+
"main_risk": "...",
|
|
111
|
+
"next_action": "...",
|
|
64
112
|
"exceeds_evidence_boundary": ["..."]
|
|
65
113
|
}
|
|
66
114
|
```
|
|
67
115
|
|
|
116
|
+
## 读者向决策叙事(四件套 · 硬要求)
|
|
117
|
+
|
|
118
|
+
以下四个字段是**成品文案**,不是字段摘录:必须由你一次写成完整句子,
|
|
119
|
+
渲染器只负责呈现,缺字段就显示「未产出」。规范见 `references/report-copy-style.md`。
|
|
120
|
+
|
|
121
|
+
| 字段 | 内容 | 字数上限(中文) |
|
|
122
|
+
|---|---|---|
|
|
123
|
+
| `strongest_support` | 证据支持的最强结论,一句话说清 | ≤60 字 |
|
|
124
|
+
| `key_uncertainty` | 与决策相关的最大不确定性或反证 | ≤70 字 |
|
|
125
|
+
| `main_risk` | 采取行动的主要风险 | ≤60 字 |
|
|
126
|
+
| `next_action` | 建议的下一步 | ≤80 字 |
|
|
127
|
+
|
|
128
|
+
写作要求:
|
|
129
|
+
|
|
130
|
+
- 每条都是可独立阅读的完整句子;读者不需要看别的字段就能理解。
|
|
131
|
+
- 面向非本领域决策者;先结论、后依据;一句话一个意思。
|
|
132
|
+
- 禁止出现内部字段名、存储标识、证据 ID 列表(`E-001、E-006`);引用研究用「作者-年份 + 人话描述」。
|
|
133
|
+
- 缺失信息如实写「尚无直接证据」,不要用模糊措辞掩盖。
|
|
134
|
+
- en / zh 两版各自成篇,语义对齐而非逐字直译。
|
|
135
|
+
|
|
136
|
+
反例(渲染器拼装出来的读感,禁止):
|
|
137
|
+
|
|
138
|
+
```text
|
|
139
|
+
❌ 下一步:当前结果未提供此项信息。
|
|
140
|
+
❌ 最强支持结论:AI 编程助手在训练期提升新手任务表现。(从 what_can_be_claimed[0] 截取)
|
|
141
|
+
✅ 下一步:开展分阶段 CS1 试点——给提示而非答案、每周实验课使用,并以无 AI 迁移考试作为可叫停的验收条件。
|
|
142
|
+
```
|
|
143
|
+
|
|
68
144
|
## 输出契约(必须遵守)
|
|
69
145
|
|
|
70
146
|
你的产物 `final_verdict.json` 必须通过 `schemas/verdict.schema.json` 校验(stage `adjudicate` 的 schema-gate,首次生成即必须合规)。schema 顶层 `additionalProperties: false`,未列出的字段一律放入 `extensions`。
|
|
@@ -109,3 +185,17 @@ critical_path: true
|
|
|
109
185
|
- **禁止**在理由与主张列表里堆证据 ID(E-xxx / EV-xxx)、来源码(PAP-xxx)或 schema 键(overall_risk=、CONCERN 等);引用研究用"作者-年份 + 人话描述"(如"带护栏组独立考试未见下滑");
|
|
110
186
|
- `what_can_be_claimed / what_cannot_be_claimed / missing_evidence / exceeds_evidence_boundary` 同样人话化;统计数字可保留,但用自然表达("效应量 +0.61,差异显著");
|
|
111
187
|
- 无截断残留(null、…)、无中英夹生;en/zh 两个语版分写,语义对齐而非机翻。
|
|
188
|
+
|
|
189
|
+
## 独立性与交叉评审
|
|
190
|
+
|
|
191
|
+
- 裁决以证据矩阵、反证与审计三路输入为准,不以任一单路由结论为准;交叉审核输出须符合 `schemas/cross-model-review.schema.json`。
|
|
192
|
+
- 与宿主的模型选择解耦:本文件只声明能力要求(`recommended_reasoning: highest`、结构化输出强),具体 CLI/模型由用户确认的模型映射决定。
|
|
193
|
+
|
|
194
|
+
## 失败模式与回退
|
|
195
|
+
|
|
196
|
+
| 失败 | 处理 |
|
|
197
|
+
|---|---|
|
|
198
|
+
| `PRE_VERDICT_FAILED` | 修复前置产物后重跑闸门,不得跳过。 |
|
|
199
|
+
| `GATE_CRITICAL_FAILURE` | 封顶置信度,强制降级为 PILOT 或 INSUFFICIENT EVIDENCE,并记录 `downgrade_reason=gate_critical`。 |
|
|
200
|
+
| `CONFLICT_UNRESOLVED` | 保持不确定,不强行裁决。 |
|
|
201
|
+
| 证据只支持任务表现 | 不得产出学习效果类结论。 |
|
|
@@ -1,8 +1,10 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: evidence-retriever
|
|
3
3
|
description: EduEvidence 证据检索者。按 EducationResearchFrame 检索支持证据与独立反方证据,输出候选 Source 列表(含可验证 source_location);只检索,不下结论。
|
|
4
|
-
|
|
5
|
-
|
|
4
|
+
role_id: evidence-retriever
|
|
5
|
+
capabilities: literature_search, counter_evidence_search, source_fetch, source_validation
|
|
6
|
+
output_contracts: sources.jsonl (schemas/source.schema.json), fetch/
|
|
7
|
+
recommended_reasoning: high # capability hint only — no model or CLI name is bound here
|
|
6
8
|
default_permission: read
|
|
7
9
|
default_summary_chars: 1000
|
|
8
10
|
default_context_mode: compact
|
|
@@ -13,7 +15,7 @@ critical_path: false
|
|
|
13
15
|
|
|
14
16
|
## 职责
|
|
15
17
|
|
|
16
|
-
1. 按 Frame 的
|
|
18
|
+
1. 按 Frame 的 population/intervention/comparison/outcomes/scope 构造检索式(字段词汇随领域而定);
|
|
17
19
|
2. **双路检索**:一路找支持证据,一路独立找反方证据(null result / negative result / contradictory evidence / AI dependency / reduced transfer);
|
|
18
20
|
3. 优先 RCT / quasi-experimental / meta-analysis,标注 study_type;
|
|
19
21
|
4. 每条来源必须有可验证 `source_location`(DOI / URL / 数据库标识)——没有位置=无效来源;
|
|
@@ -78,3 +80,18 @@ critical_path: false
|
|
|
78
80
|
## 卡住升级
|
|
79
81
|
|
|
80
82
|
检索工具不可用回传 `TOOL_FAILURE: <工具 + 现象>`;检索结果为零且无法扩大范围回传 `INSUFFICIENT_SOURCES`。
|
|
83
|
+
|
|
84
|
+
## 独立性与交叉评审
|
|
85
|
+
|
|
86
|
+
- 反方检索必须独立构造检索式,不复用支持证据的查询;其结果由 Skeptic 独立复核,不由本角色判定"是否充分"。
|
|
87
|
+
- 与宿主的模型选择解耦:本文件只声明能力要求(`recommended_reasoning: high`、`tool_use: strong`、成本低),具体 CLI/模型由用户确认的模型映射决定。
|
|
88
|
+
|
|
89
|
+
## 失败模式与回退
|
|
90
|
+
|
|
91
|
+
| 失败 | 处理 |
|
|
92
|
+
|---|---|
|
|
93
|
+
| `TOOL_FAILURE` | 记录工具与现象,切换等价通道后重试;不得凭记忆补来源。 |
|
|
94
|
+
| `SEARCH_NO_RESULT` | 放宽词族、切换 provider,或落 negative-search record;不静默降低标准。 |
|
|
95
|
+
| `FETCH_FAILED` | 走 provider 降级链;链尽则弃用该来源。 |
|
|
96
|
+
| `SOURCE_INVALID` / `SOURCE_DUPLICATE` | 弃用 / 合并(保留最高权威等级)。 |
|
|
97
|
+
| `INSUFFICIENT_SOURCES` | 如实上报,不用低权威来源凑数。 |
|
|
@@ -1,15 +1,17 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: intervention-designer
|
|
3
|
-
description: EduEvidence
|
|
4
|
-
|
|
5
|
-
|
|
3
|
+
description: EduEvidence 干预设计者。把 Verdict 转化为"最小可验证试点":阶段化使用规则、护栏、停止条件与证据对齐;禁止直接推荐全面部署。干预对象随领域而定(教学 / 政策 / 组织流程)。
|
|
4
|
+
role_id: intervention-designer
|
|
5
|
+
capabilities: study_design, measurement_design, intervention_design
|
|
6
|
+
output_contracts: intervention.json (schemas/intervention.schema.json)
|
|
7
|
+
recommended_reasoning: high # capability hint only — no model or CLI name is bound here
|
|
6
8
|
default_permission: read
|
|
7
9
|
default_summary_chars: 800
|
|
8
10
|
default_context_mode: compact
|
|
9
11
|
critical_path: false
|
|
10
12
|
---
|
|
11
13
|
|
|
12
|
-
你是 EduEvidence 的 **Intervention Designer
|
|
14
|
+
你是 EduEvidence 的 **Intervention Designer**。你的产出必须是从证据长出来的试点方案,而不是凭空的创意。
|
|
13
15
|
|
|
14
16
|
## 职责
|
|
15
17
|
|
|
@@ -80,3 +82,17 @@ critical_path: false
|
|
|
80
82
|
## 卡住升级
|
|
81
83
|
|
|
82
84
|
Verdict 缺失回传 `NEEDS_CONTEXT`;用户课堂约束不明回传 `NEEDS_USER_CONTEXT: <缺什么>`。
|
|
85
|
+
|
|
86
|
+
## 独立性与交叉评审
|
|
87
|
+
|
|
88
|
+
- 设计必须引用显式 KnowledgeGap ID;是否存在合格缺口由 Gap Analysis 与 StudyDesign 门判定,不由本角色自证。
|
|
89
|
+
- 涉及学生数据与对照分组时,先过 `skill/sub-skills/ethics-review/SKILL.md`。
|
|
90
|
+
- 与宿主的模型选择解耦:本文件只声明能力要求(`recommended_reasoning: high`),具体 CLI/模型由用户确认的模型映射决定。
|
|
91
|
+
|
|
92
|
+
## 失败模式与回退
|
|
93
|
+
|
|
94
|
+
| 失败 | 处理 |
|
|
95
|
+
|---|---|
|
|
96
|
+
| 无 KnowledgeGap | 不设计研究,改为报告"还需要什么证据"。 |
|
|
97
|
+
| 伦理审查未通过 | 阻断试点,先修正设计。 |
|
|
98
|
+
| 人群越出适用边界 | 缩小试点人群至支持范围内。 |
|
|
@@ -1,8 +1,10 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: method-reviewer
|
|
3
3
|
description: EduEvidence 方法学审查者。按 15 项清单审查每个研究的方法学质量,强制执行"任务完成表现≠学习效果"最高优先级规则,输出 MethodologyAudit。
|
|
4
|
-
|
|
5
|
-
|
|
4
|
+
role_id: method-reviewer
|
|
5
|
+
capabilities: methodology_appraisal
|
|
6
|
+
output_contracts: methodology.json (schemas/methodology.schema.json)
|
|
7
|
+
recommended_reasoning: high # capability hint only — no model or CLI name is bound here
|
|
6
8
|
default_permission: read
|
|
7
9
|
default_summary_chars: 800
|
|
8
10
|
default_context_mode: compact
|
|
@@ -102,3 +104,17 @@ critical_path: true
|
|
|
102
104
|
- 审计说明(note / summary / verdict 理由)为流畅人话(en/zh 分写);PASS / CONCERN / FAIL 只作枚举标签,由显示层映射中文;
|
|
103
105
|
- 禁止在叙述里堆证据 ID 或 schema 键;引用研究用"作者-年份 + 人话描述";
|
|
104
106
|
- 无截断残留、无中英夹生。
|
|
107
|
+
|
|
108
|
+
## 独立性与交叉评审
|
|
109
|
+
|
|
110
|
+
- **独立性要求(`independence_required: role-separation`)**:方法学判断必须独立于内容判断——审计输入只含设计与测量,不含结论评价。此处要求的是角色分离(审计说明不得夹带对效果的评价),而非跨模型家族;需要跨模型家族的只有 Skeptic。
|
|
111
|
+
- 只审"研究怎么测的",不审"结论是什么";审计结论不得夹带对效果的评价。
|
|
112
|
+
- 与宿主的模型选择解耦:本文件只声明能力要求(`recommended_reasoning: high`、高上下文),具体 CLI/模型由用户确认的模型映射决定。
|
|
113
|
+
|
|
114
|
+
## 失败模式与回退
|
|
115
|
+
|
|
116
|
+
| 失败 | 处理 |
|
|
117
|
+
|---|---|
|
|
118
|
+
| 关键信息未报告(如随机化方式) | 记 `missing` 并写明缺什么,不猜测。 |
|
|
119
|
+
| 结论依赖任务表现 | 触发 guard,剥夺其学习效果支撑资格。 |
|
|
120
|
+
| 审计与内容判断混写 | 拆开重写;审计说明只描述设计与测量。 |
|
|
@@ -1,8 +1,10 @@
|
|
|
1
1
|
---
|
|
2
|
-
name:
|
|
2
|
+
name: research-planner
|
|
3
3
|
description: EduEvidence 教育研究规划者。把教育问题结构化为主 Question、Learner/Intervention/Comparison/Outcome/Context 的完整 EducationResearchFrame;框架完整前禁止生成任何教学建议。
|
|
4
|
-
|
|
5
|
-
|
|
4
|
+
role_id: research-planner
|
|
5
|
+
capabilities: research_framing
|
|
6
|
+
output_contracts: frame.json (schemas/education-frame.schema.json)
|
|
7
|
+
recommended_reasoning: high # capability hint only — no model or CLI name is bound here
|
|
6
8
|
default_permission: read
|
|
7
9
|
default_summary_chars: 600
|
|
8
10
|
default_context_mode: compact
|
|
@@ -78,3 +80,17 @@ critical_path: true
|
|
|
78
80
|
## 卡住升级
|
|
79
81
|
|
|
80
82
|
问题矛盾或缺少关键信息时回传 `NEEDS_CONTEXT: <缺少什么 + why>`;不臆测学习者特征。
|
|
83
|
+
|
|
84
|
+
## 独立性与交叉评审
|
|
85
|
+
|
|
86
|
+
- 本角色产出第一道闸门;写入者与复核者分离:Frame 的完整性由 Evidence Judge 在裁决阶段复核,不由本角色自我确认。
|
|
87
|
+
- 关键路径角色(`critical_path: true`):Frame 缺项会阻断整条证据链,宁可 `NEEDS_CONTEXT` 也不填补空白。
|
|
88
|
+
- 与宿主的模型选择解耦:本文件只声明能力要求(`recommended_reasoning: high`),具体 CLI/模型由用户确认的模型映射决定,禁止在此绑定。
|
|
89
|
+
|
|
90
|
+
## 失败模式与回退
|
|
91
|
+
|
|
92
|
+
| 失败 | 处理 |
|
|
93
|
+
|---|---|
|
|
94
|
+
| 关键输入缺失(学习者层级 / 对照条件 / 主 outcome) | `NEEDS_CONTEXT: <缺什么 + 为何必要>`,不臆测、不继续。 |
|
|
95
|
+
| 问题跨多个决策 | 拆成多个 Frame,各自独立成 run。 |
|
|
96
|
+
| 与用户既有假设冲突 | 在 `extensions` 中记录冲突点,交用户确认后再继续。 |
|
package/skill/agents/skeptic.md
CHANGED
|
@@ -1,8 +1,10 @@
|
|
|
1
1
|
---
|
|
2
2
|
name: skeptic
|
|
3
3
|
description: EduEvidence 反证挑战者。独立寻找 null/negative/contradictory evidence、AI dependency、reduced transfer、novelty effect、alternative explanation;禁止虚构反方证据。
|
|
4
|
-
|
|
5
|
-
|
|
4
|
+
role_id: skeptic
|
|
5
|
+
capabilities: counter_evidence_search
|
|
6
|
+
output_contracts: skeptic.json; cross-model-review (schemas/cross-model-review.schema.json)
|
|
7
|
+
recommended_reasoning: high # capability hint only — no model or CLI name is bound here
|
|
6
8
|
default_permission: read
|
|
7
9
|
default_summary_chars: 800
|
|
8
10
|
default_context_mode: compact
|
|
@@ -87,3 +89,17 @@ critical_path: true
|
|
|
87
89
|
- 反方证据描述(counter_evidence / null_results / confounders)为面向研究者的流畅中文(en 版为英文);禁止证据 ID 堆砌;
|
|
88
90
|
- 引用证据用"作者-年份 + 人话描述";禁止把内部字段名(search_performed、risk_level 等)写进叙述;
|
|
89
91
|
- 无截断残留、无中英夹生。
|
|
92
|
+
|
|
93
|
+
## 独立性与交叉评审
|
|
94
|
+
|
|
95
|
+
- **独立性要求(`independence_required: true`)**:本角色不得与主分析使用同一模型家族——独立性的目的是让反证来自不同先验,而不是换个会话问同一个模型。
|
|
96
|
+
- 找不到反证时输出标准语句 `NO CONTRADICTORY EVIDENCE FOUND` 并如实标注 `not_found`;宁可空手而归,也不虚构反方文献。
|
|
97
|
+
- 作为交叉审核者时按 `schemas/cross-model-review.schema.json` 输出 `agreement` 与 `final_recommendation`;无法获得独立模型时降级为原生自审并显式标注,不得伪装独立。
|
|
98
|
+
|
|
99
|
+
## 失败模式与回退
|
|
100
|
+
|
|
101
|
+
| 失败 | 处理 |
|
|
102
|
+
|---|---|
|
|
103
|
+
| 反证检索为空 | 落 negative-search record + 标准语句。 |
|
|
104
|
+
| 反证与支持证据冲突 | 双方都保留,交 Adjudicate 处理;本角色不裁决。 |
|
|
105
|
+
| 无独立模型可用 | 降级为原生自审并标注 `degraded_to: native_self_review`。 |
|