engineering-argument-language 3.2.3__tar.gz
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- engineering_argument_language-3.2.3/.github/workflows/experiment-stage.yml +110 -0
- engineering_argument_language-3.2.3/.github/workflows/experiment.yml +114 -0
- engineering_argument_language-3.2.3/.github/workflows/publish.yml +191 -0
- engineering_argument_language-3.2.3/.github/workflows/test.yml +94 -0
- engineering_argument_language-3.2.3/.github/workflows/verify.yml +23 -0
- engineering_argument_language-3.2.3/AGENTS.md +13 -0
- engineering_argument_language-3.2.3/CONTRACT.md +89 -0
- engineering_argument_language-3.2.3/LICENSE +24 -0
- engineering_argument_language-3.2.3/MANIFEST.in +21 -0
- engineering_argument_language-3.2.3/Makefile +31 -0
- engineering_argument_language-3.2.3/PKG-INFO +188 -0
- engineering_argument_language-3.2.3/README.md +149 -0
- engineering_argument_language-3.2.3/docs/argument-model.md +99 -0
- engineering_argument_language-3.2.3/docs/argument-reconstruction.md +55 -0
- engineering_argument_language-3.2.3/docs/argument-service.md +104 -0
- engineering_argument_language-3.2.3/docs/aspic-method.md +127 -0
- engineering_argument_language-3.2.3/docs/aspic-view-v1.schema.json +688 -0
- engineering_argument_language-3.2.3/docs/aspic-view.schema.json +1065 -0
- engineering_argument_language-3.2.3/docs/design-aim.md +42 -0
- engineering_argument_language-3.2.3/docs/eal3-composition.md +72 -0
- engineering_argument_language-3.2.3/docs/eal3-design.md +78 -0
- engineering_argument_language-3.2.3/docs/eal3-experiment-methodology.md +251 -0
- engineering_argument_language-3.2.3/docs/eal3-syntax-comparison.md +113 -0
- engineering_argument_language-3.2.3/docs/grounded-reasoning.md +113 -0
- engineering_argument_language-3.2.3/docs/language.md +226 -0
- engineering_argument_language-3.2.3/docs/mcp-and-tools.md +167 -0
- engineering_argument_language-3.2.3/docs/mcp-architecture.md +53 -0
- engineering_argument_language-3.2.3/docs/packaging.md +170 -0
- engineering_argument_language-3.2.3/docs/reasoning-modes.md +269 -0
- engineering_argument_language-3.2.3/docs/sources.md +186 -0
- engineering_argument_language-3.2.3/docs/vocabulary.md +32 -0
- engineering_argument_language-3.2.3/examples/api-load-test/README.md +171 -0
- engineering_argument_language-3.2.3/examples/api-load-test/aspic-compiled.eal +85 -0
- engineering_argument_language-3.2.3/examples/api-load-test/aspic-fixture.json +26 -0
- engineering_argument_language-3.2.3/examples/api-load-test/aspic-source.eal +94 -0
- engineering_argument_language-3.2.3/examples/api-load-test/aspic-tools.toml +11 -0
- engineering_argument_language-3.2.3/examples/api-load-test/aspic_compiled_demo.py +91 -0
- engineering_argument_language-3.2.3/examples/api-load-test/aspic_demo.py +80 -0
- engineering_argument_language-3.2.3/examples/api-load-test/aspic_fixture_collector.py +44 -0
- engineering_argument_language-3.2.3/examples/api-load-test/collect_results.py +103 -0
- engineering_argument_language-3.2.3/examples/api-load-test/mcp.toml +23 -0
- engineering_argument_language-3.2.3/examples/api-load-test/recompute_synthetic.py +68 -0
- engineering_argument_language-3.2.3/examples/api-load-test/report.json +111 -0
- engineering_argument_language-3.2.3/examples/api-load-test/run.py +108 -0
- engineering_argument_language-3.2.3/examples/api-load-test/source.eal +46 -0
- engineering_argument_language-3.2.3/examples/api-load-test/tools.toml +8 -0
- engineering_argument_language-3.2.3/examples/aspic-keywords/README.md +100 -0
- engineering_argument_language-3.2.3/examples/aspic-keywords/collector.py +41 -0
- engineering_argument_language-3.2.3/examples/aspic-keywords/demo.py +220 -0
- engineering_argument_language-3.2.3/examples/aspic-keywords/fixture.json +174 -0
- engineering_argument_language-3.2.3/examples/aspic-keywords/formal.eal +191 -0
- engineering_argument_language-3.2.3/examples/aspic-keywords/output/compiled-view.json +976 -0
- engineering_argument_language-3.2.3/examples/aspic-keywords/output/formal-view.json +403 -0
- engineering_argument_language-3.2.3/examples/aspic-keywords/output/summary.json +237 -0
- engineering_argument_language-3.2.3/examples/aspic-keywords/source.eal +120 -0
- engineering_argument_language-3.2.3/examples/aspic-keywords/tools.toml +10 -0
- engineering_argument_language-3.2.3/examples/multi-environment/README.md +22 -0
- engineering_argument_language-3.2.3/examples/multi-environment/collector.py +48 -0
- engineering_argument_language-3.2.3/examples/multi-environment/fixtures.json +17 -0
- engineering_argument_language-3.2.3/examples/multi-environment/run.py +73 -0
- engineering_argument_language-3.2.3/examples/multi-environment/source.eal +39 -0
- engineering_argument_language-3.2.3/examples/multi-environment/tools.toml +7 -0
- engineering_argument_language-3.2.3/examples/scoped-review/README.md +13 -0
- engineering_argument_language-3.2.3/examples/scoped-review/collector.py +17 -0
- engineering_argument_language-3.2.3/examples/scoped-review/review.eal +51 -0
- engineering_argument_language-3.2.3/examples/scoped-review/run.py +32 -0
- engineering_argument_language-3.2.3/examples/scoped-review/tools.toml +7 -0
- engineering_argument_language-3.2.3/experiments/__init__.py +1 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/README.md +207 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/WORKFLOW.md +277 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/__init__.py +1 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/accounting.py +77 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/adoption_costs.py +144 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/allocation_assessment.py +97 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/analyse.py +53 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/annotation_provenance.py +35 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/annotations.py +189 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/answers.py +145 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/artifacts.py +50 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/batch_simulation.py +38 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/benchmark.py +89 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/bounded_statistics.py +50 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/cadence-diagnostic-plan.json +7485 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/cadence-information-design.json +199 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/cadence-plan.json +7497 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/calibration.py +164 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/cases.py +123 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/collector.py +11 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/comparisons.py +57 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/conditions.py +53 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/conventional_reasoner.py +65 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/corpus_cases.py +131 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/corpus_context.py +77 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/corpus_logic.py +146 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/corpus_reference.py +46 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/decision_statistics.py +171 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/design.py +97 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/diagnostic-plan.json +79 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/diagnostic_design.py +133 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/diagnostic_inference.py +32 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/diagnostic_measurement.py +135 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/diagnostic_preparation.py +40 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/diagnostic_reporting.py +123 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/diagnostic_resources.py +50 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/diagnostics.py +159 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/evaluation-tasks.example.json +7417 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/execution.py +73 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/information-design.json +199 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/information_config.py +86 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/information_design.py +144 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/journal.py +78 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/nuisance.py +37 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/outcomes.py +22 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/pilot_data.py +178 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/plan.json +68 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/plan_information.py +59 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/progress.py +35 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/project.py +130 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/prompts.py +38 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/protocol.json +1753 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/provenance.py +28 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/provider.py +147 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/reasoning_context.py +30 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/records.py +64 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/rehearse.py +123 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/reporting.py +138 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/requirements.lock +72 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/resources.py +64 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/results/run-36432530106.md +42 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/run_state.py +105 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/runner.py +195 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/scoring.py +93 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/session.py +150 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/session_recovery.py +35 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/statistical_quantiles.py +97 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/task_case.py +13 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/task_context.py +165 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/task_corpus.json +5896 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/task_manifest.py +114 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/threshold_method.py +35 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/trajectory_simulation.py +172 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/validate_design.py +86 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/validation/design-validation-screening.json +490 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/validation/design-validation.json +498 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/validation/eal3-language.json +92 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/validation/eal31-composition.json +85 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/validation/rehearsals.json +2059 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/validation/runtime-rehearsals.json +2020 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/validation/throughput.json +77 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/validation/wrapper.json +46 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/verification-eal3.md +51 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/verification-eal31.md +34 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/verification.md +172 -0
- engineering_argument_language-3.2.3/experiments/model_transfer/workflow.py +109 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/README.md +275 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/__init__.py +1 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/__main__.py +114 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/analysis.py +283 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/calibration.py +109 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/demo/eal-authoring-example.eal +30 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/demo/eal-authoring-tools.toml +7 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/demo/plan.json +58 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/demo/seed/README.md +7 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/demo/seed/measurement.json +1 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/demo/seed/probe.py +16 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/design.py +217 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/inference.py +70 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/mock_model.py +23 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/model_gateway.py +112 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/openai_responses.py +68 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/operations.py +125 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/protocol.json +818 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/verification/calibration.json +167 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/verification/results.md +46 -0
- engineering_argument_language-3.2.3/experiments/transfer_study/workspace.py +218 -0
- engineering_argument_language-3.2.3/grammar/EAL.g4 +151 -0
- engineering_argument_language-3.2.3/pyproject.toml +50 -0
- engineering_argument_language-3.2.3/scripts/check_distribution.py +138 -0
- engineering_argument_language-3.2.3/scripts/check_release_tag.py +97 -0
- engineering_argument_language-3.2.3/scripts/distribution_smoke.py +169 -0
- engineering_argument_language-3.2.3/scripts/experiment_pipeline.py +345 -0
- engineering_argument_language-3.2.3/scripts/generate_parser.py +61 -0
- engineering_argument_language-3.2.3/setup.cfg +4 -0
- engineering_argument_language-3.2.3/skills/eal-assessment-routing/SKILL.md +18 -0
- engineering_argument_language-3.2.3/skills/engineer-argumentation-languages/SKILL.md +78 -0
- engineering_argument_language-3.2.3/skills/engineer-argumentation-languages/agents/openai.yaml +15 -0
- engineering_argument_language-3.2.3/skills/engineer-argumentation-languages/assets/icon.svg +6 -0
- engineering_argument_language-3.2.3/skills/engineer-argumentation-languages/references/design-contract.md +58 -0
- engineering_argument_language-3.2.3/skills/engineer-argumentation-languages/references/expert-language-design.md +74 -0
- engineering_argument_language-3.2.3/skills/engineer-argumentation-languages/references/implementation-patterns.md +26 -0
- engineering_argument_language-3.2.3/skills/engineer-argumentation-languages/references/mcp-execution.md +31 -0
- engineering_argument_language-3.2.3/skills/engineer-argumentation-languages/references/reasoning-modes.md +20 -0
- engineering_argument_language-3.2.3/skills/engineer-argumentation-languages/references/reasoning-semantics.md +27 -0
- engineering_argument_language-3.2.3/skills/engineer-argumentation-languages/references/sources.md +23 -0
- engineering_argument_language-3.2.3/src/eal/__init__.py +3 -0
- engineering_argument_language-3.2.3/src/eal/abstractions.py +144 -0
- engineering_argument_language-3.2.3/src/eal/acquisition_coordination.py +80 -0
- engineering_argument_language-3.2.3/src/eal/api_load_methods.py +77 -0
- engineering_argument_language-3.2.3/src/eal/aspic.py +478 -0
- engineering_argument_language-3.2.3/src/eal/aspic_compiler.py +445 -0
- engineering_argument_language-3.2.3/src/eal/aspic_export.py +309 -0
- engineering_argument_language-3.2.3/src/eal/builtin_methods.py +115 -0
- engineering_argument_language-3.2.3/src/eal/catalogue.py +384 -0
- engineering_argument_language-3.2.3/src/eal/cli.py +204 -0
- engineering_argument_language-3.2.3/src/eal/client_transports.py +69 -0
- engineering_argument_language-3.2.3/src/eal/collection_identity.py +28 -0
- engineering_argument_language-3.2.3/src/eal/collection_scheduler.py +72 -0
- engineering_argument_language-3.2.3/src/eal/command_process.py +118 -0
- engineering_argument_language-3.2.3/src/eal/command_supervisor.py +112 -0
- engineering_argument_language-3.2.3/src/eal/composition.py +531 -0
- engineering_argument_language-3.2.3/src/eal/credentials.py +33 -0
- engineering_argument_language-3.2.3/src/eal/dialectic.py +321 -0
- engineering_argument_language-3.2.3/src/eal/discovery.py +127 -0
- engineering_argument_language-3.2.3/src/eal/evaluator.py +619 -0
- engineering_argument_language-3.2.3/src/eal/expressions.py +404 -0
- engineering_argument_language-3.2.3/src/eal/extensions.py +38 -0
- engineering_argument_language-3.2.3/src/eal/formatter.py +156 -0
- engineering_argument_language-3.2.3/src/eal/generated/EALLexer.py +434 -0
- engineering_argument_language-3.2.3/src/eal/generated/EALParser.py +6559 -0
- engineering_argument_language-3.2.3/src/eal/generated/EALVisitor.py +393 -0
- engineering_argument_language-3.2.3/src/eal/generated/__init__.py +1 -0
- engineering_argument_language-3.2.3/src/eal/host.py +179 -0
- engineering_argument_language-3.2.3/src/eal/host_redaction.py +51 -0
- engineering_argument_language-3.2.3/src/eal/knowledge.py +96 -0
- engineering_argument_language-3.2.3/src/eal/limits.py +88 -0
- engineering_argument_language-3.2.3/src/eal/mcp_guard.py +56 -0
- engineering_argument_language-3.2.3/src/eal/methods.py +463 -0
- engineering_argument_language-3.2.3/src/eal/model.py +266 -0
- engineering_argument_language-3.2.3/src/eal/model_context.py +55 -0
- engineering_argument_language-3.2.3/src/eal/modes.py +121 -0
- engineering_argument_language-3.2.3/src/eal/observation_reuse.py +128 -0
- engineering_argument_language-3.2.3/src/eal/operation_contracts.py +25 -0
- engineering_argument_language-3.2.3/src/eal/operations.py +155 -0
- engineering_argument_language-3.2.3/src/eal/packets.py +604 -0
- engineering_argument_language-3.2.3/src/eal/parser.py +407 -0
- engineering_argument_language-3.2.3/src/eal/planning.py +136 -0
- engineering_argument_language-3.2.3/src/eal/propositions.py +231 -0
- engineering_argument_language-3.2.3/src/eal/reachability.py +85 -0
- engineering_argument_language-3.2.3/src/eal/reasoning/__init__.py +25 -0
- engineering_argument_language-3.2.3/src/eal/reasoning/abductive.py +45 -0
- engineering_argument_language-3.2.3/src/eal/reasoning/analogical.py +41 -0
- engineering_argument_language-3.2.3/src/eal/reasoning/causal.py +46 -0
- engineering_argument_language-3.2.3/src/eal/reasoning/counterfactual.py +64 -0
- engineering_argument_language-3.2.3/src/eal/reasoning/deductive.py +72 -0
- engineering_argument_language-3.2.3/src/eal/reasoning/inductive.py +29 -0
- engineering_argument_language-3.2.3/src/eal/reasoning/strategy.py +18 -0
- engineering_argument_language-3.2.3/src/eal/reasoning/structured.py +11 -0
- engineering_argument_language-3.2.3/src/eal/reasoning/temporal.py +47 -0
- engineering_argument_language-3.2.3/src/eal/reasoning/validation.py +90 -0
- engineering_argument_language-3.2.3/src/eal/registered_assessment.py +181 -0
- engineering_argument_language-3.2.3/src/eal/runtime.py +454 -0
- engineering_argument_language-3.2.3/src/eal/sampled_negative.py +131 -0
- engineering_argument_language-3.2.3/src/eal/scope_transfer.py +59 -0
- engineering_argument_language-3.2.3/src/eal/semantics.py +656 -0
- engineering_argument_language-3.2.3/src/eal/server.py +79 -0
- engineering_argument_language-3.2.3/src/eal/server_auth.py +37 -0
- engineering_argument_language-3.2.3/src/eal/server_settings.py +258 -0
- engineering_argument_language-3.2.3/src/eal/source_printer.py +129 -0
- engineering_argument_language-3.2.3/src/eal/store.py +272 -0
- engineering_argument_language-3.2.3/src/eal/tool_acquisition.py +383 -0
- engineering_argument_language-3.2.3/src/engineering_argument_language.egg-info/PKG-INFO +188 -0
- engineering_argument_language-3.2.3/src/engineering_argument_language.egg-info/SOURCES.txt +357 -0
- engineering_argument_language-3.2.3/src/engineering_argument_language.egg-info/dependency_links.txt +1 -0
- engineering_argument_language-3.2.3/src/engineering_argument_language.egg-info/entry_points.txt +4 -0
- engineering_argument_language-3.2.3/src/engineering_argument_language.egg-info/requires.txt +16 -0
- engineering_argument_language-3.2.3/src/engineering_argument_language.egg-info/top_level.txt +1 -0
- engineering_argument_language-3.2.3/tests/test_acquisition_coordination.py +152 -0
- engineering_argument_language-3.2.3/tests/test_ai_annotation_provenance.py +61 -0
- engineering_argument_language-3.2.3/tests/test_api_load_methods.py +35 -0
- engineering_argument_language-3.2.3/tests/test_aspic.py +417 -0
- engineering_argument_language-3.2.3/tests/test_aspic_compiler.py +376 -0
- engineering_argument_language-3.2.3/tests/test_aspic_compiler_entrypoints.py +124 -0
- engineering_argument_language-3.2.3/tests/test_aspic_export.py +230 -0
- engineering_argument_language-3.2.3/tests/test_aspic_extensions.py +104 -0
- engineering_argument_language-3.2.3/tests/test_aspic_formal_annotations.py +281 -0
- engineering_argument_language-3.2.3/tests/test_aspic_keywords_example.py +144 -0
- engineering_argument_language-3.2.3/tests/test_aspic_pyarg.py +270 -0
- engineering_argument_language-3.2.3/tests/test_binding_contract_review.py +148 -0
- engineering_argument_language-3.2.3/tests/test_builtin_method_specs.py +56 -0
- engineering_argument_language-3.2.3/tests/test_catalogue.py +161 -0
- engineering_argument_language-3.2.3/tests/test_cli_catalogue.py +129 -0
- engineering_argument_language-3.2.3/tests/test_client_transports.py +81 -0
- engineering_argument_language-3.2.3/tests/test_collection_budgets.py +104 -0
- engineering_argument_language-3.2.3/tests/test_collection_identity.py +69 -0
- engineering_argument_language-3.2.3/tests/test_collection_scheduler.py +183 -0
- engineering_argument_language-3.2.3/tests/test_command_process.py +212 -0
- engineering_argument_language-3.2.3/tests/test_composed_dialectic.py +279 -0
- engineering_argument_language-3.2.3/tests/test_composed_language.py +176 -0
- engineering_argument_language-3.2.3/tests/test_dialectic.py +235 -0
- engineering_argument_language-3.2.3/tests/test_discovery.py +20 -0
- engineering_argument_language-3.2.3/tests/test_distribution_metadata.py +121 -0
- engineering_argument_language-3.2.3/tests/test_eal3_frontend.py +225 -0
- engineering_argument_language-3.2.3/tests/test_eal3_integration.py +69 -0
- engineering_argument_language-3.2.3/tests/test_eal3_runtime.py +211 -0
- engineering_argument_language-3.2.3/tests/test_eal3_syntax.py +156 -0
- engineering_argument_language-3.2.3/tests/test_evaluator.py +370 -0
- engineering_argument_language-3.2.3/tests/test_execution_contract_review.py +193 -0
- engineering_argument_language-3.2.3/tests/test_experiment_pipeline.py +300 -0
- engineering_argument_language-3.2.3/tests/test_extension_runtime.py +65 -0
- engineering_argument_language-3.2.3/tests/test_formatter.py +100 -0
- engineering_argument_language-3.2.3/tests/test_frontend_contract_review.py +142 -0
- engineering_argument_language-3.2.3/tests/test_host.py +168 -0
- engineering_argument_language-3.2.3/tests/test_host_credentials.py +136 -0
- engineering_argument_language-3.2.3/tests/test_host_limits.py +25 -0
- engineering_argument_language-3.2.3/tests/test_http_origin.py +113 -0
- engineering_argument_language-3.2.3/tests/test_investigation_improvements.py +190 -0
- engineering_argument_language-3.2.3/tests/test_knowledge.py +158 -0
- engineering_argument_language-3.2.3/tests/test_language.py +167 -0
- engineering_argument_language-3.2.3/tests/test_language_version.py +42 -0
- engineering_argument_language-3.2.3/tests/test_load_test_example.py +145 -0
- engineering_argument_language-3.2.3/tests/test_mcp.py +172 -0
- engineering_argument_language-3.2.3/tests/test_mcp_lifetime.py +253 -0
- engineering_argument_language-3.2.3/tests/test_mcp_transports.py +385 -0
- engineering_argument_language-3.2.3/tests/test_method_extensions.py +233 -0
- engineering_argument_language-3.2.3/tests/test_method_worker_limits.py +28 -0
- engineering_argument_language-3.2.3/tests/test_model_transfer.py +620 -0
- engineering_argument_language-3.2.3/tests/test_model_transfer_answers.py +258 -0
- engineering_argument_language-3.2.3/tests/test_model_transfer_corpus.py +127 -0
- engineering_argument_language-3.2.3/tests/test_model_transfer_design.py +54 -0
- engineering_argument_language-3.2.3/tests/test_model_transfer_diagnostics.py +301 -0
- engineering_argument_language-3.2.3/tests/test_model_transfer_execution.py +339 -0
- engineering_argument_language-3.2.3/tests/test_model_transfer_information_design.py +257 -0
- engineering_argument_language-3.2.3/tests/test_model_transfer_records.py +71 -0
- engineering_argument_language-3.2.3/tests/test_model_transfer_rehearsal.py +47 -0
- engineering_argument_language-3.2.3/tests/test_model_transfer_reporting.py +185 -0
- engineering_argument_language-3.2.3/tests/test_model_transfer_session.py +161 -0
- engineering_argument_language-3.2.3/tests/test_model_transfer_task.py +86 -0
- engineering_argument_language-3.2.3/tests/test_model_transfer_workflow.py +93 -0
- engineering_argument_language-3.2.3/tests/test_modes.py +353 -0
- engineering_argument_language-3.2.3/tests/test_multi_environment_example.py +103 -0
- engineering_argument_language-3.2.3/tests/test_observation_reuse.py +201 -0
- engineering_argument_language-3.2.3/tests/test_packets.py +313 -0
- engineering_argument_language-3.2.3/tests/test_planning.py +239 -0
- engineering_argument_language-3.2.3/tests/test_qualified_packets.py +133 -0
- engineering_argument_language-3.2.3/tests/test_reachability_companion.py +32 -0
- engineering_argument_language-3.2.3/tests/test_reasoning_contract_review.py +119 -0
- engineering_argument_language-3.2.3/tests/test_reasoning_pipeline_contract.py +167 -0
- engineering_argument_language-3.2.3/tests/test_registered_assessment.py +218 -0
- engineering_argument_language-3.2.3/tests/test_release_tag.py +141 -0
- engineering_argument_language-3.2.3/tests/test_revision_contract.py +148 -0
- engineering_argument_language-3.2.3/tests/test_runtime.py +295 -0
- engineering_argument_language-3.2.3/tests/test_sampled_negative.py +251 -0
- engineering_argument_language-3.2.3/tests/test_scope_transfer_language.py +74 -0
- engineering_argument_language-3.2.3/tests/test_scoped_composition.py +261 -0
- engineering_argument_language-3.2.3/tests/test_semantic_review.py +294 -0
- engineering_argument_language-3.2.3/tests/test_server_auth.py +65 -0
- engineering_argument_language-3.2.3/tests/test_server_settings.py +266 -0
- engineering_argument_language-3.2.3/tests/test_service_workflow.py +115 -0
- engineering_argument_language-3.2.3/tests/test_store_initialisation.py +252 -0
- engineering_argument_language-3.2.3/tests/test_store_lifecycle.py +28 -0
- engineering_argument_language-3.2.3/tests/test_transfer_inference.py +48 -0
- engineering_argument_language-3.2.3/tests/test_transfer_study.py +373 -0
- engineering_argument_language-3.2.3/tests/test_typed_propositions.py +157 -0
- engineering_argument_language-3.2.3/tools/investigation-validator/README.md +16 -0
- engineering_argument_language-3.2.3/tools/investigation-validator/assets/investigation-protocol.template.json +217 -0
- engineering_argument_language-3.2.3/tools/investigation-validator/provenance.json +15 -0
- engineering_argument_language-3.2.3/tools/investigation-validator/scripts/validate_design.py +263 -0
- engineering_argument_language-3.2.3/tools/investigation-validator/scripts/validate_investigation.py +2180 -0
- engineering_argument_language-3.2.3/tools/investigation-validator/scripts/validate_workflow.py +170 -0
|
@@ -0,0 +1,110 @@
|
|
|
1
|
+
name: EAL experiment stage
|
|
2
|
+
on:
|
|
3
|
+
workflow_call:
|
|
4
|
+
inputs:
|
|
5
|
+
phase:
|
|
6
|
+
type: string
|
|
7
|
+
required: true
|
|
8
|
+
stage_name:
|
|
9
|
+
type: string
|
|
10
|
+
required: true
|
|
11
|
+
operation:
|
|
12
|
+
type: string
|
|
13
|
+
default: start
|
|
14
|
+
plan:
|
|
15
|
+
type: string
|
|
16
|
+
default: comparison
|
|
17
|
+
plan_path:
|
|
18
|
+
type: string
|
|
19
|
+
default: ''
|
|
20
|
+
source_artifact_id:
|
|
21
|
+
type: string
|
|
22
|
+
default: ''
|
|
23
|
+
annotation_labels:
|
|
24
|
+
type: string
|
|
25
|
+
default: ''
|
|
26
|
+
information_config:
|
|
27
|
+
type: string
|
|
28
|
+
default: ''
|
|
29
|
+
workers:
|
|
30
|
+
type: number
|
|
31
|
+
default: 0
|
|
32
|
+
secrets:
|
|
33
|
+
OPENAI_API_KEY:
|
|
34
|
+
required: false
|
|
35
|
+
outputs:
|
|
36
|
+
artifact_id:
|
|
37
|
+
description: Cumulative retained data or the frozen preflight plan
|
|
38
|
+
value: ${{ jobs.stage.outputs.artifact_id }}
|
|
39
|
+
can_collect:
|
|
40
|
+
description: True only when another bounded collection job is appropriate
|
|
41
|
+
value: ${{ jobs.stage.outputs.can_collect }}
|
|
42
|
+
permissions:
|
|
43
|
+
contents: read
|
|
44
|
+
actions: read
|
|
45
|
+
jobs:
|
|
46
|
+
stage:
|
|
47
|
+
runs-on: ubuntu-24.04
|
|
48
|
+
timeout-minutes: 120
|
|
49
|
+
outputs:
|
|
50
|
+
artifact_id: ${{ steps.retain.outputs.artifact-id }}
|
|
51
|
+
can_collect: ${{ steps.execute.outputs.can_collect }}
|
|
52
|
+
steps:
|
|
53
|
+
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
|
54
|
+
with:
|
|
55
|
+
persist-credentials: false
|
|
56
|
+
ref: ${{ github.sha }}
|
|
57
|
+
- uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1
|
|
58
|
+
with:
|
|
59
|
+
python-version: '3.12.14'
|
|
60
|
+
- name: Install the pinned experiment runtime
|
|
61
|
+
run: python -m pip install -c experiments/model_transfer/requirements.lock -e .
|
|
62
|
+
- name: Execute the ordered stage
|
|
63
|
+
id: execute
|
|
64
|
+
env:
|
|
65
|
+
PIPELINE_PHASE: ${{ inputs.phase }}
|
|
66
|
+
PIPELINE_OPERATION: ${{ inputs.operation }}
|
|
67
|
+
TRANSFER_PLAN: ${{ inputs.plan }}
|
|
68
|
+
PLAN_PATH: ${{ inputs.plan_path }}
|
|
69
|
+
SOURCE_ARTIFACT_ID: ${{ inputs.source_artifact_id }}
|
|
70
|
+
ANNOTATION_LABELS: ${{ inputs.annotation_labels }}
|
|
71
|
+
INFORMATION_CONFIG: ${{ inputs.information_config }}
|
|
72
|
+
EXPERIMENT_WORKERS: ${{ inputs.workers }}
|
|
73
|
+
OPENAI_API_KEY: ${{ inputs.phase == 'collect' && secrets.OPENAI_API_KEY || '' }}
|
|
74
|
+
GH_TOKEN: ${{ github.token }}
|
|
75
|
+
run: python -m scripts.experiment_pipeline "$PIPELINE_PHASE"
|
|
76
|
+
- name: Retain this cumulative stage even after a failure
|
|
77
|
+
id: retain
|
|
78
|
+
if: ${{ always() && steps.execute.outputs.artifact_path != '' }}
|
|
79
|
+
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02
|
|
80
|
+
with:
|
|
81
|
+
name: experiment-${{ inputs.stage_name }}-${{ github.run_id }}-${{ github.run_attempt }}
|
|
82
|
+
path: ${{ steps.execute.outputs.artifact_path }}
|
|
83
|
+
include-hidden-files: true
|
|
84
|
+
if-no-files-found: warn
|
|
85
|
+
retention-days: 90
|
|
86
|
+
- name: Retain calibration and scripted rehearsal records
|
|
87
|
+
if: ${{ always() && inputs.phase == 'prepare' && inputs.operation != 'resume' }}
|
|
88
|
+
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02
|
|
89
|
+
with:
|
|
90
|
+
name: experiment-preflight-${{ github.run_id }}-${{ github.run_attempt }}
|
|
91
|
+
path: experiments/model_transfer/runs/preflight
|
|
92
|
+
include-hidden-files: true
|
|
93
|
+
if-no-files-found: warn
|
|
94
|
+
retention-days: 90
|
|
95
|
+
- name: Export only masked answer text for the independent assessor
|
|
96
|
+
id: masked
|
|
97
|
+
if: ${{ success() && inputs.phase == 'export' }}
|
|
98
|
+
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02
|
|
99
|
+
with:
|
|
100
|
+
name: experiment-annotations-${{ github.run_id }}-${{ github.run_attempt }}
|
|
101
|
+
path: experiments/model_transfer/runs/live/annotation-bundle/items.json
|
|
102
|
+
if-no-files-found: warn
|
|
103
|
+
retention-days: 90
|
|
104
|
+
- name: Explain the result and the next required input
|
|
105
|
+
if: ${{ always() && steps.execute.outputs.artifact_path != '' }}
|
|
106
|
+
env:
|
|
107
|
+
PIPELINE_ARTIFACT_PATH: ${{ steps.execute.outputs.artifact_path }}
|
|
108
|
+
PIPELINE_ARTIFACT_ID: ${{ steps.retain.outputs.artifact-id }}
|
|
109
|
+
ANNOTATION_ARTIFACT_ID: ${{ steps.masked.outputs.artifact-id }}
|
|
110
|
+
run: python -m scripts.experiment_pipeline summary
|
|
@@ -0,0 +1,114 @@
|
|
|
1
|
+
name: EAL experiment
|
|
2
|
+
run-name: EAL experiment — ${{ inputs.operation }} — ${{ inputs.plan }}
|
|
3
|
+
on:
|
|
4
|
+
workflow_dispatch:
|
|
5
|
+
inputs:
|
|
6
|
+
operation:
|
|
7
|
+
description: 'start runs the pilot; finish imports labels and processes results; evaluate starts a separately budgeted evaluation'
|
|
8
|
+
type: choice
|
|
9
|
+
options: [start, resume, finish, evaluate, rehearse]
|
|
10
|
+
default: start
|
|
11
|
+
plan:
|
|
12
|
+
description: 'Pilot question (used only for start/rehearse); bundled plans have a USD 2 ceiling'
|
|
13
|
+
type: choice
|
|
14
|
+
options: [comparison, diagnostics, cadence, cadence-diagnostics, custom]
|
|
15
|
+
default: comparison
|
|
16
|
+
source_artifact_id:
|
|
17
|
+
description: 'Latest full artefact ID for resume/finish; planning artefact for evaluate; otherwise blank'
|
|
18
|
+
type: string
|
|
19
|
+
default: ''
|
|
20
|
+
annotation_labels:
|
|
21
|
+
description: 'For finish: repository path to completed labels JSON with declared human/AI assessor provenance'
|
|
22
|
+
type: string
|
|
23
|
+
default: ''
|
|
24
|
+
plan_path:
|
|
25
|
+
description: 'For a custom start/rehearse only: repository path to pilot plan'
|
|
26
|
+
type: string
|
|
27
|
+
default: ''
|
|
28
|
+
information_config:
|
|
29
|
+
description: 'For finish: optional repository path to a matching allocation configuration'
|
|
30
|
+
type: string
|
|
31
|
+
default: ''
|
|
32
|
+
workers:
|
|
33
|
+
description: '0 uses the plan (default 4); 1–8 overrides a new pilot; resume/evaluation preserve their allocation'
|
|
34
|
+
type: number
|
|
35
|
+
default: 0
|
|
36
|
+
permissions:
|
|
37
|
+
contents: read
|
|
38
|
+
actions: read
|
|
39
|
+
concurrency:
|
|
40
|
+
group: eal-experiment-${{ inputs.source_artifact_id || github.run_id }}
|
|
41
|
+
cancel-in-progress: false
|
|
42
|
+
jobs:
|
|
43
|
+
verify:
|
|
44
|
+
uses: ./.github/workflows/verify.yml
|
|
45
|
+
prepare:
|
|
46
|
+
needs: verify
|
|
47
|
+
if: ${{ inputs.operation != 'finish' }}
|
|
48
|
+
uses: ./.github/workflows/experiment-stage.yml
|
|
49
|
+
with:
|
|
50
|
+
phase: prepare
|
|
51
|
+
stage_name: prepared
|
|
52
|
+
operation: ${{ inputs.operation }}
|
|
53
|
+
plan: ${{ inputs.plan }}
|
|
54
|
+
plan_path: ${{ inputs.plan_path }}
|
|
55
|
+
source_artifact_id: ${{ inputs.source_artifact_id }}
|
|
56
|
+
workers: ${{ inputs.workers }}
|
|
57
|
+
collect-1:
|
|
58
|
+
needs: prepare
|
|
59
|
+
if: ${{ needs.prepare.outputs.can_collect == 'true' }}
|
|
60
|
+
uses: ./.github/workflows/experiment-stage.yml
|
|
61
|
+
with:
|
|
62
|
+
phase: collect
|
|
63
|
+
stage_name: segment-1
|
|
64
|
+
source_artifact_id: ${{ needs.prepare.outputs.artifact_id }}
|
|
65
|
+
secrets:
|
|
66
|
+
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
|
|
67
|
+
collect-2:
|
|
68
|
+
needs: collect-1
|
|
69
|
+
if: ${{ needs.collect-1.outputs.can_collect == 'true' }}
|
|
70
|
+
uses: ./.github/workflows/experiment-stage.yml
|
|
71
|
+
with:
|
|
72
|
+
phase: collect
|
|
73
|
+
stage_name: segment-2
|
|
74
|
+
source_artifact_id: ${{ needs.collect-1.outputs.artifact_id }}
|
|
75
|
+
secrets:
|
|
76
|
+
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
|
|
77
|
+
collect-3:
|
|
78
|
+
needs: collect-2
|
|
79
|
+
if: ${{ needs.collect-2.outputs.can_collect == 'true' }}
|
|
80
|
+
uses: ./.github/workflows/experiment-stage.yml
|
|
81
|
+
with:
|
|
82
|
+
phase: collect
|
|
83
|
+
stage_name: segment-3
|
|
84
|
+
source_artifact_id: ${{ needs.collect-2.outputs.artifact_id }}
|
|
85
|
+
secrets:
|
|
86
|
+
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
|
|
87
|
+
collect-4:
|
|
88
|
+
needs: collect-3
|
|
89
|
+
if: ${{ needs.collect-3.outputs.can_collect == 'true' }}
|
|
90
|
+
uses: ./.github/workflows/experiment-stage.yml
|
|
91
|
+
with:
|
|
92
|
+
phase: collect
|
|
93
|
+
stage_name: segment-4
|
|
94
|
+
source_artifact_id: ${{ needs.collect-3.outputs.artifact_id }}
|
|
95
|
+
secrets:
|
|
96
|
+
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
|
|
97
|
+
export:
|
|
98
|
+
needs: [prepare, collect-1, collect-2, collect-3, collect-4]
|
|
99
|
+
if: ${{ !cancelled() && needs.prepare.result == 'success' && inputs.operation != 'rehearse' }}
|
|
100
|
+
uses: ./.github/workflows/experiment-stage.yml
|
|
101
|
+
with:
|
|
102
|
+
phase: export
|
|
103
|
+
stage_name: export
|
|
104
|
+
source_artifact_id: ${{ needs.collect-4.outputs.artifact_id || needs.collect-3.outputs.artifact_id || needs.collect-2.outputs.artifact_id || needs.collect-1.outputs.artifact_id || needs.prepare.outputs.artifact_id }}
|
|
105
|
+
finish:
|
|
106
|
+
needs: verify
|
|
107
|
+
if: ${{ inputs.operation == 'finish' }}
|
|
108
|
+
uses: ./.github/workflows/experiment-stage.yml
|
|
109
|
+
with:
|
|
110
|
+
phase: finish
|
|
111
|
+
stage_name: finish
|
|
112
|
+
source_artifact_id: ${{ inputs.source_artifact_id }}
|
|
113
|
+
annotation_labels: ${{ inputs.annotation_labels }}
|
|
114
|
+
information_config: ${{ inputs.information_config }}
|
|
@@ -0,0 +1,191 @@
|
|
|
1
|
+
name: EAL package release
|
|
2
|
+
|
|
3
|
+
on:
|
|
4
|
+
workflow_dispatch:
|
|
5
|
+
inputs:
|
|
6
|
+
release_tag:
|
|
7
|
+
description: 'Existing release tag; dispatch this workflow on the same tag'
|
|
8
|
+
type: string
|
|
9
|
+
required: true
|
|
10
|
+
default: v3.2.3
|
|
11
|
+
destination:
|
|
12
|
+
description: 'Verify artefacts, or publish the verified artefacts to the selected index'
|
|
13
|
+
type: choice
|
|
14
|
+
options: [verify, testpypi, pypi]
|
|
15
|
+
default: verify
|
|
16
|
+
|
|
17
|
+
permissions:
|
|
18
|
+
contents: read
|
|
19
|
+
|
|
20
|
+
concurrency:
|
|
21
|
+
group: package-${{ inputs.release_tag }}-${{ inputs.destination }}
|
|
22
|
+
cancel-in-progress: false
|
|
23
|
+
|
|
24
|
+
jobs:
|
|
25
|
+
resolve:
|
|
26
|
+
name: Validate release tag and version
|
|
27
|
+
runs-on: ubuntu-24.04
|
|
28
|
+
timeout-minutes: 5
|
|
29
|
+
outputs:
|
|
30
|
+
commit: ${{ steps.release.outputs.commit }}
|
|
31
|
+
version: ${{ steps.release.outputs.version }}
|
|
32
|
+
steps:
|
|
33
|
+
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
|
34
|
+
with:
|
|
35
|
+
ref: ${{ github.sha }}
|
|
36
|
+
fetch-depth: 0
|
|
37
|
+
persist-credentials: false
|
|
38
|
+
- uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
|
|
39
|
+
with:
|
|
40
|
+
python-version: '3.12'
|
|
41
|
+
- name: Resolve the matching tag to an immutable commit on main
|
|
42
|
+
id: release
|
|
43
|
+
env:
|
|
44
|
+
RELEASE_TAG: ${{ inputs.release_tag }}
|
|
45
|
+
DISPATCH_REF: ${{ github.ref }}
|
|
46
|
+
DISPATCH_SHA: ${{ github.sha }}
|
|
47
|
+
run: >-
|
|
48
|
+
python scripts/check_release_tag.py "$RELEASE_TAG"
|
|
49
|
+
--dispatch-ref "$DISPATCH_REF" --dispatch-sha "$DISPATCH_SHA"
|
|
50
|
+
--github-output "$GITHUB_OUTPUT"
|
|
51
|
+
|
|
52
|
+
full-check:
|
|
53
|
+
name: Verify parser, tests and maintained example
|
|
54
|
+
needs: resolve
|
|
55
|
+
runs-on: ubuntu-24.04
|
|
56
|
+
timeout-minutes: 30
|
|
57
|
+
steps:
|
|
58
|
+
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
|
59
|
+
with:
|
|
60
|
+
ref: ${{ needs.resolve.outputs.commit }}
|
|
61
|
+
persist-credentials: false
|
|
62
|
+
- uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
|
|
63
|
+
with:
|
|
64
|
+
python-version: '3.12.14'
|
|
65
|
+
- name: Install constrained package and test dependencies
|
|
66
|
+
run: python -m pip install -c experiments/model_transfer/requirements.lock '.[dev,release]'
|
|
67
|
+
- name: Check parser, scientific contracts, tests and maintained example
|
|
68
|
+
run: make check PYTHON=python
|
|
69
|
+
|
|
70
|
+
build:
|
|
71
|
+
name: Build one wheel and source distribution
|
|
72
|
+
needs: resolve
|
|
73
|
+
runs-on: ubuntu-24.04
|
|
74
|
+
timeout-minutes: 15
|
|
75
|
+
outputs:
|
|
76
|
+
artifact-id: ${{ steps.distributions.outputs.artifact-id }}
|
|
77
|
+
steps:
|
|
78
|
+
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
|
79
|
+
with:
|
|
80
|
+
ref: ${{ needs.resolve.outputs.commit }}
|
|
81
|
+
persist-credentials: false
|
|
82
|
+
- uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
|
|
83
|
+
with:
|
|
84
|
+
python-version: '3.12'
|
|
85
|
+
- name: Install package and release tools
|
|
86
|
+
run: python -m pip install -c experiments/model_transfer/requirements.lock '.[release]'
|
|
87
|
+
- name: Build and check distribution metadata
|
|
88
|
+
run: |
|
|
89
|
+
make build PYTHON=python
|
|
90
|
+
python -m twine check --strict dist/*.whl dist/*.tar.gz
|
|
91
|
+
sha256sum dist/*.whl dist/*.tar.gz > SHA256SUMS
|
|
92
|
+
- name: Retain immutable distributions and their checksums
|
|
93
|
+
id: distributions
|
|
94
|
+
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02 # v4
|
|
95
|
+
with:
|
|
96
|
+
name: python-distributions-${{ github.run_id }}-${{ github.run_attempt }}
|
|
97
|
+
path: |
|
|
98
|
+
dist/*.whl
|
|
99
|
+
dist/*.tar.gz
|
|
100
|
+
SHA256SUMS
|
|
101
|
+
if-no-files-found: error
|
|
102
|
+
retention-days: 14
|
|
103
|
+
|
|
104
|
+
installed-package:
|
|
105
|
+
name: Installed package (${{ matrix.os }}, Python ${{ matrix.python }})
|
|
106
|
+
needs: [resolve, build]
|
|
107
|
+
runs-on: ${{ matrix.os }}
|
|
108
|
+
timeout-minutes: 20
|
|
109
|
+
strategy:
|
|
110
|
+
fail-fast: false
|
|
111
|
+
matrix:
|
|
112
|
+
include:
|
|
113
|
+
- {os: ubuntu-24.04, python: '3.11'}
|
|
114
|
+
- {os: ubuntu-24.04, python: '3.12'}
|
|
115
|
+
- {os: ubuntu-24.04, python: '3.13'}
|
|
116
|
+
- {os: ubuntu-24.04, python: '3.14'}
|
|
117
|
+
- {os: macos-14, python: '3.12'}
|
|
118
|
+
steps:
|
|
119
|
+
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
|
120
|
+
with:
|
|
121
|
+
ref: ${{ needs.resolve.outputs.commit }}
|
|
122
|
+
persist-credentials: false
|
|
123
|
+
- uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
|
|
124
|
+
with:
|
|
125
|
+
python-version: ${{ matrix.python }}
|
|
126
|
+
- name: Retrieve the exact built distributions from this run
|
|
127
|
+
uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4
|
|
128
|
+
with:
|
|
129
|
+
artifact-ids: ${{ needs.build.outputs.artifact-id }}
|
|
130
|
+
merge-multiple: true
|
|
131
|
+
path: release-artifacts
|
|
132
|
+
- name: Check distribution bytes
|
|
133
|
+
working-directory: release-artifacts
|
|
134
|
+
run: shasum -a 256 --check SHA256SUMS
|
|
135
|
+
- name: Install each distribution outside the checkout and exercise the public interfaces
|
|
136
|
+
env:
|
|
137
|
+
RELEASE_VERSION: ${{ needs.resolve.outputs.version }}
|
|
138
|
+
run: >-
|
|
139
|
+
python scripts/check_distribution.py release-artifacts/dist
|
|
140
|
+
--expected-version "$RELEASE_VERSION"
|
|
141
|
+
|
|
142
|
+
publish-testpypi:
|
|
143
|
+
name: Publish verified distributions to TestPyPI
|
|
144
|
+
if: ${{ inputs.destination == 'testpypi' }}
|
|
145
|
+
needs: [resolve, full-check, build, installed-package]
|
|
146
|
+
runs-on: ubuntu-24.04
|
|
147
|
+
timeout-minutes: 10
|
|
148
|
+
environment:
|
|
149
|
+
name: testpypi
|
|
150
|
+
url: https://test.pypi.org/p/engineering-argument-language
|
|
151
|
+
permissions:
|
|
152
|
+
id-token: write
|
|
153
|
+
steps:
|
|
154
|
+
- name: Retrieve the exact verified distributions from this run
|
|
155
|
+
uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4
|
|
156
|
+
with:
|
|
157
|
+
artifact-ids: ${{ needs.build.outputs.artifact-id }}
|
|
158
|
+
merge-multiple: true
|
|
159
|
+
- name: Check distribution bytes before publication
|
|
160
|
+
run: sha256sum --check --strict SHA256SUMS
|
|
161
|
+
- name: Publish through the TestPyPI Trusted Publisher
|
|
162
|
+
uses: pypa/gh-action-pypi-publish@dc37677b2e1c63e2034f94d8a5b11f265b73ba33 # v1.14.2
|
|
163
|
+
with:
|
|
164
|
+
repository-url: https://test.pypi.org/legacy/
|
|
165
|
+
packages-dir: dist/
|
|
166
|
+
attestations: true
|
|
167
|
+
|
|
168
|
+
publish-pypi:
|
|
169
|
+
name: Publish verified distributions to PyPI
|
|
170
|
+
if: ${{ inputs.destination == 'pypi' }}
|
|
171
|
+
needs: [resolve, full-check, build, installed-package]
|
|
172
|
+
runs-on: ubuntu-24.04
|
|
173
|
+
timeout-minutes: 10
|
|
174
|
+
environment:
|
|
175
|
+
name: pypi.org
|
|
176
|
+
url: https://pypi.org/p/engineering-argument-language
|
|
177
|
+
permissions:
|
|
178
|
+
id-token: write
|
|
179
|
+
steps:
|
|
180
|
+
- name: Retrieve the exact verified distributions from this run
|
|
181
|
+
uses: actions/download-artifact@d3f86a106a0bac45b974a628896c90dbdf5c8093 # v4
|
|
182
|
+
with:
|
|
183
|
+
artifact-ids: ${{ needs.build.outputs.artifact-id }}
|
|
184
|
+
merge-multiple: true
|
|
185
|
+
- name: Check distribution bytes before publication
|
|
186
|
+
run: sha256sum --check --strict SHA256SUMS
|
|
187
|
+
- name: Publish through the PyPI Trusted Publisher
|
|
188
|
+
uses: pypa/gh-action-pypi-publish@dc37677b2e1c63e2034f94d8a5b11f265b73ba33 # v1.14.2
|
|
189
|
+
with:
|
|
190
|
+
packages-dir: dist/
|
|
191
|
+
attestations: true
|
|
@@ -0,0 +1,94 @@
|
|
|
1
|
+
name: EAL tests
|
|
2
|
+
on:
|
|
3
|
+
workflow_dispatch:
|
|
4
|
+
inputs:
|
|
5
|
+
model_transfer:
|
|
6
|
+
description: 'Run an experiment stage after tests (only collect calls the API)'
|
|
7
|
+
type: boolean
|
|
8
|
+
default: false
|
|
9
|
+
model_transfer_plan:
|
|
10
|
+
description: 'Principal, reasoning or evidence-cadence pilot; each has a USD 2 ceiling'
|
|
11
|
+
type: choice
|
|
12
|
+
options:
|
|
13
|
+
- comparison
|
|
14
|
+
- diagnostics
|
|
15
|
+
- cadence
|
|
16
|
+
- cadence-diagnostics
|
|
17
|
+
- custom
|
|
18
|
+
default: comparison
|
|
19
|
+
experiment_action:
|
|
20
|
+
description: 'Stage: collect/resume, free rehearsal/calibration, or offline processing'
|
|
21
|
+
type: choice
|
|
22
|
+
options: [collect, calibrate, rehearse, export-annotations, import-annotations, analyse, plan]
|
|
23
|
+
default: collect
|
|
24
|
+
plan_path:
|
|
25
|
+
description: 'Custom plan: repository path, or path within source artefact; ignored on resume'
|
|
26
|
+
type: string
|
|
27
|
+
default: ''
|
|
28
|
+
source_artifact_id:
|
|
29
|
+
description: 'Numeric artefact ID from this repository for resume or offline stages'
|
|
30
|
+
type: string
|
|
31
|
+
default: ''
|
|
32
|
+
resume_run:
|
|
33
|
+
description: 'Continue the frozen retained run; keep its cumulative budget and worker count'
|
|
34
|
+
type: boolean
|
|
35
|
+
default: false
|
|
36
|
+
segment_seconds:
|
|
37
|
+
description: 'Collection allowance for this job, 1–6000 seconds; remaining time saves artefacts'
|
|
38
|
+
type: number
|
|
39
|
+
default: 6000
|
|
40
|
+
annotation_labels:
|
|
41
|
+
description: 'Repository path to completed labels for import-annotations'
|
|
42
|
+
type: string
|
|
43
|
+
default: ''
|
|
44
|
+
information_config:
|
|
45
|
+
description: 'Optional repository path to prospective allocation configuration'
|
|
46
|
+
type: string
|
|
47
|
+
default: ''
|
|
48
|
+
workers:
|
|
49
|
+
description: '0 uses the plan (default 4); 1–8 overrides pilot workers; ignored for resume'
|
|
50
|
+
type: number
|
|
51
|
+
default: 0
|
|
52
|
+
permissions:
|
|
53
|
+
contents: read
|
|
54
|
+
actions: read
|
|
55
|
+
jobs:
|
|
56
|
+
test:
|
|
57
|
+
uses: ./.github/workflows/verify.yml
|
|
58
|
+
model-transfer:
|
|
59
|
+
if: ${{ inputs.model_transfer }}
|
|
60
|
+
needs: test
|
|
61
|
+
runs-on: ubuntu-24.04
|
|
62
|
+
timeout-minutes: 120
|
|
63
|
+
steps:
|
|
64
|
+
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262
|
|
65
|
+
with:
|
|
66
|
+
persist-credentials: false
|
|
67
|
+
- uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1
|
|
68
|
+
with:
|
|
69
|
+
python-version: '3.12.14'
|
|
70
|
+
- name: Install package
|
|
71
|
+
run: python -m pip install -c experiments/model_transfer/requirements.lock -e .
|
|
72
|
+
- name: Run the selected experiment stage
|
|
73
|
+
env:
|
|
74
|
+
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
|
|
75
|
+
TRANSFER_PLAN: ${{ inputs.model_transfer_plan }}
|
|
76
|
+
EXPERIMENT_ACTION: ${{ inputs.experiment_action }}
|
|
77
|
+
PLAN_PATH: ${{ inputs.plan_path }}
|
|
78
|
+
SOURCE_ARTIFACT_ID: ${{ inputs.source_artifact_id }}
|
|
79
|
+
RESUME_RUN: ${{ inputs.resume_run }}
|
|
80
|
+
SEGMENT_SECONDS: ${{ inputs.segment_seconds }}
|
|
81
|
+
ANNOTATION_LABELS: ${{ inputs.annotation_labels }}
|
|
82
|
+
INFORMATION_CONFIG: ${{ inputs.information_config }}
|
|
83
|
+
EXPERIMENT_WORKERS: ${{ inputs.workers }}
|
|
84
|
+
GH_TOKEN: ${{ github.token }}
|
|
85
|
+
run: python -m experiments.model_transfer.workflow
|
|
86
|
+
- name: Retain every attempt and the report
|
|
87
|
+
if: ${{ always() }}
|
|
88
|
+
uses: actions/upload-artifact@ea165f8d65b6e75b540449e92b4886f43607fa02
|
|
89
|
+
with:
|
|
90
|
+
name: model-transfer-${{ github.run_id }}-${{ github.run_attempt }}
|
|
91
|
+
path: experiments/model_transfer/runs/live
|
|
92
|
+
include-hidden-files: true
|
|
93
|
+
if-no-files-found: warn
|
|
94
|
+
retention-days: 90
|
|
@@ -0,0 +1,23 @@
|
|
|
1
|
+
name: EAL verification
|
|
2
|
+
on:
|
|
3
|
+
workflow_call:
|
|
4
|
+
permissions:
|
|
5
|
+
contents: read
|
|
6
|
+
jobs:
|
|
7
|
+
verify:
|
|
8
|
+
runs-on: ubuntu-24.04
|
|
9
|
+
timeout-minutes: 30
|
|
10
|
+
steps:
|
|
11
|
+
- uses: actions/checkout@11d5960a326750d5838078e36cf38b85af677262 # v4
|
|
12
|
+
with:
|
|
13
|
+
persist-credentials: false
|
|
14
|
+
ref: ${{ github.sha }}
|
|
15
|
+
- uses: actions/setup-python@ece7cb06caefa5fff74198d8649806c4678c61a1 # v6
|
|
16
|
+
with:
|
|
17
|
+
python-version: '3.12.14'
|
|
18
|
+
- name: Install package and test dependencies
|
|
19
|
+
run: python -m pip install -c experiments/model_transfer/requirements.lock -e '.[dev]'
|
|
20
|
+
- name: Check parser, scientific contracts, tests and maintained example
|
|
21
|
+
run: make check
|
|
22
|
+
- name: Build distribution
|
|
23
|
+
run: make build
|
|
@@ -0,0 +1,13 @@
|
|
|
1
|
+
# Project instructions
|
|
2
|
+
|
|
3
|
+
EAL/3 is the only supported language. Backwards compatibility is never a requirement for this project. Choose the clearest coherent design for the current language; remove obsolete syntax, version-dependent semantics, aliases and compatibility adapters rather than preserving them. Do not add migration machinery solely to support an earlier EAL version.
|
|
4
|
+
|
|
5
|
+
Apply each change across affected current documentation, discovery schemas and the single API load-test example. Document the present contract and implementation; keep revision history in Git rather than explanatory pages. Keep the maintained example self-contained and clearly distinguish synthetic data from measurements. Focus validation on discovering registered EAL files, reusing compatible observations across sessions and models, and reaching the right scoped output with fewer repeated calls. A source-language version, package version and observation schema version identify different contracts; declare changes to each affected contract explicitly.
|
|
6
|
+
|
|
7
|
+
Use the repository's `skills/engineer-argumentation-languages/SKILL.md` and its expert language design reference for changes to syntax, abstractions, method contracts or reasoning semantics. Separate primary-source recommendations, EAL design decisions, implemented behaviour and measured results.
|
|
8
|
+
|
|
9
|
+
Keep one versioned reasoning-method selector: `method "name/version"`. A tool declaration names an interface version; its operational characteristics belong to the trusted host binding and the recorded acquisition, not an EAL `mode` clause. Built-in and installed reasoning methods follow the same typed contracts and binding checks.
|
|
10
|
+
|
|
11
|
+
Reusable argument patterns use typed parameters with closed lexical scope. Expansion must preserve evidence, claim and assumption identity and the qualifications of an ordinary argument. Retain source locations in diagnostics and canonical parse–format–parse meaning.
|
|
12
|
+
|
|
13
|
+
Do not infer model capability, comprehension gains or cost savings from a working interpreter alone. Measurements require actual trials and retained outcomes, including failures. Keep GitHub workflows manually triggered.
|
|
@@ -0,0 +1,89 @@
|
|
|
1
|
+
# EAL/3 integration contract
|
|
2
|
+
|
|
3
|
+
EAL/3 is the supported source language. The Python distribution is `engineering-argument-language` version `3.2.3`, imported as `eal`, and its source notation uses newline-terminated fields and typed flows. The source language, persisted `EAL/observation-record/1`, typed `EAL/typed-input/1`, compact `EAL/assessment-packet/2` and registered `EAL/registered-assessment/1` results are separate contracts. See [language syntax](docs/language.md), [reasoning modes](docs/reasoning-modes.md) and the [argument service](docs/argument-service.md).
|
|
4
|
+
|
|
5
|
+
Source fields use newlines, contexts resolve metadata defaults, and typed support flows lower to the existing argument model. Package, source notation and stored observation versions remain separate. Strictness, ranks and contraries retain their opt-in ASPIC+ compilation effects.
|
|
6
|
+
|
|
7
|
+
## Distribution and supported Python imports
|
|
8
|
+
|
|
9
|
+
The distribution installs the `eal` Python package and the `eal`, `eal-mcp` and `eal-host` console commands. The supported application imports are the documented interfaces in `eal.knowledge`, `eal.runtime`, `eal.parser`, `eal.semantics`, `eal.formatter`, `eal.evaluator`, `eal.methods` and `eal.limits`; [packaging and releases](docs/packaging.md#python-integration) gives their exact import paths. Private helpers, generated-parser internals and database tables are implementation details. A package version identifies the distributed implementation; the source and record schema identifiers identify their own contracts. Integrators pin a reviewed package version and check the current contract when upgrading. The project develops the current EAL/3 contract without a promise of backwards compatibility.
|
|
10
|
+
|
|
11
|
+
Python 3.11 or later is required. Command collectors and bounded custom-method workers require POSIX process and resource facilities. The release workflow is configured to install both the wheel and source distribution outside the checkout and exercise the public imports, console commands and MCP on Linux and macOS. The [Unlicense](LICENSE) covers original EARL software; third-party terms and the scope of repository materials are described in [licence scope](docs/packaging.md#licence-scope).
|
|
12
|
+
|
|
13
|
+
|
|
14
|
+
## Authored source and observations
|
|
15
|
+
|
|
16
|
+
`parse(source: str) -> Program` parses exact EAL/3 UTF-8 text; `validate(program, *, registry=None)` returns structured diagnostics for references, types and method contracts. `format_source(source)` returns canonical source. `Program.source_digest` hashes exact source bytes, even when two programs have the same meaning.
|
|
17
|
+
|
|
18
|
+
An EAL a `tool NAME` block containing `version "VERSION"` declares an interface. An `evidence NAME` selects that tool, evidence `kind`, `environment`, `max_age` in seconds, optional JSON `input` and one or more `require "VALUE.FIELD" OP SCALAR` predicates. A `reasoning NAME` selects exactly one installed versioned method, rationale, optional evidence `backing` and optional output predicates. Claims, arguments, premises, assumptions, objections, patterns and optional formal directives state the argument graph. The host never treats prose as a mechanically proven warrant.
|
|
19
|
+
|
|
20
|
+
An observation is not an EAL declaration. The configured tool returns JSON with `value` and optional `observed_at`, `context`, `request` and `details`; a file import must provide the latter acquisition fields and original time. The host persists the observation under `EAL/observation-record/1` with evidence ID, kind, environment, tool/version, input and request identity, context, source digest, original time, result digest, binding identity and a durable run ID. The original time is never advanced by reuse. A command receives one JSON object with `evidence_id`, `environment`, `tool`, `tool_version`, `input` and `context` on stdin and returns one JSON object on stdout. EAL source cannot select a command, file path, credential or Python method implementation.
|
|
21
|
+
|
|
22
|
+
`evidence.max_age` checks observation freshness at assessment time. `assumption.valid_from` and `valid_until` define a half-open interval during which that assumption may apply; its `validate` clause names an evidence declaration. These are independent clocks: recollecting evidence cannot extend an expired assumption. An absent, stale, invalid or failed observation cannot establish a negative result. A complete negative measurement may support an explicitly authored negative route.
|
|
23
|
+
|
|
24
|
+
## Operator tool bindings and method contracts
|
|
25
|
+
|
|
26
|
+
A sibling TOML registry binds a declared name/version to `command` (`argv`) or `json_file` (`path`). Command settings include timeout, output bound, optional explicit `env`, `inherit_env`, pinned files and `parallel_safe`. Commands run with host permissions; `parallel_safe = true` asserts independent read-only calls, which the bounded scheduler may overlap while serial calls remain barriers. A configured command may authenticate or query an external system. The application controls executable code and credentials through the registry and host environment.
|
|
27
|
+
|
|
28
|
+
The selected binding has a keyed identity. Command records also carry a keyed digest of the effective process environment, including selected inherited variables and explicit overrides. A default command inherits the full host environment; `inherit_env = ["PATH", "KUBECONFIG"]` narrows it to selected names. A changed tool binding, pinned dependency or effective environment makes prior output ineligible for automatic reuse. The store-local private binding key accompanies a moved database. Current binding checks are performed when reasoning over stored collections.
|
|
29
|
+
|
|
30
|
+
The default method registry contains `structured/1`, `deductive/1`, `inductive/1`, `abductive/1`, `causal/1`, `counterfactual/1`, `analogical/1` and `temporal/1`. Each method has a typed input kind, resource bound, computation and output interpretation. Computational methods use exactly one suitable designated observation from direct evidence, reasoning backing and assumption validation; an argument premise remains a separate claim dependency. `MethodContract` and `MethodRegistry.with_method` install additional immutable host methods; `--methods package.module:function` loads a trusted factory. A source selects a method ID, never code. The optional `eal.aspic:aspic_registry` installs a bounded formal ASPIC+ method, and `compile_eal_aspic` derives a separate formal snapshot from checked EAL routes.
|
|
31
|
+
|
|
32
|
+
## Persistent developer knowledge
|
|
33
|
+
|
|
34
|
+
`EALKnowledgeBase(workspace, registry_path=None, database_path=None, *, method_registry=None)` composes `ReasoningService`, `WorkspaceKnowledgeCatalogue` and `RegisteredAssessmentHost`. The default SQLite path is `.eal/runs.sqlite3` under the workspace. `register(path, *, entry_id=None, context=None, claims=None, label=None)` stores a validated workspace `.eal` file under a stable developer ID (the relative path by default), default context and selected claim IDs. A valid file edit creates an immutable source/method revision; an invalid edit fails without replacing a historical revision. `register_tree`, `sources`, `find`, `history` and catalogue `revisions` support bounded discovery and prior-run lookup. Search results are advisory metadata; assessment takes an exact entry ID and selected claim.
|
|
35
|
+
|
|
36
|
+
`EALKnowledgeBase.assess(entry_id, claim, *, context=None, now=None, reuse="compatible")` plans the complete authored support and objection closure. For each evidence ID it selects a successful stored observation only if its request, tool binding, execution environment, context, kind and original age remain compatible. It collects missing or expired evidence, stores a new source-bound collection, evaluates the declared methods at the supplied timezone-aware time or current UTC, and returns status, identities, reused/collected counts, a bounded packet and full-explanation reference. Compatible prior observations may come from different sessions and source revisions. `reuse="fresh"` is an explicit operator request to reacquire every selected evidence ID. A prior claim status is never silently substituted for the current assessment.
|
|
37
|
+
|
|
38
|
+
`ModelContextAdapter(knowledge).prepare(question, entry_id, claim, *, context=None, now=None, reuse="compatible")` returns `EAL/model-context/2` with the checked assessment and two prompt messages: bounded host context and the developer's question. The application chooses the entry and claim; free-form model text does not choose tools. It retains the host status separately from any generated prose. Model-facing MCP registered assessment uses the entry's stored context and exact claim selection, without a caller context override.
|
|
39
|
+
|
|
40
|
+
## Lower-level Python, CLI and MCP
|
|
41
|
+
|
|
42
|
+
`ReasoningService.plan(source, claim)` returns ordered evidence IDs, method-registry and source identities, calls and graph dependencies. `collect_claim(source, context, claim)` executes that closure. `collect(source, context, evidence_ids=None)` collects an explicit subset or all declarations. `reason(source, context, collection_id=None, now=None)` evaluates a matching stored collection; `packet(assessment_id, claim=None)` returns a bounded `EAL/assessment-packet/2`; `explain(assessment_id, claim=None)` retrieves the persisted assessment. `validate`, `format`, `describe` and optional `compile_aspic` share the same contracts. Pure `evaluate(program, records, *, now, context, registry=None, binding_digests=None)` performs no implicit acquisition or clock selection.
|
|
43
|
+
|
|
44
|
+
The CLI exposes `register`, `register-tree`, `sources`, `find`, `assess-known`, `history`, `model-context`, `validate`, `format`, `plan`, `collect`, `reason`, `packet`, `explain`, `grounded`, `compile-aspic` and `export-aspic`. The generic MCP server exposes the corresponding source-level operations. Starting `eal-mcp` with one or more `--known-entry ENTRY_ID` options instead exposes only `eal_sources`, `eal_find_claims` and `eal_assess_known` for those registered entries; it does not expose raw explanation or source-level collection on that endpoint. `eal-host --known-entry ENTRY_ID` accepts a single strict JSON text request through the same MCP server for a model without native tools.
|
|
45
|
+
|
|
46
|
+
FastMCP `4.0.10` supplies stdio and Streamable HTTP transports through MCP Python SDK `2.2.0`. A shared operation contract defines names and permitted fields; the operation catalogue binds typed handlers and exposure selection. FastMCP derives their schemas, which the strict JSON host discovers and checks before invocation. Both transports reject extra fields and invalid JSON types before handler invocation and expose equivalent application results. Protocol negotiation and HTTP authentication belong to the transport adapter; EAL assessment semantics remain in the SDK-independent service. The supported protocol paths include the `2025-11-25` handshake and `2026-07-28` discovery profiles.
|
|
47
|
+
|
|
48
|
+
`eal-mcp --config FILE` loads validated `[service]`, `[server]` and `[server.http]` TOML settings. Precedence is CLI > `EAL_MCP_*` environment > TOML > defaults. File-supplied paths are relative to the configuration file's directory; CLI/environment paths are relative to the working directory. `--transport stdio|http` selects one transport per process. `--exposure operator|registered` must agree with the configured entry selection; absent an explicit exposure, a non-empty selection chooses registered exposure. Transport selection does not alter tool exposure. `--max-in-flight` defaults to one admitted tool call per process.
|
|
49
|
+
|
|
50
|
+
The HTTP runner always enables strict Host/Origin protection before MCP handling. Accepted hosts include FastMCP's loopback defaults, the configured host, the actual bound address and `allowed_hosts` additions. Supplied origins must match the request's origin, qualify as loopback on a loopback request, or appear in `allowed_origins`. Untrusted Host and Origin headers receive HTTP 421 and 403 respectively. `[server.http].allowed_hosts` and `.allowed_origins` contain up to 128 unique literal entries each; repeated `--allowed-host` / `--allowed-origin` options and JSON-array `EAL_MCP_ALLOWED_HOSTS` / `EAL_MCP_ALLOWED_ORIGINS` variables follow normal precedence. Wildcards are rejected.
|
|
51
|
+
|
|
52
|
+
Non-loopback HTTP binds and additional non-loopback trusted hosts/origins require `--token-env NAME` or `[server.http].token_env`, including public proxies forwarding to loopback. Startup reads the token from that environment variable; settings and source contain only its name. The bearer credential grants access to the endpoint's fixed operator workspace or registered-entry allowlist. The service implements one host trust domain per endpoint. HTTP clients use `eal-host --url URL [--token-env NAME]`; this connects to an already running service and treats the credential as an opaque bearer value. The host's default route launches a stdio subprocess, and its flat JSON `compile_aspic` operation uses the same checked contract as `eal_compile_aspic`. Credential sanitisation preserves JSON field names and public host metadata, while redacting application string values and diagnostics, including values in JSON-bearing content text. See [MCP and tools](docs/mcp-and-tools.md) for settings and examples.
|
|
53
|
+
|
|
54
|
+
Collections bind exact source bytes and context. The evaluator reports `supported`, `contested`, `unsupported` or `out_of_scope` for claims; it also returns named argument, evidence, assumption, objection and reasoning results. Objections and support dependencies are solved under the finite authored argument model. An `unsupported` result does not assert the opposite claim. Typed propositions check formal correspondence and result predicates within the installed method's limits; human review remains responsible for whether an authored claim and its measurement answer the intended engineering question.
|
|
55
|
+
|
|
56
|
+
## Bounds and persistence
|
|
57
|
+
|
|
58
|
+
Source/graph/construction/search budgets are operator-owned `ExecutionLimits`. CLI, MCP and JSON launchers accept `--limits FILE`; source cannot raise them. `[limits]` TOML values are positive integers. Discovery advertises the selected budgets. Compile/export captures them in the checked snapshot, and exporting under a smaller current host budget is rejected. An exhausted calculation is incomplete and supplies no accepted or rejected conclusion. See [scoped composition](docs/eal3-composition.md).
|
|
59
|
+
|
|
60
|
+
Collection preflight bounds context JSON to 16 KiB, a host-configurable evidence count (default 128) and each request to 1 MiB. The sum of configured tool output allowances is at most 128 MiB, individual command outputs at most 16 MiB, and the final collection JSON at most 32 MiB. Packet output defaults to a 16 KiB bound with explicit omission counts; it excludes raw observation values, command streams and arbitrary extension outputs. SQLite stores immutable observations, collections and assessments under private files. Stored records remain historical results; a new assessment checks current source, method, binding and observation freshness. See [MCP and tools](docs/mcp-and-tools.md) for exact adapter envelopes and failure handling.
|
|
61
|
+
|
|
62
|
+
A workspace acquisition lease coordinates MCP collection and registered-assessment reuse checks across threads and server processes on a local filesystem. It prevents overlapping acquisitions through the participating MCP services; eligible `parallel_safe` collectors may still overlap within one admitted acquisition. A POSIX command supervisor inherits the lock descriptor and retains it until its collector process group has stopped and been reaped, including after a host deadline or abrupt server termination. Normal collector completion also cleans up residual group members. Configured commands must keep descendants in the owned process group; remotely started work requires separate ownership and cancellation. Command execution requires POSIX and the host's permissions.
|
|
63
|
+
|
|
64
|
+
Processes launched with stdio and HTTP can share that workspace and database. A private sibling store lock serialises binding-key, WAL, schema and observation-index initialisation during concurrent startup. SQLite connections close after commit or rollback. This coordination contract covers a single local filesystem; distributed replicas and network filesystems require a different coordinator. Admission limits and acquisition leases constrain execution; throughput and model-cost improvements require measurements.
|
|
65
|
+
|
|
66
|
+
## Investigation contracts
|
|
67
|
+
|
|
68
|
+
Protocol 5.0.0 uses primary plan/report `/4`, diagnostic plan/report `/2`, and
|
|
69
|
+
information-design configuration/result `/2`. These experiment contracts do not
|
|
70
|
+
change EAL/3 semantics. The default population contains six evaluation cases;
|
|
71
|
+
eight threshold cases serve calibration only. A frozen
|
|
72
|
+
`EAL/evaluation-task-manifest/1` records independently supplied task provenance,
|
|
73
|
+
reference-checked answer keys and explicit evidence revisions. Revisions enter
|
|
74
|
+
assessment context so incompatible snapshots cannot be reused.
|
|
75
|
+
|
|
76
|
+
Correctness inference uses independent paired trajectories and prospectively
|
|
77
|
+
selected empirical Bernstein bounds. Unknown outcomes retain their planned
|
|
78
|
+
units and completion envelopes. Simultaneous estimation and the one-sided
|
|
79
|
+
intersection-union adoption decision are separate outputs. Resource inference
|
|
80
|
+
is approximate; allocation requires declared nuisance scenarios and fresh-seed
|
|
81
|
+
Monte Carlo calibration. Scripted runs cannot produce an empirical evaluation
|
|
82
|
+
allocation.
|
|
83
|
+
|
|
84
|
+
Matched-fact diagnostics independently vary withheld, EAL-derived and conventional
|
|
85
|
+
conclusions. Failed manipulation checks retain observations and costs but block
|
|
86
|
+
component attribution. The optional `EAL/adoption-cost-ledger/1` binds measured
|
|
87
|
+
activity to a run and plan digest; missing rates or coverage remain unknown.
|
|
88
|
+
See [methodology](docs/eal3-experiment-methodology.md) and
|
|
89
|
+
[verification](experiments/model_transfer/verification.md).
|