robotics-acceptance-harness 0.19.0__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (119) hide show
  1. robotics_acceptance_harness-0.19.0/.editorconfig +10 -0
  2. robotics_acceptance_harness-0.19.0/.gitattributes +4 -0
  3. robotics_acceptance_harness-0.19.0/.gitignore +17 -0
  4. robotics_acceptance_harness-0.19.0/.markdownlint-cli2.yaml +7 -0
  5. robotics_acceptance_harness-0.19.0/.python-version +1 -0
  6. robotics_acceptance_harness-0.19.0/.semgrep/attach-only.py +68 -0
  7. robotics_acceptance_harness-0.19.0/.semgrep/attach-only.yml +78 -0
  8. robotics_acceptance_harness-0.19.0/.yamllint.yml +15 -0
  9. robotics_acceptance_harness-0.19.0/CHANGELOG.md +27 -0
  10. robotics_acceptance_harness-0.19.0/CONTRIBUTING.md +28 -0
  11. robotics_acceptance_harness-0.19.0/LICENSE +21 -0
  12. robotics_acceptance_harness-0.19.0/PKG-INFO +394 -0
  13. robotics_acceptance_harness-0.19.0/QUALITY_DECLARATION.md +48 -0
  14. robotics_acceptance_harness-0.19.0/README.md +366 -0
  15. robotics_acceptance_harness-0.19.0/SECURITY.md +33 -0
  16. robotics_acceptance_harness-0.19.0/docs/compatibility.md +45 -0
  17. robotics_acceptance_harness-0.19.0/docs/decisions/0001-runtime-owns-the-execution-environment.md +33 -0
  18. robotics_acceptance_harness-0.19.0/docs/decisions/0004-aggregate-per-domain-results.md +32 -0
  19. robotics_acceptance_harness-0.19.0/docs/decisions/0005-consume-qualified-accelerator-evidence.md +33 -0
  20. robotics_acceptance_harness-0.19.0/docs/decisions/0009-separate-permit-and-workload-identities.md +38 -0
  21. robotics_acceptance_harness-0.19.0/docs/decisions/README.md +15 -0
  22. robotics_acceptance_harness-0.19.0/docs/evidence-files.md +28 -0
  23. robotics_acceptance_harness-0.19.0/docs/live-tests.md +127 -0
  24. robotics_acceptance_harness-0.19.0/docs/stdlib-helpers.md +19 -0
  25. robotics_acceptance_harness-0.19.0/docs/supply-chain.md +45 -0
  26. robotics_acceptance_harness-0.19.0/pyproject.toml +59 -0
  27. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/__init__.py +24 -0
  28. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/_evidence_files.py +73 -0
  29. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/_histogram_estimates.py +128 -0
  30. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/aggregate.py +612 -0
  31. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/application.py +600 -0
  32. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/authorization.py +222 -0
  33. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/campaign.py +130 -0
  34. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/cli.py +632 -0
  35. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/diagnostics.py +271 -0
  36. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/documents.py +390 -0
  37. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/errors.py +67 -0
  38. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/evaluation.py +527 -0
  39. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/evidence.py +261 -0
  40. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/extension_schemas.py +31 -0
  41. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/forbidden_graph.py +66 -0
  42. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/graph_window.py +81 -0
  43. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/hardware_timing.py +160 -0
  44. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/metrics.py +1045 -0
  45. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/otel.py +265 -0
  46. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/plugin.py +174 -0
  47. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/policy.py +329 -0
  48. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/py.typed +0 -0
  49. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/readiness.py +291 -0
  50. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/receipts.py +233 -0
  51. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/result.py +418 -0
  52. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/ros.py +386 -0
  53. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/run_context.py +80 -0
  54. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/time_authority.py +183 -0
  55. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/timing.py +463 -0
  56. robotics_acceptance_harness-0.19.0/src/robotics_acceptance_harness/traces.py +860 -0
  57. robotics_acceptance_harness-0.19.0/tests/__init__.py +1 -0
  58. robotics_acceptance_harness-0.19.0/tests/conftest.py +3 -0
  59. robotics_acceptance_harness-0.19.0/tests/fixtures/histograms/cumulative-baseline.jsonl +1 -0
  60. robotics_acceptance_harness-0.19.0/tests/fixtures/histograms/cumulative-reset.jsonl +1 -0
  61. robotics_acceptance_harness-0.19.0/tests/fixtures/histograms/delta.jsonl +1 -0
  62. robotics_acceptance_harness-0.19.0/tests/fixtures/physical/hil-permit.json +30 -0
  63. robotics_acceptance_harness-0.19.0/tests/fixtures/physical/hil-runtime.json +112 -0
  64. robotics_acceptance_harness-0.19.0/tests/fixtures/physical/hil-scenario.yaml +71 -0
  65. robotics_acceptance_harness-0.19.0/tests/fixtures/physical/hil-verification.json +36 -0
  66. robotics_acceptance_harness-0.19.0/tests/fixtures/simulation/runtime.yaml +71 -0
  67. robotics_acceptance_harness-0.19.0/tests/fixtures/simulation/scenario.yaml +68 -0
  68. robotics_acceptance_harness-0.19.0/tests/graph_types.py +32 -0
  69. robotics_acceptance_harness-0.19.0/tests/live/Dockerfile +37 -0
  70. robotics_acceptance_harness-0.19.0/tests/live/Dockerfile.dockerignore +22 -0
  71. robotics_acceptance_harness-0.19.0/tests/live/__init__.py +1 -0
  72. robotics_acceptance_harness-0.19.0/tests/live/collector.py +106 -0
  73. robotics_acceptance_harness-0.19.0/tests/live/conftest.py +64 -0
  74. robotics_acceptance_harness-0.19.0/tests/live/evidence.py +174 -0
  75. robotics_acceptance_harness-0.19.0/tests/live/fixtures/metric-golden.json +9 -0
  76. robotics_acceptance_harness-0.19.0/tests/live/fixtures/qualification-golden.json +12 -0
  77. robotics_acceptance_harness-0.19.0/tests/live/graph.py +172 -0
  78. robotics_acceptance_harness-0.19.0/tests/live/run.sh +12 -0
  79. robotics_acceptance_harness-0.19.0/tests/live/test_cli.py +361 -0
  80. robotics_acceptance_harness-0.19.0/tests/live/test_ros.py +119 -0
  81. robotics_acceptance_harness-0.19.0/tests/live/types.py +125 -0
  82. robotics_acceptance_harness-0.19.0/tests/support.py +283 -0
  83. robotics_acceptance_harness-0.19.0/tests/test_aggregate.py +986 -0
  84. robotics_acceptance_harness-0.19.0/tests/test_application.py +720 -0
  85. robotics_acceptance_harness-0.19.0/tests/test_application_graph.py +77 -0
  86. robotics_acceptance_harness-0.19.0/tests/test_application_timing.py +110 -0
  87. robotics_acceptance_harness-0.19.0/tests/test_application_timing_bounds.py +256 -0
  88. robotics_acceptance_harness-0.19.0/tests/test_application_timing_uncertainty.py +217 -0
  89. robotics_acceptance_harness-0.19.0/tests/test_authorization.py +141 -0
  90. robotics_acceptance_harness-0.19.0/tests/test_campaign.py +133 -0
  91. robotics_acceptance_harness-0.19.0/tests/test_cli.py +624 -0
  92. robotics_acceptance_harness-0.19.0/tests/test_cli_errors.py +215 -0
  93. robotics_acceptance_harness-0.19.0/tests/test_diagnostics.py +125 -0
  94. robotics_acceptance_harness-0.19.0/tests/test_documents.py +97 -0
  95. robotics_acceptance_harness-0.19.0/tests/test_evaluation.py +473 -0
  96. robotics_acceptance_harness-0.19.0/tests/test_evaluator_record_loading.py +624 -0
  97. robotics_acceptance_harness-0.19.0/tests/test_evidence.py +167 -0
  98. robotics_acceptance_harness-0.19.0/tests/test_evidence_files.py +159 -0
  99. robotics_acceptance_harness-0.19.0/tests/test_forbidden_graph.py +49 -0
  100. robotics_acceptance_harness-0.19.0/tests/test_graph_window.py +186 -0
  101. robotics_acceptance_harness-0.19.0/tests/test_hardware_timing.py +122 -0
  102. robotics_acceptance_harness-0.19.0/tests/test_histogram_estimates.py +129 -0
  103. robotics_acceptance_harness-0.19.0/tests/test_histogram_windows.py +315 -0
  104. robotics_acceptance_harness-0.19.0/tests/test_metric_window_coverage.py +211 -0
  105. robotics_acceptance_harness-0.19.0/tests/test_metrics.py +602 -0
  106. robotics_acceptance_harness-0.19.0/tests/test_otel.py +180 -0
  107. robotics_acceptance_harness-0.19.0/tests/test_plugin.py +184 -0
  108. robotics_acceptance_harness-0.19.0/tests/test_policy.py +580 -0
  109. robotics_acceptance_harness-0.19.0/tests/test_readiness.py +270 -0
  110. robotics_acceptance_harness-0.19.0/tests/test_receipt_inventory.py +233 -0
  111. robotics_acceptance_harness-0.19.0/tests/test_result.py +336 -0
  112. robotics_acceptance_harness-0.19.0/tests/test_result_io.py +80 -0
  113. robotics_acceptance_harness-0.19.0/tests/test_ros.py +461 -0
  114. robotics_acceptance_harness-0.19.0/tests/test_safety_boundary.py +24 -0
  115. robotics_acceptance_harness-0.19.0/tests/test_time_authority.py +164 -0
  116. robotics_acceptance_harness-0.19.0/tests/test_timing.py +297 -0
  117. robotics_acceptance_harness-0.19.0/tests/test_timing_windows.py +223 -0
  118. robotics_acceptance_harness-0.19.0/tests/test_traces.py +655 -0
  119. robotics_acceptance_harness-0.19.0/tests/test_version.py +7 -0
@@ -0,0 +1,10 @@
1
+ root = true
2
+
3
+ [*]
4
+ charset = utf-8
5
+ end_of_line = lf
6
+ insert_final_newline = true
7
+ trim_trailing_whitespace = true
8
+
9
+ [*.md]
10
+ trim_trailing_whitespace = false
@@ -0,0 +1,4 @@
1
+ * text=auto eol=lf
2
+
3
+ *.bat text eol=crlf
4
+ *.cmd text eol=crlf
@@ -0,0 +1,17 @@
1
+ .venv/
2
+ .wheel-venv/
3
+ build/
4
+ dist/
5
+ .pytest_cache/
6
+ .ruff_cache/
7
+ .coverage
8
+ htmlcov/
9
+ __pycache__/
10
+ *.py[cod]
11
+ artifacts/
12
+ runs/
13
+ tmp/
14
+ *.egg-info/
15
+ .env
16
+ .env.*
17
+ !.env.example
@@ -0,0 +1,7 @@
1
+ ---
2
+ config:
3
+ line-length:
4
+ line_length: 100
5
+ tables: false
6
+ code_blocks: false
7
+ headings: false
@@ -0,0 +1,68 @@
1
+ # ruff: noqa: E402, F401
2
+
3
+ import os
4
+
5
+ # ruleid: attach-only-no-process-control
6
+ import subprocess
7
+
8
+ # ruleid: attach-only-no-orchestrator-sdk
9
+ import docker
10
+
11
+ # ruleid: attach-only-no-network-client
12
+ import requests
13
+
14
+
15
+ def allowed_helper():
16
+ # ok: attach-only-no-network-client
17
+ from urllib.request import url2pathname
18
+
19
+
20
+ def allowed_helper_alias():
21
+ # ok: attach-only-no-network-client
22
+ from urllib.request import url2pathname as to_path
23
+
24
+
25
+ def forbidden_mixed_import():
26
+ # ruleid: attach-only-no-network-client
27
+ from urllib.request import build_opener, url2pathname
28
+
29
+
30
+ def forbidden_client_import():
31
+ # ruleid: attach-only-no-network-client
32
+ from urllib.request import urlopen
33
+
34
+
35
+ def forbidden_client_alias():
36
+ # ruleid: attach-only-no-network-client
37
+ from urllib.request import urlretrieve as retrieve
38
+
39
+
40
+ def forbidden_module_import():
41
+ # ruleid: attach-only-no-network-client
42
+ import urllib.request
43
+
44
+
45
+ def forbidden_module_alias():
46
+ # ruleid: attach-only-no-network-client
47
+ import urllib.request as client
48
+
49
+
50
+ # ruleid: attach-only-no-mutation-service-types
51
+ from lifecycle_msgs.srv import ChangeState
52
+
53
+
54
+ def forbidden_ros_apis(node, client):
55
+ # ruleid: attach-only-no-ros-publisher
56
+ node.create_publisher(object, "/command", 10)
57
+ # ruleid: attach-only-no-action-client
58
+ client.send_goal_async(object())
59
+
60
+
61
+ def forbidden_os_api():
62
+ # ruleid: attach-only-no-process-control
63
+ os.execv("/bin/false", ["false"])
64
+
65
+
66
+ def allowed_observer_apis(node):
67
+ # ok: attach-only-no-ros-publisher
68
+ node.create_subscription(object, "/observation", lambda message: message, 10)
@@ -0,0 +1,78 @@
1
+ ---
2
+ rules:
3
+ - id: attach-only-no-process-control
4
+ message: The acceptance harness must not start or control external processes.
5
+ severity: ERROR
6
+ languages: [python]
7
+ pattern-either:
8
+ - pattern: import subprocess
9
+ - pattern: from subprocess import $X
10
+ - pattern: os.system(...)
11
+ - pattern: os.popen(...)
12
+ - pattern-regex: '\bos\.(?:exec|spawn)[a-z]*\s*\('
13
+
14
+ - id: attach-only-no-orchestrator-sdk
15
+ message: Process orchestration and object storage belong to the runtime infrastructure.
16
+ severity: ERROR
17
+ languages: [python]
18
+ pattern-either:
19
+ - pattern: import docker
20
+ - pattern: from docker import $X
21
+ - pattern: import launch
22
+ - pattern: from launch import $X
23
+ - pattern: import launch_ros
24
+ - pattern: from launch_ros import $X
25
+ - pattern: import boto3
26
+ - pattern: from boto3 import $X
27
+ - pattern: import botocore
28
+ - pattern: from botocore import $X
29
+
30
+ - id: attach-only-no-network-client
31
+ message: The acceptance harness validates local inputs and must not fetch them over the network.
32
+ severity: ERROR
33
+ languages: [python]
34
+ patterns:
35
+ - pattern-either:
36
+ - pattern: import requests
37
+ - pattern: from requests import $X
38
+ - pattern: import httpx
39
+ - pattern: from httpx import $X
40
+ - pattern: import aiohttp
41
+ - pattern: from aiohttp import $X
42
+ - pattern: import urllib.request
43
+ - pattern: from urllib.request import $X
44
+ - pattern: urllib.request.urlopen(...)
45
+ - pattern: urllib.request.urlretrieve(...)
46
+ - pattern: import socket
47
+ - pattern: from socket import $X
48
+ # Import equivalence can also match the module-import pattern above.
49
+ # Exclude only a complete single-helper import, never a mixed import.
50
+ - pattern-not-regex: >-
51
+ (?m)^[ \t]*from urllib\.request import url2pathname(?: as [A-Za-z_]\w*)?[ \t]*$
52
+
53
+ - id: attach-only-no-ros-publisher
54
+ message: The acceptance harness may subscribe and query but must not publish ROS commands.
55
+ severity: ERROR
56
+ languages: [python]
57
+ pattern: $NODE.create_publisher(...)
58
+
59
+ - id: attach-only-no-action-client
60
+ message: The acceptance harness must not send ROS action goals.
61
+ severity: ERROR
62
+ languages: [python]
63
+ pattern-either:
64
+ - pattern: $CLIENT.send_goal(...)
65
+ - pattern: $CLIENT.send_goal_async(...)
66
+ - pattern: ActionClient(...)
67
+
68
+ - id: attach-only-no-mutation-service-types
69
+ message: The acceptance harness must not import mutation or actuation service types.
70
+ severity: ERROR
71
+ languages: [python]
72
+ pattern-either:
73
+ - pattern: from lifecycle_msgs.srv import ChangeState
74
+ - pattern: from ros_gz_interfaces.srv import ControlWorld
75
+ - pattern: from controller_manager_msgs.srv import SwitchController
76
+ - pattern: from mavros_msgs.srv import CommandBool
77
+ - pattern: from mavros_msgs.srv import CommandLong
78
+ - pattern: from mavros_msgs.srv import SetMode
@@ -0,0 +1,15 @@
1
+ ---
2
+ extends: default
3
+
4
+ ignore: |
5
+ .venv/
6
+ artifacts/
7
+ tmp/
8
+
9
+ rules:
10
+ line-length:
11
+ max: 120
12
+ level: warning
13
+ truthy:
14
+ allowed-values: ["true", "false"]
15
+ check-keys: false
@@ -0,0 +1,27 @@
1
+ # Changelog
2
+
3
+ ## 0.19.0
4
+
5
+ The first harness release from the `robotics-runtime` workspace requires
6
+ contracts 0.18 (`>=0.18,<0.19`). Its release prerequisite is
7
+ [contracts-v0.18.1](https://github.com/mmkolpakov/robotics-runtime/releases/tag/contracts-v0.18.1).
8
+ The wheel and source distribution remain independently installable.
9
+
10
+ - Evaluate causal message pairs and connected trace paths, attributing channel
11
+ spans by their declared topic. Fold acceptance and campaign verdicts through
12
+ the canonical status order.
13
+ - Recheck expected graph and lifecycle conditions during measurement. Expire
14
+ cached lifecycle state when a GetState request stops answering.
15
+ - Evaluate realtime factor in sliding windows, cumulative histogram baselines
16
+ and resets, and interior gaps in delta coverage. Retain proven violations
17
+ when other observations are unavailable.
18
+ - Read finalized evidence through bounded, contained file descriptors and retry
19
+ partial or missing inputs to the deadline. Write UTC result timestamps and
20
+ preserve an existing run context when create-run is repeated.
21
+ - Accept finalized receipt inventories and digest-pinned extension schemas.
22
+ Verify evaluator sources against receipt and RECORD data and reject external
23
+ bytecode caches.
24
+ - Expose HarnessError and HarnessInputError with stable identifiers and exit
25
+ codes. Preserve dependency causes and structured issue paths in diagnostics.
26
+ - Ship type information, support Python 3.12–3.14 and relax supported runtime
27
+ dependency ranges.
@@ -0,0 +1,28 @@
1
+ # Contributing
2
+
3
+ Open an issue before changing a public contract, safety boundary, or supported
4
+ execution mode. Keep product scenes, robot commands, model weights, and launch
5
+ orchestration in consuming repositories.
6
+
7
+ ## Local Checks
8
+
9
+ ```bash
10
+ uv sync --locked --all-groups
11
+ uv run pre-commit run --all-files
12
+ uv run pytest \
13
+ -p robotics_acceptance_harness.plugin \
14
+ --robotics-scenario tests/fixtures/simulation/scenario.yaml \
15
+ --robotics-runtime tests/fixtures/simulation/runtime.yaml
16
+ uv build --no-sources
17
+ ```
18
+
19
+ Place `robotics-runtime-contracts` next to this repository. The standard
20
+ dependency remains a Semantic Versioning range; `tool.uv.sources` uses the
21
+ sibling checkout only for development and CI.
22
+
23
+ Every behavioral change needs a focused test. Changes to Semgrep policy need a
24
+ matching `ruleid` or `ok` example in `.semgrep/attach-only.py`. Pull requests
25
+ must pass the required `test` check and resolve all review conversations.
26
+
27
+ Use Conventional Commit subjects. Do not commit generated results, evidence,
28
+ private scenarios, credentials, or hardware identities.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Maxim Kolpakov
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,394 @@
1
+ Metadata-Version: 2.5
2
+ Name: robotics-acceptance-harness
3
+ Version: 0.19.0
4
+ Summary: Attach-only acceptance observer for validated robotics executions.
5
+ Project-URL: Homepage, https://github.com/mmkolpakov/robotics-runtime/tree/main/packages/harness
6
+ Project-URL: Repository, https://github.com/mmkolpakov/robotics-runtime
7
+ Project-URL: Issues, https://github.com/mmkolpakov/robotics-runtime/issues
8
+ Author: mmkolpakov
9
+ License-Expression: MIT
10
+ License-File: LICENSE
11
+ Keywords: acceptance,robotics,ros2,simulation,testing
12
+ Classifier: Development Status :: 4 - Beta
13
+ Classifier: Environment :: Console
14
+ Classifier: License :: OSI Approved :: MIT License
15
+ Classifier: Operating System :: POSIX :: Linux
16
+ Classifier: Programming Language :: Python :: 3.12
17
+ Classifier: Programming Language :: Python :: 3.13
18
+ Classifier: Programming Language :: Python :: 3.14
19
+ Classifier: Topic :: Software Development :: Testing
20
+ Requires-Python: <3.15,>=3.12
21
+ Requires-Dist: junitparser<6,>=5.0.1
22
+ Requires-Dist: opentelemetry-proto<2,>=1.43
23
+ Requires-Dist: packaging>=24
24
+ Requires-Dist: robotics-runtime-contracts<0.19,>=0.18
25
+ Provides-Extra: pytest
26
+ Requires-Dist: pytest<10,>=8; extra == 'pytest'
27
+ Description-Content-Type: text/markdown
28
+
29
+ # Robotics Acceptance Harness
30
+
31
+ [![CI](https://github.com/mmkolpakov/robotics-runtime/actions/workflows/ci.yml/badge.svg)](https://github.com/mmkolpakov/robotics-runtime/actions/workflows/ci.yml)
32
+ [![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
33
+
34
+ Attach-only acceptance testing for existing ROS 2 executions.
35
+
36
+ The harness validates an execution bundle, observes the declared ROS graph and
37
+ OpenTelemetry data, verifies retained evidence, and writes contract-valid JSON
38
+ and JUnit results. It does not launch workloads, control simulators, change node
39
+ lifecycle states, execute cryptographic signature tools, or publish commands to equipment.
40
+ It verifies externally produced signature results and their digest chain; the
41
+ signature tool itself remains an infrastructure responsibility.
42
+
43
+ See [local evidence and timestamps](docs/evidence-files.md) for filesystem
44
+ containment, finalized-file reads and the distinction between Unix and monotonic time.
45
+
46
+ ## Architecture
47
+
48
+ ```text
49
+ product workload -> runtime infrastructure -> running ROS 2 graph
50
+ |
51
+ runtime contracts -> acceptance harness -----------+
52
+ |
53
+ +-> result JSON + JUnit + evidence links
54
+ ```
55
+
56
+ - [robotics-runtime-contracts](https://github.com/mmkolpakov/robotics-runtime/tree/main/packages/contracts)
57
+ owns document structure and verdict semantics.
58
+ - This repository owns observation and evaluation.
59
+ - [robotics-runtime-infra](https://github.com/mmkolpakov/robotics-runtime-infra)
60
+ owns runtime, simulator, middleware, recorder, and hardware provider adapters.
61
+ - Product repositories own scenes, robots, models, behavior, and business
62
+ evaluators.
63
+
64
+ The harness consumes provider-neutral runtime facts. Adding a simulator,
65
+ middleware, recorder, or accelerator does not require a new scenario or result
66
+ version.
67
+
68
+ ## Requirements
69
+
70
+ | Component | Baseline |
71
+ | --- | --- |
72
+ | Python | 3.12 through 3.14 |
73
+ | Contracts | `robotics-runtime-contracts>=0.18,<0.19` |
74
+ | ROS observation | ROS 2 Jazzy packages in the observer environment |
75
+ | Metrics | OTLP JSON Lines exported by OpenTelemetry Collector |
76
+
77
+ All public contract families currently use one canonical `v1`. Published
78
+ schemas are checked for compatible changes against the release baseline.
79
+
80
+ ## Install
81
+
82
+ Development uses the exact contracts revision recorded in `uv.lock`:
83
+
84
+ ```bash
85
+ git clone https://github.com/mmkolpakov/robotics-runtime.git
86
+ cd robotics-runtime
87
+ uv sync --locked --all-packages --all-groups
88
+ uv run robotics-acceptance --version
89
+ cd packages/harness
90
+ ```
91
+
92
+ Release consumers should install the published wheel together with the locked
93
+ contracts wheel and verify release provenance as described in
94
+ [`docs/supply-chain.md`](docs/supply-chain.md).
95
+
96
+ ## Quick Start
97
+
98
+ Validate and cross-check a known-good bundle without ROS:
99
+
100
+ ```bash
101
+ uv run robotics-acceptance explain \
102
+ --scenario tests/fixtures/simulation/scenario.yaml \
103
+ --runtime tests/fixtures/simulation/runtime.yaml
104
+ ```
105
+
106
+ Create the immutable context shared by every domain in one run:
107
+
108
+ ```bash
109
+ robotics-acceptance create-run \
110
+ --scenario scenario.yaml \
111
+ --output acceptance-run.json \
112
+ --domain primary=observer \
113
+ --time-authority sim_clock \
114
+ --time-source simulation-clock
115
+ ```
116
+
117
+ Run `robotics-acceptance COMMAND --help` for the complete option set.
118
+
119
+ ## Commands
120
+
121
+ | Command | Purpose | Controls the workload |
122
+ | --- | --- | --- |
123
+ | `create-run` | Create an immutable run context | No |
124
+ | `explain` | Validate and explain an execution bundle | No |
125
+ | `verify` | Observe a live ROS 2 execution | No |
126
+ | `evaluate` | Re-evaluate finalized evidence offline | No |
127
+ | `aggregate` | Fold all declared domain results | No |
128
+ | `transport-evaluate` | Qualify cross-domain delivery and causal traces | No |
129
+ | `campaign` | Aggregate repeated run verdicts | No |
130
+ | `doctor` | Report observer and extension metadata | No |
131
+ | `why` | Explain a result verdict | No |
132
+ | `timing-check` | Check verified clock metrics against scenario policy | No |
133
+ | `otel-summary` | Summarize normalized OTLP metric points | No |
134
+
135
+ Verdict-producing commands return `0` for `passed` and `1` for a completed
136
+ non-passing verdict, including `failed`, `incomplete`, or `error`. An input,
137
+ observation, or execution exception handled by the CLI returns `2` with a
138
+ diagnostic. A result whose status is `error` is distinct from such an exception.
139
+
140
+ A domain result takes the most severe outcome of its assertions and of the
141
+ forbidden-graph, hardware-clock and time-authority observations. Skipped
142
+ assertions and declared `unevaluated` paths make an otherwise passing result
143
+ `incomplete`; they appear in JUnit as skipped cases, not failures.
144
+
145
+ `clock_observation.real_time_factor` and `deadline_miss_ratio` are measured only
146
+ in `simulation_realtime`; other time modes report `0`.
147
+
148
+ `campaign` reports `incomplete` when fewer runs passed than the required
149
+ minimum and no failed or error runs were observed. Tolerated failed or error
150
+ runs still allow `passed` when every threshold is met.
151
+
152
+ Modeled library failures inherit from the public `HarnessError` base, with an
153
+ explicit `error_id` and `exit_code` (normally `2`). Existing exception classes
154
+ retain their identifiers and their `ValueError`, `RuntimeError`, or `TimeoutError`
155
+ compatibility. Invalid values supplied to the harness raise `HarnessInputError`
156
+ with `input.invalid`.
157
+
158
+ The command boundary translates dependency failures before the CLI handles
159
+ `HarnessError`: contract errors retain their identifier, I/O errors use
160
+ `input.io_error`, invalid dependency values use `input.invalid`, and unexpected
161
+ exceptions use `internal.error`. Diagnostics retain the original exception type,
162
+ message, and any JSON paths; exception chaining preserves the original cause.
163
+ Readiness, bundle, and evidence errors expose structured `(json_path, message)`
164
+ pairs through `diagnostic_issues`. `KeyboardInterrupt` and `SystemExit` propagate.
165
+
166
+ ## Live Observation
167
+
168
+ Runtime infrastructure starts the workload, recorder, and telemetry collector.
169
+ The harness joins the existing ROS domain:
170
+
171
+ ```bash
172
+ robotics-acceptance verify \
173
+ --scenario scenario.yaml \
174
+ --runtime runtime-manifest.json \
175
+ --run-id run-7dd792f2-4f75-4f4d-81b0-48c8c2a8f76c \
176
+ --domain-id primary \
177
+ --run-context acceptance-run.json \
178
+ --evidence-index evidence-index.json \
179
+ --otel-metrics metrics.otlp.jsonl \
180
+ --measurement-complete measurement-complete \
181
+ --output results
182
+ ```
183
+
184
+ The output directory contains `acceptance-result.json` and `junit.xml`. The
185
+ observer inherits standard ROS variables such as `ROS_DOMAIN_ID`,
186
+ `RMW_IMPLEMENTATION`, and the SROS2 environment. It has no private fallback for
187
+ document paths or execution identity.
188
+
189
+ An expected topic's `qos_profile` selects the observer subscription's QoS.
190
+ The compatibility check compares discovered publishers with that subscription;
191
+ it does not compare every application publisher/subscriber pair. The harness
192
+ excludes its own subscriptions from the observed subscriber count. When the
193
+ scenario does not declare `/clock`, its observation subscription uses depth 1,
194
+ best-effort reliability, and volatile durability.
195
+
196
+ ## Offline Evaluation
197
+
198
+ `evaluate` runs the same metric, evidence, and product evaluators without
199
+ joining ROS. Live graph, clock, safety-boundary, and shutdown observations are
200
+ marked `unevaluated`, so an offline result cannot silently claim complete live
201
+ acceptance.
202
+
203
+ If every available offline check passes, the result is still `incomplete` and
204
+ `evaluate` exits `1`. JUnit marks the missing live coverage as skipped. Evidence
205
+ of a failure or error can raise the result's severity; malformed input or an
206
+ execution exception exits `2`. A successful offline invocation therefore does
207
+ not establish the `passed` verdict required for exit `0`.
208
+
209
+ Local evidence is verified by URI, path, size, and SHA-256. Retained evidence
210
+ also requires a receipt, its typed external-verification record, and the
211
+ referenced statement, trust policy, and verification evidence:
212
+
213
+ ```bash
214
+ robotics-acceptance evaluate \
215
+ --scenario scenario.yaml --runtime runtime-manifest.json \
216
+ --run-id "$RUN_ID" --domain-id primary --run-context acceptance-run.json \
217
+ --evidence-index evidence-index.json \
218
+ --artifact-receipt artifact-receipt.json \
219
+ --artifact-verification artifact-verification.json \
220
+ --receipt-dependency statement.json \
221
+ --receipt-dependency trust-policy.json \
222
+ --receipt-dependency verification.bundle \
223
+ --otel-metrics metrics.otlp.jsonl \
224
+ --window-start-ns 1786000000000000000 \
225
+ --window-end-ns 1786000030000000000 \
226
+ --output results
227
+ ```
228
+
229
+ The external verification binds the full artifact descriptor: URI, immutable
230
+ revision, media type, size, and SHA-256. The harness needs no storage
231
+ credentials. Upload, signing, and retention lifecycle remain infrastructure
232
+ responsibilities. OTLP file-exporter streams use `application/x-ndjson`.
233
+
234
+ For recordings whose receipts appear after observation starts, pass
235
+ `--receipt-inventory /evidence/receipt-inventory.json` to `verify`, `evaluate`,
236
+ or `timing-check`. Use `--receipt-inventory DOMAIN=PATH` with
237
+ `transport-evaluate`. The inventory contains exactly these three file lists:
238
+
239
+ ```json
240
+ {
241
+ "receipts": ["receipts/recording-0.json"],
242
+ "verifications": ["provenance/recording-0/artifact-verification.json"],
243
+ "dependencies": [
244
+ "provenance/recording-0/statement.json",
245
+ "provenance/recording-0/trust-policy.pem",
246
+ "provenance/recording-0/verification-evidence.sigstore.json"
247
+ ]
248
+ }
249
+ ```
250
+
251
+ Publish the inventory atomically before publishing the finalized evidence index.
252
+ The live observer reads it during its existing evidence wait, after measurement;
253
+ it need not exist when the observer starts. Missing or invalid files retain the
254
+ same evidence timeout and diagnostic behavior. Paths use canonical relative
255
+ POSIX notation below the inventory's directory. Absolute paths, traversal,
256
+ directory links escaping that root, duplicate paths and unused provenance are
257
+ rejected. The inventory uses the contract parser's document size limit and a
258
+ maximum of 4,096 files. A shared dependency occurs once in the list. Every
259
+ referenced receipt, verification and dependency still passes the same role and
260
+ byte-digest checks. Inventory and explicit receipt inputs cannot be mixed for
261
+ the same domain.
262
+
263
+ Python callers can pass `ReceiptInventory(path)` as `receipt_paths` to
264
+ `load_evidence_index`, or as `artifact_receipt_paths` to `run_verification`
265
+ and `evaluate_from_evidence`. Existing sequences of explicit file paths remain
266
+ supported. The inventory is a CLI input list, not a new contract document.
267
+
268
+ ## Histogram windows
269
+
270
+ Explicit-bucket histogram counts and recorded sums are aggregated over their
271
+ actual contribution intervals. Cumulative evidence needs a baseline at the
272
+ window boundary. An earlier baseline is usable only when an unchanged point
273
+ after the boundary proves that the intervening interval contained no events.
274
+ Otherwise event timestamps cannot be recovered and the window is unevaluated.
275
+ A changed start timestamp delimits a reset; a decreasing count without a new
276
+ start timestamp leaves the reset boundary unknown.
277
+ A cumulative point whose start equals its observation timestamp is an
278
+ [unknown-start marker](https://opentelemetry.io/docs/specs/otel/metrics/data-model/#cumulative-streams-handling-unknown-start-time).
279
+ Its existing population is subtracted before counting subsequent window events.
280
+
281
+ After baseline subtraction, lifetime minima and maxima are discarded unless
282
+ the baseline was empty. Quantiles use the inverse empirical CDF (integer rank
283
+ `ceil(p * count)`) and report conservative bucket intervals. A threshold passes
284
+ or fails only when the entire interval proves that outcome. A straddling or
285
+ unbounded interval produces a skipped assertion and an incomplete result.
286
+ Delta histograms can also lack recorded extrema; the same bound rules apply.
287
+
288
+ Time-authority results keep the measured event count when only latency bounds
289
+ are uncertain. Their required numeric fields contain finite bound endpoints;
290
+ the `time-authority-evidence` assertion records the full intervals and outcome.
291
+ Missing statistics use explicit unevaluated markers and diagnostic placeholders,
292
+ never proof of a measured zero or a threshold breach. JUnit preserves these
293
+ skipped outcomes. Artifact digests and existing result schema fields are unchanged.
294
+
295
+ ## Realtime timing windows
296
+
297
+ Live verification evaluates the complete measurement interval and overlapping
298
+ windows of at least one second. Clock callbacks bound source-clock progress;
299
+ the recorded `real_time_factor` is a conservative lower bound. A lower bound
300
+ below the policy threshold alone does not prove a violation. If the upper bound
301
+ also lies below the threshold, the time-policy assertion fails. When the bounds
302
+ straddle the threshold or clock coverage is missing, it is skipped and the
303
+ corresponding clock fields are listed as unevaluated, producing `incomplete`
304
+ unless another observation proves a failure.
305
+
306
+ Deadline ratios are independent evidence: the greatest observed value is checked
307
+ even when clock callbacks or other deadline samples are missing. A known deadline
308
+ exceedance or a clock stall proved by recorded endpoints remains a failure.
309
+ `why` preserves the distinction between unobserved and violated clock properties.
310
+
311
+ ## Contract Inputs
312
+
313
+ | Input | Contract role |
314
+ | --- | --- |
315
+ | Scenario | `acceptance-scenario.v1` |
316
+ | Runtime facts | `runtime-manifest.v1` |
317
+ | Run context | `acceptance-run.v1` |
318
+ | Evidence and provenance | `evidence-index.v1`, `artifact-receipt.v1`, `artifact-verification.v1` |
319
+ | Model and dataset provenance | `model-artifact-manifest.v1`, `dataset-manifest.v1` |
320
+ | Physical authorization | `execution-permit.v1`, `execution-verification.v1` |
321
+ | Transport inputs | `transport-channel.v1`, `clock-relation.v1`, `causal-chain.v1` |
322
+ | Outputs | `acceptance-result.v1`, `acceptance-aggregate.v1`, `campaign-summary.v1` |
323
+
324
+ Scenario extensions are explicit and digest-pinned. Pass the same
325
+ `--extension-schema URI=PATH` mapping to every command that reads the scenario.
326
+ Extensions cannot replace common safety, timing, transport, or evidence rules.
327
+
328
+ ## Product Evaluators
329
+
330
+ Product packages register standard PyPA entry points:
331
+
332
+ ```toml
333
+ [project.entry-points."robotics_acceptance.evaluators"]
334
+ "org.example.sorting" = "sorting_acceptance:evaluate"
335
+ ```
336
+
337
+ The scenario and runtime must declare the same namespace, target, distribution,
338
+ version, wheel SHA-256, and receipt SHA-256. Before importing the target, the
339
+ harness verifies the released wheel's receipt and provenance chain. Separately,
340
+ it verifies every installed file hash declared by the environment's `RECORD`;
341
+ that installation belongs to the observed execution-subject image. PEP 610
342
+ metadata and an installed `RECORD` are not treated as proof of released wheel
343
+ identity. Unhashed bytecode and module origins outside that `RECORD` fail
344
+ closed; evaluator images should install with bytecode generation disabled.
345
+ Evaluator loading also refuses `sys.pycache_prefix` (including
346
+ `PYTHONPYCACHEPREFIX`), symlinked installed paths, and evaluator modules already
347
+ imported by an unverified loader. Start the harness in a fresh interpreter.
348
+
349
+ The harness compiles the Python source bytes checked against `RECORD` using an
350
+ explicit source loader, without reading or writing bytecode caches. This covers
351
+ the evaluator's parent packages and imports within the distribution's module
352
+ namespaces, including imports deferred until evaluation. Regular packages,
353
+ namespace packages, relative imports, and dotted entry-point attributes are
354
+ supported. Evaluator-owned modules require hashed Python source; native and
355
+ sourceless evaluator modules are rejected. Dependencies outside those namespaces
356
+ use Python's normal import machinery and remain part of the execution image's
357
+ trust boundary. This loading check is not a sandbox for malicious Python code,
358
+ and a locally editable `RECORD` does not authenticate the released wheel.
359
+
360
+ An evaluator receives an immutable `EvaluationContext` and returns
361
+ `AssertionEvaluation` objects in its own namespace. Every product assertion
362
+ must reference at least one digest from verified evidence. Duplicate assertion
363
+ IDs, undeclared packages, foreign namespaces, and unknown evidence fail closed.
364
+
365
+ `doctor --scenario scenario.yaml` checks the required evaluator metadata,
366
+ receipt chain, and installed `RECORD` without importing evaluator code.
367
+
368
+ ## Pytest Integration
369
+
370
+ ```bash
371
+ uv run pytest \
372
+ -p robotics_acceptance_harness.plugin \
373
+ --robotics-scenario scenario.yaml \
374
+ --robotics-runtime runtime-manifest.json
375
+ ```
376
+
377
+ Use `robotics_bundle` for the validated immutable bundle and
378
+ `robotics_scenario` for the scenario mapping. The plugin is not activated by
379
+ package installation and refuses physical targets.
380
+
381
+ ## Development
382
+
383
+ ```bash
384
+ uv sync --locked --all-groups
385
+ uv run pre-commit run --all-files
386
+ uv run coverage run --branch -m pytest
387
+ uv run coverage report --fail-under=80
388
+ uv build --no-sources
389
+ ```
390
+
391
+ See [compatibility](docs/compatibility.md), [architecture decisions](docs/decisions/README.md),
392
+ [supply-chain assurance](docs/supply-chain.md), and the
393
+ [REP-2004 quality declaration](QUALITY_DECLARATION.md). Security reports follow
394
+ [SECURITY.md](SECURITY.md).