pollard-jev 0.1__tar.gz

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (39) hide show
  1. pollard_jev-0.1/AGENTS.md +20 -0
  2. pollard_jev-0.1/CHANGELOG.md +18 -0
  3. pollard_jev-0.1/LICENSE +21 -0
  4. pollard_jev-0.1/MANIFEST.in +14 -0
  5. pollard_jev-0.1/PKG-INFO +258 -0
  6. pollard_jev-0.1/README.md +223 -0
  7. pollard_jev-0.1/docs/IMPLEMENTATION.md +34 -0
  8. pollard_jev-0.1/docs/POLLARD_API.md +96 -0
  9. pollard_jev-0.1/docs/RELEASING.md +79 -0
  10. pollard_jev-0.1/docs/openjev.md +99 -0
  11. pollard_jev-0.1/examples/offline_loop.py +16 -0
  12. pollard_jev-0.1/pyproject.toml +48 -0
  13. pollard_jev-0.1/scripts/release_version.py +145 -0
  14. pollard_jev-0.1/setup.cfg +4 -0
  15. pollard_jev-0.1/src/pollard_jev/__init__.py +3 -0
  16. pollard_jev-0.1/src/pollard_jev/__main__.py +3 -0
  17. pollard_jev-0.1/src/pollard_jev/contracts.py +223 -0
  18. pollard_jev-0.1/src/pollard_jev/demo.py +101 -0
  19. pollard_jev-0.1/src/pollard_jev/loop.py +283 -0
  20. pollard_jev-0.1/src/pollard_jev/policy.py +60 -0
  21. pollard_jev-0.1/src/pollard_jev/providers/__init__.py +1 -0
  22. pollard_jev-0.1/src/pollard_jev/providers/base.py +15 -0
  23. pollard_jev-0.1/src/pollard_jev/providers/fixture.py +89 -0
  24. pollard_jev-0.1/src/pollard_jev/providers/openjev.py +161 -0
  25. pollard_jev-0.1/src/pollard_jev/py.typed +0 -0
  26. pollard_jev-0.1/src/pollard_jev/records.py +61 -0
  27. pollard_jev-0.1/src/pollard_jev/simulator.py +199 -0
  28. pollard_jev-0.1/src/pollard_jev.egg-info/PKG-INFO +258 -0
  29. pollard_jev-0.1/src/pollard_jev.egg-info/SOURCES.txt +37 -0
  30. pollard_jev-0.1/src/pollard_jev.egg-info/dependency_links.txt +1 -0
  31. pollard_jev-0.1/src/pollard_jev.egg-info/entry_points.txt +2 -0
  32. pollard_jev-0.1/src/pollard_jev.egg-info/requires.txt +13 -0
  33. pollard_jev-0.1/src/pollard_jev.egg-info/top_level.txt +1 -0
  34. pollard_jev-0.1/tests/conftest.py +15 -0
  35. pollard_jev-0.1/tests/test_loop.py +258 -0
  36. pollard_jev-0.1/tests/test_openjev.py +185 -0
  37. pollard_jev-0.1/tests/test_release_version.py +117 -0
  38. pollard_jev-0.1/tests/test_review_regressions.py +147 -0
  39. pollard_jev-0.1/tests/test_simulator.py +143 -0
@@ -0,0 +1,20 @@
1
+ # Release policy
2
+
3
+ The user explicitly requires these rules for this repository:
4
+
5
+ - Start the public releases at `0.1`.
6
+ - Increment a minor release by decimal `0.01` and a major release by decimal
7
+ `0.10`. These terms describe the user's decimal policy, not SemVer.
8
+ - Keep the initial version exactly `0.1`; render subsequent versions with two
9
+ decimal digits, such as `0.11`, `0.19`, `0.20`, and `0.30`. This preserves
10
+ increasing PyPI/PEP 440 order. Do not shorten `0.20` to `0.2`.
11
+ - **Do not create, tag, or publish version `1.0` or above unless the user
12
+ explicitly approves that change in a future instruction.** Do not infer
13
+ approval from a generic request to release or make a major update.
14
+ - Use `python scripts/release_version.py check` before building, and
15
+ `python scripts/release_version.py check --tag vVERSION` before publishing.
16
+ Use the script's `bump minor` or `bump major` commands for version changes.
17
+ - Keep `pyproject.toml` and `src/pollard_jev/__init__.py` versions identical.
18
+
19
+ See [docs/RELEASING.md](docs/RELEASING.md) for the release workflow. Never commit
20
+ credentials or copy publishing tokens into documentation, issues, or logs.
@@ -0,0 +1,18 @@
1
+ # Changelog
2
+
3
+ ## 0.1 — 2026-09-23
4
+
5
+ Initial public release of the offline robotics/IoT decision companion.
6
+
7
+ - Validated observations, permitted choices, provider results, policies, and outcomes.
8
+ - Pollard 1.6.0 integration for request budgets, registered actions, bounds, and audit records.
9
+ - Dispatch-time freshness, timeout and cancellation checks; late results cannot execute actions.
10
+ - Synthetic fixture provider, bounded robot simulator, and state-machine baseline.
11
+ - Six offline demonstration scenarios, JSONL/SQLite records, read-only history,
12
+ policy simulation, and explicit reevaluation without dispatch.
13
+ - Optional OpenJev NLI adapter with independent score semantics and a pinned local-cache loader.
14
+ - Release tooling for decimal minor (+0.01) and major (+0.10) increments, with
15
+ versions at or above 1.0 prohibited until the project owner explicitly approves.
16
+
17
+ Model inference and physical hardware operation are outside this release's
18
+ verification. Optional backend tests validate mapping with test doubles.
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Muntaser Syed
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,14 @@
1
+ include README.md LICENSE CHANGELOG.md AGENTS.md
2
+ include docs/POLLARD_API.md docs/openjev.md docs/IMPLEMENTATION.md docs/RELEASING.md
3
+ recursive-include src/pollard_jev *.py py.typed
4
+ recursive-include tests *.py
5
+ recursive-include examples *.py
6
+ recursive-include scripts *.py
7
+ exclude POLLARD_JEV_START.md IMPLEMENTATION_PLAN.md
8
+ exclude docs/STATUS.md
9
+ prune artifacts
10
+ prune .venv
11
+ prune .git
12
+ prune build
13
+ prune dist
14
+ global-exclude __pycache__ *.py[cod] .env .env.* .pypirc *.pem *.key
@@ -0,0 +1,258 @@
1
+ Metadata-Version: 2.4
2
+ Name: pollard-jev
3
+ Version: 0.1
4
+ Summary: Typed, governed offline decisions for a robotics supervisor
5
+ Author: Muntaser Syed
6
+ License-Expression: MIT
7
+ Project-URL: Homepage, https://github.com/jemsbhai/pollard-jev
8
+ Project-URL: Repository, https://github.com/jemsbhai/pollard-jev
9
+ Project-URL: Documentation, https://github.com/jemsbhai/pollard-jev#readme
10
+ Project-URL: Issues, https://github.com/jemsbhai/pollard-jev/issues
11
+ Project-URL: Changelog, https://github.com/jemsbhai/pollard-jev/blob/main/CHANGELOG.md
12
+ Keywords: robotics,iot,decisions,pollard,governance,simulation
13
+ Classifier: Development Status :: 3 - Alpha
14
+ Classifier: Intended Audience :: Developers
15
+ Classifier: Programming Language :: Python :: 3
16
+ Classifier: Programming Language :: Python :: 3.11
17
+ Classifier: Programming Language :: Python :: 3.12
18
+ Classifier: Programming Language :: Python :: 3.13
19
+ Classifier: Operating System :: OS Independent
20
+ Requires-Python: >=3.11
21
+ Description-Content-Type: text/markdown
22
+ License-File: LICENSE
23
+ Requires-Dist: pollard==1.6.0
24
+ Requires-Dist: pydantic<3,>=2.13.5
25
+ Provides-Extra: test
26
+ Requires-Dist: pytest<10,>=8; extra == "test"
27
+ Requires-Dist: packaging>=24; extra == "test"
28
+ Provides-Extra: openjev
29
+ Requires-Dist: torch>=2.6; extra == "openjev"
30
+ Requires-Dist: transformers<6,>=5.15; extra == "openjev"
31
+ Requires-Dist: numpy>=1.26; extra == "openjev"
32
+ Requires-Dist: pillow>=10; extra == "openjev"
33
+ Requires-Dist: huggingface-hub>=0.36; extra == "openjev"
34
+ Dynamic: license-file
35
+
36
+ # pollard-jev
37
+
38
+ A Python companion to Pollard that turns timestamped observations and permitted
39
+ choices into a governed decision, a bounded simulated action, and an outcome
40
+ record. The first milestone runs entirely offline after installation. It needs
41
+ no credentials, model weights, or GPU.
42
+
43
+ ```powershell
44
+ python -m pip install pollard-jev
45
+ pollard-jev demo --output artifacts
46
+ ```
47
+
48
+ Published under the MIT license. The initial release is `0.1`; minor increments
49
+ are `+0.01`, major increments are `+0.10`, and releases at or above `1.0` require
50
+ the owner's explicit approval. See the
51
+ [release policy](https://github.com/jemsbhai/pollard-jev/blob/main/docs/RELEASING.md).
52
+
53
+ ## Run on Windows PowerShell
54
+
55
+ Requires Python 3.11 or newer. Tested here on Python 3.12.2.
56
+
57
+ ```powershell
58
+ git clone https://github.com/jemsbhai/pollard-jev.git
59
+ Set-Location pollard-jev
60
+ py -3 -m venv .venv
61
+ .\.venv\Scripts\python.exe -m pip install -e ".[test]"
62
+ .\.venv\Scripts\python.exe -m pollard_jev demo --output artifacts
63
+ .\.venv\Scripts\python.exe -m pytest -q
64
+ ```
65
+
66
+ Setup uses PyPI; the demo and tests make no network requests. The equivalent
67
+ installed entry point is `.\.venv\Scripts\pollard-jev.exe demo --output artifacts`.
68
+ Each demo writes uniquely named JSONL records and a Pollard SQLite ledger.
69
+
70
+ | Scenario | Synthetic proposal | Final policy | Simulated action |
71
+ | --- | --- | --- | --- |
72
+ | Fresh coherent observations | `continue` | accept | move 20 cm |
73
+ | Conflicting observations | `inspect` | need evidence | none |
74
+ | Missing evidence | `inspect` | need evidence | none |
75
+ | Evidence expires during inference | `continue` | defer | none |
76
+ | Request budget exhausted | none | defer | none |
77
+ | Provider failure | none | defer | none |
78
+
79
+ The stale scenario advances a virtual clock during inference so it is
80
+ reproducible. Separate tests exercise real waiting, timeout, and cancellation.
81
+ The state-machine baseline reads the same request at the final evaluation time.
82
+ It has no provider and therefore no provider failures or inference-request cost;
83
+ the table is a behavior comparison, not a performance benchmark.
84
+
85
+ ## Typed inputs and outputs
86
+
87
+ ```python
88
+ from datetime import datetime, timedelta, timezone
89
+ from pollard_jev.contracts import Observation
90
+ from pollard_jev.demo import example_request
91
+ from pollard_jev.loop import DecisionLoop
92
+ from pollard_jev.providers.fixture import FixtureProvider
93
+
94
+ now = datetime.now(timezone.utc)
95
+ range_reading = Observation(
96
+ observation_id="range-42", source="front-range-sensor",
97
+ feature="front_range_m", value=1.5, unit="m",
98
+ observed_at=now, valid_until=now + timedelta(seconds=5),
99
+ )
100
+ # The demo request supplies all four required sensor features and four choices.
101
+ request = example_request(now, request_id="robot-42")
102
+ request = request.model_copy(update={
103
+ "observations": (range_reading, *request.observations[1:]),
104
+ })
105
+ with DecisionLoop(FixtureProvider(), max_requests=2,
106
+ records_path="artifacts/decisions.jsonl",
107
+ pollard_path="artifacts/decisions.db") as loop:
108
+ record = loop.decide(request)
109
+ print(record.model_dump_json(indent=2))
110
+ ```
111
+
112
+ Run the complete example with
113
+ `.\.venv\Scripts\python.exe examples\offline_loop.py`.
114
+ Unknown readings use `status="unknown", value=None`; absence of a required
115
+ feature is also explicit insufficient evidence. Datetimes must have timezones,
116
+ validity windows must be positive, and numeric values must be finite.
117
+
118
+ A successful record contains fields like these (excerpt):
119
+
120
+ ```json
121
+ {
122
+ "schema_version": "1",
123
+ "provider_status": "ok",
124
+ "provider_result": {
125
+ "semantics": "synthetic_support",
126
+ "scores": {"continue": 0.94, "inspect": 0.12, "recover": 0.12, "request_help": 0.12},
127
+ "proposed_action": "continue",
128
+ "parameters": {"distance_cm": 20},
129
+ "evidence": "sufficient"
130
+ },
131
+ "policy": {"disposition": "accept", "reason": "supported_and_feasible"},
132
+ "action": {"action": "continue", "status": "completed", "simulated": true},
133
+ "budget_limit_requests": 2,
134
+ "budget_spent_requests": 1
135
+ }
136
+ ```
137
+
138
+ Full records include observation IDs, timestamps and units; the question and
139
+ allowed choices; provider identity, version and settings; policy configuration
140
+ and version; simulated skill version and observed state; timing and budget
141
+ fields; and Pollard root/model/action node IDs. Records round-trip through
142
+ validated Pydantic models. The JSONL is a convenient export; the same complete
143
+ record is also stored as a Pollard note.
144
+
145
+ ## How the loop works
146
+
147
+ `DecisionProvider.infer(tuple[DecisionRequest, ...])` returns typed results.
148
+ The interface is general; this milestone's policy and skill set are specific to
149
+ the robot demonstration. Providers return data and receive no dispatcher.
150
+
151
+ Pollard **1.6.0** supplies the real model-call ledger, custom request budget,
152
+ registered action allowlist, integer argument bounds, refusal nodes, action
153
+ records, SQLite persistence, and integrity verification. The companion supplies
154
+ contracts, evidence/support policy, sensor validity checks, deadlines,
155
+ cancellation, simulation, and linked decision records. The older sibling
156
+ Pollard checkout was inspected but is not modified or used as a dependency.
157
+ See [the API inspection](https://github.com/jemsbhai/pollard-jev/blob/main/docs/POLLARD_API.md) for exact versions and boundaries.
158
+
159
+ The policy checks four engineered sensor features: range and camera clearance
160
+ in metres, battery percentage, and a `stuck` indicator of 0 or 1. Conflicting,
161
+ missing, duplicate, invalid, expired, and future-dated readings block action.
162
+ Support threshold 0.80, score margin 0.15, clearance 0.5 m, battery 10%, and
163
+ range disagreement 0.4 m are **demonstration settings**, configurable through
164
+ `PolicyConfig`. Required physical units cannot be relabeled to change their meaning.
165
+ The policy also checks feasibility separately from evidence support. Task
166
+ utility is represented by the caller's question and choice hypotheses; there
167
+ is no trained utility or physical success predictor here.
168
+
169
+ The bounded simulated skills are:
170
+
171
+ | Skill | Permitted integer parameter | Default |
172
+ | --- | --- | --- |
173
+ | `continue` | `distance_cm`: 1–50 | 20 |
174
+ | `inspect` | `samples`: 1–5 | 2 |
175
+ | `recover` | `distance_cm`: 1–20 | 10 |
176
+ | `request_help` | `retries`: 1 | 1 |
177
+
178
+ The request's selected parameter values must also match the provider proposal.
179
+ Acceptance never expands the caller's available actions. The registered
180
+ handler rechecks policy, freshness, cancellation, and the monotonic deadline
181
+ immediately before dispatch. A second action in the same batch requires new
182
+ observations, including when the first attempted action fails.
183
+
184
+ Inference runs in one daemon worker per loop. A timeout or cancellation discards
185
+ that worker's reply. Until it finishes, later calls return `busy`; this prevents
186
+ unbounded abandoned workers. Python cannot terminate arbitrary native model
187
+ compute. A timeout bounds response/dispatch eligibility, not GPU execution or
188
+ total process resource use. Cancellation does not undo an action already
189
+ dispatched. Actual controllers, watchdogs, immediate stopping, and motion limits
190
+ remain outside this supervisor.
191
+
192
+ The resource budget counts **logical provider batch attempts**: one
193
+ `infer(requests)` invocation costs one regardless of question count. Success,
194
+ malformed output, provider failure, timeout, and cancellation after admission
195
+ each consume one. Pre-cancelled, busy, and budget-refused calls consume zero.
196
+ Actions and audit notes consume zero model requests. This is not token, HTTP
197
+ request, measured energy, or estimated-joule accounting. An adapter may perform
198
+ multiple internal forward passes per logical batch. Budgets last for one loop
199
+ instance/run; creating a new loop creates a new budget.
200
+
201
+ Independent support/entailment scores remain independent; they are never
202
+ normalized into an action distribution. Categorical results are validated as
203
+ such only when explicitly declared. Neither score type proves physical success.
204
+ The fixture scores are synthetic and uncalibrated.
205
+
206
+ ## History, policy simulation, and live reevaluation
207
+
208
+ These are separate operations:
209
+
210
+ ```powershell
211
+ $recordFile = (Get-ChildItem artifacts\decisions-*.jsonl | Sort-Object LastWriteTime -Descending | Select-Object -First 1).FullName
212
+ .\.venv\Scripts\python.exe -m pollard_jev history $recordFile
213
+ .\.venv\Scripts\python.exe -m pollard_jev simulate-policy $recordFile
214
+ $ledgerFile = (Get-ChildItem artifacts\pollard-*.db | Sort-Object LastWriteTime -Descending | Select-Object -First 1).FullName
215
+ .\.venv\Scripts\pollard.exe verify $ledgerFile --json
216
+ ```
217
+
218
+ `inspect_history(path)` only loads recorded events. `simulate_policy(record,
219
+ config, at=...)` evaluates recorded observations and provider output with a
220
+ specified policy/time. Neither has a dispatcher or a provider. Policy
221
+ simulation does not replay the resource ledger, cancellation state, or a new
222
+ physical trajectory. A changed action does not reveal what its consequences
223
+ would have been.
224
+
225
+ `live_reevaluate(record, loop)` is an explicit new provider invocation. It
226
+ consumes the loop's request budget, writes a new record linked to the original,
227
+ preserves the original timestamps, and always disables dispatch. It uses
228
+ whatever provider the caller explicitly put in that loop; with a fixture it
229
+ remains synthetic. Old evidence normally defers at today's time. Historical
230
+ records are never overwritten.
231
+
232
+ Credential fields, arbitrary malformed responses, and exception text are not
233
+ persisted. Provider authentication must remain inside the caller-owned client.
234
+ Do not put secrets in questions, observations, model identities, or settings:
235
+ these are intentional audit content, not automatically scrubbed free text.
236
+ The JSONL export and SQLite ledger are not a cross-file transaction. A crash
237
+ may leave an intent or action without the final JSONL record; use the ledger
238
+ for inspection. This is a single-process demonstrator, not crash recovery or
239
+ exactly-once actuator infrastructure.
240
+
241
+ ## Optional real backend and next milestone
242
+
243
+ The implemented adapter targets the inspected `AlexWortega/openjev`
244
+ `qwen3.5-4b-nli-v2` helper at a pinned commit. Its NLI request/result mapping and
245
+ local-cache loader are tested with doubles. **No real model inference was run.**
246
+ The default cache contains no selected checkpoint, and the project environment
247
+ has no heavy model dependencies. See [OpenJev setup and evidence](https://github.com/jemsbhai/pollard-jev/blob/main/docs/openjev.md)
248
+ for the optional dependency/download commands, API, licenses, and limitations.
249
+
250
+ The next milestone is an explicit pinned real-model benchmark: validate the
251
+ Windows/CUDA runtime, compare 0.8B and 4B candidates against the state machine
252
+ on recorded/noisy scenarios, measure decision errors, abstention and end-to-end
253
+ latency, and calibrate support thresholds on held-out data. Add process-level
254
+ inference cancellation before device integration. Measure energy with an actual
255
+ meter when available; keep TOML and other proxies separate from joules. Native
256
+ sensor encoders, controller connections, and MCU deployment are later work.
257
+
258
+ See [implementation status](https://github.com/jemsbhai/pollard-jev/blob/main/docs/IMPLEMENTATION.md) for verified behavior and remaining limitations.
@@ -0,0 +1,223 @@
1
+ # pollard-jev
2
+
3
+ A Python companion to Pollard that turns timestamped observations and permitted
4
+ choices into a governed decision, a bounded simulated action, and an outcome
5
+ record. The first milestone runs entirely offline after installation. It needs
6
+ no credentials, model weights, or GPU.
7
+
8
+ ```powershell
9
+ python -m pip install pollard-jev
10
+ pollard-jev demo --output artifacts
11
+ ```
12
+
13
+ Published under the MIT license. The initial release is `0.1`; minor increments
14
+ are `+0.01`, major increments are `+0.10`, and releases at or above `1.0` require
15
+ the owner's explicit approval. See the
16
+ [release policy](https://github.com/jemsbhai/pollard-jev/blob/main/docs/RELEASING.md).
17
+
18
+ ## Run on Windows PowerShell
19
+
20
+ Requires Python 3.11 or newer. Tested here on Python 3.12.2.
21
+
22
+ ```powershell
23
+ git clone https://github.com/jemsbhai/pollard-jev.git
24
+ Set-Location pollard-jev
25
+ py -3 -m venv .venv
26
+ .\.venv\Scripts\python.exe -m pip install -e ".[test]"
27
+ .\.venv\Scripts\python.exe -m pollard_jev demo --output artifacts
28
+ .\.venv\Scripts\python.exe -m pytest -q
29
+ ```
30
+
31
+ Setup uses PyPI; the demo and tests make no network requests. The equivalent
32
+ installed entry point is `.\.venv\Scripts\pollard-jev.exe demo --output artifacts`.
33
+ Each demo writes uniquely named JSONL records and a Pollard SQLite ledger.
34
+
35
+ | Scenario | Synthetic proposal | Final policy | Simulated action |
36
+ | --- | --- | --- | --- |
37
+ | Fresh coherent observations | `continue` | accept | move 20 cm |
38
+ | Conflicting observations | `inspect` | need evidence | none |
39
+ | Missing evidence | `inspect` | need evidence | none |
40
+ | Evidence expires during inference | `continue` | defer | none |
41
+ | Request budget exhausted | none | defer | none |
42
+ | Provider failure | none | defer | none |
43
+
44
+ The stale scenario advances a virtual clock during inference so it is
45
+ reproducible. Separate tests exercise real waiting, timeout, and cancellation.
46
+ The state-machine baseline reads the same request at the final evaluation time.
47
+ It has no provider and therefore no provider failures or inference-request cost;
48
+ the table is a behavior comparison, not a performance benchmark.
49
+
50
+ ## Typed inputs and outputs
51
+
52
+ ```python
53
+ from datetime import datetime, timedelta, timezone
54
+ from pollard_jev.contracts import Observation
55
+ from pollard_jev.demo import example_request
56
+ from pollard_jev.loop import DecisionLoop
57
+ from pollard_jev.providers.fixture import FixtureProvider
58
+
59
+ now = datetime.now(timezone.utc)
60
+ range_reading = Observation(
61
+ observation_id="range-42", source="front-range-sensor",
62
+ feature="front_range_m", value=1.5, unit="m",
63
+ observed_at=now, valid_until=now + timedelta(seconds=5),
64
+ )
65
+ # The demo request supplies all four required sensor features and four choices.
66
+ request = example_request(now, request_id="robot-42")
67
+ request = request.model_copy(update={
68
+ "observations": (range_reading, *request.observations[1:]),
69
+ })
70
+ with DecisionLoop(FixtureProvider(), max_requests=2,
71
+ records_path="artifacts/decisions.jsonl",
72
+ pollard_path="artifacts/decisions.db") as loop:
73
+ record = loop.decide(request)
74
+ print(record.model_dump_json(indent=2))
75
+ ```
76
+
77
+ Run the complete example with
78
+ `.\.venv\Scripts\python.exe examples\offline_loop.py`.
79
+ Unknown readings use `status="unknown", value=None`; absence of a required
80
+ feature is also explicit insufficient evidence. Datetimes must have timezones,
81
+ validity windows must be positive, and numeric values must be finite.
82
+
83
+ A successful record contains fields like these (excerpt):
84
+
85
+ ```json
86
+ {
87
+ "schema_version": "1",
88
+ "provider_status": "ok",
89
+ "provider_result": {
90
+ "semantics": "synthetic_support",
91
+ "scores": {"continue": 0.94, "inspect": 0.12, "recover": 0.12, "request_help": 0.12},
92
+ "proposed_action": "continue",
93
+ "parameters": {"distance_cm": 20},
94
+ "evidence": "sufficient"
95
+ },
96
+ "policy": {"disposition": "accept", "reason": "supported_and_feasible"},
97
+ "action": {"action": "continue", "status": "completed", "simulated": true},
98
+ "budget_limit_requests": 2,
99
+ "budget_spent_requests": 1
100
+ }
101
+ ```
102
+
103
+ Full records include observation IDs, timestamps and units; the question and
104
+ allowed choices; provider identity, version and settings; policy configuration
105
+ and version; simulated skill version and observed state; timing and budget
106
+ fields; and Pollard root/model/action node IDs. Records round-trip through
107
+ validated Pydantic models. The JSONL is a convenient export; the same complete
108
+ record is also stored as a Pollard note.
109
+
110
+ ## How the loop works
111
+
112
+ `DecisionProvider.infer(tuple[DecisionRequest, ...])` returns typed results.
113
+ The interface is general; this milestone's policy and skill set are specific to
114
+ the robot demonstration. Providers return data and receive no dispatcher.
115
+
116
+ Pollard **1.6.0** supplies the real model-call ledger, custom request budget,
117
+ registered action allowlist, integer argument bounds, refusal nodes, action
118
+ records, SQLite persistence, and integrity verification. The companion supplies
119
+ contracts, evidence/support policy, sensor validity checks, deadlines,
120
+ cancellation, simulation, and linked decision records. The older sibling
121
+ Pollard checkout was inspected but is not modified or used as a dependency.
122
+ See [the API inspection](https://github.com/jemsbhai/pollard-jev/blob/main/docs/POLLARD_API.md) for exact versions and boundaries.
123
+
124
+ The policy checks four engineered sensor features: range and camera clearance
125
+ in metres, battery percentage, and a `stuck` indicator of 0 or 1. Conflicting,
126
+ missing, duplicate, invalid, expired, and future-dated readings block action.
127
+ Support threshold 0.80, score margin 0.15, clearance 0.5 m, battery 10%, and
128
+ range disagreement 0.4 m are **demonstration settings**, configurable through
129
+ `PolicyConfig`. Required physical units cannot be relabeled to change their meaning.
130
+ The policy also checks feasibility separately from evidence support. Task
131
+ utility is represented by the caller's question and choice hypotheses; there
132
+ is no trained utility or physical success predictor here.
133
+
134
+ The bounded simulated skills are:
135
+
136
+ | Skill | Permitted integer parameter | Default |
137
+ | --- | --- | --- |
138
+ | `continue` | `distance_cm`: 1–50 | 20 |
139
+ | `inspect` | `samples`: 1–5 | 2 |
140
+ | `recover` | `distance_cm`: 1–20 | 10 |
141
+ | `request_help` | `retries`: 1 | 1 |
142
+
143
+ The request's selected parameter values must also match the provider proposal.
144
+ Acceptance never expands the caller's available actions. The registered
145
+ handler rechecks policy, freshness, cancellation, and the monotonic deadline
146
+ immediately before dispatch. A second action in the same batch requires new
147
+ observations, including when the first attempted action fails.
148
+
149
+ Inference runs in one daemon worker per loop. A timeout or cancellation discards
150
+ that worker's reply. Until it finishes, later calls return `busy`; this prevents
151
+ unbounded abandoned workers. Python cannot terminate arbitrary native model
152
+ compute. A timeout bounds response/dispatch eligibility, not GPU execution or
153
+ total process resource use. Cancellation does not undo an action already
154
+ dispatched. Actual controllers, watchdogs, immediate stopping, and motion limits
155
+ remain outside this supervisor.
156
+
157
+ The resource budget counts **logical provider batch attempts**: one
158
+ `infer(requests)` invocation costs one regardless of question count. Success,
159
+ malformed output, provider failure, timeout, and cancellation after admission
160
+ each consume one. Pre-cancelled, busy, and budget-refused calls consume zero.
161
+ Actions and audit notes consume zero model requests. This is not token, HTTP
162
+ request, measured energy, or estimated-joule accounting. An adapter may perform
163
+ multiple internal forward passes per logical batch. Budgets last for one loop
164
+ instance/run; creating a new loop creates a new budget.
165
+
166
+ Independent support/entailment scores remain independent; they are never
167
+ normalized into an action distribution. Categorical results are validated as
168
+ such only when explicitly declared. Neither score type proves physical success.
169
+ The fixture scores are synthetic and uncalibrated.
170
+
171
+ ## History, policy simulation, and live reevaluation
172
+
173
+ These are separate operations:
174
+
175
+ ```powershell
176
+ $recordFile = (Get-ChildItem artifacts\decisions-*.jsonl | Sort-Object LastWriteTime -Descending | Select-Object -First 1).FullName
177
+ .\.venv\Scripts\python.exe -m pollard_jev history $recordFile
178
+ .\.venv\Scripts\python.exe -m pollard_jev simulate-policy $recordFile
179
+ $ledgerFile = (Get-ChildItem artifacts\pollard-*.db | Sort-Object LastWriteTime -Descending | Select-Object -First 1).FullName
180
+ .\.venv\Scripts\pollard.exe verify $ledgerFile --json
181
+ ```
182
+
183
+ `inspect_history(path)` only loads recorded events. `simulate_policy(record,
184
+ config, at=...)` evaluates recorded observations and provider output with a
185
+ specified policy/time. Neither has a dispatcher or a provider. Policy
186
+ simulation does not replay the resource ledger, cancellation state, or a new
187
+ physical trajectory. A changed action does not reveal what its consequences
188
+ would have been.
189
+
190
+ `live_reevaluate(record, loop)` is an explicit new provider invocation. It
191
+ consumes the loop's request budget, writes a new record linked to the original,
192
+ preserves the original timestamps, and always disables dispatch. It uses
193
+ whatever provider the caller explicitly put in that loop; with a fixture it
194
+ remains synthetic. Old evidence normally defers at today's time. Historical
195
+ records are never overwritten.
196
+
197
+ Credential fields, arbitrary malformed responses, and exception text are not
198
+ persisted. Provider authentication must remain inside the caller-owned client.
199
+ Do not put secrets in questions, observations, model identities, or settings:
200
+ these are intentional audit content, not automatically scrubbed free text.
201
+ The JSONL export and SQLite ledger are not a cross-file transaction. A crash
202
+ may leave an intent or action without the final JSONL record; use the ledger
203
+ for inspection. This is a single-process demonstrator, not crash recovery or
204
+ exactly-once actuator infrastructure.
205
+
206
+ ## Optional real backend and next milestone
207
+
208
+ The implemented adapter targets the inspected `AlexWortega/openjev`
209
+ `qwen3.5-4b-nli-v2` helper at a pinned commit. Its NLI request/result mapping and
210
+ local-cache loader are tested with doubles. **No real model inference was run.**
211
+ The default cache contains no selected checkpoint, and the project environment
212
+ has no heavy model dependencies. See [OpenJev setup and evidence](https://github.com/jemsbhai/pollard-jev/blob/main/docs/openjev.md)
213
+ for the optional dependency/download commands, API, licenses, and limitations.
214
+
215
+ The next milestone is an explicit pinned real-model benchmark: validate the
216
+ Windows/CUDA runtime, compare 0.8B and 4B candidates against the state machine
217
+ on recorded/noisy scenarios, measure decision errors, abstention and end-to-end
218
+ latency, and calibrate support thresholds on held-out data. Add process-level
219
+ inference cancellation before device integration. Measure energy with an actual
220
+ meter when available; keep TOML and other proxies separate from joules. Native
221
+ sensor encoders, controller connections, and MCU deployment are later work.
222
+
223
+ See [implementation status](https://github.com/jemsbhai/pollard-jev/blob/main/docs/IMPLEMENTATION.md) for verified behavior and remaining limitations.
@@ -0,0 +1,34 @@
1
+ # Implementation status
2
+
3
+ The first milestone is implemented as an offline Python package with a bounded
4
+ mobile-robot supervisor. Pollard 1.6.0 supplies real execution governance;
5
+ the provider scores, sensors, and robot effects in the default demo are synthetic.
6
+
7
+ The original implementation passed 88 tests on Windows/Python 3.12.2. The six
8
+ demonstration scenarios produced one accepted simulated move, two requests for
9
+ more evidence, and three deferrals. Pollard verified all 26 nodes in the
10
+ resulting ledger without findings. Editable installation, wheel installation,
11
+ the documented example, history, and policy simulation were also verified.
12
+ Initial release validation passed 120 local tests, including version-policy
13
+ checks. Cross-platform GitHub CI exercises Linux and Windows installations.
14
+
15
+ Coverage includes action allowlists, parameter bounds, malformed/non-finite
16
+ outputs, units, request-batch accounting, inference failures, expiry at dispatch,
17
+ cancellation/deadline changes during validation, discarded late replies,
18
+ bounded outstanding workers, and history without actuator side effects.
19
+
20
+ The optional OpenJev adapter has 22 mapping/cache-loader tests using test doubles.
21
+ No model weights, real model inference, physical controller operation, or energy
22
+ measurement are implied by those tests. No selected checkpoint or optional model
23
+ runtime was available during implementation.
24
+
25
+ The resource budget counts logical provider batch attempts, not tokens, GPU
26
+ forward passes, or joules. Native inference may continue after timeout; late
27
+ results cannot dispatch. Records are single-process and do not provide
28
+ cross-file transactions or exactly-once actuator recovery.
29
+
30
+ Next: explicitly prepare and benchmark pinned 0.8B/4B model candidates against
31
+ the state-machine baseline on noisy/missing/stale observations. Measure latency,
32
+ decision errors and abstention, calibrate thresholds on held-out cases, and add
33
+ process cancellation before adapting the interface to devices. Keep measured
34
+ energy separate from TOML and other proxies.