specpi 0.11.2 → 0.13.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (41) hide show
  1. package/CHANGELOG.md +21 -1
  2. package/NPM_RELEASE.md +3 -1
  3. package/README.md +41 -29
  4. package/SECURITY_MODEL.md +36 -0
  5. package/THIRD_PARTY.md +18 -2
  6. package/docs/browser-testing.md +76 -0
  7. package/docs/delegation/README.md +264 -0
  8. package/docs/delegation/design-protocol.md +382 -0
  9. package/docs/delegation/design.md +525 -0
  10. package/docs/delegation/evaluation.md +307 -0
  11. package/docs/delegation/protocol.md +271 -0
  12. package/docs/delegation/research.md +216 -0
  13. package/extensions/browser/core.d.mts +64 -0
  14. package/extensions/browser/diagnostics.ts +275 -0
  15. package/extensions/browser/index.ts +349 -57
  16. package/extensions/browser/interactions.ts +118 -0
  17. package/extensions/browser/lifecycle.ts +28 -0
  18. package/extensions/command-guard/index.ts +118 -33
  19. package/extensions/delegation/core.mjs +772 -0
  20. package/extensions/delegation/errors.mjs +8 -0
  21. package/extensions/delegation/extension.mjs +475 -0
  22. package/extensions/delegation/index.ts +9 -0
  23. package/extensions/delegation/managed-files.mjs +13 -0
  24. package/extensions/delegation/native.mjs +155 -0
  25. package/extensions/delegation/presentation.mjs +315 -0
  26. package/extensions/delegation/protocol.mjs +296 -0
  27. package/extensions/delegation/provider.mjs +689 -0
  28. package/extensions/delegation/snapshot.mjs +532 -0
  29. package/extensions/delegation/worker.mjs +218 -0
  30. package/extensions/tool-wishlist/verification.mjs +1 -0
  31. package/extensions/workflow-controls/index.ts +5 -1
  32. package/package.json +17 -4
  33. package/scripts/check-package.mjs +29 -3
  34. package/scripts/check-pi-package.mjs +5 -0
  35. package/scripts/check-syntax.mjs +61 -0
  36. package/scripts/run-browser-tests.mjs +61 -0
  37. package/scripts/setup-browser-tests.mjs +38 -0
  38. package/scripts/site-browser.mjs +272 -0
  39. package/scripts/specpi.mjs +24 -0
  40. package/site/logo.svg +1 -9
  41. package/templates/AGENTS.md +1 -0
@@ -0,0 +1,307 @@
1
+ # Delegation evaluation and implementation gates
2
+
3
+ This is an experiment specification, not a report of comparative outcomes. The
4
+ [experimental implementation](README.md) uses Pi AgentSession workers under the
5
+ [bounded session protocol](protocol.md). Runtime, broker, lifecycle and synthetic-provider
6
+ fixtures require a fresh run against this SDK integration before they support a passing
7
+ contract claim. Tests do not establish better task quality, cost or latency, or parity
8
+ with every live provider. No comparative outcome experiment is reported here. The numeric criteria
9
+ below remain proposed hypotheses, not research-derived constants or achieved results.
10
+
11
+ The [archived architecture](design.md) and [target protocol](design-protocol.md) retain
12
+ stronger proof obligations. In particular, full parent request-pipeline parity,
13
+ midstream/raw-transport bounds, admission of every underlying provider attempt and
14
+ monetary admission remain unmet by the native `bounded-pi-sessions-v1` contract.
15
+
16
+ ## 1. Decide what improvement means before a run
17
+
18
+ Choose one primary objective for each cohort: quality, cost, or time. Set the quality
19
+ floor, resource envelope, allowed model routes, acceptance rubric and stop rules before
20
+ looking at outcomes. A route must not claim success by changing its objective afterward.
21
+
22
+ Acceptance is task-specific and independently adjudicated. For code, use hidden or
23
+ held-out behavior checks, appropriate regression tests, and human review of the diff.
24
+ For research, use verified factual coverage, source quality, citations, and material
25
+ omissions. For review, use confirmed defects, false positives and reviewer-induced
26
+ regressions. A worker's `complete` declaration or a host receipt is never the target metric.
27
+
28
+ Use versioned public or sanitized fixtures. Keep evaluation data separate from private
29
+ Pi sessions and user repositories unless a human explicitly provides those materials
30
+ for this purpose. No automatic collection from ordinary conversations.
31
+
32
+ ## 2. Baselines and comparisons
33
+
34
+ | Arm | Purpose |
35
+ | ------------------------------------------------------- | ------------------------------------------------------------------------------- |
36
+ | A: capable single agent with current workflow | Establish the useful baseline |
37
+ | B: the same agent with the same total additional budget | Distinguish delegation from simply buying more reasoning or another review pass |
38
+ | C: structured serial workflow, same execution context | Test workflow structure without separate identities |
39
+ | D: selective one- or two-worker delegation | Test the experimental implementation |
40
+ | E: always-delegate policy | Measure routing value and unnecessary overhead; evaluation only |
41
+ | F: approved sequential model routing | Separate model-selection benefit from child-context benefit |
42
+
43
+ Do not run all six arms on every case before establishing which question matters.
44
+ For stage 1, A/B/D on frozen reviews is sufficient. Add C/E for investigation and F
45
+ only when alternative models are authorized. Use the same exact provider/model version
46
+ and tools where the comparison calls for it. A mixed-model arm must disclose composition.
47
+
48
+ Start with the implemented `review` and `scout` purposes, not all archived profiles.
49
+ Compare frozen review against both normal parent review and an additional parent pass
50
+ under the same total envelope. Include clean changes so false positives carry a cost.
51
+ For scouts, compare parent-only and structured serial analysis before attributing gains
52
+ to fresh context or parallelism. Evaluate one versus two workers only when independent
53
+ source partitions justify that question.
54
+
55
+ Record the parent's and child's effective thinking settings and model clamping. The
56
+ child uses a fresh standard Pi ModelRuntime; parent hooks and ephemeral settings are
57
+ not inherited. Control these differences when isolating architecture, or disclose them
58
+ as part of an end-to-end workflow comparison. Same model IDs alone do not establish
59
+ equal inference configuration.
60
+
61
+ Use total task spending, not only worker spending. Include preparation, parent reasoning,
62
+ worker requests, tools, retries, synthesis and validation. When actual dollar cost is
63
+ unavailable, label the cohort as call/token/latency constrained and do not describe it
64
+ as dollar-matched. Give the parent sufficient remaining budget to consume and check a
65
+ result; a worker that spends the whole allowance before integration has not succeeded.
66
+
67
+ Match maximum resource envelopes and report actual use in both arms. Equal ceilings
68
+ do not imply equal spending. Run latency comparisons separately from quality/cost
69
+ comparisons because provider load and concurrency can change latency.
70
+
71
+ ## 3. Task strata
72
+
73
+ Use at least these separately reported classes:
74
+
75
+ - Small understood edits and short lookups, where delegation should usually be rejected.
76
+ - Coupled changes with unsettled interfaces, where one writer should retain ownership.
77
+ - Independent repository investigations, including contradictory hypotheses.
78
+ - Large source collections with verifiable coverage requirements and repeated evidence.
79
+ - Frozen reviews with naturally occurring defects, seeded defects, and clean changes.
80
+ - Difficult decisions with sufficient context, plus missing-context cases that require abstention.
81
+ - Cancellation, stale source, prompt injection, bad decomposition and provider failure cases.
82
+
83
+ A pilot of 30–50 cases is for debugging the protocol and estimating variance, not
84
+ declaring a universal routing winner. Freeze policy after the pilot. Size a separate
85
+ holdout using the minimum useful effect, expected paired disagreement and cluster
86
+ variation. A target of 200 or more holdout cases can be a planning starting point;
87
+ it is not a power guarantee. Use repeated runs where stochastic variation is material.
88
+
89
+ Partition by repository or source family when possible so close variants do not leak
90
+ between tuning and evaluation. Randomize arm order, record service conditions, define
91
+ cold/warm cache handling, and use repeated time blocks to reduce infrastructure bias.
92
+ Cluster uncertainty by task/repository as appropriate. Include all admitted trials,
93
+ timeouts and failed attempts; do not condition the headline metric on successful jobs.
94
+
95
+ ## 4. Required measurements
96
+
97
+ | Measure | Definition or interpretation |
98
+ | ---------------------- | ------------------------------------------------------------------------------------------- |
99
+ | Accepted-task rate | Accepted tasks / all attempted tasks, under the frozen rubric |
100
+ | Cost per accepted task | Sum of all task costs / accepted tasks; undefined if none pass |
101
+ | Resource use | Input/output/cache/tool categories, calls, retries, unknown accounting and model identities |
102
+ | Elapsed time | End-to-end p50/p95 plus uncertainty; separate provider wait, work and integration |
103
+ | Review precision | Confirmed material findings / all material findings presented |
104
+ | Review recall | Confirmed detected defects / adjudicated defects in the evaluated fixtures |
105
+ | Human correction | Time and actions needed to resolve findings and repair final output |
106
+ | Coverage | Required questions or requirements supported by applicable evidence |
107
+ | Context failures | Missing or distorted information that changed a conclusion or prevented acceptance |
108
+ | Coordination overhead | Preparation, repeated source work, messages, synthesis, invalidated jobs and integration |
109
+ | Recovery | Verified recoveries by fault class, including abandoned work and additional cost |
110
+ | Routing errors | Unnecessary delegation, missed useful delegation and unsupported model/capability requests |
111
+ | Policy violations | Denied capabilities attempted, actual escaped capabilities and stale result publication |
112
+
113
+ Do not count a speculative finding as a caught bug. Separate naturally occurring and
114
+ seeded defects. Do not reward verbose output, agreement among agents, smaller packets,
115
+ clean merges, model confidence, or absence of exceptions as substitutes for acceptance.
116
+
117
+ ## 5. Promotion rules
118
+
119
+ The user should choose the minimum useful change before a confirmatory run. Suggested
120
+ initial rules for discussion and pre-registration are:
121
+
122
+ | Objective | Candidate rule |
123
+ | --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
124
+ | Cost | At least 20% lower total cost per accepted task, with the lower one-sided 95% bound for the quality difference above a predeclared −2 percentage-point margin |
125
+ | Time | At least 25% lower median end-to-end latency under the same resource ceiling and quality margin, without unacceptable p95 regression |
126
+ | Quality | Positive paired improvement whose confidence interval excludes zero, under the fixed budget; no unacceptable increase in false positives or correction time |
127
+
128
+ Report confidence intervals on the primary effect, not only on each arm's mean.
129
+ Use paired task-level analyses and clustered resampling where the sampling design
130
+ requires it. A low-powered non-significant difference is not evidence of equivalence.
131
+ If the holdout cannot resolve the chosen margin, keep the route experimental rather
132
+ than moving the threshold after the result. Correct for multiple confirmatory route
133
+ comparisons or specify one primary comparison and treat the rest as exploratory.
134
+
135
+ Any actual capability escape, credential exposure, invalid acceptance across a task
136
+ revision, or concealed cost accounting blocks promotion regardless of mean quality.
137
+ Zero observed escapes in fixtures does not prove universal containment; record the
138
+ enforced interface and residual trust assumptions.
139
+
140
+ Promote one task class and model route at a time. Pin the accepted policy version and
141
+ retain a parent-only fallback. A later model/provider/harness change invalidates an
142
+ assumption and requires appropriate reevaluation, not automatic inheritance of gains.
143
+
144
+ ## 6. Targeted ablations
145
+
146
+ Run only the ablations needed to resolve a material decision:
147
+
148
+ 1. Same-context self-review versus fresh-context review with identical requirements.
149
+ 2. Same-model worker versus an authorized alternative, keeping context and tools fixed.
150
+ 3. One versus two workers on genuinely independent source partitions.
151
+ 4. Selected original context versus compressed context with original-source retrieval.
152
+ 5. Explicit incomplete results versus forced answers under missing evidence.
153
+ 6. One diagnosed retry versus blind replay on transient and semantic failures.
154
+ 7. Parent verification versus unchecked worker summaries, inside test fixtures only.
155
+ 8. Warm versus cold cache and single-provider versus allowed model switching.
156
+
157
+ This separates the value of context isolation, model capability, concurrency,
158
+ compression and validation. A gain from one must not be attributed to another.
159
+ Alternative models, compression, automatic retries and live-web research are not
160
+ implemented routes. Ablations requiring them need separate authorized fixtures or a
161
+ reviewed implementation; they are not options in the shipped model-facing tool.
162
+
163
+ ## 7. Runtime compatibility and security fixtures
164
+
165
+ The implementation includes deterministic fake-provider and broker fixtures. Run the
166
+ relevant suites and record their actual results before live inference; this checklist
167
+ is not itself a passing receipt. The supported calls/time contract requires coverage for:
168
+
169
+ - Native Pi package discovery and actual SDK AgentSession creation on recorded test
170
+ versions, including current Pi releases, with in-memory sessions and rejection of
171
+ missing capabilities and unsupported routes. New version identifiers alone must not
172
+ block activation. API presence or a CLI version check is not suite completion.
173
+ - Explicit parent model/thinking with Pi clamping; fresh standard ModelRuntime
174
+ configuration; configured global transport/thinking budgets without project settings;
175
+ rejection of runtime-only authentication, selected extension-provider overrides,
176
+ model-specific headers, startup proxy configuration and safe descriptor mismatches.
177
+ Parent hooks, ephemeral runtime settings
178
+ and session affinity are not implicitly inherited.
179
+ - No ambient child extensions, skills, AGENTS files or parent session history; the SDK
180
+ owns the agent/tool loop, with only selected-source tools exposed.
181
+ - Review/scout mode and benefit compatibility, required evidence, assigned requirement
182
+ subsets, duplicate-question rejection and parentWork required only for parallel claims.
183
+ - Normal Pi ownership of resource discovery, trust, proxy policy and authentication,
184
+ without a separate SDK host or bootstrap override.
185
+ - Native and legacy provider composition using synthetic credentials only; no secrets
186
+ returned to the extension, logged, or copied into packets.
187
+ - Errors represented as resolved terminal messages, setup failure and missing usage.
188
+ - Abort before admission, during provider setup, during inference and immediately before
189
+ a broker operation; non-cooperative provider settlement and late output rejection.
190
+ - Concurrent call-slot accounting, duplicate submissions, aggregate counters and
191
+ non-resetting follow-up limits; cancelled requests keep slots until settlement.
192
+ - Admission before every SDK model invocation, including tool continuations; provider
193
+ and session retries and automatic compaction disabled; requested output clamped to
194
+ the model maximum. Follow-up does not reset allowances.
195
+ - Pi-owned authentication preflight before model-invocation accounting, without
196
+ claiming that the inference counter bounds authentication/OAuth preparation.
197
+ - Oversized SDK-visible streaming responses, bounded reports/tool output and correct
198
+ settlement state. Hold slots through stream/result and prompt settlement; do not infer
199
+ physical remote termination or a raw-response/allocation guarantee from SDK events.
200
+ - Lost collection responses, cursor replay, idempotent follow-up and resolution,
201
+ stale result revisions, conflicting idempotency payloads and cancelled-job revival.
202
+ - Asynchronous job leases surviving normal `run` return without retaining stale
203
+ extension contexts, and revocation at the intended task/policy boundaries.
204
+ - Complete versus partial tool calls, repeated IDs, unknown tools and schema tampering.
205
+ - Path traversal, encoded paths, symlinks/reparse points, private files, source changes,
206
+ invalid line ranges and oversized input/output.
207
+ - Actual branch navigation versus ordinary leaf advancement; model, task, scope and
208
+ policy changes during queued and active work; controller and aggregate quotas
209
+ retained across reloads, session switches and off/on within the Pi process.
210
+ - Fixed canonical working root, with a Pi restart required for a new root or runtime
211
+ version rather than loading changed implementation code through `/reload`.
212
+ - Completed reports remain source-bound after the deadline; child sessions release at
213
+ the deadline and subsequent follow-up fails.
214
+ - Reused Guard instances restore exactly one state responder; invalid, declined and
215
+ stale-confirmation commands cannot revoke or downgrade policy.
216
+ - SDK errors with synthetic secret canaries never enter status, notices or tool output.
217
+ Throwing/rejecting teardown is contained across cancellation and deadline callbacks.
218
+ - Shared snapshot text survives an eligible sibling follow-up, then is destroyed;
219
+ failed jobs still expire inputs, and retired batches preserve quotas and replay keys.
220
+ - Partial usage retains known fields and per-field reporting coverage; missing fields
221
+ are distinct from reported zero, including unsuccessful calls.
222
+ - Operation-count probes cover per-event root/model checks, snapshot tool reads and
223
+ linear serialized-byte work. Overflow and source/model drift remain rejected at
224
+ protected boundaries. These probes measure implementation work, not user-visible
225
+ latency, production model quality, or dollar savings.
226
+ - Recursive source syntax checks discover new nested modules and TypeScript syntax,
227
+ including negative fixtures; they do not rely on a maintained filename list.
228
+ - No worker writes, recursion, arbitrary process execution or copied parent-tool bypass.
229
+ - Normal parent-tool-result retention through Pi versus in-memory worker state, accurate
230
+ privacy disclosure, bounded retained data, cleanup and no automatic resume.
231
+
232
+ ### Local regression measurements: September 5, 2026
233
+
234
+ The review-fix probe used identical source bytes from
235
+ `d49fe9bc227cac480fea825932795e6db57e2127:extensions/delegation/worker.mjs`
236
+ as a selected fixture, comparing the old snapshot broker with the corrected broker.
237
+ Read used `read("s1", 1, 200)`; search used `search("a", 20)`. Returned objects were equal.
238
+
239
+ | Operation | Before: bytes serialized | After: bytes serialized | Returned JSON bytes |
240
+ | -------------------------- | -----------------------: | ----------------------: | ------------------: |
241
+ | Read 183 lines | 809,549 | 8,220 | 8,220 |
242
+ | Literal search, 20 matches | 25,485 | 2,741 | 2,762 |
243
+
244
+ This counts work inside `JSON.stringify`, including intermediate prefixes, not network
245
+ traffic. The read path now encodes each line once; search encodes each match once and
246
+ adds array punctuation arithmetically. Per-tool binding checks performed zero file
247
+ reads; full freshness checks still reread and hash the source. A synthetic 1,000-delta
248
+ provider regression requires fewer than 40 full lease checks and less than 100 KB of
249
+ serialization, while every delta still receives cheap lease and bounded-data checks.
250
+ The actual Pi fixture also exercises fragmented localhost SSE and transient root lookup
251
+ recovery. These are reproducible workload checks in the snapshot/provider/native test
252
+ suites, not measured UI-latency improvements or production-provider benchmarks.
253
+
254
+ ### Stronger target gates: not met by the calls/time experiment
255
+
256
+ The following remain requirements before claiming the corresponding guarantees in
257
+ the [target protocol](design-protocol.md):
258
+
259
+ - Full parent inference-pipeline parity, including request hooks, ephemeral runtime
260
+ settings and session affinity. Explicit model/thinking and standard child configuration
261
+ do not supply that broader contract.
262
+ - Every underlying inference attempt, including transport fallback or an internal SDK
263
+ retry, receives pre-dispatch admission; opaque attempts are rejected and initial
264
+ dispatch is not double-debited.
265
+ - Raw text, reasoning, framing, compressed transport and tool arguments are bounded
266
+ before parsing or allocation, including decompression and transport buffering.
267
+ - Monetary reservations account conservatively for provider pricing, uncertain dispatch,
268
+ retries and settlement, with honest unknown-cost behavior.
269
+
270
+ The current runtime disables configurable retries, counts model invocations, enforces
271
+ logical deadlines, validates inputs before dispatch, counts observed stream deltas and
272
+ validates complete parsed responses at protected boundaries. It reports cost as unavailable. These controls do
273
+ not pass the stronger gates: adapters may buffer the whole response or make transport
274
+ attempts before SpecPi observes completion. The narrower experiment must reject policies requiring unsupported
275
+ guarantees. Its fixture tests cannot be presented as transport, invoice or process-memory
276
+ proof. The unimplemented recovery/retry policy also needs separate fault-class tests.
277
+
278
+ For the later web adapter, add connection-time public-address checks, redirects,
279
+ rebinding defenses, credential/header stripping, oversized responses, malformed text,
280
+ page instruction injection and blocked private-network destinations. Test the actual
281
+ transport boundary; a validator unit test alone does not prove connection behavior.
282
+
283
+ All installer, package and Pi integration tests use fresh temporary state and skip
284
+ unnecessary external installation. A live provider smoke test is a separate explicit
285
+ authorization with a harmless prompt and a bounded cost envelope. Fake tests cannot
286
+ prove production provider parity or network behavior.
287
+
288
+ ## 8. Shipping gate and rollback
289
+
290
+ Before release, inspect the complete diff, run focused tests and the full repository
291
+ check, exercise the packaged artifact, and obtain fresh independent review of provider,
292
+ permission, cancellation, retention and dependency changes. Update third-party and
293
+ security documentation to describe the contracts actually implemented. A host bridge
294
+ that does not exist cannot be replaced with a private-field workaround just to pass
295
+ the release gate.
296
+
297
+ Disabling the extension must stop new calls and revoke broker grants immediately while
298
+ reporting requests still settling. Uninstall must not require restoring provider
299
+ configuration or authentication files because the package never owns them. Explicit
300
+ user-exported evidence remains user data. Policy rollback returns routing to the last
301
+ validated version or parent-only operation; it never discards user changes or retries
302
+ abandoned work automatically.
303
+
304
+ An experimental implementation with passing fixtures establishes only its tested
305
+ contract. Promotion beyond experimental status requires repeatable value on a defined
306
+ task class and understandable operating boundaries. Unimplemented capabilities remain
307
+ absent, not represented as partially functioning options.
@@ -0,0 +1,271 @@
1
+ # Delegation protocol: bounded-pi-sessions-v1
2
+
3
+ This is the implemented in-process API. It has no HTTP listener, daemon, child process,
4
+ or child session store. The broader [target protocol](design-protocol.md) remains a
5
+ proposal; its stronger transport/attempt/cost gates are not supplied by this version.
6
+
7
+ The extension loads through normal `pi` package discovery and remains disabled until
8
+ the human runs `/delegate on`. Compatibility is checked through required public SDK
9
+ capabilities; there is no exact-version allowlist. Missing APIs prevent activation and
10
+ are named in the error. The runtime also verifies the created session's thinking,
11
+ tools and streaming interface. Tested versions are evidence, not an activation gate.
12
+ Workers are SDK `createAgentSession` instances using in-memory session storage and a
13
+ fresh Pi `ModelRuntime`. The parent model and thinking level are passed explicitly,
14
+ subject to Pi's clamping. Standard Pi authentication, environment and `models.json`
15
+ resolution apply. Child transport/thinking budgets come from configured global settings;
16
+ project settings are not loaded. Runtime-only authentication, selected extension-provider
17
+ overrides, model-specific headers, startup proxy configuration and safe model-descriptor
18
+ mismatches fail preflight. Parent request hooks,
19
+ ephemeral runtime settings, session affinity and ambient resources are not inherited.
20
+
21
+ Command Guard is optional. Absent and Off states permit human activation; an installed
22
+ Guard's Strict approvals and explicit locks remain enforced. Unready or duplicate
23
+ Guard responders prevent activation with a specific error. Guard state changes revoke
24
+ the current delegation generation. Snapshot tools and resource limits are enforced
25
+ by delegation itself in every mode. Status includes the observed `guard` state.
26
+
27
+ The SDK runs the conversation and selected-source tool loop. SpecPi admits each SDK
28
+ invocation before dispatch and observes the SDK stream; provider/session retries and
29
+ automatic compaction are disabled. This does not establish hard raw-transport,
30
+ hidden-provider-attempt, invoice or process-memory limits.
31
+ Pi authentication preflight precedes the model-invocation counter; its preparation
32
+ is not a model invocation or an operation bounded by that counter.
33
+
34
+ Every object is closed: unknown fields, duplicate IDs, malformed values and oversized
35
+ data are rejected. The host creates identities and receipts; workers cannot supply them.
36
+ The protocol identifier `bounded-pi-sessions-v1` and inference contract
37
+ `pi-agent-session-v1` describe the host implementation, not model-selected options.
38
+
39
+ ## Submit a batch
40
+
41
+ After the human runs `/delegate on`, the parent calls the `delegate` tool:
42
+
43
+ ```json
44
+ {
45
+ "operation": "run",
46
+ "requestId": "scout-routing-1",
47
+ "packet": {
48
+ "objective": "Explain how model selection reaches the request pipeline",
49
+ "requirements": [{ "id": "R1", "text": "Identify the route binding and its invalidation behavior" }],
50
+ "decisions": ["The parent remains the sole writer"],
51
+ "nonGoals": ["Do not implement or change providers"],
52
+ "reason": {
53
+ "benefit": "parallel_analysis",
54
+ "why": "Source analysis can proceed independently of the parent's lifecycle-test inspection",
55
+ "parentWork": "Inspect the lifecycle tests while the worker reads"
56
+ },
57
+ "jobs": [
58
+ {
59
+ "id": "route",
60
+ "mode": "scout",
61
+ "requirements": ["R1"],
62
+ "question": "Where is the model captured, and when is that capability revoked?",
63
+ "context": "Return evidence for R1, including missing or contrary evidence.",
64
+ "sources": ["extensions/delegation/provider.mjs"]
65
+ }
66
+ ]
67
+ }
68
+ }
69
+ ```
70
+
71
+ `run` returns immediately with `batchId`, `packetDigest`, `generation`, a collection
72
+ cursor and job states. It does not return invented findings while work is pending.
73
+ The digest binds the packet, source descriptors, host identity, selected model/provider
74
+ IDs and resource policy. It does not freeze the registry's endpoint, headers or
75
+ authentication configuration. Those can change behind the same public identities;
76
+ the extension cannot observe every such change or certify an unchanged provider route.
77
+ IDs are short alphanumeric/hyphen/underscore strings. The host batch and attempt IDs
78
+ are UUIDs. Source paths are exact relative filenames, not directories, globs or commands.
79
+ Modes are `review` and `scout`. Either can use the selected-source list/read/literal-search
80
+ tools when files are supplied. A scout requires at least one selected file. A review
81
+ requires nonempty inline context or selected files; with an empty `sources` array it
82
+ uses inline context without tools. Source selection never grants ambient filesystem
83
+ or web access. Questions that are identical after trimming and case normalization are
84
+ rejected; distinct text does not establish distinct reasoning work.
85
+
86
+ `reason` contains exactly `benefit`, `why` and `parentWork`. `why` is nonempty.
87
+ `independent_review` requires review jobs; `parallel_analysis` and `context_isolation`
88
+ require scout jobs. `parentWork` must describe useful concurrent work for
89
+ `parallel_analysis` and may be empty otherwise. These are structural admission checks,
90
+ not proof that delegation improves the task.
91
+
92
+ Each job's nonempty `requirements` list names unique IDs from the packet's global
93
+ requirements. The child receives only its assigned requirements, plus the global
94
+ decisions and non-goals. A batch contains at most two jobs. All modes require the same
95
+ explicit packet fields, even when an allowed array or string is empty.
96
+
97
+ ## Inspect and collect
98
+
99
+ Worker `list_sources` accepts an optional zero-based `offset` and returns
100
+ `{ "sources": [...], "nextOffset": number | null }`. Each page is at most 16 KiB,
101
+ including JSON metadata, and fits the remaining tool-byte allowance. A non-null
102
+ `nextOffset` identifies the next page; null marks the end. Pages consume the same
103
+ per-job tool-call and byte budgets as reads and searches.
104
+
105
+ ```json
106
+ { "operation": "status" }
107
+ ```
108
+
109
+ ```json
110
+ { "operation": "collect", "batchId": "HOST_BATCH_ID", "afterCursor": 0, "waitMs": 30000 }
111
+ ```
112
+
113
+ `status` is compact and available while disabled. `collect` returns reports whose
114
+ completion cursor is newer than `afterCursor`; omitting it replays retained reports.
115
+ An accepted next batch retires the previous batch's reports. Old generations retire
116
+ after their workers settle. `status` retains bounded summaries marked `retired: true`,
117
+ but retired batches reject collect and follow-up. Quotas and request fingerprints survive
118
+ retirement, so replay cannot reopen a batch or replenish the process allowance.
119
+ `waitMs` defaults to zero and cannot exceed 30 seconds. There is no destructive dequeue
120
+ and no automatic parent turn. A cursor ahead of the batch is invalid.
121
+
122
+ Collection rechecks host/task/policy generation and every selected source binding.
123
+ A changed file is stale even when its pathname is unchanged. Source or lifecycle changes
124
+ cannot be repaired by presenting an old receipt or idempotency key.
125
+
126
+ Each returned item has `receipt`, `result`, `error`, and `disposition`. A host receipt
127
+ contains `batchId`, `jobId`, `attemptId`, `packetDigest`, `generation`, `resultRevision`,
128
+ `model`, `state`, `settling`, call/tool counters, token usage, `usageReportedCalls`, `usageComplete`, and
129
+ `cost: null`. Copy the six binding fields when following up or resolving. The
130
+ `collectionCursor` belongs to delivery ordering, not result identity.
131
+ Each usage field sums its valid reported values independently; never-reported fields
132
+ are `null`. `usageReportedCalls` counts reports for each field. A positive subtotal may
133
+ still be incomplete; `usageComplete` requires all four fields for every admitted call.
134
+
135
+ ## Worker report
136
+
137
+ Workers return this exact shape, without Markdown fences:
138
+
139
+ ```json
140
+ {
141
+ "status": "partial",
142
+ "answer": "The selected context is insufficient to establish runtime behavior.",
143
+ "requirements": [{ "id": "R1", "status": "unaddressed", "evidence": [] }],
144
+ "findings": [],
145
+ "missing": ["A provider implementation or runtime fixture"],
146
+ "nextStep": "The parent should inspect the provider fixture before making a claim."
147
+ }
148
+ ```
149
+
150
+ Statuses are `complete`, `partial`, or `needs_context`. Every assigned requirement appears
151
+ exactly once as `addressed` or `unaddressed`. Each finding has `id`, `claim`, `confidence`
152
+ (`observed`, `inferred`, `unverified`), `evidence` and `contraryEvidence`. Evidence entries
153
+ contain exactly `sourceId`, `lineStart`, `lineEnd`. References must resolve within the
154
+ job selection and valid line ranges. `p1` denotes the submitted inline context; it
155
+ does not denote the parent transcript or prove repository behavior. `observed` requires
156
+ at least one reference. At most eight findings and 16 KiB are retained.
157
+
158
+ Malformed reports fail the attempt. There is no automatic retry. Provider exceptions
159
+ are reduced to a generic failure message because raw errors may contain sensitive URLs
160
+ or content. Missing usage is explicit; failed calls still consume invocation allowance.
161
+
162
+ ## Follow up and resolve
163
+
164
+ ```json
165
+ {
166
+ "operation": "follow_up",
167
+ "requestId": "route-correction-1",
168
+ "batchId": "HOST_BATCH_ID",
169
+ "jobId": "route",
170
+ "attemptId": "HOST_ATTEMPT_ID",
171
+ "packetDigest": "HOST_64_CHARACTER_SHA256",
172
+ "generation": 2,
173
+ "resultRevision": 1,
174
+ "prompt": "Reconsider the claim using the already-selected source; identify the missing condition."
175
+ }
176
+ ```
177
+
178
+ The uppercase placeholders above must be replaced with the current host receipt.
179
+ One follow-up creates a new attempt under the original job deadline, context,
180
+ selected sources, model-call counters and tool counters. It cannot add files, change
181
+ models, extend time or reset quotas. If it needs a new grant or source set, discard
182
+ the report and submit a fresh batch. A settling, cancelled, expired, stale or finally
183
+ disposed result cannot receive a follow-up.
184
+
185
+ ```json
186
+ {
187
+ "operation": "resolve",
188
+ "requestId": "route-assessment-1",
189
+ "batchId": "HOST_BATCH_ID",
190
+ "jobId": "route",
191
+ "attemptId": "HOST_ATTEMPT_ID",
192
+ "packetDigest": "HOST_64_CHARACTER_SHA256",
193
+ "generation": 2,
194
+ "resultRevision": 1,
195
+ "decision": "needs_check",
196
+ "findings": []
197
+ }
198
+ ```
199
+
200
+ Overall decisions are `accept`, `discard`, `needs_check`. Each returned finding must
201
+ appear exactly once in `findings` with `{ "id": "F1", "decision": "confirmed" }`,
202
+ `rejected`, or `needs_check`. Acceptance requires a result and no unchecked findings.
203
+ Accept/discard finalize the parent disposition; needs_check leaves it open for a later
204
+ assessment or allowed follow-up. These are parent assertions, never human permission,
205
+ automatic tool execution, verified completion, or wishlist authorization.
206
+
207
+ All mutations require `requestId`. Successful runs, follow-ups and final dispositions
208
+ retain their replay receipts for the process lifetime; the fixed batch/job ceilings
209
+ bound these to at most 20 entries. Replaying a retained request returns its stored
210
+ response without another request or transition. Reusing a retained ID with a different
211
+ payload fails. Failed requests do not reserve IDs and may be corrected or retried.
212
+ Successful cancellation and `needs_check` responses use a separate 128-entry oldest-first
213
+ cache. After eviction, those operations are revalidated against current state; they
214
+ cannot start inference or restore cancelled jobs. Generation and source checks still
215
+ apply. Neither cache eviction nor failed attempts reset quotas or block cancellation.
216
+
217
+ ## Cancellation and lifecycle
218
+
219
+ ```json
220
+ { "operation": "cancel", "requestId": "cancel-route-1", "batchId": "HOST_BATCH_ID", "jobId": "route" }
221
+ ```
222
+
223
+ Omit `jobId` to cancel the batch. Cancellation remains available after invalidation.
224
+ Job states are `queued`, `running`, the three report statuses, `failed`, `cancelled`,
225
+ `expired`, and `stale`. Terminal job state and `settling` are separate. Cancellation
226
+ revokes tool access and requests SDK abort. A worker keeps its global slot through
227
+ SDK-visible stream/result and prompt settlement. Old payloads are discarded. Neither
228
+ terminal receipt delivery nor SDK settlement proves physical remote execution has ended.
229
+
230
+ The same in-memory controller survives `/reload` and session switches within the Pi
231
+ process. Its ceilings are two active workers, four batches and 32 SDK model invocations
232
+ per process, with two jobs and 8 invocations per batch and four invocations per logical
233
+ job including one follow-up. Requested output is 8,192 tokens clamped to the model
234
+ maximum. These are experiment limits, not research-derived optimal values.
235
+ Human off/on, task changes,
236
+ branch navigation, model selection, guard changes and reloads revoke old generations;
237
+ they do not create a new resource allowance. Normal parent turns do not revoke a job.
238
+ Model and thinking selections retain the human's activation choice. The extension
239
+ preflights the latest selected host and resumes dispatch automatically, without
240
+ replaying old jobs. Unsupported selections pause dispatch and report `pauseReason`;
241
+ a compatible selection resumes it. Status exposes `requested`, `updating` and
242
+ `pauseReason` alongside the controller's `enabled` flag. Concurrent notifications
243
+ for the same host share setup, and a late setup cannot overwrite a newer selection,
244
+ explicit off command, Guard revocation or session change. A turn-start refresh also
245
+ catches a changed host before the next parent turn.
246
+ The canonical working root remains fixed for that process. Restart Pi to change the
247
+ root or load a new delegation runtime version. There is no retry on process restart
248
+ and no durable worker queue. Completed reports retain their original source bindings
249
+ after the deadline; child sessions are released at the deadline and subsequent
250
+ follow-up is rejected. Snapshot text is destroyed when every job loses continuation
251
+ eligibility. Failed first attempts retain the original expiry timer. After settlement,
252
+ packet and job-input references are dropped; only source metadata/digests remain to
253
+ validate completed reports until retirement.
254
+
255
+ Per-event lease checks use model/context identities and generation. Full canonical-root,
256
+ safe model-descriptor and provider-policy checks run at request, tool and publication
257
+ boundaries. Snapshot tools check canonical/stat bindings; full digest checks run at
258
+ capture, publication, collect, follow-up and resolve. These cheaper checks assume a
259
+ trusted local filesystem and Pi's parsed stream contract. Streaming uses incremental
260
+ recognized-delta accounting and bounded structure checks, with exact response checks
261
+ at content/terminal boundaries. It does not promise an exact per-event size bound for
262
+ inconsistent SDK partial objects. Bytes may already be allocated before an event.
263
+ Per invocation the parser permits 64 content blocks, 512 structural nodes per partial,
264
+ 65,536 events and 130 non-delta boundaries. These fixed engineering ceilings also bound
265
+ repeated whole-message validation for malformed event sequences.
266
+ Unexpected SDK errors are redacted before status/tool/UI output; teardown failures
267
+ cannot escape timer callbacks or reset settling ownership.
268
+
269
+ The enforced [resource envelope and limitations](README.md#enforced-resource-envelope)
270
+ define `bounded-pi-sessions-v1`. Unsupported hard billing, raw transport, provider
271
+ attempt and process-memory policies are not accepted through this API.