specpi 0.11.2 → 0.13.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (41) hide show
  1. package/CHANGELOG.md +21 -1
  2. package/NPM_RELEASE.md +3 -1
  3. package/README.md +41 -29
  4. package/SECURITY_MODEL.md +36 -0
  5. package/THIRD_PARTY.md +18 -2
  6. package/docs/browser-testing.md +76 -0
  7. package/docs/delegation/README.md +264 -0
  8. package/docs/delegation/design-protocol.md +382 -0
  9. package/docs/delegation/design.md +525 -0
  10. package/docs/delegation/evaluation.md +307 -0
  11. package/docs/delegation/protocol.md +271 -0
  12. package/docs/delegation/research.md +216 -0
  13. package/extensions/browser/core.d.mts +64 -0
  14. package/extensions/browser/diagnostics.ts +275 -0
  15. package/extensions/browser/index.ts +349 -57
  16. package/extensions/browser/interactions.ts +118 -0
  17. package/extensions/browser/lifecycle.ts +28 -0
  18. package/extensions/command-guard/index.ts +118 -33
  19. package/extensions/delegation/core.mjs +772 -0
  20. package/extensions/delegation/errors.mjs +8 -0
  21. package/extensions/delegation/extension.mjs +475 -0
  22. package/extensions/delegation/index.ts +9 -0
  23. package/extensions/delegation/managed-files.mjs +13 -0
  24. package/extensions/delegation/native.mjs +155 -0
  25. package/extensions/delegation/presentation.mjs +315 -0
  26. package/extensions/delegation/protocol.mjs +296 -0
  27. package/extensions/delegation/provider.mjs +689 -0
  28. package/extensions/delegation/snapshot.mjs +532 -0
  29. package/extensions/delegation/worker.mjs +218 -0
  30. package/extensions/tool-wishlist/verification.mjs +1 -0
  31. package/extensions/workflow-controls/index.ts +5 -1
  32. package/package.json +17 -4
  33. package/scripts/check-package.mjs +29 -3
  34. package/scripts/check-pi-package.mjs +5 -0
  35. package/scripts/check-syntax.mjs +61 -0
  36. package/scripts/run-browser-tests.mjs +61 -0
  37. package/scripts/setup-browser-tests.mjs +38 -0
  38. package/scripts/site-browser.mjs +272 -0
  39. package/scripts/specpi.mjs +24 -0
  40. package/site/logo.svg +1 -9
  41. package/templates/AGENTS.md +1 -0
@@ -0,0 +1,525 @@
1
+ # SpecPi delegation design: archived target proposal
2
+
3
+ Status: archived target architecture, not the implemented runtime contract.
4
+
5
+ The experimental implementation is part of `specpi` and is disabled by default.
6
+ Read the [implemented guide](README.md) and [calls/time protocol](protocol.md) for
7
+ supported commands, tested API compatibility, limits and trust assumptions. This
8
+ document preserves the original broader proposal, including unimplemented live-web,
9
+ alternative-model, monetary, raw-transport and underlying-attempt guarantees. Its
10
+ normative requirements, command sketches and delivery stages are targets, not runtime
11
+ promises or evidence that the experiment has passed outcome evaluation.
12
+
13
+ The current implementation uses actual Pi AgentSession workers with only `review` and
14
+ `scout` admission, explicit parent model/thinking, selected-source tools and in-memory
15
+ sessions. It supersedes the proposed custom loop and four-profile sequence below.
16
+
17
+ Research reviewed through 5 September 2026. The original working name
18
+ `specpi-delegation` below was a design placeholder, not a separate published or reserved
19
+ npm package.
20
+
21
+ The [target protocol](design-protocol.md) specifies this proposal's messages and state
22
+ transitions. [Research](research.md) records the evidence and its limits.
23
+ [Evaluation](evaluation.md) separates current fixture coverage from experiments and
24
+ stronger proof obligations that remain outstanding.
25
+
26
+ ## Decision
27
+
28
+ Build an optional Pi extension that gives one capable parent **bounded access to
29
+ additional reasoning and investigation**, while the parent owns changes and acceptance.
30
+ Use one scheduler, one provider adapter, one evidence broker, and one result format.
31
+ Start with a tool-free review call, then add independent investigation and research
32
+ through the same protocol. Treat consultation as another bounded call, rather than
33
+ introducing a second planner or a permanent team.
34
+
35
+ The objective is the best accepted outcome within the user's resource constraints.
36
+ There is no demonstrated universal optimum in agent count, topology, model, or context
37
+ size. The architecture therefore includes **no delegation** as a first-class result of
38
+ routing. It should make a useful two-worker investigation straightforward and make an
39
+ unnecessary ten-worker swarm difficult to create.
40
+
41
+ The evidence supports this structure more strongly than an unrestricted writing swarm.
42
+ Task-dependent gains in the [existing seven-study assessment](../../site/single-agent/index.html)
43
+ establish specific opportunities. Cognition's April 2026 production account also
44
+ describes useful review and consultation around a single writer, while reporting a
45
+ quality ceiling when the primary model was too weak to delegate effectively.
46
+ [Cognition, 22 April 2026](https://cognition.com/blog/multi-agents-working).
47
+
48
+ This design is an engineering inference. Its defaults and promotion thresholds below
49
+ are proposed operating parameters, not numbers established by those studies.
50
+
51
+ ## 1. Choose execution by task shape
52
+
53
+ | Situation | Starting route | Why another context may help | Stop condition |
54
+ | ------------------------------------------------------------------- | ---------------------------------------------------------- | ------------------------------------------------------------------------ | -------------------------------------------------------------------- |
55
+ | Small fix, known answer, one short lookup | Parent only; batch independent tools when useful | Usually insufficient benefit | Do not create a worker |
56
+ | Long sequential reasoning with all relevant context already present | Same agent, stronger workflow or supported model selection | Separate contexts can lose the reasoning chain | Keep execution serial |
57
+ | Two substantial, independent evidence questions | Two investigation or research jobs | Parallel exploration and less irrelevant material in the parent | Evidence sufficient to answer the questions |
58
+ | One large source collection | Partition by question or source ownership, then synthesize | Workers can search deeply without sending every intermediate result back | Coverage achieved, result and resource limits reached |
59
+ | Frozen implementation with meaningful review risk | One fresh reviewer | Different context and focused inspection | Findings and coverage returned; parent adjudicates |
60
+ | A specific difficult decision or failed approach | One capable consultant | A different capability or fresh interpretation may help | Answer, discriminating next experiment, or explicit missing evidence |
61
+ | Coupled implementation, evolving interfaces, shared generated files | Parent is the only writer | Concurrent writers introduce integration decisions | Finish the interface or obtain evidence before proceeding |
62
+
63
+ These are routing hypotheses. The parent can misclassify a dependency or a source
64
+ collection. The scheduler checks the declared contract; it cannot prove semantic
65
+ independence. Evaluation must count bad decomposition as a system failure.
66
+
67
+ Four modes share one engine:
68
+
69
+ - **Review:** inspect a frozen packet; report material defects and uncovered requirements.
70
+ - **Investigate:** answer a repository question using bounded snapshot reads.
71
+ - **Research:** answer an external evidence question using an explicitly enabled web adapter.
72
+ - **Consult:** examine a difficult decision using selected context and, when authorized,
73
+ a different exact model. This mode has no tools initially.
74
+
75
+ A model switch inside the parent's existing execution is a separate alternative in the
76
+ evaluation. Learned sequential routing has evidence of its own; it does not establish
77
+ the benefit of spawning children. [MTRouter](https://aclanthology.org/2026.acl-long.2045/).
78
+
79
+ ## 2. One level of delegation, one owner of decisions
80
+
81
+ ```mermaid
82
+ flowchart TD
83
+ U[Human task and resource policy] --> P[Parent: intent, decisions, writing]
84
+ P --> G{Delegation admission}
85
+ G -->|insufficient benefit or unsupported capability| S[Continue in parent]
86
+ G -->|bounded independent work| Q[Scheduler and shared budget ledger]
87
+ Q --> A[Worker A: explicit packet]
88
+ Q --> B[Worker B: explicit packet]
89
+ A <--> E[Evidence broker: granted snapshots or public sources]
90
+ B <--> E
91
+ A --> V[Validate provenance, coverage, state and usage]
92
+ B --> V
93
+ V --> P
94
+ P --> C[Actual checks and human acceptance]
95
+ ```
96
+
97
+ Workers cannot spawn workers, send sibling messages, alter policy, or write the project.
98
+ They return evidence to the parent. The parent decides whether a discovery changes the
99
+ task or another lane. A changed decision produces a new packet generation; it does not
100
+ silently replace instructions in a running job.
101
+
102
+ Initially the scheduler accepts a batch of independent jobs. Dependent work requires
103
+ the parent to accept the prerequisite and submit the next batch. This gives a useful
104
+ parallel frontier without a general workflow engine, distributed message bus, agent
105
+ registry, or arbitrary task graph interpreter.
106
+
107
+ Parallel work is admitted only when its result is needed and either the parent has
108
+ useful independent work or multiple workers can finish independent questions together.
109
+ A review may be entirely serial and still earn its cost through improved detection.
110
+ There is no requirement to keep workers busy.
111
+
112
+ ## 3. Human control without approval fatigue
113
+
114
+ Installation and activation are separate. The package starts disabled and adds no
115
+ standing model calls. Proposed command surface:
116
+
117
+ ```text
118
+ /delegate Show mode, grants, limits, running jobs and usage
119
+ /delegate on Enable bounded delegation for this session
120
+ /delegate off Stop admission and cancel outstanding jobs
121
+ /delegate limits Inspect or change the session's resource envelope
122
+ /delegate cancel <id> Cancel one job or batch
123
+ ```
124
+
125
+ Activation presents the concrete capability envelope: exact model route, review-only
126
+ or snapshot access, web access if any, concurrency, call limits, retention, and budget
127
+ accounting limitations. A human-approved envelope permits subsequent parent requests
128
+ within that envelope. It does not require another permission prompt for every worker
129
+ or file read. Human instructions already authorizing an exact envelope can satisfy this
130
+ selection. Unsupported capabilities remain unavailable.
131
+
132
+ The parent proposes a job through one model tool, `delegate`, with operations `run`,
133
+ `status`, `collect`, `follow_up`, `resolve`, and `cancel`. A `run` response returns an opaque batch ID and the
134
+ admitted job identities. Detailed worker tools and prompts are disclosed only when the
135
+ corresponding mode is used. Keep the standing description small; do not load four role
136
+ manuals into every turn.
137
+
138
+ The extension owns resolved grants, budget reservations, model identities, job IDs, and
139
+ generation tokens. The parent model cannot grant itself broader access by changing
140
+ tool arguments. A denied Command Guard operation remains denied; delegation is never
141
+ an alternate execution route for it. In Strict mode, approval of the top-level custom
142
+ tool must expose the effective delegation capabilities, rather than an opaque label.
143
+
144
+ User-facing progress is one status row, for example:
145
+
146
+ ```text
147
+ Delegation 1 running · 1 ready · 4/12 model calls · cost unavailable
148
+ ```
149
+
150
+ Only blockers, usable results, and meaningful state changes deserve notifications.
151
+ No periodic narration, hidden follow-up turns, or automatic restart after closing Pi.
152
+
153
+ ## 4. Admission and resource allocation
154
+
155
+ Admission has two layers. **The parent judges usefulness; deterministic code enforces
156
+ authority and limits.** Do not ask another model to decide whether to launch a model.
157
+
158
+ The parent supplies a short reason, deliverable, evidence needed for acceptance,
159
+ independence boundary, expected remaining work, and the best parent-only alternative.
160
+ The controller then checks:
161
+
162
+ 1. The mode, model, tools, and inputs are inside the current human-authorized policy.
163
+ 2. Original requirements and relevant settled decisions are represented in the packet.
164
+ 3. The job is ready; its prerequisites are accepted and its result has a named consumer.
165
+ 4. Its source snapshot and task generation are still current for admission.
166
+ 5. No equivalent job is running or already has an applicable result in this batch.
167
+ 6. A concurrency slot and the job's full resource reservation are available.
168
+ 7. The result has a feasible verification path. A worker's confidence is not that path.
169
+
170
+ Failure returns a reason and leaves the parent able to work. It does not launch a
171
+ smaller unrequested model, widen access, or silently weaken an exact provider rule.
172
+
173
+ Once measured task-class data exists, select routes against one declared objective:
174
+
175
+ | Objective | Route comparison |
176
+ | --------- | ------------------------------------------------------------------------------ |
177
+ | Cost | Minimize expected total cost per accepted task, with a specified quality floor |
178
+ | Time | Minimize expected elapsed time, with quality and spend constraints |
179
+ | Quality | Maximize accepted outcomes under a fixed resource envelope |
180
+
181
+ Total cost includes packet preparation, all model and tool work, retries, parent
182
+ synthesis, verification, and correction. Report human review time separately unless a
183
+ human explicitly chooses a monetary conversion. Compute cost per accepted outcome as
184
+ total cohort cost divided by accepted tasks, including the cost of failed tasks.
185
+ When none is accepted the measure is undefined, not zero.
186
+
187
+ Do not invent a numerical expected-value estimate from an LLM's self-reported
188
+ confidence. Until a route is evaluated, label it experimental and use explicit bounded
189
+ trials. Learned routing, automatic difficulty prediction, and online bandit exploration
190
+ are unnecessary for the first implementation.
191
+
192
+ ## 5. Context is a contract, not a transcript fork
193
+
194
+ Every job receives the original relevant requirements, its exact question, acceptance
195
+ criteria, non-goals, selected decisions, and reference identities. It does not inherit
196
+ the parent's tools, mutable memory, complete conversation, or unrelated instructions.
197
+ No Pi history or authentication files are scanned to construct a packet.
198
+
199
+ Use different context policies deliberately:
200
+
201
+ | Mode | Initial context | Deliberately excluded | Expansion |
202
+ | ----------- | ---------------------------------------------------------------------------- | ---------------------------------------------- | ----------------------------------------------------------------------------- |
203
+ | Review | Requirements, frozen diff, relevant source, actual check receipts | Implementation reasoning and preferred verdict | Specific missing source requested; bounded snapshot access in a later release |
204
+ | Investigate | One question, hypotheses, repository map, source manifest | Unrelated task transcript | Broker reads within the granted snapshot |
205
+ | Research | Question, source boundaries, dates, exclusions, required output | Parent's desired conclusion | Bounded search and retrieval of approved public sources |
206
+ | Consult | Decision, constraints, failed attempts, observations, competing explanations | Unrelated history and hidden model reasoning | A targeted request for additional evidence |
207
+
208
+ Fresh review still receives the human requirements. Withholding those requirements to
209
+ make a reviewer more independent can create irrelevant findings. A defect reviewer may
210
+ question a requirement but cannot redefine it. A separate code-only review is a possible
211
+ evaluation ablation, not the normal acceptance contract.
212
+
213
+ Distinguish two optimization steps. **Selection** omits irrelevant material and preserves
214
+ explicit references to what remains available. **Compression** rewrites information and
215
+ can lose crucial details. Start with selection and deterministic limits. Use model
216
+ summaries only as indexed leads, retain their source references, and measure any solve
217
+ rate loss before adopting a compression policy.
218
+
219
+ Workers return `needs_context` when missing information prevents a grounded answer.
220
+ They name the required evidence and the conclusion it would resolve. One focused
221
+ follow-up can be admitted within the original batch budget. Repeated gaps return the
222
+ task to the parent; they do not create an unlimited conversation.
223
+
224
+ Keep a stable instruction prefix and append changing evidence after it, where the
225
+ provider supports useful cache behavior. Record actual cache accounting when exposed.
226
+ Never assume a cache hit, transfer provider-specific reasoning blocks between models,
227
+ or sacrifice necessary context to satisfy a speculative token-saving target.
228
+
229
+ ## 6. The evidence broker
230
+
231
+ Workers get explicit function schemas, never a generic shell or arbitrary extension
232
+ dispatcher. The first repository tools are `list_sources`, `read_source`, and
233
+ `search_sources`, operating on opaque IDs in a bounded in-memory snapshot. A source ID
234
+ maps to captured bytes and a content hash, not a worker-supplied filesystem path.
235
+
236
+ The parent selects repository-relative inputs. The host resolves and validates them
237
+ before capture: canonical root containment, regular files, allowed types, explicit byte
238
+ limits, and exclusion of private state, credentials, binary files, links and reparse
239
+ points whose containment cannot be established. Search is a bounded literal search
240
+ initially; arbitrary regular expressions and shell-based grep are unnecessary. The
241
+ broker does not invoke project scripts, Git filters, LSP plugins, or package hooks.
242
+
243
+ Take source hashes before and after capture. This detects observed changes but is not
244
+ an atomic operating-system snapshot. Serve only the captured bytes. A manifest labels
245
+ the snapshot's coverage and capture limitations. If a required consistent source set
246
+ cannot be captured, return `source_changed` and ask the parent to establish a stable
247
+ target. Never claim to have reviewed a moving checkout as one atomic revision.
248
+
249
+ Large repositories do not justify uploading the repository. Start with a selected
250
+ bounded manifest. If the selected input exceeds its envelope, request a smaller scope
251
+ or an explicit expansion. Hidden omissions or silently shortened evidence are invalid.
252
+ This also keeps source hashes and exact line citations useful for parent inspection.
253
+
254
+ The research adapter is a separate opt-in capability with typed `search_public` and
255
+ `fetch_public` operations. It must implement a reviewed network boundary: allowed
256
+ schemes, host/IP checks at connection time and redirects, public-address restrictions,
257
+ byte/time limits, no cookies or inherited authorization, bounded downloads, and no
258
+ execution of page scripts. URL parsing alone does not prevent DNS rebinding or access
259
+ to private services. If the chosen transport cannot enforce the declared boundary,
260
+ public fetch is unavailable. Do not quietly proxy arbitrary existing browser or MCP
261
+ tools through the worker. Parent-supplied public excerpts remain a usable fallback.
262
+
263
+ Source text and worker output are untrusted evidence. Neither can create permissions,
264
+ change budgets, inject system instructions, or authorize a sibling action. Exact source
265
+ locations establish traceability; they do not establish truth.
266
+
267
+ ## 7. Resource limits that describe what is actually enforced
268
+
269
+ These are initial trial defaults, to be tuned through evaluation:
270
+
271
+ | Limit | Proposed default | Enforcement |
272
+ | -------------------------- | --------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- |
273
+ | Active workers | 2 per batch; one active batch per parent session | Scheduler semaphore |
274
+ | Session allowance | 4 batches and 48 model requests | Monotone session ledger; renewed only by a human policy change |
275
+ | Logical jobs | 4 per batch | Admission counter; no nested delegation |
276
+ | Model requests | 12 total per batch, including retries and follow-ups | Reserve before every underlying inference attempt |
277
+ | Requests per logical job | 2 for review/consult; 4 for investigate/research | Review/consult uses one initial call; the second is at most one retry or follow-up |
278
+ | Follow-up generation | At most 1 per job | Explicit parent resubmission; same batch allowance |
279
+ | Tool calls | 12 per job | Broker counter before execution |
280
+ | Job / batch deadline | 120 / 300 seconds | Monotonic deadline, cancellation, late-result rejection |
281
+ | Selected source snapshot | 8 MiB and 200 files per batch | Checked before reads; bounded allocation |
282
+ | Serialized model context | 256 KiB per request, also within model context limits | Byte cap plus provider token/context validation when supported |
283
+ | Evidence returned by tools | 64 KiB total per job | Count bytes; explicit incomplete result at limit |
284
+ | Final worker payload | 16 KiB; 8 findings | Reject oversized payloads, never silently accept a truncated object |
285
+ | Raw provider response | 256 KiB per attempt, including text, reasoning, tool arguments and metadata | Bound transport buffering and accumulation before parsing or appending content |
286
+ | Provider output limit | 2,048 tokens per request where supported | Provider request setting; record unsupported accounting |
287
+ | Retry | One classified retry per job | Counts against calls, deadline, and spend admission |
288
+
289
+ All attempts under the same logical job share its tool, evidence, elapsed-time, and
290
+ follow-up limits. A replacement does not reset the counters. Limits are configurable
291
+ inside human policy; worker tool arguments can only request less. Source and result
292
+ bytes are local memory limits, not token estimates. Hidden reasoning, provider retries,
293
+ tokenization, rate limits, and post-cancellation billing require separate accounting.
294
+ Batch completion, `/task` revision, normal user steering and model changes do not
295
+ replenish the session allowance. An on/off toggle does not erase consumed usage.
296
+ Session restart clears ephemeral state and starts disabled; no work resumes by itself.
297
+
298
+ For monetary admission, reserve the conservative cost of the next request, using
299
+ complete supported pricing and bounded billable input/output categories; do not assume
300
+ cache discounts. Atomically reserve before dispatch, reconcile with actual usage, and
301
+ keep unknown in-flight charges reserved. Every retry pays for its own reservation.
302
+ Reject concurrent requests whose combined reservations exceed the admitted envelope.
303
+
304
+ Here a counted model request means an underlying inference attempt, including an SDK
305
+ or provider-adapter retry. `InferencePort` must either disable opaque automatic retries
306
+ under the admitted child policy or obtain controller admission before every attempt.
307
+ One method call that internally retries three times consumes four attempts. Report
308
+ high-level invocations separately. An adapter that cannot expose or prevent hidden
309
+ attempts cannot satisfy a hard attempt cap and is unavailable for that policy.
310
+
311
+ The raw response cap applies while receiving every content category, before assembling
312
+ JSON, reasoning, tool arguments or a final result. The adapter must bound transport
313
+ chunks and its own buffering, not merely check the completed response size. Overflow
314
+ aborts the attempt, preserves usage uncertainty and produces an invalid/incomplete
315
+ outcome; a truncated object is never treated as a successful answer. Discard buffered
316
+ content when its job no longer needs it, while retaining bounded usage and state.
317
+
318
+ A configured spend ceiling is an **admission limit**, not a guarantee of a final invoice.
319
+ Cancellation cannot retract work already performed by a provider. If the provider cannot
320
+ supply a defensible cost bound, either decline a cost-capped route or use a separately
321
+ human-authorized calls/time-limited mode with cost explicitly marked unavailable. Never
322
+ turn missing billing or subscription usage into `$0.00`.
323
+
324
+ This extension can enforce its own requests and broker calls. It cannot cap the parent's
325
+ ordinary inference, other extensions, or a provider's unrelated account usage. Reserve
326
+ parent synthesis and verification capacity in the evaluation plan; an extension-only
327
+ spend counter must not be presented as the total task budget.
328
+
329
+ ## 8. Pi runtime and package boundary
330
+
331
+ The following records the original compatibility analysis and its stronger target.
332
+ The current native experiment instead uses public SDK `createAgentSession` with a fresh
333
+ standard Pi `ModelRuntime`, explicit parent model/thinking, in-memory sessions and
334
+ admission before each SDK invocation. Delegation checks required public SDK capabilities
335
+ instead of gating exact Pi versions; the normal installer floor and bootstrap pin are
336
+ separate contracts. See the compatibility record for tested versions and completed validation.
337
+ SDK-visible streams are checked, but full parent-hook/ephemeral-setting parity and
338
+ hard raw-transport, hidden-attempt and monetary bounds remain unclaimed.
339
+ See the [implemented guide](README.md). The target `InferencePort` below is not an
340
+ exported Pi API or a claim that the experiment supplies these stronger guarantees.
341
+
342
+ **Production is gated on a supported host request capability.** Inspection of Pi
343
+ `v0.84.4`, SpecPi's reviewed host floor, found an important distinction:
344
+
345
+ | Verified public contract | What it establishes | What it does not establish |
346
+ | ----------------------------------------------------- | --------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ |
347
+ | `ctx.modelRegistry.complete(model, context, options)` | Runtime-owned authentication and effective provider preparation | The host's full request hooks, thinking translation, settings, transport and session-affinity pipeline |
348
+ | `ctx.model`, `ctx.thinkingLevel`, `ctx.signal` | Active route metadata and parent cancellation | A reusable configured child streaming capability |
349
+ | `pi.getAllTools()` | Tool metadata and schemas | An executable dispatcher that inherits Command Guard |
350
+ | `Agent` / `agentLoop` from `pi-agent-core` | A reusable loop with explicit streaming and tool callbacks | The parent's hooks, grants or runtime automatically |
351
+ | `SessionManager.inMemory()` | No child session persistence | No configuration, authentication-service or extension discovery |
352
+
353
+ The registry facade does not expose `streamSimple`. Calling its completion method is
354
+ credential-blind, but it would skip host request layers that can carry provider policy.
355
+ Passing `reasoning: ctx.thinkingLevel` is not a portable replacement for Pi's provider
356
+ translation. Do not ship that shortcut as equivalent behavior.
357
+ [Pinned registry](https://github.com/earendil-works/pi/blob/v0.84.4/packages/coding-agent/src/core/model-registry.ts#L59),
358
+ [pinned host pipeline](https://github.com/earendil-works/pi/blob/v0.84.4/packages/coding-agent/src/core/sdk.ts#L283).
359
+
360
+ Define a host-owned `InferencePort` that accepts a bound child context and admitted
361
+ options while applying the configured request-policy pipeline. This is a **proposed
362
+ interface**, not an existing Pi API. It needs public upstream support or an explicitly
363
+ supported integration supplied by the host. Request hooks must receive the correct
364
+ child identity without copying the parent transcript or recursively triggering
365
+ delegation. Prototype with an injected fake port until parity can be demonstrated.
366
+ If that seam is unavailable, keep the package experimental rather than reconstructing
367
+ credentials or accessing private runtime fields.
368
+
369
+ Use a small in-process execution loop around Pi's supported model runtime, with only
370
+ the tools admitted by the broker. Avoid invoking `pi` subprocesses or default child
371
+ sessions, which can load unrelated extensions, commands, configuration, or state.
372
+ An in-process worker is an isolated model context, not an OS sandbox; trusted extension
373
+ code retains the invoking user's privileges.
374
+
375
+ The provider adapter must preserve the active model's effective provider behavior,
376
+ including route identity, headers, base URL, thinking configuration, cancellation,
377
+ usage, and the runtime's authentication refresh path. It must not read, copy, persist,
378
+ or log credentials, create a new authentication store, or silently switch providers.
379
+ Pi may still perform its normal credential-store I/O and OAuth rotation internally.
380
+ The promise is that SpecPi does not handle credentials, not that inference performs no
381
+ authentication-state I/O.
382
+ Inspecting public API source is allowed; resolving real credentials is unnecessary for
383
+ design or fixture tests. Production compatibility requires a separate explicit spike.
384
+
385
+ Default to the parent's exact active model. Cheaper workers or a different consultant
386
+ model require a human-approved exact route and task-class evidence. Strong synthesis
387
+ does not reliably repair poor retrieval or a weak delegate's unsupported assertions.
388
+ Keep capability selection separate from a generic cheap/expensive difficulty ladder.
389
+
390
+ An active policy or model change closes admission for the old generation and cancels
391
+ its workers. Branch navigation, task-card revision, shutdown, and scope changes that
392
+ invalidate a grant also invalidate outstanding work. A provider transport that ignores
393
+ abort cannot publish late results or obtain a new scheduler slot under a false claim
394
+ that its previous request stopped. Quarantine the unresolved request until it settles
395
+ or the session closes; state the remaining billing uncertainty.
396
+
397
+ Proposed implementation layout, not files to install yet:
398
+
399
+ ```text
400
+ specpi-delegation/
401
+ package.json Pi extension package; optional host peers
402
+ extensions/delegation/
403
+ index.ts command/tool registration and lifecycle
404
+ protocol.mjs closed schemas and protocol validation
405
+ controller.mjs admission, reservations and transitions
406
+ runner.mjs bounded worker loop and cancellation
407
+ provider.ts public Pi model-runtime adapter
408
+ evidence.mjs snapshot capture and source broker
409
+ prompts.mjs compact mode-specific instructions
410
+ tests/ fake providers, adversarial fixtures, integration
411
+ ```
412
+
413
+ The web adapter is a later isolated module, not an initial dependency. This package
414
+ needs no orchestration framework, daemon, database, agent persona marketplace, vector
415
+ store, or executable plugin discovery. Depend on Pi's public host modules without
416
+ bundling another Pi runtime. Do not deep-import SpecPi private functions: define a
417
+ small public integration seam before reusing task/scope information.
418
+
419
+ The bounded broker enforces its own capabilities. Parent Command Guard interception
420
+ does not automatically cover child calls: the host installs those events through
421
+ `AgentSession`. Reusing a built-in tool factory or `pi.exec()` is not equivalent.
422
+ Any future route to parent tools needs a supported guarded dispatch seam and explicit
423
+ context, rather than a copied tool implementation. The initial broker can be smaller
424
+ and avoid parent-tool execution entirely.
425
+ [Pinned session interception](https://github.com/earendil-works/pi/blob/v0.84.4/packages/coding-agent/src/core/agent-session.ts#L451).
426
+
427
+ Existing `@earendil-works/pi-ai` usage in SpecPi is currently the schema helper, not
428
+ delegated inference. Shipping this design would change that third-party contract and
429
+ requires updates to `THIRD_PARTY.md`, `CHANGELOG.md`, and `SECURITY_MODEL.md` together
430
+ with the adapter and its tests. None of those runtime claims is changed by this RFC.
431
+
432
+ ## 9. Acceptance and failure recovery
433
+
434
+ The host validates structure and provenance; the parent validates meaning.
435
+
436
+ | Worker situation | Required behavior |
437
+ | ----------------------------------------- | -------------------------------------------------------------------------------------------------------- |
438
+ | Missing context | Return the specific missing source and affected conclusion; one bounded follow-up is possible |
439
+ | Provider throttling before useful output | Honor bounded retry delay if the original deadline and reservation allow it |
440
+ | Connection loss after dispatch | Mark usage/outcome uncertain; do not assume retry is free or exactly once |
441
+ | Repeated unsupported tool request | Reject it; one explanatory result, then stop the job if repeated |
442
+ | A normal failed hypothesis | Preserve the observation; continue only with a discriminating next step inside the job's limits |
443
+ | Repeated failure without changed evidence | Stop; return diagnostic evidence to the parent |
444
+ | Malformed final payload | Treat as invalid output; a format correction, if admitted, consumes a request and the original allowance |
445
+ | Source, requirement, or generation drift | Mark stale; do not accept or automatically rerun it |
446
+ | Conflicting findings | Parent inspects the cited sources or runs a discriminating check; no majority vote |
447
+ | Timeout or user cancellation | Stop tools, abort inference, reject late output; report partial observations only as incomplete |
448
+ | Repeated low-value delegation | Finish the task in the parent; record outcome only if measurement is enabled |
449
+
450
+ Before incorporating a result, verify the job and packet identities, active task
451
+ generation, reference resolution, declared source coverage, and resource accounting.
452
+ Then inspect evidence for material claims. An evidence hash proves which bytes were
453
+ referenced, not whether the inference was correct. A review result with no findings
454
+ and missing required coverage is incomplete, not a clean bill of health.
455
+
456
+ For code reviews, each finding must identify a concrete trigger, a violated requirement
457
+ or behavior, a source location, and supporting evidence. Distinguish observed failure,
458
+ reasoned defect, and unverified suspicion. The parent adjudicates `confirmed`,
459
+ `rejected`, or `needs_check`; disagreement is a reason to investigate, not to run a
460
+ debate until the agents agree. Worker-suggested commands are text until the parent
461
+ chooses to run them through its normal controls.
462
+
463
+ Checks use the actual implementation and approved fixtures. A child cannot mark a
464
+ task complete, satisfy a human-selection requirement, retire a wishlist item, or issue
465
+ a trusted verification receipt. The parent retains those existing boundaries.
466
+
467
+ ## 10. Retention and integration with SpecPi
468
+
469
+ Default retention is memory for active jobs and results, with no package-owned
470
+ transcript journal. The parent explicitly collects a bounded result. Once returned as
471
+ a Pi tool response, that result may be persisted or compacted by Pi and sent in later
472
+ parent inference. Memory-only worker storage does not imply a transcript-free or
473
+ offline workflow. The UI must describe this boundary accurately.
474
+
475
+ Result bodies, paths, task descriptions, source excerpts, and model reasoning are not
476
+ written to an additional metrics journal. Opt-in local metrics retain only bounded
477
+ counts, timings, route/version identifiers, accounting availability, and adjudicated
478
+ outcome categories. Even those fields can reveal work patterns; export is explicit.
479
+ No cross-session cache, automatic transcript mining, or global background learning.
480
+
481
+ `/task` remains optional. When an active card is available through an agreed public
482
+ adapter, bind the job to its exact requirement IDs and digest. Otherwise bind to an
483
+ explicit ephemeral requirement set supplied for this task. `/scope` limits parent
484
+ change governance; it is not a worker filesystem sandbox. `/challenge` can consume
485
+ adjudicated findings as supplementary evidence, while its existing deterministic gate
486
+ and parent-run checks remain authoritative.
487
+
488
+ A routing improvement follows SpecPi's existing self-improvement contract: observed
489
+ friction, a proposed narrow change, exact human selection when sourced from the
490
+ wishlist, implementation, required validation, and later outcome evidence. Job success
491
+ does not give the extension permission to rewrite its own prompts, tools, budgets,
492
+ model routes, or routing policy. Policy changes get explicit version identities so
493
+ evaluation cohorts remain interpretable.
494
+
495
+ ## 11. Delivery sequence and decisions
496
+
497
+ This is the original proposed sequence, not an implementation-status table. The
498
+ current experiment exposes two purposes, `review` and `scout`, through actual Pi
499
+ AgentSession instances under `bounded-pi-sessions-v1`. Their implementation does not establish the comparative benefits required
500
+ by the stages below.
501
+
502
+ | Stage | Build | Required evidence before proceeding |
503
+ | -------------------- | --------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- |
504
+ | 0: compatibility | Public provider/lifecycle spike and fake-provider contract tests | Correct route/options, no credential handling or autoloaded resources, abort/usage semantics understood |
505
+ | 1: review | One tool-free call, explicit packet, parent adjudication, no persistence | Useful findings beyond a same-budget parent review, bounded false positives and correction time |
506
+ | 2: investigation | Snapshot broker, two-worker scheduler, reservations and bounded result collection | Better accepted outcomes, cost, or elapsed time on separable tasks; no source/grant leakage |
507
+ | 3: public research | Reviewed public-source adapter and provenance receipts | Verified coverage gains under the declared network and spending boundaries |
508
+ | 4: consultation | Exact approved alternative model and decision packet | Task-class gain beyond same-model fresh context and sequential model routing |
509
+ | 5: route calibration | Versioned, human-selected policy from observed outcomes | Holdout benefit and reliable rollback; no silent online experimentation |
510
+
511
+ This is an implementation order, not a claim that review has the largest expected
512
+ benefit. Review is the smallest integration that tests the provider and result
513
+ contracts. Independent research is a stronger directly supported delegation use case,
514
+ but it requires a larger evidence and network boundary.
515
+
516
+ Worker code writing is excluded from these stages. A later patch-producing experiment
517
+ would need truly isolated inputs, explicit ownership, an enforced execution boundary,
518
+ source/patch fingerprints, and parent-only application. A Git worktree alone does not
519
+ contain shell execution or provide credential isolation. Concurrent writers must beat
520
+ the same one-writer system in measured integration cost before adding that machinery.
521
+
522
+ The recommended first release is deliberately small: one capable parent, one bounded
523
+ fresh review, a protocol that already accounts for incomplete evidence and cost, and
524
+ no generic swarm infrastructure. Grow into two-worker investigation using measured
525
+ results, not a larger default agent count.