specpi 0.11.2 → 0.13.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +21 -1
- package/NPM_RELEASE.md +3 -1
- package/README.md +41 -29
- package/SECURITY_MODEL.md +36 -0
- package/THIRD_PARTY.md +18 -2
- package/docs/browser-testing.md +76 -0
- package/docs/delegation/README.md +264 -0
- package/docs/delegation/design-protocol.md +382 -0
- package/docs/delegation/design.md +525 -0
- package/docs/delegation/evaluation.md +307 -0
- package/docs/delegation/protocol.md +271 -0
- package/docs/delegation/research.md +216 -0
- package/extensions/browser/core.d.mts +64 -0
- package/extensions/browser/diagnostics.ts +275 -0
- package/extensions/browser/index.ts +349 -57
- package/extensions/browser/interactions.ts +118 -0
- package/extensions/browser/lifecycle.ts +28 -0
- package/extensions/command-guard/index.ts +118 -33
- package/extensions/delegation/core.mjs +772 -0
- package/extensions/delegation/errors.mjs +8 -0
- package/extensions/delegation/extension.mjs +475 -0
- package/extensions/delegation/index.ts +9 -0
- package/extensions/delegation/managed-files.mjs +13 -0
- package/extensions/delegation/native.mjs +155 -0
- package/extensions/delegation/presentation.mjs +315 -0
- package/extensions/delegation/protocol.mjs +296 -0
- package/extensions/delegation/provider.mjs +689 -0
- package/extensions/delegation/snapshot.mjs +532 -0
- package/extensions/delegation/worker.mjs +218 -0
- package/extensions/tool-wishlist/verification.mjs +1 -0
- package/extensions/workflow-controls/index.ts +5 -1
- package/package.json +17 -4
- package/scripts/check-package.mjs +29 -3
- package/scripts/check-pi-package.mjs +5 -0
- package/scripts/check-syntax.mjs +61 -0
- package/scripts/run-browser-tests.mjs +61 -0
- package/scripts/setup-browser-tests.mjs +38 -0
- package/scripts/site-browser.mjs +272 -0
- package/scripts/specpi.mjs +24 -0
- package/site/logo.svg +1 -9
- package/templates/AGENTS.md +1 -0
|
@@ -0,0 +1,525 @@
|
|
|
1
|
+
# SpecPi delegation design: archived target proposal
|
|
2
|
+
|
|
3
|
+
Status: archived target architecture, not the implemented runtime contract.
|
|
4
|
+
|
|
5
|
+
The experimental implementation is part of `specpi` and is disabled by default.
|
|
6
|
+
Read the [implemented guide](README.md) and [calls/time protocol](protocol.md) for
|
|
7
|
+
supported commands, tested API compatibility, limits and trust assumptions. This
|
|
8
|
+
document preserves the original broader proposal, including unimplemented live-web,
|
|
9
|
+
alternative-model, monetary, raw-transport and underlying-attempt guarantees. Its
|
|
10
|
+
normative requirements, command sketches and delivery stages are targets, not runtime
|
|
11
|
+
promises or evidence that the experiment has passed outcome evaluation.
|
|
12
|
+
|
|
13
|
+
The current implementation uses actual Pi AgentSession workers with only `review` and
|
|
14
|
+
`scout` admission, explicit parent model/thinking, selected-source tools and in-memory
|
|
15
|
+
sessions. It supersedes the proposed custom loop and four-profile sequence below.
|
|
16
|
+
|
|
17
|
+
Research reviewed through 5 September 2026. The original working name
|
|
18
|
+
`specpi-delegation` below was a design placeholder, not a separate published or reserved
|
|
19
|
+
npm package.
|
|
20
|
+
|
|
21
|
+
The [target protocol](design-protocol.md) specifies this proposal's messages and state
|
|
22
|
+
transitions. [Research](research.md) records the evidence and its limits.
|
|
23
|
+
[Evaluation](evaluation.md) separates current fixture coverage from experiments and
|
|
24
|
+
stronger proof obligations that remain outstanding.
|
|
25
|
+
|
|
26
|
+
## Decision
|
|
27
|
+
|
|
28
|
+
Build an optional Pi extension that gives one capable parent **bounded access to
|
|
29
|
+
additional reasoning and investigation**, while the parent owns changes and acceptance.
|
|
30
|
+
Use one scheduler, one provider adapter, one evidence broker, and one result format.
|
|
31
|
+
Start with a tool-free review call, then add independent investigation and research
|
|
32
|
+
through the same protocol. Treat consultation as another bounded call, rather than
|
|
33
|
+
introducing a second planner or a permanent team.
|
|
34
|
+
|
|
35
|
+
The objective is the best accepted outcome within the user's resource constraints.
|
|
36
|
+
There is no demonstrated universal optimum in agent count, topology, model, or context
|
|
37
|
+
size. The architecture therefore includes **no delegation** as a first-class result of
|
|
38
|
+
routing. It should make a useful two-worker investigation straightforward and make an
|
|
39
|
+
unnecessary ten-worker swarm difficult to create.
|
|
40
|
+
|
|
41
|
+
The evidence supports this structure more strongly than an unrestricted writing swarm.
|
|
42
|
+
Task-dependent gains in the [existing seven-study assessment](../../site/single-agent/index.html)
|
|
43
|
+
establish specific opportunities. Cognition's April 2026 production account also
|
|
44
|
+
describes useful review and consultation around a single writer, while reporting a
|
|
45
|
+
quality ceiling when the primary model was too weak to delegate effectively.
|
|
46
|
+
[Cognition, 22 April 2026](https://cognition.com/blog/multi-agents-working).
|
|
47
|
+
|
|
48
|
+
This design is an engineering inference. Its defaults and promotion thresholds below
|
|
49
|
+
are proposed operating parameters, not numbers established by those studies.
|
|
50
|
+
|
|
51
|
+
## 1. Choose execution by task shape
|
|
52
|
+
|
|
53
|
+
| Situation | Starting route | Why another context may help | Stop condition |
|
|
54
|
+
| ------------------------------------------------------------------- | ---------------------------------------------------------- | ------------------------------------------------------------------------ | -------------------------------------------------------------------- |
|
|
55
|
+
| Small fix, known answer, one short lookup | Parent only; batch independent tools when useful | Usually insufficient benefit | Do not create a worker |
|
|
56
|
+
| Long sequential reasoning with all relevant context already present | Same agent, stronger workflow or supported model selection | Separate contexts can lose the reasoning chain | Keep execution serial |
|
|
57
|
+
| Two substantial, independent evidence questions | Two investigation or research jobs | Parallel exploration and less irrelevant material in the parent | Evidence sufficient to answer the questions |
|
|
58
|
+
| One large source collection | Partition by question or source ownership, then synthesize | Workers can search deeply without sending every intermediate result back | Coverage achieved, result and resource limits reached |
|
|
59
|
+
| Frozen implementation with meaningful review risk | One fresh reviewer | Different context and focused inspection | Findings and coverage returned; parent adjudicates |
|
|
60
|
+
| A specific difficult decision or failed approach | One capable consultant | A different capability or fresh interpretation may help | Answer, discriminating next experiment, or explicit missing evidence |
|
|
61
|
+
| Coupled implementation, evolving interfaces, shared generated files | Parent is the only writer | Concurrent writers introduce integration decisions | Finish the interface or obtain evidence before proceeding |
|
|
62
|
+
|
|
63
|
+
These are routing hypotheses. The parent can misclassify a dependency or a source
|
|
64
|
+
collection. The scheduler checks the declared contract; it cannot prove semantic
|
|
65
|
+
independence. Evaluation must count bad decomposition as a system failure.
|
|
66
|
+
|
|
67
|
+
Four modes share one engine:
|
|
68
|
+
|
|
69
|
+
- **Review:** inspect a frozen packet; report material defects and uncovered requirements.
|
|
70
|
+
- **Investigate:** answer a repository question using bounded snapshot reads.
|
|
71
|
+
- **Research:** answer an external evidence question using an explicitly enabled web adapter.
|
|
72
|
+
- **Consult:** examine a difficult decision using selected context and, when authorized,
|
|
73
|
+
a different exact model. This mode has no tools initially.
|
|
74
|
+
|
|
75
|
+
A model switch inside the parent's existing execution is a separate alternative in the
|
|
76
|
+
evaluation. Learned sequential routing has evidence of its own; it does not establish
|
|
77
|
+
the benefit of spawning children. [MTRouter](https://aclanthology.org/2026.acl-long.2045/).
|
|
78
|
+
|
|
79
|
+
## 2. One level of delegation, one owner of decisions
|
|
80
|
+
|
|
81
|
+
```mermaid
|
|
82
|
+
flowchart TD
|
|
83
|
+
U[Human task and resource policy] --> P[Parent: intent, decisions, writing]
|
|
84
|
+
P --> G{Delegation admission}
|
|
85
|
+
G -->|insufficient benefit or unsupported capability| S[Continue in parent]
|
|
86
|
+
G -->|bounded independent work| Q[Scheduler and shared budget ledger]
|
|
87
|
+
Q --> A[Worker A: explicit packet]
|
|
88
|
+
Q --> B[Worker B: explicit packet]
|
|
89
|
+
A <--> E[Evidence broker: granted snapshots or public sources]
|
|
90
|
+
B <--> E
|
|
91
|
+
A --> V[Validate provenance, coverage, state and usage]
|
|
92
|
+
B --> V
|
|
93
|
+
V --> P
|
|
94
|
+
P --> C[Actual checks and human acceptance]
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
Workers cannot spawn workers, send sibling messages, alter policy, or write the project.
|
|
98
|
+
They return evidence to the parent. The parent decides whether a discovery changes the
|
|
99
|
+
task or another lane. A changed decision produces a new packet generation; it does not
|
|
100
|
+
silently replace instructions in a running job.
|
|
101
|
+
|
|
102
|
+
Initially the scheduler accepts a batch of independent jobs. Dependent work requires
|
|
103
|
+
the parent to accept the prerequisite and submit the next batch. This gives a useful
|
|
104
|
+
parallel frontier without a general workflow engine, distributed message bus, agent
|
|
105
|
+
registry, or arbitrary task graph interpreter.
|
|
106
|
+
|
|
107
|
+
Parallel work is admitted only when its result is needed and either the parent has
|
|
108
|
+
useful independent work or multiple workers can finish independent questions together.
|
|
109
|
+
A review may be entirely serial and still earn its cost through improved detection.
|
|
110
|
+
There is no requirement to keep workers busy.
|
|
111
|
+
|
|
112
|
+
## 3. Human control without approval fatigue
|
|
113
|
+
|
|
114
|
+
Installation and activation are separate. The package starts disabled and adds no
|
|
115
|
+
standing model calls. Proposed command surface:
|
|
116
|
+
|
|
117
|
+
```text
|
|
118
|
+
/delegate Show mode, grants, limits, running jobs and usage
|
|
119
|
+
/delegate on Enable bounded delegation for this session
|
|
120
|
+
/delegate off Stop admission and cancel outstanding jobs
|
|
121
|
+
/delegate limits Inspect or change the session's resource envelope
|
|
122
|
+
/delegate cancel <id> Cancel one job or batch
|
|
123
|
+
```
|
|
124
|
+
|
|
125
|
+
Activation presents the concrete capability envelope: exact model route, review-only
|
|
126
|
+
or snapshot access, web access if any, concurrency, call limits, retention, and budget
|
|
127
|
+
accounting limitations. A human-approved envelope permits subsequent parent requests
|
|
128
|
+
within that envelope. It does not require another permission prompt for every worker
|
|
129
|
+
or file read. Human instructions already authorizing an exact envelope can satisfy this
|
|
130
|
+
selection. Unsupported capabilities remain unavailable.
|
|
131
|
+
|
|
132
|
+
The parent proposes a job through one model tool, `delegate`, with operations `run`,
|
|
133
|
+
`status`, `collect`, `follow_up`, `resolve`, and `cancel`. A `run` response returns an opaque batch ID and the
|
|
134
|
+
admitted job identities. Detailed worker tools and prompts are disclosed only when the
|
|
135
|
+
corresponding mode is used. Keep the standing description small; do not load four role
|
|
136
|
+
manuals into every turn.
|
|
137
|
+
|
|
138
|
+
The extension owns resolved grants, budget reservations, model identities, job IDs, and
|
|
139
|
+
generation tokens. The parent model cannot grant itself broader access by changing
|
|
140
|
+
tool arguments. A denied Command Guard operation remains denied; delegation is never
|
|
141
|
+
an alternate execution route for it. In Strict mode, approval of the top-level custom
|
|
142
|
+
tool must expose the effective delegation capabilities, rather than an opaque label.
|
|
143
|
+
|
|
144
|
+
User-facing progress is one status row, for example:
|
|
145
|
+
|
|
146
|
+
```text
|
|
147
|
+
Delegation 1 running · 1 ready · 4/12 model calls · cost unavailable
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
Only blockers, usable results, and meaningful state changes deserve notifications.
|
|
151
|
+
No periodic narration, hidden follow-up turns, or automatic restart after closing Pi.
|
|
152
|
+
|
|
153
|
+
## 4. Admission and resource allocation
|
|
154
|
+
|
|
155
|
+
Admission has two layers. **The parent judges usefulness; deterministic code enforces
|
|
156
|
+
authority and limits.** Do not ask another model to decide whether to launch a model.
|
|
157
|
+
|
|
158
|
+
The parent supplies a short reason, deliverable, evidence needed for acceptance,
|
|
159
|
+
independence boundary, expected remaining work, and the best parent-only alternative.
|
|
160
|
+
The controller then checks:
|
|
161
|
+
|
|
162
|
+
1. The mode, model, tools, and inputs are inside the current human-authorized policy.
|
|
163
|
+
2. Original requirements and relevant settled decisions are represented in the packet.
|
|
164
|
+
3. The job is ready; its prerequisites are accepted and its result has a named consumer.
|
|
165
|
+
4. Its source snapshot and task generation are still current for admission.
|
|
166
|
+
5. No equivalent job is running or already has an applicable result in this batch.
|
|
167
|
+
6. A concurrency slot and the job's full resource reservation are available.
|
|
168
|
+
7. The result has a feasible verification path. A worker's confidence is not that path.
|
|
169
|
+
|
|
170
|
+
Failure returns a reason and leaves the parent able to work. It does not launch a
|
|
171
|
+
smaller unrequested model, widen access, or silently weaken an exact provider rule.
|
|
172
|
+
|
|
173
|
+
Once measured task-class data exists, select routes against one declared objective:
|
|
174
|
+
|
|
175
|
+
| Objective | Route comparison |
|
|
176
|
+
| --------- | ------------------------------------------------------------------------------ |
|
|
177
|
+
| Cost | Minimize expected total cost per accepted task, with a specified quality floor |
|
|
178
|
+
| Time | Minimize expected elapsed time, with quality and spend constraints |
|
|
179
|
+
| Quality | Maximize accepted outcomes under a fixed resource envelope |
|
|
180
|
+
|
|
181
|
+
Total cost includes packet preparation, all model and tool work, retries, parent
|
|
182
|
+
synthesis, verification, and correction. Report human review time separately unless a
|
|
183
|
+
human explicitly chooses a monetary conversion. Compute cost per accepted outcome as
|
|
184
|
+
total cohort cost divided by accepted tasks, including the cost of failed tasks.
|
|
185
|
+
When none is accepted the measure is undefined, not zero.
|
|
186
|
+
|
|
187
|
+
Do not invent a numerical expected-value estimate from an LLM's self-reported
|
|
188
|
+
confidence. Until a route is evaluated, label it experimental and use explicit bounded
|
|
189
|
+
trials. Learned routing, automatic difficulty prediction, and online bandit exploration
|
|
190
|
+
are unnecessary for the first implementation.
|
|
191
|
+
|
|
192
|
+
## 5. Context is a contract, not a transcript fork
|
|
193
|
+
|
|
194
|
+
Every job receives the original relevant requirements, its exact question, acceptance
|
|
195
|
+
criteria, non-goals, selected decisions, and reference identities. It does not inherit
|
|
196
|
+
the parent's tools, mutable memory, complete conversation, or unrelated instructions.
|
|
197
|
+
No Pi history or authentication files are scanned to construct a packet.
|
|
198
|
+
|
|
199
|
+
Use different context policies deliberately:
|
|
200
|
+
|
|
201
|
+
| Mode | Initial context | Deliberately excluded | Expansion |
|
|
202
|
+
| ----------- | ---------------------------------------------------------------------------- | ---------------------------------------------- | ----------------------------------------------------------------------------- |
|
|
203
|
+
| Review | Requirements, frozen diff, relevant source, actual check receipts | Implementation reasoning and preferred verdict | Specific missing source requested; bounded snapshot access in a later release |
|
|
204
|
+
| Investigate | One question, hypotheses, repository map, source manifest | Unrelated task transcript | Broker reads within the granted snapshot |
|
|
205
|
+
| Research | Question, source boundaries, dates, exclusions, required output | Parent's desired conclusion | Bounded search and retrieval of approved public sources |
|
|
206
|
+
| Consult | Decision, constraints, failed attempts, observations, competing explanations | Unrelated history and hidden model reasoning | A targeted request for additional evidence |
|
|
207
|
+
|
|
208
|
+
Fresh review still receives the human requirements. Withholding those requirements to
|
|
209
|
+
make a reviewer more independent can create irrelevant findings. A defect reviewer may
|
|
210
|
+
question a requirement but cannot redefine it. A separate code-only review is a possible
|
|
211
|
+
evaluation ablation, not the normal acceptance contract.
|
|
212
|
+
|
|
213
|
+
Distinguish two optimization steps. **Selection** omits irrelevant material and preserves
|
|
214
|
+
explicit references to what remains available. **Compression** rewrites information and
|
|
215
|
+
can lose crucial details. Start with selection and deterministic limits. Use model
|
|
216
|
+
summaries only as indexed leads, retain their source references, and measure any solve
|
|
217
|
+
rate loss before adopting a compression policy.
|
|
218
|
+
|
|
219
|
+
Workers return `needs_context` when missing information prevents a grounded answer.
|
|
220
|
+
They name the required evidence and the conclusion it would resolve. One focused
|
|
221
|
+
follow-up can be admitted within the original batch budget. Repeated gaps return the
|
|
222
|
+
task to the parent; they do not create an unlimited conversation.
|
|
223
|
+
|
|
224
|
+
Keep a stable instruction prefix and append changing evidence after it, where the
|
|
225
|
+
provider supports useful cache behavior. Record actual cache accounting when exposed.
|
|
226
|
+
Never assume a cache hit, transfer provider-specific reasoning blocks between models,
|
|
227
|
+
or sacrifice necessary context to satisfy a speculative token-saving target.
|
|
228
|
+
|
|
229
|
+
## 6. The evidence broker
|
|
230
|
+
|
|
231
|
+
Workers get explicit function schemas, never a generic shell or arbitrary extension
|
|
232
|
+
dispatcher. The first repository tools are `list_sources`, `read_source`, and
|
|
233
|
+
`search_sources`, operating on opaque IDs in a bounded in-memory snapshot. A source ID
|
|
234
|
+
maps to captured bytes and a content hash, not a worker-supplied filesystem path.
|
|
235
|
+
|
|
236
|
+
The parent selects repository-relative inputs. The host resolves and validates them
|
|
237
|
+
before capture: canonical root containment, regular files, allowed types, explicit byte
|
|
238
|
+
limits, and exclusion of private state, credentials, binary files, links and reparse
|
|
239
|
+
points whose containment cannot be established. Search is a bounded literal search
|
|
240
|
+
initially; arbitrary regular expressions and shell-based grep are unnecessary. The
|
|
241
|
+
broker does not invoke project scripts, Git filters, LSP plugins, or package hooks.
|
|
242
|
+
|
|
243
|
+
Take source hashes before and after capture. This detects observed changes but is not
|
|
244
|
+
an atomic operating-system snapshot. Serve only the captured bytes. A manifest labels
|
|
245
|
+
the snapshot's coverage and capture limitations. If a required consistent source set
|
|
246
|
+
cannot be captured, return `source_changed` and ask the parent to establish a stable
|
|
247
|
+
target. Never claim to have reviewed a moving checkout as one atomic revision.
|
|
248
|
+
|
|
249
|
+
Large repositories do not justify uploading the repository. Start with a selected
|
|
250
|
+
bounded manifest. If the selected input exceeds its envelope, request a smaller scope
|
|
251
|
+
or an explicit expansion. Hidden omissions or silently shortened evidence are invalid.
|
|
252
|
+
This also keeps source hashes and exact line citations useful for parent inspection.
|
|
253
|
+
|
|
254
|
+
The research adapter is a separate opt-in capability with typed `search_public` and
|
|
255
|
+
`fetch_public` operations. It must implement a reviewed network boundary: allowed
|
|
256
|
+
schemes, host/IP checks at connection time and redirects, public-address restrictions,
|
|
257
|
+
byte/time limits, no cookies or inherited authorization, bounded downloads, and no
|
|
258
|
+
execution of page scripts. URL parsing alone does not prevent DNS rebinding or access
|
|
259
|
+
to private services. If the chosen transport cannot enforce the declared boundary,
|
|
260
|
+
public fetch is unavailable. Do not quietly proxy arbitrary existing browser or MCP
|
|
261
|
+
tools through the worker. Parent-supplied public excerpts remain a usable fallback.
|
|
262
|
+
|
|
263
|
+
Source text and worker output are untrusted evidence. Neither can create permissions,
|
|
264
|
+
change budgets, inject system instructions, or authorize a sibling action. Exact source
|
|
265
|
+
locations establish traceability; they do not establish truth.
|
|
266
|
+
|
|
267
|
+
## 7. Resource limits that describe what is actually enforced
|
|
268
|
+
|
|
269
|
+
These are initial trial defaults, to be tuned through evaluation:
|
|
270
|
+
|
|
271
|
+
| Limit | Proposed default | Enforcement |
|
|
272
|
+
| -------------------------- | --------------------------------------------------------------------------- | ---------------------------------------------------------------------------------- |
|
|
273
|
+
| Active workers | 2 per batch; one active batch per parent session | Scheduler semaphore |
|
|
274
|
+
| Session allowance | 4 batches and 48 model requests | Monotone session ledger; renewed only by a human policy change |
|
|
275
|
+
| Logical jobs | 4 per batch | Admission counter; no nested delegation |
|
|
276
|
+
| Model requests | 12 total per batch, including retries and follow-ups | Reserve before every underlying inference attempt |
|
|
277
|
+
| Requests per logical job | 2 for review/consult; 4 for investigate/research | Review/consult uses one initial call; the second is at most one retry or follow-up |
|
|
278
|
+
| Follow-up generation | At most 1 per job | Explicit parent resubmission; same batch allowance |
|
|
279
|
+
| Tool calls | 12 per job | Broker counter before execution |
|
|
280
|
+
| Job / batch deadline | 120 / 300 seconds | Monotonic deadline, cancellation, late-result rejection |
|
|
281
|
+
| Selected source snapshot | 8 MiB and 200 files per batch | Checked before reads; bounded allocation |
|
|
282
|
+
| Serialized model context | 256 KiB per request, also within model context limits | Byte cap plus provider token/context validation when supported |
|
|
283
|
+
| Evidence returned by tools | 64 KiB total per job | Count bytes; explicit incomplete result at limit |
|
|
284
|
+
| Final worker payload | 16 KiB; 8 findings | Reject oversized payloads, never silently accept a truncated object |
|
|
285
|
+
| Raw provider response | 256 KiB per attempt, including text, reasoning, tool arguments and metadata | Bound transport buffering and accumulation before parsing or appending content |
|
|
286
|
+
| Provider output limit | 2,048 tokens per request where supported | Provider request setting; record unsupported accounting |
|
|
287
|
+
| Retry | One classified retry per job | Counts against calls, deadline, and spend admission |
|
|
288
|
+
|
|
289
|
+
All attempts under the same logical job share its tool, evidence, elapsed-time, and
|
|
290
|
+
follow-up limits. A replacement does not reset the counters. Limits are configurable
|
|
291
|
+
inside human policy; worker tool arguments can only request less. Source and result
|
|
292
|
+
bytes are local memory limits, not token estimates. Hidden reasoning, provider retries,
|
|
293
|
+
tokenization, rate limits, and post-cancellation billing require separate accounting.
|
|
294
|
+
Batch completion, `/task` revision, normal user steering and model changes do not
|
|
295
|
+
replenish the session allowance. An on/off toggle does not erase consumed usage.
|
|
296
|
+
Session restart clears ephemeral state and starts disabled; no work resumes by itself.
|
|
297
|
+
|
|
298
|
+
For monetary admission, reserve the conservative cost of the next request, using
|
|
299
|
+
complete supported pricing and bounded billable input/output categories; do not assume
|
|
300
|
+
cache discounts. Atomically reserve before dispatch, reconcile with actual usage, and
|
|
301
|
+
keep unknown in-flight charges reserved. Every retry pays for its own reservation.
|
|
302
|
+
Reject concurrent requests whose combined reservations exceed the admitted envelope.
|
|
303
|
+
|
|
304
|
+
Here a counted model request means an underlying inference attempt, including an SDK
|
|
305
|
+
or provider-adapter retry. `InferencePort` must either disable opaque automatic retries
|
|
306
|
+
under the admitted child policy or obtain controller admission before every attempt.
|
|
307
|
+
One method call that internally retries three times consumes four attempts. Report
|
|
308
|
+
high-level invocations separately. An adapter that cannot expose or prevent hidden
|
|
309
|
+
attempts cannot satisfy a hard attempt cap and is unavailable for that policy.
|
|
310
|
+
|
|
311
|
+
The raw response cap applies while receiving every content category, before assembling
|
|
312
|
+
JSON, reasoning, tool arguments or a final result. The adapter must bound transport
|
|
313
|
+
chunks and its own buffering, not merely check the completed response size. Overflow
|
|
314
|
+
aborts the attempt, preserves usage uncertainty and produces an invalid/incomplete
|
|
315
|
+
outcome; a truncated object is never treated as a successful answer. Discard buffered
|
|
316
|
+
content when its job no longer needs it, while retaining bounded usage and state.
|
|
317
|
+
|
|
318
|
+
A configured spend ceiling is an **admission limit**, not a guarantee of a final invoice.
|
|
319
|
+
Cancellation cannot retract work already performed by a provider. If the provider cannot
|
|
320
|
+
supply a defensible cost bound, either decline a cost-capped route or use a separately
|
|
321
|
+
human-authorized calls/time-limited mode with cost explicitly marked unavailable. Never
|
|
322
|
+
turn missing billing or subscription usage into `$0.00`.
|
|
323
|
+
|
|
324
|
+
This extension can enforce its own requests and broker calls. It cannot cap the parent's
|
|
325
|
+
ordinary inference, other extensions, or a provider's unrelated account usage. Reserve
|
|
326
|
+
parent synthesis and verification capacity in the evaluation plan; an extension-only
|
|
327
|
+
spend counter must not be presented as the total task budget.
|
|
328
|
+
|
|
329
|
+
## 8. Pi runtime and package boundary
|
|
330
|
+
|
|
331
|
+
The following records the original compatibility analysis and its stronger target.
|
|
332
|
+
The current native experiment instead uses public SDK `createAgentSession` with a fresh
|
|
333
|
+
standard Pi `ModelRuntime`, explicit parent model/thinking, in-memory sessions and
|
|
334
|
+
admission before each SDK invocation. Delegation checks required public SDK capabilities
|
|
335
|
+
instead of gating exact Pi versions; the normal installer floor and bootstrap pin are
|
|
336
|
+
separate contracts. See the compatibility record for tested versions and completed validation.
|
|
337
|
+
SDK-visible streams are checked, but full parent-hook/ephemeral-setting parity and
|
|
338
|
+
hard raw-transport, hidden-attempt and monetary bounds remain unclaimed.
|
|
339
|
+
See the [implemented guide](README.md). The target `InferencePort` below is not an
|
|
340
|
+
exported Pi API or a claim that the experiment supplies these stronger guarantees.
|
|
341
|
+
|
|
342
|
+
**Production is gated on a supported host request capability.** Inspection of Pi
|
|
343
|
+
`v0.84.4`, SpecPi's reviewed host floor, found an important distinction:
|
|
344
|
+
|
|
345
|
+
| Verified public contract | What it establishes | What it does not establish |
|
|
346
|
+
| ----------------------------------------------------- | --------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------ |
|
|
347
|
+
| `ctx.modelRegistry.complete(model, context, options)` | Runtime-owned authentication and effective provider preparation | The host's full request hooks, thinking translation, settings, transport and session-affinity pipeline |
|
|
348
|
+
| `ctx.model`, `ctx.thinkingLevel`, `ctx.signal` | Active route metadata and parent cancellation | A reusable configured child streaming capability |
|
|
349
|
+
| `pi.getAllTools()` | Tool metadata and schemas | An executable dispatcher that inherits Command Guard |
|
|
350
|
+
| `Agent` / `agentLoop` from `pi-agent-core` | A reusable loop with explicit streaming and tool callbacks | The parent's hooks, grants or runtime automatically |
|
|
351
|
+
| `SessionManager.inMemory()` | No child session persistence | No configuration, authentication-service or extension discovery |
|
|
352
|
+
|
|
353
|
+
The registry facade does not expose `streamSimple`. Calling its completion method is
|
|
354
|
+
credential-blind, but it would skip host request layers that can carry provider policy.
|
|
355
|
+
Passing `reasoning: ctx.thinkingLevel` is not a portable replacement for Pi's provider
|
|
356
|
+
translation. Do not ship that shortcut as equivalent behavior.
|
|
357
|
+
[Pinned registry](https://github.com/earendil-works/pi/blob/v0.84.4/packages/coding-agent/src/core/model-registry.ts#L59),
|
|
358
|
+
[pinned host pipeline](https://github.com/earendil-works/pi/blob/v0.84.4/packages/coding-agent/src/core/sdk.ts#L283).
|
|
359
|
+
|
|
360
|
+
Define a host-owned `InferencePort` that accepts a bound child context and admitted
|
|
361
|
+
options while applying the configured request-policy pipeline. This is a **proposed
|
|
362
|
+
interface**, not an existing Pi API. It needs public upstream support or an explicitly
|
|
363
|
+
supported integration supplied by the host. Request hooks must receive the correct
|
|
364
|
+
child identity without copying the parent transcript or recursively triggering
|
|
365
|
+
delegation. Prototype with an injected fake port until parity can be demonstrated.
|
|
366
|
+
If that seam is unavailable, keep the package experimental rather than reconstructing
|
|
367
|
+
credentials or accessing private runtime fields.
|
|
368
|
+
|
|
369
|
+
Use a small in-process execution loop around Pi's supported model runtime, with only
|
|
370
|
+
the tools admitted by the broker. Avoid invoking `pi` subprocesses or default child
|
|
371
|
+
sessions, which can load unrelated extensions, commands, configuration, or state.
|
|
372
|
+
An in-process worker is an isolated model context, not an OS sandbox; trusted extension
|
|
373
|
+
code retains the invoking user's privileges.
|
|
374
|
+
|
|
375
|
+
The provider adapter must preserve the active model's effective provider behavior,
|
|
376
|
+
including route identity, headers, base URL, thinking configuration, cancellation,
|
|
377
|
+
usage, and the runtime's authentication refresh path. It must not read, copy, persist,
|
|
378
|
+
or log credentials, create a new authentication store, or silently switch providers.
|
|
379
|
+
Pi may still perform its normal credential-store I/O and OAuth rotation internally.
|
|
380
|
+
The promise is that SpecPi does not handle credentials, not that inference performs no
|
|
381
|
+
authentication-state I/O.
|
|
382
|
+
Inspecting public API source is allowed; resolving real credentials is unnecessary for
|
|
383
|
+
design or fixture tests. Production compatibility requires a separate explicit spike.
|
|
384
|
+
|
|
385
|
+
Default to the parent's exact active model. Cheaper workers or a different consultant
|
|
386
|
+
model require a human-approved exact route and task-class evidence. Strong synthesis
|
|
387
|
+
does not reliably repair poor retrieval or a weak delegate's unsupported assertions.
|
|
388
|
+
Keep capability selection separate from a generic cheap/expensive difficulty ladder.
|
|
389
|
+
|
|
390
|
+
An active policy or model change closes admission for the old generation and cancels
|
|
391
|
+
its workers. Branch navigation, task-card revision, shutdown, and scope changes that
|
|
392
|
+
invalidate a grant also invalidate outstanding work. A provider transport that ignores
|
|
393
|
+
abort cannot publish late results or obtain a new scheduler slot under a false claim
|
|
394
|
+
that its previous request stopped. Quarantine the unresolved request until it settles
|
|
395
|
+
or the session closes; state the remaining billing uncertainty.
|
|
396
|
+
|
|
397
|
+
Proposed implementation layout, not files to install yet:
|
|
398
|
+
|
|
399
|
+
```text
|
|
400
|
+
specpi-delegation/
|
|
401
|
+
package.json Pi extension package; optional host peers
|
|
402
|
+
extensions/delegation/
|
|
403
|
+
index.ts command/tool registration and lifecycle
|
|
404
|
+
protocol.mjs closed schemas and protocol validation
|
|
405
|
+
controller.mjs admission, reservations and transitions
|
|
406
|
+
runner.mjs bounded worker loop and cancellation
|
|
407
|
+
provider.ts public Pi model-runtime adapter
|
|
408
|
+
evidence.mjs snapshot capture and source broker
|
|
409
|
+
prompts.mjs compact mode-specific instructions
|
|
410
|
+
tests/ fake providers, adversarial fixtures, integration
|
|
411
|
+
```
|
|
412
|
+
|
|
413
|
+
The web adapter is a later isolated module, not an initial dependency. This package
|
|
414
|
+
needs no orchestration framework, daemon, database, agent persona marketplace, vector
|
|
415
|
+
store, or executable plugin discovery. Depend on Pi's public host modules without
|
|
416
|
+
bundling another Pi runtime. Do not deep-import SpecPi private functions: define a
|
|
417
|
+
small public integration seam before reusing task/scope information.
|
|
418
|
+
|
|
419
|
+
The bounded broker enforces its own capabilities. Parent Command Guard interception
|
|
420
|
+
does not automatically cover child calls: the host installs those events through
|
|
421
|
+
`AgentSession`. Reusing a built-in tool factory or `pi.exec()` is not equivalent.
|
|
422
|
+
Any future route to parent tools needs a supported guarded dispatch seam and explicit
|
|
423
|
+
context, rather than a copied tool implementation. The initial broker can be smaller
|
|
424
|
+
and avoid parent-tool execution entirely.
|
|
425
|
+
[Pinned session interception](https://github.com/earendil-works/pi/blob/v0.84.4/packages/coding-agent/src/core/agent-session.ts#L451).
|
|
426
|
+
|
|
427
|
+
Existing `@earendil-works/pi-ai` usage in SpecPi is currently the schema helper, not
|
|
428
|
+
delegated inference. Shipping this design would change that third-party contract and
|
|
429
|
+
requires updates to `THIRD_PARTY.md`, `CHANGELOG.md`, and `SECURITY_MODEL.md` together
|
|
430
|
+
with the adapter and its tests. None of those runtime claims is changed by this RFC.
|
|
431
|
+
|
|
432
|
+
## 9. Acceptance and failure recovery
|
|
433
|
+
|
|
434
|
+
The host validates structure and provenance; the parent validates meaning.
|
|
435
|
+
|
|
436
|
+
| Worker situation | Required behavior |
|
|
437
|
+
| ----------------------------------------- | -------------------------------------------------------------------------------------------------------- |
|
|
438
|
+
| Missing context | Return the specific missing source and affected conclusion; one bounded follow-up is possible |
|
|
439
|
+
| Provider throttling before useful output | Honor bounded retry delay if the original deadline and reservation allow it |
|
|
440
|
+
| Connection loss after dispatch | Mark usage/outcome uncertain; do not assume retry is free or exactly once |
|
|
441
|
+
| Repeated unsupported tool request | Reject it; one explanatory result, then stop the job if repeated |
|
|
442
|
+
| A normal failed hypothesis | Preserve the observation; continue only with a discriminating next step inside the job's limits |
|
|
443
|
+
| Repeated failure without changed evidence | Stop; return diagnostic evidence to the parent |
|
|
444
|
+
| Malformed final payload | Treat as invalid output; a format correction, if admitted, consumes a request and the original allowance |
|
|
445
|
+
| Source, requirement, or generation drift | Mark stale; do not accept or automatically rerun it |
|
|
446
|
+
| Conflicting findings | Parent inspects the cited sources or runs a discriminating check; no majority vote |
|
|
447
|
+
| Timeout or user cancellation | Stop tools, abort inference, reject late output; report partial observations only as incomplete |
|
|
448
|
+
| Repeated low-value delegation | Finish the task in the parent; record outcome only if measurement is enabled |
|
|
449
|
+
|
|
450
|
+
Before incorporating a result, verify the job and packet identities, active task
|
|
451
|
+
generation, reference resolution, declared source coverage, and resource accounting.
|
|
452
|
+
Then inspect evidence for material claims. An evidence hash proves which bytes were
|
|
453
|
+
referenced, not whether the inference was correct. A review result with no findings
|
|
454
|
+
and missing required coverage is incomplete, not a clean bill of health.
|
|
455
|
+
|
|
456
|
+
For code reviews, each finding must identify a concrete trigger, a violated requirement
|
|
457
|
+
or behavior, a source location, and supporting evidence. Distinguish observed failure,
|
|
458
|
+
reasoned defect, and unverified suspicion. The parent adjudicates `confirmed`,
|
|
459
|
+
`rejected`, or `needs_check`; disagreement is a reason to investigate, not to run a
|
|
460
|
+
debate until the agents agree. Worker-suggested commands are text until the parent
|
|
461
|
+
chooses to run them through its normal controls.
|
|
462
|
+
|
|
463
|
+
Checks use the actual implementation and approved fixtures. A child cannot mark a
|
|
464
|
+
task complete, satisfy a human-selection requirement, retire a wishlist item, or issue
|
|
465
|
+
a trusted verification receipt. The parent retains those existing boundaries.
|
|
466
|
+
|
|
467
|
+
## 10. Retention and integration with SpecPi
|
|
468
|
+
|
|
469
|
+
Default retention is memory for active jobs and results, with no package-owned
|
|
470
|
+
transcript journal. The parent explicitly collects a bounded result. Once returned as
|
|
471
|
+
a Pi tool response, that result may be persisted or compacted by Pi and sent in later
|
|
472
|
+
parent inference. Memory-only worker storage does not imply a transcript-free or
|
|
473
|
+
offline workflow. The UI must describe this boundary accurately.
|
|
474
|
+
|
|
475
|
+
Result bodies, paths, task descriptions, source excerpts, and model reasoning are not
|
|
476
|
+
written to an additional metrics journal. Opt-in local metrics retain only bounded
|
|
477
|
+
counts, timings, route/version identifiers, accounting availability, and adjudicated
|
|
478
|
+
outcome categories. Even those fields can reveal work patterns; export is explicit.
|
|
479
|
+
No cross-session cache, automatic transcript mining, or global background learning.
|
|
480
|
+
|
|
481
|
+
`/task` remains optional. When an active card is available through an agreed public
|
|
482
|
+
adapter, bind the job to its exact requirement IDs and digest. Otherwise bind to an
|
|
483
|
+
explicit ephemeral requirement set supplied for this task. `/scope` limits parent
|
|
484
|
+
change governance; it is not a worker filesystem sandbox. `/challenge` can consume
|
|
485
|
+
adjudicated findings as supplementary evidence, while its existing deterministic gate
|
|
486
|
+
and parent-run checks remain authoritative.
|
|
487
|
+
|
|
488
|
+
A routing improvement follows SpecPi's existing self-improvement contract: observed
|
|
489
|
+
friction, a proposed narrow change, exact human selection when sourced from the
|
|
490
|
+
wishlist, implementation, required validation, and later outcome evidence. Job success
|
|
491
|
+
does not give the extension permission to rewrite its own prompts, tools, budgets,
|
|
492
|
+
model routes, or routing policy. Policy changes get explicit version identities so
|
|
493
|
+
evaluation cohorts remain interpretable.
|
|
494
|
+
|
|
495
|
+
## 11. Delivery sequence and decisions
|
|
496
|
+
|
|
497
|
+
This is the original proposed sequence, not an implementation-status table. The
|
|
498
|
+
current experiment exposes two purposes, `review` and `scout`, through actual Pi
|
|
499
|
+
AgentSession instances under `bounded-pi-sessions-v1`. Their implementation does not establish the comparative benefits required
|
|
500
|
+
by the stages below.
|
|
501
|
+
|
|
502
|
+
| Stage | Build | Required evidence before proceeding |
|
|
503
|
+
| -------------------- | --------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------- |
|
|
504
|
+
| 0: compatibility | Public provider/lifecycle spike and fake-provider contract tests | Correct route/options, no credential handling or autoloaded resources, abort/usage semantics understood |
|
|
505
|
+
| 1: review | One tool-free call, explicit packet, parent adjudication, no persistence | Useful findings beyond a same-budget parent review, bounded false positives and correction time |
|
|
506
|
+
| 2: investigation | Snapshot broker, two-worker scheduler, reservations and bounded result collection | Better accepted outcomes, cost, or elapsed time on separable tasks; no source/grant leakage |
|
|
507
|
+
| 3: public research | Reviewed public-source adapter and provenance receipts | Verified coverage gains under the declared network and spending boundaries |
|
|
508
|
+
| 4: consultation | Exact approved alternative model and decision packet | Task-class gain beyond same-model fresh context and sequential model routing |
|
|
509
|
+
| 5: route calibration | Versioned, human-selected policy from observed outcomes | Holdout benefit and reliable rollback; no silent online experimentation |
|
|
510
|
+
|
|
511
|
+
This is an implementation order, not a claim that review has the largest expected
|
|
512
|
+
benefit. Review is the smallest integration that tests the provider and result
|
|
513
|
+
contracts. Independent research is a stronger directly supported delegation use case,
|
|
514
|
+
but it requires a larger evidence and network boundary.
|
|
515
|
+
|
|
516
|
+
Worker code writing is excluded from these stages. A later patch-producing experiment
|
|
517
|
+
would need truly isolated inputs, explicit ownership, an enforced execution boundary,
|
|
518
|
+
source/patch fingerprints, and parent-only application. A Git worktree alone does not
|
|
519
|
+
contain shell execution or provide credential isolation. Concurrent writers must beat
|
|
520
|
+
the same one-writer system in measured integration cost before adding that machinery.
|
|
521
|
+
|
|
522
|
+
The recommended first release is deliberately small: one capable parent, one bounded
|
|
523
|
+
fresh review, a protocol that already accounts for incomplete evidence and cost, and
|
|
524
|
+
no generic swarm infrastructure. Grow into two-worker investigation using measured
|
|
525
|
+
results, not a larger default agent count.
|