specpi 0.11.2 → 0.12.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +13 -1
- package/NPM_RELEASE.md +3 -1
- package/README.md +38 -28
- package/SECURITY_MODEL.md +30 -0
- package/THIRD_PARTY.md +14 -0
- package/docs/delegation/README.md +264 -0
- package/docs/delegation/design-protocol.md +382 -0
- package/docs/delegation/design.md +525 -0
- package/docs/delegation/evaluation.md +307 -0
- package/docs/delegation/protocol.md +271 -0
- package/docs/delegation/research.md +216 -0
- package/extensions/command-guard/index.ts +118 -33
- package/extensions/delegation/core.mjs +772 -0
- package/extensions/delegation/errors.mjs +8 -0
- package/extensions/delegation/extension.mjs +475 -0
- package/extensions/delegation/index.ts +9 -0
- package/extensions/delegation/managed-files.mjs +13 -0
- package/extensions/delegation/native.mjs +155 -0
- package/extensions/delegation/presentation.mjs +315 -0
- package/extensions/delegation/protocol.mjs +296 -0
- package/extensions/delegation/provider.mjs +689 -0
- package/extensions/delegation/snapshot.mjs +532 -0
- package/extensions/delegation/worker.mjs +218 -0
- package/extensions/workflow-controls/index.ts +5 -1
- package/package.json +4 -2
- package/scripts/check-package.mjs +21 -3
- package/scripts/check-pi-package.mjs +4 -0
- package/scripts/check-syntax.mjs +61 -0
- package/scripts/specpi.mjs +6 -0
- package/site/logo.svg +1 -9
|
@@ -0,0 +1,307 @@
|
|
|
1
|
+
# Delegation evaluation and implementation gates
|
|
2
|
+
|
|
3
|
+
This is an experiment specification, not a report of comparative outcomes. The
|
|
4
|
+
[experimental implementation](README.md) uses Pi AgentSession workers under the
|
|
5
|
+
[bounded session protocol](protocol.md). Runtime, broker, lifecycle and synthetic-provider
|
|
6
|
+
fixtures require a fresh run against this SDK integration before they support a passing
|
|
7
|
+
contract claim. Tests do not establish better task quality, cost or latency, or parity
|
|
8
|
+
with every live provider. No comparative outcome experiment is reported here. The numeric criteria
|
|
9
|
+
below remain proposed hypotheses, not research-derived constants or achieved results.
|
|
10
|
+
|
|
11
|
+
The [archived architecture](design.md) and [target protocol](design-protocol.md) retain
|
|
12
|
+
stronger proof obligations. In particular, full parent request-pipeline parity,
|
|
13
|
+
midstream/raw-transport bounds, admission of every underlying provider attempt and
|
|
14
|
+
monetary admission remain unmet by the native `bounded-pi-sessions-v1` contract.
|
|
15
|
+
|
|
16
|
+
## 1. Decide what improvement means before a run
|
|
17
|
+
|
|
18
|
+
Choose one primary objective for each cohort: quality, cost, or time. Set the quality
|
|
19
|
+
floor, resource envelope, allowed model routes, acceptance rubric and stop rules before
|
|
20
|
+
looking at outcomes. A route must not claim success by changing its objective afterward.
|
|
21
|
+
|
|
22
|
+
Acceptance is task-specific and independently adjudicated. For code, use hidden or
|
|
23
|
+
held-out behavior checks, appropriate regression tests, and human review of the diff.
|
|
24
|
+
For research, use verified factual coverage, source quality, citations, and material
|
|
25
|
+
omissions. For review, use confirmed defects, false positives and reviewer-induced
|
|
26
|
+
regressions. A worker's `complete` declaration or a host receipt is never the target metric.
|
|
27
|
+
|
|
28
|
+
Use versioned public or sanitized fixtures. Keep evaluation data separate from private
|
|
29
|
+
Pi sessions and user repositories unless a human explicitly provides those materials
|
|
30
|
+
for this purpose. No automatic collection from ordinary conversations.
|
|
31
|
+
|
|
32
|
+
## 2. Baselines and comparisons
|
|
33
|
+
|
|
34
|
+
| Arm | Purpose |
|
|
35
|
+
| ------------------------------------------------------- | ------------------------------------------------------------------------------- |
|
|
36
|
+
| A: capable single agent with current workflow | Establish the useful baseline |
|
|
37
|
+
| B: the same agent with the same total additional budget | Distinguish delegation from simply buying more reasoning or another review pass |
|
|
38
|
+
| C: structured serial workflow, same execution context | Test workflow structure without separate identities |
|
|
39
|
+
| D: selective one- or two-worker delegation | Test the experimental implementation |
|
|
40
|
+
| E: always-delegate policy | Measure routing value and unnecessary overhead; evaluation only |
|
|
41
|
+
| F: approved sequential model routing | Separate model-selection benefit from child-context benefit |
|
|
42
|
+
|
|
43
|
+
Do not run all six arms on every case before establishing which question matters.
|
|
44
|
+
For stage 1, A/B/D on frozen reviews is sufficient. Add C/E for investigation and F
|
|
45
|
+
only when alternative models are authorized. Use the same exact provider/model version
|
|
46
|
+
and tools where the comparison calls for it. A mixed-model arm must disclose composition.
|
|
47
|
+
|
|
48
|
+
Start with the implemented `review` and `scout` purposes, not all archived profiles.
|
|
49
|
+
Compare frozen review against both normal parent review and an additional parent pass
|
|
50
|
+
under the same total envelope. Include clean changes so false positives carry a cost.
|
|
51
|
+
For scouts, compare parent-only and structured serial analysis before attributing gains
|
|
52
|
+
to fresh context or parallelism. Evaluate one versus two workers only when independent
|
|
53
|
+
source partitions justify that question.
|
|
54
|
+
|
|
55
|
+
Record the parent's and child's effective thinking settings and model clamping. The
|
|
56
|
+
child uses a fresh standard Pi ModelRuntime; parent hooks and ephemeral settings are
|
|
57
|
+
not inherited. Control these differences when isolating architecture, or disclose them
|
|
58
|
+
as part of an end-to-end workflow comparison. Same model IDs alone do not establish
|
|
59
|
+
equal inference configuration.
|
|
60
|
+
|
|
61
|
+
Use total task spending, not only worker spending. Include preparation, parent reasoning,
|
|
62
|
+
worker requests, tools, retries, synthesis and validation. When actual dollar cost is
|
|
63
|
+
unavailable, label the cohort as call/token/latency constrained and do not describe it
|
|
64
|
+
as dollar-matched. Give the parent sufficient remaining budget to consume and check a
|
|
65
|
+
result; a worker that spends the whole allowance before integration has not succeeded.
|
|
66
|
+
|
|
67
|
+
Match maximum resource envelopes and report actual use in both arms. Equal ceilings
|
|
68
|
+
do not imply equal spending. Run latency comparisons separately from quality/cost
|
|
69
|
+
comparisons because provider load and concurrency can change latency.
|
|
70
|
+
|
|
71
|
+
## 3. Task strata
|
|
72
|
+
|
|
73
|
+
Use at least these separately reported classes:
|
|
74
|
+
|
|
75
|
+
- Small understood edits and short lookups, where delegation should usually be rejected.
|
|
76
|
+
- Coupled changes with unsettled interfaces, where one writer should retain ownership.
|
|
77
|
+
- Independent repository investigations, including contradictory hypotheses.
|
|
78
|
+
- Large source collections with verifiable coverage requirements and repeated evidence.
|
|
79
|
+
- Frozen reviews with naturally occurring defects, seeded defects, and clean changes.
|
|
80
|
+
- Difficult decisions with sufficient context, plus missing-context cases that require abstention.
|
|
81
|
+
- Cancellation, stale source, prompt injection, bad decomposition and provider failure cases.
|
|
82
|
+
|
|
83
|
+
A pilot of 30–50 cases is for debugging the protocol and estimating variance, not
|
|
84
|
+
declaring a universal routing winner. Freeze policy after the pilot. Size a separate
|
|
85
|
+
holdout using the minimum useful effect, expected paired disagreement and cluster
|
|
86
|
+
variation. A target of 200 or more holdout cases can be a planning starting point;
|
|
87
|
+
it is not a power guarantee. Use repeated runs where stochastic variation is material.
|
|
88
|
+
|
|
89
|
+
Partition by repository or source family when possible so close variants do not leak
|
|
90
|
+
between tuning and evaluation. Randomize arm order, record service conditions, define
|
|
91
|
+
cold/warm cache handling, and use repeated time blocks to reduce infrastructure bias.
|
|
92
|
+
Cluster uncertainty by task/repository as appropriate. Include all admitted trials,
|
|
93
|
+
timeouts and failed attempts; do not condition the headline metric on successful jobs.
|
|
94
|
+
|
|
95
|
+
## 4. Required measurements
|
|
96
|
+
|
|
97
|
+
| Measure | Definition or interpretation |
|
|
98
|
+
| ---------------------- | ------------------------------------------------------------------------------------------- |
|
|
99
|
+
| Accepted-task rate | Accepted tasks / all attempted tasks, under the frozen rubric |
|
|
100
|
+
| Cost per accepted task | Sum of all task costs / accepted tasks; undefined if none pass |
|
|
101
|
+
| Resource use | Input/output/cache/tool categories, calls, retries, unknown accounting and model identities |
|
|
102
|
+
| Elapsed time | End-to-end p50/p95 plus uncertainty; separate provider wait, work and integration |
|
|
103
|
+
| Review precision | Confirmed material findings / all material findings presented |
|
|
104
|
+
| Review recall | Confirmed detected defects / adjudicated defects in the evaluated fixtures |
|
|
105
|
+
| Human correction | Time and actions needed to resolve findings and repair final output |
|
|
106
|
+
| Coverage | Required questions or requirements supported by applicable evidence |
|
|
107
|
+
| Context failures | Missing or distorted information that changed a conclusion or prevented acceptance |
|
|
108
|
+
| Coordination overhead | Preparation, repeated source work, messages, synthesis, invalidated jobs and integration |
|
|
109
|
+
| Recovery | Verified recoveries by fault class, including abandoned work and additional cost |
|
|
110
|
+
| Routing errors | Unnecessary delegation, missed useful delegation and unsupported model/capability requests |
|
|
111
|
+
| Policy violations | Denied capabilities attempted, actual escaped capabilities and stale result publication |
|
|
112
|
+
|
|
113
|
+
Do not count a speculative finding as a caught bug. Separate naturally occurring and
|
|
114
|
+
seeded defects. Do not reward verbose output, agreement among agents, smaller packets,
|
|
115
|
+
clean merges, model confidence, or absence of exceptions as substitutes for acceptance.
|
|
116
|
+
|
|
117
|
+
## 5. Promotion rules
|
|
118
|
+
|
|
119
|
+
The user should choose the minimum useful change before a confirmatory run. Suggested
|
|
120
|
+
initial rules for discussion and pre-registration are:
|
|
121
|
+
|
|
122
|
+
| Objective | Candidate rule |
|
|
123
|
+
| --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
|
124
|
+
| Cost | At least 20% lower total cost per accepted task, with the lower one-sided 95% bound for the quality difference above a predeclared −2 percentage-point margin |
|
|
125
|
+
| Time | At least 25% lower median end-to-end latency under the same resource ceiling and quality margin, without unacceptable p95 regression |
|
|
126
|
+
| Quality | Positive paired improvement whose confidence interval excludes zero, under the fixed budget; no unacceptable increase in false positives or correction time |
|
|
127
|
+
|
|
128
|
+
Report confidence intervals on the primary effect, not only on each arm's mean.
|
|
129
|
+
Use paired task-level analyses and clustered resampling where the sampling design
|
|
130
|
+
requires it. A low-powered non-significant difference is not evidence of equivalence.
|
|
131
|
+
If the holdout cannot resolve the chosen margin, keep the route experimental rather
|
|
132
|
+
than moving the threshold after the result. Correct for multiple confirmatory route
|
|
133
|
+
comparisons or specify one primary comparison and treat the rest as exploratory.
|
|
134
|
+
|
|
135
|
+
Any actual capability escape, credential exposure, invalid acceptance across a task
|
|
136
|
+
revision, or concealed cost accounting blocks promotion regardless of mean quality.
|
|
137
|
+
Zero observed escapes in fixtures does not prove universal containment; record the
|
|
138
|
+
enforced interface and residual trust assumptions.
|
|
139
|
+
|
|
140
|
+
Promote one task class and model route at a time. Pin the accepted policy version and
|
|
141
|
+
retain a parent-only fallback. A later model/provider/harness change invalidates an
|
|
142
|
+
assumption and requires appropriate reevaluation, not automatic inheritance of gains.
|
|
143
|
+
|
|
144
|
+
## 6. Targeted ablations
|
|
145
|
+
|
|
146
|
+
Run only the ablations needed to resolve a material decision:
|
|
147
|
+
|
|
148
|
+
1. Same-context self-review versus fresh-context review with identical requirements.
|
|
149
|
+
2. Same-model worker versus an authorized alternative, keeping context and tools fixed.
|
|
150
|
+
3. One versus two workers on genuinely independent source partitions.
|
|
151
|
+
4. Selected original context versus compressed context with original-source retrieval.
|
|
152
|
+
5. Explicit incomplete results versus forced answers under missing evidence.
|
|
153
|
+
6. One diagnosed retry versus blind replay on transient and semantic failures.
|
|
154
|
+
7. Parent verification versus unchecked worker summaries, inside test fixtures only.
|
|
155
|
+
8. Warm versus cold cache and single-provider versus allowed model switching.
|
|
156
|
+
|
|
157
|
+
This separates the value of context isolation, model capability, concurrency,
|
|
158
|
+
compression and validation. A gain from one must not be attributed to another.
|
|
159
|
+
Alternative models, compression, automatic retries and live-web research are not
|
|
160
|
+
implemented routes. Ablations requiring them need separate authorized fixtures or a
|
|
161
|
+
reviewed implementation; they are not options in the shipped model-facing tool.
|
|
162
|
+
|
|
163
|
+
## 7. Runtime compatibility and security fixtures
|
|
164
|
+
|
|
165
|
+
The implementation includes deterministic fake-provider and broker fixtures. Run the
|
|
166
|
+
relevant suites and record their actual results before live inference; this checklist
|
|
167
|
+
is not itself a passing receipt. The supported calls/time contract requires coverage for:
|
|
168
|
+
|
|
169
|
+
- Native Pi package discovery and actual SDK AgentSession creation on recorded test
|
|
170
|
+
versions, including current Pi releases, with in-memory sessions and rejection of
|
|
171
|
+
missing capabilities and unsupported routes. New version identifiers alone must not
|
|
172
|
+
block activation. API presence or a CLI version check is not suite completion.
|
|
173
|
+
- Explicit parent model/thinking with Pi clamping; fresh standard ModelRuntime
|
|
174
|
+
configuration; configured global transport/thinking budgets without project settings;
|
|
175
|
+
rejection of runtime-only authentication, selected extension-provider overrides,
|
|
176
|
+
model-specific headers, startup proxy configuration and safe descriptor mismatches.
|
|
177
|
+
Parent hooks, ephemeral runtime settings
|
|
178
|
+
and session affinity are not implicitly inherited.
|
|
179
|
+
- No ambient child extensions, skills, AGENTS files or parent session history; the SDK
|
|
180
|
+
owns the agent/tool loop, with only selected-source tools exposed.
|
|
181
|
+
- Review/scout mode and benefit compatibility, required evidence, assigned requirement
|
|
182
|
+
subsets, duplicate-question rejection and parentWork required only for parallel claims.
|
|
183
|
+
- Normal Pi ownership of resource discovery, trust, proxy policy and authentication,
|
|
184
|
+
without a separate SDK host or bootstrap override.
|
|
185
|
+
- Native and legacy provider composition using synthetic credentials only; no secrets
|
|
186
|
+
returned to the extension, logged, or copied into packets.
|
|
187
|
+
- Errors represented as resolved terminal messages, setup failure and missing usage.
|
|
188
|
+
- Abort before admission, during provider setup, during inference and immediately before
|
|
189
|
+
a broker operation; non-cooperative provider settlement and late output rejection.
|
|
190
|
+
- Concurrent call-slot accounting, duplicate submissions, aggregate counters and
|
|
191
|
+
non-resetting follow-up limits; cancelled requests keep slots until settlement.
|
|
192
|
+
- Admission before every SDK model invocation, including tool continuations; provider
|
|
193
|
+
and session retries and automatic compaction disabled; requested output clamped to
|
|
194
|
+
the model maximum. Follow-up does not reset allowances.
|
|
195
|
+
- Pi-owned authentication preflight before model-invocation accounting, without
|
|
196
|
+
claiming that the inference counter bounds authentication/OAuth preparation.
|
|
197
|
+
- Oversized SDK-visible streaming responses, bounded reports/tool output and correct
|
|
198
|
+
settlement state. Hold slots through stream/result and prompt settlement; do not infer
|
|
199
|
+
physical remote termination or a raw-response/allocation guarantee from SDK events.
|
|
200
|
+
- Lost collection responses, cursor replay, idempotent follow-up and resolution,
|
|
201
|
+
stale result revisions, conflicting idempotency payloads and cancelled-job revival.
|
|
202
|
+
- Asynchronous job leases surviving normal `run` return without retaining stale
|
|
203
|
+
extension contexts, and revocation at the intended task/policy boundaries.
|
|
204
|
+
- Complete versus partial tool calls, repeated IDs, unknown tools and schema tampering.
|
|
205
|
+
- Path traversal, encoded paths, symlinks/reparse points, private files, source changes,
|
|
206
|
+
invalid line ranges and oversized input/output.
|
|
207
|
+
- Actual branch navigation versus ordinary leaf advancement; model, task, scope and
|
|
208
|
+
policy changes during queued and active work; controller and aggregate quotas
|
|
209
|
+
retained across reloads, session switches and off/on within the Pi process.
|
|
210
|
+
- Fixed canonical working root, with a Pi restart required for a new root or runtime
|
|
211
|
+
version rather than loading changed implementation code through `/reload`.
|
|
212
|
+
- Completed reports remain source-bound after the deadline; child sessions release at
|
|
213
|
+
the deadline and subsequent follow-up fails.
|
|
214
|
+
- Reused Guard instances restore exactly one state responder; invalid, declined and
|
|
215
|
+
stale-confirmation commands cannot revoke or downgrade policy.
|
|
216
|
+
- SDK errors with synthetic secret canaries never enter status, notices or tool output.
|
|
217
|
+
Throwing/rejecting teardown is contained across cancellation and deadline callbacks.
|
|
218
|
+
- Shared snapshot text survives an eligible sibling follow-up, then is destroyed;
|
|
219
|
+
failed jobs still expire inputs, and retired batches preserve quotas and replay keys.
|
|
220
|
+
- Partial usage retains known fields and per-field reporting coverage; missing fields
|
|
221
|
+
are distinct from reported zero, including unsuccessful calls.
|
|
222
|
+
- Operation-count probes cover per-event root/model checks, snapshot tool reads and
|
|
223
|
+
linear serialized-byte work. Overflow and source/model drift remain rejected at
|
|
224
|
+
protected boundaries. These probes measure implementation work, not user-visible
|
|
225
|
+
latency, production model quality, or dollar savings.
|
|
226
|
+
- Recursive source syntax checks discover new nested modules and TypeScript syntax,
|
|
227
|
+
including negative fixtures; they do not rely on a maintained filename list.
|
|
228
|
+
- No worker writes, recursion, arbitrary process execution or copied parent-tool bypass.
|
|
229
|
+
- Normal parent-tool-result retention through Pi versus in-memory worker state, accurate
|
|
230
|
+
privacy disclosure, bounded retained data, cleanup and no automatic resume.
|
|
231
|
+
|
|
232
|
+
### Local regression measurements: September 5, 2026
|
|
233
|
+
|
|
234
|
+
The review-fix probe used identical source bytes from
|
|
235
|
+
`d49fe9bc227cac480fea825932795e6db57e2127:extensions/delegation/worker.mjs`
|
|
236
|
+
as a selected fixture, comparing the old snapshot broker with the corrected broker.
|
|
237
|
+
Read used `read("s1", 1, 200)`; search used `search("a", 20)`. Returned objects were equal.
|
|
238
|
+
|
|
239
|
+
| Operation | Before: bytes serialized | After: bytes serialized | Returned JSON bytes |
|
|
240
|
+
| -------------------------- | -----------------------: | ----------------------: | ------------------: |
|
|
241
|
+
| Read 183 lines | 809,549 | 8,220 | 8,220 |
|
|
242
|
+
| Literal search, 20 matches | 25,485 | 2,741 | 2,762 |
|
|
243
|
+
|
|
244
|
+
This counts work inside `JSON.stringify`, including intermediate prefixes, not network
|
|
245
|
+
traffic. The read path now encodes each line once; search encodes each match once and
|
|
246
|
+
adds array punctuation arithmetically. Per-tool binding checks performed zero file
|
|
247
|
+
reads; full freshness checks still reread and hash the source. A synthetic 1,000-delta
|
|
248
|
+
provider regression requires fewer than 40 full lease checks and less than 100 KB of
|
|
249
|
+
serialization, while every delta still receives cheap lease and bounded-data checks.
|
|
250
|
+
The actual Pi fixture also exercises fragmented localhost SSE and transient root lookup
|
|
251
|
+
recovery. These are reproducible workload checks in the snapshot/provider/native test
|
|
252
|
+
suites, not measured UI-latency improvements or production-provider benchmarks.
|
|
253
|
+
|
|
254
|
+
### Stronger target gates: not met by the calls/time experiment
|
|
255
|
+
|
|
256
|
+
The following remain requirements before claiming the corresponding guarantees in
|
|
257
|
+
the [target protocol](design-protocol.md):
|
|
258
|
+
|
|
259
|
+
- Full parent inference-pipeline parity, including request hooks, ephemeral runtime
|
|
260
|
+
settings and session affinity. Explicit model/thinking and standard child configuration
|
|
261
|
+
do not supply that broader contract.
|
|
262
|
+
- Every underlying inference attempt, including transport fallback or an internal SDK
|
|
263
|
+
retry, receives pre-dispatch admission; opaque attempts are rejected and initial
|
|
264
|
+
dispatch is not double-debited.
|
|
265
|
+
- Raw text, reasoning, framing, compressed transport and tool arguments are bounded
|
|
266
|
+
before parsing or allocation, including decompression and transport buffering.
|
|
267
|
+
- Monetary reservations account conservatively for provider pricing, uncertain dispatch,
|
|
268
|
+
retries and settlement, with honest unknown-cost behavior.
|
|
269
|
+
|
|
270
|
+
The current runtime disables configurable retries, counts model invocations, enforces
|
|
271
|
+
logical deadlines, validates inputs before dispatch, counts observed stream deltas and
|
|
272
|
+
validates complete parsed responses at protected boundaries. It reports cost as unavailable. These controls do
|
|
273
|
+
not pass the stronger gates: adapters may buffer the whole response or make transport
|
|
274
|
+
attempts before SpecPi observes completion. The narrower experiment must reject policies requiring unsupported
|
|
275
|
+
guarantees. Its fixture tests cannot be presented as transport, invoice or process-memory
|
|
276
|
+
proof. The unimplemented recovery/retry policy also needs separate fault-class tests.
|
|
277
|
+
|
|
278
|
+
For the later web adapter, add connection-time public-address checks, redirects,
|
|
279
|
+
rebinding defenses, credential/header stripping, oversized responses, malformed text,
|
|
280
|
+
page instruction injection and blocked private-network destinations. Test the actual
|
|
281
|
+
transport boundary; a validator unit test alone does not prove connection behavior.
|
|
282
|
+
|
|
283
|
+
All installer, package and Pi integration tests use fresh temporary state and skip
|
|
284
|
+
unnecessary external installation. A live provider smoke test is a separate explicit
|
|
285
|
+
authorization with a harmless prompt and a bounded cost envelope. Fake tests cannot
|
|
286
|
+
prove production provider parity or network behavior.
|
|
287
|
+
|
|
288
|
+
## 8. Shipping gate and rollback
|
|
289
|
+
|
|
290
|
+
Before release, inspect the complete diff, run focused tests and the full repository
|
|
291
|
+
check, exercise the packaged artifact, and obtain fresh independent review of provider,
|
|
292
|
+
permission, cancellation, retention and dependency changes. Update third-party and
|
|
293
|
+
security documentation to describe the contracts actually implemented. A host bridge
|
|
294
|
+
that does not exist cannot be replaced with a private-field workaround just to pass
|
|
295
|
+
the release gate.
|
|
296
|
+
|
|
297
|
+
Disabling the extension must stop new calls and revoke broker grants immediately while
|
|
298
|
+
reporting requests still settling. Uninstall must not require restoring provider
|
|
299
|
+
configuration or authentication files because the package never owns them. Explicit
|
|
300
|
+
user-exported evidence remains user data. Policy rollback returns routing to the last
|
|
301
|
+
validated version or parent-only operation; it never discards user changes or retries
|
|
302
|
+
abandoned work automatically.
|
|
303
|
+
|
|
304
|
+
An experimental implementation with passing fixtures establishes only its tested
|
|
305
|
+
contract. Promotion beyond experimental status requires repeatable value on a defined
|
|
306
|
+
task class and understandable operating boundaries. Unimplemented capabilities remain
|
|
307
|
+
absent, not represented as partially functioning options.
|
|
@@ -0,0 +1,271 @@
|
|
|
1
|
+
# Delegation protocol: bounded-pi-sessions-v1
|
|
2
|
+
|
|
3
|
+
This is the implemented in-process API. It has no HTTP listener, daemon, child process,
|
|
4
|
+
or child session store. The broader [target protocol](design-protocol.md) remains a
|
|
5
|
+
proposal; its stronger transport/attempt/cost gates are not supplied by this version.
|
|
6
|
+
|
|
7
|
+
The extension loads through normal `pi` package discovery and remains disabled until
|
|
8
|
+
the human runs `/delegate on`. Compatibility is checked through required public SDK
|
|
9
|
+
capabilities; there is no exact-version allowlist. Missing APIs prevent activation and
|
|
10
|
+
are named in the error. The runtime also verifies the created session's thinking,
|
|
11
|
+
tools and streaming interface. Tested versions are evidence, not an activation gate.
|
|
12
|
+
Workers are SDK `createAgentSession` instances using in-memory session storage and a
|
|
13
|
+
fresh Pi `ModelRuntime`. The parent model and thinking level are passed explicitly,
|
|
14
|
+
subject to Pi's clamping. Standard Pi authentication, environment and `models.json`
|
|
15
|
+
resolution apply. Child transport/thinking budgets come from configured global settings;
|
|
16
|
+
project settings are not loaded. Runtime-only authentication, selected extension-provider
|
|
17
|
+
overrides, model-specific headers, startup proxy configuration and safe model-descriptor
|
|
18
|
+
mismatches fail preflight. Parent request hooks,
|
|
19
|
+
ephemeral runtime settings, session affinity and ambient resources are not inherited.
|
|
20
|
+
|
|
21
|
+
Command Guard is optional. Absent and Off states permit human activation; an installed
|
|
22
|
+
Guard's Strict approvals and explicit locks remain enforced. Unready or duplicate
|
|
23
|
+
Guard responders prevent activation with a specific error. Guard state changes revoke
|
|
24
|
+
the current delegation generation. Snapshot tools and resource limits are enforced
|
|
25
|
+
by delegation itself in every mode. Status includes the observed `guard` state.
|
|
26
|
+
|
|
27
|
+
The SDK runs the conversation and selected-source tool loop. SpecPi admits each SDK
|
|
28
|
+
invocation before dispatch and observes the SDK stream; provider/session retries and
|
|
29
|
+
automatic compaction are disabled. This does not establish hard raw-transport,
|
|
30
|
+
hidden-provider-attempt, invoice or process-memory limits.
|
|
31
|
+
Pi authentication preflight precedes the model-invocation counter; its preparation
|
|
32
|
+
is not a model invocation or an operation bounded by that counter.
|
|
33
|
+
|
|
34
|
+
Every object is closed: unknown fields, duplicate IDs, malformed values and oversized
|
|
35
|
+
data are rejected. The host creates identities and receipts; workers cannot supply them.
|
|
36
|
+
The protocol identifier `bounded-pi-sessions-v1` and inference contract
|
|
37
|
+
`pi-agent-session-v1` describe the host implementation, not model-selected options.
|
|
38
|
+
|
|
39
|
+
## Submit a batch
|
|
40
|
+
|
|
41
|
+
After the human runs `/delegate on`, the parent calls the `delegate` tool:
|
|
42
|
+
|
|
43
|
+
```json
|
|
44
|
+
{
|
|
45
|
+
"operation": "run",
|
|
46
|
+
"requestId": "scout-routing-1",
|
|
47
|
+
"packet": {
|
|
48
|
+
"objective": "Explain how model selection reaches the request pipeline",
|
|
49
|
+
"requirements": [{ "id": "R1", "text": "Identify the route binding and its invalidation behavior" }],
|
|
50
|
+
"decisions": ["The parent remains the sole writer"],
|
|
51
|
+
"nonGoals": ["Do not implement or change providers"],
|
|
52
|
+
"reason": {
|
|
53
|
+
"benefit": "parallel_analysis",
|
|
54
|
+
"why": "Source analysis can proceed independently of the parent's lifecycle-test inspection",
|
|
55
|
+
"parentWork": "Inspect the lifecycle tests while the worker reads"
|
|
56
|
+
},
|
|
57
|
+
"jobs": [
|
|
58
|
+
{
|
|
59
|
+
"id": "route",
|
|
60
|
+
"mode": "scout",
|
|
61
|
+
"requirements": ["R1"],
|
|
62
|
+
"question": "Where is the model captured, and when is that capability revoked?",
|
|
63
|
+
"context": "Return evidence for R1, including missing or contrary evidence.",
|
|
64
|
+
"sources": ["extensions/delegation/provider.mjs"]
|
|
65
|
+
}
|
|
66
|
+
]
|
|
67
|
+
}
|
|
68
|
+
}
|
|
69
|
+
```
|
|
70
|
+
|
|
71
|
+
`run` returns immediately with `batchId`, `packetDigest`, `generation`, a collection
|
|
72
|
+
cursor and job states. It does not return invented findings while work is pending.
|
|
73
|
+
The digest binds the packet, source descriptors, host identity, selected model/provider
|
|
74
|
+
IDs and resource policy. It does not freeze the registry's endpoint, headers or
|
|
75
|
+
authentication configuration. Those can change behind the same public identities;
|
|
76
|
+
the extension cannot observe every such change or certify an unchanged provider route.
|
|
77
|
+
IDs are short alphanumeric/hyphen/underscore strings. The host batch and attempt IDs
|
|
78
|
+
are UUIDs. Source paths are exact relative filenames, not directories, globs or commands.
|
|
79
|
+
Modes are `review` and `scout`. Either can use the selected-source list/read/literal-search
|
|
80
|
+
tools when files are supplied. A scout requires at least one selected file. A review
|
|
81
|
+
requires nonempty inline context or selected files; with an empty `sources` array it
|
|
82
|
+
uses inline context without tools. Source selection never grants ambient filesystem
|
|
83
|
+
or web access. Questions that are identical after trimming and case normalization are
|
|
84
|
+
rejected; distinct text does not establish distinct reasoning work.
|
|
85
|
+
|
|
86
|
+
`reason` contains exactly `benefit`, `why` and `parentWork`. `why` is nonempty.
|
|
87
|
+
`independent_review` requires review jobs; `parallel_analysis` and `context_isolation`
|
|
88
|
+
require scout jobs. `parentWork` must describe useful concurrent work for
|
|
89
|
+
`parallel_analysis` and may be empty otherwise. These are structural admission checks,
|
|
90
|
+
not proof that delegation improves the task.
|
|
91
|
+
|
|
92
|
+
Each job's nonempty `requirements` list names unique IDs from the packet's global
|
|
93
|
+
requirements. The child receives only its assigned requirements, plus the global
|
|
94
|
+
decisions and non-goals. A batch contains at most two jobs. All modes require the same
|
|
95
|
+
explicit packet fields, even when an allowed array or string is empty.
|
|
96
|
+
|
|
97
|
+
## Inspect and collect
|
|
98
|
+
|
|
99
|
+
Worker `list_sources` accepts an optional zero-based `offset` and returns
|
|
100
|
+
`{ "sources": [...], "nextOffset": number | null }`. Each page is at most 16 KiB,
|
|
101
|
+
including JSON metadata, and fits the remaining tool-byte allowance. A non-null
|
|
102
|
+
`nextOffset` identifies the next page; null marks the end. Pages consume the same
|
|
103
|
+
per-job tool-call and byte budgets as reads and searches.
|
|
104
|
+
|
|
105
|
+
```json
|
|
106
|
+
{ "operation": "status" }
|
|
107
|
+
```
|
|
108
|
+
|
|
109
|
+
```json
|
|
110
|
+
{ "operation": "collect", "batchId": "HOST_BATCH_ID", "afterCursor": 0, "waitMs": 30000 }
|
|
111
|
+
```
|
|
112
|
+
|
|
113
|
+
`status` is compact and available while disabled. `collect` returns reports whose
|
|
114
|
+
completion cursor is newer than `afterCursor`; omitting it replays retained reports.
|
|
115
|
+
An accepted next batch retires the previous batch's reports. Old generations retire
|
|
116
|
+
after their workers settle. `status` retains bounded summaries marked `retired: true`,
|
|
117
|
+
but retired batches reject collect and follow-up. Quotas and request fingerprints survive
|
|
118
|
+
retirement, so replay cannot reopen a batch or replenish the process allowance.
|
|
119
|
+
`waitMs` defaults to zero and cannot exceed 30 seconds. There is no destructive dequeue
|
|
120
|
+
and no automatic parent turn. A cursor ahead of the batch is invalid.
|
|
121
|
+
|
|
122
|
+
Collection rechecks host/task/policy generation and every selected source binding.
|
|
123
|
+
A changed file is stale even when its pathname is unchanged. Source or lifecycle changes
|
|
124
|
+
cannot be repaired by presenting an old receipt or idempotency key.
|
|
125
|
+
|
|
126
|
+
Each returned item has `receipt`, `result`, `error`, and `disposition`. A host receipt
|
|
127
|
+
contains `batchId`, `jobId`, `attemptId`, `packetDigest`, `generation`, `resultRevision`,
|
|
128
|
+
`model`, `state`, `settling`, call/tool counters, token usage, `usageReportedCalls`, `usageComplete`, and
|
|
129
|
+
`cost: null`. Copy the six binding fields when following up or resolving. The
|
|
130
|
+
`collectionCursor` belongs to delivery ordering, not result identity.
|
|
131
|
+
Each usage field sums its valid reported values independently; never-reported fields
|
|
132
|
+
are `null`. `usageReportedCalls` counts reports for each field. A positive subtotal may
|
|
133
|
+
still be incomplete; `usageComplete` requires all four fields for every admitted call.
|
|
134
|
+
|
|
135
|
+
## Worker report
|
|
136
|
+
|
|
137
|
+
Workers return this exact shape, without Markdown fences:
|
|
138
|
+
|
|
139
|
+
```json
|
|
140
|
+
{
|
|
141
|
+
"status": "partial",
|
|
142
|
+
"answer": "The selected context is insufficient to establish runtime behavior.",
|
|
143
|
+
"requirements": [{ "id": "R1", "status": "unaddressed", "evidence": [] }],
|
|
144
|
+
"findings": [],
|
|
145
|
+
"missing": ["A provider implementation or runtime fixture"],
|
|
146
|
+
"nextStep": "The parent should inspect the provider fixture before making a claim."
|
|
147
|
+
}
|
|
148
|
+
```
|
|
149
|
+
|
|
150
|
+
Statuses are `complete`, `partial`, or `needs_context`. Every assigned requirement appears
|
|
151
|
+
exactly once as `addressed` or `unaddressed`. Each finding has `id`, `claim`, `confidence`
|
|
152
|
+
(`observed`, `inferred`, `unverified`), `evidence` and `contraryEvidence`. Evidence entries
|
|
153
|
+
contain exactly `sourceId`, `lineStart`, `lineEnd`. References must resolve within the
|
|
154
|
+
job selection and valid line ranges. `p1` denotes the submitted inline context; it
|
|
155
|
+
does not denote the parent transcript or prove repository behavior. `observed` requires
|
|
156
|
+
at least one reference. At most eight findings and 16 KiB are retained.
|
|
157
|
+
|
|
158
|
+
Malformed reports fail the attempt. There is no automatic retry. Provider exceptions
|
|
159
|
+
are reduced to a generic failure message because raw errors may contain sensitive URLs
|
|
160
|
+
or content. Missing usage is explicit; failed calls still consume invocation allowance.
|
|
161
|
+
|
|
162
|
+
## Follow up and resolve
|
|
163
|
+
|
|
164
|
+
```json
|
|
165
|
+
{
|
|
166
|
+
"operation": "follow_up",
|
|
167
|
+
"requestId": "route-correction-1",
|
|
168
|
+
"batchId": "HOST_BATCH_ID",
|
|
169
|
+
"jobId": "route",
|
|
170
|
+
"attemptId": "HOST_ATTEMPT_ID",
|
|
171
|
+
"packetDigest": "HOST_64_CHARACTER_SHA256",
|
|
172
|
+
"generation": 2,
|
|
173
|
+
"resultRevision": 1,
|
|
174
|
+
"prompt": "Reconsider the claim using the already-selected source; identify the missing condition."
|
|
175
|
+
}
|
|
176
|
+
```
|
|
177
|
+
|
|
178
|
+
The uppercase placeholders above must be replaced with the current host receipt.
|
|
179
|
+
One follow-up creates a new attempt under the original job deadline, context,
|
|
180
|
+
selected sources, model-call counters and tool counters. It cannot add files, change
|
|
181
|
+
models, extend time or reset quotas. If it needs a new grant or source set, discard
|
|
182
|
+
the report and submit a fresh batch. A settling, cancelled, expired, stale or finally
|
|
183
|
+
disposed result cannot receive a follow-up.
|
|
184
|
+
|
|
185
|
+
```json
|
|
186
|
+
{
|
|
187
|
+
"operation": "resolve",
|
|
188
|
+
"requestId": "route-assessment-1",
|
|
189
|
+
"batchId": "HOST_BATCH_ID",
|
|
190
|
+
"jobId": "route",
|
|
191
|
+
"attemptId": "HOST_ATTEMPT_ID",
|
|
192
|
+
"packetDigest": "HOST_64_CHARACTER_SHA256",
|
|
193
|
+
"generation": 2,
|
|
194
|
+
"resultRevision": 1,
|
|
195
|
+
"decision": "needs_check",
|
|
196
|
+
"findings": []
|
|
197
|
+
}
|
|
198
|
+
```
|
|
199
|
+
|
|
200
|
+
Overall decisions are `accept`, `discard`, `needs_check`. Each returned finding must
|
|
201
|
+
appear exactly once in `findings` with `{ "id": "F1", "decision": "confirmed" }`,
|
|
202
|
+
`rejected`, or `needs_check`. Acceptance requires a result and no unchecked findings.
|
|
203
|
+
Accept/discard finalize the parent disposition; needs_check leaves it open for a later
|
|
204
|
+
assessment or allowed follow-up. These are parent assertions, never human permission,
|
|
205
|
+
automatic tool execution, verified completion, or wishlist authorization.
|
|
206
|
+
|
|
207
|
+
All mutations require `requestId`. Successful runs, follow-ups and final dispositions
|
|
208
|
+
retain their replay receipts for the process lifetime; the fixed batch/job ceilings
|
|
209
|
+
bound these to at most 20 entries. Replaying a retained request returns its stored
|
|
210
|
+
response without another request or transition. Reusing a retained ID with a different
|
|
211
|
+
payload fails. Failed requests do not reserve IDs and may be corrected or retried.
|
|
212
|
+
Successful cancellation and `needs_check` responses use a separate 128-entry oldest-first
|
|
213
|
+
cache. After eviction, those operations are revalidated against current state; they
|
|
214
|
+
cannot start inference or restore cancelled jobs. Generation and source checks still
|
|
215
|
+
apply. Neither cache eviction nor failed attempts reset quotas or block cancellation.
|
|
216
|
+
|
|
217
|
+
## Cancellation and lifecycle
|
|
218
|
+
|
|
219
|
+
```json
|
|
220
|
+
{ "operation": "cancel", "requestId": "cancel-route-1", "batchId": "HOST_BATCH_ID", "jobId": "route" }
|
|
221
|
+
```
|
|
222
|
+
|
|
223
|
+
Omit `jobId` to cancel the batch. Cancellation remains available after invalidation.
|
|
224
|
+
Job states are `queued`, `running`, the three report statuses, `failed`, `cancelled`,
|
|
225
|
+
`expired`, and `stale`. Terminal job state and `settling` are separate. Cancellation
|
|
226
|
+
revokes tool access and requests SDK abort. A worker keeps its global slot through
|
|
227
|
+
SDK-visible stream/result and prompt settlement. Old payloads are discarded. Neither
|
|
228
|
+
terminal receipt delivery nor SDK settlement proves physical remote execution has ended.
|
|
229
|
+
|
|
230
|
+
The same in-memory controller survives `/reload` and session switches within the Pi
|
|
231
|
+
process. Its ceilings are two active workers, four batches and 32 SDK model invocations
|
|
232
|
+
per process, with two jobs and 8 invocations per batch and four invocations per logical
|
|
233
|
+
job including one follow-up. Requested output is 8,192 tokens clamped to the model
|
|
234
|
+
maximum. These are experiment limits, not research-derived optimal values.
|
|
235
|
+
Human off/on, task changes,
|
|
236
|
+
branch navigation, model selection, guard changes and reloads revoke old generations;
|
|
237
|
+
they do not create a new resource allowance. Normal parent turns do not revoke a job.
|
|
238
|
+
Model and thinking selections retain the human's activation choice. The extension
|
|
239
|
+
preflights the latest selected host and resumes dispatch automatically, without
|
|
240
|
+
replaying old jobs. Unsupported selections pause dispatch and report `pauseReason`;
|
|
241
|
+
a compatible selection resumes it. Status exposes `requested`, `updating` and
|
|
242
|
+
`pauseReason` alongside the controller's `enabled` flag. Concurrent notifications
|
|
243
|
+
for the same host share setup, and a late setup cannot overwrite a newer selection,
|
|
244
|
+
explicit off command, Guard revocation or session change. A turn-start refresh also
|
|
245
|
+
catches a changed host before the next parent turn.
|
|
246
|
+
The canonical working root remains fixed for that process. Restart Pi to change the
|
|
247
|
+
root or load a new delegation runtime version. There is no retry on process restart
|
|
248
|
+
and no durable worker queue. Completed reports retain their original source bindings
|
|
249
|
+
after the deadline; child sessions are released at the deadline and subsequent
|
|
250
|
+
follow-up is rejected. Snapshot text is destroyed when every job loses continuation
|
|
251
|
+
eligibility. Failed first attempts retain the original expiry timer. After settlement,
|
|
252
|
+
packet and job-input references are dropped; only source metadata/digests remain to
|
|
253
|
+
validate completed reports until retirement.
|
|
254
|
+
|
|
255
|
+
Per-event lease checks use model/context identities and generation. Full canonical-root,
|
|
256
|
+
safe model-descriptor and provider-policy checks run at request, tool and publication
|
|
257
|
+
boundaries. Snapshot tools check canonical/stat bindings; full digest checks run at
|
|
258
|
+
capture, publication, collect, follow-up and resolve. These cheaper checks assume a
|
|
259
|
+
trusted local filesystem and Pi's parsed stream contract. Streaming uses incremental
|
|
260
|
+
recognized-delta accounting and bounded structure checks, with exact response checks
|
|
261
|
+
at content/terminal boundaries. It does not promise an exact per-event size bound for
|
|
262
|
+
inconsistent SDK partial objects. Bytes may already be allocated before an event.
|
|
263
|
+
Per invocation the parser permits 64 content blocks, 512 structural nodes per partial,
|
|
264
|
+
65,536 events and 130 non-delta boundaries. These fixed engineering ceilings also bound
|
|
265
|
+
repeated whole-message validation for malformed event sequences.
|
|
266
|
+
Unexpected SDK errors are redacted before status/tool/UI output; teardown failures
|
|
267
|
+
cannot escape timer callbacks or reset settling ownership.
|
|
268
|
+
|
|
269
|
+
The enforced [resource envelope and limitations](README.md#enforced-resource-envelope)
|
|
270
|
+
define `bounded-pi-sessions-v1`. Unsupported hard billing, raw transport, provider
|
|
271
|
+
attempt and process-memory policies are not accepted through this API.
|