@neruok/pi-advisor 0.0.0-stage → 0.1.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/LICENSE +21 -0
- package/README.md +310 -2
- package/advisor.ts +75 -0
- package/docs/pi-advisor.md +1555 -0
- package/lib/completions.ts +88 -0
- package/lib/context.ts +17 -0
- package/lib/model-picker.ts +85 -0
- package/lib/prompt.ts +21 -0
- package/lib/protocol.ts +131 -0
- package/lib/reasoning.ts +9 -0
- package/lib/render.ts +121 -0
- package/lib/replay.ts +55 -0
- package/lib/schemas.ts +36 -0
- package/lib/session.ts +193 -0
- package/lib/settings-command.ts +105 -0
- package/lib/settings.ts +99 -0
- package/package.json +40 -4
|
@@ -0,0 +1,1555 @@
|
|
|
1
|
+
<!-- generated by workspace-docs. source: pi-advisor revision 41. do not edit; author with workspace-docs checkout. -->
|
|
2
|
+
|
|
3
|
+
# Pi advisor: isolated consultation contract and design
|
|
4
|
+
|
|
5
|
+
## Contents
|
|
6
|
+
|
|
7
|
+
- [Purpose and scope](#scope)
|
|
8
|
+
- [Design and implementation choices](#design)
|
|
9
|
+
- [Tool and lifecycle contract](#contract)
|
|
10
|
+
- [Configuration contract](#configuration)
|
|
11
|
+
- [Host integration and results](#integration)
|
|
12
|
+
- [Acceptance criteria](#acceptance)
|
|
13
|
+
- [Verification scope](#verification)
|
|
14
|
+
- [Model picker correction](#model-picker-verification)
|
|
15
|
+
- [Diagnostics, usage, and presentation improvements](#diagnostics-verification)
|
|
16
|
+
- [Live rejection investigation and safe diagnostics](#reply-diagnostics-verification)
|
|
17
|
+
- [Remove the model output protocol](#plain-advice-verification)
|
|
18
|
+
- [Full responses visible by default](#response-visibility-verification)
|
|
19
|
+
- [Collapse long replies after the display demonstration](#long-reply-preview-verification)
|
|
20
|
+
- [Reasoning effort and command autocomplete](#reasoning-autocomplete-verification)
|
|
21
|
+
- [Opt-in provider failure diagnostics (issue #1)](#provider-diagnostics-verification)
|
|
22
|
+
- [SDK timeout defect and terminal error hints](#sdk-errors-verification)
|
|
23
|
+
- [xAI HTTP 400 investigation and tool-free request correction](#xai-no-tools-verification)
|
|
24
|
+
- [Remove the tool-choice override for every provider](#provider-neutral-options-verification)
|
|
25
|
+
- [Configurable deadline verification](#deadline-configuration-verification)
|
|
26
|
+
- [Short provider caching verification](#cache-verification)
|
|
27
|
+
- [Private continuation-state verification](#opaque-replay-verification)
|
|
28
|
+
- [Model context and display verification](#context-display-verification)
|
|
29
|
+
- [Fixed consultation cap removal](#fixed-cap-removal-verification)
|
|
30
|
+
|
|
31
|
+
|
|
32
|
+
## Purpose and scope {#scope}
|
|
33
|
+
|
|
34
|
+
Pi advisor gives the main agent a separate reasoning partner. The main agent explains the problem and supplies evidence across multiple turns.
|
|
35
|
+
|
|
36
|
+
The advisor cannot inspect code, run tools, or inherit the parent conversation. The main agent owns investigation, verification, implementation, and authorization.
|
|
37
|
+
|
|
38
|
+
The user requested a standalone extension, with a specification before implementation. This draft records the initial implementation contract, not separate policy approval.
|
|
39
|
+
|
|
40
|
+
V1 excludes persistence, restoration, automatic summarization, forks, autonomous consultation loops, workspace collection, custom providers, and a dependency on pi-magic8ball.
|
|
41
|
+
|
|
42
|
+
## Design and implementation choices {#design}
|
|
43
|
+
|
|
44
|
+
Use one process-local consultation manager per extension instance. Call the host model registry directly with an explicit tool-free context. Do not create a nested Pi agent runtime.
|
|
45
|
+
|
|
46
|
+
Consultations remain available across parent turns and compaction. Clear them on session start, tree navigation, and shutdown. Switch, fork, and reload replace the runtime. Do not restore consultations from tool results.
|
|
47
|
+
|
|
48
|
+
The extension writes no consultation files. Pi can still persist tool arguments and results in its parent transcript. The configured provider receives explicit consultation messages. Ephemeral means no separate extension persistence, not guaranteed erasure by Pi or the provider.
|
|
49
|
+
|
|
50
|
+
REQ-31 removes the fixed consultation, exchange, message-byte, reply-byte, and output-token caps at the user's request.
|
|
51
|
+
REQ-29 retains the pinned model context window. REQ-26 retains the configurable deadline, with a 300000 ms default.
|
|
52
|
+
Pi controls the model output allowance. Settings-file and presentation bounds remain separate from consultation capacity.
|
|
53
|
+
These controls do not provide a monetary guarantee.
|
|
54
|
+
|
|
55
|
+
Settings use the host agent directory, not a hard-coded home path. Project settings require current Pi project trust. No default model is selected. No live provider calls are part of build verification.
|
|
56
|
+
|
|
57
|
+
## Tool and lifecycle contract {#contract}
|
|
58
|
+
|
|
59
|
+
<a id="REQ-1"></a>
|
|
60
|
+
|
|
61
|
+
**REQ-1** Isolated multi-turn consultation
|
|
62
|
+
|
|
63
|
+
The extension MUST expose advisor({message, session?}). Omitted session starts a new consultation. Present session continues only that consultation.
|
|
64
|
+
Message MUST be a nonblank string. Session MUST be a nonempty opaque identifier when present. Null and unknown fields MUST fail validation.
|
|
65
|
+
On successful creation, the extension MUST pin the explicitly configured physical chat model. Continuations MUST use that selection despite later configuration changes.
|
|
66
|
+
Every model request MUST contain only the fixed advisor system prompt and that consultation's explicit exchanges. It MUST declare zero tools.
|
|
67
|
+
The extension MUST NOT read parent messages, instructions, source files, environment model selections, or another consultation's exchanges into the request.
|
|
68
|
+
|
|
69
|
+
|
|
70
|
+
<a id="REQ-2"></a>
|
|
71
|
+
|
|
72
|
+
**REQ-2** Advice and output protocol
|
|
73
|
+
|
|
74
|
+
The advisor prompt MUST distinguish reported observations, interpretations, assumptions, and hypotheses. It MUST permit questions and challenges to the main agent's framing.
|
|
75
|
+
The prompt MUST state that advice is not evidence or authorization. It MUST ask the main agent to gather missing evidence.
|
|
76
|
+
REQ-17 defines the unstructured reply contract and replaces the former marker and status requirements.
|
|
77
|
+
Tool calls, nontext/nonthinking blocks, and non-stop completion reasons MUST fail.
|
|
78
|
+
Thinking content MUST NOT enter the returned response or retained transcript. Advice MUST NOT close the consultation or grant permission to act.
|
|
79
|
+
|
|
80
|
+
|
|
81
|
+
<a id="REQ-3"></a>
|
|
82
|
+
|
|
83
|
+
**REQ-3** Atomic exchanges and failures
|
|
84
|
+
|
|
85
|
+
Each advisor call MUST make at most one model request. Provider retries MUST be disabled. Automatic fallback MUST NOT occur.
|
|
86
|
+
Only a validated reply MUST commit its user/reply pair. Failed creation MUST leave no consultation. Failed continuation MUST preserve previous exchanges.
|
|
87
|
+
The extension MUST combine caller cancellation with the deadline defined by REQ-26, measured from call entry. Cancellation MUST prevent a later commit.
|
|
88
|
+
Timeout and cancellation MUST return without waiting indefinitely for a provider that ignores abort. Late completions MUST NOT change state.
|
|
89
|
+
Concurrent calls on one consultation MUST reject the second with busy. Different consultations MAY run concurrently. REQ-31 removes the active consultation cap.
|
|
90
|
+
Failures MUST use fixed safe error codes and messages. Raw provider exceptions MUST NOT appear in output.
|
|
91
|
+
Results MUST report usage available at return, including rejected completed replies. Usage from a late unobserved completion MAY be unavailable.
|
|
92
|
+
|
|
93
|
+
|
|
94
|
+
<a id="REQ-4"></a>
|
|
95
|
+
|
|
96
|
+
**REQ-4** Resource limits
|
|
97
|
+
|
|
98
|
+
The extension MUST enforce inclusive hard limits: 8 active consultations, 24 committed exchange pairs per consultation, and 16384 UTF-8 bytes per message or raw reply text.
|
|
99
|
+
Retained transcript bytes MUST equal the UTF-8 size of JSON.stringify([{role,text}, ...]) and MUST NOT exceed 49152 bytes.
|
|
100
|
+
The pending user entry MUST pass byte validation before model work. The full pair MUST pass validation before commit.
|
|
101
|
+
A request MUST specify maxTokens = 4096. Token limits remain subject to provider support. Byte and turn limits MUST hold locally.
|
|
102
|
+
Crossing a limit MUST return limit-exceeded without changing the prior transcript. The extension MUST NOT evict, trim, compact, or summarize implicitly.
|
|
103
|
+
Closed consultations MUST release their slot. Limits reset for a new consultation or a cleared parent lifecycle.
|
|
104
|
+
|
|
105
|
+
|
|
106
|
+
<a id="REQ-5"></a>
|
|
107
|
+
|
|
108
|
+
**REQ-5** Discovery, close, and reset
|
|
109
|
+
|
|
110
|
+
advisor_sessions({}) MUST return metadata without full transcripts. REQ-13 defines the fields in addition to session, label, turns, model, and busy.
|
|
111
|
+
Label MUST derive from the first explicit message, collapse whitespace, and contain at most 80 Unicode code points. No full transcript MUST appear.
|
|
112
|
+
advisor_close({session}) MUST remove an idle consultation and return its identifier. Unknown identifiers MUST return not-found. Busy close MUST return busy.
|
|
113
|
+
A cleared manager MUST abort all pending work and reject late commits. Old identifiers MUST become not-found.
|
|
114
|
+
The extension MUST clear state on session_start, session_tree, and session_shutdown. Pi 1.0.4 replaces the runtime on session switch and fork through shutdown/start.
|
|
115
|
+
Reload MUST create fresh state. Parent compaction MUST NOT clear consultations. State MUST NOT be reconstructed from persisted entries.
|
|
116
|
+
|
|
117
|
+
|
|
118
|
+
<a id="REQ-17"></a>
|
|
119
|
+
|
|
120
|
+
**REQ-17** Unstructured advisor text
|
|
121
|
+
|
|
122
|
+
The advisor prompt MUST request natural-language advice without required markers, JSON, or status labels.
|
|
123
|
+
After REQ-16 validation, the extension MUST join text blocks with a newline and return that text unchanged in response. It MUST retain the same text in consultation history. It MUST NOT trim nonblank text, strip marker-looking text, parse advice as JSON, classify advice, or infer a status.
|
|
124
|
+
Marker-looking strings, quotations, lists, code fences, and JSON-looking text MUST be ordinary reply content. Their presence, absence, repetition, or placement MUST NOT affect acceptance.
|
|
125
|
+
The extension MUST remove status from successful advisor results, session metadata, session state, schemas, and presentation. It MUST NOT substitute a guessed or constant classification field.
|
|
126
|
+
Structured tool envelopes MUST remain extension-generated and schema-validated. Model selection, isolation, thinking removal, provider completion validation, byte/turn/session bounds, usage, cancellation, rollback, and no automatic retry MUST remain unchanged.
|
|
127
|
+
|
|
128
|
+
|
|
129
|
+
## Configuration contract {#configuration}
|
|
130
|
+
|
|
131
|
+
<a id="REQ-6"></a>
|
|
132
|
+
|
|
133
|
+
**REQ-6** Explicit trusted settings
|
|
134
|
+
|
|
135
|
+
Settings MUST be strict JSON objects with optional model: {provider, model, reasoning?} and optional root timeoutMs.
|
|
136
|
+
REQ-20 defines reasoning. REQ-26 defines timeout validation and independent timeout precedence.
|
|
137
|
+
Both identifiers MUST be nonblank strings without control characters or whitespace.
|
|
138
|
+
Missing files and {} MUST mean unconfigured. Malformed, oversized, null, or unknown-field settings MUST fail closed.
|
|
139
|
+
The extension MUST read <agent-dir>/advisor.json first. Only a trusted project MAY add <cwd>/.pi/advisor.json.
|
|
140
|
+
A project model MUST replace the entire global model pair. An empty project object MUST retain the global pair. Invalid global settings MUST fail even with a valid project override.
|
|
141
|
+
Reads MUST accept only bounded regular UTF-8 files, at most 16384 bytes. Immediate parent or target symlinks and special files MUST be rejected.
|
|
142
|
+
A new consultation MUST fail with not-configured or model-unavailable when selection or authentication is unavailable. Virtual models MUST be rejected.
|
|
143
|
+
Normal Pi authentication MUST serve model requests. The extension MUST NOT read credentials or change the parent's active model.
|
|
144
|
+
|
|
145
|
+
|
|
146
|
+
<a id="REQ-7"></a>
|
|
147
|
+
|
|
148
|
+
**REQ-7** User-only settings command
|
|
149
|
+
|
|
150
|
+
The extension MUST register /advisor [--global|--project] [show | model <provider> <model> | reasoning [level] | timeout [milliseconds|default]].
|
|
151
|
+
Global MUST be the default save scope. REQ-20 and REQ-21 define reasoning commands and argument autocomplete.
|
|
152
|
+
In TUI mode, a bare command MUST use the inline searchable picker defined by REQ-10. Other UI modes MUST retain the standard selection dialog.
|
|
153
|
+
Without UI, it MUST emit settings and usage as a Pi custom message.
|
|
154
|
+
Every non-UI command report MUST use customType advisor-settings, display: true, and triggerTurn: false. No report MUST start a model turn.
|
|
155
|
+
show MUST display effective selection and source. Invalid commands or cancelled selection MUST write nothing and call no model.
|
|
156
|
+
An explicit pair MUST resolve to a physical chat model before save. Project saves MUST recheck trust.
|
|
157
|
+
Saves MUST touch only advisor.json and owned temporary or lock files. They MUST serialize participating writers and acquire an exclusive lock.
|
|
158
|
+
Replacement MUST be atomic through a same-directory temporary file. First creation MUST NOT replace a newly appeared file.
|
|
159
|
+
At the final comparison, changed bytes or lock contention MUST fail closed. No stale-lock reclamation MUST occur.
|
|
160
|
+
Cleanup MUST remove only owned temporary files and locks. A save error MUST warn that a save may have completed and require inspection before retry.
|
|
161
|
+
|
|
162
|
+
|
|
163
|
+
<a id="REQ-10"></a>
|
|
164
|
+
|
|
165
|
+
**REQ-10** Model picker matches /model
|
|
166
|
+
|
|
167
|
+
The TUI picker MUST copy the inline interface pattern used by pi-magic8ball and Pi 1.0.4 /model.
|
|
168
|
+
It MUST show at most 10 model rows and at most max(1, terminal.rows - 2) total rows. Every rendered line MUST fit its terminal width.
|
|
169
|
+
It MUST provide editable fuzzy search across provider, model identifier, and display name, using Pi fuzzyFilter ranking.
|
|
170
|
+
Every nonempty query edit MUST select the best match. Text keys, including j and k, MUST edit the input.
|
|
171
|
+
Arrow keys MUST wrap through filtered results. The effective configured advisor model MUST appear first before search with a current-selection marker.
|
|
172
|
+
When at least 10 total rows fit, the picker MUST show the selected display name and navigation hints.
|
|
173
|
+
An empty result set MUST show No matching models. Enter with no match MUST leave the picker open.
|
|
174
|
+
Escape and Ctrl+C MUST cancel without settings writes. Focus MUST propagate to the search input for cursor and IME placement.
|
|
175
|
+
Navigation and resize MUST keep the selected item visible. Truncation MUST NOT change the model identifier returned or saved.
|
|
176
|
+
Non-TUI UI clients MUST retain the existing provider/model selection labels and cancellation behavior.
|
|
177
|
+
The picker MUST use the current available physical model snapshot. It MUST NOT generate, change the conversation model, or refresh remote catalogs.
|
|
178
|
+
|
|
179
|
+
|
|
180
|
+
<a id="REQ-20"></a>
|
|
181
|
+
|
|
182
|
+
**REQ-20** Configured advisor reasoning effort
|
|
183
|
+
|
|
184
|
+
The model selection MAY contain reasoning: default, off, minimal, low, medium, high, xhigh, or max.
|
|
185
|
+
Null, nonstring, unknown, or incorrectly cased values MUST fail settings validation.
|
|
186
|
+
Omission and default MUST retain the prior request behavior without inheriting the parent thinking level.
|
|
187
|
+
As in pi-magic8ball, off MUST omit the simple-stream reasoning option. It MUST be accepted only when the model advertises off support.
|
|
188
|
+
Other explicit levels MUST be forwarded unchanged on every advisor request. The model and reasoning selection MUST remain pinned per consultation.
|
|
189
|
+
The extension MUST check explicit levels against Pi getSupportedThinkingLevels plus default before each provider request.
|
|
190
|
+
Unsupported levels MUST return unsupported-reasoning without generation, retry, clamping, or model fallback.
|
|
191
|
+
The user-only reasoning command without a level MUST report effective selection, configured level, and supported choices without writes or generation.
|
|
192
|
+
A reasoning <level> command MUST save a complete model selection through the existing atomic writer and trust checks.
|
|
193
|
+
Global writes MUST use the global selection, never copy a project override. Trusted project writes MAY copy the effective selection.
|
|
194
|
+
Explicit model commands and the picker MUST replace the complete selection. A new selection without reasoning MUST reset effort to legacy behavior.
|
|
195
|
+
Successful result and discovery model metadata MUST retain explicit reasoning. Expanded advisor results MUST show the pinned level or legacy default.
|
|
196
|
+
Tool input schemas MUST NOT permit caller-supplied reasoning. Thinking removal, unstructured advice, usage, limits, cancellation, and isolation MUST remain unchanged.
|
|
197
|
+
The README MUST explain configuration, support checks, pinning, and the cost/latency caveat. Effort is not a monetary budget or provider guarantee.
|
|
198
|
+
|
|
199
|
+
|
|
200
|
+
<a id="REQ-21"></a>
|
|
201
|
+
|
|
202
|
+
**REQ-21** Advisor command argument autocomplete
|
|
203
|
+
|
|
204
|
+
The registered advisor command MUST suggest show, model, reasoning, timeout, --global, and --project at its root.
|
|
205
|
+
Static choices MUST match a case-sensitive token prefix. Completion MUST preserve preceding arguments and whitespace, then append a space.
|
|
206
|
+
The grammar MUST accept one scope flag before, between, or after settings arguments. Repeated or conflicting scopes MUST receive no suggestions.
|
|
207
|
+
Model suggestions MUST use current available physical chat entries, not virtual entries or remote catalog refreshes.
|
|
208
|
+
Provider identifiers MUST be distinct and sorted. Provider matching MUST use a case-insensitive prefix.
|
|
209
|
+
Model matching MUST use Pi fuzzyFilter across identifiers and display names within the completed provider. Invalid settings identifiers MUST be excluded.
|
|
210
|
+
Reasoning suggestions MUST use fresh settings and advertised support for the selection targeted by the save scope.
|
|
211
|
+
Omitted scope and --global MUST use the global selection. Trusted --project MUST use the effective selection, including global fallback.
|
|
212
|
+
The extension MUST NOT read untrusted project settings. Missing context, missing models, invalid syntax, excess arguments, and failed lookups MUST return no dynamic suggestions.
|
|
213
|
+
Static suggestions MUST remain available without settings or catalogs. Lookup failures MUST remain silent without notifications or retries.
|
|
214
|
+
Session start and tree navigation MUST replace completion context. Shutdown MUST clear it. A context revision MUST reject pending old-context results.
|
|
215
|
+
Completion MUST NOT save settings, change the parent model, invoke tools, or generate. It MUST NOT add a dependency on pi-magic8ball.
|
|
216
|
+
The README MUST distinguish registered completion behavior from the known Pi 1.0.4 command-name Tab routing limitation.
|
|
217
|
+
|
|
218
|
+
|
|
219
|
+
## Host integration and results {#integration}
|
|
220
|
+
|
|
221
|
+
<a id="REQ-8"></a>
|
|
222
|
+
|
|
223
|
+
**REQ-8** Structured tool results
|
|
224
|
+
|
|
225
|
+
The package MUST register exactly advisor, advisor_sessions, and advisor_close as model-only tools. Each MUST define strict input and output schemas.
|
|
226
|
+
Successful advisor results MUST carry ok, advisory: true, session, response, turns, model, per-call usage, and cumulative consultation usage.
|
|
227
|
+
The extension, not the advisor model, MUST construct the structured tool result. The response field MUST carry unstructured advice as defined by REQ-17.
|
|
228
|
+
REQ-12 defines usage completeness fields and failure totals.
|
|
229
|
+
Failure results MUST carry ok: false, error: {code,message}, and per-call usage. REQ-11 defines additional limit details.
|
|
230
|
+
Existing session identifiers MUST accompany continuation failures.
|
|
231
|
+
Tool content, details, and structuredContent MUST agree. Failures MUST set isError. Host usage MUST contain only this call's usage, not cumulative usage.
|
|
232
|
+
The tool guidance MUST tell the main agent to explain objective, understanding, evidence, constraints, and uncertainty. It MUST prohibit citing advice as evidence or permission.
|
|
233
|
+
It MUST recommend continuation for the same issue and disclose provider charges. The extension MUST NOT automatically execute advice or continue consultations.
|
|
234
|
+
|
|
235
|
+
|
|
236
|
+
<a id="REQ-9"></a>
|
|
237
|
+
|
|
238
|
+
**REQ-9** Standalone package and verification
|
|
239
|
+
|
|
240
|
+
The package MUST load through Pi's resource loader without model work. Loading MUST NOT change profile settings or start background resources.
|
|
241
|
+
Runtime Pi and TypeBox packages MUST remain host-provided optional wildcard peers. pi-magic8ball MUST NOT be a dependency.
|
|
242
|
+
The package MUST provide repeatable Node test and TypeScript check commands. Acceptance tests MUST use injected mock responses and temporary configuration directories.
|
|
243
|
+
The README MUST describe loading, model setup, multi-turn use, close, limits, ephemeral state, transcript retention caveats, and advisory responsibility.
|
|
244
|
+
Paid provider checks MUST require separate authorization. Offline checks MUST NOT claim model reasoning quality or provider interoperability.
|
|
245
|
+
|
|
246
|
+
|
|
247
|
+
<a id="REQ-11"></a>
|
|
248
|
+
|
|
249
|
+
**REQ-11** Resource-specific limit diagnostics
|
|
250
|
+
|
|
251
|
+
On a resource limit failure, the extension MUST retain error.code = limit-exceeded.
|
|
252
|
+
It MUST add error.limit = {resource, maximum, actual} with nonnegative integer measurements.
|
|
253
|
+
resource MUST be input-bytes, reply-bytes, history-bytes, turns, or sessions.
|
|
254
|
+
maximum MUST be the inclusive local bound. actual MUST be the measured bytes or attempted pair/session count.
|
|
255
|
+
The extension MUST use fixed messages that identify the resource and its recovery action.
|
|
256
|
+
Input errors MUST ask for a shorter message. Reply errors MUST ask for a shorter requested answer.
|
|
257
|
+
History and turn errors MUST ask for a new consultation with an explicit summary. Session errors MUST ask to close idle consultations.
|
|
258
|
+
Diagnostics MUST NOT include message contents or raw provider errors. All REQ-4 bounds and rollback behavior MUST remain unchanged.
|
|
259
|
+
|
|
260
|
+
|
|
261
|
+
<a id="REQ-12"></a>
|
|
262
|
+
|
|
263
|
+
**REQ-12** Reported usage and completeness
|
|
264
|
+
|
|
265
|
+
Every advisor success or failure MUST include usageComplete, a boolean for this call's reported usage.
|
|
266
|
+
It MUST be true if no completion attempt began or a completed response supplied usage before return.
|
|
267
|
+
It MUST be false if a completion attempt began but no usage-bearing response was observed before return.
|
|
268
|
+
An exception, cancellation, or timeout after that attempt MUST NOT imply zero remote cost or completed remote cancellation.
|
|
269
|
+
Completed rejected replies MUST retain their reported usage and usageComplete = true.
|
|
270
|
+
Every successful consultation MUST retain totalUsageComplete, initially true, alongside known cumulative usage.
|
|
271
|
+
An attempt with usageComplete = false MUST make totalUsageComplete false for the rest of that consultation.
|
|
272
|
+
Successful results MUST return totalUsage and totalUsageComplete. Failures MUST return both when their requested consultation still exists.
|
|
273
|
+
Busy failures MUST mark cumulative usage incomplete while the other request is pending, without changing stored accounting.
|
|
274
|
+
A pending request MUST NOT add estimated usage. Late responses MUST NOT change state or repair incomplete totals.
|
|
275
|
+
Totals MUST be independent snapshots. Host usage MUST remain per-call only.
|
|
276
|
+
Completeness describes locally observed reporting, not verified provider billing or a monetary guarantee.
|
|
277
|
+
|
|
278
|
+
|
|
279
|
+
<a id="REQ-13"></a>
|
|
280
|
+
|
|
281
|
+
**REQ-13** Session capacity and known usage
|
|
282
|
+
|
|
283
|
+
Each advisor_sessions entry MUST include totalUsage, totalUsageComplete, turnsRemaining, historyBytes, and historyBytesRemaining.
|
|
284
|
+
Usage MUST be the known cumulative snapshot defined by REQ-12. A pending completion attempt MUST mark that snapshot incomplete.
|
|
285
|
+
Capacity MUST count only committed history. historyBytes MUST use the serialized UTF-8 calculation in REQ-4.
|
|
286
|
+
turnsRemaining MUST equal 24 minus committed pairs. historyBytesRemaining MUST equal 49152 minus committed history bytes.
|
|
287
|
+
The manager's injected test limits MAY replace those constants for offline boundary tests.
|
|
288
|
+
Returned nested objects MUST NOT expose mutable manager state. Metadata MUST NOT add full messages, advice, thinking, or token estimates.
|
|
289
|
+
Discovery MUST NOT run a model, change accounting, or alter consultation history.
|
|
290
|
+
|
|
291
|
+
|
|
292
|
+
<a id="REQ-14"></a>
|
|
293
|
+
|
|
294
|
+
**REQ-14** Safe tool presentation and progress
|
|
295
|
+
|
|
296
|
+
The extension MUST register call and result renderers for all three tools.
|
|
297
|
+
Advisor results MUST show an advisory label, session, and known per-call usage. REQ-19 defines collapsed previews and full expanded advice.
|
|
298
|
+
No renderer MUST infer a status from advice.
|
|
299
|
+
Expanded results MUST additionally show model, turns, and known cumulative usage with completeness labels.
|
|
300
|
+
Failures MUST show safe diagnostics and label incomplete usage. Missing details MUST use fixed fallback text, not raw result content.
|
|
301
|
+
Discovery and close renderers MUST show metadata and close outcomes without reconstructing transcripts.
|
|
302
|
+
Renderers MUST neutralize data control characters, fit terminal widths, and use the current callback theme.
|
|
303
|
+
The advisor tool MUST emit phase-only updates: preparing before new-model preparation and waiting before a completion attempt.
|
|
304
|
+
Updates MUST NOT contain advice, thinking, explicit input, raw errors, or usage. Callback errors MUST NOT change the consultation outcome.
|
|
305
|
+
No update MUST occur for invalid arguments or exhausted capacity. No late completion MUST emit updates.
|
|
306
|
+
Final model-facing JSON and structuredContent MUST remain identical. Rendering MUST NOT change data or start provider work.
|
|
307
|
+
|
|
308
|
+
|
|
309
|
+
<a id="REQ-15"></a>
|
|
310
|
+
|
|
311
|
+
**REQ-15** Opt-in live verification procedure
|
|
312
|
+
|
|
313
|
+
The README MUST provide an opt-in procedure for one successful consultation and one continuation against a designated physical model.
|
|
314
|
+
The procedure MUST require separate spending authorization before either call. It MUST disclose the absence of a local monetary cap.
|
|
315
|
+
It MUST limit the run to two requests, forbid automatic retries, and stop after any failure.
|
|
316
|
+
It MUST use nonsecret text, verify model pinning and reported usage, and close the consultation.
|
|
317
|
+
It MUST identify provider/model, Pi/Node versions, outcomes, and usage completeness in a report without storing credentials.
|
|
318
|
+
Offline checks MUST NOT invoke that procedure or claim live interoperability. This task MUST NOT perform live calls without further authorization.
|
|
319
|
+
|
|
320
|
+
|
|
321
|
+
<a id="REQ-16"></a>
|
|
322
|
+
|
|
323
|
+
**REQ-16** Safe reply rejection diagnostics
|
|
324
|
+
|
|
325
|
+
For invalid-response, the extension MUST identify the failed validation in error.message using fixed text. It MUST NOT expose reply text, thinking, raw provider errors, or provider-supplied field values. Default calls MUST NOT add result fields; REQ-22 defines the opt-in diagnostic fields.
|
|
326
|
+
Validation MUST check in order: provider error/abort, normal stop, content shape, reply byte limit, and nonblank text. REQ-17 removes marker validation.
|
|
327
|
+
Content MUST be an array of nonnull text or thinking blocks. Every text block MUST have string text. Malformed content MUST return invalid-response rather than provider-failed.
|
|
328
|
+
Messages MUST distinguish abnormal completion, unsupported or malformed content, and absent text. Every rejection message MUST state that no exchange was committed.
|
|
329
|
+
The extension MUST preserve byte limits, known usage, rollback, and no retry. It MUST NOT infer which validation failed in an earlier call whose reply is unavailable.
|
|
330
|
+
|
|
331
|
+
|
|
332
|
+
<a id="REQ-18"></a>
|
|
333
|
+
|
|
334
|
+
**REQ-18** Responses visible without expansion
|
|
335
|
+
|
|
336
|
+
REQ-19 supersedes the former requirement for full advice in the collapsed view. Wrapping and control neutralization MUST remain as defined in REQ-14.
|
|
337
|
+
The expanded view MUST show full advice and metadata. The README MUST describe the current behavior.
|
|
338
|
+
Showing advice MUST use the existing tool result and renderer. The extension MUST NOT send duplicate custom/user messages, add transcript entries, start parent model turns, or change result content/details/structuredContent for this presentation change.
|
|
339
|
+
Partial or failed replies MUST NOT appear as successful advice. Existing privacy, completion validation, lifecycle, and provider-call behavior MUST remain unchanged.
|
|
340
|
+
|
|
341
|
+
|
|
342
|
+
<a id="REQ-22"></a>
|
|
343
|
+
|
|
344
|
+
**REQ-22** Opt-in privacy-safe failure diagnostics
|
|
345
|
+
|
|
346
|
+
The advisor input MAY contain diagnostics, a boolean. Omission and false MUST preserve the previous failure envelope. Null and nonboolean values MUST fail validation. True MUST add error.diagnostics to failures of that call only; successes MUST remain unchanged. This extends REQ-1 and REQ-8 without relaxing unknown-field validation. REQ-16's prohibition on additional fields remains applicable to default calls; opted-in failures MAY add only this diagnostic object.
|
|
347
|
+
Diagnostics MUST contain phase: validation, preparation, completion, response-validation, or commit. Validation covers input, reservation, and history preflight. Preparation covers selection resolution. Completion covers the completion invocation until a response is returned. Response-validation covers returned usage, reply validation, and full-pair history bounds. Commit covers final cancellation/deadline checks and committing the pair.
|
|
348
|
+
Diagnostics MUST contain category: advisor-error for extension-defined errors, local-error for unexpected exceptions outside completion, or unknown for unclassified completion exceptions. A returned error or aborted stop reason MUST use provider-error or provider-aborted respectively, except for the terminal error hints in REQ-23. Terminal abort MUST take precedence over error hints.
|
|
349
|
+
For thrown completion exceptions, the extension MUST read only own data descriptors for status and code. Status 401/403 MUST map to authentication; 400/404/409/413/422 to provider-rejection; 429 to rate-limit; 408/500/502/503/504 to transport. Recognized status MUST take precedence over code. Codes ECONNRESET, ECONNREFUSED, ETIMEDOUT, ENOTFOUND, and EAI_AGAIN MUST map to transport when status is unrecognized. All other values, accessor fields, and descriptor inspection failures MUST yield unknown. The extension MUST NOT parse free-form provider error messages, except for the fixed SDK wrapper in REQ-23. It MUST NOT traverse causes, headers, or payloads.
|
|
350
|
+
When a selection is known, diagnostics MUST include model with configured reasoning and selectionSource: global, project, or unknown. Preparation MUST report resolved settings before checking model availability/authentication. A continuation MUST use the pinned model and original source, not current settings. Unresolved selection MUST omit both fields. Late preparation/completion MUST NOT mutate returned diagnostics or pinned state.
|
|
351
|
+
Diagnostic categories are hints, not verified root causes. Errors MUST retain existing fixed codes/messages, usage completeness, rollback, isolation, and no-retry behavior. Diagnostics MUST NOT expose credentials, headers, exception text, rejected advice, thinking, or arbitrary provider payloads. Renderers MUST display opted-in diagnostics with existing width/control protections. The README MUST explain opt-in, phases, categories, selection provenance, transcript persistence, and the zero-usage billing caveat. Enabling diagnostics MUST NOT authorize another live request.
|
|
352
|
+
|
|
353
|
+
|
|
354
|
+
<a id="REQ-23"></a>
|
|
355
|
+
|
|
356
|
+
**REQ-23** Integer timeouts and safe SDK terminal error hints
|
|
357
|
+
|
|
358
|
+
For each completion attempt, advisor MUST pass timeoutMs = max(1, floor(deadline milliseconds minus elapsed milliseconds)).
|
|
359
|
+
The timeout MUST remain an integer within the original per-call deadline budget. The local deadline and cancellation checks MUST remain unchanged.
|
|
360
|
+
For opted-in diagnostics on a terminal error, advisor MUST inspect only an own data descriptor for errorMessage.
|
|
361
|
+
It MUST NOT invoke getters. Missing, nonstring, or hostile metadata MUST retain provider-error without extra fields.
|
|
362
|
+
An exact errorMessage of "timeout must be an integer" or "timeout must be a positive integer" MUST report category sdk-error and code sdk-invalid-timeout.
|
|
363
|
+
The renderer MUST show fixed text that identifies the SDK timeout constraint.
|
|
364
|
+
Otherwise, advisor MUST check only the SDK wrapper prefix "<selected-provider> API error (<status>): ", except for the fixed rejection detail in REQ-24.
|
|
365
|
+
For provider openai, the wrapper label MUST be OpenAI. Status MUST be exactly three decimal digits from 400 through 599.
|
|
366
|
+
A matching wrapper MUST add httpStatus and use REQ-22's HTTP category mapping. Unmapped statuses MUST retain provider-error.
|
|
367
|
+
No body, free-form message, credential, header, advice, or thinking MUST enter diagnostic output.
|
|
368
|
+
These text-derived hints MUST NOT claim a verified remote root cause. Terminal abort MUST remain provider-aborted without text-derived hints.
|
|
369
|
+
Default failure envelopes and successful envelopes MUST remain unchanged. Opt-in, usage reporting, rollback, isolation, and no retry MUST remain unchanged.
|
|
370
|
+
Tests MUST exercise the installed OpenAI Responses adapter with a mock fetch and a synthetic key. They MUST make no network requests.
|
|
371
|
+
|
|
372
|
+
|
|
373
|
+
<a id="REQ-24"></a>
|
|
374
|
+
|
|
375
|
+
**REQ-24** Tool-free xAI requests and fixed rejection diagnostics
|
|
376
|
+
|
|
377
|
+
REQ-25 supersedes the provider-specific toolChoice rules. Advisor MUST continue to declare zero tools and reject tool-call replies.
|
|
378
|
+
The selected model, effort, timeout, token bound, and no-retry options MUST remain unchanged.
|
|
379
|
+
For opted-in xai terminal errors with the exact SDK HTTP 400 wrapper, advisor MUST recognize these complete rejection details:
|
|
380
|
+
- Invalid request content: A tool_choice was set on the request but no tools were specified.
|
|
381
|
+
- A tool_choice was set on the request but no tools were specified.
|
|
382
|
+
- 400 Invalid request content: A tool_choice was set on the request but no tools were specified.
|
|
383
|
+
For an exact match, diagnostics MUST add code tool-choice-without-tools to the existing provider-rejection category and httpStatus 400.
|
|
384
|
+
The renderer MUST show fixed text that identifies tool_choice without tools. It MUST NOT expose the original detail.
|
|
385
|
+
Arbitrary details, partial matches, appended content, different providers, and different statuses MUST retain existing hints without this code.
|
|
386
|
+
Default failures, successes, isolation, usage, rollback, and no retries MUST remain unchanged. No settings or credential files MUST change.
|
|
387
|
+
|
|
388
|
+
|
|
389
|
+
<a id="REQ-25"></a>
|
|
390
|
+
|
|
391
|
+
**REQ-25** Provider-neutral tool-free request options
|
|
392
|
+
|
|
393
|
+
For every creation and continuation request, advisor MUST omit the toolChoice option from the host model-registry call.
|
|
394
|
+
This rule MUST apply to every selected provider, including custom providers. Advisor MUST NOT select toolChoice behavior by provider identifier.
|
|
395
|
+
Pi MUST control provider payload construction. Advisor MUST continue to declare zero tools and reject tool-call replies.
|
|
396
|
+
The selected model, configured effort, integer timeout, cancellation signal, token bound, cache retention, session identifier, and disabled retries MUST remain unchanged.
|
|
397
|
+
Private diagnostics, atomic history, usage accounting, settings, and authentication MUST remain unchanged.
|
|
398
|
+
This requirement supersedes only REQ-24's provider-specific toolChoice rules.
|
|
399
|
+
|
|
400
|
+
|
|
401
|
+
- #REQ-25 supersedes: [#REQ-24](#REQ-24)
|
|
402
|
+
|
|
403
|
+
|
|
404
|
+
<a id="REQ-26"></a>
|
|
405
|
+
|
|
406
|
+
**REQ-26** Configurable pinned advisor deadline
|
|
407
|
+
|
|
408
|
+
Advisor MUST use a default deadline of 300000 milliseconds per call. The user approved this five-minute default.
|
|
409
|
+
Settings MAY contain root timeoutMs as an integer from 1 through 2147483647 milliseconds, inclusive.
|
|
410
|
+
The upper bound MUST match Node's timer range. Zero, negative, fractional, nonnumber, null, and larger values MUST fail settings validation.
|
|
411
|
+
Timeout MUST resolve independently of model selection: trusted project timeoutMs, then global timeoutMs, then the default.
|
|
412
|
+
A project model without timeoutMs MUST inherit the global timeout. A timeout-only project MUST retain the global model.
|
|
413
|
+
Untrusted project settings MUST remain unread. Invalid global settings MUST fail even when a project overrides them.
|
|
414
|
+
|
|
415
|
+
Each new consultation MUST pin the resolved timeout with its model. Configuration changes MUST affect new consultations only.
|
|
416
|
+
Every call MUST measure its budget from call entry, including preparation. The SDK MUST receive a positive integer remaining timeout.
|
|
417
|
+
Before settings resolve, preparation MUST use the default deadline. A constructor test override MAY replace that initial default.
|
|
418
|
+
When settings resolve, advisor MUST rearm the deadline from the original call-entry time, not start a new budget.
|
|
419
|
+
If the configured budget has already elapsed, advisor MUST fail with timeout before generation.
|
|
420
|
+
Late preparation callbacks MUST NOT change deadlines, diagnostics, state, or schedule timers after return or cancellation.
|
|
421
|
+
Cancellation, disabled retries, rollback, usage completeness, and rejection of late commits MUST remain unchanged.
|
|
422
|
+
|
|
423
|
+
The user-only command timeout without a value MUST show effective milliseconds and source without writes or generation.
|
|
424
|
+
The timeout <milliseconds> command MUST accept only decimal integer text within the settings range.
|
|
425
|
+
The timeout default command MUST remove only that scope's explicit timeout, restoring inheritance or the default.
|
|
426
|
+
Timeout commands MUST work without a configured model. They MUST preserve the scope's own model and reasoning, without copying another scope's model.
|
|
427
|
+
Model, picker, and reasoning saves MUST preserve the target scope's explicit timeout, without copying an inherited timeout into that scope.
|
|
428
|
+
Preserved fields MUST come from the target layer at the save checkpoint under the exclusive lock, not an earlier command snapshot.
|
|
429
|
+
Reset MAY atomically save an empty settings object. Existing lock, trust, file bounds, and checkpoint safety rules MUST remain unchanged.
|
|
430
|
+
show MUST include effective timeout and source. Autocomplete MUST add timeout and timeout default without settings reads or provider work.
|
|
431
|
+
Tool inputs MUST NOT accept caller-supplied timeouts. Existing tool result schemas and model metadata MUST remain unchanged.
|
|
432
|
+
The README MUST document units, range, precedence, pinning, reset, and the fact that a longer deadline is not a billing guarantee.
|
|
433
|
+
|
|
434
|
+
|
|
435
|
+
<a id="REQ-27"></a>
|
|
436
|
+
|
|
437
|
+
**REQ-27** Short provider prompt caching
|
|
438
|
+
|
|
439
|
+
For every creation and continuation request, advisor MUST set cacheRetention to short at the host model-registry boundary.
|
|
440
|
+
Advisor MUST retain one stable sessionId per consultation. Different consultations MUST use different identifiers, independent of the parent session.
|
|
441
|
+
Advisor MUST retain the fixed system prompt and append-only validated exchange text. It MUST NOT add changing cache metadata to the prompt.
|
|
442
|
+
Pi MUST control provider cache payload construction. Advisor MUST NOT add provider-specific cache fields or a new cache setting.
|
|
443
|
+
Cache support, minimum prompt size, lifetime, and cache hits depend on the provider. Short MUST NOT imply guaranteed savings or erasure.
|
|
444
|
+
Closing a consultation MUST remove local state only. It MUST NOT claim to delete provider caches or the parent transcript.
|
|
445
|
+
Model and effort pinning, zero tools, retries, timeout, cancellation, rollback, and usage reporting MUST remain unchanged.
|
|
446
|
+
The README MUST describe short caching, same-session continuation, provider-dependent results, and the limits of local close.
|
|
447
|
+
This requirement supersedes only REQ-25's instruction to preserve cache retention.
|
|
448
|
+
|
|
449
|
+
|
|
450
|
+
- #REQ-27 supersedes: [#REQ-25](#REQ-25)
|
|
451
|
+
|
|
452
|
+
|
|
453
|
+
<a id="REQ-28"></a>
|
|
454
|
+
|
|
455
|
+
**REQ-28** Private opaque continuation state
|
|
456
|
+
|
|
457
|
+
After reply validation, advisor MUST preserve allowlisted opaque replay data for the pinned provider, model, and API only.
|
|
458
|
+
Supported APIs MUST be openai-responses, openai-codex-responses, azure-openai-responses, google-generative-ai, google-vertex, anthropic-messages, and bedrock-converse-stream.
|
|
459
|
+
For Responses APIs, replay MUST preserve encrypted_content and reasoning item IDs. It MUST remove plaintext reasoning content and summaries.
|
|
460
|
+
Responses reasoning signatures MUST contain type reasoning, a nonempty string ID, and nonempty string encrypted_content. Replay summary MUST be empty.
|
|
461
|
+
Replay MAY retain status only when it is completed, incomplete, or in_progress. Other reasoning-item fields MUST NOT survive.
|
|
462
|
+
Responses text signatures MUST retain only version 1, a string ID, and commentary or final_answer phase when present.
|
|
463
|
+
For Google APIs, replay MUST preserve valid base64 text/thinking signatures on their original blocks. Replayed thinking text MUST be empty.
|
|
464
|
+
For Anthropic and Bedrock, replay MUST preserve only redacted thinking with a nonempty string signature. Ordinary signed plaintext thinking MUST NOT survive.
|
|
465
|
+
All supported APIs MUST preserve original text-block boundaries and order. Absent, malformed, and unsupported metadata MUST be omitted without changing valid reply text.
|
|
466
|
+
A reply with a different provider or model MUST NOT contribute replay data. Unsupported APIs MUST keep the existing text-only history behavior.
|
|
467
|
+
|
|
468
|
+
Opaque replay MUST remain process-local and private. Results, metadata, progress, renderers, diagnostics, parent tool results, and persisted files MUST NOT contain replay data.
|
|
469
|
+
The extension MUST NOT retain plaintext thinking, summaries, raw provider replies, arbitrary metadata, or response IDs for server-side chaining.
|
|
470
|
+
Advisor MUST NOT enable server-side storage or previous_response_id. Pi MUST continue to own provider payload construction.
|
|
471
|
+
Continuation contexts MUST use the original API identity and detached replay copies. Provider mutation MUST NOT change retained state.
|
|
472
|
+
Replay MUST commit atomically with its validated user/reply pair. Failure, cancellation, timeout, and late completion MUST leave prior replay unchanged.
|
|
473
|
+
Different consultations MUST NOT share replay data. Close and lifecycle clear MUST release local replay with the consultation.
|
|
474
|
+
|
|
475
|
+
The existing 49152-byte history limit MUST include private replay. Count UTF-8 bytes of JSON.stringify entries with role, text, and optional replay.
|
|
476
|
+
Each replay value MUST contain api and sanitized content. With no replay, the existing role/text byte measurement MUST remain unchanged.
|
|
477
|
+
Pending-user and full-pair checks MUST use this combined measurement. Overflow MUST fail with history-bytes and commit no pair or replay.
|
|
478
|
+
The README MUST describe opaque local state, combined byte accounting, supported formats, and the plaintext-thinking limitation.
|
|
479
|
+
This requirement supersedes only REQ-2's prohibition on retaining opaque thinking metadata and REQ-4's role/text-only byte formula.
|
|
480
|
+
Plaintext thinking exclusion, public advice text, isolation, short caching, request controls, usage, and failure behavior MUST remain unchanged.
|
|
481
|
+
|
|
482
|
+
|
|
483
|
+
- #REQ-28 supersedes: [#REQ-4](#REQ-4)
|
|
484
|
+
|
|
485
|
+
|
|
486
|
+
- #REQ-28 supersedes: [#REQ-2](#REQ-2)
|
|
487
|
+
|
|
488
|
+
|
|
489
|
+
<a id="REQ-29"></a>
|
|
490
|
+
|
|
491
|
+
**REQ-29** Model-sized advisor context
|
|
492
|
+
|
|
493
|
+
Advisor MUST replace the 49152-byte history cap with the pinned advisor model's contextWindow from Pi's model registry.
|
|
494
|
+
The resolved value MUST be a positive safe integer. Missing or invalid values MUST return model-unavailable before generation.
|
|
495
|
+
The model window MUST be pinned at creation. Parent model and parent context usage MUST NOT supply this value.
|
|
496
|
+
Each history entry MAY retain validated reply usage for context measurement. Raw replies and plaintext thinking MUST remain excluded.
|
|
497
|
+
Context tokens MUST follow Pi 1.0.4: latest positive validated assistant calculateContextTokens(usage), plus estimateTokens for later messages.
|
|
498
|
+
With no positive usage, advisor MUST estimate each message with Pi's public estimateTokens helper, including its fixed system prompt.
|
|
499
|
+
Opaque ciphertext length MUST NOT be treated as a token count. Context measurement is an estimate, not a tokenizer guarantee.
|
|
500
|
+
Before dispatch, estimated pending context plus the existing 4096-token output request MUST NOT exceed contextWindow.
|
|
501
|
+
After reply validation, estimated committed context MUST NOT exceed contextWindow. Equality MUST pass.
|
|
502
|
+
Overflow MUST return limit-exceeded with resource context-tokens, maximum contextWindow, and attempted token count.
|
|
503
|
+
Failed or interrupted pairs MUST NOT advance retained context. Completed rejected usage MUST still count toward cumulative usage.
|
|
504
|
+
Advisor MUST NOT trim, summarize, compact, retry, or request parent history automatically.
|
|
505
|
+
The 8-session, 24-pair, 16384-byte input/reply, timeout, and output request limits MUST remain unchanged.
|
|
506
|
+
Advisor success results and session metadata MUST expose contextUsage with tokens, contextWindow, and percent.
|
|
507
|
+
Failed advisor calls MUST include retained contextUsage when the consultation still exists with a resolved window.
|
|
508
|
+
Percent MUST equal 100 times tokens divided by contextWindow. Returned values MUST be detached snapshots.
|
|
509
|
+
Metadata MUST replace historyBytesRemaining with contextUsage. historyBytes MAY remain informational, without a fixed cap.
|
|
510
|
+
Strict schemas MUST reflect the new metadata and limit resource. Unsupported unknown fields MUST remain rejected.
|
|
511
|
+
This requirement supersedes only REQ-4 and REQ-28's retained-history byte cap and REQ-13's byte-remaining and no-token-estimate rules.
|
|
512
|
+
|
|
513
|
+
|
|
514
|
+
- #REQ-29 supersedes: [#REQ-13](#REQ-13)
|
|
515
|
+
|
|
516
|
+
|
|
517
|
+
- #REQ-29 supersedes: [#REQ-28](#REQ-28)
|
|
518
|
+
|
|
519
|
+
|
|
520
|
+
- #REQ-29 supersedes: [#REQ-4](#REQ-4)
|
|
521
|
+
|
|
522
|
+
|
|
523
|
+
<a id="REQ-30"></a>
|
|
524
|
+
|
|
525
|
+
**REQ-30** Visible advisor question and Pi-like usage line
|
|
526
|
+
|
|
527
|
+
The advisor call renderer MUST show the main agent's explicit message below the call heading, with original line breaks.
|
|
528
|
+
Display sanitization MUST neutralize terminal controls. Rendering MUST NOT change the message sent to the provider.
|
|
529
|
+
The full sanitized message MUST be visible. The renderer MUST NOT show hidden replay state, raw provider payloads, or plaintext thinking.
|
|
530
|
+
Phase updates MUST remain phase-only. No new transcript entries, provider calls, parent-footer replacement, or persisted consultation files are permitted.
|
|
531
|
+
Both collapsed and expanded advisor results MUST show a consultation-local usage line below the result heading.
|
|
532
|
+
The line MUST show cumulative uncached input with ↑, output with ↓, cache reads with R, and cache writes with W when nonzero.
|
|
533
|
+
CH MUST use the latest call's reported cacheRead divided by input plus cacheRead plus cacheWrite, times 100.
|
|
534
|
+
Unknown usage or a zero denominator MUST show CH?. Output tokens MUST NOT enter the denominator.
|
|
535
|
+
The line MUST show cumulative reported cost to three decimals and context percent/window from REQ-29.
|
|
536
|
+
Token abbreviations MUST follow Pi 1.0.4's footer thresholds. Context percent and CH MUST use one decimal.
|
|
537
|
+
Missing context from historic results MUST show ?/?. Incomplete cumulative usage MUST be labeled incomplete.
|
|
538
|
+
The line MUST NOT show sub or auto labels. Advice MUST NOT acquire a semantic status classification.
|
|
539
|
+
Known cumulative failures MUST show the line with retained context. Expanded session metadata MUST show context rather than remaining bytes.
|
|
540
|
+
Existing eight-row advice previews, expanded full advice, safe failures, current themes, width bounds, isolation, and opaque-state privacy MUST remain.
|
|
541
|
+
README MUST explain visible questions, model-sized estimated context, remaining limits, cumulative totals, per-call CH, and reported cost.
|
|
542
|
+
|
|
543
|
+
|
|
544
|
+
<a id="REQ-31"></a>
|
|
545
|
+
|
|
546
|
+
**REQ-31** Remove fixed consultation caps
|
|
547
|
+
|
|
548
|
+
Advisor MUST NOT cap active consultations, committed exchanges, message bytes, or reply bytes with fixed extension limits.
|
|
549
|
+
Advisor MUST omit maxTokens from every completion options object. Pi MUST control the model output allowance.
|
|
550
|
+
Before dispatch, advisor MUST compare estimated pending context tokens with the pinned model contextWindow, without a fixed output reserve.
|
|
551
|
+
After reply validation, advisor MUST compare estimated committed context tokens with that same window.
|
|
552
|
+
Equality MUST pass both checks. Overflow MUST return limit-exceeded with resource context-tokens and commit no exchange.
|
|
553
|
+
The fixed safe context error MUST report maximum contextWindow and actual estimated tokens.
|
|
554
|
+
Completed rejected replies MUST retain their known usage. A failed continuation MUST retain its previous context and replay.
|
|
555
|
+
|
|
556
|
+
Success and discovery schemas MUST permit exchange counts above 24. Discovery MUST permit more than eight entries.
|
|
557
|
+
Discovery MUST omit turnsRemaining. Session renderers MUST show committed exchange counts and all listed consultations, without a fixed eight-entry cutoff.
|
|
558
|
+
Public limit schemas MUST accept only context-tokens as a resource. Removed resource names MUST fail schema validation.
|
|
559
|
+
The README MUST describe the removed caps, Pi output defaults, retained context checks, and unchanged deadline.
|
|
560
|
+
|
|
561
|
+
The configurable timeout, caller cancellation, busy protection, no retries, strict inputs, reply validation, and isolation MUST remain unchanged.
|
|
562
|
+
Usage snapshots, opaque replay privacy, close, lifecycle reset, and ephemeral retention MUST remain unchanged.
|
|
563
|
+
The 16384-byte settings-file bound, 80-code-point labels, picker row bounds, and eight-row advice previews MUST remain unchanged.
|
|
564
|
+
|
|
565
|
+
This requirement supersedes REQ-4's fixed consultation, exchange, input-byte, reply-byte, and output-token bounds.
|
|
566
|
+
It supersedes REQ-11's removed limit resources and REQ-13's turnsRemaining field.
|
|
567
|
+
It supersedes REQ-29's fixed output reserve and preservation of the removed caps.
|
|
568
|
+
Earlier requirements and acceptance criteria MUST NOT restore these caps through a preservation clause.
|
|
569
|
+
|
|
570
|
+
|
|
571
|
+
- #REQ-31 supersedes: [#REQ-4](#REQ-4)
|
|
572
|
+
|
|
573
|
+
|
|
574
|
+
- #REQ-31 supersedes: [#REQ-11](#REQ-11)
|
|
575
|
+
|
|
576
|
+
|
|
577
|
+
- #REQ-31 supersedes: [#REQ-13](#REQ-13)
|
|
578
|
+
|
|
579
|
+
|
|
580
|
+
- #REQ-31 supersedes: [#REQ-29](#REQ-29)
|
|
581
|
+
|
|
582
|
+
|
|
583
|
+
## Acceptance criteria {#acceptance}
|
|
584
|
+
|
|
585
|
+
<a id="AC-31"></a>
|
|
586
|
+
|
|
587
|
+
**AC-31** Calls beyond former caps
|
|
588
|
+
|
|
589
|
+
Changed checks MUST fail against the capped baseline before implementation.
|
|
590
|
+
Send multibyte messages and valid replies above 16384 UTF-8 bytes within the model window.
|
|
591
|
+
PASS when the exact message reaches completion, the exact reply commits, and continuation retains both without truncation.
|
|
592
|
+
Create at least nine consultations, including pending creation. Continue one consultation through at least 25 successful exchanges.
|
|
593
|
+
PASS when all calls succeed, identifiers remain distinct, and prior consultations remain available without eviction.
|
|
594
|
+
Inspect options across provider selections. PASS when maxTokens is absent while reasoning, short caching, and zero retries remain.
|
|
595
|
+
Use the installed Responses adapter with synthetic authentication and mock fetch.
|
|
596
|
+
PASS when a model allowance above 4096 reaches the payload instead of the former extension cap.
|
|
597
|
+
Test pending context at the exact model window and one token above it.
|
|
598
|
+
PASS when equality dispatches without a fixed output reserve and overflow returns exact context diagnostics without dispatch.
|
|
599
|
+
Validate successful results and discovery above the former count caps.
|
|
600
|
+
PASS when turnsRemaining is absent, strict unknown fields remain rejected, and removed limit resource names fail validation.
|
|
601
|
+
Render at least nine entries in both views. PASS when every identifier appears, committed counts replace remaining counts, and lines fit.
|
|
602
|
+
Inspect the README. PASS when it describes the removed caps and retained context, timeout, and provider-cost caveats.
|
|
603
|
+
|
|
604
|
+
|
|
605
|
+
<a id="AC-32"></a>
|
|
606
|
+
|
|
607
|
+
**AC-32** Preserved safeguards after cap removal
|
|
608
|
+
|
|
609
|
+
Preserved checks MUST pass before and after implementation.
|
|
610
|
+
Inject committed-context overflow and a malformed completion during continuation.
|
|
611
|
+
PASS when previous context and private replay remain unchanged, completed rejected usage remains counted, and no retry occurs.
|
|
612
|
+
Cancel pending creation and settle its reply afterward. PASS when no late exchange commits or discarded consultation reappears.
|
|
613
|
+
Retain existing deadline, busy, close, lifecycle, settings-file bounds, strict validation, privacy, and eight-row advice-preview checks.
|
|
614
|
+
Use mock catalogs, temporary settings, synthetic replay, and mock fetch only. No paid provider call is part of verification.
|
|
615
|
+
|
|
616
|
+
|
|
617
|
+
- #REQ-31 verified-by: [#AC-31](#AC-31)
|
|
618
|
+
|
|
619
|
+
|
|
620
|
+
- #REQ-31 verified-by: [#AC-32](#AC-32)
|
|
621
|
+
|
|
622
|
+
|
|
623
|
+
<a id="AC-23"></a>
|
|
624
|
+
|
|
625
|
+
**AC-23** SDK timeout and terminal error regression coverage
|
|
626
|
+
|
|
627
|
+
Changed behavior: pass the actual manager timeout to the installed Responses adapter with a mock HTTP 401 response.
|
|
628
|
+
PASS when the timeout is a positive integer, one mock fetch occurs, and opted-in diagnostics report authentication with httpStatus 401.
|
|
629
|
+
Changed behavior: inject both exact timeout validation strings. PASS when code sdk-invalid-timeout and category sdk-error appear without raw text.
|
|
630
|
+
Changed behavior: inject valid SDK HTTP wrappers for REQ-22's status groups and an unmapped 418.
|
|
631
|
+
PASS when only httpStatus and the mapped category appear, without private body content.
|
|
632
|
+
Changed behavior: render SDK and HTTP hints at widths 1, 8, 30, and 80. PASS when fixed timeout text and HTTP status appear and lines fit.
|
|
633
|
+
Preserved behavior: inject arbitrary text, wrong provider labels, malformed statuses, accessors, hostile metadata, and terminal aborts.
|
|
634
|
+
PASS when only safe fallback categories appear and no getter or private content is exposed.
|
|
635
|
+
Preserved behavior: omission and false MUST omit hints. Success shapes MUST remain unchanged. Failure MUST preserve history and make no retry.
|
|
636
|
+
All 84 existing checks and TypeScript MUST pass before and after this change.
|
|
637
|
+
|
|
638
|
+
|
|
639
|
+
- #REQ-23 verified-by: [#AC-23](#AC-23)
|
|
640
|
+
|
|
641
|
+
|
|
642
|
+
AC-1 through AC-9 describe the initial implementation. AC-10 describes the picker correction.
|
|
643
|
+
For the diagnostics and presentation change, preserve all 16 existing checks. AC-11 through AC-15 describe changed behavior.
|
|
644
|
+
|
|
645
|
+
<a id="AC-1"></a>
|
|
646
|
+
|
|
647
|
+
**AC-1**
|
|
648
|
+
|
|
649
|
+
Start two consultations, then continue the first after changing configured selection. PASS when each request contains only its own exchanges, zero tools, and the first pinned model.
|
|
650
|
+
|
|
651
|
+
|
|
652
|
+
<a id="AC-2"></a>
|
|
653
|
+
|
|
654
|
+
**AC-2**
|
|
655
|
+
|
|
656
|
+
Exercise thinking plus text, blank or thinking-only content, tool calls, and incomplete completions. PASS when nonblank text returns without thinking and invalid completions commit nothing. AC-17 covers removal of marker and status handling.
|
|
657
|
+
Inspect both prompts and tool guidance for evidence and authorization distinctions.
|
|
658
|
+
|
|
659
|
+
|
|
660
|
+
<a id="AC-3"></a>
|
|
661
|
+
|
|
662
|
+
**AC-3**
|
|
663
|
+
|
|
664
|
+
Inject provider exceptions, completed error responses, caller abort, deadline expiration, ignored abort with late success, and overlapping continuation.
|
|
665
|
+
PASS when each failure returns its safe code, makes no automatic retry, preserves history, accounts for known usage, and prevents late commits.
|
|
666
|
+
|
|
667
|
+
|
|
668
|
+
<a id="AC-4"></a>
|
|
669
|
+
|
|
670
|
+
**AC-4**
|
|
671
|
+
|
|
672
|
+
Test limits at their inclusive boundary and one unit beyond. Use multibyte text for byte limits.
|
|
673
|
+
PASS when allowed inputs commit, forbidden inputs return limit-exceeded, state stays bounded, no implicit compaction occurs, and closing releases capacity.
|
|
674
|
+
|
|
675
|
+
|
|
676
|
+
<a id="AC-5"></a>
|
|
677
|
+
|
|
678
|
+
**AC-5**
|
|
679
|
+
|
|
680
|
+
List, close, duplicate-close, close during a request, and reset during a request. PASS when metadata excludes transcript, close rejects unknown/busy identifiers, and reset prevents late commits.
|
|
681
|
+
Fire each clearing host event and compaction. PASS when clearing invalidates identifiers and compaction retains them.
|
|
682
|
+
|
|
683
|
+
|
|
684
|
+
<a id="AC-6"></a>
|
|
685
|
+
|
|
686
|
+
**AC-6**
|
|
687
|
+
|
|
688
|
+
Use absent, valid, malformed, oversized, null, unknown-field, symlink, and special-file settings. Test trusted and untrusted projects.
|
|
689
|
+
PASS when precedence follows REQ-6, untrusted project bytes are never consumed, and invalid settings fail closed.
|
|
690
|
+
Through the tool boundary, test missing selection, missing authentication, absent model, and virtual model. PASS when no request occurs.
|
|
691
|
+
|
|
692
|
+
|
|
693
|
+
<a id="AC-7"></a>
|
|
694
|
+
|
|
695
|
+
**AC-7**
|
|
696
|
+
|
|
697
|
+
Exercise show, explicit global/project save, bare selection, cancellation, unavailable model, invalid command, lost trust, and competing locks.
|
|
698
|
+
PASS when only validated user requests save advisor.json, saves remain strict and atomic, and cancellation/errors trigger no model call or unrelated write.
|
|
699
|
+
|
|
700
|
+
|
|
701
|
+
<a id="AC-8"></a>
|
|
702
|
+
|
|
703
|
+
**AC-8**
|
|
704
|
+
|
|
705
|
+
Inspect registered tools and validate successful/failed results against their schemas. PASS when exactly three model-only tools exist and content/details/structuredContent agree.
|
|
706
|
+
PASS when per-call usage reaches the host once, cumulative usage appears separately, and argument schemas reject null or unknown fields.
|
|
707
|
+
|
|
708
|
+
|
|
709
|
+
<a id="AC-9"></a>
|
|
710
|
+
|
|
711
|
+
**AC-9**
|
|
712
|
+
|
|
713
|
+
Load the package with the installed Pi resource loader without credentials or provider calls. Run Node acceptance checks and TypeScript checking.
|
|
714
|
+
PASS when discovery reports one extension, three tools, one command, no loader errors, host-only runtime peers, and no sibling dependency.
|
|
715
|
+
Verify README coverage of the required usage and retention caveats.
|
|
716
|
+
|
|
717
|
+
|
|
718
|
+
- #REQ-1 verified-by: [#AC-1](#AC-1)
|
|
719
|
+
|
|
720
|
+
- #REQ-2 verified-by: [#AC-2](#AC-2)
|
|
721
|
+
|
|
722
|
+
- #REQ-3 verified-by: [#AC-3](#AC-3)
|
|
723
|
+
|
|
724
|
+
- #REQ-4 verified-by: [#AC-4](#AC-4)
|
|
725
|
+
|
|
726
|
+
- #REQ-5 verified-by: [#AC-5](#AC-5)
|
|
727
|
+
|
|
728
|
+
- #REQ-6 verified-by: [#AC-6](#AC-6)
|
|
729
|
+
|
|
730
|
+
- #REQ-7 verified-by: [#AC-7](#AC-7)
|
|
731
|
+
|
|
732
|
+
- #REQ-8 verified-by: [#AC-8](#AC-8)
|
|
733
|
+
|
|
734
|
+
- #REQ-9 verified-by: [#AC-9](#AC-9)
|
|
735
|
+
|
|
736
|
+
|
|
737
|
+
<a id="AC-10"></a>
|
|
738
|
+
|
|
739
|
+
**AC-10**
|
|
740
|
+
|
|
741
|
+
Changed behavior: invoke /advisor in TUI mode with 100 available models and a configured advisor selection.
|
|
742
|
+
PASS when the actual command-selected component renders inline, selects and marks the current model, and shows at most 10 model rows.
|
|
743
|
+
Exercise provider/identifier/name fuzzy queries, input editing with j and k, clearing, arrow wrap, and an empty result set.
|
|
744
|
+
PASS when each query selects the Pi-ranked best match, no-match Enter stays open, and clearing restores results.
|
|
745
|
+
Render at heights 60, 24, 10, 6, 3, and 1 and widths 30 and 80. Repeat with long Unicode identifiers.
|
|
746
|
+
PASS when rows and widths fit, selection remains visible, and saved identifiers remain complete. Verify focus reaches the search input.
|
|
747
|
+
Cancel with Escape and Ctrl+C after navigation. PASS when global and project settings stay byte-identical and no model call occurs.
|
|
748
|
+
Preserved behavior: existing AC-1 through AC-9 checks MUST pass before and after this change. Non-TUI dialogs MUST keep their existing labels and persistence behavior.
|
|
749
|
+
|
|
750
|
+
|
|
751
|
+
- #REQ-10 verified-by: [#AC-10](#AC-10)
|
|
752
|
+
|
|
753
|
+
|
|
754
|
+
<a id="AC-11"></a>
|
|
755
|
+
|
|
756
|
+
**AC-11**
|
|
757
|
+
|
|
758
|
+
Changed behavior: cross each input, reply, history, turn, and session bound by one unit.
|
|
759
|
+
PASS when error.code stays limit-exceeded and error.limit reports the resource, maximum, and actual measurement.
|
|
760
|
+
PASS when each fixed message gives its resource-specific recovery action and no rejected exchange changes history.
|
|
761
|
+
Test multibyte input, pre-request failures, and completed reply failures with reported usage.
|
|
762
|
+
|
|
763
|
+
|
|
764
|
+
<a id="AC-12"></a>
|
|
765
|
+
|
|
766
|
+
**AC-12**
|
|
767
|
+
|
|
768
|
+
Changed behavior: test preflight failure, successful reply, completed invalid reply, provider exception, cancellation, and timeout.
|
|
769
|
+
PASS when usageComplete follows REQ-12 and failures for existing consultations include independent cumulative snapshots.
|
|
770
|
+
Test ignored abort, late success, busy overlap, and success after unknown usage.
|
|
771
|
+
PASS when unknown cumulative usage stays incomplete, no usage is counted twice, and late work cannot change state.
|
|
772
|
+
Validate final schemas and agreement between content, details, and structuredContent.
|
|
773
|
+
|
|
774
|
+
|
|
775
|
+
<a id="AC-13"></a>
|
|
776
|
+
|
|
777
|
+
**AC-13**
|
|
778
|
+
|
|
779
|
+
Changed behavior: discover a successful consultation, a pending attempt, a rejected completed reply, and an aborted attempt.
|
|
780
|
+
PASS when discovery reports known usage and exact committed-history capacity, without full message or advice content.
|
|
781
|
+
Mutate returned nested objects, then discover again. PASS when manager state is unchanged.
|
|
782
|
+
Close and reset. PASS when removed identifiers and their metadata disappear.
|
|
783
|
+
|
|
784
|
+
|
|
785
|
+
<a id="AC-14"></a>
|
|
786
|
+
|
|
787
|
+
**AC-14**
|
|
788
|
+
|
|
789
|
+
Invoke registered tools and their actual renderers with mock themes at widths 1, 8, 30, and 80.
|
|
790
|
+
PASS when call, success, failure, list, close, missing-details, and partial views fit widths and neutralize data control characters.
|
|
791
|
+
PASS when collapsed and expanded advice follow AC-19. Expansion MUST add model and cumulative usage details. Theme changes MUST appear after invalidation.
|
|
792
|
+
Hold model preparation and completion with deferred mocks. PASS when updates report only phases in their defined order.
|
|
793
|
+
Cancel, fail, exceed limits, throw from an update callback, and settle a late reply.
|
|
794
|
+
PASS when updates never expose unvalidated text, no late update occurs, callback errors do not alter outcomes, and final JSON remains coherent. AC-18 covers absence of duplicate messages. AC-19 replaces its former full-response default checks.
|
|
795
|
+
|
|
796
|
+
|
|
797
|
+
<a id="AC-15"></a>
|
|
798
|
+
|
|
799
|
+
**AC-15**
|
|
800
|
+
|
|
801
|
+
Changed behavior: inspect the live verification procedure without running provider work.
|
|
802
|
+
PASS when it specifies separate spending authorization, a designated model, two requests, no retries, failure stop, pinning and usage checks, close, and reporting.
|
|
803
|
+
PASS when the procedure discloses no local monetary cap and offline verification remains provider-free.
|
|
804
|
+
This criterion checks procedure coverage, not live interoperability or reasoning quality.
|
|
805
|
+
|
|
806
|
+
|
|
807
|
+
- #REQ-11 verified-by: [#AC-11](#AC-11)
|
|
808
|
+
|
|
809
|
+
- #REQ-12 verified-by: [#AC-12](#AC-12)
|
|
810
|
+
|
|
811
|
+
- #REQ-13 verified-by: [#AC-13](#AC-13)
|
|
812
|
+
|
|
813
|
+
- #REQ-14 verified-by: [#AC-14](#AC-14)
|
|
814
|
+
|
|
815
|
+
- #REQ-15 verified-by: [#AC-15](#AC-15)
|
|
816
|
+
|
|
817
|
+
|
|
818
|
+
<a id="AC-16"></a>
|
|
819
|
+
|
|
820
|
+
**AC-16**
|
|
821
|
+
|
|
822
|
+
Inject abnormal completion, unsupported/null/nonstring content, and absent text. PASS when invalid-response has the specified fixed diagnostic and includes no reply, thinking, exception, or provider field value.
|
|
823
|
+
Inject failures during creation and continuation. PASS when no failed pair commits, no retry occurs, and completed rejected usage remains counted. Validate tool content, details, and structuredContent against the current output schema.
|
|
824
|
+
Thinking stripping, byte limits, and provider error/abort handling MUST remain unchanged. AC-17 replaces the former marker rejection checks.
|
|
825
|
+
Use no provider calls. These checks MUST NOT claim live interoperability.
|
|
826
|
+
|
|
827
|
+
|
|
828
|
+
- #REQ-16 verified-by: [#AC-16](#AC-16)
|
|
829
|
+
|
|
830
|
+
|
|
831
|
+
<a id="AC-17"></a>
|
|
832
|
+
|
|
833
|
+
**AC-17** Plain advice without status extraction
|
|
834
|
+
|
|
835
|
+
Changed behavior: inject ordinary unmarked advice and marker-looking text that is quoted, repeated, unknown, misplaced, or the only text. Include code fences, JSON-looking text, surrounding whitespace, and multiple text blocks around thinking.
|
|
836
|
+
PASS when every nonblank normally completed text reply returns and commits the same joined text without interpreting its content. Inspect a subsequent continuation to confirm thinking is absent and text is preserved.
|
|
837
|
+
PASS when success results, discovery metadata, strict schemas, prompt, guidance, collapsed/expanded renderers, and README have no required marker or inferred status. Old status fields MUST fail the new strict schemas.
|
|
838
|
+
Preserved behavior: blank/thinking-only replies, malformed or unsupported blocks, tool calls, non-stop completions, and byte overflow MUST reject without committing. Usage and no-retry assertions MUST remain. Existing isolation, settings, picker, lifecycle, limits, accounting, and phase-only presentation checks MUST pass before and after, with obsolete marker fixtures replaced by actual invalid completions.
|
|
839
|
+
Run only offline checks for this implementation. A successful live request after reload remains a separate authorized check.
|
|
840
|
+
|
|
841
|
+
|
|
842
|
+
- #REQ-17 verified-by: [#AC-17](#AC-17)
|
|
843
|
+
|
|
844
|
+
|
|
845
|
+
<a id="AC-18"></a>
|
|
846
|
+
|
|
847
|
+
**AC-18** Full user-visible response
|
|
848
|
+
|
|
849
|
+
AC-19 supersedes the former changed-behavior checks for full advice in the default collapsed view. Expanded advice MUST remain complete at widths 8, 30, and 80. At width 1, the ASCII tail MUST remain visible and all lines MUST fit. Double-width graphemes MAY be clipped by the host width helper.
|
|
850
|
+
Preserved behavior: snapshot tool content, details, structuredContent, usage, consultation count, and mock request count before rendering. PASS when repeated renders and expansion leave them unchanged and send no custom/user messages. Partial/failure/fallback views MUST exclude unvalidated text and all views MUST retain width/control protections.
|
|
851
|
+
All 55 existing checks MUST pass before the change. Afterward, all unaffected checks MUST pass, and the old bounded-preview assertion MUST be replaced by the full-response assertion. No provider call or live terminal check is part of these offline tests.
|
|
852
|
+
|
|
853
|
+
|
|
854
|
+
- #REQ-18 verified-by: [#AC-18](#AC-18)
|
|
855
|
+
|
|
856
|
+
|
|
857
|
+
<a id="REQ-19"></a>
|
|
858
|
+
|
|
859
|
+
**REQ-19** Collapse long advisor replies
|
|
860
|
+
|
|
861
|
+
The advisor renderer MUST wrap control-neutralized checked advice at the current terminal width before selecting preview rows.
|
|
862
|
+
When advice occupies at most eight wrapped rows, the collapsed view MUST show all advice without an expansion hint.
|
|
863
|
+
When advice occupies more than eight wrapped rows, the collapsed view MUST show exactly the first eight advice rows.
|
|
864
|
+
It MUST then show the separate hint: … Expand for full advice.
|
|
865
|
+
Header, usage, and hint rows MUST NOT count toward the eight advice rows. No character-count cap MUST apply.
|
|
866
|
+
The expanded view MUST show all checked advice and the existing model and cumulative usage metadata, without a preview hint.
|
|
867
|
+
Terminal resize MUST recompute the wrapped row count. REQ-14 width and control protections MUST remain unchanged.
|
|
868
|
+
REQ-18 prohibitions on duplicate messages, transcript entries, parent turns, result mutation, and unchecked advice MUST remain unchanged.
|
|
869
|
+
The README MUST state the eight-row boundary and expanded behavior. No provider request MUST be needed to expand a reply.
|
|
870
|
+
|
|
871
|
+
|
|
872
|
+
<a id="AC-19"></a>
|
|
873
|
+
|
|
874
|
+
**AC-19** Short replies and bounded long previews
|
|
875
|
+
|
|
876
|
+
Changed behavior: render nine newline-separated rows and long wrapped Unicode advice at widths 1, 8, 30, and 80.
|
|
877
|
+
PASS when collapsed advice contains exactly the first eight wrapped rows and a separate expansion hint, with the omitted tail absent.
|
|
878
|
+
PASS when expanded advice remains complete, subject to the one-column grapheme exception in AC-18, and includes existing metadata without a hint.
|
|
879
|
+
Render the same response across terminal widths. PASS when resize recomputes whether the preview is needed.
|
|
880
|
+
PASS when README coverage states the eight-row boundary and full expanded advice.
|
|
881
|
+
Preserved behavior: render one-row and exactly eight-row replies, including a reply longer than 160 characters that fits eight rows.
|
|
882
|
+
PASS when collapsed advice is complete with no hint. Repeated renders MUST preserve results, state, model request counts, and message counts as AC-18 defines.
|
|
883
|
+
Partial/failure/fallback privacy, control neutralization, width protections, and theme invalidation MUST remain unchanged.
|
|
884
|
+
Baseline MUST pass all 59 existing checks before updating the obsolete full-default assertions. All unaffected checks MUST pass afterward.
|
|
885
|
+
Use only mock replies and offline checks. No live terminal or provider behavior MUST be inferred from these checks.
|
|
886
|
+
|
|
887
|
+
|
|
888
|
+
- #REQ-19 verified-by: [#AC-19](#AC-19)
|
|
889
|
+
|
|
890
|
+
- #REQ-19 supersedes: [#REQ-18](#REQ-18)
|
|
891
|
+
|
|
892
|
+
|
|
893
|
+
<a id="AC-20"></a>
|
|
894
|
+
|
|
895
|
+
**AC-20** Reasoning settings, commands, and pinned requests
|
|
896
|
+
|
|
897
|
+
Changed behavior: parse all eight reasoning values, reject invalid values, and test complete project-model replacement of global reasoning.
|
|
898
|
+
Query and set reasoning through the actual registered command in UI and non-UI modes.
|
|
899
|
+
PASS when queries show effective and supported levels without writes, and scoped writes preserve the correct complete selection.
|
|
900
|
+
Invalid/unsupported levels, missing models, lost trust, and untrusted project writes MUST leave settings unchanged and generate nothing.
|
|
901
|
+
Start and continue actual advisor tools using mock models with capability metadata. Change settings between requests.
|
|
902
|
+
PASS when the model and effort remain pinned, explicit supported levels reach every request, and schema-valid results/discovery retain them.
|
|
903
|
+
PASS when unsupported starts/continuations generate nothing, commit no failed exchange, and return the fixed error.
|
|
904
|
+
Omission, default, and supported off MUST omit the SDK reasoning option. The tool schema MUST reject caller reasoning.
|
|
905
|
+
Expanded output MUST show configured effort or legacy default without revealing thinking.
|
|
906
|
+
Preserved behavior: all 62 existing tests and TypeScript MUST pass before implementation and all unaffected checks MUST pass afterward.
|
|
907
|
+
The existing settings locks, trust, model picker, request bounds, thinking removal, rollback, usage, and no-retry checks MUST remain unchanged.
|
|
908
|
+
|
|
909
|
+
|
|
910
|
+
<a id="AC-21"></a>
|
|
911
|
+
|
|
912
|
+
**AC-21** Contextual and side-effect-free command completions
|
|
913
|
+
|
|
914
|
+
Changed behavior: drive the registered callback through real extension lifecycle hooks with mock catalogs and temporary settings.
|
|
915
|
+
PASS when root/scoped grammar, provider prefixes, fuzzy model names/identifiers, and supported scoped reasoning match REQ-21.
|
|
916
|
+
Check fresh catalogs/settings, malformed settings, trust changes, virtual entries, unknown syntax, and lookup errors.
|
|
917
|
+
PASS when dynamic failures return null silently and static suggestions remain usable.
|
|
918
|
+
Use actual Pi CombinedAutocompleteProvider getSuggestions and applyCompletion with cursor suffixes, tabs, and trailing separators.
|
|
919
|
+
PASS when completion inserts the full argument replacement, retains the suffix, positions the cursor correctly, and appends a space.
|
|
920
|
+
Start completion, then replace the context through start/tree/shutdown hooks. PASS when pending results are discarded and shutdown retains only static choices.
|
|
921
|
+
Snapshot settings, notifications, provider/tool calls, and active-model changes. PASS when completion changes none of them.
|
|
922
|
+
These provider checks MUST NOT claim physical keyboard routing. README MUST identify the Pi 1.0.4 Tab limitation and the typing workaround.
|
|
923
|
+
Use only offline checks and temporary configuration. No paid provider request or live settings change is part of this implementation.
|
|
924
|
+
|
|
925
|
+
|
|
926
|
+
- #REQ-20 verified-by: [#AC-20](#AC-20)
|
|
927
|
+
|
|
928
|
+
- #REQ-21 verified-by: [#AC-21](#AC-21)
|
|
929
|
+
|
|
930
|
+
|
|
931
|
+
<a id="AC-22"></a>
|
|
932
|
+
|
|
933
|
+
**AC-22** Private failure phase and selection diagnostics
|
|
934
|
+
|
|
935
|
+
Inject unexpected preparation exceptions and zero-usage terminal error responses with secret-bearing messages, thinking, and payloads. PASS when diagnostics distinguish preparation/local-error from response-validation/provider-error, both retain provider-failed and usageComplete = true, no secrets appear, and creation rolls back.
|
|
936
|
+
Test omission, false, true, null, string, number, and object diagnostics inputs. PASS when only booleans are valid, omission/false retain the default failure shape, success shapes remain unchanged, and opt-in is not pinned to a consultation.
|
|
937
|
+
Throw completion exceptions with recognized status/code fields, unrecognized or string statuses, free-form messages, accessors, and throwing descriptor traps. PASS when only allowlisted categories appear, recognized status takes precedence, getters/messages are not read, metadata errors yield unknown, and one completion attempt uses maxRetries = 0.
|
|
938
|
+
Inject terminal aborts and invalid replies on continuations. PASS when pinned model/source and reported rejected usage remain visible, prior history stays intact, no failed content enters a later request, and no retry occurs.
|
|
939
|
+
Hold preparation or completion through timeout; cancel completion. PASS when the active phase and original usage-completeness semantics are reported, late settlement changes neither returned diagnostics nor consultation state, and aborted creation releases its slot.
|
|
940
|
+
Exercise actual registered tools with global settings and a trusted project override. PASS when preparation availability/authentication failures expose the resolved project selection without generation, continuations retain the original model/source after settings change, default calls omit diagnostics, and content/details/structuredContent agree with strict schemas.
|
|
941
|
+
Render opted-in failures at widths 1, 8, 30, and 80. PASS when phase/category/model/source are visible without private errors and all rows fit. Preserve all 77 existing offline checks and TypeScript. No live provider request or billing/root-cause claim is part of this criterion.
|
|
942
|
+
|
|
943
|
+
|
|
944
|
+
- #REQ-22 verified-by: [#AC-22](#AC-22)
|
|
945
|
+
|
|
946
|
+
|
|
947
|
+
<a id="AC-24"></a>
|
|
948
|
+
|
|
949
|
+
**AC-24** Tool-free xAI request compatibility
|
|
950
|
+
|
|
951
|
+
Changed behavior: use the installed grok-4.7 catalog entry and real Responses adapter with a mock fetch.
|
|
952
|
+
Test omitted and high effort, creation, and continuation. Reject tool_choice without tools in the fixture.
|
|
953
|
+
PASS when xai payloads omit tools and tool_choice, integer timeouts and maxRetries zero remain, and checked exchanges succeed.
|
|
954
|
+
Changed behavior: inject the three exact HTTP 400 rejection details in REQ-24. PASS when the fixed code appears without raw detail.
|
|
955
|
+
Changed behavior: validate and render the fixed code at widths 1, 8, 30, and 80. PASS when the fixed explanation appears and rows fit.
|
|
956
|
+
Preserved behavior: inject partial/appended rejection strings, other providers, other statuses, and terminal aborts.
|
|
957
|
+
PASS when the new code stays absent and existing private hints remain. AC-25 replaces the former other-provider toolChoice assertion.
|
|
958
|
+
All 91 existing tests and TypeScript MUST pass before and after this change. Use synthetic keys and mock fetch, with no provider requests.
|
|
959
|
+
|
|
960
|
+
|
|
961
|
+
- #REQ-24 verified-by: [#AC-24](#AC-24)
|
|
962
|
+
|
|
963
|
+
|
|
964
|
+
<a id="AC-25"></a>
|
|
965
|
+
|
|
966
|
+
**AC-25** No tool-choice override for tool-free advisor requests
|
|
967
|
+
|
|
968
|
+
Changed behavior: start and continue consultations for mock, xai, openai, anthropic, openai-codex, and custom provider selections.
|
|
969
|
+
PASS when toolChoice is not an own property of any completion options object and every request declares zero tools.
|
|
970
|
+
Change configured selection between calls. PASS when continuations retain their original model and configured effort.
|
|
971
|
+
Changed behavior: use an installed OpenAI Responses model and the real adapter with synthetic authentication and mock fetch.
|
|
972
|
+
PASS when creation and continuation payloads omit tools and tool_choice and validated replies commit.
|
|
973
|
+
Changed behavior: update the registered-tool and isolated-session assertions and README coverage to require provider-neutral omission.
|
|
974
|
+
Preserved behavior: retain xAI default/high real-adapter cases, known-good SSE, fixed diagnostics, and malformed-reply rejection.
|
|
975
|
+
PASS when integer timeout, cancellation, maxRetries zero, token bound, cache retention, usage, history rollback, and privacy remain unchanged.
|
|
976
|
+
All 97 baseline tests and TypeScript MUST pass before the change. All unaffected checks MUST pass afterward.
|
|
977
|
+
Only obsolete toolChoice assertions MAY change. No live provider request is part of these checks.
|
|
978
|
+
|
|
979
|
+
|
|
980
|
+
- #REQ-25 verified-by: [#AC-25](#AC-25)
|
|
981
|
+
|
|
982
|
+
|
|
983
|
+
<a id="AC-26"></a>
|
|
984
|
+
|
|
985
|
+
**AC-26** Deadline configuration and safe expiry
|
|
986
|
+
|
|
987
|
+
Changed behavior: parse timeout-only and model-plus-timeout settings at both inclusive bounds.
|
|
988
|
+
PASS when invalid types and out-of-range values fail closed and timeout precedence is independent of whole-model precedence.
|
|
989
|
+
Changed behavior: exercise actual timeout query, scoped saves, default reset, show, and root/default autocomplete with temporary settings.
|
|
990
|
+
PASS when query is side-effect free, writes preserve scope-owned fields, model/effort/picker writes retain explicit timeout, and untrusted writes fail.
|
|
991
|
+
Changed behavior: save model or timeout updates with preservation enabled after another field has changed in the target layer.
|
|
992
|
+
PASS when the preserved value is taken from the locked save checkpoint; a timeout reset removes only timeoutMs.
|
|
993
|
+
Changed behavior: invoke registered advisor tools with default, global, project, and timeout-only overrides.
|
|
994
|
+
Change settings between calls. PASS when new consultations use fresh settings but continuations retain pinned timeout, model, and effort.
|
|
995
|
+
Changed behavior: mock time past 120000 ms and complete before 300000 ms. PASS when a default request succeeds without early abort.
|
|
996
|
+
Changed behavior: spend part of a configured budget in preparation. PASS when the SDK receives only the integer remainder.
|
|
997
|
+
Expire the deadline during completion. PASS when no retry or failed exchange commits and usage remains incomplete without an observed reply.
|
|
998
|
+
Changed behavior: resolve a configured budget shorter than elapsed preparation. PASS when timeout occurs before generation with known zero usage.
|
|
999
|
+
Preserved behavior: invoke available preparation callbacks after failure. PASS when no new timer or returned state change occurs.
|
|
1000
|
+
Changed behavior: update only obsolete 120000 ms SDK-bound and root-completion assertions, and verify README deadline coverage.
|
|
1001
|
+
Preserved behavior: caller cancellation, ignored abort with late completion, strict tool schemas, settings locks/trust, isolation, diagnostics, and accounting MUST pass.
|
|
1002
|
+
All 100 baseline tests and TypeScript MUST pass before changes. All unaffected checks MUST pass afterward.
|
|
1003
|
+
Use mock responses, clocks, timers, and temporary settings. No paid call or live settings mutation is part of verification.
|
|
1004
|
+
|
|
1005
|
+
|
|
1006
|
+
- #REQ-26 verified-by: [#AC-26](#AC-26)
|
|
1007
|
+
|
|
1008
|
+
|
|
1009
|
+
<a id="AC-27"></a>
|
|
1010
|
+
|
|
1011
|
+
**AC-27** Cache options, stable prefixes, and provider payloads
|
|
1012
|
+
|
|
1013
|
+
Changed behavior: start and continue consultations for mock, xai, openai, anthropic, openai-codex, and custom selections.
|
|
1014
|
+
PASS when every dispatch uses cacheRetention short. Update only the obsolete none assertion in the existing provider matrix.
|
|
1015
|
+
Changed behavior: forward manager options to installed OpenAI Responses and Anthropic adapters with synthetic authentication and mock fetch.
|
|
1016
|
+
PASS when OpenAI sends the consultation prompt_cache_key and no extended retention, and Anthropic sends ephemeral cache_control without one-hour ttl.
|
|
1017
|
+
Changed behavior: inspect the README. PASS when it documents short caching, provider-dependent cache hits, and local-close retention limits.
|
|
1018
|
+
Preserved behavior: create two consultations and continue the first after changing the prepared model.
|
|
1019
|
+
PASS when the first identifier stays stable, the second differs, the model stays pinned, and prior message text remains the continuation prefix.
|
|
1020
|
+
PASS when zero tools, omitted toolChoice, zero retries, token bound, positive integer timeout, cancellation signal, and usage totals remain intact.
|
|
1021
|
+
All 112 baseline tests and TypeScript MUST pass before implementation. All unaffected checks MUST pass afterward.
|
|
1022
|
+
These checks MUST use mocks only. They MUST NOT establish claims about live cache hits, provider lifetime, or savings.
|
|
1023
|
+
|
|
1024
|
+
|
|
1025
|
+
- #REQ-27 verified-by: [#AC-27](#AC-27)
|
|
1026
|
+
|
|
1027
|
+
|
|
1028
|
+
<a id="AC-28"></a>
|
|
1029
|
+
|
|
1030
|
+
**AC-28** Opaque replay, privacy, atomicity, and bounds
|
|
1031
|
+
|
|
1032
|
+
Changed behavior: use the installed Grok Responses adapter with mock SSE containing encrypted reasoning, plaintext summaries, and distinct signed text blocks.
|
|
1033
|
+
PASS when continuations replay the encrypted reasoning and original text IDs/phases in order, with empty summary, no plaintext reasoning, stable key, store false, and no previous_response_id.
|
|
1034
|
+
Changed behavior: exercise all seven supported API identities with synthetic replies. PASS when only each API's permitted opaque fields survive.
|
|
1035
|
+
PASS when Google conversion echoes thought signatures, including empty signed parts, and Anthropic payloads echo redacted thinking but omit ordinary signed thinking.
|
|
1036
|
+
Preserved behavior: omit malformed signatures, absent encrypted content, unsupported APIs, ordinary signed plaintext thinking, and mismatched provider/model state.
|
|
1037
|
+
PASS when valid public text remains unchanged and no plaintext thinking or arbitrary metadata reaches retained state, contexts, results, or presentation.
|
|
1038
|
+
Changed behavior: measure combined serialized history and replay at the exact configured byte boundary and one byte below it.
|
|
1039
|
+
PASS when exact size commits, overflow rejects the whole pair with history-bytes, completed usage remains known, and previous state stays intact.
|
|
1040
|
+
Changed behavior: mutate returned replies and dispatched replay contexts. PASS when later replay remains unchanged and consultation-local.
|
|
1041
|
+
Changed behavior: fail or cancel a continuation with new opaque state, then continue successfully. PASS when only prior validated replay survives.
|
|
1042
|
+
Changed behavior: close or clear consultations containing replay. PASS when the removed identifiers cannot continue and new consultations have no old opaque state.
|
|
1043
|
+
Changed behavior: inspect README coverage of private opaque state, combined bounds, supported formats, and ordinary signed-thinking limitations.
|
|
1044
|
+
Preserved behavior: public advice text, strict result schemas, renderers, isolation, pinning, short caching, usage, rollback, timeout, cancellation, and request controls MUST pass.
|
|
1045
|
+
All 116 baseline tests and TypeScript MUST pass before implementation. Every changed check MUST fail for its intended missing behavior first.
|
|
1046
|
+
Verification MUST use synthetic state, mock fetch/SSE, and temporary settings only. Live cache-hit improvement requires separate spending authorization.
|
|
1047
|
+
|
|
1048
|
+
|
|
1049
|
+
- #REQ-28 verified-by: [#AC-28](#AC-28)
|
|
1050
|
+
|
|
1051
|
+
|
|
1052
|
+
<a id="AC-29"></a>
|
|
1053
|
+
|
|
1054
|
+
**AC-29** Model context, boundary checks, pinning and privacy
|
|
1055
|
+
|
|
1056
|
+
Changed checks MUST first fail with the existing byte-limited implementation.
|
|
1057
|
+
PASS when registered tools obtain contextWindow from the configured physical advisor model, not the parent model or parent getContextUsage.
|
|
1058
|
+
PASS when later registry/settings changes cannot change an existing consultation's window, while a new consultation uses the new window.
|
|
1059
|
+
PASS when missing, zero, fractional, negative, or nonfinite windows fail before generation.
|
|
1060
|
+
PASS when private replay exceeds 49152 serialized bytes and three or more pairs still commit within the model window.
|
|
1061
|
+
PASS when provider total usage supplies context tokens, including reasoning and cached input, without adding ciphertext length or cumulative totals.
|
|
1062
|
+
PASS when totalTokens zero uses Pi's input/output/cache fallback, and all-zero usage uses Pi's message estimates.
|
|
1063
|
+
PASS when pending context plus4096 passes exactly at the window and fails one token over without dispatch.
|
|
1064
|
+
PASS when committed context exactly equals the window and fails one token over without committing the pair.
|
|
1065
|
+
PASS when completion overflow retains prior context and opaque replay while accounting for rejected usage.
|
|
1066
|
+
PASS when Advice/List/failure schemas accept detached contextUsage and reject old byte-remaining or unknown fields.
|
|
1067
|
+
Preserved checks MUST pass before and after: isolated contexts, no opaque state in public results, reply validation, request controls, deadline/cancellation, atomic rollback, and existing non-history resource limits.
|
|
1068
|
+
Update the old history-byte checks only to their explicitly superseding token-window criteria. Do not weaken unchanged assertions.
|
|
1069
|
+
Baseline MUST pass all125 tests and TypeScript. Verification MUST use mock catalogs, synthetic state, temporary settings, and no paid requests.
|
|
1070
|
+
|
|
1071
|
+
|
|
1072
|
+
- #REQ-29 verified-by: [#AC-29](#AC-29)
|
|
1073
|
+
|
|
1074
|
+
|
|
1075
|
+
<a id="AC-30"></a>
|
|
1076
|
+
|
|
1077
|
+
**AC-30** Visible input and consultation status rendering
|
|
1078
|
+
|
|
1079
|
+
Changed checks MUST first fail with the old renderers or README.
|
|
1080
|
+
PASS when the call view shows a multiline Unicode message in full after control sanitization, without changing the submitted value.
|
|
1081
|
+
PASS when two calls produce cumulative ↑/↓/R/W/cost values, but CH uses only the latest call and context uses only retained advisor context.
|
|
1082
|
+
PASS when the line appears in collapsed and expanded results without sub/auto labels.
|
|
1083
|
+
PASS when zero or unobserved prompt usage shows CH?, incomplete totals have a warning, and historic results without context show ?/?.
|
|
1084
|
+
PASS when known cumulative failures show reported totals with prior context, and expanded discovery shows context instead of byte capacity.
|
|
1085
|
+
PASS when widths1/8/30/80/160, resizing, themes, and control sanitization remain safe.
|
|
1086
|
+
PASS when rendering mutates no result, sends no duplicate message, starts no model work, and exposes no private opaque data or unchecked advice.
|
|
1087
|
+
PASS when README explains questions, estimated model-sized context, per-call CH versus cumulative totals, and unchanged resource limits.
|
|
1088
|
+
Preserved checks MUST pass before and after: phase-only progress, eight advice-row previews, complete expansion, strict model-facing envelopes, and safe fallback/failure rendering.
|
|
1089
|
+
The old explicit-input omission assertion MUST be replaced with the user-authorized visibility check. No other privacy assertion may be relaxed.
|
|
1090
|
+
|
|
1091
|
+
|
|
1092
|
+
- #REQ-30 verified-by: [#AC-30](#AC-30)
|
|
1093
|
+
|
|
1094
|
+
|
|
1095
|
+
## Verification scope {#verification}
|
|
1096
|
+
|
|
1097
|
+
The initial 12 acceptance checks failed against the minimal baseline. Failures identified missing consultation behavior, settings behavior, or registration.
|
|
1098
|
+
|
|
1099
|
+
All 12 checks then passed with implementation. TypeScript checking passed. A clean source copy also passed after npm ci --ignore-scripts --offline.
|
|
1100
|
+
|
|
1101
|
+
No preserved-behavior baseline applies. This package did not exist before the task.
|
|
1102
|
+
|
|
1103
|
+
| Criterion | Check names in tests/ | Before | After |
|
|
1104
|
+
| --- | --- | --- | --- |
|
|
1105
|
+
| AC-1 | isolated starts and continuations pin the selected model; tool boundary uses explicit context | fail | pass |
|
|
1106
|
+
| AC-2 | reply protocol validates markers and does not expose thinking | fail | pass |
|
|
1107
|
+
| AC-3 | atomic failures, usage, busy, abort, timeout and late completions | fail | pass |
|
|
1108
|
+
| AC-4 | hard byte, pair, transcript, and session limits preserve state | fail | pass |
|
|
1109
|
+
| AC-5 | metadata-only discovery, busy close, deletion and reset; clears on lifecycle, not compaction | fail | pass |
|
|
1110
|
+
| AC-6 | strict settings precedence, trust, bounds, symlinks and special files; tool preflight rejects unavailable selection | fail | pass |
|
|
1111
|
+
| AC-7 | atomic settings saves, validation, trust and exclusive locks; user-only configuration | fail | pass |
|
|
1112
|
+
| AC-8 | registration, strict schemas, coherent results and usage | fail | pass |
|
|
1113
|
+
| AC-9 | standalone loader discovers extension without provider work and README covers boundaries | fail | pass |
|
|
1114
|
+
|
|
1115
|
+
Additional failing-then-passing cases verified an extra unknown marker, a missing identifier on oversized continuation errors, and non-UI command reporting.
|
|
1116
|
+
|
|
1117
|
+
Commands run in the project: npm run verify and npm pack --dry-run --ignore-scripts. Package discovery used the complete package directory, not only the entry-point file.
|
|
1118
|
+
|
|
1119
|
+
The nine new-coder profile check scripts passed. Both documented profile core TypeScript commands passed.
|
|
1120
|
+
|
|
1121
|
+
These results cover deterministic boundaries and Pi 1.0.4 package loading on Node 24, on Linux. They do not establish reasoning quality or live provider interoperability.
|
|
1122
|
+
|
|
1123
|
+
No paid model request, live terminal session, hosted CI run, Windows check, or macOS check occurred. Provider-side retention remains outside this extension's control.
|
|
1124
|
+
|
|
1125
|
+
Settings safety is checkpoint-based. An external writer can race between the final byte comparison and replacement. Parent-directory races remain outside the guarantee.
|
|
1126
|
+
|
|
1127
|
+
No unresolved implementation decision blocks this initial version. Resource defaults and marker syntax remain revisable before a public release.
|
|
1128
|
+
|
|
1129
|
+
### Model picker correction {#model-picker-verification}
|
|
1130
|
+
|
|
1131
|
+
The user requested the /model interface and named pi-magic8ball as the reference.
|
|
1132
|
+
|
|
1133
|
+
REQ-10 and AC-10 replace the generic TUI selection dialog with the sibling extension's inline searchable interface.
|
|
1134
|
+
The copied picker uses public Pi TUI components. It does not instantiate Pi's session-bound ModelSelectorComponent or refresh remote catalogs.
|
|
1135
|
+
|
|
1136
|
+
Preserved baseline: all 12 existing AC-1 through AC-9 checks passed before and after this change.
|
|
1137
|
+
Changed behavior: all four AC-10 checks failed before implementation because the command used the generic dialog.
|
|
1138
|
+
All four then passed. The complete suite now passes 16 checks. TypeScript checking also passes.
|
|
1139
|
+
|
|
1140
|
+
AC-10 check names in tests/model-picker.test.mjs:
|
|
1141
|
+
|
|
1142
|
+
- inline picker selects current model, caps rows, forwards focus and handles resize
|
|
1143
|
+
- fuzzy search matches names and provider/model tokens with editable j/k
|
|
1144
|
+
- no-match Enter stays open, clear restores results, arrows wrap and cancellation writes nothing
|
|
1145
|
+
- Unicode truncation preserves complete saved identifiers
|
|
1146
|
+
|
|
1147
|
+
A clean source copy passed npm ci --ignore-scripts --offline and npm run verify.
|
|
1148
|
+
The tests construct and exercise the actual command-selected component with mocked terminal dimensions.
|
|
1149
|
+
These checks do not establish appearance in a live terminal. No paid model calls or profile settings changes occurred.
|
|
1150
|
+
|
|
1151
|
+
The picker preserves the prior provider/model labels for RPC and other non-TUI UI clients.
|
|
1152
|
+
No additional requirement or unresolved decision controls this change.
|
|
1153
|
+
|
|
1154
|
+
### Diagnostics, usage, and presentation improvements {#diagnostics-verification}
|
|
1155
|
+
|
|
1156
|
+
The user requested a baseline commit and implementation of the five reviewed improvements.
|
|
1157
|
+
Commit c28342f records the previous extension. All 16 preserved checks and TypeScript passed before implementation.
|
|
1158
|
+
REQ-11 through REQ-15 and AC-11 through AC-15 define the change. REQ-5 and REQ-8 now reference the added metadata.
|
|
1159
|
+
The marker protocol, isolation, settings, model picker, limits, and no-retry policy remain unchanged.
|
|
1160
|
+
|
|
1161
|
+
Changed checks in tests/improvements.test.mjs failed before implementation on absent fields, renderers, progress, and procedure coverage.
|
|
1162
|
+
All 17 changed checks, including three subtests, then passed. The complete suite passes 33 checks and TypeScript.
|
|
1163
|
+
A final copy of the new tests also fails all 17 checks against archived baseline c28342f.
|
|
1164
|
+
That offline comparison shares the installed development dependencies. It is not a clean-install verification.
|
|
1165
|
+
|
|
1166
|
+
| Criterion | Test names in tests/improvements.test.mjs | Before | After |
|
|
1167
|
+
| --- | --- | --- | --- |
|
|
1168
|
+
| AC-11 | oversized multibyte input; oversized reply; history preflight and pair overflow; exhausted turns; exhausted sessions | fail | pass |
|
|
1169
|
+
| AC-12 | preflight, success and completed rejected replies; unknown usage after exception/abort/timeout; cancellation during preparation; busy accounting | fail | pass |
|
|
1170
|
+
| AC-13 | committed capacity and independent usage snapshots, not transcripts | fail | pass |
|
|
1171
|
+
| AC-14 | registered renderers and expansion; phase-only progress; failure views; cancellation during preparation | fail | pass |
|
|
1172
|
+
| AC-15 | opt-in live procedure authorization, bounded calls, failure stop and reporting | fail | pass |
|
|
1173
|
+
|
|
1174
|
+
Commands completed: npm run verify, npm pack --dry-run --ignore-scripts, and git diff --check in the project worktree.
|
|
1175
|
+
All nine new-coder profile check scripts and both profile core TypeScript commands also passed.
|
|
1176
|
+
The render tests use actual registered components with plain mock themes at widths 1, 8, 30, and 80.
|
|
1177
|
+
Pi's width helper can emit its own ANSI reset. Tests permit that reset but reject other controls with these mock themes.
|
|
1178
|
+
Deferred test requests have explicit cleanup so failed assertions cannot leave the 120-second deadline active.
|
|
1179
|
+
|
|
1180
|
+
No live provider calls, live terminal appearance check, hosted CI, Windows check, or macOS check ran for this implementation.
|
|
1181
|
+
The live procedure requires a designated model and separate spending authorization. A provider-side cap controls any required monetary ceiling.
|
|
1182
|
+
This verification does not establish provider interoperability, billing completeness, or reasoning quality.
|
|
1183
|
+
|
|
1184
|
+
### Live rejection investigation and safe diagnostics {#reply-diagnostics-verification}
|
|
1185
|
+
|
|
1186
|
+
Two user-authorized calls after the improvements returned the same generic invalid-response error. Neither created a consultation. Their raw replies were discarded, so the exact validation failures remain unknown.
|
|
1187
|
+
The installed Pi 1.0.4 runtime normalizes the explicit system message without replacing it. Its OpenAI Codex adapter uses that leading message as request instructions. Source inspection found no prompt-loss cause.
|
|
1188
|
+
|
|
1189
|
+
The user then requested diagnosis and a fix. REQ-16 and AC-16 add fixed rejection messages without result fields, raw reply exposure, weaker markers, or automatic retry.
|
|
1190
|
+
The offline checks reproduced the missing diagnostics. They also confirmed that a null content block threw during validation and was misclassified as provider-failed.
|
|
1191
|
+
That malformed-content defect is fixed. It is not established as the cause of either live failure.
|
|
1192
|
+
|
|
1193
|
+
All 33 preserved checks passed before implementation. All 13 AC-16 checks in tests/reply-diagnostics.test.mjs failed before implementation for generic diagnostics or the wrong failure code. All 13 then passed. The complete suite passes 46 checks and TypeScript through npm run verify. git diff --check also passed.
|
|
1194
|
+
The changed checks cover abnormal completion, unsupported/null/nonstring content, absent text, missing/repeated/unknown/misplaced/quoted markers, empty advice, failed continuation accounting, and coherent schema-valid tool results.
|
|
1195
|
+
|
|
1196
|
+
No new provider request ran during this investigation. A reloaded live call with separate spending authorization is still needed to identify the actual rejection and decide whether a provider, prompt, or parser fix is justified. No successful live consultation, continuation, or terminal appearance is claimed.
|
|
1197
|
+
|
|
1198
|
+
### Remove the model output protocol {#plain-advice-verification}
|
|
1199
|
+
|
|
1200
|
+
After reload, one separately authorized diagnostic call returned invalid-response with a marker-placement message. It reported 356 tokens, cost 0.0000528 dollars, and usageComplete = true. No exchange committed and no retry ran.
|
|
1201
|
+
The user then rejected reliance on model-generated structured output and authorized removal of the marker protocol. REQ-17 and AC-17 replace the former marker and status clauses. The prompt, reply extraction, state, schemas, renderers, and README now treat advice as unstructured text. The extension still constructs the strict tool envelope.
|
|
1202
|
+
|
|
1203
|
+
Baseline: all 46 checks and TypeScript passed before this change. All 15 AC-17 checks in tests/plain-advice.test.mjs failed before implementation for marker rejection, status fields, or old prompt/documentation. All 15 then passed.
|
|
1204
|
+
Six obsolete marker-rejection checks were replaced by acceptance cases in AC-17. The retained completion/content/privacy/rollback/accounting checks still pass. Accounting tests now inject truncation rather than marker omission as the rejection cause; default mock replies are unmarked. The revised suite passes all 55 checks and TypeScript through npm run verify. git diff --check passes.
|
|
1205
|
+
|
|
1206
|
+
AC-17 checks cover nine text forms, exact history retention, whitespace and multiple text blocks, thinking removal, absence of classifications, strict status-free schemas, actual registered tools and renderers, prompt/guidance, and README coverage.
|
|
1207
|
+
These offline results establish removal of the marker-placement failure path, not successful live interoperability. No provider request ran after removal. Reload is required to exercise the new runtime, and another live check requires separate spending authorization. No commit, settings change, or live terminal verification is claimed.
|
|
1208
|
+
|
|
1209
|
+
### Full responses visible by default {#response-visibility-verification}
|
|
1210
|
+
|
|
1211
|
+
A separately authorized post-removal live check succeeded for a start and continuation using openai-codex/gpt-6-luna. The model selection stayed the same across both calls, turns reached two, known cumulative usage was 674 tokens and 0.0001058 dollars, and completeness flags were true. The idle session closed and discovery returned no active sessions. No retry occurred. This establishes that one live two-call flow worked, not universal provider support or pinning under a settings change.
|
|
1212
|
+
|
|
1213
|
+
The user then asked to make advisor responses visible to the user. The existing default renderer capped advice at 160 characters and three rows. REQ-18/AC-18 remove those caps; full checked advice appears in both default and expanded views. Expansion adds metadata only. The existing tool result carries the response; no duplicate custom message, parent turn, or extra transcript entry is added.
|
|
1214
|
+
|
|
1215
|
+
All 55 existing checks and TypeScript passed before implementation. In tests/response-visibility.test.mjs, the two changed checks failed before for omitted advice and outdated README coverage. The two preserved checks passed before. All four pass after. The old bounded-preview assertion now checks the full response. The complete suite passes 59 checks and TypeScript via npm run verify; git diff --check passes.
|
|
1216
|
+
Changed checks: full response visible beyond character/row caps; README full-response coverage. Preserved checks: no duplicate messages, state/result mutation, or new model work; no unchecked text in partial/failure/fallback views.
|
|
1217
|
+
Tests exercise actual registered renderers at widths 1, 8, 30, and 80 with Unicode, long lines, and a unique tail. One-column terminals cannot show double-width graphemes; the host helper clips them while the ASCII tail remains visible. No live terminal or HTML export check was performed for this display change, and no provider call ran during its implementation.
|
|
1218
|
+
|
|
1219
|
+
### Collapse long replies after the display demonstration {#long-reply-preview-verification}
|
|
1220
|
+
|
|
1221
|
+
The user confirmed that a short live reply appeared in full without expansion after reload. A separately authorized eight-section reply then demonstrated the full-default renderer. Both consultations closed. These demonstrations used the previous renderer, not the new preview behavior.
|
|
1222
|
+
|
|
1223
|
+
The user accepted the proposed split: full short replies, collapsed previews for long replies, and full expanded advice. The implementation uses eight wrapped advice rows as its cutoff. REQ-19 supersedes only the earlier collapsed full-response rule. REQ-14, REQ-18, AC-14, and AC-18 now reference that change. Earlier verification sections describe their historical implementations.
|
|
1224
|
+
|
|
1225
|
+
Baseline: all 59 checks and TypeScript passed before changes. Five changed checks failed before implementation and passed afterward. Three focused preserved checks passed before and after. The revised full suite passes 62 checks and TypeScript with npm run verify.
|
|
1226
|
+
|
|
1227
|
+
| Criterion | Check names | Before | After |
|
|
1228
|
+
| --- | --- | --- | --- |
|
|
1229
|
+
| AC-19 | long Unicode and wrapped replies preview eight rows with full expansion | fail | pass |
|
|
1230
|
+
| AC-19 | ninth advice row triggers preview without counting headers or hint | fail | pass |
|
|
1231
|
+
| AC-19 | resize recomputes whether advice needs a preview | fail | pass |
|
|
1232
|
+
| AC-19 | README documents eight-row previews and full expanded advice | fail | pass |
|
|
1233
|
+
| AC-14, AC-19 | registered renderers show previews, full expansion, metadata and safe fallback | fail | pass |
|
|
1234
|
+
| AC-19 | one-row, eight-row and over-160-character short replies stay fully visible | pass | pass |
|
|
1235
|
+
| AC-18 | rendering does not duplicate messages, mutate results, or run model work | pass | pass |
|
|
1236
|
+
| AC-18 | partial, failed, and missing details never reveal unchecked advice | pass | pass |
|
|
1237
|
+
|
|
1238
|
+
The tests use actual registered renderers and mock replies at widths 1, 8, 30, and 80. Expansion preserves the full checked response. The one-column grapheme limitation remains unchanged.
|
|
1239
|
+
|
|
1240
|
+
Project whitespace and packaging dry-run checks passed. All nine profile check scripts and both profile core TypeScript commands passed. No paid provider request, settings change, or commit ran for this implementation. The new preview behavior still needs reload and a live terminal check. HTML export and other operating systems remain unverified.
|
|
1241
|
+
|
|
1242
|
+
|
|
1243
|
+
|
|
1244
|
+
After reload, one separately authorized live check succeeded using openai-codex/gpt-6-luna. It returned eight numbered sections as unstructured advice. The call reported 1304 tokens and 0.0005216 dollars. Both usage completeness flags were true. The session closed without a retry.
|
|
1245
|
+
The user confirmed that the collapsed preview and full expanded response looked correct in the terminal. This is user-observed evidence for that check, not an automated terminal capture or universal provider guarantee. HTML export and other operating systems remain unverified.
|
|
1246
|
+
|
|
1247
|
+
### Reasoning effort and command autocomplete {#reasoning-autocomplete-verification}
|
|
1248
|
+
|
|
1249
|
+
The user requested advisor reasoning effort and autocomplete, with pi-magic8ball as the reference. The advisor copies the reference patterns without a runtime dependency or changes to that worktree.
|
|
1250
|
+
|
|
1251
|
+
REQ-20 and AC-20 define strict configured effort, capability checks, scoped user commands, pinned requests, metadata, and expanded labels. REQ-21 and AC-21 define registered argument completion, fresh scoped lookups, replacement text, and context invalidation. REQ-6 and REQ-7 now reference these additions.
|
|
1252
|
+
|
|
1253
|
+
Baseline: all 62 existing checks and TypeScript passed. All 15 new checks in tests/configuration-features.test.mjs failed before implementation. They failed for rejected reasoning fields, missing commands/options/labels, absent completion callbacks, or missing README coverage. All 15 pass afterward. The 62 existing checks remain unchanged and pass. The complete suite passes 77 checks and TypeScript with npm run verify.
|
|
1254
|
+
|
|
1255
|
+
| Criterion | Changed checks | Before | After |
|
|
1256
|
+
| --- | --- | --- | --- |
|
|
1257
|
+
| AC-20 | strict reasoning values and whole-selection precedence | fail | pass |
|
|
1258
|
+
| AC-20 | reasoning query reports effective capabilities in UI and non-UI without writes or spending | fail | pass |
|
|
1259
|
+
| AC-20 | reasoning writes use the save scope and model selection resets effort | fail | pass |
|
|
1260
|
+
| AC-20 | unsupported, missing, invalid and untrusted reasoning commands never write or generate | fail | pass |
|
|
1261
|
+
| AC-20 | effort reaches every request and remains pinned with schema-valid metadata | fail | pass |
|
|
1262
|
+
| AC-20 | legacy/default/off omit SDK effort and supported explicit levels are forwarded unchanged | fail | pass |
|
|
1263
|
+
| AC-20 | unsupported effort fails before generation and continuation commits nothing | fail | pass |
|
|
1264
|
+
| AC-20 | expanded effort labels and README explain compatibility and costs | fail | pass |
|
|
1265
|
+
| AC-21 | command/scoped grammar completes without catalogs and rejects excess/conflicting arguments | fail | pass |
|
|
1266
|
+
| AC-21 | current physical providers and fuzzy model names preserve command identifiers | fail | pass |
|
|
1267
|
+
| AC-21 | reasoning completions use fresh save-scope capabilities and current trust | fail | pass |
|
|
1268
|
+
| AC-21 | actual Pi argument replacement retains whitespace, suffix and cursor with trailing space | fail | pass |
|
|
1269
|
+
| AC-21 | lookup failures stay silent and completions never write, spend or notify | fail | pass |
|
|
1270
|
+
| AC-21 | start, tree and shutdown discard stale pending results and context | fail | pass |
|
|
1271
|
+
| AC-21 | README distinguishes argument completions from the Pi 1.0.4 Tab routing limit | fail | pass |
|
|
1272
|
+
|
|
1273
|
+
The tests use actual extension registrations, real temporary settings I/O, and the installed CombinedAutocompleteProvider. Catalogs, model responses, and terminal dimensions are mocks. Capability checks use the public Pi AI 1.0.4 helper. The entry point uses the general ExtensionContext for lifecycle hooks and the tool context for tool requests.
|
|
1274
|
+
|
|
1275
|
+
The pi-magic8ball file memory records a prior installed-editor reproduction of a Pi 1.0.4 Tab-routing defect. Current editor source confirms that space-containing prefixes leave command-name Tab handling. Provider-level completion checks do not establish physical keyboard routing. The README states the typing workaround. At the offline implementation stage, no installed editor patch or live-terminal check ran.
|
|
1276
|
+
|
|
1277
|
+
Project whitespace and packaging dry-run checks passed. All nine profile check scripts and both profile core TypeScript commands passed. At the offline implementation stage, no paid provider request, live profile setting change, commit, or push ran for this feature. Provider effort compliance, live terminal behavior, HTML export, and other operating systems were still unverified.
|
|
1278
|
+
|
|
1279
|
+
|
|
1280
|
+
|
|
1281
|
+
After reload, the user reported completing the autocomplete and high-effort setup steps. A separately authorized live start and continuation then succeeded on openai-codex/gpt-6-luna. Both returned model.reasoning = high, turns reached two, and both usage completeness flags were true.
|
|
1282
|
+
The start reported 322 tokens and 0.0000582 dollars. The continuation reported 516 tokens and 0.0001064 dollars. The total was 838 tokens and 0.0001646 dollars, including 86 reported reasoning tokens. No retry ran. The idle session closed, and discovery returned no active sessions.
|
|
1283
|
+
This establishes one successful configured-effort flow, not provider-native high-effort compliance, universal support, or live pinning under a settings change. The user-reported setup is not an automated terminal capture. The known Pi 1.0.4 Tab limitation remains documented.
|
|
1284
|
+
A final npm run verify passed all 77 checks and TypeScript before commit. No additional paid request or agent-written live setting change ran.
|
|
1285
|
+
|
|
1286
|
+
### Opt-in provider failure diagnostics (issue #1) {#provider-diagnostics-verification}
|
|
1287
|
+
|
|
1288
|
+
The user requested implementation of GitHub issue neruok/pi-advisor#1. The inspected baseline was 32908ab and stored pi-advisor revision 19. The existing TypeScript check and all 77 offline tests passed before changes.
|
|
1289
|
+
|
|
1290
|
+
REQ-22 and AC-22 add a per-call diagnostics boolean, finite failure phases/categories, and effective or pinned model/source metadata. REQ-16 now explicitly permits those fields only for opted-in calls. Default error envelopes and successful results remain unchanged. The existing REQ-3 rollback/no-retry and REQ-12 usage-reporting semantics remain unchanged.
|
|
1291
|
+
|
|
1292
|
+
All seven new checks in tests/provider-diagnostics.test.mjs pass. They cover preparation versus zero-usage terminal failure, strict per-call opt-in, structured-field allowlists and hostile metadata, terminal abort/reply rejection with retained history, timeout/cancellation and late settlement, actual schema-valid tools with project overrides and pinned provenance, and safe rendering/README coverage. All 77 existing checks remain unchanged and pass. npm run verify passes TypeScript and all 84 tests.
|
|
1293
|
+
|
|
1294
|
+
Terminal provider responses can contain only a free-form errorMessage. The implementation deliberately does not parse it. Such failures report provider-error or provider-aborted, not a guessed authentication or transport cause. Recognized structured fields on thrown completion exceptions can provide narrower hints. The original live failure's root cause remains unknown.
|
|
1295
|
+
|
|
1296
|
+
No provider request, credential inspection, live settings change, live terminal check, commit, or push ran. Offline results do not establish billing completeness, provider interoperability, terminal appearance, HTML exports, or other operating systems.
|
|
1297
|
+
|
|
1298
|
+
### SDK timeout defect and terminal error hints {#sdk-errors-verification}
|
|
1299
|
+
|
|
1300
|
+
|
|
1301
|
+
Investigation found a local defect at baseline 7253a7e. Advisor computed a fractional timeout from performance.now().
|
|
1302
|
+
The installed Pi AI 1.0.4 Responses adapter passes that timeout to OpenAI SDK 7.19.0.
|
|
1303
|
+
The SDK rejects it with "timeout must be an integer" before fetch. A mock fetch reproduced this without network access.
|
|
1304
|
+
Rounding only the injected timeout allowed the same fixture to reach its HTTP 401 response.
|
|
1305
|
+
This confirms the defect, not the cause of the discarded historical xAI response.
|
|
1306
|
+
|
|
1307
|
+
REQ-23 and AC-23 require integer timeouts and fixed SDK terminal error hints. REQ-22 now permits those narrow hints.
|
|
1308
|
+
The prior verification section describes the original no-text-parsing implementation. The current classifier checks exact timeout strings and the SDK wrapper prefix only.
|
|
1309
|
+
It never returns or parses the provider body. Arbitrary terminal errors can still remain provider-error.
|
|
1310
|
+
|
|
1311
|
+
Baseline: all 84 existing tests and TypeScript passed. Five changed checks in tests/sdk-errors.test.mjs failed before implementation.
|
|
1312
|
+
Two new preserved checks passed before implementation. All seven now pass. All 84 existing tests remain unchanged and pass.
|
|
1313
|
+
npm run verify passes TypeScript and all 91 tests.
|
|
1314
|
+
|
|
1315
|
+
| AC-23 check | Before | After |
|
|
1316
|
+
| --- | --- | --- |
|
|
1317
|
+
| integer manager timeout reaches the real Responses adapter mock fetch | fail | pass |
|
|
1318
|
+
| SDK timeout validation surfaces a fixed code without network or raw text | fail | pass |
|
|
1319
|
+
| SDK wrapper statuses expose only allowlisted hints | fail | pass |
|
|
1320
|
+
| schemas and renderers show fixed SDK and HTTP diagnostics | fail | pass |
|
|
1321
|
+
| README documents SDK hints without raw provider payloads | fail | pass |
|
|
1322
|
+
| arbitrary messages, malformed wrappers, getters and terminal abort stay private | pass | pass |
|
|
1323
|
+
| default envelopes, accounting, rollback and no retry remain unchanged | pass | pass |
|
|
1324
|
+
|
|
1325
|
+
The real adapter tests use synthetic keys and mock fetch functions. No credentials or live provider calls were used.
|
|
1326
|
+
All nine profile check scripts and both profile core TypeScript commands passed. git diff --check passed.
|
|
1327
|
+
No commit, push, live terminal check, HTML export check, or external settings change ran.
|
|
1328
|
+
The running extension needs reload before it uses the changed code. Live xAI verification still requires separate spending authorization.
|
|
1329
|
+
|
|
1330
|
+
### xAI HTTP 400 investigation and tool-free request correction {#xai-no-tools-verification}
|
|
1331
|
+
|
|
1332
|
+
|
|
1333
|
+
After reload, two separately user-authorized xAI/grok-4.7 requests returned HTTP 400 with provider-rejection diagnostics.
|
|
1334
|
+
The first omitted configured reasoning. The second used high. Both reported zero usage with usageComplete true and committed nothing.
|
|
1335
|
+
Those reports do not establish zero billing. The failed provider details were discarded, so the exact rejection remains unconfirmed.
|
|
1336
|
+
|
|
1337
|
+
Offline capture with the installed grok-4.7 catalog entry and Responses adapter found tool_choice none without tools.
|
|
1338
|
+
Both omitted and high effort emitted that combination. Independent reports identify this combination as an xAI HTTP 400 trigger:
|
|
1339
|
+
- https://github.com/NousResearch/hermes-agent/issues/20590 reports the same Responses endpoint and a successful omission workaround.
|
|
1340
|
+
- https://github.com/router-for-me/CLIProxyAPI/issues/4491 reports the same rule through a Chat Completions relay.
|
|
1341
|
+
xAI's function-calling documentation defines none as a tool-choice mode but does not document this empty-tools validation rule.
|
|
1342
|
+
These reports support a likely cause. They do not prove the cause of this workspace's discarded responses.
|
|
1343
|
+
|
|
1344
|
+
REQ-24 and AC-24 omit toolChoice for xai only. Requests still declare no tools and reject tool-call replies.
|
|
1345
|
+
Other providers keep toolChoice none. Integer timeouts, configured effort, no retries, isolation, and settings remain unchanged.
|
|
1346
|
+
Exact known xAI no-tools rejection details can now produce code tool-choice-without-tools without raw detail exposure.
|
|
1347
|
+
|
|
1348
|
+
Baseline: all 91 tests and TypeScript passed. Four changed checks in tests/xai-no-tools.test.mjs failed before this correction.
|
|
1349
|
+
Two preserved checks passed before. All six now pass. All 91 existing checks remain unchanged and pass.
|
|
1350
|
+
npm run verify passes TypeScript and all 97 tests. git diff --check passed.
|
|
1351
|
+
|
|
1352
|
+
| AC-24 check | Before | After |
|
|
1353
|
+
| --- | --- | --- |
|
|
1354
|
+
| xAI real-adapter requests omit tool_choice with default effort | fail | pass |
|
|
1355
|
+
| xAI real-adapter requests omit tool_choice with high effort | fail | pass |
|
|
1356
|
+
| exact xAI no-tools rejection details return a fixed diagnostic code | fail | pass |
|
|
1357
|
+
| schemas and renderers expose the fixed no-tools explanation | fail | pass |
|
|
1358
|
+
| known-good SSE fixture succeeds and other providers keep toolChoice none | pass | pass |
|
|
1359
|
+
| unrecognized details, providers, statuses and aborts add no rejection code | pass | pass |
|
|
1360
|
+
|
|
1361
|
+
The SDK tests use the real adapter and catalog with synthetic keys and mock fetch. The fixture models the reported validation rule.
|
|
1362
|
+
A separate real-adapter mock HTTP 400 also confirmed that the fixed code survives SDK error normalization.
|
|
1363
|
+
All nine profile check scripts and both profile core TypeScript commands passed.
|
|
1364
|
+
No additional paid request, credential inspection, settings change, commit, or push ran during this correction.
|
|
1365
|
+
Reload and one separately authorized live check are still needed to establish whether this resolves the observed xAI failure.
|
|
1366
|
+
|
|
1367
|
+
### Remove the tool-choice override for every provider {#provider-neutral-options-verification}
|
|
1368
|
+
|
|
1369
|
+
|
|
1370
|
+
The user requested omission of unnecessary request options instead of provider-specific exceptions.
|
|
1371
|
+
REQ-25 and AC-25 replace only the toolChoice rules in REQ-24 and AC-24.
|
|
1372
|
+
Advisor now omits toolChoice for every provider and leaves payload construction to Pi.
|
|
1373
|
+
Zero tools, tool-call reply rejection, integer deadlines, cancellation, disabled retries, configured effort, and other request options remain unchanged.
|
|
1374
|
+
|
|
1375
|
+
Baseline: npm run verify passed TypeScript and all 97 tests before this change.
|
|
1376
|
+
Five AC-25 changed checks failed before implementation because of explicit toolChoice or outdated README coverage.
|
|
1377
|
+
The 24 preserved checks in the focused run passed before implementation. All 29 focused checks now pass.
|
|
1378
|
+
The revised full suite passes TypeScript and all 100 tests through npm run verify.
|
|
1379
|
+
|
|
1380
|
+
| AC-25 check | Before | After |
|
|
1381
|
+
| --- | --- | --- |
|
|
1382
|
+
| every provider omits toolChoice on creation and pinned continuation | fail | pass |
|
|
1383
|
+
| OpenAI real-adapter creation and continuation omit tools and tool_choice | fail | pass |
|
|
1384
|
+
| README documents Pi-owned tool choice for every provider | fail | pass |
|
|
1385
|
+
| isolated starts and continuations pin the selected model without toolChoice | fail | pass |
|
|
1386
|
+
| effort reaches every request and remains pinned without toolChoice with schema-valid metadata | fail | pass |
|
|
1387
|
+
| known-good SSE fixture, xAI default/high payloads, private diagnostics, and malformed reply rejection | pass | pass |
|
|
1388
|
+
|
|
1389
|
+
The provider matrix covers mock, xai, openai, anthropic, openai-codex, and custom selections with injected completions.
|
|
1390
|
+
The real-adapter checks use installed grok-4.7 and gpt-4.1 catalog entries, synthetic API keys, and mock fetch.
|
|
1391
|
+
The OpenAI fixture needed a dummy sk-prefixed key to exercise API-key rather than ChatGPT credential handling.
|
|
1392
|
+
Its first catalog selection also triggered Pi's context-budget clamp. The fixture now uses large-context gpt-4.1 without weakening the token-bound assertion.
|
|
1393
|
+
These fixture corrections do not change production request limits or authentication.
|
|
1394
|
+
|
|
1395
|
+
All nine profile check scripts, both profile core TypeScript commands, docs_validate, and git diff --check passed.
|
|
1396
|
+
No live model call, credential inspection, settings change, commit, or push occurred.
|
|
1397
|
+
Live xAI compatibility remains unconfirmed. Reload and separate spending authorization are required before a live check.
|
|
1398
|
+
|
|
1399
|
+
|
|
1400
|
+
After reload, the user authorized one live connectivity check. That request succeeded on xai/grok-4.7 with configured high reasoning.
|
|
1401
|
+
The checked reply was "I can respond." One exchange committed, and both usage completeness flags were true.
|
|
1402
|
+
Reported usage was 1653 total tokens, including 150 reasoning tokens. Reported cost was 0.002194 dollars, not independently verified billing.
|
|
1403
|
+
The idle consultation closed successfully. Discovery then returned no active consultations. No retry or continuation ran.
|
|
1404
|
+
This establishes one successful live request after removing the override, not universal compatibility or the exact cause of earlier HTTP 400 responses.
|
|
1405
|
+
No credential inspection, settings change, commit, or push occurred. Live continuation remains untested for this xAI configuration.
|
|
1406
|
+
|
|
1407
|
+
|
|
1408
|
+
The user then changed the advisor model to GPT and authorized one live connectivity check.
|
|
1409
|
+
That request succeeded on openai-codex/gpt-6-luna with no explicit reasoning selection in the returned model metadata.
|
|
1410
|
+
One checked exchange committed. Reported usage was 285 total tokens, zero reasoning tokens, and 0.0000365 dollars.
|
|
1411
|
+
Both usage completeness flags were true. The idle consultation closed successfully, and discovery returned no active consultations.
|
|
1412
|
+
No retry or continuation ran. The result establishes one successful GPT request, not universal compatibility or verified provider billing.
|
|
1413
|
+
The reply's caveat about verifying connectivity is model prose. Success is based on the observed tool outcome, not the model's self-assessment.
|
|
1414
|
+
The user authorized commit and push after success. No further paid request is part of pre-commit verification.
|
|
1415
|
+
|
|
1416
|
+
## Configurable deadline verification {#deadline-configuration-verification}
|
|
1417
|
+
|
|
1418
|
+
The user approved a five-minute default and global/project timeout configuration. REQ-26 and AC-26 define the behavior.
|
|
1419
|
+
|
|
1420
|
+
Before implementation, the first focused run failed 11 changed checks and passed 30 preserved checks. Failures showed the old 120000 ms default, missing timeout settings and commands, missing autocomplete, and absent pinning. Two additional checks later failed before their respective fixes: locked-checkpoint field preservation and elapsed preparation without timer dispatch.
|
|
1421
|
+
|
|
1422
|
+
After implementation, all 43 focused checks passed. The full suite passed 112 tests, and TypeScript passed. Verification used temporary settings, mock provider replies, synthetic authentication with mock fetch, and controlled clocks/timers. No paid provider call or live settings mutation occurred.
|
|
1423
|
+
|
|
1424
|
+
The deadline tests cover success after two minutes under the five-minute default, preparation consuming the configured budget, timeout during completion, no retry, ignored abort with a late reply, caller cancellation, known-zero usage before generation, and no timer or state change from late preparation callbacks.
|
|
1425
|
+
|
|
1426
|
+
The settings tests cover both inclusive timer bounds, invalid values, independent model and timeout precedence, ignored untrusted projects, scoped queries and saves, reset, field preservation from the locked checkpoint, static completion without settings/catalog reads, pinned continuations, fresh new consultations, and unchanged strict tool envelopes.
|
|
1427
|
+
|
|
1428
|
+
All nine profile verification scripts and both profile core TypeScript checks passed. These are offline observations, not proof that a slow remote provider will complete within the longer deadline or stop billing after cancellation.
|
|
1429
|
+
|
|
1430
|
+
## Short provider caching verification {#cache-verification}
|
|
1431
|
+
|
|
1432
|
+
The user authorized short provider caching and regression tests. REQ-27 replaces only the cache-retention preservation rule in REQ-25.
|
|
1433
|
+
|
|
1434
|
+
Before implementation, all 112 baseline tests and TypeScript passed. The AC-27 focused run then failed four changed checks and passed nine checks. Failures identified cacheRetention none, absent OpenAI cache keys, absent Anthropic cache markers, and missing README guidance. The new preserved check passed before and after implementation.
|
|
1435
|
+
|
|
1436
|
+
After implementation, all 13 focused checks and all 116 full-suite tests passed. TypeScript passed. Mock fetch exercised the installed Pi AI 1.0.4 OpenAI Responses and Anthropic adapters on creation and continuation. Synthetic SSE replies committed normally. OpenAI sent a stable consultation cache key without extended retention. Anthropic sent ephemeral system and last-message cache markers without one-hour TTL, even with a synthetic long environment default.
|
|
1437
|
+
|
|
1438
|
+
The provider matrix verified short caching for built-in and custom selections. The preserved check verified distinct consultation identifiers, unchanged continuation prefixes, model pinning, request controls, local close, and cache usage totals. Existing cancellation, rollback, timeout, schema, and privacy tests remained intact.
|
|
1439
|
+
|
|
1440
|
+
Verification used mocks only. No paid provider request, live settings mutation, commit, or push occurred. These checks do not prove live cache hits, savings, provider retention duration, or deletion of provider caches.
|
|
1441
|
+
|
|
1442
|
+
## Private continuation-state verification {#opaque-replay-verification}
|
|
1443
|
+
|
|
1444
|
+
REQ-28 and AC-28 implement the user-authorized caching fix without retaining plaintext thinking or enabling server-side conversation storage. Support is API-aware, not specific to Grok.
|
|
1445
|
+
|
|
1446
|
+
The baseline of 116 tests and TypeScript passed before this change. In tests/replay.test.mjs, eight changed-behavior checks failed before implementation for missing opaque replay, API identity, combined accounting, or README coverage. The malformed/unsupported/cross-model privacy check passed before and after.
|
|
1447
|
+
|
|
1448
|
+
| Offline check | Result after implementation |
|
|
1449
|
+
| --- | --- |
|
|
1450
|
+
| Installed Grok Responses adapter with synthetic SSE: encrypted reasoning, original text IDs/phases, empty summaries, stable key, store false, no previous_response_id | pass |
|
|
1451
|
+
| All seven supported API identities; Google conversion echoes signed empty parts | pass |
|
|
1452
|
+
| Installed Anthropic adapter with synthetic SSE: opaque redacted block replay, ordinary signed thinking omitted | pass |
|
|
1453
|
+
| Missing, malformed, unsupported, cross-provider/model, and accessor metadata omitted | pass |
|
|
1454
|
+
| Combined byte boundary, overflow rollback, known completed usage | pass |
|
|
1455
|
+
| Returned-reply and dispatch-context mutations cannot alter retained state | pass |
|
|
1456
|
+
| Failed and cancelled pairs, including ignored-abort late replies, preserve prior state | pass |
|
|
1457
|
+
| Consultation isolation, close, clear, and clean creation | pass |
|
|
1458
|
+
| README private-state and signed-plaintext limitation | pass |
|
|
1459
|
+
|
|
1460
|
+
Focused replay verification passed all nine checks. npm run verify passed all 125 tests and TypeScript. The nine new-coder profile checks and both profile core TypeScript commands passed. git diff --check passed.
|
|
1461
|
+
|
|
1462
|
+
Mocks use synthetic keys and provider state. They establish local payload construction and privacy/resource invariants, not live cache-hit improvement, provider acceptance, savings, billing, or remote deletion.
|
|
1463
|
+
Responses replay is allowlisted and sanitized. Google retains only valid base64 thought signatures with empty thinking text. Anthropic and Bedrock retain only opaque redacted thinking. Ordinary signed plaintext thinking remains unsupported when a provider requires it for continuation or caching.
|
|
1464
|
+
Private state counts against the existing 49152-byte limit; text blocks are duplicated in the private replay representation, so supported formats can reach the limit sooner.
|
|
1465
|
+
A further live probe requires separate spending authorization and a reload. No new paid provider request, live setting change, commit, push, or deployment occurred during this fix. Existing worktree changes were preserved.
|
|
1466
|
+
|
|
1467
|
+
## Model context and display verification {#context-display-verification}
|
|
1468
|
+
|
|
1469
|
+
REQ-29 and REQ-30 implement the user's three requests: visible agent messages, model-sized isolated context, and a Pi-like usage line.
|
|
1470
|
+
The user confirmed that parent context remains excluded. The user requested no subscription or auto-compaction labels.
|
|
1471
|
+
|
|
1472
|
+
The 125-test baseline and TypeScript passed before implementation. In tests/context-display.test.mjs, eleven changed checks failed for the intended missing behavior.
|
|
1473
|
+
The preserved no-side-effects and private/partial-state display check passed before and after. All twelve checks then passed.
|
|
1474
|
+
|
|
1475
|
+
| Criterion | Offline check | Before | After |
|
|
1476
|
+
| --- | --- | --- | --- |
|
|
1477
|
+
| AC-30 | Full sanitized multiline question, original dispatched input | fail | pass |
|
|
1478
|
+
| AC-29 | Advisor registry window, independent parent state, pinned window | fail | pass |
|
|
1479
|
+
| AC-29 | Missing and invalid window metadata rejects before generation | fail | pass |
|
|
1480
|
+
| AC-29 | Four pairs with replay exceeding49152 bytes, provider token measurement | fail | pass |
|
|
1481
|
+
| AC-29 | Exact output-reserved preflight and one-token overflow without dispatch | fail | pass |
|
|
1482
|
+
| AC-29 | Inclusive committed-context boundary, rollback and known rejected usage | fail | pass |
|
|
1483
|
+
| AC-29 | Pi usage fallback and all-zero visible-message estimate | fail | pass |
|
|
1484
|
+
| AC-29 | Detached context snapshots and strict schema changes | fail | pass |
|
|
1485
|
+
| AC-30 | Cumulative totals, latest-call CH, Pi abbreviations, both views and discovery | fail | pass |
|
|
1486
|
+
| AC-30 | Unknown usage, incomplete totals, historic context, and cumulative failures | fail | pass |
|
|
1487
|
+
| AC-30 | Display isolation, private/partial exclusions, no mutation or extra calls | pass | pass |
|
|
1488
|
+
| AC-30 | README questions, registry context, and latest-call CH | fail | pass |
|
|
1489
|
+
|
|
1490
|
+
Pi 1.0.4 reads contextWindow from its limits model. Advisor reads the configured physical advisor model through resolveModel and pins that value.
|
|
1491
|
+
It does not read the parent's model window, context, footer, or session history.
|
|
1492
|
+
Pi's public calculateContextTokens and estimateTokens helpers supply the estimate. Private encrypted payload length is not a token measurement.
|
|
1493
|
+
The fixed49152-byte cap is removed. Preflight reserves the existing4096 output request within the model window. A validated pair must also fit before commit.
|
|
1494
|
+
Other limits remain:8 sessions,24 pairs,16384 bytes per input/reply, and the pinned deadline. No automatic compaction or retry was added.
|
|
1495
|
+
|
|
1496
|
+
The full suite initially identified outdated mock model metadata and checks for explicitly superseded byte capacity and input omission.
|
|
1497
|
+
Fixtures now supply model contextWindow/getContextWindow. Only superseded byte-cap, byte-remaining, and hidden-call-input assertions changed.
|
|
1498
|
+
The remaining reply, progress, privacy, renderer preview, lifecycle, request-control, and timeout assertions were retained.
|
|
1499
|
+
|
|
1500
|
+
npm run verify passed all137 tests and TypeScript. git diff --check passed.
|
|
1501
|
+
The nine new-coder profile checks and both profile core TypeScript commands passed.
|
|
1502
|
+
Verification used temporary settings, mock catalogs, synthetic replay, and mock provider responses only.
|
|
1503
|
+
No new paid request, live settings change, commit, push, deployment, or live terminal check occurred.
|
|
1504
|
+
These results do not prove exact provider tokenization, universal compatibility, or live UI appearance. Reload is required to load the changed extension.
|
|
1505
|
+
|
|
1506
|
+
## Fixed consultation cap removal {#fixed-cap-removal-verification}
|
|
1507
|
+
|
|
1508
|
+
|
|
1509
|
+
The user authorized removal of the active consultation, exchange, message-byte, reply-byte, and output-token caps.
|
|
1510
|
+
The user retained the model context limit and configurable timeout. REQ-31, AC-31, and AC-32 define this change.
|
|
1511
|
+
The inspected baseline was commit 69c64a3 and stored document revision 38. TypeScript and all 137 baseline tests passed.
|
|
1512
|
+
|
|
1513
|
+
Changed behavior: eleven AC-31 checks failed before implementation for the removed limits, schemas, renderer cutoff, or README.
|
|
1514
|
+
Ten failures came from the first focused run. The separate reply-byte check also failed before implementation.
|
|
1515
|
+
Preserved behavior: the AC-32 rollback, private replay, rejected usage, and cancelled late-work check passed before and after.
|
|
1516
|
+
All twelve focused checks passed after implementation.
|
|
1517
|
+
|
|
1518
|
+
| Criterion | Checks in tests/removed-limits.test.mjs | Before | After |
|
|
1519
|
+
| --- | --- | --- | --- |
|
|
1520
|
+
| AC-31 | Multibyte messages and replies above 16 KiB, including a short-input reply case | fail | pass |
|
|
1521
|
+
| AC-31 | Nine active consultations and nine pending creations, without eviction | fail | pass |
|
|
1522
|
+
| AC-31 | Twenty-five exchanges with the complete consultation prefix | fail | pass |
|
|
1523
|
+
| AC-31 | Provider matrix without maxTokens and installed Responses model output allowance | fail | pass |
|
|
1524
|
+
| AC-31 | Exact pending-context equality and one-token overflow without dispatch | fail | pass |
|
|
1525
|
+
| AC-31 | Strict schemas above count caps, no turnsRemaining, and only context-tokens | fail | pass |
|
|
1526
|
+
| AC-31 | Every discovery entry in both views at widths 1, 8, 30, 80, and 160 | fail | pass |
|
|
1527
|
+
| AC-31 | README removed caps and retained controls | fail | pass |
|
|
1528
|
+
| AC-32 | Context overflow, malformed replies, private replay, accounting, and cancelled late work | pass | pass |
|
|
1529
|
+
|
|
1530
|
+
Advisor now omits maxTokens. Installed Pi AI 1.0.4 uses model defaults and its own context clamp.
|
|
1531
|
+
A synthetic-authentication, mock-fetch Responses test used the installed gpt-4.1 output allowance of 32768 tokens.
|
|
1532
|
+
This is payload evidence, not proof of live provider acceptance or reasoning quality.
|
|
1533
|
+
Advisor compares pending context directly with the pinned model window. It also checks committed context before each commit.
|
|
1534
|
+
No fixed output reserve remains in advisor. Pi can still apply its own reserve or provider output limits.
|
|
1535
|
+
|
|
1536
|
+
The first full run identified 21 obsolete expectations for removed caps, reserve tokens, or turnsRemaining.
|
|
1537
|
+
Those expectations now trace to AC-31. Context overflow, deadlines, privacy, accounting, replay, and cancellation assertions remain.
|
|
1538
|
+
A renderer recovery-text check required whitespace-aware matching because the longer context error wraps onto two lines.
|
|
1539
|
+
The change did not suppress that diagnostic or relax the required recovery text.
|
|
1540
|
+
|
|
1541
|
+
npm run verify passed TypeScript and all 149 tests. npm run packcheck passed package loading and archive checks.
|
|
1542
|
+
The nine new-coder profile scripts and both profile core TypeScript checks passed. Project git diff --check passed.
|
|
1543
|
+
Verification used mock catalogs, synthetic replies, mock fetch, controlled clocks, and temporary settings only.
|
|
1544
|
+
No paid provider call, live settings change, live terminal check, commit, push, or deployment ran.
|
|
1545
|
+
Reload is required before the running extension uses this change.
|
|
1546
|
+
|
|
1547
|
+
|
|
1548
|
+
|
|
1549
|
+
After reload, the user separately authorized one live consultation and one continuation with the configured advisor model.
|
|
1550
|
+
Both calls succeeded on xai/grok-4.7. The same consultation reached two exchanges with the same model selection.
|
|
1551
|
+
Both usage-completeness flags were true on each call. The reported total was 3595 tokens and 0.00471 dollars.
|
|
1552
|
+
Reported reasoning totaled 251 tokens. The pinned context window was 500000 tokens.
|
|
1553
|
+
The idle consultation closed successfully. Discovery then returned no active consultations. No retry occurred.
|
|
1554
|
+
This confirms one live two-call flow, not live behavior beyond the removed caps or independently verified provider billing.
|
|
1555
|
+
The user then authorized a version bump, commit, push, and npm publication.
|