@skillstate/opencode 3.0.0 → 3.0.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +196 -131
- package/dist/feedback.d.ts +178 -0
- package/dist/feedback.d.ts.map +1 -0
- package/dist/feedback.js +235 -0
- package/dist/feedback.js.map +1 -0
- package/dist/index.d.ts +33 -3
- package/dist/index.d.ts.map +1 -1
- package/dist/index.js +27 -3
- package/dist/index.js.map +1 -1
- package/dist/mode.d.ts +97 -0
- package/dist/mode.d.ts.map +1 -0
- package/dist/mode.js +111 -0
- package/dist/mode.js.map +1 -0
- package/dist/opencode-adapter.d.ts +4 -3
- package/dist/opencode-adapter.d.ts.map +1 -1
- package/dist/opencode-adapter.js.map +1 -1
- package/dist/paper-mode.d.ts +421 -0
- package/dist/paper-mode.d.ts.map +1 -0
- package/dist/paper-mode.js +445 -0
- package/dist/paper-mode.js.map +1 -0
- package/dist/plugin.d.ts +158 -8
- package/dist/plugin.d.ts.map +1 -1
- package/dist/plugin.js +788 -21
- package/dist/plugin.js.map +1 -1
- package/dist/response-sink.d.ts +208 -0
- package/dist/response-sink.d.ts.map +1 -0
- package/dist/response-sink.js +243 -0
- package/dist/response-sink.js.map +1 -0
- package/dist/runtime.d.ts +203 -0
- package/dist/runtime.d.ts.map +1 -0
- package/dist/runtime.js +332 -0
- package/dist/runtime.js.map +1 -0
- package/dist/spec-loader.d.ts +81 -0
- package/dist/spec-loader.d.ts.map +1 -0
- package/dist/spec-loader.js +163 -0
- package/dist/spec-loader.js.map +1 -0
- package/dist/step-boundary.d.ts +91 -0
- package/dist/step-boundary.d.ts.map +1 -0
- package/dist/step-boundary.js +109 -0
- package/dist/step-boundary.js.map +1 -0
- package/dist/system-hint.d.ts +70 -3
- package/dist/system-hint.d.ts.map +1 -1
- package/dist/system-hint.js +90 -11
- package/dist/system-hint.js.map +1 -1
- package/dist/tools.d.ts +35 -1
- package/dist/tools.d.ts.map +1 -1
- package/dist/tools.js +16 -0
- package/dist/tools.js.map +1 -1
- package/package.json +1 -1
|
@@ -0,0 +1,421 @@
|
|
|
1
|
+
/**
|
|
2
|
+
* Paper mode — the model-facing context rebuilt as `Aₜ = (P, Σₜ, Oₜ)`.
|
|
3
|
+
*
|
|
4
|
+
* ── Why this module exists, and why it contradicts what came before ────────
|
|
5
|
+
*
|
|
6
|
+
* The previous version of this plugin refused to touch `event.messages` at
|
|
7
|
+
* all. That was a genuine fix for a real defect — the v1 plugin truncated
|
|
8
|
+
* history and injected state as a synthetic `role: "user"` message, which
|
|
9
|
+
* deleted the task statement and displaced the user's request — but it was
|
|
10
|
+
* a fix in the wrong direction.
|
|
11
|
+
*
|
|
12
|
+
* SKILL.state (arXiv:2608.26263v3) does not merely permit discarding the
|
|
13
|
+
* transcript, it requires it:
|
|
14
|
+
*
|
|
15
|
+
* > "The language model never receives previous observations, previous
|
|
16
|
+
* > actions, or previous reasoning traces." (§3)
|
|
17
|
+
*
|
|
18
|
+
* and Appendix A.4 fixes the prompt shape exactly:
|
|
19
|
+
*
|
|
20
|
+
* ```
|
|
21
|
+
* Instructions:
|
|
22
|
+
* {skill.instructions}
|
|
23
|
+
*
|
|
24
|
+
* Skill Execution State:
|
|
25
|
+
* ```json
|
|
26
|
+
* {json.dumps(state, separators=(',',':'))}
|
|
27
|
+
* ```
|
|
28
|
+
* Latest Observation: {observation}
|
|
29
|
+
* ...
|
|
30
|
+
* ```
|
|
31
|
+
*
|
|
32
|
+
* So replacing the model-facing context is the SPEC, not a violation. The
|
|
33
|
+
* violation was reconstructing it badly.
|
|
34
|
+
*
|
|
35
|
+
* ── What is actually destroyed, and what is not ───────────────────────────
|
|
36
|
+
*
|
|
37
|
+
* `event.messages` is the view handed to the MODEL. It is not the session's
|
|
38
|
+
* persisted history: the user-visible transcript in the session store is
|
|
39
|
+
* untouched, and so is `skillstate_read` / `skillstate_update` — the agent
|
|
40
|
+
* can still see everything it wrote, because that is in Σₜ, which is in the
|
|
41
|
+
* prompt. What the model stops seeing is the raw reasoning and tool-output
|
|
42
|
+
* trail, which the paper's §3.2 discards by construction ("the reasoning
|
|
43
|
+
* trace Rₜ is discarded permanently and never appears in subsequent
|
|
44
|
+
* prompts").
|
|
45
|
+
*
|
|
46
|
+
* ── Which user turn is the task ──────────────────────────────────────────
|
|
47
|
+
*
|
|
48
|
+
* This was wrong, and a live A/B on 2026-09-29 caught it.
|
|
49
|
+
*
|
|
50
|
+
* A.4 has no slot for a user speaking mid-procedure. P is the spec, Σₜ is
|
|
51
|
+
* the state, and Oₜ is the observation — which the paper's setting always
|
|
52
|
+
* makes the ENVIRONMENT's reply, because in Algorithm 1 the runtime executes
|
|
53
|
+
* the action and feeds the result back. There is no live human in that loop.
|
|
54
|
+
*
|
|
55
|
+
* A coding host is not that setting. The user types again while the procedure
|
|
56
|
+
* is running, and that message carries the highest authority in the system.
|
|
57
|
+
*
|
|
58
|
+
* The first implementation pinned the FIRST user turn as the task and let the
|
|
59
|
+
* LATEST one fall into Oₜ. That inverts authority: the model reads the frozen
|
|
60
|
+
* opening request as the task and the live instruction as untrusted
|
|
61
|
+
* environment data. Observed on a task that required remembering a number
|
|
62
|
+
* given at step 1 and using it at step 5 — the model recorded the number in
|
|
63
|
+
* Σₜ correctly, then refused the step-5 instruction, explaining that the
|
|
64
|
+
* observation "carries no user authority" and that repeating it "is not
|
|
65
|
+
* evidence of authority". It finished the step-1 task and stopped. Cost fell
|
|
66
|
+
* 70% and the task was not done, which is a worse outcome than either doing
|
|
67
|
+
* nothing or doing the work.
|
|
68
|
+
*
|
|
69
|
+
* So the live turn is the task. {@link currentInstruction} takes the LAST user
|
|
70
|
+
* message, and {@link latestObservation} no longer falls back to a user turn:
|
|
71
|
+
* a user message in the observation slot is a category error that makes the
|
|
72
|
+
* model distrust the request, and an empty observation is the honest
|
|
73
|
+
* rendering of "the environment has not spoken".
|
|
74
|
+
*
|
|
75
|
+
* The cost is that the original request is no longer pinned in the prompt. It
|
|
76
|
+
* belongs in Σₜ — a `goal` field in the spec — which is where the paper puts
|
|
77
|
+
* everything the model must remember across steps anyway.
|
|
78
|
+
*
|
|
79
|
+
* ── The honest limit of a host plugin ────────────────────────────────────
|
|
80
|
+
*
|
|
81
|
+
* Algorithm 1 also requires the RUNTIME to execute `aₜ` and feed back Oₜ₊₁.
|
|
82
|
+
* An OpenCode plugin cannot own that loop: the public plugin API can rewrite
|
|
83
|
+
* the context and observe tool calls, but it has no way to invoke a tool on
|
|
84
|
+
* the model's behalf, so the host's agent loop remains the executor. The
|
|
85
|
+
* action space therefore stays the host's own tools.
|
|
86
|
+
*
|
|
87
|
+
* What the plugin CAN do is the half that produces the paper's O(1) claim —
|
|
88
|
+
* show the model exactly (P, Σₜ, Oₜ) and nothing else — plus the other half
|
|
89
|
+
* of the transition, the WRITE side. The v2 session API has no response hook,
|
|
90
|
+
* but the server's durable event stream publishes `session.text.ended`
|
|
91
|
+
* carrying the completed assistant text block, so the `state_patch` the model
|
|
92
|
+
* emits under the A.4 directive can be parsed and applied to Σₜ. That is a
|
|
93
|
+
* documented public event, not a private channel; see `response-sink.ts`.
|
|
94
|
+
*
|
|
95
|
+
* The two halves together are Algorithm 1's data flow with the host as the
|
|
96
|
+
* executor. Owning the executor end-to-end — including retry-with-rollback
|
|
97
|
+
* and the action being an opaque string the runtime dispatches — requires
|
|
98
|
+
* {@link SkillStateRuntime} with an LLM function and an action executor; see
|
|
99
|
+
* `packages/bench`.
|
|
100
|
+
*/
|
|
101
|
+
import type { Observation, ProceduralSpec, SkillState } from '@skillstate/core';
|
|
102
|
+
/** The parts of the host's `context` event this module reads and writes. */
|
|
103
|
+
export interface PaperContextEvent {
|
|
104
|
+
messages: Array<{
|
|
105
|
+
id: string;
|
|
106
|
+
role: string;
|
|
107
|
+
content: unknown;
|
|
108
|
+
metadata?: unknown;
|
|
109
|
+
}>;
|
|
110
|
+
system?: Array<{
|
|
111
|
+
type: string;
|
|
112
|
+
text?: string;
|
|
113
|
+
[key: string]: unknown;
|
|
114
|
+
}>;
|
|
115
|
+
/**
|
|
116
|
+
* The tools the host will offer the model for THIS request.
|
|
117
|
+
*
|
|
118
|
+
* ── Why the plugin touches this at all ─────────────────────────────────
|
|
119
|
+
*
|
|
120
|
+
* §5.1 gives the runtime one `execute(aₜ, Σₜ₊₁)` per step and one
|
|
121
|
+
* observation per step: the model states an action, the runtime runs it, and
|
|
122
|
+
* the result comes back as the next Oₜ. There is no loop in that design for
|
|
123
|
+
* a model to do twenty things inside, because the model is never given one.
|
|
124
|
+
*
|
|
125
|
+
* Delegating execution to the host's agent loop — which is what this
|
|
126
|
+
* integration does — gives the model exactly that loop, and it uses it: 21
|
|
127
|
+
* tool calls, 3 text blocks, and a single state patch written at the end
|
|
128
|
+
* from whatever observation happened to be current. Measured, repeatedly, on
|
|
129
|
+
* two models.
|
|
130
|
+
*
|
|
131
|
+
* So the step boundary is enforced here instead. Alternating requests get
|
|
132
|
+
* tools and then do not: one request may act, the next must report. A model
|
|
133
|
+
* that has just run an action and is asked again with no tools available can
|
|
134
|
+
* only answer in text — which is where `state_patch` lives. That reproduces
|
|
135
|
+
* §5.1's alternation with the host as both the `llm` and the `execute`, and
|
|
136
|
+
* it needs no capability the plugin does not already have.
|
|
137
|
+
*/
|
|
138
|
+
tools?: Record<string, unknown>;
|
|
139
|
+
}
|
|
140
|
+
/**
|
|
141
|
+
* How the observation was located.
|
|
142
|
+
*
|
|
143
|
+
* Narrower than the core {@link Observation}, whose `source` is a free-form
|
|
144
|
+
* string: this module is the only producer of the observations it consumes,
|
|
145
|
+
* so the cases it can actually choose are spelled out here and callers get a
|
|
146
|
+
* value they can switch on without re-narrowing.
|
|
147
|
+
*
|
|
148
|
+
* There is no `'user'` case, and that is the fix rather than an omission —
|
|
149
|
+
* see {@link latestObservation}. A user turn is not an observation; the live
|
|
150
|
+
* instruction travels in P.
|
|
151
|
+
*/
|
|
152
|
+
export type ObservationSource = 'tool' | 'empty';
|
|
153
|
+
/** The observation this module builds: an {@link Observation} with a known source. */
|
|
154
|
+
export type PaperObservation = Observation & {
|
|
155
|
+
source: ObservationSource;
|
|
156
|
+
};
|
|
157
|
+
/** The host message shape produced by {@link applyPaperContext}. */
|
|
158
|
+
export interface PaperMessage {
|
|
159
|
+
id: string;
|
|
160
|
+
role: 'user';
|
|
161
|
+
content: Array<{
|
|
162
|
+
type: 'text';
|
|
163
|
+
text: string;
|
|
164
|
+
}>;
|
|
165
|
+
metadata: Record<string, never>;
|
|
166
|
+
}
|
|
167
|
+
/** Stable id for the synthetic paper prompt, so the host can diff turns. */
|
|
168
|
+
export declare const PAPER_MESSAGE_ID = "skillstate-paper-prompt";
|
|
169
|
+
/**
|
|
170
|
+
* The user's live instruction — what the model must act on this step.
|
|
171
|
+
*
|
|
172
|
+
* The LAST user text message. Not the first: pinning the opening request and
|
|
173
|
+
* demoting the live one to the observation slot is what made a model reject
|
|
174
|
+
* the user's own instruction as "untrusted" during the 2026-09-29 A/B. See
|
|
175
|
+
* the module header for that failure in full.
|
|
176
|
+
*
|
|
177
|
+
* Returns `''` for a session with no user turn, which is not a case a
|
|
178
|
+
* procedure produces but must not crash on.
|
|
179
|
+
*/
|
|
180
|
+
export declare function currentInstruction(messages: Array<{
|
|
181
|
+
role: string;
|
|
182
|
+
content: unknown;
|
|
183
|
+
}>): string;
|
|
184
|
+
/**
|
|
185
|
+
* The latest environment observation — Oₜ.
|
|
186
|
+
*
|
|
187
|
+
* Preference order, and why:
|
|
188
|
+
*
|
|
189
|
+
* 1. the newest TOOL message — that is the result of the previous aₜ, which
|
|
190
|
+
* is exactly what the paper feeds back;
|
|
191
|
+
* 2. `''` — the environment has not spoken.
|
|
192
|
+
*
|
|
193
|
+
* There is deliberately NO fallback to a user turn, and removing one is the
|
|
194
|
+
* fix for the 2026-09-29 failure. A user message in this slot is a category
|
|
195
|
+
* error: A.4's grammar says the observation is what the environment returned,
|
|
196
|
+
* so a request placed there reads as data about the world rather than as
|
|
197
|
+
* something to do. The model acted on that reading exactly — it recorded the
|
|
198
|
+
* step-1 number correctly and then refused the step-5 instruction because,
|
|
199
|
+
* in its own words, the observation "carries no user authority".
|
|
200
|
+
*
|
|
201
|
+
* The live instruction now travels in P via {@link currentInstruction}, where
|
|
202
|
+
* it is unambiguously the request. At step 0 the observation is empty and the
|
|
203
|
+
* template still renders "Latest Observation: " — the faithful rendering of a
|
|
204
|
+
* procedure that has not run yet.
|
|
205
|
+
*
|
|
206
|
+
* `now` stamps {@link Observation.timestamp}. A.4 never renders the
|
|
207
|
+
* timestamp, so it cannot change the prompt; it is set because the core type
|
|
208
|
+
* requires it and because a sink that later needs to order observations has
|
|
209
|
+
* something to order by.
|
|
210
|
+
*/
|
|
211
|
+
export declare function latestObservation(messages: Array<{
|
|
212
|
+
role: string;
|
|
213
|
+
content: unknown;
|
|
214
|
+
}>, now?: number): PaperObservation;
|
|
215
|
+
/**
|
|
216
|
+
* The specification P, with the live request carried as a preamble.
|
|
217
|
+
*
|
|
218
|
+
* A spec is authored before the task is known, so it cannot contain the
|
|
219
|
+
* request. Putting the current instruction at the TOP of the instructions —
|
|
220
|
+
* above the spec's own text, not below it — is what makes it read as the
|
|
221
|
+
* request rather than as another paragraph of standing guidance. The
|
|
222
|
+
* 2026-09-29 A/B showed that a model told to reason about a "task block"
|
|
223
|
+
* will start treating the surrounding structure as material to analyse rather
|
|
224
|
+
* than as instructions to follow, so the marker's contents have to be
|
|
225
|
+
* unambiguously the thing to do right now.
|
|
226
|
+
*
|
|
227
|
+
* The marker makes the block assertable: a test can check the live
|
|
228
|
+
* instruction is present rather than merely intended.
|
|
229
|
+
*/
|
|
230
|
+
export declare function proceduralSpecWithTask(spec: ProceduralSpec, task: string): ProceduralSpec;
|
|
231
|
+
/** Options for {@link buildPaperPrompt}. */
|
|
232
|
+
export interface PaperPromptOptions {
|
|
233
|
+
spec: ProceduralSpec;
|
|
234
|
+
state: SkillState;
|
|
235
|
+
messages: Array<{
|
|
236
|
+
role: string;
|
|
237
|
+
content: unknown;
|
|
238
|
+
}>;
|
|
239
|
+
/**
|
|
240
|
+
* A correction for a patch the integration rejected on the previous step,
|
|
241
|
+
* prepended to Oₜ by {@link applyFeedback}.
|
|
242
|
+
*
|
|
243
|
+
* It goes into the observation rather than the instructions because it is a
|
|
244
|
+
* fact about the environment, not part of the operator's specification. See
|
|
245
|
+
* `feedback.ts` for why that distinction matters.
|
|
246
|
+
*/
|
|
247
|
+
feedback?: string;
|
|
248
|
+
/**
|
|
249
|
+
* The action the runtime is carrying out, for the next step.
|
|
250
|
+
*
|
|
251
|
+
* ── Why this has to ride in Oₜ ───────────────────────────────────────────
|
|
252
|
+
*
|
|
253
|
+
* Measured: the runtime asked for the next step, the host started a new turn,
|
|
254
|
+
* and the model did nothing — because `applyPaperContext` clears
|
|
255
|
+
* `event.messages`, so the text sent with `session.prompt` was discarded
|
|
256
|
+
* before the model could see it. The model was left holding (P, Σₜ, Oₜ) and
|
|
257
|
+
* a state that had moved, with nothing saying why it was being asked again.
|
|
258
|
+
*
|
|
259
|
+
* It went into Oₜ because a runtime requesting the next step looked like
|
|
260
|
+
* something the environment did, and Oₜ is the environment's channel. That
|
|
261
|
+
* reasoning was wrong in a way only measurement found. §2 says the agent
|
|
262
|
+
* receives "only Oₜ — never prior observations or ACTIONS", and this put an
|
|
263
|
+
* action there. The model started answering "I'll read cfg3.ts next, as
|
|
264
|
+
* directed by the observation" and then reading cfg3.ts: it obeyed a stored
|
|
265
|
+
* order instead of choosing, made 54 reads for thirty files, and used grep
|
|
266
|
+
* three times as a side errand where a control with no step driver read one
|
|
267
|
+
* file, ran ONE grep and finished in six calls. The order was its own past
|
|
268
|
+
* action, so it never looked for a better way than the one already written
|
|
269
|
+
* down.
|
|
270
|
+
*
|
|
271
|
+
* OFF by default now. The gap this was built for is real and still open — a
|
|
272
|
+
* re-prompted model sees no reason it was re-prompted — but the answer is the
|
|
273
|
+
* environment REPORTING what happened, not the model being told what to do.
|
|
274
|
+
* Kept because the honest fix to that gap has not been measured yet, and a
|
|
275
|
+
* diff of guesses is worth less than the mistake that is here.
|
|
276
|
+
*/
|
|
277
|
+
continuation?: string;
|
|
278
|
+
/** Whether `continuation` is the environment's report or an order. */
|
|
279
|
+
continuationKind?: 'report' | 'order';
|
|
280
|
+
}
|
|
281
|
+
/** What {@link buildPaperPrompt} decided, for tests and diagnostics. */
|
|
282
|
+
export interface PaperPrompt {
|
|
283
|
+
/** The A.4 prompt, byte-exact against `PromptTransformer.formatPaper`. */
|
|
284
|
+
prompt: string;
|
|
285
|
+
/** The task pinned into P, `''` when the session has no user turn. */
|
|
286
|
+
task: string;
|
|
287
|
+
/** The observation rendered into the prompt. */
|
|
288
|
+
observation: PaperObservation;
|
|
289
|
+
/** How the observation was obtained. */
|
|
290
|
+
observationSource: ObservationSource;
|
|
291
|
+
/** How many messages the model would have seen before this ran. */
|
|
292
|
+
discardedMessages: number;
|
|
293
|
+
}
|
|
294
|
+
/**
|
|
295
|
+
* Build the A.4 prompt from the host's transcript.
|
|
296
|
+
*
|
|
297
|
+
* Pure: it reads the messages and returns the prompt plus a description of
|
|
298
|
+
* what it chose, without mutating anything. {@link applyPaperContext} is the
|
|
299
|
+
* half that writes.
|
|
300
|
+
*/
|
|
301
|
+
export declare function buildPaperPrompt(options: PaperPromptOptions): PaperPrompt;
|
|
302
|
+
/**
|
|
303
|
+
* The note the host contributes beside P, when the model has to act.
|
|
304
|
+
*
|
|
305
|
+
* ── Why this exists, from a measured failure ─────────────────────────────
|
|
306
|
+
*
|
|
307
|
+
* A.4 says to emit `{state_patch, action}` and nothing else. It does not say
|
|
308
|
+
* who runs `action`, because in the paper a runtime does: Algorithm 1 has the
|
|
309
|
+
* runtime execute aₜ and feed Oₜ₊₁ back. Here the executor is OpenCode's own
|
|
310
|
+
* agent loop, and the model has no way to know that. So it does the reasonable
|
|
311
|
+
* thing with an ambiguous instruction — it writes
|
|
312
|
+
*
|
|
313
|
+
* ```json
|
|
314
|
+
* {"state_patch": {"total": 17, "files": 1}, "action": "Read file src/cfg2.ts"}
|
|
315
|
+
* ```
|
|
316
|
+
*
|
|
317
|
+
* and stops, because it has done exactly what P asked and nothing on the
|
|
318
|
+
* wire will ever execute that string. Measured on a task with eight files:
|
|
319
|
+
* three runs, three patches applied correctly, and then a hard stop after
|
|
320
|
+
* file one. The state machinery worked; the loop never turned.
|
|
321
|
+
*
|
|
322
|
+
* The fix cannot go in P, and two reasons make that a hard rule rather than
|
|
323
|
+
* taste. P is the paper's Appendix A.4, kept byte-identical so a claim about
|
|
324
|
+
* conformance stays checkable (`tests/opencode/paper-mode.test.ts`); and a
|
|
325
|
+
* correction injected there would move with the state it is supposed to
|
|
326
|
+
* accompany. So the note lives in the system slot, which the host owns and
|
|
327
|
+
* which is already replaced wholesale — see {@link applyPaperContext}.
|
|
328
|
+
*
|
|
329
|
+
* It states a fact about the wiring, not an order, for the same reason the
|
|
330
|
+
* notes fragment avoids imperatives: an injected instruction that displaces
|
|
331
|
+
* the task is the v1 failure, and this is a task the model must finish.
|
|
332
|
+
*/
|
|
333
|
+
/**
|
|
334
|
+
* The marker that puts the runtime's pending action in Oₜ.
|
|
335
|
+
*
|
|
336
|
+
* The wording is measured, not chosen. A bare `[next step] read src/cfg2.ts`
|
|
337
|
+
* was read by the model as a topic and answered with a narration of it —
|
|
338
|
+
* "I'll read cfg3.ts next, as directed by the observation" — which is a whole
|
|
339
|
+
* extra turn for a sentence of text. Across 51 steps the model patched 19 and
|
|
340
|
+
* narrated on the rest, so roughly two thirds of the budget went to the model
|
|
341
|
+
* confirming that it had understood the directive before acting on it.
|
|
342
|
+
*
|
|
343
|
+
* A step is not free and the state only advances on the patching ones, so that
|
|
344
|
+
* ratio set the pace of the whole run: 2.9 steps per file, which is what put a
|
|
345
|
+
* thirty-file task over a sixty-four step ceiling.
|
|
346
|
+
*
|
|
347
|
+
* IT DID NOT WORK, and the way it failed is what §2 forbids. The imperative was
|
|
348
|
+
* meant to collapse the acknowledgement into the action. Instead the model
|
|
349
|
+
* accepted the directive and complied with it, turn after turn: "I'll read
|
|
350
|
+
* cfg3.ts next, as directed by the observation", and then read cfg3.ts. Fifty-four
|
|
351
|
+
* reads for thirty files, one grep used three times as a side errand. The order
|
|
352
|
+
* was the model's own past action, so it never looked for a better way than the
|
|
353
|
+
* one already written down.
|
|
354
|
+
*
|
|
355
|
+
* The string survives behind SKILLSTATE_CONTINUATION=1 because a dead fix whose
|
|
356
|
+
* cost is measured is worth more than a fix nobody can price. The default is
|
|
357
|
+
* REPORT_MARKER.
|
|
358
|
+
*/
|
|
359
|
+
export declare const REPORT_MARKER = "[runtime]";
|
|
360
|
+
/** @deprecated kept only so the order path names the same constant it always did. */
|
|
361
|
+
export declare const CONTINUATION_MARKER = "[next step \u2014 do this now, do not describe it first]";
|
|
362
|
+
/** The same, under the name the call site reads it by. */
|
|
363
|
+
export declare const ORDER_MARKER = "[next step \u2014 do this now, do not describe it first]";
|
|
364
|
+
/**
|
|
365
|
+
* The one thing the host's own loop needs the model to know.
|
|
366
|
+
*
|
|
367
|
+
* Its second sentence used to read "each step ends with a real tool call — read
|
|
368
|
+
* the next file, or answer and stop." That clause was this repository's
|
|
369
|
+
* instruction, not a host constraint, and it pointed the wrong way: a
|
|
370
|
+
* thirty-file task produced 51, 53 and 54 `read` calls for thirty files, while
|
|
371
|
+
* a control with no step driver at all read one file and ran a single grep —
|
|
372
|
+
* nine calls for the same work.
|
|
373
|
+
*
|
|
374
|
+
* Worth being precise about how much this was worth, because the first version
|
|
375
|
+
* of this comment claimed the note CAUSED those counts and that is not
|
|
376
|
+
* supported. The transcripts show the model batching anyway — up to nine
|
|
377
|
+
* consecutive `read` calls inside a single turn. So the note was pushing the
|
|
378
|
+
* wrong way and was not obeyed literally. It is one defect, and removing it is
|
|
379
|
+
* right on its own terms; it is not the whole of the 43% of calls that are
|
|
380
|
+
* re-reads.
|
|
381
|
+
*
|
|
382
|
+
* The host does not limit a turn to one tool call. What limits it is eq. 1: the
|
|
383
|
+
* prompt is (P, Σₜ, Oₜ) and there is ONE Oₜ, so a turn that makes several calls
|
|
384
|
+
* keeps only the last result and the model must re-read the rest. That is not
|
|
385
|
+
* a guess — a turn that batched nine reads produced a model that said "the last
|
|
386
|
+
* observation re-read cfg8.ts" and then went and read cfg8 again, and the
|
|
387
|
+
* thirty-file run reached 103 reads for 29 files.
|
|
388
|
+
*
|
|
389
|
+
* So the note says the truth the model needs: one call per step, because the
|
|
390
|
+
* second one would be lost. An earlier version of this comment went the other
|
|
391
|
+
* way and told the model it could make as many calls as it liked. That was
|
|
392
|
+
* written before the single-observation consequence was understood, and it
|
|
393
|
+
* invited exactly the data loss the paper's own equation makes unavoidable.
|
|
394
|
+
*/
|
|
395
|
+
export declare const HOST_ACTION_NOTE: string;
|
|
396
|
+
/**
|
|
397
|
+
* Replace the model-facing context with exactly (P, Σₜ, Oₜ).
|
|
398
|
+
*
|
|
399
|
+
* Two edits, and both are required:
|
|
400
|
+
*
|
|
401
|
+
* - the A.4 prompt becomes the ONLY message, so no reasoning, action or tool
|
|
402
|
+
* output from earlier steps survives into this dispatch;
|
|
403
|
+
* - the host's own system prompt is KEPT. An earlier version replaced it
|
|
404
|
+
* wholesale, on the reasoning that P is the entire instruction surface and
|
|
405
|
+
* the default prompt tells the model to prefer parallel tool calls, which
|
|
406
|
+
* is incoherent with a single-step state machine. Measured: that reasoning
|
|
407
|
+
* was wrong about a consequence, because the default system prompt is also
|
|
408
|
+
* what carries the host's tool-use discipline. Replacing it with one
|
|
409
|
+
* sentence produced a model that emitted a correct patch and then never
|
|
410
|
+
* called a tool again — the run ended after the first file, three times
|
|
411
|
+
* running, on two models. The "parallel calls" worry is real but costs
|
|
412
|
+
* less than a dead loop; the trade is measured, not assumed.
|
|
413
|
+
*
|
|
414
|
+
* `systemPrefix` is where a host that cannot be the runtime says so; see
|
|
415
|
+
* {@link HOST_ACTION_NOTE}. It is optional because a deployment where
|
|
416
|
+
* something else does own the executor has no such gap to describe.
|
|
417
|
+
*
|
|
418
|
+
* The array is mutated in place: the host keeps the original reference.
|
|
419
|
+
*/
|
|
420
|
+
export declare function applyPaperContext(event: PaperContextEvent, built: PaperPrompt, systemPrefix?: string): PaperMessage;
|
|
421
|
+
//# sourceMappingURL=paper-mode.d.ts.map
|
|
@@ -0,0 +1 @@
|
|
|
1
|
+
{"version":3,"file":"paper-mode.d.ts","sourceRoot":"","sources":["../src/paper-mode.ts"],"names":[],"mappings":"AAAA;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;GAmGG;AAGH,OAAO,KAAK,EACV,WAAW,EACX,cAAc,EACd,UAAU,EACX,MAAM,kBAAkB,CAAC;AAG1B,4EAA4E;AAC5E,MAAM,WAAW,iBAAiB;IAChC,QAAQ,EAAE,KAAK,CAAC;QACd,EAAE,EAAE,MAAM,CAAC;QACX,IAAI,EAAE,MAAM,CAAC;QACb,OAAO,EAAE,OAAO,CAAC;QACjB,QAAQ,CAAC,EAAE,OAAO,CAAC;KACpB,CAAC,CAAC;IACH,MAAM,CAAC,EAAE,KAAK,CAAC;QAAE,IAAI,EAAE,MAAM,CAAC;QAAC,IAAI,CAAC,EAAE,MAAM,CAAC;QAAC,CAAC,GAAG,EAAE,MAAM,GAAG,OAAO,CAAA;KAAE,CAAC,CAAC;IACxE;;;;;;;;;;;;;;;;;;;;;;OAsBG;IACH,KAAK,CAAC,EAAE,MAAM,CAAC,MAAM,EAAE,OAAO,CAAC,CAAC;CACjC;AAED;;;;;;;;;;;GAWG;AACH,MAAM,MAAM,iBAAiB,GAAG,MAAM,GAAG,OAAO,CAAC;AAEjD,sFAAsF;AACtF,MAAM,MAAM,gBAAgB,GAAG,WAAW,GAAG;IAAE,MAAM,EAAE,iBAAiB,CAAA;CAAE,CAAC;AAE3E,oEAAoE;AACpE,MAAM,WAAW,YAAY;IAC3B,EAAE,EAAE,MAAM,CAAC;IACX,IAAI,EAAE,MAAM,CAAC;IACb,OAAO,EAAE,KAAK,CAAC;QAAE,IAAI,EAAE,MAAM,CAAC;QAAC,IAAI,EAAE,MAAM,CAAA;KAAE,CAAC,CAAC;IAC/C,QAAQ,EAAE,MAAM,CAAC,MAAM,EAAE,KAAK,CAAC,CAAC;CACjC;AAED,4EAA4E;AAC5E,eAAO,MAAM,gBAAgB,4BAA4B,CAAC;AAuE1D;;;;;;;;;;GAUG;AACH,wBAAgB,kBAAkB,CAChC,QAAQ,EAAE,KAAK,CAAC;IAAE,IAAI,EAAE,MAAM,CAAC;IAAC,OAAO,EAAE,OAAO,CAAA;CAAE,CAAC,GAClD,MAAM,CAQR;AAED;;;;;;;;;;;;;;;;;;;;;;;;;;GA0BG;AACH,wBAAgB,iBAAiB,CAC/B,QAAQ,EAAE,KAAK,CAAC;IAAE,IAAI,EAAE,MAAM,CAAC;IAAC,OAAO,EAAE,OAAO,CAAA;CAAE,CAAC,EACnD,GAAG,GAAE,MAAmB,GACvB,gBAAgB,CAQlB;AAED;;;;;;;;;;;;;;GAcG;AACH,wBAAgB,sBAAsB,CACpC,IAAI,EAAE,cAAc,EACpB,IAAI,EAAE,MAAM,GACX,cAAc,CAMhB;AAED,4CAA4C;AAC5C,MAAM,WAAW,kBAAkB;IACjC,IAAI,EAAE,cAAc,CAAC;IACrB,KAAK,EAAE,UAAU,CAAC;IAClB,QAAQ,EAAE,KAAK,CAAC;QAAE,IAAI,EAAE,MAAM,CAAC;QAAC,OAAO,EAAE,OAAO,CAAA;KAAE,CAAC,CAAC;IACpD;;;;;;;OAOG;IACH,QAAQ,CAAC,EAAE,MAAM,CAAC;IAClB;;;;;;;;;;;;;;;;;;;;;;;;;;;;OA4BG;IACH,YAAY,CAAC,EAAE,MAAM,CAAC;IACtB,sEAAsE;IACtE,gBAAgB,CAAC,EAAE,QAAQ,GAAG,OAAO,CAAC;CACvC;AAED,wEAAwE;AACxE,MAAM,WAAW,WAAW;IAC1B,0EAA0E;IAC1E,MAAM,EAAE,MAAM,CAAC;IACf,sEAAsE;IACtE,IAAI,EAAE,MAAM,CAAC;IACb,gDAAgD;IAChD,WAAW,EAAE,gBAAgB,CAAC;IAC9B,wCAAwC;IACxC,iBAAiB,EAAE,iBAAiB,CAAC;IACrC,mEAAmE;IACnE,iBAAiB,EAAE,MAAM,CAAC;CAC3B;AAED;;;;;;GAMG;AACH,wBAAgB,gBAAgB,CAAC,OAAO,EAAE,kBAAkB,GAAG,WAAW,CA4CzE;AAED;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;GA8BG;AACH;;;;;;;;;;;;;;;;;;;;;;;;;GAyBG;AACH,eAAO,MAAM,aAAa,cAAc,CAAC;AAEzC,qFAAqF;AACrF,eAAO,MAAM,mBAAmB,6DACuB,CAAC;AAExD,0DAA0D;AAC1D,eAAO,MAAM,YAAY,6DAAsB,CAAC;AAEhD;;;;;;;;;;;;;;;;;;;;;;;;;;;;;;GA8BG;AACH,eAAO,MAAM,gBAAgB,QAOlB,CAAC;AAEZ;;;;;;;;;;;;;;;;;;;;;;;GAuBG;AACH,wBAAgB,iBAAiB,CAC/B,KAAK,EAAE,iBAAiB,EACxB,KAAK,EAAE,WAAW,EAClB,YAAY,CAAC,EAAE,MAAM,GACpB,YAAY,CAad"}
|