copilotkit 4.9.4 → 4.9.17

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (47) hide show
  1. package/README.md +13 -12
  2. package/cli-build-info.json +8 -8
  3. package/index.js +1752 -672
  4. package/onboarding/index.json +15 -1
  5. package/onboarding/prompts/authenticate/start.md +129 -65
  6. package/onboarding/prompts/conversion/plan.md +103 -0
  7. package/onboarding/prompts/credentials/finalize-plan.md +97 -122
  8. package/onboarding/prompts/credentials/plan.md +26 -21
  9. package/onboarding/prompts/fallback/best-effort.md +99 -19
  10. package/onboarding/prompts/framework/ag2.md +7 -7
  11. package/onboarding/prompts/framework/agno.md +11 -8
  12. package/onboarding/prompts/framework/built-in.md +2 -2
  13. package/onboarding/prompts/framework/claude-sdk-python.md +7 -7
  14. package/onboarding/prompts/framework/claude-sdk-typescript.md +9 -9
  15. package/onboarding/prompts/framework/crewai-flows.md +27 -11
  16. package/onboarding/prompts/framework/deep-agents.md +8 -7
  17. package/onboarding/prompts/framework/google-adk.md +3 -3
  18. package/onboarding/prompts/framework/langgraph-fastapi.md +3 -3
  19. package/onboarding/prompts/framework/langgraph-python.md +3 -3
  20. package/onboarding/prompts/framework/langgraph-typescript.md +3 -3
  21. package/onboarding/prompts/framework/llamaindex.md +6 -6
  22. package/onboarding/prompts/framework/mastra.md +3 -3
  23. package/onboarding/prompts/framework/ms-agent-dotnet.md +3 -3
  24. package/onboarding/prompts/framework/ms-agent-harness-dotnet.md +7 -9
  25. package/onboarding/prompts/framework/ms-agent-python.md +3 -3
  26. package/onboarding/prompts/framework/pydantic-ai.md +26 -15
  27. package/onboarding/prompts/framework/strands-python.md +7 -6
  28. package/onboarding/prompts/framework/strands-typescript.md +9 -7
  29. package/onboarding/prompts/frontend/angular.md +16 -3
  30. package/onboarding/prompts/frontend/nextjs.md +3 -3
  31. package/onboarding/prompts/frontend/plan.md +6 -6
  32. package/onboarding/prompts/frontend/react-native.md +7 -2
  33. package/onboarding/prompts/frontend/react-spa.md +2 -2
  34. package/onboarding/prompts/frontend/vue.md +7 -2
  35. package/onboarding/prompts/implementation/build-and-validate.md +45 -79
  36. package/onboarding/prompts/proof/complete.md +8 -6
  37. package/onboarding/prompts/proof/oss-baseline.md +15 -50
  38. package/onboarding/prompts/proof/round-trip.md +83 -263
  39. package/onboarding/prompts/starter/clone.md +4 -4
  40. package/onboarding/prompts/subagent/create-plan.md +45 -0
  41. package/onboarding/prompts/subagent/implement-and-validate.md +70 -44
  42. package/onboarding/prompts/subagent/inspect-repository.md +25 -3
  43. package/onboarding/prompts/subagent/prove-oss-baseline.md +2 -1
  44. package/onboarding/prompts/subagent/prove-round-trip.md +143 -34
  45. package/onboarding/prompts/unsupported/no-validated-path.md +10 -2
  46. package/package.json +1 -1
  47. package/release/release-tool.js +1 -1
@@ -1,10 +1,13 @@
1
1
  # Prove the user journey
2
2
 
3
- Do not do the proof work yourself. Spawn one proof subagent.
4
-
5
- Give the subagent the full text of the proof brief at the end of this prompt, the selected
6
- framework, frontend, model, approved plan, selected documentation URLs, and validation
7
- evidence. Give it the documentation policy recorded during framework selection.
3
+ Do not do the proof work yourself. Spawn one proof subagent. Tell it to run
4
+ `npx --yes copilotkit@4.9.17 onboard read subagent/prove-round-trip` first and follow the
5
+ prompt it returns. If that read fails because the subagent cannot use the shell, stop that
6
+ subagent. Run the same command yourself, then spawn a fresh subagent with the returned prompt
7
+ and the same handoff. Give it the selected framework, frontend, model, approved plan, selected
8
+ documentation URLs, and validation evidence. Give it the documentation policy recorded during
9
+ framework selection or conversion planning. On a conversion, also give it the frozen
10
+ criterion. Give it the protected path list.
8
11
 
9
12
  Give it the browser or device control you recorded in the preflight as well. The subagent
10
13
  drives the surface and cannot see your environment, so without that finding it spends the
@@ -16,10 +19,20 @@ https://docs.copilotkit.ai/build-with-agents.md
16
19
 
17
20
  Wait for the subagent to finish.
18
21
 
19
- If the subagent proves the complete round trip, run
20
- `npx --yes copilotkit@4.9.4 onboard read proof/complete`. The round trip proves core success
21
- even if a continued-development tool fails. Keep the Skills and MCP results separate from
22
- the proof result.
22
+ For every protected-path audit in this prompt, run
23
+ `npx --yes copilotkit@4.9.17 onboard audit` from the target app directory. If its result
24
+ starts with `Status: blocked`, report the printed reason and use the route-out rules below.
25
+ A blocked audit proved nothing changed and is not a preservation failure. If a
26
+ protected-path audit reports a changed path, use the route-out rules below. Never repair,
27
+ reset, or revert a protected path.
28
+
29
+ If the proof result starts with `Status: passed`, run the protected-path audit. Continue to
30
+ `proof/complete` only if that audit passes. After the audit passes, run
31
+ `npx --yes copilotkit@4.9.17 onboard read proof/complete`. A performed surface outcome with
32
+ the full round trip is core success even if a continued-development tool fails. A skipped
33
+ surface outcome still enters `proof/complete` so the CLI records the blocked result. Do not
34
+ describe a skipped surface as proved. Keep the Skills and MCP results separate from the proof
35
+ result.
23
36
 
24
37
  If the round trip proves and something after it blocks this run anyway -- the managed
25
38
  Intelligence dashboard, the debugging surface, or a capability the approved plan already
@@ -30,266 +43,73 @@ not for a journey that worked and then hit a wall: routing a proved run there te
30
43
  developer that a framework and frontend that just worked are unsupported, and hands them a
31
44
  stop report instead of the application they now have.
32
45
 
33
- If the round trip fails, decide which kind of failure it is before you route. A failure
34
- caused by a file this run created or changed is a defect in the new work. Send it back to
35
- the proof subagent to fix and prove again, at most three attempts.
46
+ If the round trip fails, decide which kind of failure it is before you route. The proof
47
+ subagent can repair only project-owned processes, ports, and request options. A failure caused
48
+ by a source, configuration, dependency, or tracked-file change is a defect in the new work.
49
+
50
+ Retry the proof subagent only for a project-owned process, port, or request-option failure.
51
+ Give it the failure evidence and the failed attempt's pinned Step 1 record. Wait for the
52
+ proof subagent after each operational repair.
53
+ If the result passes, use the passed-result route above. Retry a failed result at most three
54
+ times. Use the route-out rules below for a blocked or third failed result.
55
+
56
+ Enter the code-repair branch only for a source, configuration, dependency, or tracked-file
57
+ defect.
58
+
59
+ If the starter shortcut created the selected path, spawn one repair subagent. Give it the
60
+ failed step, proof evidence, generated path list, selected framework, frontend, model, exact
61
+ target app directory, documentation, and policy.
62
+ Require it to change only files that the starter command generated. Do not read or return
63
+ secret values. Require its result to start with `Status: passed`, `Status: failed`, or
64
+ `Status: blocked`, followed by Files changed, Validation, and Blockers. Otherwise, send the
65
+ defect to the implementation subagent that validated the planned implementation path. Give it
66
+ the approved plan and proof evidence. Treat the worker that gets the defect as the repair worker.
67
+ Give the repair worker the protected path list. Do not let a generated or repaired path
68
+ overlap a protected path.
69
+
70
+ Require the full validation list to pass again. Wait for the repair worker to finish.
71
+ Continue only if its result starts with `Status: passed`. If the repair worker returns
72
+ `Status: failed`, retry the repair with its evidence. Run the protected-path audit after the
73
+ repair passes. Send the full validation list to the validator for the selected path. Use the
74
+ implementation worker for a planned implementation path. Use the repair worker for a starter
75
+ path. Wait for the validator to finish. If its result starts with `Status: passed`, continue.
76
+ If the validator returns `Status: failed`, return its evidence to the repair
77
+ worker. Repeat repair and validation at most three times. If either worker returns
78
+ `Status: blocked`, use the route-out rules below.
79
+
80
+ Run the protected-path audit again after validation passes. Continue only if its result
81
+ starts with `Status: passed`.
82
+
83
+ Restart each project-owned process changed by the repair. Then spawn a fresh proof subagent
84
+ with the full original proof handoff, failed proof evidence, and new validation evidence.
85
+ This handoff includes the plan, documentation, policy, current process IDs, ports,
86
+ surface-control state, the continued-tools guide, and the failed attempt's pinned Step 1
87
+ record. Carry that record so the second attempt sends the request the first one failed on.
88
+ A fresh subagent given the plan alone writes its own request where the plan named none, and
89
+ its grounding check then speaks about a question the failed attempt never asked.
90
+ Wait for the fresh proof subagent to finish.
91
+ If the fresh proof result starts with `Status: passed`, run the protected-path audit and use
92
+ the passed-result route above. If it starts with `Status: failed`, begin the next bounded
93
+ repair cycle. Use the route-out rules below for `Status: blocked`. Stop after three repair
94
+ and proof cycles.
36
95
 
37
96
  Route out only when the failure is not yours to fix, when the same proof still fails after
38
97
  three attempts, or when no evidence of the round trip can be produced. In those cases run
39
- `npx --yes copilotkit@4.9.4 onboard read unsupported/no-validated-path`. All three are
98
+ `npx --yes copilotkit@4.9.17 onboard read unsupported/no-validated-path`. All three are
40
99
  about the round trip itself. A round trip that proved is not one of them, whatever failed
41
100
  after it.
42
101
 
43
- ## Proof subagent brief
44
-
45
- Everything below the rule is the subagent's prompt. Give it verbatim.
46
-
47
- ---
48
-
49
- # Prove the complete round trip
50
-
51
- Use the proof steps and documentation URLs from the approved plan. Follow the documentation
52
- policy the main coding agent gives you before you start.
53
-
54
- Run the steps below in the order they appear. They are the proof. Do not design a
55
- different sequence of your own, and do not drop a step because an earlier one looked
56
- convincing. Step 6 has a web form and a React Native form: run the one that matches this
57
- journey's frontend, and run only that one.
58
-
59
- Record the result of every step as you go. A step with nothing recorded did not happen.
60
- Write each captured file to `.copilotkit/proof/` inside the project and name the path in
61
- the record. That directory holds regenerable evidence rather than application code. Where
62
- a browser or device tool writes to a location of its own, keep that location and record
63
- it instead.
64
-
65
- Gather what you need in as few commands as possible. Combine independent reads into one
66
- command rather than running them one at a time. Split a command only when its result decides
67
- what you run next.
68
-
69
- ## Step 1 -- Read what the proof needs
70
-
71
- Read all of this from the approved plan in one pass, before you start anything:
72
-
73
- - the exact request to send through the frontend, in the words the plan gave it,
74
- - the user-visible result that request has to produce,
75
- - the expected agent id,
76
- - where the project's own data lives, and which entities the answer has to name,
77
- - the start command for the agent and for the frontend, from the plan where it named
78
- one and from the project's own scripts otherwise,
79
- - the runtime URL.
80
-
81
- Where the plan named no request, write one that produces the outcome the plan named, and
82
- record the request you wrote. Every later step uses these words unchanged, so that the
83
- browser, the device, and the grounding check all speak about one request.
84
-
85
- ## Step 2 -- Start the agent and the frontend
102
+ If you stop here without taking that route -- a repair cycle you cannot finish, a limit in
103
+ your own harness, a run that has run out -- send one short report before you stop.
104
+ Run the feedback command without another developer question. The CLI telemetry gate decides
105
+ whether the report is sent.
86
106
 
87
- Read the port the developer's agent already serves from this project's own configuration.
88
- Do not assume a default, and do not start a second copy of an agent this project is
89
- already running. Before you bind any new server, check that the port is free and pick
90
- another one if it is not. Record every port you used.
91
-
92
- The table names what each frontend's own scaffolder writes, for recognizing a port in a
93
- started process's output. The project's configuration wins over the table.
94
-
95
- | Frontend | Dev server the scaffolder writes |
96
- | ------------ | -------------------------------- |
97
- | Next.js | `next dev`, port 3000 |
98
- | React SPA | Vite, port 5173 |
99
- | Vue 3 | Vite, port 5173 |
100
- | Angular | `ng serve`, port 4200 |
101
- | React Native | Metro, port 8081 |
102
-
103
- Where this project's runtime runs as a process of its own, the frontend documentation
104
- puts it on port 8200. Use the runtime URL from step 1 rather than that number.
105
-
106
- Start each server in the background with the project's own script. Then wait for it to
107
- answer rather than for a fixed number of seconds:
108
-
109
- ```bash
110
- ready=
111
- for _ in $(seq 90); do
112
- curl -fs -o /dev/null "<url>" && ready=1 && break
113
- sleep 1
114
- done
115
- [ "$ready" = 1 ] && echo "up" || echo "no answer from <url> after 90 seconds"
107
+ ```text
108
+ npx --yes copilotkit@4.9.17 onboard feedback
116
109
  ```
117
110
 
118
- Read the loop's own last line rather than assuming it ended because the server answered.
119
- A server that never answered has written the reason to its own output, and reading that
120
- output is faster than starting it again.
121
-
122
- Leave the agent and frontend servers running after proof. Record each process ID and a
123
- safe command that stops that process. Record the frontend URL and the commands that start
124
- both servers again.
125
-
126
- ## Step 3 -- Identify the process that answered
127
-
128
- Before you trust the agent, confirm that the process answering is the one in this
129
- repository. `verify` reports which agents the runtime declares and has nothing to compare
130
- them against, and `--round-trip` proves an agent answers under the declared id without
131
- proving which deployment did, so this comparison is yours. A health endpoint that returns
132
- success proves only that something listens on
133
- that port. An agent from earlier work often still holds it, and a stale process answers
134
- as though it were the new one. Ask the running agent which graph or agent id it serves
135
- and compare that with the id declared in this project.
136
-
137
- If they do not match, find out whose process it is before you signal anything.
138
- `lsof -ti :<port> -sTCP:LISTEN` gives the process id, and `lsof -a -p <pid> -d cwd` gives
139
- the directory it runs in. Stop it only when that directory is inside this project, and
140
- stop its children before the parent so nothing survives by reparenting. A holder outside
141
- this project belongs to other work: leave it running, report it, and bind to another port.
142
- Never stop a process because its command line matches a name. One `pkill` pattern reaches
143
- every project on the machine and takes down work that has nothing to do with this run.
144
- Never continue against a process you cannot identify, and never report a round trip proven
145
- by one.
146
-
147
- Address a local agent by host name rather than by an IP literal. Some local agents bind
148
- IPv6 only, so an IPv4 literal fails against the correct port.
149
-
150
- ## Step 4 -- Check the wiring
151
-
152
- With both running, check the wiring in one command before you open a browser:
153
- `npx --yes copilotkit@4.9.4 verify --json`. Add `--runtime-url` when the runtime is not at
154
- `http://localhost:3000/api/copilotkit`. Read the individual checks rather than the summary
155
- alone: a check reported `undetermined` did not run, and that is not a pass. Fix anything
156
- that is not a pass before the browser, because a browser failure stacked on broken wiring
157
- costs a round of debugging to reach an answer this command already gave.
158
-
159
- Treat `intelligence_consumed` as the check that matters most here. A journey that finishes
160
- with the Intelligence credential never read looks complete and proves nothing about the
161
- paid surface. `api_key_authenticates` passing beside it says the key is real and the
162
- runtime never used it.
163
-
164
- `intelligence_thread_routes` fails when the runtime reports a license but serves no
165
- thread routes, which means saved Threads cannot load in a browser. The usual cause is a
166
- handler mounted `mode: "single-route"`: remove that option so the handler serves its full
167
- route set, and mount it at a catch-all route. If instead that check is `undetermined`
168
- because the runtime reports no thread-endpoint state, the runtime predates the field.
169
- Record that and move on -- there is nothing to repair.
170
-
171
- ## Step 5 -- Prove that the agent runs
172
-
173
- Run `npx --yes copilotkit@4.9.4 verify --round-trip --json`. It sends one request through
174
- the runtime and reads the answer back from the thread, so it separates an agent that is
175
- configured from an agent that works. Use `--agent <id>` when the runtime declares more
176
- than one. If it reports `user-not-identified`, this project's `identifyUser` reads a
177
- session the CLI does not carry: pass what it reads with `--header "Name: value"` and run
178
- it again, because an auth-gated app refusing an unauthenticated caller is that app
179
- working. Do not continue until this passes, and never report a round trip proven without
180
- it.
181
-
182
- ## Step 6 -- Drive the surface
183
-
184
- Now send one real request through the running frontend, on this journey's own surface.
185
- Make sure that the request passes through CopilotKit and reaches the selected agent, and
186
- that the frontend receives working generative UI from the agent. This is the step that
187
- covers realtime delivery, the frontend provider being wired to this runtime, and a
188
- generative UI component actually rendering, and no command-line check reaches any of them.
189
- It is not optional polish: a run that skips it has proven the agent and not the journey,
190
- and the graph ends such a run as blocked rather than complete.
191
-
192
- Use the surface control the main coding agent recorded for your environment. It either had
193
- one already or registered one before this step, so that finding is the answer and there is
194
- nothing here for you to go looking for. Do not add a browser driver or a device tool to this
195
- project: a devDependency and a browser download land in the diff and tax a repository that
196
- never asked for one, which is a different thing from the server registered against the
197
- coding agent. If nothing in your environment can drive the surface this journey needs, skip
198
- this step rather than installing one, and report the skip outcome named below.
199
-
200
- Never report a result you did not see, on either surface.
201
-
202
- ### Step 6a -- Web frontends: React SPA, Next.js, Angular, Vue 3
203
-
204
- The surface is a browser, and it also covers browser-origin CORS and CSP, which a CLI
205
- request never exercises. Drive it with the browser control step 6 named.
206
-
207
- 1. Open the frontend URL from step 2 and wait for the page to finish loading.
208
- 2. Take one page snapshot. Record whether the CopilotKit surface is on the page. A page
209
- that renders without it is a wiring failure rather than a proof to retry.
210
- 3. Read the browser console before you type anything, and record every error already
211
- there. An error at this point belongs to page load rather than to the request.
212
- 4. Enter the step 1 request into the CopilotKit input, in the words step 1 recorded, and
213
- submit it.
214
- 5. Wait for the assistant turn to finish rather than for a fixed number of seconds. The
215
- turn is finished when the streamed text stops growing and the generative UI component
216
- has rendered.
217
- 6. Take one screenshot of the finished turn and record where you wrote it.
218
- 7. Read the browser console a second time, and record every error that step 3 did not
219
- already list. Those belong to the request.
220
- 8. Read the network requests. Record every request the page made to the runtime endpoint
221
- and the status each returned. An answer on the page with no successful request to this
222
- runtime behind it came from something else, and that is a failed proof rather than a
223
- passing one.
224
- 9. Record one line for the finished turn: the time, the page URL, the element you read
225
- the answer from, and what that element showed.
226
-
227
- Report exactly one of `performed`, `skipped-no-browser-tool`, or `failed` for a web
228
- frontend.
229
-
230
- ### Step 6b -- React Native
231
-
232
- The surface is a device or emulator, and a browser cannot stand in for it.
233
-
234
- 1. List the booted devices in one command: `adb devices -l`. With no booted device this
235
- step is `skipped-no-device`. Do not substitute a browser.
236
- 2. Build and install the app on the booted device with the project's own script, which
237
- an Expo or React Native CLI project names in `package.json`.
238
- 3. Clear the log buffer before the request: `adb logcat -c`.
239
- 4. Open the CopilotKit surface in the running app and enter the step 1 request, in the
240
- words step 1 recorded.
241
- 5. Wait for the assistant turn to finish rather than for a fixed number of seconds.
242
- 6. Capture the terminal state with
243
- `adb exec-out screencap -p > .copilotkit/proof/surface.png`, and record that path.
244
- 7. Read the log for the request with `adb logcat -d -t 500`. A redbox is a runtime failure
245
- the terminal state never shows.
246
- 8. Record one line for the finished turn: the time, the platform, the device id, the
247
- screen you were on, and what that screen showed.
248
-
249
- An Android emulator reaches a runtime on the host machine at `10.0.2.2` rather than at
250
- `localhost`. A request that fails against the runtime with nothing in the runtime's own
251
- log is that, rather than a broken runtime.
252
-
253
- Report exactly one of `performed`, `skipped-no-device`, or `failed` for React Native.
254
-
255
- ## Step 7 -- Check the answer against the project's data
256
-
257
- Where the answer is meant to be about data the project holds, check it against that data.
258
- Read the entities the project holds -- the ids, names, or records the answer claims to
259
- describe -- and confirm the answer names those and no others. Record the entities you
260
- compared. Where the outcome of this journey references no project data, record that
261
- instead, and do not invent a comparison to pass this step.
262
-
263
- An answer that renders correctly over entities the project does not hold looks the same as
264
- a correct one in a browser, in a screenshot, and in a video, so this comparison is the only
265
- stage that separates them. An answer that names an entity the project does not hold is a
266
- failed proof, not a passing one. Find which of these it is before you change anything: the
267
- project's data never reaches the agent, the agent receives it and its instructions ignore
268
- it, the page loads its data after the context was registered, or the run wired a different
269
- source than the page renders. Fix that cause, then prove again.
270
-
271
- ## Step 8 -- Compare with the recorded OSS baseline
272
-
273
- Run this step only for a recorded `both-oss` starting state. Compare the final round trip
274
- with the recorded OSS baseline. The same frontend request must still reach the same agent
275
- and produce the same kind of user-visible result. The runtime must now report
276
- `licenseStatus`, and the authenticated Intelligence checks must pass. Record both before
277
- and after evidence.
278
-
279
- ## Step 9 -- Set up the continued-development tools
280
-
281
- Fetch the continued-development guide from the main coding agent with the proof
282
- documentation. Try the continued-development tools after the application passes proof.
283
- Use it to install the project-scoped CopilotKit Skills.
284
- Use it to configure the CopilotKit documentation MCP server for the current coding agent.
285
-
286
- Do not validate whether the Skills or MCP server installed correctly. Record the command
287
- result for each attempt. Report each tool result separately. A tool error does not change
288
- the proof result.
289
-
290
- ## Step 10 -- Return the result
291
-
292
- Return the proof or the exact failed step to the main coding agent, together with the
293
- input, the visible result, the relevant process status, the evidence locations, the
294
- surface-check outcome, and which surface that outcome speaks for. Do not return secret
295
- values. Stop after you return the result.
111
+ Write at most four lines to standard input: the step you stopped at and what stopped it.
112
+ Send no secrets, source code, logs, or command output. The command refuses a report that
113
+ carries any of those, prints the reason, and exits zero. A refused report is not a failed
114
+ step. Report friction only from a run that finished, never from a stop. A run that dies in
115
+ this phase is the one this graph most needs to hear about and the one it hears from least.
@@ -36,7 +36,7 @@ instruction in this prompt.
36
36
 
37
37
  If the developer did not name an Intelligence project, list the available projects:
38
38
 
39
- `npx --yes copilotkit@4.9.4 project list --json`
39
+ `npx --yes copilotkit@4.9.17 project list --json`
40
40
 
41
41
  Ask the developer to select a project or give a name for a new project. If the developer
42
42
  already gave this answer, do not ask again. Do not read a secret value. Do not show or
@@ -46,7 +46,7 @@ Run the command from the parent directory. Do not inspect another entry in the p
46
46
  directory. Replace each placeholder with the recorded value. Do not run a placeholder as
47
47
  a shell argument.
48
48
 
49
- `npx --yes copilotkit@4.9.4 init --name <project-name> --framework <framework-id> --channel none --no-banner --project <slug-or-id> --install`
49
+ `npx --yes copilotkit@4.9.17 init --name <project-name> --framework <framework-id> --channel none --no-banner --project <slug-or-id> --install`
50
50
 
51
51
  If the developer wants a new project, use `--create <name>` instead of
52
52
  `--project <slug-or-id>`. If the developer asked to skip the dependency install, use
@@ -60,8 +60,8 @@ account. The command does not need terminal input.
60
60
  If the command succeeds, do not rebuild the starter by hand. Inspect only the generated
61
61
  paths inside the target directory. Record the files, install result, project connection,
62
62
  and validation commands. Then run
63
- `npx --yes copilotkit@4.9.4 onboard read proof/round-trip`.
63
+ `npx --yes copilotkit@4.9.17 onboard read proof/round-trip`.
64
64
 
65
65
  If the command fails, report its exact error and do not claim that the starter is ready.
66
66
  Then run
67
- `npx --yes copilotkit@4.9.4 onboard read unsupported/no-validated-path`.
67
+ `npx --yes copilotkit@4.9.17 onboard read unsupported/no-validated-path`.
@@ -73,6 +73,50 @@ Create one plan for implementation, validation, and proof. Name each file or are
73
73
  change. Name the commands that can validate the result.
74
74
  Name the commands or user path that prove a real generative-UI round trip.
75
75
 
76
+ The terminal is one component rendered through this frontend's own registration -- the
77
+ `useComponent` hook, or Angular's `registerComponent` -- over entities the project holds.
78
+ Plan that component: name it, name the fields it shows, and name which of the project's own
79
+ records fill them. Do not plan a declarative A2UI card unless the developer asks for one:
80
+ the terminal needs no renderer package and nothing added to the agent, because the tool is
81
+ declared by the frontend and forwarded over AG-UI.
82
+
83
+ A prose answer is not the terminal. The structure is what the grounding comparison reads: a
84
+ component with named fields can be checked field by field against the records the project
85
+ holds, and that comparison is the only thing separating a correct answer from a confident
86
+ invented one.
87
+
88
+ For a recorded `both-oss` starting state, plan no generative-UI component. The terminal is
89
+ the baseline request still working over the same frontend, the runtime now constructed
90
+ with `intelligence`, and the thread that request creates listed in this frontend's threads
91
+ drawer. A `both-oss` project already has its user-visible result. The conversion preserves
92
+ that result rather than replacing it with a new one, and a generative-UI component added
93
+ here is work the developer did not ask for.
94
+
95
+ Plan the threads drawer itself: add it from the selected drawer page, where this frontend
96
+ does not already render one. Where this journey's frontend framework ships no threads
97
+ drawer -- React Native --, plan that the thread is proved in the managed Intelligence
98
+ dashboard instead.
99
+
100
+ ## Order the plan into steps
101
+
102
+ Create one ordered list of implementation steps. For each step, name the outcome, the exact
103
+ files or directories it changes, its validation commands, and the steps it depends on.
104
+
105
+ One implementation subagent runs every step in that order. Do not split the steps across
106
+ concurrent subagents, and do not plan a separate step to reconcile a split.
107
+
108
+ Use the protected path list you were given. It is the baseline the CLI captured, not a
109
+ list to re-derive: a path you leave out of it is not protected, and a path you add to it
110
+ is not in the baseline the audit reads. Do not name a protected path as a path a step
111
+ changes. Do not name a path that overlaps a protected path. The paths overlap when they are
112
+ equal or either path is an ancestor directory on a path-segment boundary. `apps/a` overlaps
113
+ `apps/a/src`, but not `apps/ab`. `app.ts` does not overlap `app.tsx`. If the plan needs
114
+ one, return `Status: blocked` before approval.
115
+ Return the protected path list as part of the plan.
116
+
117
+ Keep final validation and proof outside the implementation steps. The developer approves
118
+ one plan.
119
+
76
120
  Keep a production build out of the validation commands. A type check plus the real round
77
121
  trip is the proof, and the round trip runs in development mode. Name the production build
78
122
  as a follow-up for the developer instead. Never raise a bundle budget or relax a lint rule
@@ -88,5 +132,6 @@ Gather what you need in as few commands as possible. Combine independent reads i
88
132
  command rather than running them one at a time. Split a command only when its result decides
89
133
  what you run next.
90
134
 
135
+ Start with `Status: passed`, `Status: failed`, or `Status: blocked`.
91
136
  Return the plan, the URLs that you read, and each documentation gap. Stop after you return
92
137
  the findings to the main coding agent.
@@ -1,59 +1,85 @@
1
1
  # Implement and validate the approved plan
2
2
 
3
+ Implement every step of the approved plan in plan order, and run the full validation list.
4
+ Do not repeat the full repository inspection.
5
+
6
+ Do not change a protected path.
7
+ Do not change a path that overlaps a protected path. Paths overlap when they are equal or
8
+ either path is an ancestor directory on a path-segment boundary. `apps/a` overlaps
9
+ `apps/a/src`, but not `apps/ab`. `app.ts` does not overlap `app.tsx`.
10
+
11
+ Change only the files and directories the approved plan names as changeable. If the work
12
+ requires a file the plan does not name, stop and return the blocker. Do not run a
13
+ repository-wide formatter.
14
+
15
+ Before validation, read the protected path list. Inspect the current changed and untracked
16
+ paths. Keep protected paths out of the run path set. Compare every changed path with the
17
+ paths the plan named. Here, a changed path means one in the run path set. If a changed path
18
+ is outside them, return `Status: blocked` before validation. Run the approved validation
19
+ commands one at a time. Fix validation and proof defects only in the paths the plan named.
20
+
3
21
  Use only the approved plan and the selected documentation URLs. Follow the documentation
4
22
  policy the main coding agent gives you before you change the project.
5
23
 
6
24
  Add only the props, options, and imports that appear in the fetched documentation. A
7
25
  remembered API from an earlier CopilotKit version will fail type checking against this
8
- release, so do not decorate a documented example with anything it does not show.
26
+ release. Do not add details that the fetched documentation omits.
27
+
28
+ The documentation supplies the wiring. The project supplies the application. Use that
29
+ wiring in the project's domain. Do not copy an example's agent name, tools, or data.
9
30
 
10
- The documentation supplies the wiring. The project supplies the application. That rule
11
- governs props, options, and imports, and it stops there. A documentation example names a
12
- domain of its own to make itself readable. Build what the plan names, in the project's
13
- own domain, from the wiring the page shows. Do not carry the page's example domain into
14
- the project: not its agent name, not its tools, not its data.
31
+ Preserve each existing agent and frontend. Add only missing parts and the CopilotKit
32
+ connection. Do not read, show, store, or return secrets. Do not read a file outside the
33
+ project directory: not for a credential, and not for an API question the fetched
34
+ documentation answers. Ask the main coding agent for a missing credential.
15
35
 
16
- Preserve each agent or frontend that already exists. Add only the missing parts and the
17
- CopilotKit connection. Do not read, show, store, or return secret values. Do not read a
18
- file outside the project directory: not for a credential, and not for an API question the
19
- fetched documentation answers. A missing credential is the main coding agent's to ask
20
- for.
36
+ An existing `README.md` is the developer's, not this run's. Leave it exactly as you found
37
+ it: not the title, not a section, not a line. Where this run has documentation to write,
38
+ write it to a new file and name that file when you report the result. On an empty project
39
+ the README is the whole of the brief -- the only place the developer said what they
40
+ wanted, and the thing they read this run's result against -- so the rule binds hardest
41
+ where the file looks emptiest.
21
42
 
22
43
  For a recorded `both-oss` baseline, preserve the existing agent, frontend, CopilotKit
23
- integration, user-visible request, and runtime behavior. Do not replace the existing
24
- persistence. Change only the approved project selection, the approved CopilotKit
25
- dependency upgrade, and the Intelligence runtime wiring.
26
- Run the recorded baseline checks again after the Intelligence wiring and report any
27
- regression.
28
-
29
- Apply the planned CopilotKit dependency upgrade before you wire Intelligence, as its own
30
- step, and record each version it changed. Then re-run the recorded baseline checks. A
31
- regression at this point can only be the upgrade, which is why it is checked here rather
32
- than after the wiring.
33
-
34
- If the baseline regresses, restore the manifest and lockfile to their recorded state,
35
- leave Intelligence unwired, and report the regression with the failing check. A
36
- half-upgraded project with no working chat is worse than the one this run was given, and
37
- finishing the wiring on top of a broken baseline hides which change broke it.
38
-
39
- Report each `@copilotkit/*` version before and after the upgrade. Where the plan named no
40
- dependency change, say that the installed versions already met the floor.
41
-
42
- When the documentation creates the frontend with the framework's own scaffolder,
44
+ integration, user-visible request, runtime behavior, and persistence. Do not replace the
45
+ existing persistence. Change only the approved project selection, CopilotKit dependency
46
+ upgrade, Intelligence runtime wiring, and the approved threads drawer: the drawer belongs
47
+ in that list because the approved plan already named it, not because this step invents a
48
+ new one. After wiring, run the recorded baseline checks and report any regression.
49
+
50
+ Where the plan names a CopilotKit dependency upgrade:
51
+
52
+ 1. Apply the planned CopilotKit dependency upgrade before you wire Intelligence.
53
+ 2. Report each `@copilotkit/*` version before and after the upgrade.
54
+ 3. Then re-run the recorded baseline checks.
55
+
56
+ If the baseline regresses, restore the manifest and lockfile to their recorded state, leave
57
+ Intelligence unwired, and report the regression with the failing check. Stop. Where the plan
58
+ names no dependency change, report that the installed versions met the floor.
59
+
60
+ When the documentation uses the framework's scaffolder,
43
61
  run that scaffolder rather than hand-authoring what it emits.
44
- Do not re-declare a compiler option it set. Every real project of that framework
45
- inherits those defaults, so a hand-written replacement measures a configuration no
46
- developer has. Where the plan needs configuration the scaffolder did not write,
47
- extend the scaffolder's file and set only what you add.
62
+ Do not re-declare a compiler option it set.
63
+ If the plan needs more configuration, extend the scaffolder's file and set only what you add.
64
+
65
+ Where the plan names data the project already holds, share that data with the agent through
66
+ the documented context API. Tell the agent to refuse to answer about entities the shared
67
+ context does not carry and to name what is missing.
68
+
69
+ An empty context set is a condition to state rather than an input to work around. Keep
70
+ these two cases apart in the instructions you write. The context set holds no records:
71
+ the answer says the page sent none. The context set holds records and none of them match
72
+ the request: the answer says the request names something the page does not hold, and
73
+ names what the page did send. An agent that runs the two together reads an empty page as
74
+ a page to fill in.
48
75
 
49
- Where the plan names data the project already holds, share that data with the agent
50
- through the context API the selected documentation names for this frontend. Then write the agent's instructions to
51
- refuse to answer about entities the shared context does not carry, and to name what is
52
- missing instead. An agent with no context and no refusal invents plausible entities, and
53
- no part of the run errors.
76
+ Where a tool renders one of those records, take every field of the record from the
77
+ context set. A field the model is allowed to fill is a field it fills, so an optional
78
+ field on that tool's schema turns an empty context set into a well-formed record the
79
+ project does not hold, and the rendered result looks the same either way.
54
80
 
55
- Run the validation commands from the approved plan. Record the changed files, command
56
- results, and each error. Do not claim that the real user journey works in this phase.
81
+ Run the full validation list. Record changed files, command results, and errors. Do not
82
+ claim that the real user journey works in this phase.
57
83
 
58
84
  Report the runtime constructor you wrote. Name whether it passes `intelligence` or
59
85
  `runner`. A runtime built with `runner` is the SSE runtime and never reads the credential,
@@ -63,5 +89,5 @@ Gather what you need in as few commands as possible. Combine independent reads i
63
89
  command rather than running them one at a time. Split a command only when its result decides
64
90
  what you run next.
65
91
 
66
- Return the implementation result and validation evidence to the main coding agent. Stop
67
- after you return the result.
92
+ Start with `Status: passed`, `Status: failed`, or `Status: blocked`.
93
+ Return exactly four sections: Status, Files changed, Validation, and Blockers. Stop.