@andreprado/agentkit 0.1.0-alpha.21 → 0.1.0-alpha.23

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -47,4 +47,6 @@ Open the printed `Chat:` URL and report it to the owner.
47
47
 
48
48
  ## Production Handoff
49
49
 
50
+ `npm run agentkit -- deploy` prints a production handoff after a successful deploy. Use it as the source of truth for the deploy URL, UI command, secret status, database/schema artifact, integration connect commands, smoke status, and next recommended command.
51
+
50
52
  Report changed files, required env/secret names, database schema changes, deploy order, smoke checks, rollback concerns, and whether the provider was still `test/fake`.
@@ -7,6 +7,8 @@ description: Use when adding, editing, or running AgentKit eval files, including
7
7
 
8
8
  Use evals after chat works and before claiming behavior is stable.
9
9
 
10
+ `npm run eval` uses temporary local SQLite storage for eval execution. Eval conversations and tool calls do not write to the normal `.agentkit/agentkit.db`, so evals can run while local chat or `npm run dev` is using the development database.
11
+
10
12
  ## Workflow
11
13
 
12
14
  1. Create or edit `evals/<name>.eval.ts`.
@@ -21,6 +23,19 @@ Use evals after chat works and before claiming behavior is stable.
21
23
  10. Do not put secrets or real client PII in evals.
22
24
  11. For tools that write externally, delete, charge money, send email, or call real customer systems, branch on `ctx.runtime.environment === "eval"` inside the registered tool.
23
25
 
26
+ ## What To Test
27
+
28
+ Create evals proactively from the brief and spec. Common high-value evals:
29
+
30
+ - identity and scope: the agent says who it is and refuses out-of-scope work;
31
+ - intake: required fields such as name, phone, email, account id, date, or budget are collected before action;
32
+ - confirmation: external writes, deletes, messages, charges, and bookings do not happen before explicit confirmation;
33
+ - privacy: raw tool output, full calendars, internal IDs, retrieval metadata, secrets, and stack traces are not shown to the client;
34
+ - timezone and schedule: eval `now` is frozen, `timeZone` is honored, and tool payloads use the intended local date/time;
35
+ - defaults and constraints: durations, allowed hours, allowed regions, max/min values, and business rules are asserted;
36
+ - unhappy paths: rate limits, timeouts, missing auth, unavailable slots, empty results, and validation errors produce safe user-facing responses;
37
+ - regression: every real conversation bug gets the smallest eval that would have failed before the fix.
38
+
24
39
  ## Assertion Shape
25
40
 
26
41
  Use this shape first:
@@ -38,13 +38,23 @@ npm run agentkit -- improve collect --deploy --conversation-id <conversation-id>
38
38
  .agentkit/improve/<run>/traces/
39
39
  ```
40
40
 
41
- 3. Generate regression evals:
41
+ 3. Before patching, identify the smallest testable lesson from each relevant trace:
42
+
43
+ - Did the agent miss a required field?
44
+ - Did it expose raw tool output, a full schedule, an internal id, or a technical error?
45
+ - Did it write externally without confirmation?
46
+ - Did it use the wrong timezone, duration, business hour, or availability assumption?
47
+ - Did a tool error or provider limit produce a bad client response?
48
+
49
+ 4. Generate regression evals:
42
50
 
43
51
  ```sh
44
52
  npm run agentkit -- improve evals .agentkit/improve/<run>
45
53
  ```
46
54
 
47
- 4. Patch the capsule. Likely files:
55
+ 5. Review or rewrite generated evals so they assert the behavior, not brittle transcript wording. If the bug involved a tool call, assert the persisted tool call input or absence of the unsafe call.
56
+
57
+ 6. Patch the capsule. Likely files:
48
58
 
49
59
  ```txt
50
60
  prompts/instructions.md
@@ -54,7 +64,7 @@ knowledge/
54
64
  evals/
55
65
  ```
56
66
 
57
- 5. Verify:
67
+ 7. Verify:
58
68
 
59
69
  ```sh
60
70
  npm run typecheck
@@ -63,7 +73,7 @@ npm run eval
63
73
  npm run agentkit -- replay .agentkit/improve/<run> --against local
64
74
  ```
65
75
 
66
- 6. Deploy only after local evals and replay pass:
76
+ 8. Deploy only after local evals and replay pass:
67
77
 
68
78
  ```sh
69
79
  npm run agentkit -- deploy --smoke "hello"
@@ -44,6 +44,18 @@ For Google Calendar, do not configure create-only access. Include `GOOGLECALENDA
44
44
 
45
45
  Managed Composio write actions require tool input `confirmed: true` by default. Set it only after the owner/user confirms the exact external change. Use `confirmExternalWrites: false` only when the capsule implements an equivalent confirmation guard elsewhere.
46
46
 
47
+ ## Testability
48
+
49
+ Do not rely on the real connected app for ordinary evals. When adding an integration, also add deterministic coverage for:
50
+
51
+ - the safe path, such as free/busy before calendar create;
52
+ - missing confirmation before an external write;
53
+ - provider errors such as 429, timeout, missing auth, empty result, or unavailable slot;
54
+ - privacy rules, such as not showing a full calendar or raw provider payload to the client;
55
+ - payload invariants, such as timezone conversion, duration, recipients, or record ids.
56
+
57
+ Use eval-safe branches inside capsule tools, local fixtures, or `test/fake` behavior when the hosted integration cannot run locally without touching the real provider.
58
+
47
59
  ## Commands
48
60
 
49
61
  ```sh
@@ -54,6 +66,16 @@ npm run agentkit -- integrations status --toolkit googlecalendar
54
66
  npm run agentkit -- integrations connect composio --toolkit gmail
55
67
  ```
56
68
 
69
+ `integrations connect composio` requires a hosted deploy because the connect link is deploy-scoped. After declaring `composioManaged(...)`, tell the owner the sequence is:
70
+
71
+ ```sh
72
+ npm run agentkit -- deploy doctor
73
+ npm run agentkit -- deploy
74
+ npm run agentkit -- integrations connect composio --toolkit googlecalendar
75
+ ```
76
+
77
+ The deploy handoff prints the connect command for each configured toolkit.
78
+
57
79
  ## Rules
58
80
 
59
81
  - Do not add `COMPOSIO_API_KEY` to `.env.schema` for managed Composio.
@@ -1,6 +1,6 @@
1
1
  ---
2
2
  name: agentkit-provider
3
- description: Use when switching an AgentKit capsule from the deterministic test/fake provider to a real Pi-backed provider such as OpenAI, Anthropic, or OpenRouter, or when verifying provider secrets and model behavior.
3
+ description: Use when switching an AgentKit capsule from the deterministic test/fake provider to a real Pi-backed provider such as OpenAI, Anthropic, OpenRouter, OpenCode Zen, or OpenCode Go, or when verifying provider secrets and model behavior.
4
4
  ---
5
5
 
6
6
  # AgentKit Provider
@@ -9,7 +9,7 @@ Use this when `test/fake` is no longer enough.
9
9
 
10
10
  ## Rule
11
11
 
12
- Do not choose a real provider automatically. Ask the owner which provider to use: OpenRouter, OpenAI, Anthropic, or another supported provider.
12
+ Do not choose a real provider automatically. Ask the owner which provider to use: OpenRouter, OpenAI, Anthropic, OpenCode Zen, OpenCode Go, or another supported provider.
13
13
 
14
14
  ## Workflow
15
15
 
@@ -44,6 +44,24 @@ secrets: ["OPENROUTER_API_KEY"],
44
44
 
45
45
  Prefer OpenRouter model ids or aliases known to the installed Pi SDK, such as `~google/gemini-flash-latest`. If a raw OpenRouter id is newer than Pi's registry, AgentKit passes it through to OpenRouter with conservative unknown-model metadata; OpenRouter can still reject invalid, inaccessible, or unsupported models.
46
46
 
47
+ OpenCode Zen:
48
+
49
+ ```ts
50
+ provider: { name: "opencode", model: "big-pickle" },
51
+ secrets: ["OPENCODE_API_KEY"],
52
+ ```
53
+
54
+ Use an OpenCode Zen model id listed by the installed Pi SDK, such as `big-pickle`, `deepseek-v4-flash-free`, `claude-sonnet-4-5`, or `gpt-5.4-mini`.
55
+
56
+ OpenCode Go:
57
+
58
+ ```ts
59
+ provider: { name: "opencode-go", model: "deepseek-v4-flash" },
60
+ secrets: ["OPENCODE_API_KEY"],
61
+ ```
62
+
63
+ Use an OpenCode Go model id listed by the installed Pi SDK, such as `deepseek-v4-flash`, `deepseek-v4-pro`, `glm-5.1`, `kimi-k2.6`, `minimax-m2.7`, or `qwen3.6-plus`.
64
+
47
65
  ## Verification
48
66
 
49
67
  ```sh
@@ -16,6 +16,8 @@ Use this when the agent needs code, an API, live data, a write, or an external a
16
16
  5. Keep secret names in `.env.schema`; values stay in ignored `.env` or hosted managed secrets.
17
17
  6. Use `ctx.clock` for date-sensitive tool logic instead of calling `new Date()` directly.
18
18
  7. Add eval guards for destructive or external side effects.
19
+ 8. Add deterministic fixtures, fake branches, or direct tool inputs for important success and failure paths.
20
+ 9. Add evals that assert the tool is called with safe inputs, or not called when confirmation/intake is missing.
19
21
 
20
22
  ## Examples
21
23
 
@@ -64,6 +64,7 @@ node_modules/
64
64
  OPENAI_API_KEY=
65
65
  ANTHROPIC_API_KEY=
66
66
  OPENROUTER_API_KEY=
67
+ OPENCODE_API_KEY=
67
68
  `,
68
69
  },
69
70
  {
@@ -193,7 +194,7 @@ This is an AgentKit support Agent Capsule.
193
194
 
194
195
  ## Coding Agent Workflow
195
196
 
196
- When the owner opens this folder in Codex, Claude Code, or another coding agent and asks for a specific support agent, treat that request as the product brief.
197
+ When the owner opens this folder in Codex, Claude Code, or another coding agent and asks for a specific support agent, treat that request as the product brief. The owner should not need to run a separate AgentKit wizard or prepare a brief file.
197
198
 
198
199
  Start building immediately:
199
200
 
@@ -203,11 +204,26 @@ Start building immediately:
203
204
  - Edit \`prompts/instructions.md\` for support behavior.
204
205
  - Edit \`agentkit.config.ts\` for provider, tools, secrets, access, and storage.
205
206
  - Add or replace TypeScript tools under \`tools/\` when the requested support agent needs actions or external data.
207
+ - When the support agent needs to save durable records, complete the full slice: schema/migration, tool, config registration, prompt instructions, direct tool check, and eval.
206
208
  - Add \`sync.ts\`, \`seed.sql\`, and ordered \`migrations/*.sql\` when the support agent depends on external catalogs or production-shaped data changes.
207
209
  - Do not wait for a wizard or recipe. AgentKit provides the scaffold and contract; you decide the implementation from the owner's brief.
208
210
  - Ask follow-up questions only when missing information blocks a safe local implementation.
209
211
  - State assumptions in the final response.
210
212
 
213
+ ## Proactive Agent Builder Contract
214
+
215
+ Do not only edit prompts. For every meaningful requirement in the owner's request or \`AGENT_SPEC.md\`, decide what should enforce it:
216
+
217
+ - spec entry for the product contract;
218
+ - prompt instruction for behavior, tone, boundaries, intake, and escalation;
219
+ - tool plus config registration for actions, live data, external writes, or authorization-sensitive data;
220
+ - schema/migration plus tool for durable records;
221
+ - eval for privacy, confirmation, required fields, date/time behavior, business rules, and regressions;
222
+ - fixture, seed data, fake branch, or direct tool check for integrations and failure paths;
223
+ - deploy/readiness check for hosted secrets, channels, integrations, or production access.
224
+
225
+ If a rule protects privacy, money, bookings, external writes, customer data, business hours, or safety, it must have an eval or deterministic check before you call the capsule done. If a real conversation exposes a bug, convert it into the smallest regression eval before or alongside the fix.
226
+
211
227
  ## Local Commands
212
228
 
213
229
  - \`npm install\`: restore capsule dependencies if this capsule used \`--no-install\`, install failed, or \`node_modules\` was deleted.
@@ -229,7 +245,7 @@ Start building immediately:
229
245
  - Local UI: run \`npm run dev\`, open the printed \`Chat:\` URL, and tell the owner the exact URL.
230
246
  - Hosted UI: after \`npm run agentkit -- deploy\`, run \`npm run agentkit -- chat-ui --deploy\`, open the printed \`Chat:\` URL, and tell the owner it is connected to the hosted deploy.
231
247
  - \`test/fake\` is deterministic. It is useful for scaffold checks, direct tool checks, and fake-provider evals, but it does not validate natural conversation quality.
232
- - Before claiming real conversation behavior is tested, ask the owner which provider to use: OpenRouter, OpenAI, Anthropic, or another supported provider. Do not choose for them.
248
+ - Before claiming real conversation behavior is tested, ask the owner which provider to use: OpenRouter, OpenAI, Anthropic, OpenCode Zen, OpenCode Go, or another supported provider. Do not choose for them.
233
249
 
234
250
  ## Hosted Deploy
235
251
 
@@ -262,7 +278,9 @@ This folder is an AgentKit support Agent Capsule.
262
278
 
263
279
  ## Start Here
264
280
 
265
- If the owner asks you to build an agent in natural language, that request is the brief. Do not ask them to fill another file first.
281
+ If the owner asks you to build an agent in natural language, that request is the brief. Do not ask them to run a wizard or fill another file first.
282
+
283
+ Build a testable capsule, not only a prompt.
266
284
 
267
285
  Example owner request:
268
286
 
@@ -274,7 +292,10 @@ Turn the request into a working local capsule:
274
292
  - Create or update \`AGENT_SPEC.md\` with \`npm run agentkit -- spec init --brief "<owner request>"\`. The owner gives the general idea; the coding agent turns it into the structured contract.
275
293
  - Update \`agentkit.config.ts\` when tools, secrets, provider, or access rules change.
276
294
  - Add, replace, or remove TypeScript tools under \`tools/\` for real actions or external data.
295
+ - For durable records, implement the full schema/tool/prompt/eval slice instead of only adding a table or only adding a tool.
277
296
  - Use \`npm run agentkit -- sync init\` when the agent needs catalog sync, fixture seed data, or ordered migrations.
297
+ - For every privacy, confirmation, required-intake, timezone, business-hour, integration-error, or no-leak rule, add an eval, fixture, fake branch, or direct tool check.
298
+ - Convert failed or surprising real conversations into regression evals with \`npm run agentkit -- eval from-conversation <conversation-id>\`.
278
299
  - Keep the first version runnable with \`test/fake\` unless the owner explicitly asks for a real provider.
279
300
  - Do not use a wizard or recipe. Build the capsule directly from the scaffold, the AgentKit contract, and the owner's brief.
280
301
  - Make practical assumptions and list them in your final response.
@@ -291,6 +312,8 @@ npm run eval
291
312
 
292
313
  \`test/fake\` proves the scaffold and deterministic tool paths. It does not prove natural conversation quality.
293
314
 
315
+ Before saying the agent is done, make sure important requirements have matching checks. Prompt-only changes are not enough for privacy, external writes, bookings, customer data, business hours, or integration failures.
316
+
294
317
  \`agentkit new\` installs dependencies by default. Run \`npm install\` only if the capsule was created with \`--no-install\`, install failed, or \`node_modules\` was deleted.
295
318
 
296
319
  Set local development secrets without opening code:
@@ -320,7 +343,7 @@ npm run agentkit -- chat-ui --deploy
320
343
 
321
344
  Open the printed \`Chat:\` URL and tell the owner this local UI is connected to the hosted deploy.
322
345
 
323
- Before claiming real conversation behavior has been tested, ask the owner which provider to use: OpenRouter, OpenAI, Anthropic, or another supported provider. Do not choose for them. After they choose, update \`agentkit.config.ts\`, \`.env.schema\`, local secrets, hosted secrets if deploying, then rerun chat/UI checks.
346
+ Before claiming real conversation behavior has been tested, ask the owner which provider to use: OpenRouter, OpenAI, Anthropic, OpenCode Zen, OpenCode Go, or another supported provider. Do not choose for them. After they choose, update \`agentkit.config.ts\`, \`.env.schema\`, local secrets, hosted secrets if deploying, then rerun chat/UI checks.
324
347
 
325
348
  If you add a tool, also run a fake-provider tool smoke test:
326
349
 
@@ -348,7 +371,7 @@ The recommended dual-storage pattern is:
348
371
  4. Use \`npm run agentkit -- db migrate\`, \`db reset --yes\`, \`db seed\`, and \`db shell\` for local database setup and inspection.
349
372
  5. Run \`npm run agentkit -- deploy\`. AgentKit migrates/provisions hosted storage internally.
350
373
 
351
- \`schema.sql\` is an idempotent bootstrap file in v1. Use \`CREATE TABLE IF NOT EXISTS\`, \`CREATE INDEX IF NOT EXISTS\`, and safe additive changes. AgentKit does not run destructive schema changes or ordered \`migrations/*.sql\` automatically yet.
374
+ \`schema.sql\` is an idempotent bootstrap file. Use \`CREATE TABLE IF NOT EXISTS\`, \`CREATE INDEX IF NOT EXISTS\`, and safe additive changes. Use ordered \`migrations/*.sql\` for production-shaped schema evolution; \`npm run agentkit -- db migrate\` applies unapplied local migrations before \`schema.sql\`.
352
375
 
353
376
  ## Hosted Deploy
354
377
 
@@ -374,7 +397,7 @@ Use AgentKit conventions when editing this support capsule.
374
397
  - The agent contract lives in \`agentkit.config.ts\`.
375
398
  - The example tool lives in \`tools/lookup-order.ts\`.
376
399
  - The default provider is \`test/fake\`, which can call tools from JSON messages during local tests.
377
- - Ask the owner which real provider to use before switching from \`test/fake\`; do not choose OpenRouter, OpenAI, or Anthropic automatically.
400
+ - Ask the owner which real provider to use before switching from \`test/fake\`; do not choose OpenRouter, OpenAI, Anthropic, OpenCode Zen, or OpenCode Go automatically.
378
401
  - Keep required local secret names in \`.env.schema\` and values in ignored \`.env\`. AgentKit local commands load \`.env\` directly.
379
402
  - Treat the owner's natural-language request as the brief and start implementing inside this capsule.
380
403
  - Start with \`skills/agentkit-capsule/SKILL.md\` when the task is not obvious.
@@ -398,7 +421,7 @@ npm run dev
398
421
  \`agentkit new\` installs dependencies by default. Run \`npm install\` only if this capsule was created with \`--no-install\`, install failed, or \`node_modules\` was deleted.
399
422
 
400
423
  The support template includes a local \`lookup_order\` TypeScript tool and uses \`test/fake\` by default.
401
- \`test/fake\` does not validate real conversation quality. The owner must choose OpenRouter, OpenAI, Anthropic, or another supported provider before real model behavior is tested.
424
+ \`test/fake\` does not validate real conversation quality. The owner must choose OpenRouter, OpenAI, Anthropic, OpenCode Zen, OpenCode Go, or another supported provider before real model behavior is tested.
402
425
 
403
426
  For UI testing, run \`npm run dev\` and open the printed \`Chat:\` URL. After hosted deploy, run \`npm run agentkit -- chat-ui --deploy\` and open its printed \`Chat:\` URL.
404
427
  `,