@genex-ai/cli-demo 1.35.1-dev.748 → 1.35.2-dev.749

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@genex-ai/cli-demo",
3
- "version": "1.35.1-dev.748",
3
+ "version": "1.35.2-dev.749",
4
4
  "description": "Set up your project's agent workspace (.claude/.codex/.cursor in the game folder), authorize, create a game project, generate AI assets, and publish (genex CLI).",
5
5
  "type": "module",
6
6
  "bin": {
@@ -36,6 +36,12 @@ budget, so a check every turn is affordable where a generate every turn is not.
36
36
  `generate()` for words, `judge()` for decisions; a feature often wants both —
37
37
  the judge picks the branch, generate writes the line.
38
38
 
39
+ **Words the player watches appear? Stream them.** Dialogue, narration, an NPC's
40
+ reply in a speech bubble: under a grant, `generateStream()` is the same call as
41
+ `generate({ grantId })`, but the text arrives while the model writes it instead
42
+ of all at once after it. Keep `generate()` for answers the game only uses whole
43
+ (a quest object, a list of decisions) — nobody watches those being written.
44
+
39
45
  ## Step 0 — is the lane live on this stand?
40
46
 
41
47
  ```bash
@@ -89,6 +95,13 @@ identity must be resolved first — `$genex-threejs-embed-auth`.
89
95
  (`{ instructions, criteria: [levels, lowest first] }`) or `noul`
90
96
  (`{ instructions }`); each answer has its question's type, and `noul` is a
91
97
  probability — pick your own threshold. Needs a budget requested with `judge`.
98
+ - `generateStream({ …the generate() options, grantId, onDelta?(text, soFar),
99
+ onPartialField?(name, partialText), signal? })` → the same result as
100
+ `generate()`. Grant-only. `onDelta` hands over each piece and all the text so
101
+ far; for `outputFormat: 'json'`, `onPartialField` reports the schema's FIRST
102
+ top-level string property as far as it is written — put the spoken line first
103
+ in the schema. `signal` stops the reading, never the call (it still finishes
104
+ and is charged).
92
105
  - `getSpendGrant(grantId)` — live state and counters; the ONE source for an
93
106
  in-game budget readout.
94
107
  - `stopSpendGrant(grantId)` — the game's own stop door. Prospective: no further
@@ -216,6 +229,27 @@ only, and judge calls share the budget's `maxConcurrent` and
216
229
  `maxCallsPerMinute`: batch questions into one call (up to 64) rather than one
217
230
  call per question.
218
231
 
232
+ **Streaming a reply under the budget:**
233
+
234
+ ```ts
235
+ const res = await generateStream({
236
+ grantId, modelId, outputFormat: 'json', schema: LINE_SCHEMA, // `line` is its first property
237
+ prompt: npcPrompt(npc, playerLine), estimateCoins: NPC_CALL_PRICE,
238
+ idempotencyKey: `npc:${npc.id}:${turnId}`,
239
+ onPartialField: (_name, soFar) => npc.bubble.show(soFar), // provisional
240
+ });
241
+ applyGeneration(res); // the answer is THIS
242
+ ```
243
+
244
+ The pieces are provisional: a call can still fail after text has shown (the
245
+ model broke off, the JSON missed the schema), so the bubble shows them and
246
+ `applyGeneration` decides what the game keeps — the same one writer as always.
247
+ Price, benchmark and receipt are `generate()`'s: a started stream is charged its
248
+ declared price even when the player walks away mid-sentence. A stream holds one
249
+ of the budget's `maxConcurrent` slots until it has settled. Sometimes the reply
250
+ arrives whole — a busy stand, a retried call, a budget on the player's own plan —
251
+ and `onDelta` then gets the whole text once; the game needs no second path.
252
+
219
253
  **Then keep the burn low, because you wrote the loop:** batch those five NPCs
220
254
  into ONE call returning five decisions, cache a decision until the situation
221
255
  that caused it changes, pick the cheapest model that passes your own check, and
@@ -314,6 +348,7 @@ These are source contracts, not a claim that every stand runs this lane —
314
348
  - [ ] `generate()` / `requestSpendGrant()` is the first statement of a click handler
315
349
  - [ ] A repeated-call feature uses a grant; a one-off uses `generate()`
316
350
  - [ ] A decision (a label, a yes/no, a level) is a `judge()` question under the grant, not a `generate()` asked for a word
351
+ - [ ] Dialogue the player watches appear uses `generateStream()` under the grant; the resolved result, not the pieces, goes to `applyGeneration()`
317
352
  - [ ] Disclosure numbers derive from the real loop and are written in `DESIGN.md`
318
353
  - [ ] Calls are batched and cached; nothing fires on an invisible timer
319
354
  - [ ] Every grant-ending code has in-fiction copy and a playable fallback
@@ -354,6 +389,10 @@ the period (your own `estimatedCallsPerPeriod` at the largest size): wait, and
354
389
  if it keeps happening your declared rate is too low — fix the loop or the
355
390
  estimate, never retry in a tight loop.
356
391
 
392
+ **`stream_requires_grant`** — `generateStream()` was called without a
393
+ `grantId`. Streaming is grant-only: a one-time call waits on the player's
394
+ approval popup, so use `generate()` for it.
395
+
357
396
  **`generation_limit`** — three ad-hoc calls are already in flight for this
358
397
  account, or recently stopped with their bill still pending; a pending call
359
398
  stops counting ten minutes after it was dispatched. From the bench: run