@genex-ai/cli-demo 1.34.9-dev.735 → 1.35.0-dev.736
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@genex-ai/cli-demo",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.35.0-dev.736",
|
|
4
4
|
"description": "Set up your project's agent workspace (.claude/.codex/.cursor in the game folder), authorize, create a game project, generate AI assets, and publish (genex CLI).",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
|
@@ -21,7 +21,7 @@
|
|
|
21
21
|
"start": "node src/index.ts",
|
|
22
22
|
"dev": "node --watch src/index.ts",
|
|
23
23
|
"typecheck": "tsc --noEmit",
|
|
24
|
-
"test": "node --test --test-concurrency=1 test/*.test.ts"
|
|
24
|
+
"test": "GENEX_NO_BROWSER=1 node --test --test-concurrency=1 test/*.test.ts"
|
|
25
25
|
},
|
|
26
26
|
"keywords": [
|
|
27
27
|
"cli",
|
|
@@ -156,6 +156,7 @@ vendored code from memory of another engine.
|
|
|
156
156
|
| sound effect, one looping music bed, or a short spoken line | `$genex-ai-sfx`, `$genex-ai-music`, or `$genex-ai-voice` |
|
|
157
157
|
| requested UI/HUD/menu/interface work, a visible UI problem, or an interface you decided this game wants built with generated art | `$genex-threejs-game-ui` |
|
|
158
158
|
| selling anything for platform coin: a shop, an item catalog, boosts, cosmetics, "make it earn"; also any request for a loot box, gacha, wager, casino mechanic or donation prompt, which that skill refuses and replaces | `$genex-monetization` |
|
|
159
|
+
| the game calls a language model AT RUNTIME on the player's money: NPCs that answer in their own words, dialogue or quests written per save, a judge reading what the player typed — one-time calls, or a standing budget the player approves once. Check the lane with `npx genex llm models` before designing it in | `$genex-llm-in-games` |
|
|
159
160
|
| cinematic menu/title/pause/victory/defeat/lobby/credits video treatment | `$genex-ai-menu` |
|
|
160
161
|
| drawn HUD chrome the game's style wants—one element or a matched set of frames, masks, and icons | `$genex-ai-hud` |
|
|
161
162
|
| the game works but feels flat, floaty, or unresponsive: input response, camera, impacts, cooldowns, difficulty, fail/retry | `$genex-threejs-game-feel` |
|
|
@@ -0,0 +1,271 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: genex-llm-in-games
|
|
3
|
+
description: Call a language model from inside a running game — an NPC that answers in its own words, a quest written for this save, a judge that reads what the player typed. The PLAYER pays and approves, on a Genex surface the game cannot forge. Covers the two modes (a popup per call, or one standing budget then many silent calls), benchmarking the price before declaring it, the receiver pattern, and honest handling of every refusal.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Genex LLM in Games
|
|
7
|
+
|
|
8
|
+
A Genex game can call a language model **while the player is playing** and get
|
|
9
|
+
back text or JSON. Nothing else: no images, no code execution, no tools.
|
|
10
|
+
|
|
11
|
+
**The player pays, and the player approves.** Funding is coin from their Genex
|
|
12
|
+
wallet or their own Claude / ChatGPT plan, chosen on a Genex-drawn surface your
|
|
13
|
+
game cannot render, skin or bypass. The game holds no provider key, sees no
|
|
14
|
+
credential, and never talks to a model vendor.
|
|
15
|
+
|
|
16
|
+
Two modes. Picking the wrong one is the most expensive mistake on this lane:
|
|
17
|
+
|
|
18
|
+
| | One-time | Standing budget |
|
|
19
|
+
| --- | --- | --- |
|
|
20
|
+
| Shape | `generate()` — one approval popup per call | `requestSpendGrant()` once, then `generate({ grantId })` many times, no popup |
|
|
21
|
+
| Fits | a rare, deliberate moment the player asked for | a loop — NPCs thinking, a director reacting, anything per-wave or per-minute |
|
|
22
|
+
| Ends | when that call settles | at the player's limit, a Stop, or 24h |
|
|
23
|
+
| Gesture | must run inside a click handler | none needed once the grant is active |
|
|
24
|
+
|
|
25
|
+
A popup per NPC turn is not a feature, it is an interruption. More than a call
|
|
26
|
+
or two per session means a grant.
|
|
27
|
+
|
|
28
|
+
## Step 0 — is the lane live on this stand?
|
|
29
|
+
|
|
30
|
+
```bash
|
|
31
|
+
npx genex llm models
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
- **live** — it prints the models this stand serves. Build the feature.
|
|
35
|
+
- **off on this stand** — the routes answer 404. Build the feature behind a
|
|
36
|
+
graceful `unavailable` state (the NPC uses its authored lines, the quest falls
|
|
37
|
+
back to the written one) and **say so plainly in the handoff**. Never promise
|
|
38
|
+
the player something that 404s.
|
|
39
|
+
- **misconfigured** — say that too; it is an operator fix, not a game bug.
|
|
40
|
+
|
|
41
|
+
Model ids come from that command and from `getGenerationModels()` at runtime.
|
|
42
|
+
Never write one into the game's source: they differ per stand, and a hardcoded
|
|
43
|
+
id is a feature that dies on somebody else's environment.
|
|
44
|
+
|
|
45
|
+
## The SDK surface (exact — do not invent methods)
|
|
46
|
+
|
|
47
|
+
From `@genex-ai/embed-sdk`, already installed. `initEmbed()` must have run and
|
|
48
|
+
identity must be resolved first — `$genex-threejs-embed-auth`.
|
|
49
|
+
|
|
50
|
+
- `getGenerationModels()` — the models this stand serves. Any picker renders
|
|
51
|
+
from this, never from a list you wrote.
|
|
52
|
+
- `generate({ modelId, prompt, outputFormat, schema?, estimateCoins,
|
|
53
|
+
allowExternal?, idempotencyKey?, grantId?, timeoutMs? })` →
|
|
54
|
+
`{ status, generationId, output, source, error }` plus billing fields
|
|
55
|
+
(`billingStatus`, `reservedCoins`, `chargedCoins`, their display-USD twins).
|
|
56
|
+
- `requestSpendGrant({ models, perCallMaxCoins, perCallEstimateCoins,
|
|
57
|
+
disclosure: { periodLabel, estimatedCallsPerPeriod, estimatedCoinsPerPeriod },
|
|
58
|
+
maxConcurrent?, maxCallsPerMinute?, allowExternal?, idempotencyKey? })` →
|
|
59
|
+
`{ status, grantId, … the limits the player approved }`; status is `active` |
|
|
60
|
+
`canceled` | `expired` | `failed` | `pending`.
|
|
61
|
+
- `getSpendGrant(grantId)` — live state and counters; the ONE source for an
|
|
62
|
+
in-game budget readout.
|
|
63
|
+
- `stopSpendGrant(grantId)` — the game's own stop door. Prospective: no further
|
|
64
|
+
calls are admitted, anything in flight drains and settles.
|
|
65
|
+
- `waitForGeneration(id)` / `getGeneration(id)` — re-attach to a call already
|
|
66
|
+
started, including after a reload.
|
|
67
|
+
- `generationErrorMessage(code)` — one player-facing sentence for an error code.
|
|
68
|
+
|
|
69
|
+
```ts
|
|
70
|
+
askButton.addEventListener('click', async () => { // a real click
|
|
71
|
+
const res = await generate({ // FIRST statement, no await before it
|
|
72
|
+
modelId, outputFormat: 'json', schema: ANSWER_SCHEMA,
|
|
73
|
+
prompt: askPrompt(npc, playerLine),
|
|
74
|
+
estimateCoins: NPC_CALL_PRICE, // benchmarked — see below
|
|
75
|
+
idempotencyKey: `npc:${npc.id}:${turnId}`,
|
|
76
|
+
});
|
|
77
|
+
applyGeneration(res); // the one writer — see below
|
|
78
|
+
});
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
**`generate()` and `requestSpendGrant()` are the first statement of the click
|
|
82
|
+
handler, before any `await`.** The approval popup is reserved synchronously off
|
|
83
|
+
the gesture; an `await` in front of it loses the gesture and nothing opens.
|
|
84
|
+
`generate({ grantId })` needs no gesture at all — that is what a grant buys.
|
|
85
|
+
|
|
86
|
+
## `estimateCoins` is a price, not an estimate
|
|
87
|
+
|
|
88
|
+
You declare it; the platform charges it. Declare 5 and 5 is charged — on a
|
|
89
|
+
success, a failure, a cancel, and when the model stops at its budget. Only an
|
|
90
|
+
attempt with no model work at all costs nothing. A number picked by feel is
|
|
91
|
+
money taken from your players for nothing, or a call that cannot fund itself.
|
|
92
|
+
|
|
93
|
+
**Benchmark, then declare:**
|
|
94
|
+
|
|
95
|
+
```bash
|
|
96
|
+
npx genex llm bench "<the real prompt, with a real example filled in>" \
|
|
97
|
+
--schema ./answer.schema.json --samples 3 --max-coins <n> --user-approved
|
|
98
|
+
```
|
|
99
|
+
|
|
100
|
+
It runs on **your own coins**, on the development lane, and prints what each
|
|
101
|
+
sample actually charged plus the recommendation to declare: p95 of the charged
|
|
102
|
+
coins with the server's own recommended headroom already applied. Declare that
|
|
103
|
+
printed number. Never guess it, never work it out from a vendor's price list,
|
|
104
|
+
never add a margin of your own. For a standing budget the run prints a second
|
|
105
|
+
line, `Grant perCallMaxCoins`, and that one is `perCallMaxCoins` — declare it
|
|
106
|
+
verbatim as well rather than deriving a ceiling from `max`, which lands under
|
|
107
|
+
the price and makes `requestSpendGrant()` refuse before it reaches the network. Full procedure — reading p50/p95, turning the
|
|
108
|
+
loop into disclosure numbers, re-benchmarking after a prompt change — is in
|
|
109
|
+
[references/pricing.md](references/pricing.md).
|
|
110
|
+
|
|
111
|
+
## Standing budgets
|
|
112
|
+
|
|
113
|
+
```ts
|
|
114
|
+
const grant = await requestSpendGrant({ // inside the click handler
|
|
115
|
+
models: [modelId],
|
|
116
|
+
perCallMaxCoins: NPC_CALL_CEILING,
|
|
117
|
+
perCallEstimateCoins: NPC_CALL_PRICE,
|
|
118
|
+
disclosure: {
|
|
119
|
+
periodLabel: 'minute',
|
|
120
|
+
estimatedCallsPerPeriod: 10, // 5 NPCs, one decision each per 30s
|
|
121
|
+
estimatedCoinsPerPeriod: 10 * NPC_CALL_PRICE,
|
|
122
|
+
},
|
|
123
|
+
maxConcurrent: 2,
|
|
124
|
+
maxCallsPerMinute: 30,
|
|
125
|
+
});
|
|
126
|
+
if (grant.status !== 'active') { runWithAuthoredLines(); return; }
|
|
127
|
+
await savePlayerState({ ...state, grantId: grant.grantId });
|
|
128
|
+
```
|
|
129
|
+
|
|
130
|
+
**The disclosure is computed from this game's own loop, never wished for.** Five
|
|
131
|
+
NPCs deciding once every thirty seconds is ten calls a minute — write that
|
|
132
|
+
arithmetic into `DESIGN.md` beside the feature. The player sees your estimate
|
|
133
|
+
attributed to the game, beside the platform's own worst case; an estimate that
|
|
134
|
+
is transparently low is a grant that dies mid-session.
|
|
135
|
+
|
|
136
|
+
**Then keep the burn low, because you wrote the loop:** batch those five NPCs
|
|
137
|
+
into ONE call returning five decisions, cache a decision until the situation
|
|
138
|
+
that caused it changes, pick the cheapest model that passes your own check, and
|
|
139
|
+
never fire on a timer the player cannot see.
|
|
140
|
+
|
|
141
|
+
Grant endings are ordinary game states with in-fiction copy, never an error toast:
|
|
142
|
+
|
|
143
|
+
| code | what happened | what the game does |
|
|
144
|
+
| --- | --- | --- |
|
|
145
|
+
| `grant_limit_reached` | the approved limit is spent | authored behaviour returns; a button offers to re-request |
|
|
146
|
+
| `grant_stopped` | the player pressed Stop | accept silently, keep playing |
|
|
147
|
+
| `grant_expired` | 24h passed, or the session ended | as stopped; re-request on the next deliberate click |
|
|
148
|
+
| `waiting_for_plan` | their own plan is rate-limited | wait out the stated time — not a failure, and there is no paid fallback |
|
|
149
|
+
| `grant_insufficient_funds` | the wallet cannot fund the next call | pause the thinking NPCs, say it once, stay playable |
|
|
150
|
+
|
|
151
|
+
Draw the readout from `getSpendGrant(grantId)` — calls made, coins settled, what
|
|
152
|
+
remains — never from a counter the game keeps itself. A finished grant may be
|
|
153
|
+
re-requested, but only from a **fresh deliberate click**: a silent auto-renew is
|
|
154
|
+
the exact shape a standing approval exists to prevent.
|
|
155
|
+
|
|
156
|
+
## The receiver pattern — one writer, two entry points
|
|
157
|
+
|
|
158
|
+
A generation outlives the frame that asked for it; reloads and closed tabs land
|
|
159
|
+
in the middle of one.
|
|
160
|
+
|
|
161
|
+
```ts
|
|
162
|
+
function applyGeneration(res) { // THE only place output becomes game state
|
|
163
|
+
if (res.status !== 'succeeded') return showLine(generationErrorMessage(res.error));
|
|
164
|
+
const parsed = ANSWER.safeParse(res.output); // validated against YOUR expectation
|
|
165
|
+
if (!parsed.success) return showLine("The voice trails off.");
|
|
166
|
+
speak(parsed.data.line);
|
|
167
|
+
savePlayerState({ ...state, pendingGenerationId: null });
|
|
168
|
+
}
|
|
169
|
+
```
|
|
170
|
+
|
|
171
|
+
- The click path writes `generationId` into player state **before** awaiting.
|
|
172
|
+
- Boot reads any stored id, calls `waitForGeneration(savedId)`, and passes the
|
|
173
|
+
result to the **same** `applyGeneration`. One writer, two entry points, is the
|
|
174
|
+
difference between "works once" and "survives a reload".
|
|
175
|
+
- Output is **data, never authority**: it may not grant coin, items,
|
|
176
|
+
entitlements, scores or progression by saying so. `source: 'external'` carries
|
|
177
|
+
`modelProvenance: 'unverified'` because it is user-supplied — check it exactly
|
|
178
|
+
as you would check typed player input.
|
|
179
|
+
|
|
180
|
+
## Errors land on the player's wallet
|
|
181
|
+
|
|
182
|
+
There is no compensation lane, so this is all work you do before the call:
|
|
183
|
+
|
|
184
|
+
- **Validate inputs first** — a malformed prompt is still charged.
|
|
185
|
+
- **Always set `schema` for `outputFormat: 'json'`** — unschema'd JSON is the
|
|
186
|
+
commonest way a call is charged and the result is unusable.
|
|
187
|
+
- **Keep prompts short.** Long context is the price.
|
|
188
|
+
- **Never loop `generate()` without a grant**, and never retry in a loop — each
|
|
189
|
+
attempt is a separate charge.
|
|
190
|
+
- **Map every code through `generationErrorMessage(code)`** into in-fiction
|
|
191
|
+
copy. A player should never read a raw error code inside your game.
|
|
192
|
+
- **`status: 'unknown'` is not a failure.** It means the charge is not known
|
|
193
|
+
yet, billing pending. Say "still settling", keep the reserved figure in the
|
|
194
|
+
readout, re-read with `getGeneration(id)` — never call it failed, never retry.
|
|
195
|
+
|
|
196
|
+
## What the Genex side already does — do not rebuild it
|
|
197
|
+
|
|
198
|
+
The approval sheet shows the model, the prompt, the price and the terms; the
|
|
199
|
+
game renders no price sheet. **Subscription funding is chosen only there** —
|
|
200
|
+
never add a "Your plan" row to the game's model picker, because a game cannot
|
|
201
|
+
offer a funding source. When the player's own watcher is online the personal-plan
|
|
202
|
+
answer arrives by itself and the game just waits, exactly as it waits for a
|
|
203
|
+
coin-funded call. The Genex dashboard header shows progress, active grants with
|
|
204
|
+
their spend, and a Stop; a Stop pressed there reaches the game as `grant_stopped`.
|
|
205
|
+
|
|
206
|
+
## Never
|
|
207
|
+
|
|
208
|
+
- **Never execute returned output** — no `eval`, no dynamic import, no scene
|
|
209
|
+
graph or shader built from model text, no URL fetched because the output said so.
|
|
210
|
+
- **Never bundle a creator credential in a game.** Benchmarking is a CLI action
|
|
211
|
+
on your machine, never something a shipped build does.
|
|
212
|
+
- **Never hand-roll fetch to the runtime API** — the SDK owns the approval
|
|
213
|
+
handshake, and a hand-rolled call cannot obtain one.
|
|
214
|
+
- **Never hardcode a price, a model id, a stand URL, or a margin.**
|
|
215
|
+
- **Never let the model be an authority over money, items or rewards**
|
|
216
|
+
(`$genex-monetization` owns what may move a wallet).
|
|
217
|
+
|
|
218
|
+
These are source contracts, not a claim that every stand runs this lane —
|
|
219
|
+
`npx genex llm models` is what tells you.
|
|
220
|
+
|
|
221
|
+
## Checklist
|
|
222
|
+
|
|
223
|
+
- [ ] `npx genex llm models` was run and its verdict is in the handoff
|
|
224
|
+
- [ ] Model ids come from `getGenerationModels()`, never from source
|
|
225
|
+
- [ ] `estimateCoins` came from `npx genex llm bench`, not from judgement
|
|
226
|
+
- [ ] `generate()` / `requestSpendGrant()` is the first statement of a click handler
|
|
227
|
+
- [ ] A repeated-call feature uses a grant; a one-off uses `generate()`
|
|
228
|
+
- [ ] Disclosure numbers derive from the real loop and are written in `DESIGN.md`
|
|
229
|
+
- [ ] Calls are batched and cached; nothing fires on an invisible timer
|
|
230
|
+
- [ ] Every grant-ending code has in-fiction copy and a playable fallback
|
|
231
|
+
- [ ] The in-game readout comes from `getSpendGrant()`
|
|
232
|
+
- [ ] Exactly one `applyGeneration()` writer; boot re-attaches with `waitForGeneration()`
|
|
233
|
+
- [ ] `generationId` is saved BEFORE the await
|
|
234
|
+
- [ ] Output is schema-validated and grants nothing by itself
|
|
235
|
+
- [ ] `unknown` reads as "still settling", never as a failure
|
|
236
|
+
- [ ] The game renders no price sheet and no funding picker
|
|
237
|
+
|
|
238
|
+
## Troubleshooting
|
|
239
|
+
|
|
240
|
+
**Everything on this lane 404s** — runtime generation is off on this stand.
|
|
241
|
+
Nothing to fix in the game: ship the fallback and say so.
|
|
242
|
+
|
|
243
|
+
**Nothing opens when the player clicks** — an `await` ran before `generate()`
|
|
244
|
+
and the gesture was lost. Move the call to the first line of the handler.
|
|
245
|
+
|
|
246
|
+
**`guest_no_wallet`** — guests play but hold no wallet. Show the feature as
|
|
247
|
+
sign-in-to-use rather than hiding it (`$genex-threejs-embed-auth`).
|
|
248
|
+
|
|
249
|
+
**`staging_no_generation`** — a `genex preview` build cannot spend real coin.
|
|
250
|
+
Check the layout on the draft; check the call itself after `genex promote`.
|
|
251
|
+
|
|
252
|
+
**`grant_price_unreasonable`** — the declared per-call price is far above what
|
|
253
|
+
that prompt can cost on that model. Re-benchmark and declare what it prints.
|
|
254
|
+
|
|
255
|
+
**`grant_concurrency` / `grant_rate_limited`** — the game calls faster than the
|
|
256
|
+
grant's own limits. Batch and cache; do not raise the limits to hide it.
|
|
257
|
+
|
|
258
|
+
**`grant_not_active`** — the saved `grantId` is finished. Clear the stored id
|
|
259
|
+
and re-request from a fresh click.
|
|
260
|
+
|
|
261
|
+
**`external_request_active`** — that player already has one personal-plan
|
|
262
|
+
request running. Wait for it; never fall back to charging coin instead.
|
|
263
|
+
|
|
264
|
+
**The call is charged but the result is unusable** — `outputFormat: 'json'`
|
|
265
|
+
without a `schema`. Add one; the charge already happened.
|
|
266
|
+
|
|
267
|
+
**A reload lost the answer** — `generationId` was not saved before the await, or
|
|
268
|
+
boot never calls `waitForGeneration()`. Both halves are required.
|
|
269
|
+
|
|
270
|
+
**The in-game readout disagrees with the Genex header** — the game is counting
|
|
271
|
+
calls itself. Read `getSpendGrant()` instead.
|
|
@@ -0,0 +1,153 @@
|
|
|
1
|
+
# Calibrate, then declare
|
|
2
|
+
|
|
3
|
+
`estimateCoins` is the **fixed price of a started attempt**, not a guess about
|
|
4
|
+
one. The platform charges exactly what you declare, whether the call succeeds,
|
|
5
|
+
fails, is canceled, or stops at its own budget. There is one honest way to pick
|
|
6
|
+
it: run the real prompt on your own coins, read what it charged, and declare the
|
|
7
|
+
number the benchmark recommends.
|
|
8
|
+
|
|
9
|
+
Never derive it from a model vendor's published rates. Three layers sit between
|
|
10
|
+
that rate and what the player is charged — the provider's cost, the platform's
|
|
11
|
+
tariff, and the headroom a declared price needs — and only the benchmark sees
|
|
12
|
+
all three. A number worked out from the vendor's page silently drops the middle
|
|
13
|
+
layer and under-prices every call you will ever make.
|
|
14
|
+
|
|
15
|
+
## 1. Freeze the prompt first
|
|
16
|
+
|
|
17
|
+
Benchmark the prompt you are actually shipping, with a realistic example filled
|
|
18
|
+
in: the longest NPC memory you will pass, the fullest world snapshot, a player
|
|
19
|
+
line of the length people really type. Short test prompts produce a cheap number
|
|
20
|
+
that the real game then cannot fund.
|
|
21
|
+
|
|
22
|
+
If the feature returns structured data, write the schema to a file now
|
|
23
|
+
(`./answer.schema.json`) and benchmark with it. Schema'd JSON and free text do
|
|
24
|
+
not cost the same, and shipping without a schema is the commonest way a call is
|
|
25
|
+
charged for an unusable result.
|
|
26
|
+
|
|
27
|
+
## 2. Run the benchmark
|
|
28
|
+
|
|
29
|
+
```bash
|
|
30
|
+
npx genex llm models # which models this stand serves; pick candidates
|
|
31
|
+
|
|
32
|
+
npx genex llm bench "<the frozen prompt, one real example filled in>" \
|
|
33
|
+
--model <id from the line above> \
|
|
34
|
+
--schema ./answer.schema.json \
|
|
35
|
+
--samples 3 \
|
|
36
|
+
--max-coins <n> --user-approved
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
- It spends **your** coins, on the development lane, through the CLI's own
|
|
40
|
+
credential. Nothing here runs in a shipped build.
|
|
41
|
+
- `--max-coins <n> --user-approved` is a hard gate, refused before any network
|
|
42
|
+
call. That is the same shape as every other spend approval in the CLI: you
|
|
43
|
+
state the ceiling for this run, out loud, once.
|
|
44
|
+
- `--samples 3` is the floor. The same prompt costs different amounts on
|
|
45
|
+
different runs, because the model's own output length varies.
|
|
46
|
+
- `npx genex llm price` re-prints the last run's recommendation without spending
|
|
47
|
+
anything again.
|
|
48
|
+
|
|
49
|
+
## 3. Read the output
|
|
50
|
+
|
|
51
|
+
Each sample prints what it actually charged. The aggregate prints p50, p95 and
|
|
52
|
+
max of the charged coins, plus one recommendation.
|
|
53
|
+
|
|
54
|
+
- **p50** is what a typical call costs. It is the number to reason about when
|
|
55
|
+
you ask "can the game afford this loop?" — multiply it by the calls per
|
|
56
|
+
minute you are about to disclose.
|
|
57
|
+
- **p95** is what a bad-but-normal call costs: a long answer, a model that
|
|
58
|
+
reasons its way around. It is the number to **declare**, because a declared
|
|
59
|
+
price below it means the unlucky calls cannot fund themselves and get refused
|
|
60
|
+
mid-session.
|
|
61
|
+
- **max** is diagnostic. When max sits far above p95, the prompt has an
|
|
62
|
+
unbounded branch in it — usually an unconstrained list or a missing schema.
|
|
63
|
+
Fix the prompt rather than declaring a bigger number.
|
|
64
|
+
|
|
65
|
+
The recommendation line already applies the **server's own recommended
|
|
66
|
+
headroom** on top of p95. Declare that figure verbatim:
|
|
67
|
+
|
|
68
|
+
```ts
|
|
69
|
+
const NPC_CALL_PRICE = <the recommended figure>; // from `npx genex llm bench`, <date>
|
|
70
|
+
const NPC_CALL_CEILING = <the printed ceiling>; // same run — grants only
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
Write the benchmark date and the model id beside it in `DESIGN.md`, so the next
|
|
74
|
+
person knows what the number describes. Do not add a margin of your own on top
|
|
75
|
+
of the recommendation, and do not round it down to look cheaper.
|
|
76
|
+
|
|
77
|
+
For a grant, the run prints a **second** number beside the price: the ceiling,
|
|
78
|
+
`perCallMaxCoins`. Declare that one verbatim too, and do not work it out by
|
|
79
|
+
hand. It is emphatically NOT the benchmark's `max`: p95 is nearest-rank, so at
|
|
80
|
+
the sample counts a benchmark actually takes, p95 and max are the same figure —
|
|
81
|
+
a ceiling set from `max` therefore lands *below* the recommended price, and
|
|
82
|
+
`requestSpendGrant()` refuses that pair before the request ever leaves the page.
|
|
83
|
+
|
|
84
|
+
The invariant, which the SDK and the server both enforce:
|
|
85
|
+
|
|
86
|
+
```
|
|
87
|
+
perCallEstimateCoins ≤ perCallMaxCoins ≤ the limit the player approves
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
The ceiling is the price's room to be wrong, not a second price: no call is ever
|
|
91
|
+
charged more than the price it declares, and the platform separately refuses any
|
|
92
|
+
declared price out of proportion to what the model could really cost.
|
|
93
|
+
|
|
94
|
+
## 4. Turn the game's loop into the disclosure
|
|
95
|
+
|
|
96
|
+
A standing budget asks the player to approve a rate, so the numbers have to come
|
|
97
|
+
from the loop you wrote, counted honestly:
|
|
98
|
+
|
|
99
|
+
1. **Count the callers.** How many things call the model at once? Five thinking
|
|
100
|
+
NPCs, one director, one narrator.
|
|
101
|
+
2. **Count each one's cadence.** How often does each decide? Once every thirty
|
|
102
|
+
seconds of play.
|
|
103
|
+
3. **Multiply, and pick the period that makes the number legible.** Five NPCs at
|
|
104
|
+
one call per thirty seconds is ten calls a minute, so `periodLabel: 'minute'`
|
|
105
|
+
and `estimatedCallsPerPeriod: 10`.
|
|
106
|
+
4. **Coins per period is calls × the declared price** —
|
|
107
|
+
`estimatedCoinsPerPeriod: 10 * NPC_CALL_PRICE`. Not the p50, not a hope: the
|
|
108
|
+
price you declare is the price charged.
|
|
109
|
+
|
|
110
|
+
```ts
|
|
111
|
+
disclosure: {
|
|
112
|
+
periodLabel: 'minute',
|
|
113
|
+
estimatedCallsPerPeriod: 10,
|
|
114
|
+
estimatedCoinsPerPeriod: 10 * NPC_CALL_PRICE,
|
|
115
|
+
}
|
|
116
|
+
```
|
|
117
|
+
|
|
118
|
+
The player sees this attributed to your game — "the game estimates about …" —
|
|
119
|
+
beside the platform's own worst case computed from `perCallMaxCoins`. The two
|
|
120
|
+
being far apart is normal; the estimate being far below what the game really
|
|
121
|
+
does is what breaks trust and burns the grant mid-session.
|
|
122
|
+
|
|
123
|
+
**Batching changes this arithmetic more than any price tuning can.** Five NPCs
|
|
124
|
+
answered by one call that returns five decisions is two calls a minute, not ten,
|
|
125
|
+
and one benchmarked price for the batched prompt replaces five of the unbatched
|
|
126
|
+
one. Do that before you reach for a cheaper model.
|
|
127
|
+
|
|
128
|
+
## Worked example
|
|
129
|
+
|
|
130
|
+
A tavern with five NPCs who react to what the player says. Each NPC decides once
|
|
131
|
+
every thirty seconds; a decision is a short JSON object (a mood, one line of
|
|
132
|
+
speech). The prompt carries the NPC's memory and the last two player lines.
|
|
133
|
+
|
|
134
|
+
1. Freeze the prompt with a full memory and a long player line. Write
|
|
135
|
+
`./answer.schema.json` with the two fields.
|
|
136
|
+
2. `npx genex llm bench "<that prompt>" --schema ./answer.schema.json
|
|
137
|
+
--samples 3 --max-coins <ceiling> --user-approved`.
|
|
138
|
+
3. Read: p50 `<p50>`, p95 `<p95>`, max `<max>`, and the two printed
|
|
139
|
+
declarations — `Declare estimateCoins: <recommended>` and
|
|
140
|
+
`Grant perCallMaxCoins: <ceiling>`.
|
|
141
|
+
4. Declare `NPC_CALL_PRICE = <recommended>` and
|
|
142
|
+
`NPC_CALL_CEILING = <ceiling>`, each verbatim from the line that printed it.
|
|
143
|
+
5. Unbatched, the loop is ten calls a minute, so the disclosure is
|
|
144
|
+
`10` and `10 * <recommended>` coins per minute.
|
|
145
|
+
6. Batch the five NPCs into one call — benchmark the batched prompt separately,
|
|
146
|
+
because it is a different prompt — and the disclosure becomes two calls a
|
|
147
|
+
minute at the batched price.
|
|
148
|
+
7. Re-run steps 1–4 whenever the prompt, the schema or the model changes. A
|
|
149
|
+
prompt edit is a price change; treat it like one.
|
|
150
|
+
|
|
151
|
+
Every `<placeholder>` above is read off your own benchmark run. None of these
|
|
152
|
+
figures is a platform constant, and none of them should be copied from another
|
|
153
|
+
game — a different prompt has a different price.
|