@genex-ai/cli-demo 1.34.9 → 1.35.0-dev.737
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/{blender-mcp-6TEAE65U.js → blender-mcp-HKT7N5BQ.js} +2 -2
- package/dist/{blender-serve-66OHC4ML.js → blender-serve-EH2MOW3U.js} +1 -1
- package/dist/{chunk-5WBSWMH5.js → chunk-FYHYZYBC.js} +1 -1
- package/dist/{chunk-QI7FIYBY.js → chunk-K4EHXOZK.js} +1 -1
- package/dist/index.js +2502 -128
- package/package.json +2 -2
- package/templates/skills/genex-game-director/SKILL.md +1 -0
- package/templates/skills/genex-getting-started/SKILL.md +2 -2
- package/templates/skills/genex-llm-in-games/SKILL.md +275 -0
- package/templates/skills/genex-llm-in-games/references/pricing.md +153 -0
- package/templates/skills/genex-tool-llm/SKILL.md +96 -0
- package/templates/skills/genex-tool-workflow/SKILL.md +10 -5
- package/templates/skills/genex-updates/SKILL.md +1 -1
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@genex-ai/cli-demo",
|
|
3
|
-
"version": "1.
|
|
3
|
+
"version": "1.35.0-dev.737",
|
|
4
4
|
"description": "Set up your project's agent workspace (.claude/.codex/.cursor in the game folder), authorize, create a game project, generate AI assets, and publish (genex CLI).",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
|
@@ -21,7 +21,7 @@
|
|
|
21
21
|
"start": "node src/index.ts",
|
|
22
22
|
"dev": "node --watch src/index.ts",
|
|
23
23
|
"typecheck": "tsc --noEmit",
|
|
24
|
-
"test": "node --test --test-concurrency=1 test/*.test.ts"
|
|
24
|
+
"test": "GENEX_NO_BROWSER=1 node --test --test-concurrency=1 test/*.test.ts"
|
|
25
25
|
},
|
|
26
26
|
"keywords": [
|
|
27
27
|
"cli",
|
|
@@ -156,6 +156,7 @@ vendored code from memory of another engine.
|
|
|
156
156
|
| sound effect, one looping music bed, or a short spoken line | `$genex-ai-sfx`, `$genex-ai-music`, or `$genex-ai-voice` |
|
|
157
157
|
| requested UI/HUD/menu/interface work, a visible UI problem, or an interface you decided this game wants built with generated art | `$genex-threejs-game-ui` |
|
|
158
158
|
| selling anything for platform coin: a shop, an item catalog, boosts, cosmetics, "make it earn"; also any request for a loot box, gacha, wager, casino mechanic or donation prompt, which that skill refuses and replaces | `$genex-monetization` |
|
|
159
|
+
| the game calls a language model AT RUNTIME on the player's money: NPCs that answer in their own words, dialogue or quests written per save, a judge reading what the player typed — one-time calls, or a standing budget the player approves once. Check the lane with `npx genex llm models` before designing it in | `$genex-llm-in-games` |
|
|
159
160
|
| cinematic menu/title/pause/victory/defeat/lobby/credits video treatment | `$genex-ai-menu` |
|
|
160
161
|
| drawn HUD chrome the game's style wants—one element or a matched set of frames, masks, and icons | `$genex-ai-hud` |
|
|
161
162
|
| the game works but feels flat, floaty, or unresponsive: input response, camera, impacts, cooldowns, difficulty, fail/retry | `$genex-threejs-game-feel` |
|
|
@@ -161,7 +161,7 @@ downloads the game too, binary assets and all:
|
|
|
161
161
|
|
|
162
162
|
```bash
|
|
163
163
|
mkdir my-game && cd my-game
|
|
164
|
-
npx @genex-ai/cli-demo@
|
|
164
|
+
npx @genex-ai/cli-demo@dev link <slug> # slug = the name in the play URL
|
|
165
165
|
npm install
|
|
166
166
|
```
|
|
167
167
|
|
|
@@ -225,7 +225,7 @@ Safe to run any time — genex-owned skills are refreshed to the latest version,
|
|
|
225
225
|
and your own files are never touched:
|
|
226
226
|
|
|
227
227
|
```bash
|
|
228
|
-
npx @genex-ai/cli-demo@
|
|
228
|
+
npx @genex-ai/cli-demo@dev init
|
|
229
229
|
```
|
|
230
230
|
|
|
231
231
|
Use `--force` only if you intentionally want your own existing files overwritten
|
|
@@ -0,0 +1,275 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: genex-llm-in-games
|
|
3
|
+
description: Call a language model from inside a running game — an NPC that answers in its own words, a quest written for this save, a judge that reads what the player typed. The PLAYER pays and approves, on a Genex surface the game cannot forge. Covers the two modes (a popup per call, or one standing budget then many silent calls), benchmarking the price before declaring it, the receiver pattern, and honest handling of every refusal.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Genex LLM in Games
|
|
7
|
+
|
|
8
|
+
A Genex game can call a language model **while the player is playing** and get
|
|
9
|
+
back text or JSON. Nothing else: no images, no code execution, no tools.
|
|
10
|
+
|
|
11
|
+
**The player pays, and the player approves.** Funding is coin from their Genex
|
|
12
|
+
wallet or their own Claude / ChatGPT plan, chosen on a Genex-drawn surface your
|
|
13
|
+
game cannot render, skin or bypass. The game holds no provider key, sees no
|
|
14
|
+
credential, and never talks to a model vendor.
|
|
15
|
+
|
|
16
|
+
Two modes. Picking the wrong one is the most expensive mistake on this lane:
|
|
17
|
+
|
|
18
|
+
| | One-time | Standing budget |
|
|
19
|
+
| --- | --- | --- |
|
|
20
|
+
| Shape | `generate()` — one approval popup per call | `requestSpendGrant()` once, then `generate({ grantId })` many times, no popup |
|
|
21
|
+
| Fits | a rare, deliberate moment the player asked for | a loop — NPCs thinking, a director reacting, anything per-wave or per-minute |
|
|
22
|
+
| Ends | when that call settles | at the player's limit, a Stop, or 24h |
|
|
23
|
+
| Gesture | must run inside a click handler | none needed once the grant is active |
|
|
24
|
+
|
|
25
|
+
A popup per NPC turn is not a feature, it is an interruption. More than a call
|
|
26
|
+
or two per session means a grant.
|
|
27
|
+
|
|
28
|
+
## Step 0 — is the lane live on this stand?
|
|
29
|
+
|
|
30
|
+
```bash
|
|
31
|
+
npx genex llm models
|
|
32
|
+
```
|
|
33
|
+
|
|
34
|
+
- **live** — it prints the models this stand serves. Build the feature.
|
|
35
|
+
- **off on this stand** — the routes answer 404. Build the feature behind a
|
|
36
|
+
graceful `unavailable` state (the NPC uses its authored lines, the quest falls
|
|
37
|
+
back to the written one) and **say so plainly in the handoff**. Never promise
|
|
38
|
+
the player something that 404s.
|
|
39
|
+
- **misconfigured** — say that too; it is an operator fix, not a game bug.
|
|
40
|
+
|
|
41
|
+
Model ids come from that command and from `getGenerationModels()` at runtime.
|
|
42
|
+
Never write one into the game's source: they differ per stand, and a hardcoded
|
|
43
|
+
id is a feature that dies on somebody else's environment.
|
|
44
|
+
|
|
45
|
+
## The SDK surface (exact — do not invent methods)
|
|
46
|
+
|
|
47
|
+
From `@genex-ai/embed-sdk`, already installed. `initEmbed()` must have run and
|
|
48
|
+
identity must be resolved first — `$genex-threejs-embed-auth`.
|
|
49
|
+
|
|
50
|
+
- `getGenerationModels()` — the models this stand serves. Any picker renders
|
|
51
|
+
from this, never from a list you wrote.
|
|
52
|
+
- `generate({ modelId, prompt, outputFormat, schema?, estimateCoins,
|
|
53
|
+
allowExternal?, idempotencyKey?, grantId?, timeoutMs? })` →
|
|
54
|
+
`{ status, generationId, output, source, error }` plus billing fields
|
|
55
|
+
(`billingStatus`, `reservedCoins`, `chargedCoins`, their display-USD twins).
|
|
56
|
+
- `requestSpendGrant({ models, perCallMaxCoins, perCallEstimateCoins,
|
|
57
|
+
disclosure: { periodLabel, estimatedCallsPerPeriod, estimatedCoinsPerPeriod },
|
|
58
|
+
maxConcurrent?, maxCallsPerMinute?, allowExternal?, idempotencyKey? })` →
|
|
59
|
+
`{ status, grantId, … the limits the player approved }`; status is `active` |
|
|
60
|
+
`canceled` | `expired` | `failed` | `pending`.
|
|
61
|
+
- `getSpendGrant(grantId)` — live state and counters; the ONE source for an
|
|
62
|
+
in-game budget readout.
|
|
63
|
+
- `stopSpendGrant(grantId)` — the game's own stop door. Prospective: no further
|
|
64
|
+
calls are admitted, anything in flight drains and settles.
|
|
65
|
+
- `waitForGeneration(id)` / `getGeneration(id)` — re-attach to a call already
|
|
66
|
+
started, including after a reload.
|
|
67
|
+
- `generationErrorMessage(code)` — one player-facing sentence for an error code.
|
|
68
|
+
|
|
69
|
+
```ts
|
|
70
|
+
askButton.addEventListener('click', async () => { // a real click
|
|
71
|
+
const res = await generate({ // FIRST statement, no await before it
|
|
72
|
+
modelId, outputFormat: 'json', schema: ANSWER_SCHEMA,
|
|
73
|
+
prompt: askPrompt(npc, playerLine),
|
|
74
|
+
estimateCoins: NPC_CALL_PRICE, // benchmarked — see below
|
|
75
|
+
idempotencyKey: `npc:${npc.id}:${turnId}`,
|
|
76
|
+
});
|
|
77
|
+
applyGeneration(res); // the one writer — see below
|
|
78
|
+
});
|
|
79
|
+
```
|
|
80
|
+
|
|
81
|
+
**`generate()` and `requestSpendGrant()` are the first statement of the click
|
|
82
|
+
handler, before any `await`.** The approval popup is reserved synchronously off
|
|
83
|
+
the gesture; an `await` in front of it loses the gesture and nothing opens.
|
|
84
|
+
`generate({ grantId })` needs no gesture at all — that is what a grant buys.
|
|
85
|
+
|
|
86
|
+
## `estimateCoins` is a price, not an estimate
|
|
87
|
+
|
|
88
|
+
You declare it; the platform charges it. Declare 5 and 5 is charged — on a
|
|
89
|
+
success, a failure, a cancel, and when the model stops at its budget. Only an
|
|
90
|
+
attempt with no model work at all costs nothing. A number picked by feel is
|
|
91
|
+
money taken from your players for nothing, or a call that cannot fund itself.
|
|
92
|
+
|
|
93
|
+
**Benchmark, then declare:**
|
|
94
|
+
|
|
95
|
+
```bash
|
|
96
|
+
npx genex llm bench "<the real prompt, with a real example filled in>" \
|
|
97
|
+
--schema ./answer.schema.json --samples 3 --max-coins <n> --user-approved
|
|
98
|
+
```
|
|
99
|
+
|
|
100
|
+
It runs on **your own coins**, on the development lane, and prints what each
|
|
101
|
+
sample actually charged plus the recommendation to declare: p95 of the charged
|
|
102
|
+
coins with the server's own recommended headroom already applied. Declare that
|
|
103
|
+
printed number. Never guess it, never work it out from a vendor's price list,
|
|
104
|
+
never add a margin of your own. For a standing budget the run prints a second
|
|
105
|
+
line, `Grant perCallMaxCoins`, and that one is `perCallMaxCoins` — declare it
|
|
106
|
+
verbatim as well rather than deriving a ceiling from `max`, which lands under
|
|
107
|
+
the price and makes `requestSpendGrant()` refuse before it reaches the network. Full procedure — reading p50/p95, turning the
|
|
108
|
+
loop into disclosure numbers, re-benchmarking after a prompt change — is in
|
|
109
|
+
[references/pricing.md](references/pricing.md).
|
|
110
|
+
|
|
111
|
+
## Standing budgets
|
|
112
|
+
|
|
113
|
+
```ts
|
|
114
|
+
const grant = await requestSpendGrant({ // inside the click handler
|
|
115
|
+
models: [modelId],
|
|
116
|
+
perCallMaxCoins: NPC_CALL_CEILING,
|
|
117
|
+
perCallEstimateCoins: NPC_CALL_PRICE,
|
|
118
|
+
disclosure: {
|
|
119
|
+
periodLabel: 'minute',
|
|
120
|
+
estimatedCallsPerPeriod: 10, // 5 NPCs, one decision each per 30s
|
|
121
|
+
estimatedCoinsPerPeriod: 10 * NPC_CALL_PRICE,
|
|
122
|
+
},
|
|
123
|
+
maxConcurrent: 2,
|
|
124
|
+
maxCallsPerMinute: 30,
|
|
125
|
+
});
|
|
126
|
+
if (grant.status !== 'active') { runWithAuthoredLines(); return; }
|
|
127
|
+
await savePlayerState({ ...state, grantId: grant.grantId });
|
|
128
|
+
```
|
|
129
|
+
|
|
130
|
+
**The disclosure is computed from this game's own loop, never wished for.** Five
|
|
131
|
+
NPCs deciding once every thirty seconds is ten calls a minute — write that
|
|
132
|
+
arithmetic into `DESIGN.md` beside the feature. The player sees your estimate
|
|
133
|
+
attributed to the game, beside the platform's own worst case; an estimate that
|
|
134
|
+
is transparently low is a grant that dies mid-session.
|
|
135
|
+
|
|
136
|
+
**Then keep the burn low, because you wrote the loop:** batch those five NPCs
|
|
137
|
+
into ONE call returning five decisions, cache a decision until the situation
|
|
138
|
+
that caused it changes, pick the cheapest model that passes your own check, and
|
|
139
|
+
never fire on a timer the player cannot see.
|
|
140
|
+
|
|
141
|
+
Grant endings are ordinary game states with in-fiction copy, never an error toast:
|
|
142
|
+
|
|
143
|
+
| code | what happened | what the game does |
|
|
144
|
+
| --- | --- | --- |
|
|
145
|
+
| `grant_limit_reached` | the approved limit is spent | authored behaviour returns; a button offers to re-request |
|
|
146
|
+
| `grant_stopped` | the player pressed Stop | accept silently, keep playing |
|
|
147
|
+
| `grant_expired` | 24h passed, or the session ended | as stopped; re-request on the next deliberate click |
|
|
148
|
+
| `waiting_for_plan` | their own plan is rate-limited | wait out the stated time — not a failure, and there is no paid fallback |
|
|
149
|
+
| `grant_insufficient_funds` | the wallet cannot fund the next call | pause the thinking NPCs, say it once, stay playable |
|
|
150
|
+
|
|
151
|
+
Draw the readout from `getSpendGrant(grantId)` — calls made, coins settled, what
|
|
152
|
+
remains — never from a counter the game keeps itself. A finished grant may be
|
|
153
|
+
re-requested, but only from a **fresh deliberate click**: a silent auto-renew is
|
|
154
|
+
the exact shape a standing approval exists to prevent.
|
|
155
|
+
|
|
156
|
+
## The receiver pattern — one writer, two entry points
|
|
157
|
+
|
|
158
|
+
A generation outlives the frame that asked for it; reloads and closed tabs land
|
|
159
|
+
in the middle of one.
|
|
160
|
+
|
|
161
|
+
```ts
|
|
162
|
+
function applyGeneration(res) { // THE only place output becomes game state
|
|
163
|
+
if (res.status !== 'succeeded') return showLine(generationErrorMessage(res.error));
|
|
164
|
+
const parsed = ANSWER.safeParse(res.output); // validated against YOUR expectation
|
|
165
|
+
if (!parsed.success) return showLine("The voice trails off.");
|
|
166
|
+
speak(parsed.data.line);
|
|
167
|
+
savePlayerState({ ...state, pendingGenerationId: null });
|
|
168
|
+
}
|
|
169
|
+
```
|
|
170
|
+
|
|
171
|
+
- The click path writes `generationId` into player state **before** awaiting.
|
|
172
|
+
- Boot reads any stored id, calls `waitForGeneration(savedId)`, and passes the
|
|
173
|
+
result to the **same** `applyGeneration`. One writer, two entry points, is the
|
|
174
|
+
difference between "works once" and "survives a reload".
|
|
175
|
+
- Output is **data, never authority**: it may not grant coin, items,
|
|
176
|
+
entitlements, scores or progression by saying so. `source: 'external'` carries
|
|
177
|
+
`modelProvenance: 'unverified'` because it is user-supplied — check it exactly
|
|
178
|
+
as you would check typed player input.
|
|
179
|
+
|
|
180
|
+
## Errors land on the player's wallet
|
|
181
|
+
|
|
182
|
+
There is no compensation lane, so this is all work you do before the call:
|
|
183
|
+
|
|
184
|
+
- **Validate inputs first** — a malformed prompt is still charged.
|
|
185
|
+
- **Always set `schema` for `outputFormat: 'json'`** — unschema'd JSON is the
|
|
186
|
+
commonest way a call is charged and the result is unusable.
|
|
187
|
+
- **Keep prompts short.** Long context is the price.
|
|
188
|
+
- **Never loop `generate()` without a grant**, and never retry in a loop — each
|
|
189
|
+
attempt is a separate charge.
|
|
190
|
+
- **Map every code through `generationErrorMessage(code)`** into in-fiction
|
|
191
|
+
copy. A player should never read a raw error code inside your game.
|
|
192
|
+
- **`status: 'unknown'` is not a failure.** It means the charge is not known
|
|
193
|
+
yet, billing pending. Say "still settling", keep the reserved figure in the
|
|
194
|
+
readout, re-read with `getGeneration(id)` — never call it failed, never retry.
|
|
195
|
+
|
|
196
|
+
## What the Genex side already does — do not rebuild it
|
|
197
|
+
|
|
198
|
+
The approval sheet shows the model, the prompt, the price and the terms; the
|
|
199
|
+
game renders no price sheet. **Subscription funding is chosen only there** —
|
|
200
|
+
never add a "Your plan" row to the game's model picker, because a game cannot
|
|
201
|
+
offer a funding source. When the player's own watcher is online the personal-plan
|
|
202
|
+
answer arrives by itself and the game just waits, exactly as it waits for a
|
|
203
|
+
coin-funded call. The Genex dashboard header shows progress, active grants with
|
|
204
|
+
their spend, and a Stop; a Stop pressed there reaches the game as `grant_stopped`.
|
|
205
|
+
|
|
206
|
+
## Never
|
|
207
|
+
|
|
208
|
+
- **Never execute returned output** — no `eval`, no dynamic import, no scene
|
|
209
|
+
graph or shader built from model text, no URL fetched because the output said so.
|
|
210
|
+
- **Never bundle a creator credential in a game.** Benchmarking is a CLI action
|
|
211
|
+
on your machine, never something a shipped build does.
|
|
212
|
+
- **Never hand-roll fetch to the runtime API** — the SDK owns the approval
|
|
213
|
+
handshake, and a hand-rolled call cannot obtain one.
|
|
214
|
+
- **Never hardcode a price, a model id, a stand URL, or a margin.**
|
|
215
|
+
- **Never let the model be an authority over money, items or rewards**
|
|
216
|
+
(`$genex-monetization` owns what may move a wallet).
|
|
217
|
+
|
|
218
|
+
These are source contracts, not a claim that every stand runs this lane —
|
|
219
|
+
`npx genex llm models` is what tells you.
|
|
220
|
+
|
|
221
|
+
## Checklist
|
|
222
|
+
|
|
223
|
+
- [ ] `npx genex llm models` was run and its verdict is in the handoff
|
|
224
|
+
- [ ] Model ids come from `getGenerationModels()`, never from source
|
|
225
|
+
- [ ] `estimateCoins` came from `npx genex llm bench`, not from judgement
|
|
226
|
+
- [ ] `generate()` / `requestSpendGrant()` is the first statement of a click handler
|
|
227
|
+
- [ ] A repeated-call feature uses a grant; a one-off uses `generate()`
|
|
228
|
+
- [ ] Disclosure numbers derive from the real loop and are written in `DESIGN.md`
|
|
229
|
+
- [ ] Calls are batched and cached; nothing fires on an invisible timer
|
|
230
|
+
- [ ] Every grant-ending code has in-fiction copy and a playable fallback
|
|
231
|
+
- [ ] The in-game readout comes from `getSpendGrant()`
|
|
232
|
+
- [ ] Exactly one `applyGeneration()` writer; boot re-attaches with `waitForGeneration()`
|
|
233
|
+
- [ ] `generationId` is saved BEFORE the await
|
|
234
|
+
- [ ] Output is schema-validated and grants nothing by itself
|
|
235
|
+
- [ ] `unknown` reads as "still settling", never as a failure
|
|
236
|
+
- [ ] The game renders no price sheet and no funding picker
|
|
237
|
+
|
|
238
|
+
## Troubleshooting
|
|
239
|
+
|
|
240
|
+
**Everything on this lane 404s** — runtime generation is off on this stand.
|
|
241
|
+
Nothing to fix in the game: ship the fallback and say so.
|
|
242
|
+
|
|
243
|
+
**Nothing opens when the player clicks** — an `await` ran before `generate()`
|
|
244
|
+
and the gesture was lost. Move the call to the first line of the handler.
|
|
245
|
+
|
|
246
|
+
**`player_wallet_required`** — ONE code for the two early dead ends: the lane
|
|
247
|
+
refuses a guest and a PREVIEW build on the same line. `waitForPlayer()` tells
|
|
248
|
+
them apart. `guest: true` — guests play but hold no wallet, so show the feature
|
|
249
|
+
as sign-in-to-use rather than hiding it (`$genex-threejs-embed-auth`). Signed
|
|
250
|
+
in and still refused — this is a `genex preview` draft, which never spends:
|
|
251
|
+
check the layout there, and the call itself only after `genex promote`. The
|
|
252
|
+
SDK's stock sentence for this code is "Sign in to Genex to use this", which is
|
|
253
|
+
right for the guest and wrong on a draft, so write the in-fiction line per
|
|
254
|
+
cause rather than showing it for both.
|
|
255
|
+
|
|
256
|
+
**`grant_price_unreasonable`** — the declared per-call price is far above what
|
|
257
|
+
that prompt can cost on that model. Re-benchmark and declare what it prints.
|
|
258
|
+
|
|
259
|
+
**`grant_concurrency` / `grant_rate_limited`** — the game calls faster than the
|
|
260
|
+
grant's own limits. Batch and cache; do not raise the limits to hide it.
|
|
261
|
+
|
|
262
|
+
**`grant_not_active`** — the saved `grantId` is finished. Clear the stored id
|
|
263
|
+
and re-request from a fresh click.
|
|
264
|
+
|
|
265
|
+
**`external_request_active`** — that player already has one personal-plan
|
|
266
|
+
request running. Wait for it; never fall back to charging coin instead.
|
|
267
|
+
|
|
268
|
+
**The call is charged but the result is unusable** — `outputFormat: 'json'`
|
|
269
|
+
without a `schema`. Add one; the charge already happened.
|
|
270
|
+
|
|
271
|
+
**A reload lost the answer** — `generationId` was not saved before the await, or
|
|
272
|
+
boot never calls `waitForGeneration()`. Both halves are required.
|
|
273
|
+
|
|
274
|
+
**The in-game readout disagrees with the Genex header** — the game is counting
|
|
275
|
+
calls itself. Read `getSpendGrant()` instead.
|
|
@@ -0,0 +1,153 @@
|
|
|
1
|
+
# Calibrate, then declare
|
|
2
|
+
|
|
3
|
+
`estimateCoins` is the **fixed price of a started attempt**, not a guess about
|
|
4
|
+
one. The platform charges exactly what you declare, whether the call succeeds,
|
|
5
|
+
fails, is canceled, or stops at its own budget. There is one honest way to pick
|
|
6
|
+
it: run the real prompt on your own coins, read what it charged, and declare the
|
|
7
|
+
number the benchmark recommends.
|
|
8
|
+
|
|
9
|
+
Never derive it from a model vendor's published rates. Three layers sit between
|
|
10
|
+
that rate and what the player is charged — the provider's cost, the platform's
|
|
11
|
+
tariff, and the headroom a declared price needs — and only the benchmark sees
|
|
12
|
+
all three. A number worked out from the vendor's page silently drops the middle
|
|
13
|
+
layer and under-prices every call you will ever make.
|
|
14
|
+
|
|
15
|
+
## 1. Freeze the prompt first
|
|
16
|
+
|
|
17
|
+
Benchmark the prompt you are actually shipping, with a realistic example filled
|
|
18
|
+
in: the longest NPC memory you will pass, the fullest world snapshot, a player
|
|
19
|
+
line of the length people really type. Short test prompts produce a cheap number
|
|
20
|
+
that the real game then cannot fund.
|
|
21
|
+
|
|
22
|
+
If the feature returns structured data, write the schema to a file now
|
|
23
|
+
(`./answer.schema.json`) and benchmark with it. Schema'd JSON and free text do
|
|
24
|
+
not cost the same, and shipping without a schema is the commonest way a call is
|
|
25
|
+
charged for an unusable result.
|
|
26
|
+
|
|
27
|
+
## 2. Run the benchmark
|
|
28
|
+
|
|
29
|
+
```bash
|
|
30
|
+
npx genex llm models # which models this stand serves; pick candidates
|
|
31
|
+
|
|
32
|
+
npx genex llm bench "<the frozen prompt, one real example filled in>" \
|
|
33
|
+
--model <id from the line above> \
|
|
34
|
+
--schema ./answer.schema.json \
|
|
35
|
+
--samples 3 \
|
|
36
|
+
--max-coins <n> --user-approved
|
|
37
|
+
```
|
|
38
|
+
|
|
39
|
+
- It spends **your** coins, on the development lane, through the CLI's own
|
|
40
|
+
credential. Nothing here runs in a shipped build.
|
|
41
|
+
- `--max-coins <n> --user-approved` is a hard gate, refused before any network
|
|
42
|
+
call. That is the same shape as every other spend approval in the CLI: you
|
|
43
|
+
state the ceiling for this run, out loud, once.
|
|
44
|
+
- `--samples 3` is the floor. The same prompt costs different amounts on
|
|
45
|
+
different runs, because the model's own output length varies.
|
|
46
|
+
- `npx genex llm price` re-prints the last run's recommendation without spending
|
|
47
|
+
anything again.
|
|
48
|
+
|
|
49
|
+
## 3. Read the output
|
|
50
|
+
|
|
51
|
+
Each sample prints what it actually charged. The aggregate prints p50, p95 and
|
|
52
|
+
max of the charged coins, plus one recommendation.
|
|
53
|
+
|
|
54
|
+
- **p50** is what a typical call costs. It is the number to reason about when
|
|
55
|
+
you ask "can the game afford this loop?" — multiply it by the calls per
|
|
56
|
+
minute you are about to disclose.
|
|
57
|
+
- **p95** is what a bad-but-normal call costs: a long answer, a model that
|
|
58
|
+
reasons its way around. It is the number to **declare**, because a declared
|
|
59
|
+
price below it means the unlucky calls cannot fund themselves and get refused
|
|
60
|
+
mid-session.
|
|
61
|
+
- **max** is diagnostic. When max sits far above p95, the prompt has an
|
|
62
|
+
unbounded branch in it — usually an unconstrained list or a missing schema.
|
|
63
|
+
Fix the prompt rather than declaring a bigger number.
|
|
64
|
+
|
|
65
|
+
The recommendation line already applies the **server's own recommended
|
|
66
|
+
headroom** on top of p95. Declare that figure verbatim:
|
|
67
|
+
|
|
68
|
+
```ts
|
|
69
|
+
const NPC_CALL_PRICE = <the recommended figure>; // from `npx genex llm bench`, <date>
|
|
70
|
+
const NPC_CALL_CEILING = <the printed ceiling>; // same run — grants only
|
|
71
|
+
```
|
|
72
|
+
|
|
73
|
+
Write the benchmark date and the model id beside it in `DESIGN.md`, so the next
|
|
74
|
+
person knows what the number describes. Do not add a margin of your own on top
|
|
75
|
+
of the recommendation, and do not round it down to look cheaper.
|
|
76
|
+
|
|
77
|
+
For a grant, the run prints a **second** number beside the price: the ceiling,
|
|
78
|
+
`perCallMaxCoins`. Declare that one verbatim too, and do not work it out by
|
|
79
|
+
hand. It is emphatically NOT the benchmark's `max`: p95 is nearest-rank, so at
|
|
80
|
+
the sample counts a benchmark actually takes, p95 and max are the same figure —
|
|
81
|
+
a ceiling set from `max` therefore lands *below* the recommended price, and
|
|
82
|
+
`requestSpendGrant()` refuses that pair before the request ever leaves the page.
|
|
83
|
+
|
|
84
|
+
The invariant, which the SDK and the server both enforce:
|
|
85
|
+
|
|
86
|
+
```
|
|
87
|
+
perCallEstimateCoins ≤ perCallMaxCoins ≤ the limit the player approves
|
|
88
|
+
```
|
|
89
|
+
|
|
90
|
+
The ceiling is the price's room to be wrong, not a second price: no call is ever
|
|
91
|
+
charged more than the price it declares, and the platform separately refuses any
|
|
92
|
+
declared price out of proportion to what the model could really cost.
|
|
93
|
+
|
|
94
|
+
## 4. Turn the game's loop into the disclosure
|
|
95
|
+
|
|
96
|
+
A standing budget asks the player to approve a rate, so the numbers have to come
|
|
97
|
+
from the loop you wrote, counted honestly:
|
|
98
|
+
|
|
99
|
+
1. **Count the callers.** How many things call the model at once? Five thinking
|
|
100
|
+
NPCs, one director, one narrator.
|
|
101
|
+
2. **Count each one's cadence.** How often does each decide? Once every thirty
|
|
102
|
+
seconds of play.
|
|
103
|
+
3. **Multiply, and pick the period that makes the number legible.** Five NPCs at
|
|
104
|
+
one call per thirty seconds is ten calls a minute, so `periodLabel: 'minute'`
|
|
105
|
+
and `estimatedCallsPerPeriod: 10`.
|
|
106
|
+
4. **Coins per period is calls × the declared price** —
|
|
107
|
+
`estimatedCoinsPerPeriod: 10 * NPC_CALL_PRICE`. Not the p50, not a hope: the
|
|
108
|
+
price you declare is the price charged.
|
|
109
|
+
|
|
110
|
+
```ts
|
|
111
|
+
disclosure: {
|
|
112
|
+
periodLabel: 'minute',
|
|
113
|
+
estimatedCallsPerPeriod: 10,
|
|
114
|
+
estimatedCoinsPerPeriod: 10 * NPC_CALL_PRICE,
|
|
115
|
+
}
|
|
116
|
+
```
|
|
117
|
+
|
|
118
|
+
The player sees this attributed to your game — "the game estimates about …" —
|
|
119
|
+
beside the platform's own worst case computed from `perCallMaxCoins`. The two
|
|
120
|
+
being far apart is normal; the estimate being far below what the game really
|
|
121
|
+
does is what breaks trust and burns the grant mid-session.
|
|
122
|
+
|
|
123
|
+
**Batching changes this arithmetic more than any price tuning can.** Five NPCs
|
|
124
|
+
answered by one call that returns five decisions is two calls a minute, not ten,
|
|
125
|
+
and one benchmarked price for the batched prompt replaces five of the unbatched
|
|
126
|
+
one. Do that before you reach for a cheaper model.
|
|
127
|
+
|
|
128
|
+
## Worked example
|
|
129
|
+
|
|
130
|
+
A tavern with five NPCs who react to what the player says. Each NPC decides once
|
|
131
|
+
every thirty seconds; a decision is a short JSON object (a mood, one line of
|
|
132
|
+
speech). The prompt carries the NPC's memory and the last two player lines.
|
|
133
|
+
|
|
134
|
+
1. Freeze the prompt with a full memory and a long player line. Write
|
|
135
|
+
`./answer.schema.json` with the two fields.
|
|
136
|
+
2. `npx genex llm bench "<that prompt>" --schema ./answer.schema.json
|
|
137
|
+
--samples 3 --max-coins <ceiling> --user-approved`.
|
|
138
|
+
3. Read: p50 `<p50>`, p95 `<p95>`, max `<max>`, and the two printed
|
|
139
|
+
declarations — `Declare estimateCoins: <recommended>` and
|
|
140
|
+
`Grant perCallMaxCoins: <ceiling>`.
|
|
141
|
+
4. Declare `NPC_CALL_PRICE = <recommended>` and
|
|
142
|
+
`NPC_CALL_CEILING = <ceiling>`, each verbatim from the line that printed it.
|
|
143
|
+
5. Unbatched, the loop is ten calls a minute, so the disclosure is
|
|
144
|
+
`10` and `10 * <recommended>` coins per minute.
|
|
145
|
+
6. Batch the five NPCs into one call — benchmark the batched prompt separately,
|
|
146
|
+
because it is a different prompt — and the disclosure becomes two calls a
|
|
147
|
+
minute at the batched price.
|
|
148
|
+
7. Re-run steps 1–4 whenever the prompt, the schema or the model changes. A
|
|
149
|
+
prompt edit is a price change; treat it like one.
|
|
150
|
+
|
|
151
|
+
Every `<placeholder>` above is read off your own benchmark run. None of these
|
|
152
|
+
figures is a platform constant, and none of them should be copied from another
|
|
153
|
+
game — a different prompt has a different price.
|
|
@@ -0,0 +1,96 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: genex-tool-llm
|
|
3
|
+
description: A language model running while somebody PLAYS the finished game — an NPC that answers in its own words, a quest written from what the player typed, a prompt box in the game. On Genex the player pays and approves it, so this is a platform feature, never a key on the author's meter. Read this the moment a request implies a model at play time, before building anything.
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
# Genex Tools · Models at play time
|
|
7
|
+
|
|
8
|
+
This card is about a model that runs **while somebody plays the finished game**
|
|
9
|
+
— not about generating art while you build it. "Make me a picture, a model, a
|
|
10
|
+
voice line" is the asset belt: use the lane card for it and stay here only if a
|
|
11
|
+
model has to run for the player.
|
|
12
|
+
|
|
13
|
+
## What the platform does
|
|
14
|
+
|
|
15
|
+
A game hosted on Genex can call a language model from inside the running game
|
|
16
|
+
and get back text or JSON. **The player pays and the player approves**: coins
|
|
17
|
+
from their Genex wallet, or their own Claude / ChatGPT subscription, chosen on
|
|
18
|
+
an approval sheet Genex draws that the game cannot render, skin or bypass.
|
|
19
|
+
Either one call at a time, or one standing budget the player approves once and
|
|
20
|
+
the game then spends against without another popup. The game holds no provider
|
|
21
|
+
key, sees no credential, and never talks to a model vendor.
|
|
22
|
+
|
|
23
|
+
That is the whole reason this is a platform feature and not something you wire
|
|
24
|
+
up here: an unhosted folder has nowhere to put a key that is not the author's
|
|
25
|
+
own.
|
|
26
|
+
|
|
27
|
+
## Recognise the request
|
|
28
|
+
|
|
29
|
+
Any of these means a model at PLAY time, however it is phrased:
|
|
30
|
+
|
|
31
|
+
- NPCs that talk, answer, argue, or decide in their own words
|
|
32
|
+
- anything written from what the player types — a name, a wish, a command, a question
|
|
33
|
+
- a prompt box, a chat panel, an "ask the oracle" widget inside the game
|
|
34
|
+
- quests, items, dialogue, or levels generated per save or per run
|
|
35
|
+
- "let the player choose a model", "use my API key in the game", "hook a chatbot into it"
|
|
36
|
+
- a judge, a grader, or a referee that reads free-form player input
|
|
37
|
+
|
|
38
|
+
## Offer it in one line, then ASK
|
|
39
|
+
|
|
40
|
+
Say this and stop:
|
|
41
|
+
|
|
42
|
+
> A model running while people play is built into the Genex platform — the
|
|
43
|
+
> player pays, with Genex coins or their own Claude/ChatGPT subscription, and
|
|
44
|
+
> approves it on a Genex sheet; your game just calls `generate()`. Want it
|
|
45
|
+
> that way?
|
|
46
|
+
|
|
47
|
+
Wait for the answer. Do not start building either version first, and do not
|
|
48
|
+
expand the offer into a pitch — one line, one question.
|
|
49
|
+
|
|
50
|
+
## On a yes
|
|
51
|
+
|
|
52
|
+
1. `npx genex llm models` — whether this stand serves the lane at all, and
|
|
53
|
+
which models. It answers in this folder as it is, before anything is
|
|
54
|
+
converted, so it comes first: if it says the lane is off, in-game calls
|
|
55
|
+
answer 404 here — tell the user so plainly, build the graceful fallback, and
|
|
56
|
+
do not convert a folder for a feature the stand does not serve.
|
|
57
|
+
2. `npx genex init --convert` — it connects this folder to a hosted Genex game
|
|
58
|
+
in place. The code, the files and this toolkit stay exactly as they are, and
|
|
59
|
+
generations still land in `./assets`. It is the user's yes that runs it, so
|
|
60
|
+
ask before you do if you have not already.
|
|
61
|
+
3. Load `$genex-llm-in-games` — it arrives with the conversion and owns the
|
|
62
|
+
build: the SDK surface, one-time calls versus a standing budget, measuring
|
|
63
|
+
the price before the game declares it, and honest handling of every refusal.
|
|
64
|
+
|
|
65
|
+
**The game has to be a static browser build.** A hosted game is files served
|
|
66
|
+
from the edge; there is no server of yours inside it. So the familiar pattern —
|
|
67
|
+
a small local Express/Flask app that holds a key and proxies the model — works
|
|
68
|
+
on your machine and can never ship. Calling `generate()` from the browser is
|
|
69
|
+
the shipping shape of that idea.
|
|
70
|
+
|
|
71
|
+
## On a no
|
|
72
|
+
|
|
73
|
+
Then build nothing that calls a model at play time. Not a key in `.env`, not a
|
|
74
|
+
local proxy, not the author's own account behind a fetch. A shipped game
|
|
75
|
+
carrying the author's credential means every visitor spends the author's money,
|
|
76
|
+
with nobody approving anything and no limit on it — and the credential is
|
|
77
|
+
readable in the bundle. Say that in one plain sentence, then build the authored
|
|
78
|
+
version instead: a written dialogue tree, a table of lines, a rule-based
|
|
79
|
+
director. Those are not consolation prizes; they are what most good games use.
|
|
80
|
+
|
|
81
|
+
## The honest boundary — text and JSON only
|
|
82
|
+
|
|
83
|
+
Player-funded generation returns **text or JSON**. Nothing else.
|
|
84
|
+
|
|
85
|
+
- **"The player types anything and gets a 3D model, paid by them"** is not a
|
|
86
|
+
thing on this platform. Say so plainly instead of half-building it.
|
|
87
|
+
- 3D models, images, textures, video, music, voice and characters are the ASSET
|
|
88
|
+
lanes of this toolkit: you generate them while you build, on the user's own
|
|
89
|
+
meter, and they download into `./assets` and ship inside the game. What is
|
|
90
|
+
live and what it costs: `npx genex doctor`.
|
|
91
|
+
- The shape that does work is **JSON parameters, then render**: the model
|
|
92
|
+
returns a structured description and the game builds it from assets and code
|
|
93
|
+
you already shipped — a creature assembled from parts you generated, a room
|
|
94
|
+
laid out from a list of prefab ids, a palette, a stat block, a line delivered
|
|
95
|
+
from pre-generated voice clips. Offer that when somebody asks for the
|
|
96
|
+
impossible version.
|
|
@@ -8,6 +8,9 @@ description: How Genex Tools works across every lane — `npx genex doctor` to c
|
|
|
8
8
|
The rules that apply to every lane. The per-lane cards are
|
|
9
9
|
`$genex-tool-model`, `$genex-tool-image`, `$genex-tool-video`,
|
|
10
10
|
`$genex-tool-texture`, `$genex-tool-audio`, `$genex-tool-character`.
|
|
11
|
+
`$genex-tool-llm` is the odd one out: it is not a generation lane at all but
|
|
12
|
+
the door for a model that runs while somebody PLAYS the finished game — read it
|
|
13
|
+
the moment a request implies one.
|
|
11
14
|
|
|
12
15
|
## Check before you promise
|
|
13
16
|
|
|
@@ -92,8 +95,10 @@ tiling floor and three paid assets nobody loaded.
|
|
|
92
95
|
|
|
93
96
|
Hosting, publishing, multiplayer, remixing and custom domains are the Genex
|
|
94
97
|
platform, not this toolkit — those commands are refused in this folder by
|
|
95
|
-
design, and the refusal says where they live.
|
|
96
|
-
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
98
|
+
design, and the refusal says where they live. So is a model that runs while
|
|
99
|
+
somebody plays the finished game: that one is player-funded and needs a hosted
|
|
100
|
+
game, and `$genex-tool-llm` owns how to offer it. This workspace generates
|
|
101
|
+
assets for a game you build and ship yourself. When the user wants that game
|
|
102
|
+
live on Genex with its own URL, the AGENTS.md rules say how to offer it; on a
|
|
103
|
+
yes the folder is connected to a hosted game in place — same cards, same rules
|
|
104
|
+
— and the `$genex-tool-publish` card arrives with the publishing commands.
|
|
@@ -38,7 +38,7 @@ update, so update immediately.)
|
|
|
38
38
|
Run exactly the command the nudge printed, from the game project root:
|
|
39
39
|
|
|
40
40
|
```bash
|
|
41
|
-
npm i -D @genex-ai/cli-demo@
|
|
41
|
+
npm i -D @genex-ai/cli-demo@dev # the genex CLI (a dev dependency)
|
|
42
42
|
npm i @genex-ai/embed-sdk@latest # identity/saves SDK (ships inside the game)
|
|
43
43
|
npm i @genex-ai/multiplayer@latest # multiplayer SDK (only if the game uses it)
|
|
44
44
|
```
|