@genex-ai/cli-demo 1.36.0-dev.772 → 1.36.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,28 +1,16 @@
1
1
  # Calibrate, then declare
2
2
 
3
- An in-game call is paid from the player's Genex credits and **billed as used**:
4
- its real provider cost plus the platform fee. The game declares no price. It
5
- declares a per-call **ceiling** — `maxCredits` on `generate()`,
6
- `perCallMaxCredits` on a standing budget — and no call is ever billed past it.
7
-
8
- The ceiling is still a number with two jobs, and both punish a guess:
9
-
10
- - **It bounds the cost.** A budget's worst case, shown to the player beside
11
- your own estimate, is built from `perCallMaxCredits`. A ceiling far above what
12
- a call really costs asks the player to approve money the game will never
13
- spend, and makes the sheet look like a withdrawal.
14
- - **It sizes the answer.** Each call's room to answer is funded from its
15
- ceiling, so a ceiling below what the answer needs cuts it off — and a cut-off
16
- attempt is still billed for the work it did.
17
-
18
- There is one honest way to pick it: run the real prompt on your own credits,
19
- read what it really cost, and declare the numbers the benchmark prints.
20
-
21
- Never derive them from a model vendor's published rates. Three layers sit
22
- between that rate and what the player is billed — the provider's cost for YOUR
23
- prompt, the platform fee, and the headroom a ceiling needs — and only the
24
- benchmark sees the first. The other two come from the server; a number worked
25
- out from the vendor's page drops at least one of them.
3
+ `estimateCoins` is the **fixed price of a started attempt**, not a guess about
4
+ one. The platform charges exactly what you declare, whether the call succeeds,
5
+ fails, is canceled, or stops at its own budget. There is one honest way to pick
6
+ it: run the real prompt on your own coins, read what it charged, and declare the
7
+ number the benchmark recommends.
8
+
9
+ Never derive it from a model vendor's published rates. Three layers sit between
10
+ that rate and what the player is charged — the provider's cost, the platform's
11
+ tariff, and the headroom a declared price needs — and only the benchmark sees
12
+ all three. A number worked out from the vendor's page silently drops the middle
13
+ layer and under-prices every call you will ever make.
26
14
 
27
15
  ## 1. Freeze the prompt first
28
16
 
@@ -34,7 +22,7 @@ that the real game then cannot fund.
34
22
  If the feature returns structured data, write the schema to a file now
35
23
  (`./answer.schema.json`) and benchmark with it. Schema'd JSON and free text do
36
24
  not cost the same, and shipping without a schema is the commonest way a call is
37
- billed for an unusable result.
25
+ charged for an unusable result.
38
26
 
39
27
  ## 2. Run the benchmark
40
28
 
@@ -45,105 +33,89 @@ npx genex llm bench "<the frozen prompt, one real example filled in>" \
45
33
  --model <id from the line above> \
46
34
  --schema ./answer.schema.json \
47
35
  --samples 3 \
48
- --max-credits <n> --user-approved
36
+ --max-coins <n> --user-approved
49
37
  ```
50
38
 
51
- - It spends **your** credits, on the development lane, through the CLI's own
52
- credential. Nothing here runs in a shipped build. The build's asset allowance
53
- does not count it, which is why it has its own gate.
54
- - `--max-credits <n> --user-approved` is a hard gate, refused before any network
39
+ - It spends **your** coins, on the development lane, through the CLI's own
40
+ credential. Nothing here runs in a shipped build.
41
+ - `--max-coins <n> --user-approved` is a hard gate, refused before any network
55
42
  call. That is the same shape as every other spend approval in the CLI: you
56
- state the ceiling for this run, out loud, once. Your spendable balance has to
57
- cover one attempt at it, or the run is refused before it starts.
58
- `--max-coins` is the old name of the same flag and still works.
43
+ state the ceiling for this run, out loud, once.
59
44
  - `--samples 3` is the floor. The same prompt costs different amounts on
60
45
  different runs, because the model's own output length varies.
61
46
  - `npx genex llm price` re-prints the last run's recommendation without spending
62
- anything again. A file from before credits is reprinted as history, with a
63
- request to re-run.
47
+ anything again.
64
48
 
65
49
  ## 3. Read the output
66
50
 
67
- Each sample prints its real cost and the whole credits it was charged. The
68
- aggregate prints p50, p95 and max of both **over the samples that succeeded and
69
- settled**. A sample that failed is printed with its code and stays out of the
70
- numbers; a sample the provider refused at its door (`provider_http_<status>`)
71
- ran no inference, cost nothing, and prints the provider's own message under its
72
- row — on a 401 or 403 that is the stand's provider configuration refusing the
73
- model, which is the operator's to fix. A sample refused as `generation_limit`
74
- never started and ends the run: the account's three ad-hoc calls are open or
75
- recently stopped with a pending bill. `npx genex llm status` lists them with
76
- what each holds; `npx genex llm cancel <id>` stops an active one; a stopped one
77
- frees on its own once its bill resolves, and stops holding a slot ten minutes
78
- after dispatch. Re-running the bench into the same refusal spends nothing and
51
+ Each sample prints what it actually charged. The aggregate prints p50, p95 and
52
+ max of the charged coins **over the samples that succeeded and settled**, plus
53
+ one recommendation. A sample that failed is printed with its code and stays
54
+ out of the numbers; a sample the provider refused at its door
55
+ (`provider_http_<status>`) ran no inference, cost nothing, and prints the
56
+ provider's own message under its row — on a 401 or 403 that is the stand's
57
+ provider configuration refusing the model, which is the operator's to fix. A
58
+ sample refused as `generation_limit` never started and ends the run: the
59
+ account's three ad-hoc calls are open or recently stopped with a pending
60
+ bill. `npx genex llm status` lists them with what each holds;
61
+ `npx genex llm cancel <id>` stops an active one; a stopped one frees on its
62
+ own once its bill resolves, and stops holding a slot ten minutes after
63
+ dispatch. Re-running the bench into the same refusal spends nothing and
79
64
  learns nothing.
80
65
 
81
66
  A sample the model had to stop writing (`provider_token_limit`) was cut off at
82
- your `--max-credits`: it was billed, it is not a sample, and its real length is
83
- unknown — so the run recommends no ceiling and asks you to re-run with a higher
84
- `--max-credits`. When the stand itself cannot hold the answer, the run says the
85
- answer is longer than one call there may produce; no ceiling fixes that, a
67
+ your `--max-coins`: it was charged, it is not a sample, and its real length is
68
+ unknown — so the run recommends no price and asks you to re-run with a higher
69
+ `--max-coins`. When the stand itself cannot hold the answer, the run says the
70
+ answer is longer than one call there may produce; no price fixes that, a
86
71
  shorter answer does.
87
72
 
88
73
  - **p50** is what a typical call costs. It is the number to reason about when
89
- you ask "can the game afford this loop?".
74
+ you ask "can the game afford this loop?" — multiply it by the calls per
75
+ minute you are about to disclose.
90
76
  - **p95** is what a bad-but-normal call costs: a long answer, a model that
91
- reasons its way around. It is what the **ceiling** has to cover, because a
92
- ceiling below it cuts the unlucky calls off.
77
+ reasons its way around. It is the number to **declare**, because a declared
78
+ price below it means the unlucky calls cannot fund themselves and get refused
79
+ mid-session.
93
80
  - **max** is diagnostic. When max sits far above p95, the prompt has an
94
81
  unbounded branch in it — usually an unconstrained list or a missing schema.
95
82
  Fix the prompt rather than declaring a bigger number.
96
83
 
97
- Then the run prints what to declare:
98
-
99
- - **`Declare maxCredits`** — the p95 real cost with the platform fee and the
100
- **server's own recommended headroom** applied, in whole credits. It also
101
- covers the answer's **length**: it is never below the smallest ceiling that
102
- leaves room for the benchmarked answer, and when the length is what set it
103
- the run says so in one sentence. A ceiling built from cost alone can cut the
104
- answer off in the game while the bench — run under a larger `--max-credits`
105
- — never saw it.
106
- - **`Grant perCallMaxCredits`** — the same over the worst sample, never below
107
- the line above it. It is emphatically NOT the benchmark's `max` of charged
108
- credits: that number carries neither the headroom nor the answer's length,
109
- so a ceiling set from it lands under the call ceiling and the budget's calls
110
- are refused (`grant_price_unreasonable`).
111
- - **`Per-call estimate: about … credits per call`** — the measured average,
112
- fee included, no headroom: a fraction of a credit for a short call. It is the
113
- honest figure for the disclosure.
114
- - **`Grant perCallEstimateCredits`** — that average rounded up to a whole
115
- credit.
116
-
117
- Declare each verbatim:
84
+ The recommendation line already applies the **server's own recommended
85
+ headroom** on top of p95. It also covers the answer's **length**: the declared
86
+ price decides how long each call's answer may be, because the room to answer is
87
+ funded from it, so a price built from charged coins alone can cut the answer
88
+ off in the game while the bench — run under a larger `--max-coins` — never saw
89
+ it. The recommendation is never below the smallest price that leaves room for
90
+ the benchmarked answer, and when the length is what set it the run says so in
91
+ one sentence. Either way, declare that figure verbatim — never the bare charged
92
+ number:
118
93
 
119
94
  ```ts
120
- const NPC_CALL_MAX = <Declare maxCredits>; // from `npx genex llm bench`, <date>
121
- const NPC_GRANT_MAX = <Grant perCallMaxCredits>; // same run — budgets only
122
- const NPC_CALL_ESTIMATE = <Grant perCallEstimateCredits>; // same run — budgets only
123
- const NPC_CALL_AVERAGE = <the "about … credits per call" figure>; // same run
95
+ const NPC_CALL_PRICE = <the recommended figure>; // from `npx genex llm bench`, <date>
96
+ const NPC_CALL_CEILING = <the printed ceiling>; // same run — grants only
124
97
  ```
125
98
 
126
- Write the benchmark date and the model id beside them in `DESIGN.md`, so the
127
- next person knows what the numbers describe. Do not add a margin of your own on
128
- top, and do not round them down to look cheaper.
99
+ Write the benchmark date and the model id beside it in `DESIGN.md`, so the next
100
+ person knows what the number describes. Do not add a margin of your own on top
101
+ of the recommendation, and do not round it down to look cheaper.
129
102
 
130
- The invariants, which the SDK and the server both enforce:
103
+ For a grant, the run prints a **second** number beside the price: the ceiling,
104
+ `perCallMaxCoins`. Declare that one verbatim too, and do not work it out by
105
+ hand. It is emphatically NOT the benchmark's `max`: p95 is nearest-rank, so at
106
+ the sample counts a benchmark actually takes, p95 and max are the same figure —
107
+ a ceiling set from `max` therefore lands *below* the recommended price, and
108
+ `requestSpendGrant()` refuses that pair before the request ever leaves the page.
109
+
110
+ The invariant, which the SDK and the server both enforce:
131
111
 
132
112
  ```
133
- perCallEstimateCredits ≤ perCallMaxCredits ≤ the limit the player approves
134
- maxCredits on each call ≤ perCallMaxCredits
113
+ perCallEstimateCoins ≤ perCallMaxCoins ≤ the limit the player approves
135
114
  ```
136
115
 
137
- A ceiling is not a second price: a call is billed the real cost of the work it
138
- did plus the fee (a one-time call rounded up to a whole credit), and never past
139
- its ceiling.
140
-
141
- **One-time calls round up; budgets do not.** A one-time `generate()` is captured
142
- rounded UP to a whole credit when it settles, so a call that really costs a
143
- fraction of a credit still costs the player one whole credit each time. Under a
144
- budget each call adds its exact fraction and the player's credits move one
145
- whole credit at a time. That is one more reason a repeated call belongs under a
146
- budget.
116
+ The ceiling is the price's room to be wrong, not a second price: no call is ever
117
+ charged more than the price it declares, and the platform separately refuses any
118
+ declared price out of proportion to what the model could really cost.
147
119
 
148
120
  ## 4. Turn the game's loop into the disclosure
149
121
 
@@ -157,39 +129,27 @@ from the loop you wrote, counted honestly:
157
129
  3. **Multiply, and pick the period that makes the number legible.** Five NPCs at
158
130
  one call per thirty seconds is ten calls a minute, so `periodLabel: 'minute'`
159
131
  and `estimatedCallsPerPeriod: 10`.
160
- 4. **Credits per period is calls × the measured average, rounded up** —
161
- `estimatedCreditsPerPeriod: Math.ceil(10 * NPC_CALL_AVERAGE)`. Not the
162
- ceiling: a call is billed what it uses, so the honest estimate is what calls
163
- really average.
132
+ 4. **Coins per period is calls × the declared price** —
133
+ `estimatedCoinsPerPeriod: 10 * NPC_CALL_PRICE`. Not the p50, not a hope: the
134
+ price you declare is the price charged.
164
135
 
165
136
  ```ts
166
- perCallMaxCredits: NPC_GRANT_MAX,
167
- perCallEstimateCredits: NPC_CALL_ESTIMATE,
168
137
  disclosure: {
169
138
  periodLabel: 'minute',
170
139
  estimatedCallsPerPeriod: 10,
171
- estimatedCreditsPerPeriod: Math.ceil(10 * NPC_CALL_AVERAGE),
140
+ estimatedCoinsPerPeriod: 10 * NPC_CALL_PRICE,
172
141
  }
173
142
  ```
174
143
 
175
144
  The player sees this attributed to your game — "the game estimates about …" —
176
- beside the platform's own worst case computed from `perCallMaxCredits`. The two
145
+ beside the platform's own worst case computed from `perCallMaxCoins`. The two
177
146
  being far apart is normal; the estimate being far below what the game really
178
147
  does is what breaks trust and burns the grant mid-session.
179
148
 
180
- **Batching changes this arithmetic more than any tuning can.** Five NPCs
149
+ **Batching changes this arithmetic more than any price tuning can.** Five NPCs
181
150
  answered by one call that returns five decisions is two calls a minute, not ten,
182
- and one benchmark of the batched prompt replaces five of the unbatched one. Do
183
- that before you reach for a cheaper model.
184
-
185
- ## Old code
186
-
187
- `estimateCoins`, `perCallMaxCoins`, `perCallEstimateCoins` and
188
- `disclosure.estimatedCoinsPerPeriod` are the names from before credits. The
189
- server still accepts them and reads them as the same numbers in credits, but
190
- they are deprecated: rename them to `maxCredits`, `perCallMaxCredits`,
191
- `perCallEstimateCredits` and `estimatedCreditsPerPeriod` when you touch the
192
- code, and re-benchmark — a number measured as a coin price is not a ceiling.
151
+ and one benchmarked price for the batched prompt replaces five of the unbatched
152
+ one. Do that before you reach for a cheaper model.
193
153
 
194
154
  ## Worked example
195
155
 
@@ -200,21 +160,20 @@ speech). The prompt carries the NPC's memory and the last two player lines.
200
160
  1. Freeze the prompt with a full memory and a long player line. Write
201
161
  `./answer.schema.json` with the two fields.
202
162
  2. `npx genex llm bench "<that prompt>" --schema ./answer.schema.json
203
- --samples 3 --max-credits <ceiling> --user-approved`.
204
- 3. Read: p50 `<p50>`, p95 `<p95>`, max `<max>`, and the printed lines —
205
- `Declare maxCredits: <call ceiling>`, `Grant perCallMaxCredits: <grant ceiling>`,
206
- `Per-call estimate: about <average> credits per call` and
207
- `Grant perCallEstimateCredits: <estimate>`.
208
- 4. Declare `NPC_CALL_MAX`, `NPC_GRANT_MAX`, `NPC_CALL_ESTIMATE` and
209
- `NPC_CALL_AVERAGE`, each verbatim from the line that printed it.
210
- 5. Unbatched, the loop is ten calls a minute, so the disclosure is `10` calls
211
- and `Math.ceil(10 * NPC_CALL_AVERAGE)` credits per minute.
163
+ --samples 3 --max-coins <ceiling> --user-approved`.
164
+ 3. Read: p50 `<p50>`, p95 `<p95>`, max `<max>`, and the two printed
165
+ declarations — `Declare estimateCoins: <recommended>` and
166
+ `Grant perCallMaxCoins: <ceiling>`.
167
+ 4. Declare `NPC_CALL_PRICE = <recommended>` and
168
+ `NPC_CALL_CEILING = <ceiling>`, each verbatim from the line that printed it.
169
+ 5. Unbatched, the loop is ten calls a minute, so the disclosure is
170
+ `10` and `10 * <recommended>` coins per minute.
212
171
  6. Batch the five NPCs into one call — benchmark the batched prompt separately,
213
172
  because it is a different prompt — and the disclosure becomes two calls a
214
- minute at the batched average.
173
+ minute at the batched price.
215
174
  7. Re-run steps 1–4 whenever the prompt, the schema or the model changes. A
216
- prompt edit changes the ceiling; treat it like one.
175
+ prompt edit is a price change; treat it like one.
217
176
 
218
177
  Every `<placeholder>` above is read off your own benchmark run. None of these
219
178
  figures is a platform constant, and none of them should be copied from another
220
- game — a different prompt has a different cost.
179
+ game — a different prompt has a different price.
@@ -13,12 +13,11 @@ model has to run for the player.
13
13
  ## What the platform does
14
14
 
15
15
  A game hosted on Genex can call a language model from inside the running game
16
- and get back text or JSON. **The player pays and the player approves**: credits
17
- from their Genex account, billed as each call is used and never past the
18
- per-call ceiling the game declares, or their own Claude / ChatGPT subscription,
19
- chosen on an approval sheet Genex draws that the game cannot render, skin or
20
- bypass. Either one call at a time, or one standing budget the player approves
21
- once and the game then spends against without another popup. The game holds no provider
16
+ and get back text or JSON. **The player pays and the player approves**: coins
17
+ from their Genex wallet, or their own Claude / ChatGPT subscription, chosen on
18
+ an approval sheet Genex draws that the game cannot render, skin or bypass.
19
+ Either one call at a time, or one standing budget the player approves once and
20
+ the game then spends against without another popup. The game holds no provider
22
21
  key, sees no credential, and never talks to a model vendor.
23
22
 
24
23
  That is the whole reason this is a platform feature and not something you wire
@@ -41,9 +40,9 @@ Any of these means a model at PLAY time, however it is phrased:
41
40
  Say this and stop:
42
41
 
43
42
  > A model running while people play is built into the Genex platform — the
44
- > player pays, with their Genex credits or their own Claude/ChatGPT
45
- > subscription, and approves it on a Genex sheet; your game just calls
46
- > `generate()`. Want it that way?
43
+ > player pays, with Genex coins or their own Claude/ChatGPT subscription, and
44
+ > approves it on a Genex sheet; your game just calls `generate()`. Want it
45
+ > that way?
47
46
 
48
47
  Wait for the answer. Do not start building either version first, and do not
49
48
  expand the offer into a pitch — one line, one question.
@@ -65,8 +64,7 @@ expand the offer into a pitch — one line, one question.
65
64
  ask before you do if you have not already.
66
65
  3. Load `$genex-llm-in-games` — it arrives with the conversion and owns the
67
66
  build: the SDK surface, one-time calls versus a standing budget, measuring
68
- the per-call ceiling before the game declares it, and honest handling of
69
- every refusal.
67
+ the price before the game declares it, and honest handling of every refusal.
70
68
 
71
69
  **The game has to be a static browser build.** A hosted game is files served
72
70
  from the edge; there is no server of yours inside it. So the familiar pattern —
@@ -38,7 +38,7 @@ update, so update immediately.)
38
38
  Run exactly the command the nudge printed, from the game project root:
39
39
 
40
40
  ```bash
41
- npm i -D @genex-ai/cli-demo@dev # the genex CLI (a dev dependency)
41
+ npm i -D @genex-ai/cli-demo@latest # the genex CLI (a dev dependency)
42
42
  npm i @genex-ai/embed-sdk@latest # identity/saves SDK (ships inside the game)
43
43
  npm i @genex-ai/multiplayer@latest # multiplayer SDK (only if the game uses it)
44
44
  ```