@genex-ai/cli-demo 1.36.0 → 1.36.1-dev.776

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,16 +1,28 @@
1
1
  # Calibrate, then declare
2
2
 
3
- `estimateCoins` is the **fixed price of a started attempt**, not a guess about
4
- one. The platform charges exactly what you declare, whether the call succeeds,
5
- fails, is canceled, or stops at its own budget. There is one honest way to pick
6
- it: run the real prompt on your own coins, read what it charged, and declare the
7
- number the benchmark recommends.
8
-
9
- Never derive it from a model vendor's published rates. Three layers sit between
10
- that rate and what the player is charged — the provider's cost, the platform's
11
- tariff, and the headroom a declared price needs — and only the benchmark sees
12
- all three. A number worked out from the vendor's page silently drops the middle
13
- layer and under-prices every call you will ever make.
3
+ An in-game call is paid from the player's Genex credits and **billed as used**:
4
+ its real provider cost plus the platform fee. The game declares no price. It
5
+ declares a per-call **ceiling** — `maxCredits` on `generate()`,
6
+ `perCallMaxCredits` on a standing budget — and no call is ever billed past it.
7
+
8
+ The ceiling is still a number with two jobs, and both punish a guess:
9
+
10
+ - **It bounds the cost.** A budget's worst case, shown to the player beside
11
+ your own estimate, is built from `perCallMaxCredits`. A ceiling far above what
12
+ a call really costs asks the player to approve money the game will never
13
+ spend, and makes the sheet look like a withdrawal.
14
+ - **It sizes the answer.** Each call's room to answer is funded from its
15
+ ceiling, so a ceiling below what the answer needs cuts it off — and a cut-off
16
+ attempt is still billed for the work it did.
17
+
18
+ There is one honest way to pick it: run the real prompt on your own credits,
19
+ read what it really cost, and declare the numbers the benchmark prints.
20
+
21
+ Never derive them from a model vendor's published rates. Three layers sit
22
+ between that rate and what the player is billed — the provider's cost for YOUR
23
+ prompt, the platform fee, and the headroom a ceiling needs — and only the
24
+ benchmark sees the first. The other two come from the server; a number worked
25
+ out from the vendor's page drops at least one of them.
14
26
 
15
27
  ## 1. Freeze the prompt first
16
28
 
@@ -22,7 +34,7 @@ that the real game then cannot fund.
22
34
  If the feature returns structured data, write the schema to a file now
23
35
  (`./answer.schema.json`) and benchmark with it. Schema'd JSON and free text do
24
36
  not cost the same, and shipping without a schema is the commonest way a call is
25
- charged for an unusable result.
37
+ billed for an unusable result.
26
38
 
27
39
  ## 2. Run the benchmark
28
40
 
@@ -33,89 +45,105 @@ npx genex llm bench "<the frozen prompt, one real example filled in>" \
33
45
  --model <id from the line above> \
34
46
  --schema ./answer.schema.json \
35
47
  --samples 3 \
36
- --max-coins <n> --user-approved
48
+ --max-credits <n> --user-approved
37
49
  ```
38
50
 
39
- - It spends **your** coins, on the development lane, through the CLI's own
40
- credential. Nothing here runs in a shipped build.
41
- - `--max-coins <n> --user-approved` is a hard gate, refused before any network
51
+ - It spends **your** credits, on the development lane, through the CLI's own
52
+ credential. Nothing here runs in a shipped build. The build's asset allowance
53
+ does not count it, which is why it has its own gate.
54
+ - `--max-credits <n> --user-approved` is a hard gate, refused before any network
42
55
  call. That is the same shape as every other spend approval in the CLI: you
43
- state the ceiling for this run, out loud, once.
56
+ state the ceiling for this run, out loud, once. Your spendable balance has to
57
+ cover one attempt at it, or the run is refused before it starts.
58
+ `--max-coins` is the old name of the same flag and still works.
44
59
  - `--samples 3` is the floor. The same prompt costs different amounts on
45
60
  different runs, because the model's own output length varies.
46
61
  - `npx genex llm price` re-prints the last run's recommendation without spending
47
- anything again.
62
+ anything again. A file from before credits is reprinted as history, with a
63
+ request to re-run.
48
64
 
49
65
  ## 3. Read the output
50
66
 
51
- Each sample prints what it actually charged. The aggregate prints p50, p95 and
52
- max of the charged coins **over the samples that succeeded and settled**, plus
53
- one recommendation. A sample that failed is printed with its code and stays
54
- out of the numbers; a sample the provider refused at its door
55
- (`provider_http_<status>`) ran no inference, cost nothing, and prints the
56
- provider's own message under its row — on a 401 or 403 that is the stand's
57
- provider configuration refusing the model, which is the operator's to fix. A
58
- sample refused as `generation_limit` never started and ends the run: the
59
- account's three ad-hoc calls are open or recently stopped with a pending
60
- bill. `npx genex llm status` lists them with what each holds;
61
- `npx genex llm cancel <id>` stops an active one; a stopped one frees on its
62
- own once its bill resolves, and stops holding a slot ten minutes after
63
- dispatch. Re-running the bench into the same refusal spends nothing and
67
+ Each sample prints its real cost and the whole credits it was charged. The
68
+ aggregate prints p50, p95 and max of both **over the samples that succeeded and
69
+ settled**. A sample that failed is printed with its code and stays out of the
70
+ numbers; a sample the provider refused at its door (`provider_http_<status>`)
71
+ ran no inference, cost nothing, and prints the provider's own message under its
72
+ row — on a 401 or 403 that is the stand's provider configuration refusing the
73
+ model, which is the operator's to fix. A sample refused as `generation_limit`
74
+ never started and ends the run: the account's three ad-hoc calls are open or
75
+ recently stopped with a pending bill. `npx genex llm status` lists them with
76
+ what each holds; `npx genex llm cancel <id>` stops an active one; a stopped one
77
+ frees on its own once its bill resolves, and stops holding a slot ten minutes
78
+ after dispatch. Re-running the bench into the same refusal spends nothing and
64
79
  learns nothing.
65
80
 
66
81
  A sample the model had to stop writing (`provider_token_limit`) was cut off at
67
- your `--max-coins`: it was charged, it is not a sample, and its real length is
68
- unknown — so the run recommends no price and asks you to re-run with a higher
69
- `--max-coins`. When the stand itself cannot hold the answer, the run says the
70
- answer is longer than one call there may produce; no price fixes that, a
82
+ your `--max-credits`: it was billed, it is not a sample, and its real length is
83
+ unknown — so the run recommends no ceiling and asks you to re-run with a higher
84
+ `--max-credits`. When the stand itself cannot hold the answer, the run says the
85
+ answer is longer than one call there may produce; no ceiling fixes that, a
71
86
  shorter answer does.
72
87
 
73
88
  - **p50** is what a typical call costs. It is the number to reason about when
74
- you ask "can the game afford this loop?" — multiply it by the calls per
75
- minute you are about to disclose.
89
+ you ask "can the game afford this loop?".
76
90
  - **p95** is what a bad-but-normal call costs: a long answer, a model that
77
- reasons its way around. It is the number to **declare**, because a declared
78
- price below it means the unlucky calls cannot fund themselves and get refused
79
- mid-session.
91
+ reasons its way around. It is what the **ceiling** has to cover, because a
92
+ ceiling below it cuts the unlucky calls off.
80
93
  - **max** is diagnostic. When max sits far above p95, the prompt has an
81
94
  unbounded branch in it — usually an unconstrained list or a missing schema.
82
95
  Fix the prompt rather than declaring a bigger number.
83
96
 
84
- The recommendation line already applies the **server's own recommended
85
- headroom** on top of p95. It also covers the answer's **length**: the declared
86
- price decides how long each call's answer may be, because the room to answer is
87
- funded from it, so a price built from charged coins alone can cut the answer
88
- off in the game while the bench — run under a larger `--max-coins` — never saw
89
- it. The recommendation is never below the smallest price that leaves room for
90
- the benchmarked answer, and when the length is what set it the run says so in
91
- one sentence. Either way, declare that figure verbatim — never the bare charged
92
- number:
97
+ Then the run prints what to declare:
98
+
99
+ - **`Declare maxCredits`** — the p95 real cost with the platform fee and the
100
+ **server's own recommended headroom** applied, in whole credits. It also
101
+ covers the answer's **length**: it is never below the smallest ceiling that
102
+ leaves room for the benchmarked answer, and when the length is what set it
103
+ the run says so in one sentence. A ceiling built from cost alone can cut the
104
+ answer off in the game while the bench — run under a larger `--max-credits`
105
+ — never saw it.
106
+ - **`Grant perCallMaxCredits`** — the same over the worst sample, never below
107
+ the line above it. It is emphatically NOT the benchmark's `max` of charged
108
+ credits: that number carries neither the headroom nor the answer's length,
109
+ so a ceiling set from it lands under the call ceiling and the budget's calls
110
+ are refused (`grant_price_unreasonable`).
111
+ - **`Per-call estimate: about … credits per call`** — the measured average,
112
+ fee included, no headroom: a fraction of a credit for a short call. It is the
113
+ honest figure for the disclosure.
114
+ - **`Grant perCallEstimateCredits`** — that average rounded up to a whole
115
+ credit.
116
+
117
+ Declare each verbatim:
93
118
 
94
119
  ```ts
95
- const NPC_CALL_PRICE = <the recommended figure>; // from `npx genex llm bench`, <date>
96
- const NPC_CALL_CEILING = <the printed ceiling>; // same run — grants only
120
+ const NPC_CALL_MAX = <Declare maxCredits>; // from `npx genex llm bench`, <date>
121
+ const NPC_GRANT_MAX = <Grant perCallMaxCredits>; // same run — budgets only
122
+ const NPC_CALL_ESTIMATE = <Grant perCallEstimateCredits>; // same run — budgets only
123
+ const NPC_CALL_AVERAGE = <the "about … credits per call" figure>; // same run
97
124
  ```
98
125
 
99
- Write the benchmark date and the model id beside it in `DESIGN.md`, so the next
100
- person knows what the number describes. Do not add a margin of your own on top
101
- of the recommendation, and do not round it down to look cheaper.
126
+ Write the benchmark date and the model id beside them in `DESIGN.md`, so the
127
+ next person knows what the numbers describe. Do not add a margin of your own on
128
+ top, and do not round them down to look cheaper.
102
129
 
103
- For a grant, the run prints a **second** number beside the price: the ceiling,
104
- `perCallMaxCoins`. Declare that one verbatim too, and do not work it out by
105
- hand. It is emphatically NOT the benchmark's `max`: p95 is nearest-rank, so at
106
- the sample counts a benchmark actually takes, p95 and max are the same figure —
107
- a ceiling set from `max` therefore lands *below* the recommended price, and
108
- `requestSpendGrant()` refuses that pair before the request ever leaves the page.
109
-
110
- The invariant, which the SDK and the server both enforce:
130
+ The invariants, which the SDK and the server both enforce:
111
131
 
112
132
  ```
113
- perCallEstimateCoins ≤ perCallMaxCoins ≤ the limit the player approves
133
+ perCallEstimateCredits ≤ perCallMaxCredits ≤ the limit the player approves
134
+ maxCredits on each call ≤ perCallMaxCredits
114
135
  ```
115
136
 
116
- The ceiling is the price's room to be wrong, not a second price: no call is ever
117
- charged more than the price it declares, and the platform separately refuses any
118
- declared price out of proportion to what the model could really cost.
137
+ A ceiling is not a second price: a call is billed the real cost of the work it
138
+ did plus the fee (a one-time call rounded up to a whole credit), and never past
139
+ its ceiling.
140
+
141
+ **One-time calls round up; budgets do not.** A one-time `generate()` is captured
142
+ rounded UP to a whole credit when it settles, so a call that really costs a
143
+ fraction of a credit still costs the player one whole credit each time. Under a
144
+ budget each call adds its exact fraction and the player's credits move one
145
+ whole credit at a time. That is one more reason a repeated call belongs under a
146
+ budget.
119
147
 
120
148
  ## 4. Turn the game's loop into the disclosure
121
149
 
@@ -129,27 +157,39 @@ from the loop you wrote, counted honestly:
129
157
  3. **Multiply, and pick the period that makes the number legible.** Five NPCs at
130
158
  one call per thirty seconds is ten calls a minute, so `periodLabel: 'minute'`
131
159
  and `estimatedCallsPerPeriod: 10`.
132
- 4. **Coins per period is calls × the declared price** —
133
- `estimatedCoinsPerPeriod: 10 * NPC_CALL_PRICE`. Not the p50, not a hope: the
134
- price you declare is the price charged.
160
+ 4. **Credits per period is calls × the measured average, rounded up** —
161
+ `estimatedCreditsPerPeriod: Math.ceil(10 * NPC_CALL_AVERAGE)`. Not the
162
+ ceiling: a call is billed what it uses, so the honest estimate is what calls
163
+ really average.
135
164
 
136
165
  ```ts
166
+ perCallMaxCredits: NPC_GRANT_MAX,
167
+ perCallEstimateCredits: NPC_CALL_ESTIMATE,
137
168
  disclosure: {
138
169
  periodLabel: 'minute',
139
170
  estimatedCallsPerPeriod: 10,
140
- estimatedCoinsPerPeriod: 10 * NPC_CALL_PRICE,
171
+ estimatedCreditsPerPeriod: Math.ceil(10 * NPC_CALL_AVERAGE),
141
172
  }
142
173
  ```
143
174
 
144
175
  The player sees this attributed to your game — "the game estimates about …" —
145
- beside the platform's own worst case computed from `perCallMaxCoins`. The two
176
+ beside the platform's own worst case computed from `perCallMaxCredits`. The two
146
177
  being far apart is normal; the estimate being far below what the game really
147
178
  does is what breaks trust and burns the grant mid-session.
148
179
 
149
- **Batching changes this arithmetic more than any price tuning can.** Five NPCs
180
+ **Batching changes this arithmetic more than any tuning can.** Five NPCs
150
181
  answered by one call that returns five decisions is two calls a minute, not ten,
151
- and one benchmarked price for the batched prompt replaces five of the unbatched
152
- one. Do that before you reach for a cheaper model.
182
+ and one benchmark of the batched prompt replaces five of the unbatched one. Do
183
+ that before you reach for a cheaper model.
184
+
185
+ ## Old code
186
+
187
+ `estimateCoins`, `perCallMaxCoins`, `perCallEstimateCoins` and
188
+ `disclosure.estimatedCoinsPerPeriod` are the names from before credits. The
189
+ server still accepts them and reads them as the same numbers in credits, but
190
+ they are deprecated: rename them to `maxCredits`, `perCallMaxCredits`,
191
+ `perCallEstimateCredits` and `estimatedCreditsPerPeriod` when you touch the
192
+ code, and re-benchmark — a number measured as a coin price is not a ceiling.
153
193
 
154
194
  ## Worked example
155
195
 
@@ -160,20 +200,21 @@ speech). The prompt carries the NPC's memory and the last two player lines.
160
200
  1. Freeze the prompt with a full memory and a long player line. Write
161
201
  `./answer.schema.json` with the two fields.
162
202
  2. `npx genex llm bench "<that prompt>" --schema ./answer.schema.json
163
- --samples 3 --max-coins <ceiling> --user-approved`.
164
- 3. Read: p50 `<p50>`, p95 `<p95>`, max `<max>`, and the two printed
165
- declarations — `Declare estimateCoins: <recommended>` and
166
- `Grant perCallMaxCoins: <ceiling>`.
167
- 4. Declare `NPC_CALL_PRICE = <recommended>` and
168
- `NPC_CALL_CEILING = <ceiling>`, each verbatim from the line that printed it.
169
- 5. Unbatched, the loop is ten calls a minute, so the disclosure is
170
- `10` and `10 * <recommended>` coins per minute.
203
+ --samples 3 --max-credits <ceiling> --user-approved`.
204
+ 3. Read: p50 `<p50>`, p95 `<p95>`, max `<max>`, and the printed lines —
205
+ `Declare maxCredits: <call ceiling>`, `Grant perCallMaxCredits: <grant ceiling>`,
206
+ `Per-call estimate: about <average> credits per call` and
207
+ `Grant perCallEstimateCredits: <estimate>`.
208
+ 4. Declare `NPC_CALL_MAX`, `NPC_GRANT_MAX`, `NPC_CALL_ESTIMATE` and
209
+ `NPC_CALL_AVERAGE`, each verbatim from the line that printed it.
210
+ 5. Unbatched, the loop is ten calls a minute, so the disclosure is `10` calls
211
+ and `Math.ceil(10 * NPC_CALL_AVERAGE)` credits per minute.
171
212
  6. Batch the five NPCs into one call — benchmark the batched prompt separately,
172
213
  because it is a different prompt — and the disclosure becomes two calls a
173
- minute at the batched price.
214
+ minute at the batched average.
174
215
  7. Re-run steps 1–4 whenever the prompt, the schema or the model changes. A
175
- prompt edit is a price change; treat it like one.
216
+ prompt edit changes the ceiling; treat it like one.
176
217
 
177
218
  Every `<placeholder>` above is read off your own benchmark run. None of these
178
219
  figures is a platform constant, and none of them should be copied from another
179
- game — a different prompt has a different price.
220
+ game — a different prompt has a different cost.
@@ -13,11 +13,12 @@ model has to run for the player.
13
13
  ## What the platform does
14
14
 
15
15
  A game hosted on Genex can call a language model from inside the running game
16
- and get back text or JSON. **The player pays and the player approves**: coins
17
- from their Genex wallet, or their own Claude / ChatGPT subscription, chosen on
18
- an approval sheet Genex draws that the game cannot render, skin or bypass.
19
- Either one call at a time, or one standing budget the player approves once and
20
- the game then spends against without another popup. The game holds no provider
16
+ and get back text or JSON. **The player pays and the player approves**: credits
17
+ from their Genex account, billed as each call is used and never past the
18
+ per-call ceiling the game declares, or their own Claude / ChatGPT subscription,
19
+ chosen on an approval sheet Genex draws that the game cannot render, skin or
20
+ bypass. Either one call at a time, or one standing budget the player approves
21
+ once and the game then spends against without another popup. The game holds no provider
21
22
  key, sees no credential, and never talks to a model vendor.
22
23
 
23
24
  That is the whole reason this is a platform feature and not something you wire
@@ -40,9 +41,9 @@ Any of these means a model at PLAY time, however it is phrased:
40
41
  Say this and stop:
41
42
 
42
43
  > A model running while people play is built into the Genex platform — the
43
- > player pays, with Genex coins or their own Claude/ChatGPT subscription, and
44
- > approves it on a Genex sheet; your game just calls `generate()`. Want it
45
- > that way?
44
+ > player pays, with their Genex credits or their own Claude/ChatGPT
45
+ > subscription, and approves it on a Genex sheet; your game just calls
46
+ > `generate()`. Want it that way?
46
47
 
47
48
  Wait for the answer. Do not start building either version first, and do not
48
49
  expand the offer into a pitch — one line, one question.
@@ -64,7 +65,8 @@ expand the offer into a pitch — one line, one question.
64
65
  ask before you do if you have not already.
65
66
  3. Load `$genex-llm-in-games` — it arrives with the conversion and owns the
66
67
  build: the SDK surface, one-time calls versus a standing budget, measuring
67
- the price before the game declares it, and honest handling of every refusal.
68
+ the per-call ceiling before the game declares it, and honest handling of
69
+ every refusal.
68
70
 
69
71
  **The game has to be a static browser build.** A hosted game is files served
70
72
  from the edge; there is no server of yours inside it. So the familiar pattern —
@@ -38,7 +38,7 @@ update, so update immediately.)
38
38
  Run exactly the command the nudge printed, from the game project root:
39
39
 
40
40
  ```bash
41
- npm i -D @genex-ai/cli-demo@latest # the genex CLI (a dev dependency)
41
+ npm i -D @genex-ai/cli-demo@dev # the genex CLI (a dev dependency)
42
42
  npm i @genex-ai/embed-sdk@latest # identity/saves SDK (ships inside the game)
43
43
  npm i @genex-ai/multiplayer@latest # multiplayer SDK (only if the game uses it)
44
44
  ```