@genex-ai/cli-demo 1.36.0 → 1.36.1-dev.776
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/{blender-mcp-NRJ5FI22.js → blender-mcp-ZRCMYTXI.js} +2 -2
- package/dist/{blender-serve-66OHC4ML.js → blender-serve-EH2MOW3U.js} +1 -1
- package/dist/{chunk-QI7FIYBY.js → chunk-K4EHXOZK.js} +1 -1
- package/dist/{chunk-GOP5FRSJ.js → chunk-Y6WO5FJZ.js} +2 -2
- package/dist/index.js +454 -503
- package/package.json +1 -1
- package/templates/skills/genex-ai-model/SKILL.md +4 -1
- package/templates/skills/genex-game-director/SKILL.md +1 -1
- package/templates/skills/genex-getting-started/SKILL.md +2 -2
- package/templates/skills/genex-llm-in-games/SKILL.md +209 -72
- package/templates/skills/genex-llm-in-games/references/pricing.md +126 -85
- package/templates/skills/genex-tool-llm/SKILL.md +11 -9
- package/templates/skills/genex-updates/SKILL.md +1 -1
|
@@ -1,16 +1,28 @@
|
|
|
1
1
|
# Calibrate, then declare
|
|
2
2
|
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
3
|
+
An in-game call is paid from the player's Genex credits and **billed as used**:
|
|
4
|
+
its real provider cost plus the platform fee. The game declares no price. It
|
|
5
|
+
declares a per-call **ceiling** — `maxCredits` on `generate()`,
|
|
6
|
+
`perCallMaxCredits` on a standing budget — and no call is ever billed past it.
|
|
7
|
+
|
|
8
|
+
The ceiling is still a number with two jobs, and both punish a guess:
|
|
9
|
+
|
|
10
|
+
- **It bounds the cost.** A budget's worst case, shown to the player beside
|
|
11
|
+
your own estimate, is built from `perCallMaxCredits`. A ceiling far above what
|
|
12
|
+
a call really costs asks the player to approve money the game will never
|
|
13
|
+
spend, and makes the sheet look like a withdrawal.
|
|
14
|
+
- **It sizes the answer.** Each call's room to answer is funded from its
|
|
15
|
+
ceiling, so a ceiling below what the answer needs cuts it off — and a cut-off
|
|
16
|
+
attempt is still billed for the work it did.
|
|
17
|
+
|
|
18
|
+
There is one honest way to pick it: run the real prompt on your own credits,
|
|
19
|
+
read what it really cost, and declare the numbers the benchmark prints.
|
|
20
|
+
|
|
21
|
+
Never derive them from a model vendor's published rates. Three layers sit
|
|
22
|
+
between that rate and what the player is billed — the provider's cost for YOUR
|
|
23
|
+
prompt, the platform fee, and the headroom a ceiling needs — and only the
|
|
24
|
+
benchmark sees the first. The other two come from the server; a number worked
|
|
25
|
+
out from the vendor's page drops at least one of them.
|
|
14
26
|
|
|
15
27
|
## 1. Freeze the prompt first
|
|
16
28
|
|
|
@@ -22,7 +34,7 @@ that the real game then cannot fund.
|
|
|
22
34
|
If the feature returns structured data, write the schema to a file now
|
|
23
35
|
(`./answer.schema.json`) and benchmark with it. Schema'd JSON and free text do
|
|
24
36
|
not cost the same, and shipping without a schema is the commonest way a call is
|
|
25
|
-
|
|
37
|
+
billed for an unusable result.
|
|
26
38
|
|
|
27
39
|
## 2. Run the benchmark
|
|
28
40
|
|
|
@@ -33,89 +45,105 @@ npx genex llm bench "<the frozen prompt, one real example filled in>" \
|
|
|
33
45
|
--model <id from the line above> \
|
|
34
46
|
--schema ./answer.schema.json \
|
|
35
47
|
--samples 3 \
|
|
36
|
-
--max-
|
|
48
|
+
--max-credits <n> --user-approved
|
|
37
49
|
```
|
|
38
50
|
|
|
39
|
-
- It spends **your**
|
|
40
|
-
credential. Nothing here runs in a shipped build.
|
|
41
|
-
|
|
51
|
+
- It spends **your** credits, on the development lane, through the CLI's own
|
|
52
|
+
credential. Nothing here runs in a shipped build. The build's asset allowance
|
|
53
|
+
does not count it, which is why it has its own gate.
|
|
54
|
+
- `--max-credits <n> --user-approved` is a hard gate, refused before any network
|
|
42
55
|
call. That is the same shape as every other spend approval in the CLI: you
|
|
43
|
-
state the ceiling for this run, out loud, once.
|
|
56
|
+
state the ceiling for this run, out loud, once. Your spendable balance has to
|
|
57
|
+
cover one attempt at it, or the run is refused before it starts.
|
|
58
|
+
`--max-coins` is the old name of the same flag and still works.
|
|
44
59
|
- `--samples 3` is the floor. The same prompt costs different amounts on
|
|
45
60
|
different runs, because the model's own output length varies.
|
|
46
61
|
- `npx genex llm price` re-prints the last run's recommendation without spending
|
|
47
|
-
anything again.
|
|
62
|
+
anything again. A file from before credits is reprinted as history, with a
|
|
63
|
+
request to re-run.
|
|
48
64
|
|
|
49
65
|
## 3. Read the output
|
|
50
66
|
|
|
51
|
-
Each sample prints
|
|
52
|
-
max of
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
56
|
-
|
|
57
|
-
|
|
58
|
-
|
|
59
|
-
|
|
60
|
-
|
|
61
|
-
|
|
62
|
-
|
|
63
|
-
dispatch. Re-running the bench into the same refusal spends nothing and
|
|
67
|
+
Each sample prints its real cost and the whole credits it was charged. The
|
|
68
|
+
aggregate prints p50, p95 and max of both **over the samples that succeeded and
|
|
69
|
+
settled**. A sample that failed is printed with its code and stays out of the
|
|
70
|
+
numbers; a sample the provider refused at its door (`provider_http_<status>`)
|
|
71
|
+
ran no inference, cost nothing, and prints the provider's own message under its
|
|
72
|
+
row — on a 401 or 403 that is the stand's provider configuration refusing the
|
|
73
|
+
model, which is the operator's to fix. A sample refused as `generation_limit`
|
|
74
|
+
never started and ends the run: the account's three ad-hoc calls are open or
|
|
75
|
+
recently stopped with a pending bill. `npx genex llm status` lists them with
|
|
76
|
+
what each holds; `npx genex llm cancel <id>` stops an active one; a stopped one
|
|
77
|
+
frees on its own once its bill resolves, and stops holding a slot ten minutes
|
|
78
|
+
after dispatch. Re-running the bench into the same refusal spends nothing and
|
|
64
79
|
learns nothing.
|
|
65
80
|
|
|
66
81
|
A sample the model had to stop writing (`provider_token_limit`) was cut off at
|
|
67
|
-
your `--max-
|
|
68
|
-
unknown — so the run recommends no
|
|
69
|
-
`--max-
|
|
70
|
-
answer is longer than one call there may produce; no
|
|
82
|
+
your `--max-credits`: it was billed, it is not a sample, and its real length is
|
|
83
|
+
unknown — so the run recommends no ceiling and asks you to re-run with a higher
|
|
84
|
+
`--max-credits`. When the stand itself cannot hold the answer, the run says the
|
|
85
|
+
answer is longer than one call there may produce; no ceiling fixes that, a
|
|
71
86
|
shorter answer does.
|
|
72
87
|
|
|
73
88
|
- **p50** is what a typical call costs. It is the number to reason about when
|
|
74
|
-
you ask "can the game afford this loop?"
|
|
75
|
-
minute you are about to disclose.
|
|
89
|
+
you ask "can the game afford this loop?".
|
|
76
90
|
- **p95** is what a bad-but-normal call costs: a long answer, a model that
|
|
77
|
-
reasons its way around. It is the
|
|
78
|
-
|
|
79
|
-
mid-session.
|
|
91
|
+
reasons its way around. It is what the **ceiling** has to cover, because a
|
|
92
|
+
ceiling below it cuts the unlucky calls off.
|
|
80
93
|
- **max** is diagnostic. When max sits far above p95, the prompt has an
|
|
81
94
|
unbounded branch in it — usually an unconstrained list or a missing schema.
|
|
82
95
|
Fix the prompt rather than declaring a bigger number.
|
|
83
96
|
|
|
84
|
-
|
|
85
|
-
|
|
86
|
-
|
|
87
|
-
|
|
88
|
-
|
|
89
|
-
|
|
90
|
-
the
|
|
91
|
-
|
|
92
|
-
|
|
97
|
+
Then the run prints what to declare:
|
|
98
|
+
|
|
99
|
+
- **`Declare maxCredits`** — the p95 real cost with the platform fee and the
|
|
100
|
+
**server's own recommended headroom** applied, in whole credits. It also
|
|
101
|
+
covers the answer's **length**: it is never below the smallest ceiling that
|
|
102
|
+
leaves room for the benchmarked answer, and when the length is what set it
|
|
103
|
+
the run says so in one sentence. A ceiling built from cost alone can cut the
|
|
104
|
+
answer off in the game while the bench — run under a larger `--max-credits`
|
|
105
|
+
— never saw it.
|
|
106
|
+
- **`Grant perCallMaxCredits`** — the same over the worst sample, never below
|
|
107
|
+
the line above it. It is emphatically NOT the benchmark's `max` of charged
|
|
108
|
+
credits: that number carries neither the headroom nor the answer's length,
|
|
109
|
+
so a ceiling set from it lands under the call ceiling and the budget's calls
|
|
110
|
+
are refused (`grant_price_unreasonable`).
|
|
111
|
+
- **`Per-call estimate: about … credits per call`** — the measured average,
|
|
112
|
+
fee included, no headroom: a fraction of a credit for a short call. It is the
|
|
113
|
+
honest figure for the disclosure.
|
|
114
|
+
- **`Grant perCallEstimateCredits`** — that average rounded up to a whole
|
|
115
|
+
credit.
|
|
116
|
+
|
|
117
|
+
Declare each verbatim:
|
|
93
118
|
|
|
94
119
|
```ts
|
|
95
|
-
const
|
|
96
|
-
const
|
|
120
|
+
const NPC_CALL_MAX = <Declare maxCredits>; // from `npx genex llm bench`, <date>
|
|
121
|
+
const NPC_GRANT_MAX = <Grant perCallMaxCredits>; // same run — budgets only
|
|
122
|
+
const NPC_CALL_ESTIMATE = <Grant perCallEstimateCredits>; // same run — budgets only
|
|
123
|
+
const NPC_CALL_AVERAGE = <the "about … credits per call" figure>; // same run
|
|
97
124
|
```
|
|
98
125
|
|
|
99
|
-
Write the benchmark date and the model id beside
|
|
100
|
-
person knows what the
|
|
101
|
-
|
|
126
|
+
Write the benchmark date and the model id beside them in `DESIGN.md`, so the
|
|
127
|
+
next person knows what the numbers describe. Do not add a margin of your own on
|
|
128
|
+
top, and do not round them down to look cheaper.
|
|
102
129
|
|
|
103
|
-
|
|
104
|
-
`perCallMaxCoins`. Declare that one verbatim too, and do not work it out by
|
|
105
|
-
hand. It is emphatically NOT the benchmark's `max`: p95 is nearest-rank, so at
|
|
106
|
-
the sample counts a benchmark actually takes, p95 and max are the same figure —
|
|
107
|
-
a ceiling set from `max` therefore lands *below* the recommended price, and
|
|
108
|
-
`requestSpendGrant()` refuses that pair before the request ever leaves the page.
|
|
109
|
-
|
|
110
|
-
The invariant, which the SDK and the server both enforce:
|
|
130
|
+
The invariants, which the SDK and the server both enforce:
|
|
111
131
|
|
|
112
132
|
```
|
|
113
|
-
|
|
133
|
+
perCallEstimateCredits ≤ perCallMaxCredits ≤ the limit the player approves
|
|
134
|
+
maxCredits on each call ≤ perCallMaxCredits
|
|
114
135
|
```
|
|
115
136
|
|
|
116
|
-
|
|
117
|
-
|
|
118
|
-
|
|
137
|
+
A ceiling is not a second price: a call is billed the real cost of the work it
|
|
138
|
+
did plus the fee (a one-time call rounded up to a whole credit), and never past
|
|
139
|
+
its ceiling.
|
|
140
|
+
|
|
141
|
+
**One-time calls round up; budgets do not.** A one-time `generate()` is captured
|
|
142
|
+
rounded UP to a whole credit when it settles, so a call that really costs a
|
|
143
|
+
fraction of a credit still costs the player one whole credit each time. Under a
|
|
144
|
+
budget each call adds its exact fraction and the player's credits move one
|
|
145
|
+
whole credit at a time. That is one more reason a repeated call belongs under a
|
|
146
|
+
budget.
|
|
119
147
|
|
|
120
148
|
## 4. Turn the game's loop into the disclosure
|
|
121
149
|
|
|
@@ -129,27 +157,39 @@ from the loop you wrote, counted honestly:
|
|
|
129
157
|
3. **Multiply, and pick the period that makes the number legible.** Five NPCs at
|
|
130
158
|
one call per thirty seconds is ten calls a minute, so `periodLabel: 'minute'`
|
|
131
159
|
and `estimatedCallsPerPeriod: 10`.
|
|
132
|
-
4. **
|
|
133
|
-
`
|
|
134
|
-
|
|
160
|
+
4. **Credits per period is calls × the measured average, rounded up** —
|
|
161
|
+
`estimatedCreditsPerPeriod: Math.ceil(10 * NPC_CALL_AVERAGE)`. Not the
|
|
162
|
+
ceiling: a call is billed what it uses, so the honest estimate is what calls
|
|
163
|
+
really average.
|
|
135
164
|
|
|
136
165
|
```ts
|
|
166
|
+
perCallMaxCredits: NPC_GRANT_MAX,
|
|
167
|
+
perCallEstimateCredits: NPC_CALL_ESTIMATE,
|
|
137
168
|
disclosure: {
|
|
138
169
|
periodLabel: 'minute',
|
|
139
170
|
estimatedCallsPerPeriod: 10,
|
|
140
|
-
|
|
171
|
+
estimatedCreditsPerPeriod: Math.ceil(10 * NPC_CALL_AVERAGE),
|
|
141
172
|
}
|
|
142
173
|
```
|
|
143
174
|
|
|
144
175
|
The player sees this attributed to your game — "the game estimates about …" —
|
|
145
|
-
beside the platform's own worst case computed from `
|
|
176
|
+
beside the platform's own worst case computed from `perCallMaxCredits`. The two
|
|
146
177
|
being far apart is normal; the estimate being far below what the game really
|
|
147
178
|
does is what breaks trust and burns the grant mid-session.
|
|
148
179
|
|
|
149
|
-
**Batching changes this arithmetic more than any
|
|
180
|
+
**Batching changes this arithmetic more than any tuning can.** Five NPCs
|
|
150
181
|
answered by one call that returns five decisions is two calls a minute, not ten,
|
|
151
|
-
and one
|
|
152
|
-
|
|
182
|
+
and one benchmark of the batched prompt replaces five of the unbatched one. Do
|
|
183
|
+
that before you reach for a cheaper model.
|
|
184
|
+
|
|
185
|
+
## Old code
|
|
186
|
+
|
|
187
|
+
`estimateCoins`, `perCallMaxCoins`, `perCallEstimateCoins` and
|
|
188
|
+
`disclosure.estimatedCoinsPerPeriod` are the names from before credits. The
|
|
189
|
+
server still accepts them and reads them as the same numbers in credits, but
|
|
190
|
+
they are deprecated: rename them to `maxCredits`, `perCallMaxCredits`,
|
|
191
|
+
`perCallEstimateCredits` and `estimatedCreditsPerPeriod` when you touch the
|
|
192
|
+
code, and re-benchmark — a number measured as a coin price is not a ceiling.
|
|
153
193
|
|
|
154
194
|
## Worked example
|
|
155
195
|
|
|
@@ -160,20 +200,21 @@ speech). The prompt carries the NPC's memory and the last two player lines.
|
|
|
160
200
|
1. Freeze the prompt with a full memory and a long player line. Write
|
|
161
201
|
`./answer.schema.json` with the two fields.
|
|
162
202
|
2. `npx genex llm bench "<that prompt>" --schema ./answer.schema.json
|
|
163
|
-
--samples 3 --max-
|
|
164
|
-
3. Read: p50 `<p50>`, p95 `<p95>`, max `<max>`, and the
|
|
165
|
-
|
|
166
|
-
`
|
|
167
|
-
|
|
168
|
-
|
|
169
|
-
|
|
170
|
-
|
|
203
|
+
--samples 3 --max-credits <ceiling> --user-approved`.
|
|
204
|
+
3. Read: p50 `<p50>`, p95 `<p95>`, max `<max>`, and the printed lines —
|
|
205
|
+
`Declare maxCredits: <call ceiling>`, `Grant perCallMaxCredits: <grant ceiling>`,
|
|
206
|
+
`Per-call estimate: about <average> credits per call` and
|
|
207
|
+
`Grant perCallEstimateCredits: <estimate>`.
|
|
208
|
+
4. Declare `NPC_CALL_MAX`, `NPC_GRANT_MAX`, `NPC_CALL_ESTIMATE` and
|
|
209
|
+
`NPC_CALL_AVERAGE`, each verbatim from the line that printed it.
|
|
210
|
+
5. Unbatched, the loop is ten calls a minute, so the disclosure is `10` calls
|
|
211
|
+
and `Math.ceil(10 * NPC_CALL_AVERAGE)` credits per minute.
|
|
171
212
|
6. Batch the five NPCs into one call — benchmark the batched prompt separately,
|
|
172
213
|
because it is a different prompt — and the disclosure becomes two calls a
|
|
173
|
-
minute at the batched
|
|
214
|
+
minute at the batched average.
|
|
174
215
|
7. Re-run steps 1–4 whenever the prompt, the schema or the model changes. A
|
|
175
|
-
prompt edit
|
|
216
|
+
prompt edit changes the ceiling; treat it like one.
|
|
176
217
|
|
|
177
218
|
Every `<placeholder>` above is read off your own benchmark run. None of these
|
|
178
219
|
figures is a platform constant, and none of them should be copied from another
|
|
179
|
-
game — a different prompt has a different
|
|
220
|
+
game — a different prompt has a different cost.
|
|
@@ -13,11 +13,12 @@ model has to run for the player.
|
|
|
13
13
|
## What the platform does
|
|
14
14
|
|
|
15
15
|
A game hosted on Genex can call a language model from inside the running game
|
|
16
|
-
and get back text or JSON. **The player pays and the player approves**:
|
|
17
|
-
from their Genex
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
16
|
+
and get back text or JSON. **The player pays and the player approves**: credits
|
|
17
|
+
from their Genex account, billed as each call is used and never past the
|
|
18
|
+
per-call ceiling the game declares, or their own Claude / ChatGPT subscription,
|
|
19
|
+
chosen on an approval sheet Genex draws that the game cannot render, skin or
|
|
20
|
+
bypass. Either one call at a time, or one standing budget the player approves
|
|
21
|
+
once and the game then spends against without another popup. The game holds no provider
|
|
21
22
|
key, sees no credential, and never talks to a model vendor.
|
|
22
23
|
|
|
23
24
|
That is the whole reason this is a platform feature and not something you wire
|
|
@@ -40,9 +41,9 @@ Any of these means a model at PLAY time, however it is phrased:
|
|
|
40
41
|
Say this and stop:
|
|
41
42
|
|
|
42
43
|
> A model running while people play is built into the Genex platform — the
|
|
43
|
-
> player pays, with Genex
|
|
44
|
-
> approves it on a Genex sheet; your game just calls
|
|
45
|
-
> that way?
|
|
44
|
+
> player pays, with their Genex credits or their own Claude/ChatGPT
|
|
45
|
+
> subscription, and approves it on a Genex sheet; your game just calls
|
|
46
|
+
> `generate()`. Want it that way?
|
|
46
47
|
|
|
47
48
|
Wait for the answer. Do not start building either version first, and do not
|
|
48
49
|
expand the offer into a pitch — one line, one question.
|
|
@@ -64,7 +65,8 @@ expand the offer into a pitch — one line, one question.
|
|
|
64
65
|
ask before you do if you have not already.
|
|
65
66
|
3. Load `$genex-llm-in-games` — it arrives with the conversion and owns the
|
|
66
67
|
build: the SDK surface, one-time calls versus a standing budget, measuring
|
|
67
|
-
the
|
|
68
|
+
the per-call ceiling before the game declares it, and honest handling of
|
|
69
|
+
every refusal.
|
|
68
70
|
|
|
69
71
|
**The game has to be a static browser build.** A hosted game is files served
|
|
70
72
|
from the edge; there is no server of yours inside it. So the familiar pattern —
|
|
@@ -38,7 +38,7 @@ update, so update immediately.)
|
|
|
38
38
|
Run exactly the command the nudge printed, from the game project root:
|
|
39
39
|
|
|
40
40
|
```bash
|
|
41
|
-
npm i -D @genex-ai/cli-demo@
|
|
41
|
+
npm i -D @genex-ai/cli-demo@dev # the genex CLI (a dev dependency)
|
|
42
42
|
npm i @genex-ai/embed-sdk@latest # identity/saves SDK (ships inside the game)
|
|
43
43
|
npm i @genex-ai/multiplayer@latest # multiplayer SDK (only if the game uses it)
|
|
44
44
|
```
|