@genex-ai/cli-demo 1.36.0-dev.773 → 1.36.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/{blender-mcp-ZRCMYTXI.js → blender-mcp-NRJ5FI22.js} +2 -2
- package/dist/{blender-serve-EH2MOW3U.js → blender-serve-66OHC4ML.js} +1 -1
- package/dist/{chunk-Y6WO5FJZ.js → chunk-GOP5FRSJ.js} +2 -2
- package/dist/{chunk-K4EHXOZK.js → chunk-QI7FIYBY.js} +1 -1
- package/dist/index.js +216 -448
- package/package.json +1 -1
- package/templates/skills/genex-game-director/SKILL.md +1 -1
- package/templates/skills/genex-getting-started/SKILL.md +2 -2
- package/templates/skills/genex-llm-in-games/SKILL.md +72 -209
- package/templates/skills/genex-llm-in-games/references/pricing.md +85 -126
- package/templates/skills/genex-tool-llm/SKILL.md +9 -11
- package/templates/skills/genex-updates/SKILL.md +1 -1
|
@@ -1,28 +1,16 @@
|
|
|
1
1
|
# Calibrate, then declare
|
|
2
2
|
|
|
3
|
-
|
|
4
|
-
|
|
5
|
-
|
|
6
|
-
|
|
7
|
-
|
|
8
|
-
|
|
9
|
-
|
|
10
|
-
|
|
11
|
-
|
|
12
|
-
|
|
13
|
-
|
|
14
|
-
- **It sizes the answer.** Each call's room to answer is funded from its
|
|
15
|
-
ceiling, so a ceiling below what the answer needs cuts it off — and a cut-off
|
|
16
|
-
attempt is still billed for the work it did.
|
|
17
|
-
|
|
18
|
-
There is one honest way to pick it: run the real prompt on your own credits,
|
|
19
|
-
read what it really cost, and declare the numbers the benchmark prints.
|
|
20
|
-
|
|
21
|
-
Never derive them from a model vendor's published rates. Three layers sit
|
|
22
|
-
between that rate and what the player is billed — the provider's cost for YOUR
|
|
23
|
-
prompt, the platform fee, and the headroom a ceiling needs — and only the
|
|
24
|
-
benchmark sees the first. The other two come from the server; a number worked
|
|
25
|
-
out from the vendor's page drops at least one of them.
|
|
3
|
+
`estimateCoins` is the **fixed price of a started attempt**, not a guess about
|
|
4
|
+
one. The platform charges exactly what you declare, whether the call succeeds,
|
|
5
|
+
fails, is canceled, or stops at its own budget. There is one honest way to pick
|
|
6
|
+
it: run the real prompt on your own coins, read what it charged, and declare the
|
|
7
|
+
number the benchmark recommends.
|
|
8
|
+
|
|
9
|
+
Never derive it from a model vendor's published rates. Three layers sit between
|
|
10
|
+
that rate and what the player is charged — the provider's cost, the platform's
|
|
11
|
+
tariff, and the headroom a declared price needs — and only the benchmark sees
|
|
12
|
+
all three. A number worked out from the vendor's page silently drops the middle
|
|
13
|
+
layer and under-prices every call you will ever make.
|
|
26
14
|
|
|
27
15
|
## 1. Freeze the prompt first
|
|
28
16
|
|
|
@@ -34,7 +22,7 @@ that the real game then cannot fund.
|
|
|
34
22
|
If the feature returns structured data, write the schema to a file now
|
|
35
23
|
(`./answer.schema.json`) and benchmark with it. Schema'd JSON and free text do
|
|
36
24
|
not cost the same, and shipping without a schema is the commonest way a call is
|
|
37
|
-
|
|
25
|
+
charged for an unusable result.
|
|
38
26
|
|
|
39
27
|
## 2. Run the benchmark
|
|
40
28
|
|
|
@@ -45,105 +33,89 @@ npx genex llm bench "<the frozen prompt, one real example filled in>" \
|
|
|
45
33
|
--model <id from the line above> \
|
|
46
34
|
--schema ./answer.schema.json \
|
|
47
35
|
--samples 3 \
|
|
48
|
-
--max-
|
|
36
|
+
--max-coins <n> --user-approved
|
|
49
37
|
```
|
|
50
38
|
|
|
51
|
-
- It spends **your**
|
|
52
|
-
credential. Nothing here runs in a shipped build.
|
|
53
|
-
|
|
54
|
-
- `--max-credits <n> --user-approved` is a hard gate, refused before any network
|
|
39
|
+
- It spends **your** coins, on the development lane, through the CLI's own
|
|
40
|
+
credential. Nothing here runs in a shipped build.
|
|
41
|
+
- `--max-coins <n> --user-approved` is a hard gate, refused before any network
|
|
55
42
|
call. That is the same shape as every other spend approval in the CLI: you
|
|
56
|
-
state the ceiling for this run, out loud, once.
|
|
57
|
-
cover one attempt at it, or the run is refused before it starts.
|
|
58
|
-
`--max-coins` is the old name of the same flag and still works.
|
|
43
|
+
state the ceiling for this run, out loud, once.
|
|
59
44
|
- `--samples 3` is the floor. The same prompt costs different amounts on
|
|
60
45
|
different runs, because the model's own output length varies.
|
|
61
46
|
- `npx genex llm price` re-prints the last run's recommendation without spending
|
|
62
|
-
anything again.
|
|
63
|
-
request to re-run.
|
|
47
|
+
anything again.
|
|
64
48
|
|
|
65
49
|
## 3. Read the output
|
|
66
50
|
|
|
67
|
-
Each sample prints
|
|
68
|
-
|
|
69
|
-
|
|
70
|
-
numbers; a sample the provider refused at its door
|
|
71
|
-
ran no inference, cost nothing, and prints the
|
|
72
|
-
row — on a 401 or 403 that is the stand's
|
|
73
|
-
model, which is the operator's to fix. A
|
|
74
|
-
never started and ends the run: the
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
-
|
|
51
|
+
Each sample prints what it actually charged. The aggregate prints p50, p95 and
|
|
52
|
+
max of the charged coins **over the samples that succeeded and settled**, plus
|
|
53
|
+
one recommendation. A sample that failed is printed with its code and stays
|
|
54
|
+
out of the numbers; a sample the provider refused at its door
|
|
55
|
+
(`provider_http_<status>`) ran no inference, cost nothing, and prints the
|
|
56
|
+
provider's own message under its row — on a 401 or 403 that is the stand's
|
|
57
|
+
provider configuration refusing the model, which is the operator's to fix. A
|
|
58
|
+
sample refused as `generation_limit` never started and ends the run: the
|
|
59
|
+
account's three ad-hoc calls are open or recently stopped with a pending
|
|
60
|
+
bill. `npx genex llm status` lists them with what each holds;
|
|
61
|
+
`npx genex llm cancel <id>` stops an active one; a stopped one frees on its
|
|
62
|
+
own once its bill resolves, and stops holding a slot ten minutes after
|
|
63
|
+
dispatch. Re-running the bench into the same refusal spends nothing and
|
|
79
64
|
learns nothing.
|
|
80
65
|
|
|
81
66
|
A sample the model had to stop writing (`provider_token_limit`) was cut off at
|
|
82
|
-
your `--max-
|
|
83
|
-
unknown — so the run recommends no
|
|
84
|
-
`--max-
|
|
85
|
-
answer is longer than one call there may produce; no
|
|
67
|
+
your `--max-coins`: it was charged, it is not a sample, and its real length is
|
|
68
|
+
unknown — so the run recommends no price and asks you to re-run with a higher
|
|
69
|
+
`--max-coins`. When the stand itself cannot hold the answer, the run says the
|
|
70
|
+
answer is longer than one call there may produce; no price fixes that, a
|
|
86
71
|
shorter answer does.
|
|
87
72
|
|
|
88
73
|
- **p50** is what a typical call costs. It is the number to reason about when
|
|
89
|
-
you ask "can the game afford this loop?"
|
|
74
|
+
you ask "can the game afford this loop?" — multiply it by the calls per
|
|
75
|
+
minute you are about to disclose.
|
|
90
76
|
- **p95** is what a bad-but-normal call costs: a long answer, a model that
|
|
91
|
-
reasons its way around. It is
|
|
92
|
-
|
|
77
|
+
reasons its way around. It is the number to **declare**, because a declared
|
|
78
|
+
price below it means the unlucky calls cannot fund themselves and get refused
|
|
79
|
+
mid-session.
|
|
93
80
|
- **max** is diagnostic. When max sits far above p95, the prompt has an
|
|
94
81
|
unbounded branch in it — usually an unconstrained list or a missing schema.
|
|
95
82
|
Fix the prompt rather than declaring a bigger number.
|
|
96
83
|
|
|
97
|
-
|
|
98
|
-
|
|
99
|
-
|
|
100
|
-
|
|
101
|
-
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
105
|
-
|
|
106
|
-
- **`Grant perCallMaxCredits`** — the same over the worst sample, never below
|
|
107
|
-
the line above it. It is emphatically NOT the benchmark's `max` of charged
|
|
108
|
-
credits: that number carries neither the headroom nor the answer's length,
|
|
109
|
-
so a ceiling set from it lands under the call ceiling and the budget's calls
|
|
110
|
-
are refused (`grant_price_unreasonable`).
|
|
111
|
-
- **`Per-call estimate: about … credits per call`** — the measured average,
|
|
112
|
-
fee included, no headroom: a fraction of a credit for a short call. It is the
|
|
113
|
-
honest figure for the disclosure.
|
|
114
|
-
- **`Grant perCallEstimateCredits`** — that average rounded up to a whole
|
|
115
|
-
credit.
|
|
116
|
-
|
|
117
|
-
Declare each verbatim:
|
|
84
|
+
The recommendation line already applies the **server's own recommended
|
|
85
|
+
headroom** on top of p95. It also covers the answer's **length**: the declared
|
|
86
|
+
price decides how long each call's answer may be, because the room to answer is
|
|
87
|
+
funded from it, so a price built from charged coins alone can cut the answer
|
|
88
|
+
off in the game while the bench — run under a larger `--max-coins` — never saw
|
|
89
|
+
it. The recommendation is never below the smallest price that leaves room for
|
|
90
|
+
the benchmarked answer, and when the length is what set it the run says so in
|
|
91
|
+
one sentence. Either way, declare that figure verbatim — never the bare charged
|
|
92
|
+
number:
|
|
118
93
|
|
|
119
94
|
```ts
|
|
120
|
-
const
|
|
121
|
-
const
|
|
122
|
-
const NPC_CALL_ESTIMATE = <Grant perCallEstimateCredits>; // same run — budgets only
|
|
123
|
-
const NPC_CALL_AVERAGE = <the "about … credits per call" figure>; // same run
|
|
95
|
+
const NPC_CALL_PRICE = <the recommended figure>; // from `npx genex llm bench`, <date>
|
|
96
|
+
const NPC_CALL_CEILING = <the printed ceiling>; // same run — grants only
|
|
124
97
|
```
|
|
125
98
|
|
|
126
|
-
Write the benchmark date and the model id beside
|
|
127
|
-
|
|
128
|
-
|
|
99
|
+
Write the benchmark date and the model id beside it in `DESIGN.md`, so the next
|
|
100
|
+
person knows what the number describes. Do not add a margin of your own on top
|
|
101
|
+
of the recommendation, and do not round it down to look cheaper.
|
|
129
102
|
|
|
130
|
-
|
|
103
|
+
For a grant, the run prints a **second** number beside the price: the ceiling,
|
|
104
|
+
`perCallMaxCoins`. Declare that one verbatim too, and do not work it out by
|
|
105
|
+
hand. It is emphatically NOT the benchmark's `max`: p95 is nearest-rank, so at
|
|
106
|
+
the sample counts a benchmark actually takes, p95 and max are the same figure —
|
|
107
|
+
a ceiling set from `max` therefore lands *below* the recommended price, and
|
|
108
|
+
`requestSpendGrant()` refuses that pair before the request ever leaves the page.
|
|
109
|
+
|
|
110
|
+
The invariant, which the SDK and the server both enforce:
|
|
131
111
|
|
|
132
112
|
```
|
|
133
|
-
|
|
134
|
-
maxCredits on each call ≤ perCallMaxCredits
|
|
113
|
+
perCallEstimateCoins ≤ perCallMaxCoins ≤ the limit the player approves
|
|
135
114
|
```
|
|
136
115
|
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
|
|
140
|
-
|
|
141
|
-
**One-time calls round up; budgets do not.** A one-time `generate()` is captured
|
|
142
|
-
rounded UP to a whole credit when it settles, so a call that really costs a
|
|
143
|
-
fraction of a credit still costs the player one whole credit each time. Under a
|
|
144
|
-
budget each call adds its exact fraction and the player's credits move one
|
|
145
|
-
whole credit at a time. That is one more reason a repeated call belongs under a
|
|
146
|
-
budget.
|
|
116
|
+
The ceiling is the price's room to be wrong, not a second price: no call is ever
|
|
117
|
+
charged more than the price it declares, and the platform separately refuses any
|
|
118
|
+
declared price out of proportion to what the model could really cost.
|
|
147
119
|
|
|
148
120
|
## 4. Turn the game's loop into the disclosure
|
|
149
121
|
|
|
@@ -157,39 +129,27 @@ from the loop you wrote, counted honestly:
|
|
|
157
129
|
3. **Multiply, and pick the period that makes the number legible.** Five NPCs at
|
|
158
130
|
one call per thirty seconds is ten calls a minute, so `periodLabel: 'minute'`
|
|
159
131
|
and `estimatedCallsPerPeriod: 10`.
|
|
160
|
-
4. **
|
|
161
|
-
`
|
|
162
|
-
|
|
163
|
-
really average.
|
|
132
|
+
4. **Coins per period is calls × the declared price** —
|
|
133
|
+
`estimatedCoinsPerPeriod: 10 * NPC_CALL_PRICE`. Not the p50, not a hope: the
|
|
134
|
+
price you declare is the price charged.
|
|
164
135
|
|
|
165
136
|
```ts
|
|
166
|
-
perCallMaxCredits: NPC_GRANT_MAX,
|
|
167
|
-
perCallEstimateCredits: NPC_CALL_ESTIMATE,
|
|
168
137
|
disclosure: {
|
|
169
138
|
periodLabel: 'minute',
|
|
170
139
|
estimatedCallsPerPeriod: 10,
|
|
171
|
-
|
|
140
|
+
estimatedCoinsPerPeriod: 10 * NPC_CALL_PRICE,
|
|
172
141
|
}
|
|
173
142
|
```
|
|
174
143
|
|
|
175
144
|
The player sees this attributed to your game — "the game estimates about …" —
|
|
176
|
-
beside the platform's own worst case computed from `
|
|
145
|
+
beside the platform's own worst case computed from `perCallMaxCoins`. The two
|
|
177
146
|
being far apart is normal; the estimate being far below what the game really
|
|
178
147
|
does is what breaks trust and burns the grant mid-session.
|
|
179
148
|
|
|
180
|
-
**Batching changes this arithmetic more than any tuning can.** Five NPCs
|
|
149
|
+
**Batching changes this arithmetic more than any price tuning can.** Five NPCs
|
|
181
150
|
answered by one call that returns five decisions is two calls a minute, not ten,
|
|
182
|
-
and one
|
|
183
|
-
that before you reach for a cheaper model.
|
|
184
|
-
|
|
185
|
-
## Old code
|
|
186
|
-
|
|
187
|
-
`estimateCoins`, `perCallMaxCoins`, `perCallEstimateCoins` and
|
|
188
|
-
`disclosure.estimatedCoinsPerPeriod` are the names from before credits. The
|
|
189
|
-
server still accepts them and reads them as the same numbers in credits, but
|
|
190
|
-
they are deprecated: rename them to `maxCredits`, `perCallMaxCredits`,
|
|
191
|
-
`perCallEstimateCredits` and `estimatedCreditsPerPeriod` when you touch the
|
|
192
|
-
code, and re-benchmark — a number measured as a coin price is not a ceiling.
|
|
151
|
+
and one benchmarked price for the batched prompt replaces five of the unbatched
|
|
152
|
+
one. Do that before you reach for a cheaper model.
|
|
193
153
|
|
|
194
154
|
## Worked example
|
|
195
155
|
|
|
@@ -200,21 +160,20 @@ speech). The prompt carries the NPC's memory and the last two player lines.
|
|
|
200
160
|
1. Freeze the prompt with a full memory and a long player line. Write
|
|
201
161
|
`./answer.schema.json` with the two fields.
|
|
202
162
|
2. `npx genex llm bench "<that prompt>" --schema ./answer.schema.json
|
|
203
|
-
--samples 3 --max-
|
|
204
|
-
3. Read: p50 `<p50>`, p95 `<p95>`, max `<max>`, and the printed
|
|
205
|
-
|
|
206
|
-
`
|
|
207
|
-
|
|
208
|
-
|
|
209
|
-
|
|
210
|
-
|
|
211
|
-
and `Math.ceil(10 * NPC_CALL_AVERAGE)` credits per minute.
|
|
163
|
+
--samples 3 --max-coins <ceiling> --user-approved`.
|
|
164
|
+
3. Read: p50 `<p50>`, p95 `<p95>`, max `<max>`, and the two printed
|
|
165
|
+
declarations — `Declare estimateCoins: <recommended>` and
|
|
166
|
+
`Grant perCallMaxCoins: <ceiling>`.
|
|
167
|
+
4. Declare `NPC_CALL_PRICE = <recommended>` and
|
|
168
|
+
`NPC_CALL_CEILING = <ceiling>`, each verbatim from the line that printed it.
|
|
169
|
+
5. Unbatched, the loop is ten calls a minute, so the disclosure is
|
|
170
|
+
`10` and `10 * <recommended>` coins per minute.
|
|
212
171
|
6. Batch the five NPCs into one call — benchmark the batched prompt separately,
|
|
213
172
|
because it is a different prompt — and the disclosure becomes two calls a
|
|
214
|
-
minute at the batched
|
|
173
|
+
minute at the batched price.
|
|
215
174
|
7. Re-run steps 1–4 whenever the prompt, the schema or the model changes. A
|
|
216
|
-
prompt edit
|
|
175
|
+
prompt edit is a price change; treat it like one.
|
|
217
176
|
|
|
218
177
|
Every `<placeholder>` above is read off your own benchmark run. None of these
|
|
219
178
|
figures is a platform constant, and none of them should be copied from another
|
|
220
|
-
game — a different prompt has a different
|
|
179
|
+
game — a different prompt has a different price.
|
|
@@ -13,12 +13,11 @@ model has to run for the player.
|
|
|
13
13
|
## What the platform does
|
|
14
14
|
|
|
15
15
|
A game hosted on Genex can call a language model from inside the running game
|
|
16
|
-
and get back text or JSON. **The player pays and the player approves**:
|
|
17
|
-
from their Genex
|
|
18
|
-
|
|
19
|
-
|
|
20
|
-
|
|
21
|
-
once and the game then spends against without another popup. The game holds no provider
|
|
16
|
+
and get back text or JSON. **The player pays and the player approves**: coins
|
|
17
|
+
from their Genex wallet, or their own Claude / ChatGPT subscription, chosen on
|
|
18
|
+
an approval sheet Genex draws that the game cannot render, skin or bypass.
|
|
19
|
+
Either one call at a time, or one standing budget the player approves once and
|
|
20
|
+
the game then spends against without another popup. The game holds no provider
|
|
22
21
|
key, sees no credential, and never talks to a model vendor.
|
|
23
22
|
|
|
24
23
|
That is the whole reason this is a platform feature and not something you wire
|
|
@@ -41,9 +40,9 @@ Any of these means a model at PLAY time, however it is phrased:
|
|
|
41
40
|
Say this and stop:
|
|
42
41
|
|
|
43
42
|
> A model running while people play is built into the Genex platform — the
|
|
44
|
-
> player pays, with
|
|
45
|
-
>
|
|
46
|
-
>
|
|
43
|
+
> player pays, with Genex coins or their own Claude/ChatGPT subscription, and
|
|
44
|
+
> approves it on a Genex sheet; your game just calls `generate()`. Want it
|
|
45
|
+
> that way?
|
|
47
46
|
|
|
48
47
|
Wait for the answer. Do not start building either version first, and do not
|
|
49
48
|
expand the offer into a pitch — one line, one question.
|
|
@@ -65,8 +64,7 @@ expand the offer into a pitch — one line, one question.
|
|
|
65
64
|
ask before you do if you have not already.
|
|
66
65
|
3. Load `$genex-llm-in-games` — it arrives with the conversion and owns the
|
|
67
66
|
build: the SDK surface, one-time calls versus a standing budget, measuring
|
|
68
|
-
the
|
|
69
|
-
every refusal.
|
|
67
|
+
the price before the game declares it, and honest handling of every refusal.
|
|
70
68
|
|
|
71
69
|
**The game has to be a static browser build.** A hosted game is files served
|
|
72
70
|
from the edge; there is no server of yours inside it. So the familiar pattern —
|
|
@@ -38,7 +38,7 @@ update, so update immediately.)
|
|
|
38
38
|
Run exactly the command the nudge printed, from the game project root:
|
|
39
39
|
|
|
40
40
|
```bash
|
|
41
|
-
npm i -D @genex-ai/cli-demo@
|
|
41
|
+
npm i -D @genex-ai/cli-demo@latest # the genex CLI (a dev dependency)
|
|
42
42
|
npm i @genex-ai/embed-sdk@latest # identity/saves SDK (ships inside the game)
|
|
43
43
|
npm i @genex-ai/multiplayer@latest # multiplayer SDK (only if the game uses it)
|
|
44
44
|
```
|