@genex-ai/cli-demo 1.35.0-dev.746 → 1.35.1-dev.748

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/dist/index.js CHANGED
@@ -23564,8 +23564,9 @@ async function reportModels(apiUrl, token, opts, log, toolsOnly = false) {
23564
23564
  );
23565
23565
  for (const m of shown) {
23566
23566
  const plan = m.personalPlan ? ` \xB7 personal plan: ${m.personalPlan}` : "";
23567
+ const kind = m.kind === "typesafe" ? " \xB7 judge (classifier): judge() under a budget, not generate() or bench" : "";
23567
23568
  log.plain(
23568
- ` ${c.cyan(m.id)} ${m.label} \u2014 ${usdPerMillion(m.inputUsdPerMillion)} in / ${usdPerMillion(m.outputUsdPerMillion)} out per million tokens${plan}`
23569
+ ` ${c.cyan(m.id)} ${m.label} \u2014 ${usdPerMillion(m.inputUsdPerMillion)} in / ${usdPerMillion(m.outputUsdPerMillion)} out per million tokens${plan}${kind}`
23569
23570
  );
23570
23571
  }
23571
23572
  if (hiddenCount > 0) {
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@genex-ai/cli-demo",
3
- "version": "1.35.0-dev.746",
3
+ "version": "1.35.1-dev.748",
4
4
  "description": "Set up your project's agent workspace (.claude/.codex/.cursor in the game folder), authorize, create a game project, generate AI assets, and publish (genex CLI).",
5
5
  "type": "module",
6
6
  "bin": {
@@ -25,6 +25,17 @@ Two modes. Picking the wrong one is the most expensive mistake on this lane:
25
25
  A popup per NPC turn is not a feature, it is an interruption. More than a call
26
26
  or two per session means a grant.
27
27
 
28
+ **Deciding, not writing? Use the judge.** When the game needs to KNOW something
29
+ — is this NPC lying, which of five moods is the player in, how threatening is
30
+ the scene on a scale — and not to produce words, ask the grant's classifier with
31
+ `judge()` instead of asking `generate()` for a label. It answers typed questions
32
+ (`choice`, `score`, `noul`) with calibrated probabilities in about a third of a
33
+ second, writes no prose and cannot drift out of your options. It is grant-only
34
+ and billed as used: each call is a small fraction of a coin and accrues on the
35
+ budget, so a check every turn is affordable where a generate every turn is not.
36
+ `generate()` for words, `judge()` for decisions; a feature often wants both —
37
+ the judge picks the branch, generate writes the line.
38
+
28
39
  ## Step 0 — is the lane live on this stand?
29
40
 
30
41
  ```bash
@@ -67,9 +78,17 @@ identity must be resolved first — `$genex-threejs-embed-auth`.
67
78
  (`billingStatus`, `reservedCoins`, `chargedCoins`, their display-USD twins).
68
79
  - `requestSpendGrant({ models, perCallMaxCoins, perCallEstimateCoins,
69
80
  disclosure: { periodLabel, estimatedCallsPerPeriod, estimatedCoinsPerPeriod },
70
- maxConcurrent?, maxCallsPerMinute?, allowExternal?, idempotencyKey? })` →
81
+ maxConcurrent?, maxCallsPerMinute?, allowExternal?, judge?: { enabled,
82
+ estimatedCallsPerPeriod }, idempotencyKey? })` →
71
83
  `{ status, grantId, … the limits the player approved }`; status is `active` |
72
84
  `canceled` | `expired` | `failed` | `pending`.
85
+ - `judge({ grantId, state, questions, idempotencyKey?, timeoutMs? })` →
86
+ `{ status, generationId, answers, error, billing }` — the grant's classifier,
87
+ answered in the same response. `questions` is 1–64 named `choice`
88
+ (`{ instructions, criteria: { option: description } }`), `score`
89
+ (`{ instructions, criteria: [levels, lowest first] }`) or `noul`
90
+ (`{ instructions }`); each answer has its question's type, and `noul` is a
91
+ probability — pick your own threshold. Needs a budget requested with `judge`.
73
92
  - `getSpendGrant(grantId)` — live state and counters; the ONE source for an
74
93
  in-game budget readout.
75
94
  - `stopSpendGrant(grantId)` — the game's own stop door. Prospective: no further
@@ -188,6 +207,15 @@ arithmetic into `DESIGN.md` beside the feature. The player sees your estimate
188
207
  attributed to the game, beside the platform's own worst case; an estimate that
189
208
  is transparently low is a grant that dies mid-session.
190
209
 
210
+ **Using the judge too?** Add `judge: { enabled: true, estimatedCallsPerPeriod }`
211
+ to the same request — the checks you really expect per `periodLabel`, from the
212
+ same loop arithmetic. The sheet tells the player the game also uses a fast
213
+ classifier billed as used, with the platform's own worst case beside your count.
214
+ A judge-only feature passes `models: []`. A budget with the judge is coin-funded
215
+ only, and judge calls share the budget's `maxConcurrent` and
216
+ `maxCallsPerMinute`: batch questions into one call (up to 64) rather than one
217
+ call per question.
218
+
191
219
  **Then keep the burn low, because you wrote the loop:** batch those five NPCs
192
220
  into ONE call returning five decisions, cache a decision until the situation
193
221
  that caused it changes, pick the cheapest model that passes your own check, and
@@ -285,6 +313,7 @@ These are source contracts, not a claim that every stand runs this lane —
285
313
  - [ ] A bench refused as `generation_limit` was answered with `npx genex llm status`, never a retry
286
314
  - [ ] `generate()` / `requestSpendGrant()` is the first statement of a click handler
287
315
  - [ ] A repeated-call feature uses a grant; a one-off uses `generate()`
316
+ - [ ] A decision (a label, a yes/no, a level) is a `judge()` question under the grant, not a `generate()` asked for a word
288
317
  - [ ] Disclosure numbers derive from the real loop and are written in `DESIGN.md`
289
318
  - [ ] Calls are batched and cached; nothing fires on an invisible timer
290
319
  - [ ] Every grant-ending code has in-fiction copy and a playable fallback
@@ -316,6 +345,15 @@ cause rather than showing it for both.
316
345
  **`grant_price_unreasonable`** — the declared per-call price is far above what
317
346
  that prompt can cost on that model. Re-benchmark and declare what it prints.
318
347
 
348
+ **`judge_not_enabled`** — the budget was approved without the judge. Request a
349
+ new one with `judge` from the next deliberate click; never auto-renew.
350
+ **`judge_requires_grant`** — the judge's model was sent to `generate()` with no
351
+ budget; the judge is grant-only and is called with `judge()`.
352
+ **`judge_period_limit`** — the checks reached the ceiling the player approved for
353
+ the period (your own `estimatedCallsPerPeriod` at the largest size): wait, and
354
+ if it keeps happening your declared rate is too low — fix the loop or the
355
+ estimate, never retry in a tight loop.
356
+
319
357
  **`generation_limit`** — three ad-hoc calls are already in flight for this
320
358
  account, or recently stopped with their bill still pending; a pending call
321
359
  stops counting ten minutes after it was dispatched. From the bench: run