@genex-ai/cli-demo 1.35.2-dev.749 → 1.35.3-dev.750
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/{blender-mcp-HKT7N5BQ.js → blender-mcp-ATHMMRDU.js} +1 -1
- package/dist/{chunk-FYHYZYBC.js → chunk-2SCXGPZY.js} +1 -1
- package/dist/index.js +444 -213
- package/package.json +1 -1
- package/templates/skills/genex-game-director/SKILL.md +1 -1
- package/templates/skills/genex-llm-in-games/SKILL.md +127 -68
- package/templates/skills/genex-llm-in-games/references/pricing.md +126 -85
- package/templates/skills/genex-tool-llm/SKILL.md +11 -9
package/dist/index.js
CHANGED
|
@@ -42,7 +42,7 @@ import {
|
|
|
42
42
|
writeSecretFile,
|
|
43
43
|
writeUserToken,
|
|
44
44
|
writeWorkspace
|
|
45
|
-
} from "./chunk-
|
|
45
|
+
} from "./chunk-2SCXGPZY.js";
|
|
46
46
|
import {
|
|
47
47
|
CLI_CHANNEL,
|
|
48
48
|
DEFAULT_API_URL,
|
|
@@ -902,7 +902,7 @@ Important note: put soul into your creations, with many details and love. Aim to
|
|
|
902
902
|
21. A turn ends in exactly one of two ways: on a question the player must answer before the next step can be chosen, or on a handoff \u2014 the draft link, what changed, what to try, and what is still open. Never close a turn on a promise. "I'll keep building", "while I keep working", "adding that now" are things you say and then DO before you stop; if you are stopping, say that you have stopped and what remains. While the Build plan's \`Now:\` line has open work and no answer is needed from the player, the turn is not over: preview, hand off, and start the next milestone in the same turn, until the requested outcome is reached or the platform's own budget ends the session. If a stop was forced on you anyway, the first line of your next turn is the promise you left and where it stands.
|
|
903
903
|
22. **What the player sees must read as the thing it is**, in this game's own style: a building as that building, a person as a person, a prop as that prop, a surface as its material \u2014 and a coloured box, a capsule, or a flat grey block standing in for one of them is never its finished version. Generation is the default route to that bar wherever code will not honestly reach it: the player's character, whatever the request names, whatever the player walks up to, enters, or interacts with, and the music and sound underneath \u2014 reach for the belt on your own judgment, without stopping to ask; a status line saying what you queued is enough, and \`--no-wait\` keeps you building while it lands. Procedural code is a first-class engine for what is structural, repeated, distant, or parametric \u2014 terrain, sky, fences, paving, modular kits, filler \u2014 and for anything else only when the result meets the same bar and you have looked at it in a capture before its row says \`landed\`. Start with one reusable implementation per distinct required gameplay or visual role. Background populations default to reusable procedural bodies and motion or compatible existing rigged assets. Catalog clips need a compatible skeleton and may be paid. Reserve paid custom bodies for the player and characters examined or interacted with closely, while honoring explicit user requests. Add variants only for an unmet requirement; unspent allowance is not unfinished work. Allowance questions use timeoutPolicy no-consent: silence does not raise the allowance. Record role, reuse and completion criteria in the existing Assets table. Decide per object, not per category, and write the route on each Assets row in \`DESIGN.md\` \u2014 the paid flow for generated pieces, the procedural flow with its local path for code-built ones. An Assets table with no generation in it is a decision, not a default: record it as \`Generation: none \u2014 <why code alone reaches the bar here>\` or generate.
|
|
904
904
|
23. **The screenshots you take to look at your own work go in \`.genex/scratch/\`.** Captures, render comparisons, before/after strips, traces and metrics dumps are how you SEE what you built \u2014 they are not part of what you built, and \`genex preview\` pushes the whole folder to the game's source, every time, forever. \`.genex/scratch/\` already exists for this and is already ignored, so there is nothing to set up: write the capture as \`.genex/scratch/arena-before.png\` and read it back from there. Never put them in a folder of your own at the top level: \`progress/\`, \`reports/\`, \`shots/\` and their kind are pushed like source and become permanent weight in the game and in every remix of it \u2014 one real game reached 11 GB and 10,332 committed screenshots exactly this way. If you find such a folder already there, add it to \`.gitignore\`; never delete the player's files (law 18).
|
|
905
|
-
24. **Benchmark before declaring an in-game LLM
|
|
905
|
+
24. **Benchmark before declaring an in-game LLM ceiling.** If the game calls a language model while the player plays \u2014 an NPC answering in its own words, a quest written for this save, a judge reading what the player typed \u2014 the PLAYER pays for it from their Genex credits, billed as used, and the game declares no price, only a per-call CEILING: \`maxCredits\` on \`generate()\`, \`perCallMaxCredits\` on a standing budget. That ceiling bounds what one call may cost AND funds how long its answer may be, so never pick it by judgement or from a vendor's rate card. Run the real prompt on your own credits \u2014 \`npx genex llm bench "<the prompt>" --schema <file> --samples 3 --max-credits <n> --user-approved\` \u2014 and declare the figures it prints, recorded in \`DESIGN.md\` with its model and date; re-benchmark whenever the prompt, the schema or the model changes, because a prompt edit changes the ceiling. Check the lane first with \`npx genex llm models\`: where it is off the routes 404, so build the feature behind a graceful unavailable state and say so in the handoff. A repeated call is a standing budget the player approves once (\`requestSpendGrant()\`, then \`generate({ grantId })\`), never a popup per turn, and its disclosed call rate and credits per period are computed from the game's own loop and the bench's measured average. The coin spellings (\`estimateCoins\`, \`perCallMaxCoins\`) are deprecated \u2014 write the credit names. Model output is data the game validates against its own expectation \u2014 never executed, and never authority over coin, credits, items, entitlements or rewards. Load \`$genex-llm-in-games\` before writing any of it.
|
|
906
906
|
${CONTRACT_END}
|
|
907
907
|
`;
|
|
908
908
|
var TOOLS_CONTRACT_HEAD = `# Genex Tools (always in effect in this folder)
|
|
@@ -916,7 +916,7 @@ var TOOLS_CONTRACT_OFFER = `5. **When the game is built and playable \u2014 neve
|
|
|
916
916
|
var TOOLS_CONTRACT_HOSTED_LAWS = `5. **This game is hosted on Genex now, and publishing is glue \u2014 never a rebuild.** The game ships exactly as the user built it: \`initEmbed()\` first in the boot code (the \`genex-threejs-embed-auth\` card has the call; the renderer draws before any \`await waitForPlayer()\`), a static build into \`dist/\` with relative asset paths, then \`npx genex preview\`. Never restructure, reformat or "improve" the game to ship it, never scaffold a new app around it, and never start a design document or a build plan for it \u2014 the user's own process is theirs. Load the \`genex-tool-publish\` card before the first preview; it owns the vocabulary, the links and the limits.
|
|
917
917
|
6. **Ship first, then report.** \`preview\` prints a preflight \u2014 phone memory, a missing volume slider, an asset nobody wired, a viewport line. It never blocks a deploy, so push the build as it is, hand over the link, THEN relay each preflight line to the user in one plain sentence with an offer to fix it, and fix one only on their yes. The preflight is a report for the user, not a to-do list for you.
|
|
918
918
|
7. **The link is the game's page, and there are two versions.** After every preview give the user \`<dashboard>/draft/<slug>\` \u2014 \`dashboardOrigins[0]\` and \`slug\` from \`.genex/project.json\` \u2014 never localhost, a file path or the bare play origin. \`preview\` updates the draft and never touches what players are on; the first release is \`npx genex publish\` (it lists the game); after that "publish it", "update it" and "yes" all mean \`npx genex promote\` \u2014 the exact draft build, no rebuild. Ask once per round of work, in one line, and keep working while you wait.`;
|
|
919
|
-
var TOOLS_CONTRACT_LLM_OFFER = `A model running while people play is built into the Genex platform \u2014 the player pays, with Genex
|
|
919
|
+
var TOOLS_CONTRACT_LLM_OFFER = `A model running while people play is built into the Genex platform \u2014 the player pays, with their Genex credits or their own Claude/ChatGPT subscription, and approves it on a Genex sheet; your game just calls \`generate()\`. Want it that way?`;
|
|
920
920
|
function toolsContractLlmLaw(n, hosted) {
|
|
921
921
|
const onYes = hosted ? "On a yes, load `$genex-tool-llm` and then `$genex-llm-in-games`, which owns the build." : "On a yes, load `$genex-tool-llm`, which owns the rest: it checks the lane with `npx genex llm models` first, then runs `npx genex init --convert` (a hosted game is a static browser build \u2014 a local server holding a key can never ship).";
|
|
922
922
|
return `${n}. **A model running while people PLAY is a platform feature \u2014 never something you wire onto the user's own meter.** When a request implies one (NPCs that talk or decide in their own words, content written from what the player types, a prompt box in the game, "let the player pick a model"), recognise it, say in ONE line: "${TOOLS_CONTRACT_LLM_OFFER}" \u2014 then ASK and wait. ${onYes} On a no, build the authored version instead. Either way, never ship a game that calls a model on a key or an account of the user's own: every visitor would spend their money with nobody approving it, and the credential is readable in the bundle. Player-funded generation is text and JSON only \u2014 3D, images, video and audio stay the asset lanes above, on the user's meter.`;
|
|
@@ -22781,7 +22781,7 @@ async function runBlender(opts) {
|
|
|
22781
22781
|
return serveLocalBlender({ port, log });
|
|
22782
22782
|
}
|
|
22783
22783
|
if (sub === "mcp") {
|
|
22784
|
-
const { runBlenderMcp } = await import("./blender-mcp-
|
|
22784
|
+
const { runBlenderMcp } = await import("./blender-mcp-ATHMMRDU.js");
|
|
22785
22785
|
return runBlenderMcp();
|
|
22786
22786
|
}
|
|
22787
22787
|
if (sub === "seat") {
|
|
@@ -23289,7 +23289,12 @@ import fs33 from "fs/promises";
|
|
|
23289
23289
|
import path34 from "path";
|
|
23290
23290
|
import { randomUUID as randomUUID2 } from "crypto";
|
|
23291
23291
|
var SUBS3 = ["models", "bench", "price", "status", "cancel"];
|
|
23292
|
-
var LLM_BENCH_APPROVAL_REQUIRED = "STOP: `genex llm bench` runs the real model and spends YOUR OWN
|
|
23292
|
+
var LLM_BENCH_APPROVAL_REQUIRED = "STOP: `genex llm bench` runs the real model and spends YOUR OWN credits \u2014 your balance, not a player's, and a spend this build's asset allowance does not count. Re-run with --max-credits <n> --user-approved only after the person at the keyboard agreed to the number.";
|
|
23293
|
+
var LLM_MAX_COINS_DEPRECATED = "--max-coins is deprecated: it is read as --max-credits, the same number \u2014 a bench spends credits now.";
|
|
23294
|
+
function benchBalanceShortLine(spendable, maxCredits) {
|
|
23295
|
+
return `Your spendable credits (${spendable}) cannot cover one attempt at --max-credits ${maxCredits} \u2014 nothing was sent, nothing was spent. Add credits, or re-run with a lower --max-credits once the person at the keyboard agrees to it.`;
|
|
23296
|
+
}
|
|
23297
|
+
var LLM_BENCH_UNVERIFIED_LINE = "Credits pay only for a verified email, and this account's is not verified yet (credits_unverified) \u2014 nothing was sent, nothing was spent. Verify it, then re-run.";
|
|
23293
23298
|
var LLM_LANE_OFF_LINE = "The runtime LLM lane is off on this stand \u2014 in-game generate() answers 404 here, and nothing can be benchmarked.";
|
|
23294
23299
|
var LLM_TOOLS_CONVERT_HINT = [
|
|
23295
23300
|
"In-game model calls need a hosted game. This folder is a Genex Tools workspace, so nothing here",
|
|
@@ -23302,7 +23307,7 @@ function serverSlotCounts(body) {
|
|
|
23302
23307
|
return { held: body.slotsHeld, limit: body.slotLimit };
|
|
23303
23308
|
}
|
|
23304
23309
|
var PENDING_SLOT_RELEASE_MINUTES = 10;
|
|
23305
|
-
var LLM_STATUS_LEGEND = `"awaiting bill" = the call has stopped but the provider's bill is not final, so
|
|
23310
|
+
var LLM_STATUS_LEGEND = `"awaiting bill" = the call has stopped but the provider's bill is not final, so what it reserved stays held until the bill resolves; such a row stops holding a slot ${PENDING_SLOT_RELEASE_MINUTES} minutes after dispatch.`;
|
|
23306
23311
|
var LLM_STATUS_NO_COUNT_LINE = "This stand does not report how many slots are held across your projects, so no count is shown.";
|
|
23307
23312
|
var LLM_CANCEL_AWAITING_BILL = "this call already stopped; its bill is awaiting the provider";
|
|
23308
23313
|
var LLM_CANCEL_ALREADY_FINAL = "this call is already final";
|
|
@@ -23316,11 +23321,33 @@ var DEFAULT_SAMPLE_TIMEOUT_SEC = 180;
|
|
|
23316
23321
|
function parseProviderKey(value) {
|
|
23317
23322
|
return value === "ok" || value === "rejected" || value === "unknown" ? value : null;
|
|
23318
23323
|
}
|
|
23324
|
+
function bpsOrNull(value) {
|
|
23325
|
+
return typeof value === "number" && Number.isSafeInteger(value) && value >= 0 ? value : null;
|
|
23326
|
+
}
|
|
23327
|
+
function standingCreditGrantOrNull(value) {
|
|
23328
|
+
if (!value || typeof value !== "object") return null;
|
|
23329
|
+
const v = value;
|
|
23330
|
+
const whole = (x) => typeof x === "number" && Number.isSafeInteger(x) && x >= 0;
|
|
23331
|
+
if (!whole(v.minCredits) || !whole(v.maxCredits) || !whole(v.stepCredits) || !whole(v.defaultCredits)) return null;
|
|
23332
|
+
return {
|
|
23333
|
+
minCredits: v.minCredits,
|
|
23334
|
+
maxCredits: v.maxCredits,
|
|
23335
|
+
stepCredits: v.stepCredits,
|
|
23336
|
+
defaultCredits: v.defaultCredits,
|
|
23337
|
+
...whole(v.minUsdCents) ? { minUsdCents: v.minUsdCents } : {},
|
|
23338
|
+
...whole(v.maxUsdCents) ? { maxUsdCents: v.maxUsdCents } : {},
|
|
23339
|
+
...whole(v.stepUsdCents) ? { stepUsdCents: v.stepUsdCents } : {}
|
|
23340
|
+
};
|
|
23341
|
+
}
|
|
23319
23342
|
async function readRuntimeLane(apiUrl, token) {
|
|
23320
23343
|
const empty = (state) => ({
|
|
23321
23344
|
state,
|
|
23322
23345
|
models: null,
|
|
23323
23346
|
recommendedDeclaredHeadroomBps: null,
|
|
23347
|
+
currency: null,
|
|
23348
|
+
feeBps: null,
|
|
23349
|
+
usdCentsPerCredit: null,
|
|
23350
|
+
standingCreditGrant: null,
|
|
23324
23351
|
externalProviders: null,
|
|
23325
23352
|
total: null,
|
|
23326
23353
|
featuredCount: null,
|
|
@@ -23339,7 +23366,11 @@ async function readRuntimeLane(apiUrl, token) {
|
|
|
23339
23366
|
return {
|
|
23340
23367
|
state: "live",
|
|
23341
23368
|
models: body.models,
|
|
23342
|
-
recommendedDeclaredHeadroomBps:
|
|
23369
|
+
recommendedDeclaredHeadroomBps: bpsOrNull(body.recommendedDeclaredHeadroomBps),
|
|
23370
|
+
currency: body.currency === "credits" || body.currency === "coins" ? body.currency : null,
|
|
23371
|
+
feeBps: bpsOrNull(body.feeBps),
|
|
23372
|
+
usdCentsPerCredit: typeof body.usdCentsPerCredit === "number" && Number.isSafeInteger(body.usdCentsPerCredit) && body.usdCentsPerCredit >= 1 ? body.usdCentsPerCredit : null,
|
|
23373
|
+
standingCreditGrant: standingCreditGrantOrNull(body.standingCreditGrant),
|
|
23343
23374
|
externalProviders: Array.isArray(body.externalProviders) ? body.externalProviders : null,
|
|
23344
23375
|
total: typeof body.total === "number" && Number.isFinite(body.total) ? body.total : null,
|
|
23345
23376
|
featuredCount: typeof body.featuredCount === "number" && Number.isFinite(body.featuredCount) ? body.featuredCount : null,
|
|
@@ -23355,18 +23386,45 @@ function percentile2(values, p) {
|
|
|
23355
23386
|
const rank2 = Math.ceil(p / 100 * sorted.length);
|
|
23356
23387
|
return sorted[Math.min(sorted.length - 1, Math.max(0, rank2 - 1))];
|
|
23357
23388
|
}
|
|
23358
|
-
function
|
|
23359
|
-
|
|
23389
|
+
function percentilePicos(values, p) {
|
|
23390
|
+
if (values.length === 0) return null;
|
|
23391
|
+
const sorted = [...values].sort((a, b) => a < b ? -1 : a > b ? 1 : 0);
|
|
23392
|
+
const rank2 = Math.ceil(p / 100 * sorted.length);
|
|
23393
|
+
return sorted[Math.min(sorted.length - 1, Math.max(0, rank2 - 1))];
|
|
23360
23394
|
}
|
|
23361
|
-
|
|
23362
|
-
|
|
23395
|
+
var PICOS_PER_CENT = 10000000000n;
|
|
23396
|
+
function ceilingCreditsFor(costPicos, terms) {
|
|
23397
|
+
if (costPicos <= 0n) return 1;
|
|
23398
|
+
const num = costPicos * BigInt(1e4 + terms.feeBps) * BigInt(1e4 + terms.headroomBps);
|
|
23399
|
+
const den = 100000000n * BigInt(terms.usdCentsPerCredit) * PICOS_PER_CENT;
|
|
23400
|
+
return Math.max(1, Number((num + den - 1n) / den));
|
|
23401
|
+
}
|
|
23402
|
+
function averageCreditsPerCall(costsPicos, terms) {
|
|
23403
|
+
if (costsPicos.length === 0) return null;
|
|
23404
|
+
const sum = costsPicos.reduce((a, b) => a + b, 0n);
|
|
23405
|
+
return Number(sum * BigInt(1e4 + terms.feeBps)) / Number(BigInt(costsPicos.length) * 10000n * BigInt(terms.usdCentsPerCredit) * PICOS_PER_CENT);
|
|
23406
|
+
}
|
|
23407
|
+
function perCallEstimateCreditsFor(costsPicos, terms) {
|
|
23408
|
+
if (costsPicos.length === 0) return null;
|
|
23409
|
+
const sum = costsPicos.reduce((a, b) => a + b, 0n);
|
|
23410
|
+
const num = sum * BigInt(1e4 + terms.feeBps);
|
|
23411
|
+
const den = BigInt(costsPicos.length) * 10000n * BigInt(terms.usdCentsPerCredit) * PICOS_PER_CENT;
|
|
23412
|
+
return Math.max(1, Number((num + den - 1n) / den));
|
|
23413
|
+
}
|
|
23414
|
+
function formatCredits(value) {
|
|
23415
|
+
if (!Number.isFinite(value)) return "\u2014";
|
|
23416
|
+
if (value >= 100) return String(Math.round(value));
|
|
23417
|
+
return String(Number(value.toPrecision(value < 1 ? 2 : 3)));
|
|
23418
|
+
}
|
|
23419
|
+
function usdFromPicos(picos) {
|
|
23420
|
+
if (picos === null || picos === void 0) return "unknown";
|
|
23421
|
+
const value = typeof picos === "bigint" ? picos : /^\d+$/.test(picos) ? BigInt(picos) : null;
|
|
23422
|
+
if (value === null) return "unknown";
|
|
23423
|
+
return `$${(Number(value) / 1e12).toFixed(6)}`;
|
|
23363
23424
|
}
|
|
23364
23425
|
function usdPerMillion(value) {
|
|
23365
23426
|
return `$${value.toFixed(value < 1 ? 4 : 2)}`;
|
|
23366
23427
|
}
|
|
23367
|
-
function usd(value) {
|
|
23368
|
-
return value === null ? "unknown" : `$${value.toFixed(6)}`;
|
|
23369
|
-
}
|
|
23370
23428
|
async function runLlm(opts = {}) {
|
|
23371
23429
|
const log = createLogger({ quiet: opts.quiet || opts.json });
|
|
23372
23430
|
const cwd = opts.cwd ?? process.cwd();
|
|
@@ -23392,16 +23450,23 @@ async function runLlm(opts = {}) {
|
|
|
23392
23450
|
return;
|
|
23393
23451
|
}
|
|
23394
23452
|
if (sub === "bench") {
|
|
23395
|
-
if (opts.maxCoins
|
|
23453
|
+
if (opts.maxCoins !== void 0) process.stderr.write(`${c.yellow("!")} ${LLM_MAX_COINS_DEPRECATED}
|
|
23454
|
+
`);
|
|
23455
|
+
if (opts.maxCoins !== void 0 && opts.maxCredits !== void 0 && opts.maxCoins !== opts.maxCredits) {
|
|
23456
|
+
fail4("llm bench", `Pass --max-credits alone \u2014 --max-coins is its deprecated name, and the two disagree (${opts.maxCredits} vs ${opts.maxCoins}).`);
|
|
23457
|
+
return;
|
|
23458
|
+
}
|
|
23459
|
+
const approvedCredits = opts.maxCredits ?? opts.maxCoins;
|
|
23460
|
+
if (approvedCredits === void 0 || opts.userApproved !== true) {
|
|
23396
23461
|
fail4("llm bench", LLM_BENCH_APPROVAL_REQUIRED);
|
|
23397
23462
|
return;
|
|
23398
23463
|
}
|
|
23399
|
-
if (!Number.isInteger(
|
|
23400
|
-
fail4("llm bench", `--max-
|
|
23464
|
+
if (!Number.isInteger(approvedCredits) || approvedCredits < 1) {
|
|
23465
|
+
fail4("llm bench", `--max-credits takes a whole number of credits, 1 or more (got ${String(approvedCredits)}).`);
|
|
23401
23466
|
return;
|
|
23402
23467
|
}
|
|
23403
23468
|
if (!opts.benchPrompt?.trim()) {
|
|
23404
|
-
fail4("llm bench", '`genex llm bench` needs the prompt your game would send, e.g. genex llm bench "<prompt>" --max-
|
|
23469
|
+
fail4("llm bench", '`genex llm bench` needs the prompt your game would send, e.g. genex llm bench "<prompt>" --max-credits <n> --user-approved.');
|
|
23405
23470
|
return;
|
|
23406
23471
|
}
|
|
23407
23472
|
if (opts.jsonOutput && opts.textOutput) {
|
|
@@ -23493,6 +23558,10 @@ async function reportModels(apiUrl, token, opts, log, toolsOnly = false) {
|
|
|
23493
23558
|
models: null,
|
|
23494
23559
|
externalProviders: null,
|
|
23495
23560
|
recommendedDeclaredHeadroomBps: null,
|
|
23561
|
+
currency: null,
|
|
23562
|
+
feeBps: null,
|
|
23563
|
+
usdCentsPerCredit: null,
|
|
23564
|
+
standingCreditGrant: null,
|
|
23496
23565
|
total: null,
|
|
23497
23566
|
featuredCount: null,
|
|
23498
23567
|
// The stand was never asked, so the verdict is unknown — but the key is
|
|
@@ -23518,6 +23587,12 @@ async function reportModels(apiUrl, token, opts, log, toolsOnly = false) {
|
|
|
23518
23587
|
models: lane.models,
|
|
23519
23588
|
externalProviders: lane.externalProviders,
|
|
23520
23589
|
recommendedDeclaredHeadroomBps: lane.recommendedDeclaredHeadroomBps,
|
|
23590
|
+
// The credit terms a call is billed on, the server's — null on a stand
|
|
23591
|
+
// that predates credit billing, never a number the CLI filled in.
|
|
23592
|
+
currency: lane.currency,
|
|
23593
|
+
feeBps: lane.feeBps,
|
|
23594
|
+
usdCentsPerCredit: lane.usdCentsPerCredit,
|
|
23595
|
+
standingCreditGrant: lane.standingCreditGrant,
|
|
23521
23596
|
total: lane.total ?? lane.models?.length ?? null,
|
|
23522
23597
|
featuredCount: lane.featuredCount ?? (lane.models ? lane.models.filter((m) => m.featured).length : null),
|
|
23523
23598
|
// `null` when the lane is not live or the server predates the probe.
|
|
@@ -23582,74 +23657,104 @@ async function reportModels(apiUrl, token, opts, log, toolsOnly = false) {
|
|
|
23582
23657
|
}
|
|
23583
23658
|
log.plain("");
|
|
23584
23659
|
if (lane.recommendedDeclaredHeadroomBps === null) {
|
|
23585
|
-
log.dim(" This stand serves no recommended headroom, so no
|
|
23660
|
+
log.dim(" This stand serves no recommended headroom, so no ceiling can be recommended from a bench.");
|
|
23586
23661
|
} else {
|
|
23587
|
-
log.dim(` Recommended headroom over benchmarked
|
|
23662
|
+
log.dim(` Recommended headroom over a benchmarked call's cost on this stand: ${lane.recommendedDeclaredHeadroomBps} bps.`);
|
|
23663
|
+
}
|
|
23664
|
+
if (lane.feeBps === null) {
|
|
23665
|
+
log.dim(" This stand serves no platform fee \u2014 it predates credit billing, so no ceiling can be recommended from a bench.");
|
|
23666
|
+
} else {
|
|
23667
|
+
log.dim(` In-game calls are billed in the player's credits, as used: the provider's cost plus this stand's platform fee (${lane.feeBps} bps).`);
|
|
23588
23668
|
}
|
|
23589
23669
|
if (toolsOnly) {
|
|
23590
23670
|
printConvertHint();
|
|
23591
23671
|
return;
|
|
23592
23672
|
}
|
|
23593
|
-
log.dim(" Those per-million rates are the PROVIDER's; what
|
|
23594
|
-
log.dim(`
|
|
23673
|
+
log.dim(" Those per-million rates are the PROVIDER's list; what one call of YOUR prompt really costs,");
|
|
23674
|
+
log.dim(` and the per-call ceiling to declare, only a real call measures: ${c.cyan('genex llm bench "<prompt>" --max-credits <n> --user-approved')}`);
|
|
23595
23675
|
}
|
|
23596
23676
|
function wholeOrNull(value) {
|
|
23597
23677
|
return typeof value === "number" && Number.isSafeInteger(value) && value >= 0 ? value : null;
|
|
23598
23678
|
}
|
|
23679
|
+
function costPicosOf(usage) {
|
|
23680
|
+
if (typeof usage?.costUsdPicos === "string" && /^\d+$/.test(usage.costUsdPicos)) return usage.costUsdPicos;
|
|
23681
|
+
if (typeof usage?.costUsd === "number" && Number.isFinite(usage.costUsd) && usage.costUsd >= 0) {
|
|
23682
|
+
return String(BigInt(Math.round(usage.costUsd * 1e12)));
|
|
23683
|
+
}
|
|
23684
|
+
return null;
|
|
23685
|
+
}
|
|
23599
23686
|
function sampleFrom(id, settled) {
|
|
23600
23687
|
const usage = settled?.usage;
|
|
23688
|
+
const picos = costPicosOf(usage);
|
|
23601
23689
|
return {
|
|
23602
23690
|
id,
|
|
23603
23691
|
status: settled?.status ?? null,
|
|
23604
23692
|
billingStatus: settled?.billingStatus ?? null,
|
|
23605
|
-
|
|
23606
|
-
costUsd: typeof usage?.costUsd === "number" ? usage.costUsd : null,
|
|
23693
|
+
chargedCredits: wholeOrNull(settled?.chargedCredits),
|
|
23694
|
+
costUsd: typeof usage?.costUsd === "number" ? usage.costUsd : picos !== null ? Number(picos) / 1e12 : null,
|
|
23695
|
+
costUsdPicos: picos,
|
|
23607
23696
|
error: settled?.error ?? (settled ? null : "timed_out"),
|
|
23608
23697
|
providerMessage: typeof settled?.providerMessage === "string" && settled.providerMessage.trim() ? settled.providerMessage.trim() : null,
|
|
23609
23698
|
inputTokens: wholeOrNull(usage?.inputTokens),
|
|
23610
23699
|
outputTokens: wholeOrNull(usage?.outputTokens),
|
|
23611
23700
|
maxOutputTokens: wholeOrNull(usage?.maxOutputTokens),
|
|
23612
|
-
|
|
23701
|
+
fitCredits: typeof usage?.fitCredits === "number" && Number.isSafeInteger(usage.fitCredits) && usage.fitCredits >= 1 ? usage.fitCredits : null,
|
|
23613
23702
|
fitExceedsCap: typeof usage?.fitExceedsCap === "boolean" ? usage.fitExceedsCap : null,
|
|
23614
23703
|
truncated: typeof usage?.truncated === "boolean" ? usage.truncated : settled?.error === "provider_token_limit" ? true : null
|
|
23615
23704
|
};
|
|
23616
23705
|
}
|
|
23617
23706
|
var PROVIDER_TOKEN_LIMIT = "provider_token_limit";
|
|
23618
|
-
function benchRecommendation(results,
|
|
23619
|
-
const settled = results.filter(
|
|
23620
|
-
|
|
23621
|
-
|
|
23622
|
-
const
|
|
23623
|
-
const
|
|
23707
|
+
function benchRecommendation(results, terms) {
|
|
23708
|
+
const settled = results.filter(
|
|
23709
|
+
(r) => r.status === "succeeded" && r.billingStatus === "final" && r.chargedCredits !== null && r.costUsdPicos !== null
|
|
23710
|
+
);
|
|
23711
|
+
const costs = settled.map((r) => BigInt(r.costUsdPicos));
|
|
23712
|
+
const charged = settled.map((r) => r.chargedCredits);
|
|
23713
|
+
const fits = settled.map((r) => r.fitCredits).filter((v) => typeof v === "number");
|
|
23714
|
+
const costP95 = percentilePicos(costs, 95);
|
|
23715
|
+
const costMax = costs.length ? costs.reduce((a, b) => b > a ? b : a) : null;
|
|
23624
23716
|
const fitP95 = percentile2(fits, 95);
|
|
23625
23717
|
const fitMax = fits.length ? Math.max(...fits) : null;
|
|
23626
23718
|
const exceedsStandCap = results.some((r) => r.fitExceedsCap === true);
|
|
23627
23719
|
const cutOffSamples = results.filter((r) => r.error === PROVIDER_TOKEN_LIMIT && r.fitExceedsCap !== true).length;
|
|
23628
23720
|
const noRecommendationReason = exceedsStandCap ? "answer_exceeds_stand_cap" : cutOffSamples > 0 ? "answer_cut_off" : null;
|
|
23629
|
-
const
|
|
23630
|
-
const
|
|
23631
|
-
const
|
|
23721
|
+
const { headroomBps, feeBps, usdCentsPerCredit } = terms;
|
|
23722
|
+
const missingTerms = headroomBps === null ? "headroom" : feeBps === null || usdCentsPerCredit === null ? "credit_terms" : null;
|
|
23723
|
+
const full = headroomBps !== null && feeBps !== null && usdCentsPerCredit !== null ? { headroomBps, feeBps, usdCentsPerCredit } : null;
|
|
23724
|
+
const billing = feeBps !== null && usdCentsPerCredit !== null ? { feeBps, usdCentsPerCredit } : null;
|
|
23725
|
+
const costBased = full !== null && costP95 !== null ? ceilingCreditsFor(costP95, full) : null;
|
|
23726
|
+
const recommended = costBased === null || noRecommendationReason !== null ? null : Math.max(costBased, fitP95 ?? 0);
|
|
23727
|
+
const ceiling = recommended === null || costMax === null || full === null ? null : Math.max(ceilingCreditsFor(costMax, full), recommended, fitMax ?? 0);
|
|
23728
|
+
const average = billing !== null ? averageCreditsPerCall(costs, billing) : null;
|
|
23729
|
+
const estimate = ceiling === null || billing === null ? null : Math.min(ceiling, perCallEstimateCreditsFor(costs, billing) ?? 1);
|
|
23632
23730
|
return {
|
|
23731
|
+
costs,
|
|
23633
23732
|
charged,
|
|
23634
|
-
|
|
23635
|
-
|
|
23636
|
-
|
|
23733
|
+
costP50: percentilePicos(costs, 50),
|
|
23734
|
+
costP95,
|
|
23735
|
+
costMax,
|
|
23736
|
+
chargedP50: percentile2(charged, 50),
|
|
23737
|
+
chargedP95: percentile2(charged, 95),
|
|
23738
|
+
chargedMax: charged.length ? Math.max(...charged) : null,
|
|
23637
23739
|
fits,
|
|
23638
23740
|
fitP95,
|
|
23639
23741
|
fitMax,
|
|
23640
|
-
|
|
23641
|
-
|
|
23642
|
-
|
|
23643
|
-
|
|
23742
|
+
costBasedMaxCredits: costBased,
|
|
23743
|
+
recommendedMaxCredits: recommended,
|
|
23744
|
+
recommendedPerCallMaxCredits: ceiling,
|
|
23745
|
+
averageCreditsPerCall: average,
|
|
23746
|
+
recommendedPerCallEstimateCredits: estimate,
|
|
23747
|
+
lengthRaisedCeiling: recommended !== null && costBased !== null && recommended > costBased,
|
|
23644
23748
|
cutOffSamples,
|
|
23645
23749
|
exceedsStandCap,
|
|
23646
|
-
noRecommendationReason
|
|
23750
|
+
noRecommendationReason,
|
|
23751
|
+
missingTerms
|
|
23647
23752
|
};
|
|
23648
23753
|
}
|
|
23649
|
-
var LLM_LENGTH_RAISED_LINE = "The
|
|
23650
|
-
var LLM_EXCEEDS_STAND_CAP_LINE = "No
|
|
23651
|
-
function benchCutOffLine(cutOff,
|
|
23652
|
-
return `No
|
|
23754
|
+
var LLM_LENGTH_RAISED_LINE = "The ceiling also decides how long the answer may be: at the cost-based ceiling the answer would be cut off, so declare this one.";
|
|
23755
|
+
var LLM_EXCEEDS_STAND_CAP_LINE = "No ceiling recommended: this answer is longer than one call on this stand may produce. Ask for a shorter answer \u2014 fewer fields, shorter strings, a length the prompt states \u2014 and benchmark again.";
|
|
23756
|
+
function benchCutOffLine(cutOff, maxCredits) {
|
|
23757
|
+
return `No ceiling recommended: ${cutOff} sample${cutOff === 1 ? " was" : "s were"} cut off at this ceiling (--max-credits ${maxCredits}) before the answer was finished, so its real length is unknown. Re-run with a higher --max-credits.`;
|
|
23653
23758
|
}
|
|
23654
23759
|
var ACTIVE_STATUSES = /* @__PURE__ */ new Set(["requires_confirmation", "queued", "dispatching", "awaiting_external"]);
|
|
23655
23760
|
function rowState(row) {
|
|
@@ -23660,7 +23765,22 @@ function rowState(row) {
|
|
|
23660
23765
|
function rowHoldsSlot(row) {
|
|
23661
23766
|
if (typeof row.slotHeld === "boolean") return row.slotHeld;
|
|
23662
23767
|
const state = rowState(row);
|
|
23663
|
-
return state === "active" || state === "awaiting bill" && (row.reservedCoins ?? 0) > 0;
|
|
23768
|
+
return state === "active" || state === "awaiting bill" && ((row.reservedCoins ?? 0) > 0 || (row.reservedCredits ?? 0) > 0);
|
|
23769
|
+
}
|
|
23770
|
+
function rowMoney(row) {
|
|
23771
|
+
const credits = row.currency === "credits" || row.currency === void 0 && typeof row.reservedCredits === "number" && typeof row.reservedCoins !== "number";
|
|
23772
|
+
if (credits) {
|
|
23773
|
+
return {
|
|
23774
|
+
unit: "credits",
|
|
23775
|
+
reserved: typeof row.reservedCredits === "number" ? row.reservedCredits : null,
|
|
23776
|
+
charged: typeof row.chargedCredits === "number" ? row.chargedCredits : null
|
|
23777
|
+
};
|
|
23778
|
+
}
|
|
23779
|
+
return {
|
|
23780
|
+
unit: "coin",
|
|
23781
|
+
reserved: typeof row.reservedCoins === "number" ? row.reservedCoins : null,
|
|
23782
|
+
charged: typeof row.chargedCoins === "number" ? row.chargedCoins : null
|
|
23783
|
+
};
|
|
23664
23784
|
}
|
|
23665
23785
|
function providerRefusalStatus(error) {
|
|
23666
23786
|
const m = /^provider_http_(\d{3})$/.exec(error ?? "");
|
|
@@ -23679,7 +23799,7 @@ function formatAge(iso, now = Date.now()) {
|
|
|
23679
23799
|
}
|
|
23680
23800
|
async function runBench(args) {
|
|
23681
23801
|
const { apiUrl, token, projectId, cwd, opts, log } = args;
|
|
23682
|
-
const
|
|
23802
|
+
const maxCredits = opts.maxCredits ?? opts.maxCoins;
|
|
23683
23803
|
const samples = opts.samples ?? DEFAULT_BENCH_SAMPLES;
|
|
23684
23804
|
const prompt = opts.benchPrompt.trim();
|
|
23685
23805
|
const auth = { Authorization: `Bearer ${token}`, "Content-Type": "application/json" };
|
|
@@ -23701,7 +23821,7 @@ async function runBench(args) {
|
|
|
23701
23821
|
const lane = await readRuntimeLane(apiUrl, token);
|
|
23702
23822
|
if (lane.state !== "live" || !lane.models || lane.models.length === 0) {
|
|
23703
23823
|
if (lane.state === "off") {
|
|
23704
|
-
if (opts.json) writeJsonLine({ ...emptyBenchJson(apiUrl, projectId, prompt, outputFormat, samples,
|
|
23824
|
+
if (opts.json) writeJsonLine({ ...emptyBenchJson(apiUrl, projectId, prompt, outputFormat, samples, maxCredits, opts), status: "off", error: null });
|
|
23705
23825
|
else {
|
|
23706
23826
|
log.plain(LLM_LANE_OFF_LINE);
|
|
23707
23827
|
log.dim(" Nothing was spent.");
|
|
@@ -23728,17 +23848,27 @@ async function runBench(args) {
|
|
|
23728
23848
|
);
|
|
23729
23849
|
return;
|
|
23730
23850
|
}
|
|
23731
|
-
const
|
|
23851
|
+
const walletBefore = await fetchCreditsSnapshot(apiUrl, token);
|
|
23852
|
+
if (walletBefore?.emailVerified === false) {
|
|
23853
|
+
benchFailed(opts, log, LLM_BENCH_UNVERIFIED_LINE);
|
|
23854
|
+
return;
|
|
23855
|
+
}
|
|
23856
|
+
if (walletBefore !== null && walletBefore.spendable < maxCredits) {
|
|
23857
|
+
benchFailed(opts, log, benchBalanceShortLine(walletBefore.spendable, maxCredits));
|
|
23858
|
+
return;
|
|
23859
|
+
}
|
|
23860
|
+
const balanceBefore = walletBefore?.spendable ?? null;
|
|
23732
23861
|
if (!opts.json) {
|
|
23733
23862
|
log.plain(c.bold("genex llm bench"));
|
|
23734
23863
|
log.dim(` ${apiUrl}`);
|
|
23735
23864
|
log.plain("");
|
|
23736
23865
|
log.plain(` Model ${c.cyan(modelId)}`);
|
|
23737
23866
|
log.plain(` Samples ${samples} real attempt${samples === 1 ? "" : "s"}, ${outputFormat} output`);
|
|
23738
|
-
log.plain(` Approved ${
|
|
23867
|
+
log.plain(` Approved ${maxCredits} credit${maxCredits === 1 ? "" : "s"} per attempt \u2014 at worst ${maxCredits * samples} credits for this run, from your own balance`);
|
|
23739
23868
|
log.plain(
|
|
23740
|
-
` Balance ${balanceBefore === null ? "couldn't be read" : `${balanceBefore}
|
|
23869
|
+
` Balance ${balanceBefore === null ? "couldn't be read" : `${balanceBefore} credits spendable`}`
|
|
23741
23870
|
);
|
|
23871
|
+
log.dim(" Billed as used: each attempt is charged its real cost plus the platform fee, rounded up to a whole credit.");
|
|
23742
23872
|
log.plain("");
|
|
23743
23873
|
}
|
|
23744
23874
|
const base = `${apiUrl}/api/runtime/development/projects/${encodeURIComponent(projectId)}/generations`;
|
|
@@ -23764,7 +23894,7 @@ async function runBench(args) {
|
|
|
23764
23894
|
const res = await call(base, {
|
|
23765
23895
|
method: "POST",
|
|
23766
23896
|
body: JSON.stringify({
|
|
23767
|
-
|
|
23897
|
+
maxCredits,
|
|
23768
23898
|
request: {
|
|
23769
23899
|
idempotencyKey: `bench-${randomUUID2()}`,
|
|
23770
23900
|
modelId,
|
|
@@ -23794,13 +23924,14 @@ async function runBench(args) {
|
|
|
23794
23924
|
if (refusedAt !== null) providerRefused++;
|
|
23795
23925
|
if (!opts.json) {
|
|
23796
23926
|
const length = row.outputTokens !== null ? ` \xB7 ${row.outputTokens}${row.maxOutputTokens !== null ? ` of ${row.maxOutputTokens}` : ""} tokens out` : "";
|
|
23927
|
+
const charged = row.chargedCredits === null ? "\u2014" : String(row.chargedCredits);
|
|
23797
23928
|
log.plain(
|
|
23798
|
-
` ${row.status === "succeeded" ? c.green("\u2713") : c.yellow("!")} sample ${i + 1} ${
|
|
23929
|
+
` ${row.status === "succeeded" ? c.green("\u2713") : c.yellow("!")} sample ${i + 1} ${charged.padStart(4)} credit${row.chargedCredits === 1 ? "" : "s"} charged \xB7 cost ${usdFromPicos(row.costUsdPicos)}${length}${row.error ? ` \xB7 ${row.error}` : ""}`
|
|
23799
23930
|
);
|
|
23800
23931
|
if (refusedAt !== null) printProviderRefusal(log, refusedAt, row.providerMessage, row.status === "unknown");
|
|
23801
23932
|
if (row.error === PROVIDER_TOKEN_LIMIT) {
|
|
23802
23933
|
log.dim(
|
|
23803
|
-
row.fitExceedsCap === true ? " The answer was cut off at this stand's own output limit \u2014 no
|
|
23934
|
+
row.fitExceedsCap === true ? " The answer was cut off at this stand's own output limit \u2014 no ceiling makes room for it; it is not a sample." : ` The answer was cut off at this ceiling (--max-credits ${maxCredits}) before it was finished \u2014 it was charged, and it is not a sample.`
|
|
23804
23935
|
);
|
|
23805
23936
|
}
|
|
23806
23937
|
}
|
|
@@ -23808,16 +23939,19 @@ async function runBench(args) {
|
|
|
23808
23939
|
} finally {
|
|
23809
23940
|
process.removeListener("SIGINT", onSigint);
|
|
23810
23941
|
}
|
|
23811
|
-
const
|
|
23812
|
-
|
|
23813
|
-
|
|
23814
|
-
|
|
23815
|
-
|
|
23816
|
-
const
|
|
23942
|
+
const terms = {
|
|
23943
|
+
headroomBps: lane.recommendedDeclaredHeadroomBps,
|
|
23944
|
+
feeBps: lane.feeBps,
|
|
23945
|
+
usdCentsPerCredit: lane.usdCentsPerCredit
|
|
23946
|
+
};
|
|
23947
|
+
const rec = benchRecommendation(results, terms);
|
|
23948
|
+
const settledCount = rec.charged.length;
|
|
23817
23949
|
const lengthBlocked = rec.noRecommendationReason !== null;
|
|
23818
|
-
const balanceAfter = await
|
|
23950
|
+
const balanceAfter = (await fetchCreditsSnapshot(apiUrl, token))?.spendable ?? null;
|
|
23951
|
+
const picosOrNull = (v) => v === null ? null : String(v);
|
|
23819
23952
|
const record = {
|
|
23820
|
-
v:
|
|
23953
|
+
v: 2,
|
|
23954
|
+
currency: "credits",
|
|
23821
23955
|
ranAt: (/* @__PURE__ */ new Date()).toISOString(),
|
|
23822
23956
|
apiUrl,
|
|
23823
23957
|
projectId,
|
|
@@ -23826,28 +23960,34 @@ async function runBench(args) {
|
|
|
23826
23960
|
outputFormat,
|
|
23827
23961
|
samples,
|
|
23828
23962
|
completed: settledCount,
|
|
23829
|
-
|
|
23830
|
-
|
|
23831
|
-
|
|
23832
|
-
max,
|
|
23833
|
-
|
|
23834
|
-
|
|
23835
|
-
recommendedPerCallMaxCoins: ceiling,
|
|
23836
|
-
maxCoinsPerSample: maxCoins,
|
|
23837
|
-
fitCoins: rec.fits,
|
|
23963
|
+
maxCreditsPerSample: maxCredits,
|
|
23964
|
+
costUsdPicos: rec.costs.map(String),
|
|
23965
|
+
chargedCredits: rec.charged,
|
|
23966
|
+
cost: { p50: picosOrNull(rec.costP50), p95: picosOrNull(rec.costP95), max: picosOrNull(rec.costMax) },
|
|
23967
|
+
charged: { p50: rec.chargedP50, p95: rec.chargedP95, max: rec.chargedMax },
|
|
23968
|
+
fitCredits: rec.fits,
|
|
23838
23969
|
fitP95: rec.fitP95,
|
|
23839
|
-
|
|
23840
|
-
|
|
23970
|
+
fitMax: rec.fitMax,
|
|
23971
|
+
recommendedDeclaredHeadroomBps: terms.headroomBps,
|
|
23972
|
+
feeBps: terms.feeBps,
|
|
23973
|
+
usdCentsPerCredit: terms.usdCentsPerCredit,
|
|
23974
|
+
costBasedMaxCredits: rec.costBasedMaxCredits,
|
|
23975
|
+
recommendedMaxCredits: rec.recommendedMaxCredits,
|
|
23976
|
+
recommendedPerCallMaxCredits: rec.recommendedPerCallMaxCredits,
|
|
23977
|
+
averageCreditsPerCall: rec.averageCreditsPerCall,
|
|
23978
|
+
recommendedPerCallEstimateCredits: rec.recommendedPerCallEstimateCredits,
|
|
23979
|
+
lengthRaisedCeiling: rec.lengthRaisedCeiling,
|
|
23841
23980
|
cutOffSamples: rec.cutOffSamples,
|
|
23842
23981
|
noRecommendationReason: rec.noRecommendationReason
|
|
23843
23982
|
};
|
|
23844
23983
|
const savedTo = await saveBench(cwd, record);
|
|
23845
23984
|
if (opts.json) {
|
|
23846
|
-
const ok =
|
|
23985
|
+
const ok = settledCount > 0 && !lengthBlocked;
|
|
23847
23986
|
writeJsonLine({
|
|
23848
23987
|
command: "llm bench",
|
|
23849
23988
|
status: ok ? "ok" : "failed",
|
|
23850
23989
|
error: ok ? null : rec.noRecommendationReason ?? "no_sample_settled",
|
|
23990
|
+
currency: "credits",
|
|
23851
23991
|
apiUrl,
|
|
23852
23992
|
projectId,
|
|
23853
23993
|
modelId,
|
|
@@ -23856,35 +23996,54 @@ async function runBench(args) {
|
|
|
23856
23996
|
schemaPath: opts.schemaPath ?? null,
|
|
23857
23997
|
samples,
|
|
23858
23998
|
completed: settledCount,
|
|
23859
|
-
|
|
23999
|
+
maxCreditsPerSample: maxCredits,
|
|
24000
|
+
// Spendable credits before and after the run; null when unreadable.
|
|
23860
24001
|
balanceBefore,
|
|
23861
24002
|
balanceAfter,
|
|
23862
24003
|
results,
|
|
23863
24004
|
// The create-time refusal that ended the run, or null. Beside the
|
|
23864
24005
|
// samples, never among them: no attempt existed and nothing was spent.
|
|
23865
24006
|
refusal,
|
|
23866
|
-
|
|
23867
|
-
//
|
|
24007
|
+
// Over the settled samples: the real provider cost (pico-dollar strings
|
|
24008
|
+
// and USD) and the whole credits charged.
|
|
24009
|
+
costUsdPicos: record.cost,
|
|
24010
|
+
costUsd: {
|
|
24011
|
+
p50: rec.costP50 === null ? null : Number(rec.costP50) / 1e12,
|
|
24012
|
+
p95: rec.costP95 === null ? null : Number(rec.costP95) / 1e12,
|
|
24013
|
+
max: rec.costMax === null ? null : Number(rec.costMax) / 1e12
|
|
24014
|
+
},
|
|
24015
|
+
chargedCredits: record.charged,
|
|
24016
|
+
// The answer's LENGTH, priced by the server per sample (`fitCredits`,
|
|
23868
24017
|
// headroom already on the tokens), over the settled samples.
|
|
23869
|
-
|
|
23870
|
-
|
|
23871
|
-
|
|
24018
|
+
fitCredits: { p95: rec.fitP95, max: rec.fitMax },
|
|
24019
|
+
recommendedDeclaredHeadroomBps: terms.headroomBps,
|
|
24020
|
+
feeBps: terms.feeBps,
|
|
24021
|
+
usdCentsPerCredit: terms.usdCentsPerCredit,
|
|
24022
|
+
missingTerms: rec.missingTerms,
|
|
24023
|
+
costBasedMaxCredits: rec.costBasedMaxCredits,
|
|
24024
|
+
lengthRaisedCeiling: rec.lengthRaisedCeiling,
|
|
23872
24025
|
cutOffSamples: rec.cutOffSamples,
|
|
23873
24026
|
exceedsStandCap: rec.exceedsStandCap,
|
|
23874
24027
|
noRecommendationReason: rec.noRecommendationReason,
|
|
23875
|
-
|
|
23876
|
-
|
|
23877
|
-
|
|
24028
|
+
recommendedMaxCredits: rec.recommendedMaxCredits,
|
|
24029
|
+
recommendedPerCallMaxCredits: rec.recommendedPerCallMaxCredits,
|
|
24030
|
+
averageCreditsPerCall: rec.averageCreditsPerCall,
|
|
24031
|
+
recommendedPerCallEstimateCredits: rec.recommendedPerCallEstimateCredits,
|
|
23878
24032
|
savedTo
|
|
23879
24033
|
});
|
|
23880
24034
|
if (!ok) process.exitCode = 1;
|
|
23881
24035
|
return;
|
|
23882
24036
|
}
|
|
23883
24037
|
log.plain("");
|
|
24038
|
+
const printMeasured = () => {
|
|
24039
|
+
log.plain(c.bold(" Real cost per call"));
|
|
24040
|
+
log.plain(` p50 ${usdFromPicos(rec.costP50)} p95 ${usdFromPicos(rec.costP95)} max ${usdFromPicos(rec.costMax)} (${settledCount} of ${samples} settled)`);
|
|
24041
|
+
log.plain(c.bold(" Charged credits"));
|
|
24042
|
+
log.plain(` p50 ${rec.chargedP50} p95 ${rec.chargedP95} max ${rec.chargedMax} (whole credits, rounded up per call)`);
|
|
24043
|
+
};
|
|
23884
24044
|
if (lengthBlocked) {
|
|
23885
|
-
if (
|
|
23886
|
-
|
|
23887
|
-
log.plain(` p50 ${p50} p95 ${p95} max ${max} (${charged.length} of ${samples} settled)`);
|
|
24045
|
+
if (settledCount > 0) {
|
|
24046
|
+
printMeasured();
|
|
23888
24047
|
log.plain("");
|
|
23889
24048
|
}
|
|
23890
24049
|
printRecommendation(log, record);
|
|
@@ -23892,17 +24051,16 @@ async function runBench(args) {
|
|
|
23892
24051
|
process.exitCode = 1;
|
|
23893
24052
|
return;
|
|
23894
24053
|
}
|
|
23895
|
-
if (
|
|
24054
|
+
if (settledCount === 0) {
|
|
23896
24055
|
log.error(
|
|
23897
|
-
providerRefused > 0 && providerRefused === results.length ? " No sample ran \u2014 the provider refused every attempt at its door \u2014 so there is nothing to
|
|
24056
|
+
providerRefused > 0 && providerRefused === results.length ? " No sample ran \u2014 the provider refused every attempt at its door \u2014 so there is nothing to measure from." : " No sample settled, so there is nothing to measure from."
|
|
23898
24057
|
);
|
|
23899
24058
|
process.exitCode = 1;
|
|
23900
24059
|
return;
|
|
23901
24060
|
}
|
|
23902
|
-
|
|
23903
|
-
log.plain(` p50 ${p50} p95 ${p95} max ${max} (${charged.length} of ${samples} settled)`);
|
|
24061
|
+
printMeasured();
|
|
23904
24062
|
if (balanceBefore !== null && balanceAfter !== null) {
|
|
23905
|
-
log.plain(` Balance ${balanceBefore} \u2192 ${balanceAfter}
|
|
24063
|
+
log.plain(` Balance ${balanceBefore} \u2192 ${balanceAfter} credits spendable`);
|
|
23906
24064
|
}
|
|
23907
24065
|
log.plain("");
|
|
23908
24066
|
printRecommendation(log, record);
|
|
@@ -23914,43 +24072,58 @@ function printRecommendation(log, record) {
|
|
|
23914
24072
|
return;
|
|
23915
24073
|
}
|
|
23916
24074
|
if (record.noRecommendationReason === "answer_cut_off") {
|
|
23917
|
-
log.warn(` ${benchCutOffLine(record.cutOffSamples
|
|
24075
|
+
log.warn(` ${benchCutOffLine(record.cutOffSamples || 1, record.maxCreditsPerSample)}`);
|
|
23918
24076
|
return;
|
|
23919
24077
|
}
|
|
23920
|
-
|
|
23921
|
-
|
|
24078
|
+
const attempts = `${record.completed} settled attempt${record.completed === 1 ? "" : "s"} on ${record.modelId}`;
|
|
24079
|
+
if (record.recommendedMaxCredits === null || record.cost.p95 === null) {
|
|
24080
|
+
log.warn(" No ceiling recommended.");
|
|
23922
24081
|
log.dim(
|
|
23923
|
-
record.recommendedDeclaredHeadroomBps === null ? " This stand served no recommended headroom, and the multiplier is the server's to set \u2014" : "
|
|
24082
|
+
record.cost.p95 === null ? " No sample settled, so there is no p95 to build a ceiling on \u2014" : record.recommendedDeclaredHeadroomBps === null ? " This stand served no recommended headroom, and the multiplier is the server's to set \u2014" : " This stand served no platform fee or credit value, and those are the server's to set \u2014"
|
|
23924
24083
|
);
|
|
23925
24084
|
log.dim(" declaring a number from this run would be a guess dressed as a measurement.");
|
|
24085
|
+
printAverage(log, record, attempts);
|
|
23926
24086
|
return;
|
|
23927
24087
|
}
|
|
23928
|
-
log.plain(c.bold(` Declare
|
|
23929
|
-
if (record.
|
|
24088
|
+
log.plain(c.bold(` Declare maxCredits: ${record.recommendedMaxCredits}`));
|
|
24089
|
+
if (record.lengthRaisedCeiling) {
|
|
23930
24090
|
log.plain(` ${LLM_LENGTH_RAISED_LINE}`);
|
|
23931
24091
|
log.dim(
|
|
23932
|
-
` = the smallest
|
|
24092
|
+
` = the smallest ceiling whose answer allowance holds the answer plus this stand's headroom (p95 ${record.fitP95 ?? "\u2014"}), above the`
|
|
23933
24093
|
);
|
|
23934
24094
|
log.dim(
|
|
23935
|
-
`
|
|
24095
|
+
` cost-based ${record.costBasedMaxCredits ?? "\u2014"} (p95 real cost ${usdFromPicos(record.cost.p95)} with the platform fee, ${record.feeBps} bps, and ${record.recommendedDeclaredHeadroomBps} bps of headroom), over ${attempts}.`
|
|
23936
24096
|
);
|
|
23937
24097
|
} else {
|
|
23938
24098
|
log.dim(
|
|
23939
|
-
` = p95
|
|
24099
|
+
` = p95 real cost (${usdFromPicos(record.cost.p95)}) with this stand's platform fee (${record.feeBps} bps) and recommended headroom (${record.recommendedDeclaredHeadroomBps} bps), in whole credits,`
|
|
23940
24100
|
);
|
|
23941
|
-
log.dim(` measured over ${
|
|
24101
|
+
log.dim(` measured over ${attempts}.`);
|
|
23942
24102
|
}
|
|
23943
|
-
log.dim(" That number is a
|
|
23944
|
-
|
|
23945
|
-
|
|
24103
|
+
log.dim(" That number is a CEILING, not a price: a call is billed its real cost as used, never more than it \u2014");
|
|
24104
|
+
log.dim(" and it also sizes how long the answer may be.");
|
|
24105
|
+
if (record.recommendedPerCallMaxCredits !== null) {
|
|
24106
|
+
log.plain(c.bold(` Grant perCallMaxCredits: ${record.recommendedPerCallMaxCredits}`));
|
|
23946
24107
|
log.dim(
|
|
23947
|
-
record.
|
|
24108
|
+
record.lengthRaisedCeiling ? ` = the same over the worst sample (${usdFromPicos(record.cost.max)}) or the longest answer's ceiling, whichever is larger, and never below the ceiling above it.` : ` = the same over the worst sample (${usdFromPicos(record.cost.max)}), and never below the ceiling above it.`
|
|
23948
24109
|
);
|
|
23949
|
-
|
|
24110
|
+
}
|
|
24111
|
+
printAverage(log, record, attempts);
|
|
24112
|
+
}
|
|
24113
|
+
function printAverage(log, record, attempts) {
|
|
24114
|
+
if (record.averageCreditsPerCall === null) return;
|
|
24115
|
+
log.plain(` Per-call estimate: about ${formatCredits(record.averageCreditsPerCall)} credits per call`);
|
|
24116
|
+
log.dim(` = the average real cost with the platform fee, no headroom, over ${attempts}.`);
|
|
24117
|
+
if (record.recommendedPerCallEstimateCredits !== null) {
|
|
24118
|
+
log.plain(c.bold(` Grant perCallEstimateCredits: ${record.recommendedPerCallEstimateCredits}`));
|
|
24119
|
+
log.dim(` = that average rounded up to a whole credit. disclosure.estimatedCreditsPerPeriod = your calls per period \xD7 ${formatCredits(record.averageCreditsPerCall)}, rounded up.`);
|
|
23950
24120
|
}
|
|
23951
24121
|
}
|
|
23952
24122
|
async function explainBenchRefusal(res, call, base, log, opts, index, settledSoFar) {
|
|
23953
|
-
if (printedStructuredError(res))
|
|
24123
|
+
if (printedStructuredError(res)) {
|
|
24124
|
+
const printed2 = await res.json().catch(() => ({}));
|
|
24125
|
+
return { stop: true, refusal: { status: res.status, error: typeof printed2.error === "string" ? printed2.error : null, slotsHeld: null, slotLimit: null } };
|
|
24126
|
+
}
|
|
23954
24127
|
const body = await res.json().catch(() => ({}));
|
|
23955
24128
|
const refusal = { status: res.status, error: body.error ?? null, slotsHeld: null, slotLimit: null };
|
|
23956
24129
|
if (res.status === 429 && body.error === "generation_limit") {
|
|
@@ -23960,6 +24133,10 @@ async function explainBenchRefusal(res, call, base, log, opts, index, settledSoF
|
|
|
23960
24133
|
if (!opts.json) log.error(` Sample ${index + 1} ${benchSlotsHeldSentence(counts)}`);
|
|
23961
24134
|
return { stop: true, refusal };
|
|
23962
24135
|
}
|
|
24136
|
+
if (res.status === 403 && body.error === "credits_unverified") {
|
|
24137
|
+
if (!opts.json) log.error(` ${LLM_BENCH_UNVERIFIED_LINE}`);
|
|
24138
|
+
return { stop: true, refusal };
|
|
24139
|
+
}
|
|
23963
24140
|
if (opts.json) return { stop: res.status !== 429, refusal };
|
|
23964
24141
|
if (res.status === 404 && body.error === "not_found") {
|
|
23965
24142
|
log.error(` ${LLM_LANE_OFF_LINE}`);
|
|
@@ -23975,7 +24152,7 @@ async function explainBenchRefusal(res, call, base, log, opts, index, settledSoF
|
|
|
23975
24152
|
return { stop: true, refusal };
|
|
23976
24153
|
}
|
|
23977
24154
|
if (res.status === 402) {
|
|
23978
|
-
log.error(" Not enough
|
|
24155
|
+
log.error(" Not enough credits to start the attempt \u2014 a bench is paid from your own balance.");
|
|
23979
24156
|
return { stop: true, refusal };
|
|
23980
24157
|
}
|
|
23981
24158
|
log.error(
|
|
@@ -23984,7 +24161,7 @@ async function explainBenchRefusal(res, call, base, log, opts, index, settledSoF
|
|
|
23984
24161
|
return { stop: res.status >= 500 || res.status === 401 || res.status === 403, refusal };
|
|
23985
24162
|
}
|
|
23986
24163
|
function printProviderRefusal(log, status2, message, awaitingBill = false) {
|
|
23987
|
-
if (awaitingBill) log.dim(` The provider answered ${status2} after routing \u2014 its bill is not final, so this call
|
|
24164
|
+
if (awaitingBill) log.dim(` The provider answered ${status2} after routing \u2014 its bill is not final, so what this call reserved stays held until it resolves (see \`genex llm status\`); it is not a sample.`);
|
|
23988
24165
|
else log.dim(" The provider refused this call at its door \u2014 no inference ran, it cost nothing, and it is not a sample.");
|
|
23989
24166
|
if (message) log.dim(` Provider said: ${message}`);
|
|
23990
24167
|
if (status2 === 401 || status2 === 403) {
|
|
@@ -24010,19 +24187,6 @@ async function pollSettled(call, url, timeoutSec) {
|
|
|
24010
24187
|
}
|
|
24011
24188
|
return last;
|
|
24012
24189
|
}
|
|
24013
|
-
async function readCoinBalance(apiUrl, token) {
|
|
24014
|
-
try {
|
|
24015
|
-
const res = await apiFetch(`${apiUrl}/api/coin/balance`, {
|
|
24016
|
-
headers: { Authorization: `Bearer ${token}` },
|
|
24017
|
-
signal: AbortSignal.timeout(6e3)
|
|
24018
|
-
});
|
|
24019
|
-
if (!res.ok) return null;
|
|
24020
|
-
const body = await res.json().catch(() => null);
|
|
24021
|
-
return typeof body?.spendable === "number" ? body.spendable : null;
|
|
24022
|
-
} catch {
|
|
24023
|
-
return null;
|
|
24024
|
-
}
|
|
24025
|
-
}
|
|
24026
24190
|
async function saveBench(cwd, record) {
|
|
24027
24191
|
try {
|
|
24028
24192
|
const file = path34.join(cwd, LLM_BENCH_FILE);
|
|
@@ -24038,9 +24202,10 @@ function benchFailed(opts, log, message) {
|
|
|
24038
24202
|
else log.error(message);
|
|
24039
24203
|
process.exitCode = 1;
|
|
24040
24204
|
}
|
|
24041
|
-
function emptyBenchJson(apiUrl, projectId, prompt, outputFormat, samples,
|
|
24205
|
+
function emptyBenchJson(apiUrl, projectId, prompt, outputFormat, samples, maxCredits, opts) {
|
|
24042
24206
|
return {
|
|
24043
24207
|
command: "llm bench",
|
|
24208
|
+
currency: "credits",
|
|
24044
24209
|
apiUrl,
|
|
24045
24210
|
projectId,
|
|
24046
24211
|
modelId: opts.modelId ?? null,
|
|
@@ -24049,21 +24214,28 @@ function emptyBenchJson(apiUrl, projectId, prompt, outputFormat, samples, maxCoi
|
|
|
24049
24214
|
schemaPath: opts.schemaPath ?? null,
|
|
24050
24215
|
samples,
|
|
24051
24216
|
completed: 0,
|
|
24052
|
-
|
|
24217
|
+
maxCreditsPerSample: maxCredits,
|
|
24053
24218
|
balanceBefore: null,
|
|
24054
24219
|
balanceAfter: null,
|
|
24055
24220
|
results: [],
|
|
24056
24221
|
refusal: null,
|
|
24057
|
-
|
|
24058
|
-
|
|
24059
|
-
|
|
24060
|
-
|
|
24222
|
+
costUsdPicos: { p50: null, p95: null, max: null },
|
|
24223
|
+
costUsd: { p50: null, p95: null, max: null },
|
|
24224
|
+
chargedCredits: { p50: null, p95: null, max: null },
|
|
24225
|
+
fitCredits: { p95: null, max: null },
|
|
24226
|
+
recommendedDeclaredHeadroomBps: null,
|
|
24227
|
+
feeBps: null,
|
|
24228
|
+
usdCentsPerCredit: null,
|
|
24229
|
+
missingTerms: null,
|
|
24230
|
+
costBasedMaxCredits: null,
|
|
24231
|
+
lengthRaisedCeiling: false,
|
|
24061
24232
|
cutOffSamples: 0,
|
|
24062
24233
|
exceedsStandCap: false,
|
|
24063
24234
|
noRecommendationReason: null,
|
|
24064
|
-
|
|
24065
|
-
|
|
24066
|
-
|
|
24235
|
+
recommendedMaxCredits: null,
|
|
24236
|
+
recommendedPerCallMaxCredits: null,
|
|
24237
|
+
averageCreditsPerCall: null,
|
|
24238
|
+
recommendedPerCallEstimateCredits: null,
|
|
24067
24239
|
savedTo: null
|
|
24068
24240
|
};
|
|
24069
24241
|
}
|
|
@@ -24126,7 +24298,8 @@ async function reportStatus(args) {
|
|
|
24126
24298
|
for (const row of rows) {
|
|
24127
24299
|
const state = rowState(row);
|
|
24128
24300
|
const mark = state === "active" ? c.cyan("\u25CF") : state === "awaiting bill" ? c.yellow("\u25D0") : c.green("\u25CB");
|
|
24129
|
-
const
|
|
24301
|
+
const money2 = rowMoney(row);
|
|
24302
|
+
const reserved = money2.reserved === null ? "reserved \u2014" : `${money2.reserved} ${money2.unit} reserved`;
|
|
24130
24303
|
log.plain(
|
|
24131
24304
|
` ${mark} ${row.id ?? "?"} ${row.modelId ?? "?"} ${state.padEnd(13)} ${formatAge(row.createdAt, now).padStart(4)} old ${reserved}${row.error ? ` ${row.error}` : ""}${rowHoldsSlot(row) ? "" : c.dim(" (no slot)")}`
|
|
24132
24305
|
);
|
|
@@ -24198,10 +24371,11 @@ async function cancelCall(args) {
|
|
|
24198
24371
|
if (!opts.json) {
|
|
24199
24372
|
log.warn(` ${row.id} (${row.status ?? "?"}) \u2014 ${sentence}.`);
|
|
24200
24373
|
if (state === "awaiting bill") {
|
|
24201
|
-
log.dim(` Nothing to cancel:
|
|
24374
|
+
log.dim(` Nothing to cancel: what it reserved releases when the bill resolves, and it stops holding a slot ${PENDING_SLOT_RELEASE_MINUTES} minutes after dispatch.`);
|
|
24202
24375
|
if (row.providerMessage?.trim()) log.dim(` Provider said: ${row.providerMessage.trim()}`);
|
|
24203
24376
|
} else {
|
|
24204
|
-
|
|
24377
|
+
const money2 = rowMoney(row);
|
|
24378
|
+
log.dim(` ${money2.charged === null ? "charge unknown" : `${money2.charged} ${money2.unit} charged`}${row.error ? ` \xB7 ${row.error}` : ""}.`);
|
|
24205
24379
|
}
|
|
24206
24380
|
}
|
|
24207
24381
|
return done("ok", null, state === "awaiting bill" ? "awaiting_bill" : "already_final", row);
|
|
@@ -24223,78 +24397,126 @@ async function cancelCall(args) {
|
|
|
24223
24397
|
const after = await res.json().catch(() => null) ?? row;
|
|
24224
24398
|
if (!opts.json) {
|
|
24225
24399
|
log.success(` Cancelled ${row.id} (was ${row.status ?? "active"}).`);
|
|
24226
|
-
log.dim(`
|
|
24400
|
+
log.dim(` What it reserved releases once the bill settles; ${c.cyan("genex llm status")} shows it until then.`);
|
|
24227
24401
|
}
|
|
24228
24402
|
return done("ok", null, "cancelled", after);
|
|
24229
24403
|
}
|
|
24230
|
-
async function
|
|
24231
|
-
let
|
|
24404
|
+
async function readSavedBench(cwd) {
|
|
24405
|
+
let parsed;
|
|
24232
24406
|
try {
|
|
24233
|
-
|
|
24407
|
+
parsed = JSON.parse(await fs33.readFile(path34.join(cwd, LLM_BENCH_FILE), "utf8"));
|
|
24234
24408
|
} catch {
|
|
24235
|
-
|
|
24409
|
+
return null;
|
|
24236
24410
|
}
|
|
24237
|
-
if (!
|
|
24411
|
+
if (!parsed || typeof parsed !== "object" || Array.isArray(parsed)) return null;
|
|
24412
|
+
const v = parsed;
|
|
24413
|
+
if (v.v === 2 && v.currency === "credits" && v.cost && v.charged) return { kind: "credits", record: parsed };
|
|
24414
|
+
return { kind: "coins", record: parsed };
|
|
24415
|
+
}
|
|
24416
|
+
var LLM_PRICE_COIN_TERMS_LINE = "This bench was measured under the old COIN terms, when a game declared a fixed price per call. In-game calls are now billed in the player's credits, as used, and a game declares a per-call CEILING (maxCredits) instead \u2014 a coin price is not that number. Re-run the bench for the figures to declare.";
|
|
24417
|
+
async function reportSavedPrice(cwd, opts, log) {
|
|
24418
|
+
const saved = await readSavedBench(cwd);
|
|
24419
|
+
const rerun = 'genex llm bench "<prompt>" --max-credits <n> --user-approved';
|
|
24420
|
+
const creditKeys = (record2) => ({
|
|
24421
|
+
costUsdPicos: record2?.cost ?? { p50: null, p95: null, max: null },
|
|
24422
|
+
chargedCredits: record2?.charged ?? { p50: null, p95: null, max: null },
|
|
24423
|
+
fitP95: record2?.fitP95 ?? null,
|
|
24424
|
+
fitMax: record2?.fitMax ?? null,
|
|
24425
|
+
recommendedDeclaredHeadroomBps: record2?.recommendedDeclaredHeadroomBps ?? null,
|
|
24426
|
+
feeBps: record2?.feeBps ?? null,
|
|
24427
|
+
usdCentsPerCredit: record2?.usdCentsPerCredit ?? null,
|
|
24428
|
+
costBasedMaxCredits: record2?.costBasedMaxCredits ?? null,
|
|
24429
|
+
lengthRaisedCeiling: record2?.lengthRaisedCeiling ?? false,
|
|
24430
|
+
cutOffSamples: record2?.cutOffSamples ?? 0,
|
|
24431
|
+
noRecommendationReason: record2?.noRecommendationReason ?? null,
|
|
24432
|
+
recommendedMaxCredits: record2?.recommendedMaxCredits ?? null,
|
|
24433
|
+
recommendedPerCallMaxCredits: record2?.recommendedPerCallMaxCredits ?? null,
|
|
24434
|
+
averageCreditsPerCall: record2?.averageCreditsPerCall ?? null,
|
|
24435
|
+
recommendedPerCallEstimateCredits: record2?.recommendedPerCallEstimateCredits ?? null
|
|
24436
|
+
});
|
|
24437
|
+
if (!saved) {
|
|
24238
24438
|
if (opts.json) {
|
|
24239
24439
|
writeJsonLine({
|
|
24240
24440
|
command: "llm price",
|
|
24241
24441
|
status: "failed",
|
|
24242
24442
|
error: "no_bench",
|
|
24443
|
+
currency: null,
|
|
24243
24444
|
ranAt: null,
|
|
24244
24445
|
modelId: null,
|
|
24245
24446
|
samples: null,
|
|
24246
24447
|
completed: null,
|
|
24247
|
-
|
|
24248
|
-
|
|
24249
|
-
max: null,
|
|
24250
|
-
fitP95: null,
|
|
24251
|
-
chargedBasedEstimateCoins: null,
|
|
24252
|
-
lengthRaisedPrice: false,
|
|
24253
|
-
cutOffSamples: 0,
|
|
24254
|
-
noRecommendationReason: null,
|
|
24255
|
-
recommendedDeclaredHeadroomBps: null,
|
|
24256
|
-
recommendedEstimateCoins: null,
|
|
24257
|
-
recommendedPerCallMaxCoins: null,
|
|
24448
|
+
...creditKeys(null),
|
|
24449
|
+
legacyCoinTerms: null,
|
|
24258
24450
|
savedTo: null
|
|
24259
24451
|
});
|
|
24260
24452
|
} else {
|
|
24261
24453
|
log.error("Nothing has been benchmarked in this folder yet.");
|
|
24262
|
-
log.dim(` Run ${c.cyan(
|
|
24454
|
+
log.dim(` Run ${c.cyan(rerun)} first.`);
|
|
24455
|
+
}
|
|
24456
|
+
process.exitCode = 1;
|
|
24457
|
+
return;
|
|
24458
|
+
}
|
|
24459
|
+
if (saved.kind === "coins") {
|
|
24460
|
+
const old = saved.record;
|
|
24461
|
+
const legacy = {
|
|
24462
|
+
p50: old.p50 ?? null,
|
|
24463
|
+
p95: old.p95 ?? null,
|
|
24464
|
+
max: old.max ?? null,
|
|
24465
|
+
recommendedEstimateCoins: old.recommendedEstimateCoins ?? null,
|
|
24466
|
+
recommendedPerCallMaxCoins: old.recommendedPerCallMaxCoins ?? null,
|
|
24467
|
+
maxCoinsPerSample: old.maxCoinsPerSample ?? null,
|
|
24468
|
+
noRecommendationReason: old.noRecommendationReason ?? null
|
|
24469
|
+
};
|
|
24470
|
+
if (opts.json) {
|
|
24471
|
+
writeJsonLine({
|
|
24472
|
+
command: "llm price",
|
|
24473
|
+
status: "failed",
|
|
24474
|
+
error: "coin_terms",
|
|
24475
|
+
currency: "coins",
|
|
24476
|
+
ranAt: old.ranAt ?? null,
|
|
24477
|
+
modelId: old.modelId ?? null,
|
|
24478
|
+
samples: old.samples ?? null,
|
|
24479
|
+
completed: old.completed ?? null,
|
|
24480
|
+
...creditKeys(null),
|
|
24481
|
+
legacyCoinTerms: legacy,
|
|
24482
|
+
savedTo: LLM_BENCH_FILE
|
|
24483
|
+
});
|
|
24484
|
+
} else {
|
|
24485
|
+
log.plain(c.bold("genex llm price"));
|
|
24486
|
+
log.dim(` from ${LLM_BENCH_FILE}, benchmarked ${old.ranAt ?? "at an unknown time"}${old.modelId ? ` on ${old.modelId}` : ""}`);
|
|
24487
|
+
log.plain("");
|
|
24488
|
+
log.warn(` ${LLM_PRICE_COIN_TERMS_LINE}`);
|
|
24489
|
+
const measured = legacy.p95 === null ? "no settled sample" : `charged coins p50 ${legacy.p50} \xB7 p95 ${legacy.p95} \xB7 max ${legacy.max}`;
|
|
24490
|
+
const recommended = legacy.recommendedEstimateCoins === null ? "no price recommended" : `it recommended estimateCoins ${legacy.recommendedEstimateCoins}${legacy.recommendedPerCallMaxCoins != null ? ` and perCallMaxCoins ${legacy.recommendedPerCallMaxCoins}` : ""}`;
|
|
24491
|
+
log.dim(` What it measured then: ${measured}; ${recommended}.`);
|
|
24492
|
+
log.dim(` Re-run: ${c.cyan(rerun)}`);
|
|
24263
24493
|
}
|
|
24264
24494
|
process.exitCode = 1;
|
|
24265
24495
|
return;
|
|
24266
24496
|
}
|
|
24497
|
+
const record = saved.record;
|
|
24267
24498
|
if (opts.json) {
|
|
24268
24499
|
writeJsonLine({
|
|
24269
24500
|
command: "llm price",
|
|
24270
|
-
status: record.
|
|
24271
|
-
error: record.
|
|
24501
|
+
status: record.recommendedMaxCredits === null ? "failed" : "ok",
|
|
24502
|
+
error: record.recommendedMaxCredits === null ? record.noRecommendationReason ?? "no_recommendation" : null,
|
|
24503
|
+
currency: "credits",
|
|
24272
24504
|
ranAt: record.ranAt ?? null,
|
|
24273
24505
|
modelId: record.modelId ?? null,
|
|
24274
24506
|
samples: record.samples ?? null,
|
|
24275
24507
|
completed: record.completed ?? null,
|
|
24276
|
-
|
|
24277
|
-
|
|
24278
|
-
max: record.max ?? null,
|
|
24279
|
-
// Absent from a file an older CLI wrote; null / false / 0 then, never missing.
|
|
24280
|
-
fitP95: record.fitP95 ?? null,
|
|
24281
|
-
chargedBasedEstimateCoins: record.chargedBasedEstimateCoins ?? null,
|
|
24282
|
-
lengthRaisedPrice: record.lengthRaisedPrice ?? false,
|
|
24283
|
-
cutOffSamples: record.cutOffSamples ?? 0,
|
|
24284
|
-
noRecommendationReason: record.noRecommendationReason ?? null,
|
|
24285
|
-
recommendedDeclaredHeadroomBps: record.recommendedDeclaredHeadroomBps ?? null,
|
|
24286
|
-
recommendedEstimateCoins: record.recommendedEstimateCoins ?? null,
|
|
24287
|
-
recommendedPerCallMaxCoins: record.recommendedPerCallMaxCoins ?? null,
|
|
24508
|
+
...creditKeys(record),
|
|
24509
|
+
legacyCoinTerms: null,
|
|
24288
24510
|
savedTo: LLM_BENCH_FILE
|
|
24289
24511
|
});
|
|
24290
|
-
if (record.
|
|
24512
|
+
if (record.recommendedMaxCredits === null) process.exitCode = 1;
|
|
24291
24513
|
return;
|
|
24292
24514
|
}
|
|
24293
24515
|
log.plain(c.bold("genex llm price"));
|
|
24294
24516
|
log.dim(` from ${LLM_BENCH_FILE}, benchmarked ${record.ranAt}`);
|
|
24295
24517
|
log.plain("");
|
|
24296
24518
|
printRecommendation(log, record);
|
|
24297
|
-
if (record.
|
|
24519
|
+
if (record.recommendedMaxCredits === null) process.exitCode = 1;
|
|
24298
24520
|
}
|
|
24299
24521
|
function writeJsonLine(value) {
|
|
24300
24522
|
process.stdout.write(`${JSON.stringify(value)}
|
|
@@ -24554,7 +24776,7 @@ function runtimeRow(state, lane, toolsOnly = false) {
|
|
|
24554
24776
|
return {
|
|
24555
24777
|
label: "Runtime",
|
|
24556
24778
|
value: `in-game LLM lane live \xB7 ${n} model${n === 1 ? "" : "s"}${keyRefused ? ` \xB7 ${c.yellow("provider key refused")}` : ""}`,
|
|
24557
|
-
fix: keyRefused ? LLM_PROVIDER_KEY_REJECTED_LINE : toolsOnly ? convertFix : '
|
|
24779
|
+
fix: keyRefused ? LLM_PROVIDER_KEY_REJECTED_LINE : toolsOnly ? convertFix : 'Measure a call before a game declares its per-call ceiling: `npx genex llm bench "<prompt>" --max-credits <n> --user-approved`.'
|
|
24558
24780
|
};
|
|
24559
24781
|
}
|
|
24560
24782
|
case "off":
|
|
@@ -26492,16 +26714,16 @@ ${c.bold("Usage")}
|
|
|
26492
26714
|
status | cancel. "models" says whether this
|
|
26493
26715
|
stand serves the runtime LLM lane at all, and
|
|
26494
26716
|
at what provider rates. "bench" runs the real
|
|
26495
|
-
model on YOUR OWN
|
|
26496
|
-
attempt
|
|
26717
|
+
model on YOUR OWN credits and reports what each
|
|
26718
|
+
attempt really cost; it needs --max-credits <n>
|
|
26497
26719
|
--user-approved, like any spend the player has
|
|
26498
26720
|
to agree to. "price" reprints the last run's
|
|
26499
|
-
recommended
|
|
26721
|
+
recommended maxCredits. "status" lists this
|
|
26500
26722
|
project's open calls and the slots they hold;
|
|
26501
26723
|
"cancel <id>" stops an active one. Declare a
|
|
26502
|
-
game's
|
|
26503
|
-
a
|
|
26504
|
-
|
|
26724
|
+
game's per-call ceiling from a bench, never
|
|
26725
|
+
from a guess: the player is billed as used, and
|
|
26726
|
+
the ceiling also sizes how long an answer may be.
|
|
26505
26727
|
genex player <sub> [options] Answer the in-game model requests you approve, on your
|
|
26506
26728
|
OWN Claude or ChatGPT subscription: install | run |
|
|
26507
26729
|
status | stop | uninstall. "install" signs this machine
|
|
@@ -26936,7 +27158,7 @@ ${c.bold("Examples")}
|
|
|
26936
27158
|
genex shop remove sku_123
|
|
26937
27159
|
genex llm models
|
|
26938
27160
|
genex llm models --all
|
|
26939
|
-
genex llm bench "Reply with one short taunt." --samples 5 --max-
|
|
27161
|
+
genex llm bench "Reply with one short taunt." --samples 5 --max-credits 20 --user-approved
|
|
26940
27162
|
genex llm price
|
|
26941
27163
|
genex llm status
|
|
26942
27164
|
genex llm cancel <id>
|
|
@@ -27173,22 +27395,24 @@ ${c.bold("How the allowance works")}
|
|
|
27173
27395
|
|
|
27174
27396
|
Every number printed is live from your account; prices can change without a deploy.
|
|
27175
27397
|
`,
|
|
27176
|
-
llm: `${c.bold("genex llm")} \u2014 in-game model calls: is the lane live, what does one call cost, what should the game declare?
|
|
27398
|
+
llm: `${c.bold("genex llm")} \u2014 in-game model calls: is the lane live, what does one call cost, what ceiling should the game declare?
|
|
27177
27399
|
|
|
27178
27400
|
${c.bold("Usage")}
|
|
27179
27401
|
genex llm models [--all] [--json] Which models this stand serves, their provider rates per
|
|
27180
|
-
million tokens,
|
|
27181
|
-
|
|
27182
|
-
|
|
27183
|
-
|
|
27184
|
-
on this stand: it says so
|
|
27185
|
-
|
|
27186
|
-
|
|
27187
|
-
|
|
27188
|
-
|
|
27189
|
-
|
|
27190
|
-
|
|
27191
|
-
|
|
27402
|
+
million tokens, the platform fee a call is billed at, and
|
|
27403
|
+
the headroom it recommends over a benchmark. The catalog
|
|
27404
|
+
is synced from OpenRouter, so the FEATURED short list is
|
|
27405
|
+
printed by default and --all prints every row (--json
|
|
27406
|
+
always carries every row). Off on this stand: it says so
|
|
27407
|
+
and exits 0 \u2014 build the feature without a model rather
|
|
27408
|
+
than promising a 404.
|
|
27409
|
+
genex llm bench "<prompt>" --max-credits <n> --user-approved [options]
|
|
27410
|
+
Run the real model through the development lane and
|
|
27411
|
+
report what each attempt really COST and was charged.
|
|
27412
|
+
Refused without both flags, before any network call:
|
|
27413
|
+
this spends YOUR OWN CREDITS, a spend the build's asset
|
|
27414
|
+
allowance does not count. Needs a folder linked to a
|
|
27415
|
+
game you own, and a balance that covers one attempt.
|
|
27192
27416
|
genex llm price [--json] The recommendation from the last bench in this folder.
|
|
27193
27417
|
Reads a file; spends nothing.
|
|
27194
27418
|
genex llm status [--json] This project's open development calls: active, or
|
|
@@ -27205,29 +27429,32 @@ ${c.bold("Bench options")}
|
|
|
27205
27429
|
--samples <n> How many real attempts (default ${DEFAULT_BENCH_SAMPLES}). More samples, better p95.
|
|
27206
27430
|
--schema <file> A JSON file holding the output schema; implies JSON output.
|
|
27207
27431
|
--json-output Ask for JSON without a schema. --text Ask for text (the default).
|
|
27208
|
-
--max-
|
|
27432
|
+
--max-credits <n> The per-attempt ceiling YOU approved \u2014 a bench runs on your own credits,
|
|
27209
27433
|
never a player's. The run's worst case is that number times --samples,
|
|
27210
|
-
and it is printed before anything starts.
|
|
27434
|
+
and it is printed before anything starts. (--max-coins is its deprecated
|
|
27435
|
+
name, read as the same number.)
|
|
27211
27436
|
--timeout <seconds> Bounds the wait for each attempt.
|
|
27212
27437
|
--json One machine-readable object; every key is always present, null when
|
|
27213
27438
|
the CLI could not learn it.
|
|
27214
27439
|
|
|
27215
27440
|
${c.bold("Why a benchmark and not an estimate")}
|
|
27216
|
-
|
|
27217
|
-
|
|
27218
|
-
|
|
27219
|
-
|
|
27220
|
-
|
|
27221
|
-
|
|
27222
|
-
|
|
27223
|
-
|
|
27224
|
-
|
|
27441
|
+
An in-game call is paid from the PLAYER's credits and billed as used: its real provider cost
|
|
27442
|
+
plus the platform fee, never more than the per-call CEILING the game declares (\`maxCredits\`,
|
|
27443
|
+
or \`perCallMaxCredits\` on a standing budget). The ceiling is not a price, but it is still a
|
|
27444
|
+
number with two jobs: it bounds what one call may cost, and it funds how LONG the answer may
|
|
27445
|
+
be. Three layers sit between the provider's rate card and it \u2014 the provider's cost for YOUR
|
|
27446
|
+
prompt, the platform fee, and your headroom \u2014 and only a real call measures the first. The
|
|
27447
|
+
bench is billed to YOU, on the development lane; the player-funded lane is what the
|
|
27448
|
+
published game uses. The recommendation is the p95 real cost with the server's fee and
|
|
27449
|
+
headroom applied, in whole credits; a stand that serves no headroom or no fee gets no
|
|
27450
|
+
recommendation, because those multipliers are not the CLI's to invent.
|
|
27225
27451
|
|
|
27226
|
-
|
|
27227
|
-
|
|
27228
|
-
|
|
27229
|
-
|
|
27230
|
-
|
|
27452
|
+
Because the ceiling sizes each call's output allowance, the recommendation is never below
|
|
27453
|
+
the smallest ceiling that leaves room for the benchmarked answer, as the server computes it.
|
|
27454
|
+
A sample cut off at your --max-credits (provider_token_limit) is not a sample \u2014 re-run with a
|
|
27455
|
+
higher --max-credits; an answer longer than one call on this stand may produce gets no
|
|
27456
|
+
ceiling at all \u2014 ask for a shorter one. The run also prints the measured average per call,
|
|
27457
|
+
fee included, which is the honest estimate a budget's disclosure is built on.
|
|
27231
27458
|
|
|
27232
27459
|
Results land in ${LLM_BENCH_FILE} \u2014 gitignored, machine-local, re-read by \`genex llm price\`.
|
|
27233
27460
|
`,
|
|
@@ -27442,11 +27669,13 @@ function parseArgs(argv) {
|
|
|
27442
27669
|
"--limit",
|
|
27443
27670
|
// `genex budget --assets <credits>` — the allowance the player approved.
|
|
27444
27671
|
"--assets",
|
|
27445
|
-
// `genex llm bench` value flags. `--max-
|
|
27446
|
-
//
|
|
27672
|
+
// `genex llm bench` value flags. `--max-credits` is the same approval
|
|
27673
|
+
// shape as `--assets`, but its own gate: the asset allowance does not
|
|
27674
|
+
// count a bench. `--max-coins` is its deprecated name (same number).
|
|
27447
27675
|
"--model",
|
|
27448
27676
|
"--samples",
|
|
27449
27677
|
"--schema",
|
|
27678
|
+
"--max-credits",
|
|
27450
27679
|
"--max-coins",
|
|
27451
27680
|
// `genex player install --label <name>` — what this machine is called in
|
|
27452
27681
|
// the dashboard. Its own key rather than the generic 2nd positional,
|
|
@@ -27961,12 +28190,14 @@ function applyValueFlag(options, flag, value) {
|
|
|
27961
28190
|
options.samples = n;
|
|
27962
28191
|
break;
|
|
27963
28192
|
}
|
|
28193
|
+
case "--max-credits":
|
|
27964
28194
|
case "--max-coins": {
|
|
27965
28195
|
const n = Number(value);
|
|
27966
28196
|
if (!Number.isInteger(n) || n < 1) {
|
|
27967
|
-
throw new Error(`Invalid
|
|
28197
|
+
throw new Error(`Invalid ${flag} value: ${value} (whole credits, 1 or more)`);
|
|
27968
28198
|
}
|
|
27969
|
-
options.
|
|
28199
|
+
if (flag === "--max-credits") options.maxCredits = n;
|
|
28200
|
+
else options.maxCoins = n;
|
|
27970
28201
|
break;
|
|
27971
28202
|
}
|
|
27972
28203
|
case "--timeout": {
|
|
@@ -28258,9 +28489,9 @@ async function main() {
|
|
|
28258
28489
|
await runBudget(parsed.options);
|
|
28259
28490
|
break;
|
|
28260
28491
|
// The in-game model lane: is it live here, what does one call really
|
|
28261
|
-
// cost, what should the game declare. `bench`
|
|
28262
|
-
//
|
|
28263
|
-
//
|
|
28492
|
+
// cost, what ceiling should the game declare. `bench` spends the
|
|
28493
|
+
// builder's own credits OUTSIDE the asset allowance, so it carries its
|
|
28494
|
+
// own --max-credits/--user-approved gate.
|
|
28264
28495
|
case "llm":
|
|
28265
28496
|
await runLlm(parsed.options);
|
|
28266
28497
|
break;
|