@wardby/cli 0.5.1 → 0.5.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -307,7 +307,9 @@ export class NativeEngine {
307
307
  error: "LLM stream ended without a usage summary.",
308
308
  };
309
309
  }
310
- const calibration = checkTokenCalibration(inputTokens, usage.inputTokens);
310
+ // The estimate covers the whole prompt; LlmUsage.inputTokens excludes
311
+ // cache-write tokens (billed separately), so add them back to compare.
312
+ const calibration = checkTokenCalibration(inputTokens, usage.inputTokens + (usage.cacheWriteTokens ?? 0));
311
313
  if (calibration.diverged) {
312
314
  const direction = calibration.deltaRatio < 0
313
315
  ? "UNDERestimated (dangerous — the pre-flight refuse/gate could admit an over-budget run)"
@@ -1185,8 +1185,8 @@
1185
1185
  ],
1186
1186
  "appliesTo": "\">=0.4.0\"",
1187
1187
  "sourcePath": "models.md",
1188
- "markdown": "\n# Models and pricing\n\nWardby prices and routes every model call from one catalog: the models this\nrelease ships, overlaid with this deployment's own additions and overrides.\nAn agent can use a model only when it's in the catalog and its provider has\ncredentials configured here.\n\n## The five tools\n\n| Tool | Scope | What it does |\n| --------------- | -------------- | ----------------------------------------------------------------------------------------- |\n| `list_models` | `agents:read` | Every active catalog entry (or, with `includeDisabled: true`, disabled ones too). |\n| `get_model` | `agents:read` | One entry by `modelId`; for an override, also the shipped entry it shadows. |\n| `set_model` | `models:admin` | Adds or completely replaces one entry. Every field is required. |\n| `disable_model` | `models:admin` | Removes a model from routing without deleting its pricing history. |\n| `reset_model` | `models:admin` | Removes every row for a model id, reverting to the shipped entry (if any) or removing it. |\n\nReading the catalog needs only `agents:read` — no secrets live in an entry.\nChanging it needs `models:admin`, honored only for a caller whose Wardby role\ngrants it: `admin`, or the narrower `model-manager` role.\n\nAn entry's `origin` is `shipped` or `override`; `routable` says whether this\ndeployment can actually route to it right now; `shippedDiffers` (overrides of\na shipped model only) says whether your override has drifted from the\ncurrent shipped values. See [`docs/models.md`](../docs/models.md) for the\nfull field reference.\n\n## Adding or overriding a model\n\n```json\n{\n \"provider\": \"anthropic\",\n \"modelId\": \"claude-example-model\",\n \"encoding\": \"o200k_base\",\n \"inputPerMTok\": 0.0,\n \"outputPerMTok\": 0.0,\n \"cachedInputPerMTok\": 0.0,\n \"cacheWritePerMTok\": 0.0,\n \"efforts\": [\"low\", \"medium\", \"high\"],\n \"thinkingMode\": \"adaptive\",\n \"sourceUrl\": \"https://example.com/replace-with-the-providers-own-pricing-page\"\n}\n```\n\nThe rates above are placeholders. Copy the provider's own published rates for\nthat exact model — including its cache read and cache write rates — from its\npricing page, and point `sourceUrl` at that page; never compute cache rates\nfrom `inputPerMTok` with a multiplier.\n\n`set_model` refuses (409) a `modelId` another provider already owns. A\nshipped id always belongs to its shipped provider — permanently; no other\nprovider can ever claim it, not even by disabling or resetting the override.\nA non-shipped id already claimed by another provider (its row active or\ndisabled) is freed only by `reset_model` — it takes only the `modelId` and\nclears every row for it, whichever provider owns it; `disable_model` alone\nnever frees it, since the disabled row still reserves the id.\n\n`thinkingMode` (`adaptive`, `manual`, or `none`) must match what the exact\nmodel accepts. Getting it wrong doesn't fail at `set_model` — it fails later,\nwhen a run calls the model, with `unsupported_anthropic_feature`.\nFor Claude Code coding runs, `efforts` is also the exact set of levels a run\nmay send, so always include the model's default effort level.\n\nA newly released Claude model may also need a newer Claude Code than your\nClaude Code worker image has: coding runs on it then fail as\n`provider_rejected` (no cost) while native runs work. Upgrade wardby and\nrebuild the worker images before using the model in Claude Code agents.\n\nCatalog changes take effect on the writing process immediately, and on every\nother wardby process within `WARDBY_MODEL_CATALOG_REFRESH_SECONDS` (default\n45). A run already in progress keeps the catalog entry it started with, so\ndisabling or repricing a model never changes a run already under way — only\nnew runs.\n\nIf this deployment delegates to an identity provider, define `models:admin`\nthere before relying on it, and map the `model-manager` role (or `admin`) to\nthe people who maintain pricing — see\n[Configure identity and privileged access](identity-and-access.md).\n\nIf a run can't use a model, see\n[Model not available](errors/model-unavailable.md).\n",
1189
- "plainText": "Models and pricing Wardby prices and routes every model call from one catalog: the models this release ships, overlaid with this deployment's own additions and overrides. An agent can use a model only when it's in the catalog and its provider has credentials configured here. The five tools | Tool | Scope | What it does | | --------------- | -------------- | ----------------------------------------------------------------------------------------- | | listmodels | agents:read | Every active catalog entry (or, with includeDisabled: true, disabled ones too). | | getmodel | agents:read | One entry by modelId; for an override, also the shipped entry it shadows. | | setmodel | models:admin | Adds or completely replaces one entry. Every field is required. | | disablemodel | models:admin | Removes a model from routing without deleting its pricing history. | | resetmodel | models:admin | Removes every row for a model id, reverting to the shipped entry (if any) or removing it. | Reading the catalog needs only agents:read — no secrets live in an entry. Changing it needs models:admin, honored only for a caller whose Wardby role grants it: admin, or the narrower model-manager role. An entry's origin is shipped or override; routable says whether this deployment can actually route to it right now; shippedDiffers (overrides of a shipped model only) says whether your override has drifted from the current shipped values. See docs/models.md for the full field reference. Adding or overriding a model { \"provider\": \"anthropic\", \"modelId\": \"claude-example-model\", \"encoding\": \"o200kbase\", \"inputPerMTok\": 0.0, \"outputPerMTok\": 0.0, \"cachedInputPerMTok\": 0.0, \"cacheWritePerMTok\": 0.0, \"efforts\": [\"low\", \"medium\", \"high\"], \"thinkingMode\": \"adaptive\", \"sourceUrl\": \"https://example.com/replace-with-the-providers-own-pricing-page\" } The rates above are placeholders. Copy the provider's own published rates for that exact model — including its cache read and cache write rates — from its pricing page, and point sourceUrl at that page; never compute cache rates from inputPerMTok with a multiplier. setmodel refuses (409) a modelId another provider already owns. A shipped id always belongs to its shipped provider — permanently; no other provider can ever claim it, not even by disabling or resetting the override. A non-shipped id already claimed by another provider (its row active or disabled) is freed only by resetmodel — it takes only the modelId and clears every row for it, whichever provider owns it; disablemodel alone never frees it, since the disabled row still reserves the id. thinkingMode (adaptive, manual, or none) must match what the exact model accepts. Getting it wrong doesn't fail at setmodel — it fails later, when a run calls the model, with unsupportedanthropicfeature. For Claude Code coding runs, efforts is also the exact set of levels a run may send, so always include the model's default effort level. A newly released Claude model may also need a newer Claude Code than your Claude Code worker image has: coding runs on it then fail as providerrejected (no cost) while native runs work. Upgrade wardby and rebuild the worker images before using the model in Claude Code agents. Catalog changes take effect on the writing process immediately, and on every other wardby process within WARDBYMODELCATALOGREFRESHSECONDS (default 45). A run already in progress keeps the catalog entry it started with, so disabling or repricing a model never changes a run already under way — only new runs. If this deployment delegates to an identity provider, define models:admin there before relying on it, and map the model-manager role (or admin) to the people who maintain pricing — see Configure identity and privileged access. If a run can't use a model, see Model not available.",
1188
+ "markdown": "\n# Models and pricing\n\nWardby prices and routes every model call from one catalog: the models this\nrelease ships, overlaid with this deployment's own additions and overrides.\nAn agent can use a model only when it's in the catalog and its provider has\ncredentials configured here.\n\n## The five tools\n\n| Tool | Scope | What it does |\n| --------------- | -------------- | ----------------------------------------------------------------------------------------- |\n| `list_models` | `agents:read` | Every active catalog entry (or, with `includeDisabled: true`, disabled ones too). |\n| `get_model` | `agents:read` | One entry by `modelId`; for an override, also the shipped entry it shadows. |\n| `set_model` | `models:admin` | Adds or completely replaces one entry. Every field is required. |\n| `disable_model` | `models:admin` | Removes a model from routing without deleting its pricing history. |\n| `reset_model` | `models:admin` | Removes every row for a model id, reverting to the shipped entry (if any) or removing it. |\n\nReading the catalog needs only `agents:read` — no secrets live in an entry.\nChanging it needs `models:admin`, honored only for a caller whose Wardby role\ngrants it: `admin`, or the narrower `model-manager` role.\n\nAn entry's `origin` is `shipped` or `override`; `routable` says whether this\ndeployment can actually route to it right now; `shippedDiffers` (overrides of\na shipped model only) says whether your override has drifted from the\ncurrent shipped values. See [`docs/models.md`](../docs/models.md) for the\nfull field reference.\n\n## Adding or overriding a model\n\n```json\n{\n \"provider\": \"anthropic\",\n \"modelId\": \"claude-example-model\",\n \"encoding\": \"o200k_base\",\n \"inputPerMTok\": 0.0,\n \"outputPerMTok\": 0.0,\n \"cachedInputPerMTok\": 0.0,\n \"cacheWritePerMTok\": 0.0,\n \"efforts\": [\"low\", \"medium\", \"high\"],\n \"thinkingMode\": \"adaptive\",\n \"sourceUrl\": \"https://example.com/replace-with-the-providers-own-pricing-page\"\n}\n```\n\nThe rates above are placeholders. Copy the provider's own published rates for\nthat exact model — including its cache read and cache write rates — from its\npricing page, and point `sourceUrl` at that page; never compute cache rates\nfrom `inputPerMTok` with a multiplier.\n\n`set_model` refuses (409) a `modelId` another provider already owns. A\nshipped id always belongs to its shipped provider — permanently; no other\nprovider can ever claim it, not even by disabling or resetting the override.\nA non-shipped id already claimed by another provider (its row active or\ndisabled) is freed only by `reset_model` — it takes only the `modelId` and\nclears every row for it, whichever provider owns it; `disable_model` alone\nnever frees it, since the disabled row still reserves the id.\n\n`thinkingMode` (`adaptive`, `manual`, or `none`) must match what the exact\nmodel accepts. Getting it wrong doesn't fail at `set_model` — it fails later,\nwhen a run calls the model, with `unsupported_anthropic_feature`.\nFor Claude Code coding runs, `efforts` is also the exact set of levels a run\nmay send, so always include the model's default effort level.\n\nOpenAI models keep `thinkingMode` `none`; `efforts` alone drives\n`reasoning.effort` on native agents, which call OpenAI's Responses API\n(`store: false`). `gpt-5.6-sol`/`terra`/`luna` and `gpt-6-astra` list `low`\nthrough `max`; `gpt-4o`, `gpt-4o-mini` and the `gpt-4.1` family list none.\nReasoning models reject sampling parameters such as `temperature`; wardby\ndoesn't send them. Reasoning tokens bill as output tokens, charged when the\ncall ends, so one call can take a run past its budget by its reasoning; higher\neffort raises that ceiling.\n\nAfter an upgrade, a `set_model` override of a shipped model keeps its stored\n`efforts` (it replaces the shipped entry wholesale). If those are empty,\n`create_agent`/`update_agent` reject an `effort` for the model: run\n`reset_model`, or `set_model` again with the new efforts, to pick up shipped\nchanges.\n\nA newly released Claude model may also need a newer Claude Code than your\nClaude Code worker image has: coding runs on it then fail as\n`provider_rejected` (no cost) while native runs work. Upgrade wardby and\nrebuild the worker images before using the model in Claude Code agents.\n\nCatalog changes take effect on the writing process immediately, and on every\nother wardby process within `WARDBY_MODEL_CATALOG_REFRESH_SECONDS` (default\n45). A run already in progress keeps the catalog entry it started with, so\ndisabling or repricing a model never changes a run already under way — only\nnew runs.\n\nIf this deployment delegates to an identity provider, define `models:admin`\nthere before relying on it, and map the `model-manager` role (or `admin`) to\nthe people who maintain pricing — see\n[Configure identity and privileged access](identity-and-access.md).\n\nIf a run can't use a model, see\n[Model not available](errors/model-unavailable.md).\n",
1189
+ "plainText": "Models and pricing Wardby prices and routes every model call from one catalog: the models this release ships, overlaid with this deployment's own additions and overrides. An agent can use a model only when it's in the catalog and its provider has credentials configured here. The five tools | Tool | Scope | What it does | | --------------- | -------------- | ----------------------------------------------------------------------------------------- | | listmodels | agents:read | Every active catalog entry (or, with includeDisabled: true, disabled ones too). | | getmodel | agents:read | One entry by modelId; for an override, also the shipped entry it shadows. | | setmodel | models:admin | Adds or completely replaces one entry. Every field is required. | | disablemodel | models:admin | Removes a model from routing without deleting its pricing history. | | resetmodel | models:admin | Removes every row for a model id, reverting to the shipped entry (if any) or removing it. | Reading the catalog needs only agents:read — no secrets live in an entry. Changing it needs models:admin, honored only for a caller whose Wardby role grants it: admin, or the narrower model-manager role. An entry's origin is shipped or override; routable says whether this deployment can actually route to it right now; shippedDiffers (overrides of a shipped model only) says whether your override has drifted from the current shipped values. See docs/models.md for the full field reference. Adding or overriding a model { \"provider\": \"anthropic\", \"modelId\": \"claude-example-model\", \"encoding\": \"o200kbase\", \"inputPerMTok\": 0.0, \"outputPerMTok\": 0.0, \"cachedInputPerMTok\": 0.0, \"cacheWritePerMTok\": 0.0, \"efforts\": [\"low\", \"medium\", \"high\"], \"thinkingMode\": \"adaptive\", \"sourceUrl\": \"https://example.com/replace-with-the-providers-own-pricing-page\" } The rates above are placeholders. Copy the provider's own published rates for that exact model — including its cache read and cache write rates — from its pricing page, and point sourceUrl at that page; never compute cache rates from inputPerMTok with a multiplier. setmodel refuses (409) a modelId another provider already owns. A shipped id always belongs to its shipped provider — permanently; no other provider can ever claim it, not even by disabling or resetting the override. A non-shipped id already claimed by another provider (its row active or disabled) is freed only by resetmodel — it takes only the modelId and clears every row for it, whichever provider owns it; disablemodel alone never frees it, since the disabled row still reserves the id. thinkingMode (adaptive, manual, or none) must match what the exact model accepts. Getting it wrong doesn't fail at setmodel — it fails later, when a run calls the model, with unsupportedanthropicfeature. For Claude Code coding runs, efforts is also the exact set of levels a run may send, so always include the model's default effort level. OpenAI models keep thinkingMode none; efforts alone drives reasoning.effort on native agents, which call OpenAI's Responses API (store: false). gpt-5.6-sol/terra/luna and gpt-6-astra list low through max; gpt-4o, gpt-4o-mini and the gpt-4.1 family list none. Reasoning models reject sampling parameters such as temperature; wardby doesn't send them. Reasoning tokens bill as output tokens, charged when the call ends, so one call can take a run past its budget by its reasoning; higher effort raises that ceiling. After an upgrade, a setmodel override of a shipped model keeps its stored efforts (it replaces the shipped entry wholesale). If those are empty, createagent/updateagent reject an effort for the model: run resetmodel, or setmodel again with the new efforts, to pick up shipped changes. A newly released Claude model may also need a newer Claude Code than your Claude Code worker image has: coding runs on it then fail as providerrejected (no cost) while native runs work. Upgrade wardby and rebuild the worker images before using the model in Claude Code agents. Catalog changes take effect on the writing process immediately, and on every other wardby process within WARDBYMODELCATALOGREFRESHSECONDS (default 45). A run already in progress keeps the catalog entry it started with, so disabling or repricing a model never changes a run already under way — only new runs. If this deployment delegates to an identity provider, define models:admin there before relying on it, and map the model-manager role (or admin) to the people who maintain pricing — see Configure identity and privileged access. If a run can't use a model, see Model not available.",
1190
1190
  "headings": [
1191
1191
  {
1192
1192
  "level": 1,
@@ -14,5 +14,5 @@
14
14
  * literally.
15
15
  */
16
16
  import type { CatalogEntry } from "./catalog-types.js";
17
- export declare const SHIPPED_CATALOG_VERSION = "2026-10-03";
17
+ export declare const SHIPPED_CATALOG_VERSION = "2026-10-08";
18
18
  export declare const SHIPPED_CATALOG: readonly CatalogEntry[];
@@ -1,4 +1,4 @@
1
- export const SHIPPED_CATALOG_VERSION = "2026-10-03";
1
+ export const SHIPPED_CATALOG_VERSION = "2026-10-08";
2
2
  const ALL_EFFORTS = ["low", "medium", "high", "xhigh", "max"];
3
3
  export const SHIPPED_CATALOG = [
4
4
  // OpenAI
@@ -65,7 +65,7 @@ export const SHIPPED_CATALOG = [
65
65
  outputPerMTok: 50.0,
66
66
  cachedInputPerMTok: 1.0,
67
67
  cacheWritePerMTok: 12.5,
68
- efforts: [],
68
+ efforts: ALL_EFFORTS,
69
69
  thinkingMode: "none",
70
70
  },
71
71
  {
@@ -76,7 +76,7 @@ export const SHIPPED_CATALOG = [
76
76
  outputPerMTok: 20.0,
77
77
  cachedInputPerMTok: 0.4,
78
78
  cacheWritePerMTok: 5.0,
79
- efforts: [],
79
+ efforts: ALL_EFFORTS,
80
80
  thinkingMode: "none",
81
81
  },
82
82
  {
@@ -87,7 +87,7 @@ export const SHIPPED_CATALOG = [
87
87
  outputPerMTok: 12.0,
88
88
  cachedInputPerMTok: 0.2,
89
89
  cacheWritePerMTok: 2.5,
90
- efforts: [],
90
+ efforts: ALL_EFFORTS,
91
91
  thinkingMode: "none",
92
92
  },
93
93
  {
@@ -98,7 +98,7 @@ export const SHIPPED_CATALOG = [
98
98
  outputPerMTok: 1.2,
99
99
  cachedInputPerMTok: 0.02,
100
100
  cacheWritePerMTok: 0.25,
101
- efforts: [],
101
+ efforts: ALL_EFFORTS,
102
102
  thinkingMode: "none",
103
103
  },
104
104
  // Anthropic (direct API)
@@ -1,7 +1,30 @@
1
1
  /**
2
- * OpenAI adapter for the `LlmProvider` seam — Phase 1's only concrete LLM
3
- * adapter. The `openai` npm package is used only inside this file; core
4
- * never imports it.
2
+ * OpenAI adapter for the `LlmProvider` seam. The `openai` npm package is used
3
+ * only inside this file; core never imports it.
4
+ *
5
+ * It speaks the Responses API, not Chat Completions: Chat Completions rejects
6
+ * any request carrying tools on gpt-5.6-* and gpt-6-*, while Responses takes
7
+ * tools on every OpenAI model and reasoning effort on the reasoning ones.
8
+ *
9
+ * Mapping notes (what the Responses API cannot carry, and why that's fine):
10
+ * - `stopSequences` is dropped — Responses has no stop parameter ("Unknown
11
+ * parameter" on every model). No caller sets it today.
12
+ * - `LlmMessage.name` is dropped — Responses has no per-message name. The
13
+ * engine sets it only on tool results, where `call_id` already correlates.
14
+ * - `store: false` is always sent — Responses stores every response by
15
+ * default, and wardby replays the whole conversation itself each turn.
16
+ * - Reasoning items are not replayed between turns: with `store: false` the
17
+ * next turn starts without the previous turn's hidden reasoning, only the
18
+ * visible text and function_call/function_call_output items. That costs
19
+ * some reasoning continuity, not correctness.
20
+ * - Reasoning tokens never stream as deltas, so the engine's mid-stream budget
21
+ * cap (driven by text deltas) can't see them. They arrive inside
22
+ * `output_tokens` on the final usage and are charged when the call ends, at
23
+ * the output rate, which is how OpenAI bills them. There is no per-call
24
+ * reservation: one call can take a run past its budget by that call's
25
+ * reasoning, and higher effort levels raise that ceiling. The pre-turn
26
+ * input-estimate gate and the run-level budget-group reservation still
27
+ * apply.
5
28
  */
6
29
  import OpenAI from "openai";
7
30
  import type { LlmMessage, LlmRequest, LlmStreamEvent, LlmToolDef } from "./types.js";
@@ -22,6 +45,14 @@ export declare class OpenAiLlmProvider implements CatalogLlmAdapter {
22
45
  constructor(apiKey?: string, client?: OpenAI, lookup?: CatalogLookup);
23
46
  withEntry(entry: CatalogEntry): OpenAiLlmProvider;
24
47
  stream(req: LlmRequest, signal?: AbortSignal): AsyncIterable<LlmStreamEvent>;
48
+ /**
49
+ * OpenAI reports both `cached_tokens` and `cache_write_tokens` as subsets
50
+ * of `input_tokens`. wardby's LlmUsage convention (claude-messages.ts,
51
+ * computeCost) is that `inputTokens` includes cache reads but NOT cache
52
+ * writes, which are billed separately at the write rate — so the writes
53
+ * come out of `inputTokens` here, or computeCost would charge them twice.
54
+ */
55
+ private toUsage;
25
56
  countTokens(model: string, messages: LlmMessage[], tools?: LlmToolDef[]): Promise<number>;
26
57
  priceUsd(model: string, usage: {
27
58
  inputTokens: number;
@@ -1,7 +1,30 @@
1
1
  /**
2
- * OpenAI adapter for the `LlmProvider` seam — Phase 1's only concrete LLM
3
- * adapter. The `openai` npm package is used only inside this file; core
4
- * never imports it.
2
+ * OpenAI adapter for the `LlmProvider` seam. The `openai` npm package is used
3
+ * only inside this file; core never imports it.
4
+ *
5
+ * It speaks the Responses API, not Chat Completions: Chat Completions rejects
6
+ * any request carrying tools on gpt-5.6-* and gpt-6-*, while Responses takes
7
+ * tools on every OpenAI model and reasoning effort on the reasoning ones.
8
+ *
9
+ * Mapping notes (what the Responses API cannot carry, and why that's fine):
10
+ * - `stopSequences` is dropped — Responses has no stop parameter ("Unknown
11
+ * parameter" on every model). No caller sets it today.
12
+ * - `LlmMessage.name` is dropped — Responses has no per-message name. The
13
+ * engine sets it only on tool results, where `call_id` already correlates.
14
+ * - `store: false` is always sent — Responses stores every response by
15
+ * default, and wardby replays the whole conversation itself each turn.
16
+ * - Reasoning items are not replayed between turns: with `store: false` the
17
+ * next turn starts without the previous turn's hidden reasoning, only the
18
+ * visible text and function_call/function_call_output items. That costs
19
+ * some reasoning continuity, not correctness.
20
+ * - Reasoning tokens never stream as deltas, so the engine's mid-stream budget
21
+ * cap (driven by text deltas) can't see them. They arrive inside
22
+ * `output_tokens` on the final usage and are charged when the call ends, at
23
+ * the output rate, which is how OpenAI bills them. There is no per-call
24
+ * reservation: one call can take a run past its budget by that call's
25
+ * reasoning, and higher effort levels raise that ceiling. The pre-turn
26
+ * input-estimate gate and the run-level budget-group reservation still
27
+ * apply.
5
28
  */
6
29
  import OpenAI from "openai";
7
30
  import { encode as encodeCl100kBase } from "gpt-tokenizer/encoding/cl100k_base";
@@ -29,13 +52,46 @@ function encodeForModel(model, text, lookup) {
29
52
  const { encoding } = lookup(model);
30
53
  return encoding === "o200k_base" ? encodeO200kBase(text) : encodeCl100kBase(text);
31
54
  }
32
- /** Same shape sent to the API in `stream()` — kept as one function so the estimate can never drift from what's actually serialized. */
55
+ /**
56
+ * Same shape sent to the API in `stream()` — kept as one function so the
57
+ * estimate can never drift from what's actually serialized. `strict: false`
58
+ * because wardby's tool schemas aren't written to strict mode's rules (every
59
+ * property required, additionalProperties false), and Responses defaults
60
+ * strict on.
61
+ */
33
62
  function toOpenAiTools(tools) {
34
63
  return tools.map((t) => ({
35
64
  type: "function",
36
- function: { name: t.name, description: t.description, parameters: t.parameters },
65
+ name: t.name,
66
+ description: t.description,
67
+ parameters: t.parameters,
68
+ strict: false,
37
69
  }));
38
70
  }
71
+ // Responses rejects max_output_tokens below 16 with a 400 ("Expected a value
72
+ // >= 16"). Raising a smaller cap to 16 costs at most a few output tokens;
73
+ // failing the call outright would cost the whole turn.
74
+ const MIN_MAX_OUTPUT_TOKENS = 16;
75
+ /** One `LlmMessage` becomes zero or more Responses input items, in order. */
76
+ function toResponsesInput(messages) {
77
+ const input = [];
78
+ for (const m of messages) {
79
+ if (m.role === "tool") {
80
+ input.push({ type: "function_call_output", call_id: m.toolCallId ?? "", output: m.content });
81
+ continue;
82
+ }
83
+ const toolCalls = m.role === "assistant" ? (m.toolCalls ?? []) : [];
84
+ // An assistant turn that only made tool calls has empty text; an empty
85
+ // assistant message item adds nothing, so omit it.
86
+ if (!(m.content === "" && toolCalls.length > 0)) {
87
+ input.push({ role: m.role, content: m.content });
88
+ }
89
+ for (const tc of toolCalls) {
90
+ input.push({ type: "function_call", call_id: tc.id, name: tc.name, arguments: tc.argsJson });
91
+ }
92
+ }
93
+ return input;
94
+ }
39
95
  // Fixed: this used to count message tokens only. When a request carries
40
96
  // `tools`, OpenAI also tokenizes the serialized tool JSON schemas into the
41
97
  // prompt — measured live at a 44-52% under-count for a small tool roster,
@@ -83,80 +139,83 @@ export class OpenAiLlmProvider {
83
139
  return new OpenAiLlmProvider("", this.client, pinnedLookup(entry));
84
140
  }
85
141
  async *stream(req, signal) {
86
- const stream = await this.client.chat.completions.create({
142
+ const entry = this.lookup(req.model);
143
+ // Effort is sent only at a level the catalog says the model accepts and
144
+ // dropped otherwise (the LlmRequest.effort contract). Temperature only
145
+ // goes to non-reasoning models: reasoning models reject it ("not
146
+ // supported with this model").
147
+ const sendEffort = req.effort !== undefined && entry.efforts.includes(req.effort);
148
+ const sendTemperature = req.temperature !== undefined && entry.efforts.length === 0;
149
+ const params = {
87
150
  model: req.model,
88
- messages: req.messages.map((m) => ({
89
- role: m.role,
90
- content: m.content,
91
- ...(m.name ? { name: m.name } : {}),
92
- ...(m.toolCallId ? { tool_call_id: m.toolCallId } : {}),
93
- ...(m.toolCalls && m.toolCalls.length > 0
94
- ? {
95
- tool_calls: m.toolCalls.map((tc) => ({
96
- id: tc.id,
97
- type: "function",
98
- function: { name: tc.name, arguments: tc.argsJson },
99
- })),
100
- }
101
- : {}),
102
- })),
103
- tools: req.tools ? toOpenAiTools(req.tools) : undefined,
104
- max_tokens: req.maxTokens,
105
- temperature: req.temperature,
106
- stop: req.stopSequences,
151
+ input: toResponsesInput(req.messages),
107
152
  stream: true,
108
- stream_options: { include_usage: true },
109
- }, { signal });
110
- let stopReason = "stop";
111
- // OpenAI streams tool-call arguments fragmented across chunks, keyed by
112
- // index — buffer per index until the turn's finish_reason confirms the
113
- // call is complete, then emit one `tool_call` event per call.
114
- const toolCallBuffers = new Map();
115
- for await (const chunk of stream) {
116
- const choice = chunk.choices[0];
117
- if (choice?.delta?.content) {
118
- yield { type: "text", delta: choice.delta.content };
119
- }
120
- if (choice?.delta?.tool_calls) {
121
- for (const fragment of choice.delta.tool_calls) {
122
- const buffered = toolCallBuffers.get(fragment.index) ?? { id: "", name: "", argsJson: "" };
123
- if (fragment.id)
124
- buffered.id = fragment.id;
125
- if (fragment.function?.name)
126
- buffered.name += fragment.function.name;
127
- if (fragment.function?.arguments)
128
- buffered.argsJson += fragment.function.arguments;
129
- toolCallBuffers.set(fragment.index, buffered);
130
- }
131
- }
132
- if (choice?.finish_reason) {
133
- stopReason = choice.finish_reason;
134
- if (stopReason === "tool_calls") {
135
- for (const toolCall of toolCallBuffers.values()) {
136
- yield { type: "tool_call", id: toolCall.id, name: toolCall.name, argsJson: toolCall.argsJson };
153
+ store: false,
154
+ ...(req.tools ? { tools: toOpenAiTools(req.tools) } : {}),
155
+ ...(req.maxTokens !== undefined ? { max_output_tokens: Math.max(req.maxTokens, MIN_MAX_OUTPUT_TOKENS) } : {}),
156
+ // The SDK's ReasoningEffort type predates xhigh/max; the API accepts them.
157
+ ...(sendEffort ? { reasoning: { effort: req.effort } } : {}),
158
+ ...(sendTemperature ? { temperature: req.temperature } : {}),
159
+ };
160
+ const stream = await this.client.responses.create(params, { signal });
161
+ let sawToolCall = false;
162
+ for await (const event of stream) {
163
+ switch (event.type) {
164
+ case "response.output_text.delta":
165
+ yield { type: "text", delta: event.delta };
166
+ break;
167
+ case "response.output_item.done":
168
+ // Arguments also stream as function_call_arguments.delta fragments,
169
+ // but the done item carries the complete call — emit from that.
170
+ // An "incomplete" item was cut off (e.g. by max_output_tokens) and
171
+ // its arguments are truncated JSON, so it is not a call to run.
172
+ if (event.item.type === "function_call" && event.item.status !== "incomplete") {
173
+ sawToolCall = true;
174
+ yield { type: "tool_call", id: event.item.call_id, name: event.item.name, argsJson: event.item.arguments };
137
175
  }
138
- toolCallBuffers.clear();
176
+ break;
177
+ case "response.completed":
178
+ case "response.incomplete": {
179
+ const stopReason = sawToolCall
180
+ ? "tool_calls"
181
+ : event.response.incomplete_details?.reason === "max_output_tokens"
182
+ ? "length"
183
+ : "stop";
184
+ yield { type: "done", stopReason, usage: this.toUsage(req.model, event.response.usage) };
185
+ break;
139
186
  }
140
- }
141
- if (chunk.usage) {
142
- // OpenAI's prompt caching is automatic and read-only — no billed
143
- // "cache write" step, so cacheWriteTokens stays unset here. Other
144
- // future adapters (e.g. Bedrock/Claude) may report and bill one.
145
- const cachedInputTokens = chunk.usage.prompt_tokens_details?.cached_tokens ?? 0;
146
- const usage = {
147
- inputTokens: chunk.usage.prompt_tokens,
148
- outputTokens: chunk.usage.completion_tokens,
149
- cachedInputTokens,
150
- costUsd: this.priceUsd(req.model, {
151
- inputTokens: chunk.usage.prompt_tokens,
152
- outputTokens: chunk.usage.completion_tokens,
153
- cachedInputTokens,
154
- }),
155
- };
156
- yield { type: "done", stopReason, usage };
187
+ case "response.failed":
188
+ throw new Error(`OpenAI response failed: ${event.response.error?.message ?? "no error message"}`);
189
+ case "error":
190
+ // The SDK usually throws an APIError when an SSE error arrives,
191
+ // before yielding it; this branch is a backstop.
192
+ throw new Error(`OpenAI stream error: ${event.message}`);
193
+ default:
194
+ break;
157
195
  }
158
196
  }
159
197
  }
198
+ /**
199
+ * OpenAI reports both `cached_tokens` and `cache_write_tokens` as subsets
200
+ * of `input_tokens`. wardby's LlmUsage convention (claude-messages.ts,
201
+ * computeCost) is that `inputTokens` includes cache reads but NOT cache
202
+ * writes, which are billed separately at the write rate — so the writes
203
+ * come out of `inputTokens` here, or computeCost would charge them twice.
204
+ */
205
+ toUsage(model, raw) {
206
+ const cachedInputTokens = raw?.input_tokens_details?.cached_tokens ?? 0;
207
+ const cacheWriteTokens = raw?.input_tokens_details?.cache_write_tokens ?? 0;
208
+ const inputTokens = Math.max((raw?.input_tokens ?? 0) - cacheWriteTokens, 0);
209
+ // output_tokens already includes reasoning tokens.
210
+ const outputTokens = raw?.output_tokens ?? 0;
211
+ const tokens = {
212
+ inputTokens,
213
+ outputTokens,
214
+ cachedInputTokens,
215
+ ...(cacheWriteTokens > 0 ? { cacheWriteTokens } : {}),
216
+ };
217
+ return { ...tokens, costUsd: this.priceUsd(model, tokens) };
218
+ }
160
219
  async countTokens(model, messages, tools) {
161
220
  return estimateTokens(model, messages, tools, this.lookup);
162
221
  }
@@ -1 +1 @@
1
- {"runtime":"ghcr.io/wardby/wardby/wardby-runtime@sha256:1b6646eccf13dbe534257f641b757cd11a127ef54741cf7b7a0a56fc53de60da","worker":"ghcr.io/wardby/wardby/wardby-coding-worker@sha256:0a9d4a8f737b4b2badde4dfad5646db3cb26ead4f3783ee0c13819daf0f44b1a","claudeWorker":"ghcr.io/wardby/wardby/wardby-claude-coding-worker@sha256:36fdf3b28d60a332b170ec5bd52d7158d11cab70562d7cadffe2fa286ccff974","claudeToolRunner":"ghcr.io/wardby/wardby/wardby-claude-tool-runner@sha256:68c276dce5341bf5e8dd3723ae8ad45ae01c81065a99ee2c5ca9847ce1afe0d7"}
1
+ {"runtime":"ghcr.io/wardby/wardby/wardby-runtime@sha256:2747fa03de5b94f03a3cb493888e2a03d9eec59b59e1ac28f55a91c63ed073a9","worker":"ghcr.io/wardby/wardby/wardby-coding-worker@sha256:29908387f38d1307b6050cd322b9abb5576ac7b403a482b8dd8a2dc4fb8eea47","claudeWorker":"ghcr.io/wardby/wardby/wardby-claude-coding-worker@sha256:b8495a54f24a69142d99faf534dc6471526ae999008f5807b1e9bf24d4eab957","claudeToolRunner":"ghcr.io/wardby/wardby/wardby-claude-tool-runner@sha256:cb125592be91ebc660b342a0ecfed6744b9b0626445607737f3deb6a5801faaf"}
@@ -221,9 +221,11 @@ sent on every model call and trades depth of reasoning against latency and
221
221
  output-token cost; lower levels are faster and cheaper per turn. Leave it unset
222
222
  to use the provider's default. Wardby rejects a level the agent's model does not
223
223
  accept, including when you later change the model (clear it with
224
- `effort: null`). Effort currently applies to direct Anthropic API models that
225
- support it; OpenAI and Bedrock models accept no effort setting, and coding
226
- agents do not use it.
224
+ `effort: null`). Effort applies to direct Anthropic API models that support
225
+ it and to the OpenAI reasoning models (`gpt-5.6-sol`, `gpt-5.6-terra`,
226
+ `gpt-5.6-luna`, `gpt-6-astra`), which accept `low` through `max`; the older
227
+ OpenAI models (`gpt-4o`, `gpt-4o-mini`, and the `gpt-4.1` family) and Bedrock
228
+ models accept no effort setting. Coding agents do not use it.
227
229
 
228
230
  ## Coding agents
229
231
 
package/docs/models.md CHANGED
@@ -130,7 +130,24 @@ and must match what that exact model actually accepts:
130
130
  - `manual` — a fixed thinking budget (`{"type": "enabled", "budget_tokens": …}`)
131
131
  and no effort level at all. Some smaller Claude models reject `adaptive`
132
132
  entirely.
133
- - `none` — no thinking parameter is sent. Non-Claude models.
133
+ - `none` — no thinking parameter is sent. Non-Claude models, including OpenAI.
134
+
135
+ OpenAI models always use `none`; their reasoning is controlled by `efforts`
136
+ alone. Native agents call OpenAI's Responses API, and when the agent has an
137
+ `effort` that the model's `efforts` lists, wardby sends it as
138
+ `reasoning.effort`. An empty `efforts` means no reasoning parameter is sent,
139
+ and an agent with no `effort` gets OpenAI's default. The shipped reasoning
140
+ models (`gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-5.6-luna`, `gpt-6-astra`) list
141
+ `low`, `medium`, `high`, `xhigh`, and `max`; the older models (`gpt-4o`,
142
+ `gpt-4o-mini`, the `gpt-4.1` family) list none. OpenAI reasoning models reject
143
+ sampling parameters such as `temperature`; wardby doesn't send them. Requests
144
+ are sent with `store: false`, so OpenAI doesn't store the responses for later
145
+ retrieval. Reasoning tokens are billed as output tokens. They don't stream, so
146
+ a call's reasoning is charged when the call ends, and one call can take a run
147
+ past its budget by that call's reasoning; higher effort levels raise that
148
+ ceiling. The pre-turn input estimate and the budget group's run reservation
149
+ still apply. OpenAI cache-write tokens are billed at the catalog's
150
+ `cacheWritePerMTok`.
134
151
 
135
152
  Setting the wrong `thinkingMode` (or listing `efforts` a model doesn't
136
153
  actually accept) does not fail at `set_model` time — it fails when a coding
@@ -219,3 +236,10 @@ shipped values automatically. `list_models`/`get_model`'s `shippedDiffers`
219
236
  field tells you when your override and the current shipped entry disagree,
220
237
  so you can decide whether to `reset_model` back to the shipped values or
221
238
  leave your override in place.
239
+
240
+ This matters when a release changes `efforts`: a database override replaces
241
+ the shipped entry wholesale, so an override of an OpenAI reasoning model
242
+ created before the release keeps its stored `efforts` (typically empty).
243
+ While it does, `create_agent` and `update_agent` reject any `effort` for that
244
+ model, and an existing agent's effort is not sent. Run `reset_model` on the id,
245
+ or `set_model` again with the new `efforts`, to pick up the shipped levels.
package/help/models.md CHANGED
@@ -70,6 +70,21 @@ when a run calls the model, with `unsupported_anthropic_feature`.
70
70
  For Claude Code coding runs, `efforts` is also the exact set of levels a run
71
71
  may send, so always include the model's default effort level.
72
72
 
73
+ OpenAI models keep `thinkingMode` `none`; `efforts` alone drives
74
+ `reasoning.effort` on native agents, which call OpenAI's Responses API
75
+ (`store: false`). `gpt-5.6-sol`/`terra`/`luna` and `gpt-6-astra` list `low`
76
+ through `max`; `gpt-4o`, `gpt-4o-mini` and the `gpt-4.1` family list none.
77
+ Reasoning models reject sampling parameters such as `temperature`; wardby
78
+ doesn't send them. Reasoning tokens bill as output tokens, charged when the
79
+ call ends, so one call can take a run past its budget by its reasoning; higher
80
+ effort raises that ceiling.
81
+
82
+ After an upgrade, a `set_model` override of a shipped model keeps its stored
83
+ `efforts` (it replaces the shipped entry wholesale). If those are empty,
84
+ `create_agent`/`update_agent` reject an `effort` for the model: run
85
+ `reset_model`, or `set_model` again with the new efforts, to pick up shipped
86
+ changes.
87
+
73
88
  A newly released Claude model may also need a newer Claude Code than your
74
89
  Claude Code worker image has: coding runs on it then fail as
75
90
  `provider_rejected` (no cost) while native runs work. Upgrade wardby and
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@wardby/cli",
3
- "version": "0.5.1",
3
+ "version": "0.5.2",
4
4
  "description": "Self-hosted control plane for budget-guarded AI agents.",
5
5
  "license": "Apache-2.0",
6
6
  "homepage": "https://github.com/wardby/wardby#readme",