rightmodeler 0.2.1 → 0.4.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,109 @@
1
+ # Gateways
2
+
3
+ Rightmodeler replays and judges through any OpenAI-compatible base URL, so a gateway in front of your models can be the replay route. This guide covers open-source gateways verified at pinned releases.
4
+
5
+ ## What a replay route needs
6
+
7
+ - `--base-url` is the gateway's OpenAI-compatible root, ending in `/v1`, and `--api-key-env` names the variable holding the key the gateway expects.
8
+ - The gateway's `/v1/models` lists the replay models. When it lists ids without prices or capabilities, pass `--catalog-reference` with the upstream's public model list; see [Catalogs without pricing or capabilities](getting-started.md#catalogs-without-pricing-or-capabilities).
9
+ - Headers the gateway reads per request are passed with `--header 'name: value'`; see [Gateways that route by header](getting-started.md#gateways-that-route-by-header).
10
+ - The replay route runs no fallbacks, model aliases, response caching or request plugins for the replay models. Rightmodeler never asks a gateway for them, and it checks each replayed response (in Mode B, each response to a step whose model it sets): one that names another model, reports a cache hit, or reports a changed request is left out of the evidence as `attribution_substituted`; see [Which model answered](getting-started.md#which-model-answered).
11
+ - Replay latency includes the gateway hop; rightmodeler does not separate the gateway's share.
12
+ - Mode B on the cloud backend calls the base URL from a remote sandbox, so it cannot reach a gateway on localhost or a private network. Use the Docker backend for a local gateway.
13
+
14
+ ## Portkey
15
+
16
+ Verified on the open-source Portkey gateway 1.15.2 (MIT, `portkeyai/gateway:1.15.2`), the latest tagged release as of 2026-09-22. Portkey picks the upstream for each request from two headers and needs no server configuration:
17
+
18
+ ```sh
19
+ docker run -d --name portkey -p 127.0.0.1:8787:8787 portkeyai/gateway:1.15.2
20
+ npx rightmodeler init --traces ./traces.jsonl --repo . \
21
+ --base-url http://127.0.0.1:8787/v1 --api-key-env AI_GATEWAY_API_KEY \
22
+ --header 'x-portkey-provider: openai' \
23
+ --header 'x-portkey-custom-host: https://ai-gateway.vercel.sh/v1' \
24
+ --max-cost-usd 25
25
+ ```
26
+
27
+ - Use the `openai` provider with a custom host for any OpenAI-compatible upstream (Vercel AI Gateway, OpenRouter, a LiteLLM proxy). Portkey's `openrouter` provider rewrites requests and does not serve the model catalog.
28
+ - `--api-key-env` holds the upstream's key: Portkey forwards `Authorization` to the upstream and has no key of its own.
29
+ - The catalog, per-call cost, and rate-limit and credit errors come from the upstream through Portkey unchanged.
30
+ - Rightmodeler sends no `x-portkey-config`, so fallbacks, load balancing, retries, caching and guardrail or mutator hooks stay off. If you pass one with `--header`, a replayed response whose model differs (`override_params`, targets), that reports `x-portkey-cache-status: HIT`, or whose hook results show `transformed: true` is left out as `attribution_substituted`.
31
+ - Bind the container to 127.0.0.1, or start it headless (`docker run ... portkeyai/gateway:1.15.2 run start:node -- --headless`): the 1.15.2 console and live log stream are served without authentication and show provider keys.
32
+ - Portkey 1.15.2 keeps no traces or logs that rightmodeler can read. Export traces from your application: OpenTelemetry GenAI, AI SDK telemetry, or the OpenAI SDK JSONL shape.
33
+
34
+ ## Envoy AI Gateway (Agent Router)
35
+
36
+ Envoy AI Gateway was renamed Agent Router on 2026-09-09; its resources, `aigw` CLI and images keep their names. Verified on v1.1.0 (Apache 2.0) running standalone with `aigw run` (`envoyproxy/ai-gateway-cli:v1.1.0`); always use a release tag, because the `latest` image follows the development branch.
37
+
38
+ As a replay route:
39
+
40
+ - Declare each replay model as an `Exact` `x-ai-eg-model` match under the upstream's own id, one `backendRef`, no `modelNameOverride`, no priority fallback and no `BackendTrafficPolicy` retries on the replay route. A fallback or an override answers with another model, so a replayed response it answers is left out as `attribution_substituted`.
41
+ - Raise the Gateway's `ClientTrafficPolicy` `bufferLimit` (Envoy Gateway's 32 KiB default is too small for real prompts) and the route's `timeouts.request` for slow models.
42
+ - The gateway lists only the declared ids at `/v1/models`, so pass `--catalog-reference` with the upstream's public list.
43
+ - The gateway replaces `Authorization` with the route's key, so `--api-key-env` may name any non-empty variable.
44
+
45
+ ```sh
46
+ docker run --rm -p 127.0.0.1:1975:1975 -e AI_GATEWAY_API_KEY \
47
+ -e OTEL_EXPORTER_OTLP_ENDPOINT=http://collector:4318 \
48
+ -e 'OTEL_AIGW_SPAN_REQUEST_HEADER_ATTRIBUTES=agent-session-id:session.id,x-rightmodeler-family:rightmodeler.family,x-rightmodeler-replay:rightmodeler.replay' \
49
+ -v "$PWD/aigw.yaml:/config.yaml:ro" envoyproxy/ai-gateway-cli:v1.1.0 run /config.yaml
50
+ export AIGW_CLIENT_KEY=unused
51
+ npx rightmodeler init --traces ./spans.jsonl --repo . \
52
+ --base-url http://127.0.0.1:1975/v1 --api-key-env AIGW_CLIENT_KEY \
53
+ --catalog-reference https://ai-gateway.vercel.sh/v1/models \
54
+ --header 'x-rightmodeler-replay: 1' --max-cost-usd 25
55
+ ```
56
+
57
+ As a trace source, rightmodeler reads the gateway's default OpenInference spans from an OpenTelemetry collector's file export:
58
+
59
+ - Set `OTEL_EXPORTER_OTLP_ENDPOINT`. Map headers to span attributes with `OTEL_AIGW_SPAN_REQUEST_HEADER_ATTRIBUTES` as comma-separated `header:attribute` pairs; setting it replaces the default `agent-session-id:session.id`, so repeat that pair.
60
+ - Send `agent-session-id` from your application so the steps of one conversation form one ordered run, and `x-rightmodeler-family: <name>` so each call has its family. Send `x-rightmodeler-replay: 1` with rightmodeler's replays (`--header`) so a later export leaves them out.
61
+ - Each span's request body gives the model your application asked for and the conversation as sent; the response gives the output; the token counts give usage.
62
+ - Failed calls, replay-tagged calls, and calls whose prompt or output the gateway hid (`OPENINFERENCE_HIDE_INPUTS`, `OPENINFERENCE_HIDE_OUTPUTS`) are left out of the corpus with a `trace_steps_excluded` warning that names each reason.
63
+ - A span records the model that answered but not whether a priority fallback chose it (a `modelNameOverride` alias looks the same), so a call that fell back is read as an answer to the model your application asked for, with the fallback's output and usage. Keep fallback routes off the traffic you export, or expect those answers among the recorded outputs rightmodeler compares candidates against.
64
+ - Steps whose recorded conversation contains tool calls are read but not replayed yet (`recorded_messages_not_replayable`).
65
+ - With `AI_GATEWAY_TRACING_SEMCONV=gen_ai`, the spans are read by the OTel GenAI reader instead. Also set `OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=true`, because that convention records no messages without it and rightmodeler cannot read a span with none. Those spans are grouped by trace (propagate `traceparent` from your client), carry no `response_format` or `tool_choice`, and record images by type only.
66
+ - Access logs carry no message content and are not a trace source.
67
+
68
+ On Kubernetes (Kubernetes 1.32 or newer, Envoy Gateway 1.8.1 or newer, Helm charts `ai-gateway-crds-helm` and `ai-gateway-helm` v1.1.0), the same resources apply; set `OTEL_EXPORTER_OTLP_ENDPOINT` through the `ai-gateway-helm` chart's `extProc.extraEnvVars` and the header mapping through its `controller.spanRequestHeaderAttributes`. Mode B on the cloud backend cannot reach an in-cluster gateway.
69
+
70
+ ## Bifrost
71
+
72
+ Verified on the open-source Bifrost gateway transports/v2.2.1 (Apache 2.0, `maximhq/bifrost:v2.2.1`); pin the image, because releases arrive weekly. Enterprise features ship in a separate image; this guide uses the open-source one.
73
+
74
+ As a replay route:
75
+
76
+ - Model ids are `<provider>/<upstream id>`, for example `vercel/openai/gpt-4o-mini` for Vercel AI Gateway configured as an OpenAI-typed custom provider. Bifrost answers with the upstream id, which rightmodeler accepts as the requested model.
77
+ - Set every `compat` flag to `false` in the `client` block: a `client` block that omits them turns them all on, and the compat plugin can drop parameters such as `response_format` while answering 200.
78
+ - Configure no key `aliases` or routing rules for replay models; an alias answers with another model and is left out as `attribution_substituted`.
79
+ - Pass `--catalog-reference` with the upstream's public list. A custom provider lists ids and context only, with no prices or capabilities, and through Vercel AI Gateway every model carries the same `created` date: Bifrost copies Vercel's `created`, one placeholder for all models, and drops the real `released` date. Rightmodeler ranks judges partly by release date, and the reference's release dates win over the gateway's, so judges rank as they do on Vercel itself instead of the most expensive first. Bifrost's own pricing sheet is not used.
80
+ - The billed cost comes from the upstream through Bifrost's `usage.cost.total_cost`.
81
+ - Send `x-bf-cache-no-store: true` and `x-bf-dim-rightmodeler: replay` with rightmodeler's replays (`--header`): the first keeps replays out of a semantic cache (a cache hit is still detected and left out), the second tags them so a later log export leaves them out.
82
+
83
+ ```sh
84
+ docker volume create bifrost-data
85
+ docker create --name bifrost -p 127.0.0.1:8080:8080 -e AI_GATEWAY_API_KEY -v bifrost-data:/app/data maximhq/bifrost:v2.2.1
86
+ docker cp config.json bifrost:/app/data/config.json && docker start bifrost
87
+ export BIFROST_KEY=unused
88
+ npx rightmodeler init --traces ./bifrost-logs.jsonl --repo . \
89
+ --base-url http://127.0.0.1:8080/v1 --api-key-env BIFROST_KEY \
90
+ --catalog-reference https://ai-gateway.vercel.sh/v1/models \
91
+ --header 'x-bf-cache-no-store: true' --header 'x-bf-dim-rightmodeler: replay' \
92
+ --max-cost-usd 25
93
+ ```
94
+
95
+ As a trace source, rightmodeler reads Bifrost's own request logs, exported from its management API (add credentials once an admin exists):
96
+
97
+ ```sh
98
+ curl -s 'http://127.0.0.1:8080/api/logs?objects=chat_completion,chat_completion_stream&order=asc&limit=1000&offset=0' | jq -r '.logs[].id' \
99
+ | while read -r id; do curl -s "http://127.0.0.1:8080/api/logs/$id"; echo; done > bifrost-logs.jsonl
100
+ ```
101
+
102
+ With more than 1000 rows, repeat with `offset=1000`, `2000` and so on, appending (`>>`) to the same file, until a page returns fewer than 1000 rows; `order=asc` keeps earlier pages in place while new calls are logged. Do not pass `roots_only=true`: it hides fallback rows.
103
+
104
+ - Send `x-bf-session-id` from your application so the steps of one conversation form one ordered run, and `x-bf-dim-rightmodeler-family: <name>` so each call has its family.
105
+ - Each row gives the model the application asked for (`provider/model`, or the alias it sent), the conversation as sent (`input_history`), the output (`output_message`), usage, cost, latency and retries.
106
+ - Failed calls, answers from a configured fallback, rightmodeler's own replays and calls whose content was not logged are left out of the corpus with a `trace_steps_excluded` warning that names each reason. Steps whose conversation contains tool calls are read but not replayed yet.
107
+ - Bifrost's OpenTelemetry plugin export reads through the OTel GenAI reader, but it summarizes messages and loses tool-call ids: use the log export for anything but plain text calls.
108
+ - Streamed calls through Mode B are checked on the stream's model and headers only; Bifrost reports a semantic-cache hit only in the response body, so such a hit on a streamed Mode B call is not detected.
109
+ - Rightmodeler records every replay's latency through Bifrost, gateway hop included, and does not separate Bifrost's own overhead. Bifrost's published benchmark reports 11 microseconds of overhead on a t3.xlarge at 5,000 requests per second, excluding JSON marshalling and the HTTP call; measure it on your own traffic.
@@ -1,16 +1,23 @@
1
1
  # Getting started
2
2
 
3
- Rightmodeler analyzes recorded model calls, replays them against cheaper candidates, evaluates the outputs, and writes a recommendation report. The intended public package name is `rightmodeler`.
3
+ Rightmodeler analyzes recorded model calls, replays them against cheaper candidates, evaluates the outputs, and writes a recommendation report. Published on npm as `rightmodeler`.
4
4
 
5
5
  ## Requirements
6
6
 
7
7
  - Node.js 24 or newer.
8
8
  - A Git repository to analyze.
9
9
  - Trace input in a supported format.
10
- - An OpenAI-compatible provider base URL and the name of an environment variable containing its API key before replay begins.
10
+ - An OpenAI-compatible provider base URL and the name of an environment variable containing its API key before replay begins. Its `/v1/models` catalog should publish per-token pricing. OpenRouter and Vercel AI Gateway do. For a LiteLLM endpoint, Rightmodeler can fall back to `GET /model/info`; for a gateway that lists bare model ids or a direct OpenAI or Anthropic key, pass `--catalog-reference`; for another unpriced endpoint, pass `--pricing-file`.
11
11
 
12
- Supported trace sources are OTel GenAI, OpenAI JSONL, Langfuse, Braintrust,
13
- LangSmith, OpenInference, Helicone, W&B Weave, Claude Code, and Codex.
12
+ Instead of an API key, candidates or the judge can run through the `claude`
13
+ CLI you are signed in to on this machine, under your Claude plan:
14
+ `--route claude-login` with a `--judge-route` from another vendor. Run
15
+ `rightmodeler docs model-routes` for what a plan route measures, what it costs
16
+ your plan, and its safeguards.
17
+
18
+ Supported trace sources are OTel GenAI, AI SDK telemetry, OpenAI JSONL,
19
+ Langfuse, Braintrust, LangSmith, OpenInference, Helicone, Bifrost, W&B Weave,
20
+ Claude Code, and Codex.
14
21
 
15
22
  ## Start with automatic discovery
16
23
 
@@ -20,6 +27,13 @@ Run this from the repository you want to analyze:
20
27
  npx rightmodeler init
21
28
  ```
22
29
 
30
+ Before it looks for traces, an interactive `init` or `estimate` that will reach
31
+ replay asks how to call models: through the `claude` or `codex` CLI signed in on
32
+ this machine, or through an API key for OpenRouter, Vercel AI Gateway, OpenAI,
33
+ Anthropic or another OpenAI-compatible endpoint. It prints the equivalent flags
34
+ and saves the answer as the next default, which `--yes` applies without asking.
35
+ `npx rightmodeler docs model-routes` describes the choice.
36
+
23
37
  Rightmodeler checks conventional local trace files, Claude Code transcripts for
24
38
  the repository, and Codex sessions whose recorded working directory matches the
25
39
  repository. In an interactive terminal it lists matches newest-first with an
@@ -40,6 +54,27 @@ npx rightmodeler init --plan --output json --repo /path/to/repository
40
54
  npx rightmodeler init --through corpus --traces /path/to/traces.json --output json --repo /path/to/repository
41
55
  ```
42
56
 
57
+ `--traces` accepts a single file or a directory. A directory is read non-recursively as its `.json` and `.jsonl` files in name order; every file must use the same trace format.
58
+
59
+ Each family is replayed only against the call sites its own traces came from. Traced cases that cannot be tied to one such call site, because several call sites use the traced model or none matches it, are left out of the replay sample with a `family_cases_left_out` warning. A family with no case left abstains before any spend with `ambiguous_call_site_binding` or `unmatched_call_site_binding`. The AI SDK telemetry `functionId` (see below) is the way to tie an AI SDK call site to its family.
60
+
61
+ ## AI SDK telemetry
62
+
63
+ The AI SDK emits telemetry in two dialects, and Rightmodeler reads both:
64
+
65
+ - The `ai.*` dialect comes from AI SDK 5 and 6 with `experimental_telemetry: { isEnabled: true }` on each call, and from AI SDK 7 with `registerTelemetry(new LegacyOpenTelemetry())`. The AI SDK reader reads it.
66
+ - The GenAI semantic conventions dialect comes from AI SDK 7 with `registerTelemetry(new OpenTelemetry())`. The OTel GenAI reader reads it and treats the agent, step and tool spans as structure, so each model call counts once.
67
+
68
+ `registerTelemetry` comes from `ai`; `LegacyOpenTelemetry` and `OpenTelemetry` come from `@ai-sdk/otel`. Register only one of them: an export that holds both dialects is ambiguous. Keep `recordInputs` and `recordOutputs` on, which is the default, because a model call without its prompt or output cannot become a corpus case. Set a string-literal `functionId` on every call (`telemetry: { functionId: "summarize" }` in AI SDK 7, `experimental_telemetry: { isEnabled: true, functionId: "summarize" }` before it); it becomes the call's family.
69
+
70
+ The scanner records that `functionId` on the call site, and a family binds to exactly the call sites whose `functionId` equals its name: its evidence and any swap stay on those call sites, and no other family borrows them. Two call sites that do the same job may share one `functionId`. A family bound to a single call site can still be recommended. The `functionId` must be a string literal in the call; a variable or a template literal is not read, and the call site then binds by model id only. Replay cannot run a call site that needs tools or structured output: its cases are left out of the family's replay sample with a `family_cases_left_out` warning, and a family whose traced cases all come from such call sites abstains with `bound_call_sites_not_replayable` before any spend.
71
+
72
+ Export the spans through the OpenTelemetry NodeSDK or `@vercel/otel` to an OTLP collector, and pass the collector's file exporter output with `--traces`. A model call that ended without a finish reason, because it was aborted or errored, is left out of the corpus with a `trace_steps_excluded` warning, and the rest of the input is read. Token usage from AI SDK 4 exports (`ai.usage.promptTokens`) is not read, so those calls carry no usage.
73
+
74
+ ## OpenInference spans
75
+
76
+ When an OpenInference span carries the request body in `input.value` (Envoy AI Gateway, and the OpenAI instrumentation), rightmodeler reads the model your application asked for and the conversation exactly as sent from it, and the output from `output.value`. Spans that share a `session.id` form one run ordered by start time. A `rightmodeler.family` attribute names the family; otherwise the span name does. Failed calls, calls tagged `rightmodeler.replay`, and calls whose content was hidden are left out with a `trace_steps_excluded` warning.
77
+
43
78
  ## Run the complete pipeline
44
79
 
45
80
  ```sh
@@ -51,6 +86,8 @@ npx rightmodeler init --traces /path/to/traces.json --base-url https://provider.
51
86
  `--api-key-env <name>` to use a different exported variable. The CLI does not ask
52
87
  for a secret value.
53
88
 
89
+ Replay resends each recorded conversation as text. A recorded case whose conversation contains tool calls, non-text parts, or tool definitions is left out of the replay sample with a `recorded_messages_not_replayable` warning; the family's other cases replay, and a family left with too few cases abstains under the usual sample-size reasons.
90
+
54
91
  ## Estimate replay spend
55
92
 
56
93
  ```sh
@@ -59,7 +96,134 @@ npx rightmodeler estimate --traces /path/to/traces.json --base-url https://provi
59
96
 
60
97
  Estimate projects candidate replay spend from recorded token usage and the current
61
98
  model catalog before paid model calls begin.
99
+ `--max-cost-usd` caps candidate replays and judge calls together: each call reserves
100
+ its worst case before it is sent, and a call the cap cannot cover is not sent.
101
+
102
+ ## Static code context (Graphify)
103
+
104
+ rightmodeler can read a code graph built by the open-source Graphify CLI (PyPI package `graphifyy`, Apache-2.0, tested with 0.9.65). Graphify builds it locally from your source, with no account and no model call.
105
+
106
+ ```sh
107
+ uv tool install graphifyy
108
+ graphify update .
109
+ npx rightmodeler report --code-graph graphify-out/graph.json --repo .
110
+ ```
111
+
112
+ `init --code-graph <path>` renders the same section at the end of a run. For each call site the scanner found, the section lists the enclosing symbol, its callers, the tests that reach it, and the owners of those files, which are listed only. It also lists files that import an AI SDK where the scanner found no call site.
113
+
114
+ Graph edges are never replay trials, runtime proof, or quality evidence, and the flag never changes a stage before the report, a verdict, a gate, or confirmation. Each finding is labelled EXTRACTED, or INFERRED or AMBIGUOUS to verify, by its weakest hop. A graph built at another commit is shown file-level with a stale note. An unusable graph produces one warning, the section says why it is not shown, and the rest of the report is unchanged. The scan ignores `graphify-out/`, so building a graph never makes finished stages stale. Only `graphify update` and `graphify extract --code-only` are needed; other Graphify commands can call a language model.
115
+
116
+ `apply --code-graph <path>` appends the same section to the draft pull request body, limited to the call sites the pull request swaps and to five findings of each kind per call site. Owners there are listed only and are never requested as reviewers; reviewers still come from CODEOWNERS and blame. The graph never changes the swap, its digest, or its reviewers. `apply --dry-run` prints the exact body it would post, with or without `--code-graph`.
117
+
118
+ ## Release policy
119
+
120
+ `--policy <path>` is accepted by `init`, `estimate`, `replay`, and `confirm`. The JSON object can set the quality floor, shortlist size, and model allow and deny lists:
121
+
122
+ ```json
123
+ {
124
+ "qualityFloor": 0.9,
125
+ "shortlistTop": 5,
126
+ "allowModels": ["acme/small-1"],
127
+ "denyModels": ["acme/large-1"]
128
+ }
129
+ ```
130
+
131
+ `qualityFloor` must be greater than 0.8 and less than 1, and `shortlistTop` must be a positive integer. Changing the policy changes the stamped gate policy version, so shortlist and replay run again instead of pooling evidence gathered under the old policy.
132
+
133
+ ## Catalogs without pricing or capabilities
134
+
135
+ Rightmodeler reads per-token pricing from the model catalog. When every catalog
136
+ entry has null pricing and no `--pricing-file` is set, it requests LiteLLM
137
+ `GET /model/info` on the same host. Use `--pricing-file` when `/model/info` has
138
+ no usable per-token costs; file entries override provider pricing.
139
+
140
+ Some gateways answer `/v1/models` with bare model ids: Envoy AI Gateway lists
141
+ the models a route declares, and Bifrost lists custom providers with ids and
142
+ context only. Pass `--catalog-reference <url|path>` to name the upstream's own
143
+ model list, for example `https://ai-gateway.vercel.sh/v1/models` or
144
+ `https://openrouter.ai/api/v1/models`. Rightmodeler reads it without your key or
145
+ headers and joins it to the gateway's entries by id, or by the reference id a
146
+ gateway id ends with (`openai/gpt-4o-mini` for `vercel/openai/gpt-4o-mini`). A
147
+ gateway entry takes from its match only what it does not declare itself (price,
148
+ context window, output ceiling, tool and structured-output support); a match
149
+ that is not a language model removes the entry; `--pricing-file` values win over
150
+ both. Entries still unpriced after the join are named in a
151
+ `catalog_reference_unmatched` warning, and a reference that cannot be read stops
152
+ the run with `invalid_catalog_reference`.
153
+
154
+ The release date is the one field where the reference wins over the gateway:
155
+ when the matched reference entry has a release date, it replaces the date the
156
+ gateway declares, and the gateway's date is kept only when the reference gives
157
+ none. Rightmodeler ranks judges partly by how recent a model is, and a gateway
158
+ can give every model the same placeholder date (Bifrost does for Vercel AI
159
+ Gateway's models), which would leave context and price to rank the judges and
160
+ favor the most expensive.
161
+
162
+ ```sh
163
+ npx rightmodeler estimate --base-url https://provider.example/v1 --pricing-file /path/to/pricing.json --repo /path/to/repository
164
+ ```
165
+
166
+ The pricing file maps each model id to input and output USD per token and may
167
+ include the model's output ceiling:
168
+
169
+ ```json
170
+ {
171
+ "acme/model": {
172
+ "input": 0.000001,
173
+ "output": 0.000002,
174
+ "maxOutputTokens": 4096
175
+ }
176
+ }
177
+ ```
178
+
179
+ Without usable pricing from the catalog, a catalog reference, LiteLLM
180
+ `/model/info`, or a pricing file, the run refuses with `no_priced_candidates`
181
+ instead of reporting zero cost. The judge must be priced too, so price at least
182
+ one model from a family other than the current model's and the candidate's.
183
+
184
+ ### Direct OpenAI and Anthropic keys
185
+
186
+ Point `--base-url` at `https://api.openai.com/v1` or
187
+ `https://api.anthropic.com/v1` and name the variable that holds the key with
188
+ `--api-key-env`. Neither vendor's model list publishes prices, so also pass
189
+ `--catalog-reference https://ai-gateway.vercel.sh/v1/models`. Rightmodeler takes
190
+ the vendor from these two hosts: their bare ids are that vendor's models and
191
+ join the reference as `openai/<id>` or `anthropic/<id>`. Dated and dotted forms
192
+ of a name match within one vendor, in the reference and in your traces, so
193
+ `claude-haiku-4-5-20251001` matches `anthropic/claude-haiku-4.5`. For Anthropic,
194
+ rightmodeler sends `anthropic-version: 2023-06-01` (a
195
+ `--header 'anthropic-version: ...'` value wins), reads every page of the model
196
+ list, and never asks for structured output, which Anthropic's OpenAI-compatible
197
+ endpoint ignores. For OpenAI, rightmodeler caps a reply's length with
198
+ `max_completion_tokens`, which OpenAI's chat reference names in place of the
199
+ deprecated `max_tokens`; every other host still gets `max_tokens`.
200
+
201
+ The built-in judge must come from a vendor other than both the candidate's and
202
+ the recorded model's, and rightmodeler checks this before any paid call, also
203
+ when an unreachable `--evaluator` falls back to the built-in judge. One vendor's
204
+ key lists only that vendor's models, so a run on it alone stops with
205
+ `no_neutral_judge` unless the judge runs through the other vendor's CLI you are
206
+ signed in to (`--judge-route`, see `rightmodeler docs model-routes`) or your own
207
+ evaluator grades the replays. Bare ids from any other host name no vendor, and
208
+ a run on them stops with `judge_family_unknown`; use a gateway whose ids carry
209
+ their vendor (`vendor/model`), or `--evaluator`.
210
+
211
+ ## Which model answered
212
+
213
+ Rightmodeler checks every replay and judge response, and in Mode B every response to a step whose model it sets, before it counts. It records the model the response names, and it leaves a response out of the evidence when that model is not the one it asked for, when a gateway reports that it answered from its cache (Portkey's `x-portkey-cache-status: HIT`, Bifrost's `cache_debug.cache_hit`), or when a gateway reports that it changed the request (a Portkey hook with `transformed: true`, Bifrost's compat plugin dropping parameters). Such a response is recorded with `attribution: "substituted"`, is never graded, and counts as `attribution_substituted`; more than 5% of a family's replays substituted abstains the family. A `replay_responses_substituted` warning counts them.
214
+
215
+ A response names the requested model when it echoes it, when it drops a gateway provider prefix (`openai/gpt-4o-mini` for `vercel/openai/gpt-4o-mini`), or when it adds a dated snapshot (`gpt-4o-mini-2024-07-18` for `gpt-4o-mini`). An alias whose answering model has another name counts as substituted, so name replay models by their upstream ids and turn off fallbacks, model aliases, response caching and request plugins for the replay route. Completed replay cells are reused, so after fixing the route rerun with a fresh store (`--store <directory>`). For streamed Mode B calls only the model a stream names and the response headers are checked.
216
+
217
+ A judge whose response names another model is retired at that first response: rightmodeler starts no new call to it, cancels its calls still waiting for budget, and re-judges the affected replays with the next-ranked judge once the calls already under way return; the `judge_unusable` warning names the model that answered. Some catalogs also list a faster service tier of a model under its own id, which the gateway answers as the base model: on Vercel AI Gateway, `<id>-fast` is `<id>` at its fast tier. Rightmodeler never ranks such an id as a judge or shortlists it as a candidate while `<id>` is in the catalog too; an id that ends in `-fast` with no base model listed, such as `morph/morph-v3-fast`, is treated like any other model.
218
+
219
+ ## Gateways that route by header
220
+
221
+ Some gateways choose the upstream, the cache policy, or a trace tag from request headers. Pass each one with `--header 'name: value'`; repeat the option for more. Rightmodeler sends them with every request it makes to the provider base URL: the model catalog, candidate replays, judge calls, and the calls Mode B makes from your application, where a configured header replaces one your application sends. `authorization` comes only from `--api-key-env`, and `content-type`, `content-length` and `host` are set by rightmodeler, so none of them can be passed as a header. Completed replay cells are reused when the headers change, so after changing the route rerun with a fresh store (`--store <directory>`). A detached replay run is keyed to the header values by their SHA-256 digests; the values themselves are passed to the detached worker on its command line and are not written to the store. Do not put secrets in headers.
222
+
223
+ The default store is `.rightmodeler/` inside the analyzed repository. Completed stages resume when their inputs and outputs are still current. A complete run writes `.rightmodeler/project/reports/report.md`. The JSON report is kept inside the versioned store and is never written as a plain file, so read the final `result` event from `--output json` or `--output jsonl` for the machine-readable outcome.
224
+
225
+ Read the generated [command reference](commands.md), the [evaluator guide](evaluators.md), [Mode B configuration](modeb.md), the [gateway guide](gateways.md), and the [exit-code convention](exit-codes.md) before automating a full run.
62
226
 
63
- The default store is `.rightmodeler/` inside the analyzed repository. Completed stages resume when their inputs and outputs are still current. A complete run writes `.rightmodeler/project/reports/report.md` and `.rightmodeler/project/reports/report.json`.
227
+ To open the draft pull request and keep it reconciled, read the [GitHub guide](github.md).
64
228
 
65
- Read the generated [command reference](commands.md), the [evaluator guide](evaluators.md), [Mode B configuration](modeb.md), and the [exit-code convention](exit-codes.md) before automating a full run.
229
+ Run `rightmodeler docs <name>` to print any of these documents from the installed package.