@unotest/judge 0.24.0 → 0.25.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,116 @@
1
1
  # @unotest/judge
2
2
 
3
+ ## [0.25.0] - 2026-08-29
4
+
5
+ ### Minor Changes
6
+
7
+ - 70f2e11: Judge: provider edge cases, a startup preflight, and logs.
8
+
9
+ **Vertex `GOOGLE_CLOUD_LOCATION=global` could not work.** The endpoint was
10
+ built as `https://<location>-aiplatform.googleapis.com`, so `global` asked
11
+ `global-aiplatform.googleapis.com` — not a Vertex host. Google's edge
12
+ answered 404 with an HTML page, which arrived as "model not found". The
13
+ global endpoint has no region prefix; it is now special-cased. `global` is
14
+ the default in Google's own docs, so this was the first thing a Vertex user
15
+ hit. A model id containing a slash
16
+ (`publishers/anthropic/models/claude-haiku-4-5`) is now passed through
17
+ instead of being forced under `publishers/google`, and a pasted
18
+ `models/gemini-2.5-flash` no longer doubles the prefix on the Gemini API.
19
+
20
+ **Backend replies are classified instead of lumped together.**
21
+
22
+ - `401`/`403` → `config` with the variable to fix (`OPENAI_API_KEY`, ADC,
23
+ an expired `UNOTEST_JUDGE_ACCESS_TOKEN`), no longer indistinguishable
24
+ from a model fault.
25
+ - `429` and `5xx` → retried transparently with a short backoff that
26
+ honours `Retry-After`. A single rate limit used to kill a whole scenario.
27
+ The retry policy is a decorator over `JudgeProvider`, not a loop inside
28
+ the HTTP helper, so it covers every provider — `claude` spawns a process
29
+ and therefore had no re-send at all. A `claude` usage limit is
30
+ deliberately still reported rather than retried: the window is seconds,
31
+ the limit is not.
32
+ - A non-JSON error body says the endpoint is wrong rather than quoting an
33
+ HTML page back as the backend's opinion of the model. The body is kept
34
+ for `debug` logging.
35
+ - OpenAI `finish_reason: length` and Anthropic `stop_reason: max_tokens`
36
+ are named as token-limit truncation instead of surfacing as "did not
37
+ reply with the requested JSON verdict".
38
+ - The verdict parser accepts `"PASS"` and a JSON object embedded in prose —
39
+ providers without a JSON mode do both.
40
+
41
+ **`UNOTEST_JUDGE_VOTE=N` no longer fails on one bad ballot.** The N calls
42
+ ran under `Promise.all`, so voting — the reliability feature — tripled the
43
+ chance that one transient error killed the request. Errored ballots are now
44
+ dropped as long as a majority of the N calls still voted; a tie among the
45
+ survivors re-raises the failure rather than picking a side.
46
+
47
+ **The service checks its credentials before it says "listening".** Each
48
+ provider exposes a free probe (an ADC token exchange, `GET /models`,
49
+ `claude --version`) — no model call, no cost. A failing probe prints the
50
+ reason and exits non-zero (`UNOTEST_JUDGE_SKIP_PREFLIGHT=1` waives it), and
51
+ `npx @unotest/judge --check` runs just the probe for start scripts.
52
+ `GET /health` reports the same thing instead of an unconditional
53
+ `{"ok":true}`: `503` + `{"ok":false,"code":"config","error":"..."}`, cached
54
+ for a few seconds. The waiver covers the startup gate only — a service
55
+ started under it still answers `/health` from a real probe, and `--check`
56
+ still answers honestly, because the operator who reaches for the flag is
57
+ precisely the one whose credentials are dead. A static
58
+ `UNOTEST_JUDGE_ACCESS_TOKEN` is validated too, against Google's free
59
+ `tokeninfo`: those tokens live about an hour and refresh themselves never,
60
+ so "valid at startup" is a different question from "valid now".
61
+
62
+ ADC that expires on a daily reauth policy is a normal morning state, not a
63
+ one-off misconfiguration — it should not cost a five-minute run to
64
+ discover.
65
+
66
+ **New: `UNOTEST_JUDGE_LOG_LEVEL`** (falls back to the `UNOTEST_LOG_LEVEL`
67
+ `@unotest/web` already reads). At the default `info`, one line per judged
68
+ request — verdict, attempts, model, latency — with errors on stderr; a run
69
+ of 21 assertions with `VOTE=3` is 63 billed model calls that previously
70
+ left no trace. At `debug`, per call: the effective prompt (rubric
71
+ **including** any `UNOTEST_JUDGE_PREAMBLE`), each ballot with its
72
+ reasoning, and the raw provider reply. **The judged text is application
73
+ content** — debug writes it to stdout verbatim, which is why it is opt-in
74
+ and never the default.
75
+
76
+ An oversized request body now returns `413` with a reason; the socket used
77
+ to be destroyed while the client was still writing, so the client saw
78
+ `socket hang up`.
79
+
80
+ **`@unotest/web`:** the `judge:verdict` run artifact records `preamble`
81
+ beside `rubric` when one is set. The model saw the concatenation, but only
82
+ the split says which half came from the scenario and which from the
83
+ environment — without it, "a red step is explainable without a re-run" held
84
+ only for whoever also knew the env file. A judge request that outlives
85
+ `UNOTEST_JUDGE_TIMEOUT_MS` now says so, and names the service-side
86
+ `UNOTEST_JUDGE_CALL_TIMEOUT_MS` it may be racing.
87
+
88
+ The Claude CLI provider also stopped leaking a timer: the per-call budget
89
+ is now our own timer, cleared on every exit path. `child_process.spawn`'s
90
+ `timeout` option does not clear itself when the spawn fails, so a missing
91
+ binary held the event loop for the whole budget — 15s for a `--version`
92
+ probe, 120s for a verdict.
93
+
94
+ **Breaking (0.x minor):** `@unotest/judge` no longer exports `ExecFileFn` /
95
+ `ExecFileResult`; the Claude CLI provider takes `runImpl: RunProcess`
96
+ instead of `execImpl`, because the prompt moved from `argv` to stdin — a
97
+ long judged text used to overflow `ARG_MAX` and fail as a bare `E2BIG`.
98
+
99
+ ### Patch Changes
100
+
101
+ - Judge: the last `JUDGE_*` names in user-facing strings.
102
+
103
+ 0.24.0 renamed every service variable to `UNOTEST_JUDGE_*`, but three
104
+ messages kept telling people the old name: the service's `401`, the fake
105
+ provider's "switch to a real provider" hint, and the Claude CLI's
106
+ "binary not found" error. They now name the variable that actually
107
+ exists. Also landed after that release: `google-auth-library` was
108
+ declared but missing from the lockfile.
109
+
110
+ - Updated dependencies [70f2e11]
111
+ - Updated dependencies [adb9d72]
112
+ - @unotest/protocol@0.25.0
113
+
3
114
  ## [0.24.0] - 2026-08-26
4
115
 
5
116
  ### Minor Changes
package/README.md CHANGED
@@ -17,10 +17,22 @@ output.
17
17
  ```
18
18
  UNOTEST_JUDGE_PROVIDER=vertex \
19
19
  GOOGLE_CLOUD_PROJECT=my-project \
20
- GOOGLE_CLOUD_LOCATION=europe-west1 \
20
+ GOOGLE_CLOUD_LOCATION=us-central1 \
21
21
  npx @unotest/judge
22
22
  ```
23
23
 
24
+ On start the service checks the provider's credentials (an ADC token
25
+ exchange, a free `GET /models`, a `claude --version`) and refuses to
26
+ listen if they are dead — a judge that cannot judge should not report
27
+ itself as ready. `npx @unotest/judge --check` runs only that check and
28
+ exits `0`/`1`, for start scripts and CI.
29
+
30
+ `UNOTEST_JUDGE_SKIP_PREFLIGHT=1` waives the **startup gate** — the service
31
+ warns and listens anyway. It does not waive the truth: `/health` keeps
32
+ probing and keeps reporting the failure, and `--check` still answers
33
+ honestly. A green light on a judge that cannot judge is the fault this
34
+ whole check exists to remove.
35
+
24
36
  The service listens on `127.0.0.1:8790` by default. Point `@unotest/web`
25
37
  at it with `UNOTEST_JUDGE_MODE=remote` + `UNOTEST_JUDGE_URL=http://127.0.0.1:8790`,
26
38
  or skip the HTTP hop entirely with `UNOTEST_JUDGE_MODE=local` (in-process,
@@ -35,13 +47,28 @@ unotest through `npx`, with nothing installed locally, can only use
35
47
  workload identity, or `GOOGLE_APPLICATION_CREDENTIALS`). Requires the
36
48
  `google-auth-library` (a dependency of this package, loaded only when
37
49
  this provider runs) — or a static `UNOTEST_JUDGE_ACCESS_TOKEN` from
38
- `gcloud auth print-access-token`. Temperature is pinned to 0.
50
+ `gcloud auth print-access-token` (those live about an hour and are not
51
+ refreshed for you — the preflight validates one against Google's free
52
+ `tokeninfo` endpoint, so an expired token is caught at startup rather
53
+ than mid-run). Temperature is pinned to 0.
54
+
55
+ **Location and model are chosen together.** `GOOGLE_CLOUD_LOCATION` is
56
+ either `global` or a region (`us-central1`), and model availability
57
+ differs between them — `gemini-2.5-flash` is regional, newer flash-lite
58
+ models may be global-only. A model id containing a slash is used as-is
59
+ (`publishers/anthropic/models/claude-haiku-4-5`), which is how Vertex
60
+ addresses third-party models.
39
61
  - **`claude`** — the local Claude Code CLI (`claude -p`), spawned per
40
62
  judgement without shell interpretation. Auth comes from your Claude Code
41
63
  session, so a subscription works with no API key — this is the provider
42
64
  to reach for when no cloud credentials are at hand. `UNOTEST_JUDGE_MODEL`
43
65
  is passed as `--model` (aliases like `sonnet` work); unset uses the CLI's
44
- default.
66
+ default. **Budget for latency, not for tokens:** every verdict spawns the
67
+ CLI and loads a session, which measured ~6s per call here against ~1s for
68
+ an HTTP provider — 20 assertions is a couple of minutes. `UNOTEST_JUDGE_VOTE`
69
+ costs almost nothing extra (ballots run concurrently: 3 votes measured
70
+ 6.8s), and the per-call budget defaults to 120s for this provider for the
71
+ same reason. The prompt goes in on stdin, so a large judged text is fine.
45
72
  - **`gemini`** — the Gemini API with `GEMINI_API_KEY`. Same wire format as
46
73
  `vertex`, temperature pinned to 0.
47
74
  - **`openai`** — OpenAI chat completions with `OPENAI_API_KEY`
@@ -53,11 +80,35 @@ unotest through `npx`, with nothing installed locally, can only use
53
80
  - **`fake`** — deterministic, no model. For CI and smoke-testing the
54
81
  wiring. The rubric is a micro-grammar, one constraint per line:
55
82
  `must contain: <substring>` / `must not contain: <substring>`
56
- (case-insensitive). Anything else fails loudly.
83
+ (case-insensitive). Anything else fails loudly. Note that a
84
+ `UNOTEST_JUDGE_PREAMBLE` set on the client side arrives glued to the
85
+ rubric, so constraint lines in a preamble apply to every judgement.
57
86
 
58
87
  All model providers are raw HTTP or a local process — this package ships
59
88
  zero provider SDKs.
60
89
 
90
+ ## Latency
91
+
92
+ A verdict costs wall-clock time in the scenario that waits for it, so the
93
+ choice of provider is a choice of pace:
94
+
95
+ | Provider | Per verdict | Why |
96
+ | --- | --- | --- |
97
+ | `vertex` / `gemini` / `openai` / `anthropic` | one HTTP round-trip (~1s for a flash-class model) | a single JSON POST |
98
+ | `claude` | ~6s, measured on a local CLI | every call spawns the binary and loads a session |
99
+ | `fake` | none | no model behind it |
100
+
101
+ `UNOTEST_JUDGE_VOTE=N` costs almost nothing on top: the N ballots run
102
+ concurrently, so three votes take about as long as one call (measured 6.8s
103
+ for `claude`). `UNOTEST_JUDGE_RETRIES`, by contrast, is sequential — it
104
+ only spends a call when a verdict came back `fail`.
105
+
106
+ Budgets follow the same shape: `UNOTEST_JUDGE_CALL_TIMEOUT_MS` defaults to
107
+ 30s, and to 120s for `claude`. On the client side `@unotest/web` has its
108
+ own whole-request budget (`UNOTEST_JUDGE_TIMEOUT_MS`, 60s) — with `claude`
109
+ plus a slow model the client can give up before the service does, and it
110
+ now says so when that happens.
111
+
61
112
  ## Env
62
113
 
63
114
  | Variable | Meaning | Default |
@@ -68,7 +119,7 @@ zero provider SDKs.
68
119
  | `UNOTEST_JUDGE_VOTE` | odd N; N independent calls, the majority decides. Replaces retries rather than stacking with them | `1` (off) |
69
120
  | `UNOTEST_JUDGE_CALL_TIMEOUT_MS` | budget for ONE provider call, ms | `30000` (`120000` for claude) |
70
121
  | `GOOGLE_CLOUD_PROJECT` | vertex: ADC project | required for vertex |
71
- | `GOOGLE_CLOUD_LOCATION` | vertex: ADC location | required for vertex |
122
+ | `GOOGLE_CLOUD_LOCATION` | vertex: `global` or a region (`us-central1`) | required for vertex |
72
123
  | `UNOTEST_JUDGE_ACCESS_TOKEN` | vertex: static bearer override (skips ADC) | — |
73
124
  | `GEMINI_API_KEY` | gemini: API key | required for gemini |
74
125
  | `OPENAI_API_KEY` | openai: API key | required for openai |
@@ -76,6 +127,32 @@ zero provider SDKs.
76
127
  | `UNOTEST_JUDGE_CLAUDE_BIN` | claude: binary override | `claude` |
77
128
  | `UNOTEST_JUDGE_HOST` / `UNOTEST_JUDGE_PORT` | bind address | `127.0.0.1` / `8790` |
78
129
  | `UNOTEST_JUDGE_TOKEN` | bearer token required on every `/judge` call. `@unotest/web` reads the SAME name to send it | off |
130
+ | `UNOTEST_JUDGE_LOG_LEVEL` | `silent` \| `error` \| `warn` \| `info` \| `debug`; falls back to `UNOTEST_LOG_LEVEL` | `info` |
131
+ | `UNOTEST_JUDGE_SKIP_PREFLIGHT` | `1` starts even when the credential check fails (`/health` still reports it) | off |
132
+
133
+ ## Logging
134
+
135
+ At the default `info` level the service prints one line per judged
136
+ request — verdict, attempts spent, model, latency:
137
+
138
+ ```
139
+ POST /judge pass attempts 3 gemini-2.5-flash 1284ms
140
+ POST /judge ERROR config Application Default Credentials need re-authentication (invalid_rapt) — run `gcloud auth application-default login`
141
+ ```
142
+
143
+ Errors go to stderr. A run with 21 assertions and `UNOTEST_JUDGE_VOTE=3`
144
+ is 63 billed model calls; this is the only place their cost and latency
145
+ are visible.
146
+
147
+ `UNOTEST_JUDGE_LOG_LEVEL=debug` additionally prints, per call, the
148
+ effective prompt (rubric **including** any `UNOTEST_JUDGE_PREAMBLE` the
149
+ client prepended), each ballot with its reasoning, and the provider's raw
150
+ reply.
151
+
152
+ > **The judged text is application content** — support answers, user
153
+ > data, whatever the scenario put under the assertion. Debug logging
154
+ > writes it to stdout verbatim. It is opt-in for that reason; do not
155
+ > enable it on a shared runner without deciding that is acceptable.
79
156
 
80
157
  ## HTTP API
81
158
 
@@ -83,11 +160,31 @@ zero provider SDKs.
83
160
  `JudgeVerdict` (above). Errors: `{"error": "...", "code": "unauthorized" | "bad-request" | "provider-error" | "config" | "internal"}`.
84
161
  `config` means the operator must fix this service's setup (bad env,
85
162
  expired ADC) and the message says what to run.
86
- - `GET /health` — `{"ok": true}`.
163
+ - `GET /health` — `{"ok": true}` when the provider's credentials work,
164
+ otherwise `503` + `{"ok": false, "code": "config", "error": "..."}` with
165
+ the same message and code a verdict call would have failed with. The
166
+ check is cached for a few seconds, so polling is cheap, and it never
167
+ calls the model.
87
168
 
88
169
  Wire types are shared through `@unotest/protocol` (`JudgeRequest`,
89
- `JudgeVerdict`, `JUDGE_ROUTES`), so the service and `@unotest/web`'s
90
- client cannot drift.
170
+ `JudgeVerdict`, `JudgeHealthResponse`, `JUDGE_ROUTES`), so the service and
171
+ `@unotest/web`'s client cannot drift.
172
+
173
+ Faults that nobody chose are re-sent: rate limits (`429`), backend `5xx`,
174
+ a dropped connection, a spawn the OS could not grant. Up to three attempts,
175
+ 500ms then 1500ms apart, `Retry-After` honoured over that curve when the
176
+ backend names one. The policy sits on the provider contract, so it covers
177
+ **every** provider — including `claude`, which speaks to a process rather
178
+ than to HTTP.
179
+
180
+ What is never re-sent: a rejected credential (`401`/`403`) comes back as
181
+ `config` with the variable to fix, an unknown model is not a hiccup, and a
182
+ `claude` usage limit is reported rather than retried — the retry window is
183
+ seconds while a subscription limit lasts far longer, so re-spawning a ~6s
184
+ CLI would only make the same failure slower.
185
+
186
+ With `UNOTEST_JUDGE_VOTE=N`, a ballot that still errors is dropped rather
187
+ than failing the request, as long as a majority of the N calls voted.
91
188
 
92
189
  ## License
93
190