johns-harness 2026.9.28 → 2026.9.30
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +10 -0
- package/README.md +2 -2
- package/dist/launchd-lifecycle.js +18 -0
- package/docs/PATCHES.md +39 -0
- package/docs/PROVIDER-AUDIT-2026-09-20.md +120 -0
- package/docs/PROVIDER-CATALOG.md +7 -4
- package/node_modules/@mariozechner/pi-ai/dist/providers/google.js +2 -1
- package/node_modules/@mariozechner/pi-ai/dist/providers/openai-completions.js +4 -1
- package/node_modules/@mariozechner/pi-ai/dist/providers/openai-responses.js +4 -3
- package/package.json +1 -1
- package/vendor/patched-deps/@mariozechner/pi-ai/dist/providers/google.js +2 -1
- package/vendor/patched-deps/@mariozechner/pi-ai/dist/providers/openai-completions.js +4 -1
- package/vendor/patched-deps/@mariozechner/pi-ai/dist/providers/openai-responses.js +4 -3
package/CHANGELOG.md
CHANGED
|
@@ -1,5 +1,15 @@
|
|
|
1
1
|
# Changelog
|
|
2
2
|
|
|
3
|
+
## 2026.9.30
|
|
4
|
+
|
|
5
|
+
- Family 77: weekly official six-provider audit. DeepSeek V4 Flash/Pro now transmit native low for minimal/low instead of silently using high. Meta Standard Muse 1.2/1.3 preserve provider-default effort when the caller omits it instead of forcing minimal. Gemini 3.8 Flash and 3.5 Flash-Lite omit deprecated/ignored sampling temperature. Preserve all aliases, saved model IDs, owner context caps and default selections.
|
|
6
|
+
- Astra native Responses max, Fable 5.1 max and Grok 4.6 xhigh were already correct and were independently wire-verified. Meta inference remains account-dependent; no configured account was available for this audit. New Gemini Live/Interactions models are not mislabeled as working generateContent chat models.
|
|
7
|
+
- Exact preimage/postimage replay, three actual-adapter regression surfaces and installed-artifact verification. See `docs/PROVIDER-AUDIT-2026-09-20.md`. Roll back with `npm i -g johns-harness@2026.9.29`; focused adapter installs use their recorded pre-patch backups.
|
|
8
|
+
|
|
9
|
+
## 2026.9.29
|
|
10
|
+
|
|
11
|
+
- Family 76.1: the registration-safe reload worker waits for `launchctl bootout` teardown to finish before bootstrapping the staged plist. A gateway that shuts down gracefully on SIGTERM stays registered for a moment after `bootout`; 2026.9.28 mistook that lingering registration for the replacement, skipped `bootstrap`, and reported `Service removed while waiting for health` with the job left unloaded (observed on the first live fleet reload). The wait is bounded at 30 seconds and a timeout is a retryable worker error, not an advanced checkpoint. Ordinary `kickstart -k` restarts were unaffected. The real macOS self-restart fixture now exits gracefully so the reload cases reproduce the race; two unit cases cover the wait and the timeout.
|
|
12
|
+
|
|
3
13
|
## 2026.9.28
|
|
4
14
|
|
|
5
15
|
- Family 76: `johnness gateway restart` from inside a gateway's own launchd process group no longer unregisters its supervisor. The old `bootout -> wait -> bootstrap` sequence killed the caller before it could bootstrap, which caused a multi-hour outage when a cron issued the restart. All three service implementations, both internal restart bundles and both updater bundles now share `dist/launchd-lifecycle.js`: a loaded service restarts with `kickstart -k` only (no bootout, unload or bootstrap fallback), a missing service may bootstrap its explicitly selected plist, and real plist replacement (`gateway install --force`) stages validated bytes and hands off to an independently registered, durably armed completer with an fsynced phase journal, retry under launchd, and prior-plist restore after five failed bootstraps. Explicit stop/uninstall cancels any pending completer. Updaters snapshot the helper before npm replaces the installed tree. Success requires a changed launchd-owned PID and, when a port is declared, `/healthz` plus listener ownership by that PID; receipts live in `<stateDir>/lifecycle/<uuid>/`. Ten browser timeout messages now suggest a non-browser alternative instead of a whole-gateway restart. Covered by lifecycle unit, fault-injection, replay and packed-artifact tests plus five real disposable LaunchAgent self-restart/reload cases (`JOHNNESS_TEST_LAUNCHD=1`); verifier check 76.1.
|
package/README.md
CHANGED
|
@@ -85,7 +85,7 @@ node johnness.mjs --help
|
|
|
85
85
|
### Verify
|
|
86
86
|
|
|
87
87
|
```bash
|
|
88
|
-
npm test #
|
|
88
|
+
npm test # 992 tests against the installed dependency tree
|
|
89
89
|
bash scripts/verify-patches.sh # 468 checks across every patch family
|
|
90
90
|
```
|
|
91
91
|
|
|
@@ -199,7 +199,7 @@ The log is a [`RULES.md`](RULES.md) workspace convention enforced through the ac
|
|
|
199
199
|
|
|
200
200
|
## Patch families
|
|
201
201
|
|
|
202
|
-
The runtime carries 76 patch families plus the 65.1 safe-update amendment. Each is a production fix applied at the source with a regression test, a durable patch marker, and an explicit recovery path. A verifier runs 468 checks against the package tree and installed dependencies, and
|
|
202
|
+
The runtime carries 76 patch families plus the 65.1 safe-update amendment. Each is a production fix applied at the source with a regression test, a durable patch marker, and an explicit recovery path. A verifier runs 468 checks against the package tree and installed dependencies, and 992 tests run against the installed dependency tree.
|
|
203
203
|
|
|
204
204
|
Every family follows the same six steps: reproduce the failure, trace the exact runtime path, make the smallest source-level change that restores the invariant, add a regression test and a patch marker, run the verifier, and retain rollback artifacts. The full index: [`docs/PATCHES.md`](docs/PATCHES.md).
|
|
205
205
|
|
|
@@ -169,6 +169,17 @@ async function healthy(p, current, deps) {
|
|
|
169
169
|
if (!pids.length || !pids.every((pid) => descendant(pid, current.pid))) return false;
|
|
170
170
|
try { return (await fetch(`http://127.0.0.1:${p.port}/healthz`, { signal: AbortSignal.timeout(2000) })).ok; } catch { return false; }
|
|
171
171
|
}
|
|
172
|
+
/** After bootout, launchd finishes SIGTERM teardown asynchronously; only the previous pid can
|
|
173
|
+
* still be registered here. A different registered pid means the plist was already re-bootstrapped. */
|
|
174
|
+
async function unloaded(run, p, deps) {
|
|
175
|
+
const deadline = Date.now() + (deps.unloadTimeoutMs ?? 30000);
|
|
176
|
+
for (;;) {
|
|
177
|
+
const current = runtime(run, p.target);
|
|
178
|
+
if (!current.loaded || (current.pid && p.previousPid && current.pid !== p.previousPid)) return current;
|
|
179
|
+
if (Date.now() >= deadline) throw new Error("Previous process still registered after bootout; independent worker will retry");
|
|
180
|
+
await (deps.pause || pause)(200);
|
|
181
|
+
}
|
|
182
|
+
}
|
|
172
183
|
/** Restartable state machine. Checkpoints precede effects; real supervisor state
|
|
173
184
|
* reconciles a missing effect acknowledgement after SIGKILL. Exported for fault tests. */
|
|
174
185
|
export async function executeLaunchdOperation(p, receiptFile, deps = {}) {
|
|
@@ -195,6 +206,11 @@ export async function executeLaunchdOperation(p, receiptFile, deps = {}) {
|
|
|
195
206
|
const r = effect(["bootout", p.target]);
|
|
196
207
|
if (r.code && !missing(r)) throw new Error(`launchctl bootout failed (${r.code})`);
|
|
197
208
|
}
|
|
209
|
+
// bootout is asynchronous for a job that handles SIGTERM: launchd keeps it registered (old pid)
|
|
210
|
+
// until graceful shutdown completes. Wait for the teardown before recording the intent to
|
|
211
|
+
// bootstrap, or the still-registered job would be mistaken for the replacement and later
|
|
212
|
+
// vanish, producing "Service removed while waiting for health" with no gateway loaded.
|
|
213
|
+
await unloaded(run, p, deps);
|
|
198
214
|
save("bootstrap-intent"); receipt = read(receiptFile);
|
|
199
215
|
current = runtime(run, p.target);
|
|
200
216
|
}
|
|
@@ -220,6 +236,8 @@ export async function executeLaunchdOperation(p, receiptFile, deps = {}) {
|
|
|
220
236
|
save("health-pending", { rolledBack: true, healthStartedAt: Date.now() }); receipt = read(receiptFile);
|
|
221
237
|
}
|
|
222
238
|
if (receipt.status === "bootstrap-intent") {
|
|
239
|
+
// Resume after a crash between bootout and this checkpoint: the old job may still be tearing down.
|
|
240
|
+
if (p.mode === "reload") await unloaded(run, p, deps);
|
|
223
241
|
current = runtime(run, p.target);
|
|
224
242
|
if (!current.loaded) {
|
|
225
243
|
const r = effect(["bootstrap", p.domain, p.plistPath]);
|
package/docs/PATCHES.md
CHANGED
|
@@ -2741,6 +2741,18 @@ self-restart can disappear; the independent receipt is the source of truth.
|
|
|
2741
2741
|
Ten browser error copies now report a browser blocker and suggest a non-browser
|
|
2742
2742
|
alternative, never a whole-gateway restart. Exec process-group behavior is unchanged.
|
|
2743
2743
|
|
|
2744
|
+
**76.1 (2026.9.29), graceful unload before bootstrap.** The first live reload of a real
|
|
2745
|
+
gateway (Trisha, `gateway install --force` semantics) failed with `Service removed while
|
|
2746
|
+
waiting for health` and left the job unloaded. `launchctl bootout` returns before a job
|
|
2747
|
+
that handles SIGTERM has exited; the worker's next `print` still saw the old registered
|
|
2748
|
+
PID, treated the job as loaded, skipped `bootstrap`, and then watched launchd finish the
|
|
2749
|
+
teardown. The worker now waits after `bootout` (and again on resume at `bootstrap-intent`)
|
|
2750
|
+
until the target is unregistered or a different PID is registered, with a 30 s bound that
|
|
2751
|
+
surfaces as a retryable worker error instead of an advanced checkpoint. The real macOS
|
|
2752
|
+
fixture now performs a 700 ms graceful SIGTERM shutdown so the reload cases exercise the
|
|
2753
|
+
race; two unit cases cover the wait and the timeout. `kickstart -k` restarts were never
|
|
2754
|
+
affected.
|
|
2755
|
+
|
|
2744
2756
|
**Verification:** `tests/launchd-safe-restart.test.mjs` covers lifecycle phases,
|
|
2745
2757
|
fault injection, error classification, scoping, all copies and exact patch replay.
|
|
2746
2758
|
`JOHNNESS_TEST_LAUNCHD=1 node --test --test-concurrency=1 tests/launchd-self-restart.macos.test.mjs`
|
|
@@ -2761,3 +2773,30 @@ restart` uses the safe worker. For a real plist change, use the new version's
|
|
|
2761
2773
|
`gateway install --force`, which stages the change before handoff. Rolling back to
|
|
2762
2774
|
2026.9.27 or earlier restores the old restart defect; if necessary install it externally and
|
|
2763
2775
|
use kickstart, never its gateway restart CLI from a gateway child.
|
|
2776
|
+
|
|
2777
|
+
|
|
2778
|
+
## 77. Weekly provider reasoning corrections (2026-09-20)
|
|
2779
|
+
|
|
2780
|
+
The six-provider audit changes only the DeepSeek, Meta Responses and Gemini adapters.
|
|
2781
|
+
Both DeepSeek V4 models now send native `low` for harness `minimal`/`low`, rather
|
|
2782
|
+
than silently spending at `high`. The existing `xhigh` to `max` compatibility
|
|
2783
|
+
mapping, explicit off, server-default omission, limits and aliases are preserved.
|
|
2784
|
+
Meta Standard Muse 1.2/1.3 now omit `reasoning.effort` when the caller omits it,
|
|
2785
|
+
instead of forcing `minimal`; explicit efforts, validation, encrypted replay and
|
|
2786
|
+
`store: false` are unchanged. Meta account access was not available for live inference.
|
|
2787
|
+
|
|
2788
|
+
`patches/provider-audit-20260920/changes.json` owns the exact three adapter pre/post
|
|
2789
|
+
images plus Meta's historical replay snapshot. Gemini 3.8 Flash and 3.5 Flash-Lite also omit deprecated/ignored temperature.
|
|
2790
|
+
The patcher preflights every target
|
|
2791
|
+
before writing, refuses drift, supports alternate roots and a dependency-only
|
|
2792
|
+
focused refresh. Family 67 reverses only these exact edits during historical hash
|
|
2793
|
+
validation. Both the modern and old Meta replay tests remain byte-identical.
|
|
2794
|
+
|
|
2795
|
+
No model IDs, aliases, config, credentials, default model or owner context caps
|
|
2796
|
+
change. Astra Responses already supports native `max`; Grok native `xhigh` and
|
|
2797
|
+
modern Claude native `max` were already correct. The global `/think max` synonym
|
|
2798
|
+
still means harness `xhigh`; it is not a new distinct UI level in this patch.
|
|
2799
|
+
|
|
2800
|
+
Evidence and official sources: [weekly audit](PROVIDER-AUDIT-2026-09-20.md).
|
|
2801
|
+
Tests: `tests/provider-audit-20260920.test.mjs`, provider catalog, Muse transport
|
|
2802
|
+
and historical replay; verifier 77.1 checks installed adapter hashes exactly.
|
|
@@ -0,0 +1,120 @@
|
|
|
1
|
+
# Six-provider official API audit: 2026-09-20
|
|
2
|
+
|
|
3
|
+
This weekly audit builds on the successful September 13 catalog release, not on
|
|
4
|
+
upstream OpenClaw. Registry IDs, aliases, saved-session IDs, owner context caps,
|
|
5
|
+
credentials and the selected default are unchanged. Family 77 is a focused
|
|
6
|
+
provider-parameter patch, released as `johns-harness@2026.9.30`.
|
|
7
|
+
|
|
8
|
+
## Changes
|
|
9
|
+
|
|
10
|
+
1. **DeepSeek:** `minimal`/`low` now serialize as native `low` on `deepseek-flash`
|
|
11
|
+
and `deepseek-v4-pro`, rather than silently increasing effort to `high`.
|
|
12
|
+
Omission remains omitted; explicit low-level off disables thinking. Preserve
|
|
13
|
+
the existing harness `xhigh`/`max` -> native `max` compatibility mapping for
|
|
14
|
+
saved sessions. DeepSeek's own raw `xhigh` synonym maps to `high`; that is not
|
|
15
|
+
a reason to silently lower the harness's established highest-effort behavior.
|
|
16
|
+
2. **Meta Muse:** omitted effort stays absent instead of forcing `minimal`.
|
|
17
|
+
Explicit effort validation, `store:false`, encrypted reasoning replay and
|
|
18
|
+
tool commentary phases are unchanged. No Contributor/training-consent models.
|
|
19
|
+
3. **Gemini:** omit deprecated/ignored temperature for `gemini-3.8-flash` and
|
|
20
|
+
`gemini-3.5-flash-lite`. Older model behavior is unchanged. The live API still
|
|
21
|
+
accepts this field, but the official migration guide says to remove it and
|
|
22
|
+
warns that future models reject it. No topP/topK are generated by this adapter.
|
|
23
|
+
|
|
24
|
+
No new supported general-purpose model ID was identified since the prior audit.
|
|
25
|
+
No model was removed and no alias was upgraded to a preview.
|
|
26
|
+
|
|
27
|
+
## Independent provider results
|
|
28
|
+
|
|
29
|
+
| Provider | Official/account result | Decision |
|
|
30
|
+
|---|---|---|
|
|
31
|
+
| DeepSeek | `/models` HTTP 200, two IDs: `deepseek-flash`, `deepseek-v4-pro`. Latest changelog release September 10, V4.1 Flash. Native low/high/max documented. Flash omission/low/high/xhigh/max and Pro low each HTTP 200. | Correct low serialization; retain all limits and prices. |
|
|
32
|
+
| Meta | Official Standard API confirms Muse 1.3 and retained 1.2 on `https://api.meta.ai/v1/responses`; 1.3 adds native max. No configured Meta credential. | Correct omission from official schema/reasoning docs and actual offline adapter tests; do not claim authenticated inference. |
|
|
33
|
+
| Google | Native model listing HTTP 200 (62 records); 3.8 Flash and 3.5 Flash-Lite confirmed, each 1,048,576 input / 65,536 output. 3.8 minimal smoke HTTP 200. | Omit deprecated temperature; preserve thinking-level mapping. New September 15 Live audio models require a Live adapter, not generateContent chat registration. September 17 Antigravity preview is an agent/Interactions API, not a drop-in model. |
|
|
34
|
+
| OpenAI | Native `/models` HTTP 200. Astra Responses omission/xhigh/max HTTP 200; none HTTP 400 explicitly reports supported low/medium/high/xhigh/max. | Existing native Responses max support is correct. Do not import another application's Chat Completions limitations. No catalog change. |
|
|
35
|
+
| Anthropic | Native `/v1/models` HTTP 200 (10 records) confirms Fable 5.1, Opus 5 and Sonnet 5. Fable adaptive thinking + output_config.effort=max HTTP 200. | Native max was already preserved by the adapter; no change. Account-gated models are not inferred from announcements. |
|
|
36
|
+
| xAI | Native `/v1/models` HTTP 200 (34 records) includes Grok 4.6. Omitted effort and xhigh Chat Completions requests HTTP 200. | xhigh was already supported. Keep exactly 200K owner context cap, 64K output, existing aliases. |
|
|
37
|
+
|
|
38
|
+
The model listing counts are account/time-specific, not availability promises.
|
|
39
|
+
No API list had an unconsumed pagination indicator. No 429 occurred and no
|
|
40
|
+
credential was printed or persisted in evidence. Minimal inference probes used
|
|
41
|
+
synthetic "Reply with OK only" prompts and at most 64 output tokens. Successful
|
|
42
|
+
HTTP parameter acceptance is not a long-output, all-modality or full tool-loop
|
|
43
|
+
inference certification. Offline existing tests separately cover streaming,
|
|
44
|
+
reasoning/signature/encrypted replay, tool calls and errors.
|
|
45
|
+
|
|
46
|
+
## Metadata and compatibility
|
|
47
|
+
|
|
48
|
+
- DeepSeek Flash is V4.1 Flash (text/image), Pro is V4 Pro 0813 (text).
|
|
49
|
+
Both have 1,048,576 context / 393,216 maximum output. Native default is
|
|
50
|
+
thinking enabled/high. Nonthinking default max output is 8K; thinking high is
|
|
51
|
+
64K and max is 128K. Temperature has no effect in thinking mode and remains
|
|
52
|
+
omitted. The adapter sends `max_tokens`, not `max_completion_tokens`, and
|
|
53
|
+
does not invent developer-role, store, or strict-tool support.
|
|
54
|
+
- Meta Standard 1.3/1.2 retain 1,048,576 context / 131,072 output and text/image
|
|
55
|
+
input in the harness. Models always reason. Omission lets the model choose
|
|
56
|
+
effort; none/off are invalid. Max is exclusive to Standard 1.3. Standard
|
|
57
|
+
pricing remains $1.25 input / $4.25 output / $0.15 cached input per million.
|
|
58
|
+
- Google current Flash/Lite use native thinking levels rather than a synthetic
|
|
59
|
+
budget. 3.8 MINIMAL remains mapped to LOW by the previously verified adapter;
|
|
60
|
+
3.5 Lite supports MINIMAL. Omission stays omitted. Native API supports more
|
|
61
|
+
media than the pinned text/image agent interface; this patch does not claim
|
|
62
|
+
Live audio, Interactions, image/video output or unrelated tool adapters.
|
|
63
|
+
- OpenAI native context is 1,050,000 with 128K max output; conservative existing
|
|
64
|
+
harness budgets remain unchanged. Astra cannot disable reasoning. The native
|
|
65
|
+
Responses adapter already passes max exactly. The global `/think max` spelling
|
|
66
|
+
still normalizes to harness xhigh; adding a separate UI enum is not this patch.
|
|
67
|
+
- Modern Claude native effort levels include distinct xhigh and max. Adaptive
|
|
68
|
+
thinking omits manual budgets and sampling. Fable/Opus retain 300K owner caps;
|
|
69
|
+
Sonnet's conservative existing budget remains untouched.
|
|
70
|
+
- Grok 4.6 documents 500K native context / 64K output; the harness deliberately
|
|
71
|
+
retains 200K. Its reasoning_effort xhigh is transmitted exactly. Prices remain
|
|
72
|
+
$2 input / $6 output / $0.50 cached input per million; native long-prompt
|
|
73
|
+
multipliers are not represented by the basic catalog cost estimator.
|
|
74
|
+
|
|
75
|
+
Prices/limits not changed in this release remain at their last verified values;
|
|
76
|
+
there is no blanket claim of re-pricing every legacy registry entry. DeepSeek
|
|
77
|
+
current pricing and Meta Standard pricing were independently rechecked. No price,
|
|
78
|
+
context or output-limit diff was required on the changed entries.
|
|
79
|
+
|
|
80
|
+
## First-party sources checked
|
|
81
|
+
|
|
82
|
+
- DeepSeek: [catalog/pricing](https://api-docs.deepseek.com/quick_start/pricing/),
|
|
83
|
+
[thinking](https://api-docs.deepseek.com/guides/thinking_mode/),
|
|
84
|
+
[complete Chat Completions reference](https://api-docs.deepseek.com/api/create-chat-completion/),
|
|
85
|
+
[changelog](https://api-docs.deepseek.com/updates).
|
|
86
|
+
- Meta: [models](https://ai.developer.meta.com/docs/models.md),
|
|
87
|
+
[reasoning](https://ai.developer.meta.com/docs/reasoning.md),
|
|
88
|
+
[Responses protocol](https://ai.developer.meta.com/docs/protocols/responses.md),
|
|
89
|
+
[create response](https://ai.developer.meta.com/docs/api-reference/responses/create-response.md),
|
|
90
|
+
[pricing](https://ai.developer.meta.com/docs/pricing-rate-limits.md).
|
|
91
|
+
HTML indexes intermittently returned HTTP 500; official `.md` pages succeeded.
|
|
92
|
+
- Google: [model catalog](https://ai.google.dev/gemini-api/docs/models),
|
|
93
|
+
[thinking](https://ai.google.dev/gemini-api/docs/thinking.md.txt),
|
|
94
|
+
[latest-model migration](https://ai.google.dev/gemini-api/docs/generate-content/latest-model),
|
|
95
|
+
[changelog](https://ai.google.dev/gemini-api/docs/changelog).
|
|
96
|
+
Some index requests redirected repeatedly; native model listing and direct
|
|
97
|
+
content fetch supplied the data. Locale variants do not alter exact API IDs.
|
|
98
|
+
- OpenAI: [catalog](https://developers.openai.com/api/docs/models/),
|
|
99
|
+
[Astra model reference](https://developers.openai.com/api/docs/models/gpt-6-astra.md).
|
|
100
|
+
- Anthropic: [overview](https://platform.claude.com/docs/en/models/overview),
|
|
101
|
+
[Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview),
|
|
102
|
+
[effort](https://platform.claude.com/docs/en/build-with-claude/effort).
|
|
103
|
+
- xAI: [catalog](https://docs.x.ai/developers/models),
|
|
104
|
+
[Grok 4.6](https://docs.x.ai/developers/models/grok-4.6).
|
|
105
|
+
|
|
106
|
+
## Verification and recovery
|
|
107
|
+
|
|
108
|
+
Family 77 has exact preimage/postimage hashes, all-file preflight before writes,
|
|
109
|
+
read-only checks, alternate-root and dependency-only modes, and idempotent replay.
|
|
110
|
+
Family 67 validates its historical surface after reversing only the exact new
|
|
111
|
+
edits. Meta's Family 53 snapshot is synchronized. Payload tests exercise actual
|
|
112
|
+
installed adapters, not string-only fixtures. Existing alias contract tests and
|
|
113
|
+
all prior catalog/transport replay tests remain part of the release gates.
|
|
114
|
+
|
|
115
|
+
The release is a three-adapter patch on 2026.9.29. Roll back a standard package
|
|
116
|
+
installation to `johns-harness@2026.9.29`; focused installations must restore their
|
|
117
|
+
recorded pre-patch adapter files, not overwrite state or migrate directories.
|
|
118
|
+
A running gateway caches provider modules. Disk validation is not activation:
|
|
119
|
+
coordinate a safe restart after active work finishes, never interrupt jobs merely
|
|
120
|
+
to report a refreshed version.
|
package/docs/PROVIDER-CATALOG.md
CHANGED
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
# Current provider catalog — checked 2026-09-
|
|
1
|
+
# Current provider catalog — checked 2026-09-20
|
|
2
2
|
|
|
3
3
|
Family 67 audits general-purpose agent models against first-party documentation and
|
|
4
4
|
provider-authenticated model lists. It is a targeted registry/adapter refresh on the
|
|
@@ -31,7 +31,7 @@ Registry presence does not promise account entitlement.
|
|
|
31
31
|
|
|
32
32
|
- **DeepSeek V4:** omitted effort leaves `thinking` and `reasoning_effort` absent,
|
|
33
33
|
preserving the server's default enabled/high. Explicit low-level `off`/`none`
|
|
34
|
-
sets `thinking: {type: "disabled"}`. Minimal/low
|
|
34
|
+
sets `thinking: {type: "disabled"}`. Minimal/low map to native `low`; medium/high map to native `high`;
|
|
35
35
|
xhigh/max map to native `max`. Unknown efforts fail before HTTP. Disabled thinking
|
|
36
36
|
must not carry `reasoning_effort`. In enabled/default mode temperature is omitted
|
|
37
37
|
because the API ignores sampling controls. Use `max_tokens`, not
|
|
@@ -43,10 +43,12 @@ Registry presence does not promise account entitlement.
|
|
|
43
43
|
reasoning replay, and tool-loop commentary `phase`. Always-reasoning;
|
|
44
44
|
minimal/low/medium/high/xhigh, with native max only on Standard 1.3. The pinned
|
|
45
45
|
high-level `/think max` spelling still normalizes to xhigh, not a new global enum.
|
|
46
|
+
Omitted effort stays absent and provider-controlled rather than forcing minimal.
|
|
46
47
|
- **Gemini:** Flash-Lite now enters the native thinking-level branch instead of
|
|
47
48
|
the legacy thinking-budget branch. It supports MINIMAL/LOW/MEDIUM/HIGH; omission
|
|
48
49
|
preserves its native minimal default. 3.8 Flash still maps minimal to LOW because
|
|
49
50
|
the live API rejects MINIMAL (Family 50). No synthetic thinkingBudget is sent.
|
|
51
|
+
Deprecated/ignored temperature is omitted on 3.8 Flash and 3.5 Flash-Lite.
|
|
50
52
|
- **Claude:** current Fable 5/5.1, Opus 5 and Sonnet 5 preserve native xhigh rather
|
|
51
53
|
than silently downgrading it to high or promoting it to max. Low-level max remains
|
|
52
54
|
distinct. Minimal maps to low. Adaptive thinking has no manual token budget.
|
|
@@ -60,7 +62,7 @@ Registry presence does not promise account entitlement.
|
|
|
60
62
|
public thinking enum still tops out at xhigh. Transport token ceilings are not
|
|
61
63
|
raised merely because the provider advertises a larger native window.
|
|
62
64
|
- **Grok:** existing native effort mapping and 200K owner cap retained; live low
|
|
63
|
-
effort
|
|
65
|
+
and xhigh effort requests accepted by `grok-4.6`.
|
|
64
66
|
|
|
65
67
|
The pinned high-level API represents `/think off` as an absent reasoning option on
|
|
66
68
|
some routes. Consequently omission and explicit off cannot be distinguished there;
|
|
@@ -96,7 +98,8 @@ Subscription cost entries remain zero (subscription-billed), not API-price estim
|
|
|
96
98
|
|
|
97
99
|
## Evidence and sources
|
|
98
100
|
|
|
99
|
-
|
|
101
|
+
Baseline sources checked 2026-09-13; six-provider re-audit 2026-09-20 is in
|
|
102
|
+
[the dated audit](PROVIDER-AUDIT-2026-09-20.md). First-party pages unless noted. Credential-bearing API
|
|
100
103
|
requests were sent only to the relevant official host, with redirects disabled.
|
|
101
104
|
|
|
102
105
|
- DeepSeek: <https://api-docs.deepseek.com/quick_start/pricing/>,
|
|
@@ -252,7 +252,8 @@ function createClient(model, apiKey, optionsHeaders) {
|
|
|
252
252
|
function buildParams(model, context, options = {}) {
|
|
253
253
|
const contents = convertMessages(model, context);
|
|
254
254
|
const generationConfig = {};
|
|
255
|
-
|
|
255
|
+
// JOHNNESS_PATCH_PROVIDER_AUDIT_77: sampling is deprecated/ignored on these models.
|
|
256
|
+
if (options.temperature !== undefined && !["gemini-3.8-flash", "gemini-3.5-flash-lite"].includes(model.id)) {
|
|
256
257
|
generationConfig.temperature = options.temperature;
|
|
257
258
|
}
|
|
258
259
|
if (options.maxTokens !== undefined) {
|
|
@@ -364,7 +364,10 @@ function buildParams(model, context, options) {
|
|
|
364
364
|
} else {
|
|
365
365
|
if (effort !== undefined) {
|
|
366
366
|
params.thinking = { type: "enabled" };
|
|
367
|
-
|
|
367
|
+
// JOHNNESS_PATCH_PROVIDER_AUDIT_77: low is native on both V4 routes (2026-09-20).
|
|
368
|
+
// Preserve the harness's existing xhigh -> max compatibility mapping.
|
|
369
|
+
params.reasoning_effort = ["minimal", "low"].includes(effort) ? "low"
|
|
370
|
+
: ["xhigh", "max"].includes(effort) ? "max" : "high";
|
|
368
371
|
}
|
|
369
372
|
// Sampling controls have no effect in thinking mode; never imply they do.
|
|
370
373
|
delete params.temperature;
|
|
@@ -176,11 +176,12 @@ function buildParams(model, context, options) {
|
|
|
176
176
|
}
|
|
177
177
|
// JOHNNESS_PATCH_MUSE_SPARK: always-reasoning, stateless encrypted replay.
|
|
178
178
|
if (model.provider === "meta") {
|
|
179
|
-
|
|
180
|
-
const
|
|
179
|
+
// JOHNNESS_PATCH_PROVIDER_AUDIT_77: omission keeps Meta's model-selected effort.
|
|
180
|
+
const effort = options?.reasoningEffort;
|
|
181
|
+
const allowed = [undefined, "minimal", "low", "medium", "high", "xhigh"];
|
|
181
182
|
if (model.id === "muse-spark-1.3") allowed.push("max");
|
|
182
183
|
if (!allowed.includes(effort)) throw new Error("Muse Spark always reasons. Use minimal, low, medium, high or xhigh (max only on Standard 1.3); none/off is not supported.");
|
|
183
|
-
params.reasoning = { effort, summary: options?.reasoningSummary || "auto" };
|
|
184
|
+
params.reasoning = { ...(effort === undefined ? {} : { effort }), summary: options?.reasoningSummary || "auto" };
|
|
184
185
|
params.include = ["reasoning.encrypted_content"];
|
|
185
186
|
params.store = false;
|
|
186
187
|
// Avoid inventing Meta pricing multipliers from OpenAI service tiers.
|
package/package.json
CHANGED
|
@@ -252,7 +252,8 @@ function createClient(model, apiKey, optionsHeaders) {
|
|
|
252
252
|
function buildParams(model, context, options = {}) {
|
|
253
253
|
const contents = convertMessages(model, context);
|
|
254
254
|
const generationConfig = {};
|
|
255
|
-
|
|
255
|
+
// JOHNNESS_PATCH_PROVIDER_AUDIT_77: sampling is deprecated/ignored on these models.
|
|
256
|
+
if (options.temperature !== undefined && !["gemini-3.8-flash", "gemini-3.5-flash-lite"].includes(model.id)) {
|
|
256
257
|
generationConfig.temperature = options.temperature;
|
|
257
258
|
}
|
|
258
259
|
if (options.maxTokens !== undefined) {
|
|
@@ -364,7 +364,10 @@ function buildParams(model, context, options) {
|
|
|
364
364
|
} else {
|
|
365
365
|
if (effort !== undefined) {
|
|
366
366
|
params.thinking = { type: "enabled" };
|
|
367
|
-
|
|
367
|
+
// JOHNNESS_PATCH_PROVIDER_AUDIT_77: low is native on both V4 routes (2026-09-20).
|
|
368
|
+
// Preserve the harness's existing xhigh -> max compatibility mapping.
|
|
369
|
+
params.reasoning_effort = ["minimal", "low"].includes(effort) ? "low"
|
|
370
|
+
: ["xhigh", "max"].includes(effort) ? "max" : "high";
|
|
368
371
|
}
|
|
369
372
|
// Sampling controls have no effect in thinking mode; never imply they do.
|
|
370
373
|
delete params.temperature;
|
|
@@ -176,11 +176,12 @@ function buildParams(model, context, options) {
|
|
|
176
176
|
}
|
|
177
177
|
// JOHNNESS_PATCH_MUSE_SPARK: always-reasoning, stateless encrypted replay.
|
|
178
178
|
if (model.provider === "meta") {
|
|
179
|
-
|
|
180
|
-
const
|
|
179
|
+
// JOHNNESS_PATCH_PROVIDER_AUDIT_77: omission keeps Meta's model-selected effort.
|
|
180
|
+
const effort = options?.reasoningEffort;
|
|
181
|
+
const allowed = [undefined, "minimal", "low", "medium", "high", "xhigh"];
|
|
181
182
|
if (model.id === "muse-spark-1.3") allowed.push("max");
|
|
182
183
|
if (!allowed.includes(effort)) throw new Error("Muse Spark always reasons. Use minimal, low, medium, high or xhigh (max only on Standard 1.3); none/off is not supported.");
|
|
183
|
-
params.reasoning = { effort, summary: options?.reasoningSummary || "auto" };
|
|
184
|
+
params.reasoning = { ...(effort === undefined ? {} : { effort }), summary: options?.reasoningSummary || "auto" };
|
|
184
185
|
params.include = ["reasoning.encrypted_content"];
|
|
185
186
|
params.store = false;
|
|
186
187
|
// Avoid inventing Meta pricing multipliers from OpenAI service tiers.
|