johns-harness 2026.9.28 → 2026.9.30

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -1,5 +1,15 @@
1
1
  # Changelog
2
2
 
3
+ ## 2026.9.30
4
+
5
+ - Family 77: weekly official six-provider audit. DeepSeek V4 Flash/Pro now transmit native low for minimal/low instead of silently using high. Meta Standard Muse 1.2/1.3 preserve provider-default effort when the caller omits it instead of forcing minimal. Gemini 3.8 Flash and 3.5 Flash-Lite omit deprecated/ignored sampling temperature. Preserve all aliases, saved model IDs, owner context caps and default selections.
6
+ - Astra native Responses max, Fable 5.1 max and Grok 4.6 xhigh were already correct and were independently wire-verified. Meta inference remains account-dependent; no configured account was available for this audit. New Gemini Live/Interactions models are not mislabeled as working generateContent chat models.
7
+ - Exact preimage/postimage replay, three actual-adapter regression surfaces and installed-artifact verification. See `docs/PROVIDER-AUDIT-2026-09-20.md`. Roll back with `npm i -g johns-harness@2026.9.29`; focused adapter installs use their recorded pre-patch backups.
8
+
9
+ ## 2026.9.29
10
+
11
+ - Family 76.1: the registration-safe reload worker waits for `launchctl bootout` teardown to finish before bootstrapping the staged plist. A gateway that shuts down gracefully on SIGTERM stays registered for a moment after `bootout`; 2026.9.28 mistook that lingering registration for the replacement, skipped `bootstrap`, and reported `Service removed while waiting for health` with the job left unloaded (observed on the first live fleet reload). The wait is bounded at 30 seconds and a timeout is a retryable worker error, not an advanced checkpoint. Ordinary `kickstart -k` restarts were unaffected. The real macOS self-restart fixture now exits gracefully so the reload cases reproduce the race; two unit cases cover the wait and the timeout.
12
+
3
13
  ## 2026.9.28
4
14
 
5
15
  - Family 76: `johnness gateway restart` from inside a gateway's own launchd process group no longer unregisters its supervisor. The old `bootout -> wait -> bootstrap` sequence killed the caller before it could bootstrap, which caused a multi-hour outage when a cron issued the restart. All three service implementations, both internal restart bundles and both updater bundles now share `dist/launchd-lifecycle.js`: a loaded service restarts with `kickstart -k` only (no bootout, unload or bootstrap fallback), a missing service may bootstrap its explicitly selected plist, and real plist replacement (`gateway install --force`) stages validated bytes and hands off to an independently registered, durably armed completer with an fsynced phase journal, retry under launchd, and prior-plist restore after five failed bootstraps. Explicit stop/uninstall cancels any pending completer. Updaters snapshot the helper before npm replaces the installed tree. Success requires a changed launchd-owned PID and, when a port is declared, `/healthz` plus listener ownership by that PID; receipts live in `<stateDir>/lifecycle/<uuid>/`. Ten browser timeout messages now suggest a non-browser alternative instead of a whole-gateway restart. Covered by lifecycle unit, fault-injection, replay and packed-artifact tests plus five real disposable LaunchAgent self-restart/reload cases (`JOHNNESS_TEST_LAUNCHD=1`); verifier check 76.1.
package/README.md CHANGED
@@ -85,7 +85,7 @@ node johnness.mjs --help
85
85
  ### Verify
86
86
 
87
87
  ```bash
88
- npm test # 990 tests against the installed dependency tree
88
+ npm test # 992 tests against the installed dependency tree
89
89
  bash scripts/verify-patches.sh # 468 checks across every patch family
90
90
  ```
91
91
 
@@ -199,7 +199,7 @@ The log is a [`RULES.md`](RULES.md) workspace convention enforced through the ac
199
199
 
200
200
  ## Patch families
201
201
 
202
- The runtime carries 76 patch families plus the 65.1 safe-update amendment. Each is a production fix applied at the source with a regression test, a durable patch marker, and an explicit recovery path. A verifier runs 468 checks against the package tree and installed dependencies, and 990 tests run against the installed dependency tree.
202
+ The runtime carries 76 patch families plus the 65.1 safe-update amendment. Each is a production fix applied at the source with a regression test, a durable patch marker, and an explicit recovery path. A verifier runs 468 checks against the package tree and installed dependencies, and 992 tests run against the installed dependency tree.
203
203
 
204
204
  Every family follows the same six steps: reproduce the failure, trace the exact runtime path, make the smallest source-level change that restores the invariant, add a regression test and a patch marker, run the verifier, and retain rollback artifacts. The full index: [`docs/PATCHES.md`](docs/PATCHES.md).
205
205
 
@@ -169,6 +169,17 @@ async function healthy(p, current, deps) {
169
169
  if (!pids.length || !pids.every((pid) => descendant(pid, current.pid))) return false;
170
170
  try { return (await fetch(`http://127.0.0.1:${p.port}/healthz`, { signal: AbortSignal.timeout(2000) })).ok; } catch { return false; }
171
171
  }
172
+ /** After bootout, launchd finishes SIGTERM teardown asynchronously; only the previous pid can
173
+ * still be registered here. A different registered pid means the plist was already re-bootstrapped. */
174
+ async function unloaded(run, p, deps) {
175
+ const deadline = Date.now() + (deps.unloadTimeoutMs ?? 30000);
176
+ for (;;) {
177
+ const current = runtime(run, p.target);
178
+ if (!current.loaded || (current.pid && p.previousPid && current.pid !== p.previousPid)) return current;
179
+ if (Date.now() >= deadline) throw new Error("Previous process still registered after bootout; independent worker will retry");
180
+ await (deps.pause || pause)(200);
181
+ }
182
+ }
172
183
  /** Restartable state machine. Checkpoints precede effects; real supervisor state
173
184
  * reconciles a missing effect acknowledgement after SIGKILL. Exported for fault tests. */
174
185
  export async function executeLaunchdOperation(p, receiptFile, deps = {}) {
@@ -195,6 +206,11 @@ export async function executeLaunchdOperation(p, receiptFile, deps = {}) {
195
206
  const r = effect(["bootout", p.target]);
196
207
  if (r.code && !missing(r)) throw new Error(`launchctl bootout failed (${r.code})`);
197
208
  }
209
+ // bootout is asynchronous for a job that handles SIGTERM: launchd keeps it registered (old pid)
210
+ // until graceful shutdown completes. Wait for the teardown before recording the intent to
211
+ // bootstrap, or the still-registered job would be mistaken for the replacement and later
212
+ // vanish, producing "Service removed while waiting for health" with no gateway loaded.
213
+ await unloaded(run, p, deps);
198
214
  save("bootstrap-intent"); receipt = read(receiptFile);
199
215
  current = runtime(run, p.target);
200
216
  }
@@ -220,6 +236,8 @@ export async function executeLaunchdOperation(p, receiptFile, deps = {}) {
220
236
  save("health-pending", { rolledBack: true, healthStartedAt: Date.now() }); receipt = read(receiptFile);
221
237
  }
222
238
  if (receipt.status === "bootstrap-intent") {
239
+ // Resume after a crash between bootout and this checkpoint: the old job may still be tearing down.
240
+ if (p.mode === "reload") await unloaded(run, p, deps);
223
241
  current = runtime(run, p.target);
224
242
  if (!current.loaded) {
225
243
  const r = effect(["bootstrap", p.domain, p.plistPath]);
package/docs/PATCHES.md CHANGED
@@ -2741,6 +2741,18 @@ self-restart can disappear; the independent receipt is the source of truth.
2741
2741
  Ten browser error copies now report a browser blocker and suggest a non-browser
2742
2742
  alternative, never a whole-gateway restart. Exec process-group behavior is unchanged.
2743
2743
 
2744
+ **76.1 (2026.9.29), graceful unload before bootstrap.** The first live reload of a real
2745
+ gateway (Trisha, `gateway install --force` semantics) failed with `Service removed while
2746
+ waiting for health` and left the job unloaded. `launchctl bootout` returns before a job
2747
+ that handles SIGTERM has exited; the worker's next `print` still saw the old registered
2748
+ PID, treated the job as loaded, skipped `bootstrap`, and then watched launchd finish the
2749
+ teardown. The worker now waits after `bootout` (and again on resume at `bootstrap-intent`)
2750
+ until the target is unregistered or a different PID is registered, with a 30 s bound that
2751
+ surfaces as a retryable worker error instead of an advanced checkpoint. The real macOS
2752
+ fixture now performs a 700 ms graceful SIGTERM shutdown so the reload cases exercise the
2753
+ race; two unit cases cover the wait and the timeout. `kickstart -k` restarts were never
2754
+ affected.
2755
+
2744
2756
  **Verification:** `tests/launchd-safe-restart.test.mjs` covers lifecycle phases,
2745
2757
  fault injection, error classification, scoping, all copies and exact patch replay.
2746
2758
  `JOHNNESS_TEST_LAUNCHD=1 node --test --test-concurrency=1 tests/launchd-self-restart.macos.test.mjs`
@@ -2761,3 +2773,30 @@ restart` uses the safe worker. For a real plist change, use the new version's
2761
2773
  `gateway install --force`, which stages the change before handoff. Rolling back to
2762
2774
  2026.9.27 or earlier restores the old restart defect; if necessary install it externally and
2763
2775
  use kickstart, never its gateway restart CLI from a gateway child.
2776
+
2777
+
2778
+ ## 77. Weekly provider reasoning corrections (2026-09-20)
2779
+
2780
+ The six-provider audit changes only the DeepSeek, Meta Responses and Gemini adapters.
2781
+ Both DeepSeek V4 models now send native `low` for harness `minimal`/`low`, rather
2782
+ than silently spending at `high`. The existing `xhigh` to `max` compatibility
2783
+ mapping, explicit off, server-default omission, limits and aliases are preserved.
2784
+ Meta Standard Muse 1.2/1.3 now omit `reasoning.effort` when the caller omits it,
2785
+ instead of forcing `minimal`; explicit efforts, validation, encrypted replay and
2786
+ `store: false` are unchanged. Meta account access was not available for live inference.
2787
+
2788
+ `patches/provider-audit-20260920/changes.json` owns the exact three adapter pre/post
2789
+ images plus Meta's historical replay snapshot. Gemini 3.8 Flash and 3.5 Flash-Lite also omit deprecated/ignored temperature.
2790
+ The patcher preflights every target
2791
+ before writing, refuses drift, supports alternate roots and a dependency-only
2792
+ focused refresh. Family 67 reverses only these exact edits during historical hash
2793
+ validation. Both the modern and old Meta replay tests remain byte-identical.
2794
+
2795
+ No model IDs, aliases, config, credentials, default model or owner context caps
2796
+ change. Astra Responses already supports native `max`; Grok native `xhigh` and
2797
+ modern Claude native `max` were already correct. The global `/think max` synonym
2798
+ still means harness `xhigh`; it is not a new distinct UI level in this patch.
2799
+
2800
+ Evidence and official sources: [weekly audit](PROVIDER-AUDIT-2026-09-20.md).
2801
+ Tests: `tests/provider-audit-20260920.test.mjs`, provider catalog, Muse transport
2802
+ and historical replay; verifier 77.1 checks installed adapter hashes exactly.
@@ -0,0 +1,120 @@
1
+ # Six-provider official API audit: 2026-09-20
2
+
3
+ This weekly audit builds on the successful September 13 catalog release, not on
4
+ upstream OpenClaw. Registry IDs, aliases, saved-session IDs, owner context caps,
5
+ credentials and the selected default are unchanged. Family 77 is a focused
6
+ provider-parameter patch, released as `johns-harness@2026.9.30`.
7
+
8
+ ## Changes
9
+
10
+ 1. **DeepSeek:** `minimal`/`low` now serialize as native `low` on `deepseek-flash`
11
+ and `deepseek-v4-pro`, rather than silently increasing effort to `high`.
12
+ Omission remains omitted; explicit low-level off disables thinking. Preserve
13
+ the existing harness `xhigh`/`max` -> native `max` compatibility mapping for
14
+ saved sessions. DeepSeek's own raw `xhigh` synonym maps to `high`; that is not
15
+ a reason to silently lower the harness's established highest-effort behavior.
16
+ 2. **Meta Muse:** omitted effort stays absent instead of forcing `minimal`.
17
+ Explicit effort validation, `store:false`, encrypted reasoning replay and
18
+ tool commentary phases are unchanged. No Contributor/training-consent models.
19
+ 3. **Gemini:** omit deprecated/ignored temperature for `gemini-3.8-flash` and
20
+ `gemini-3.5-flash-lite`. Older model behavior is unchanged. The live API still
21
+ accepts this field, but the official migration guide says to remove it and
22
+ warns that future models reject it. No topP/topK are generated by this adapter.
23
+
24
+ No new supported general-purpose model ID was identified since the prior audit.
25
+ No model was removed and no alias was upgraded to a preview.
26
+
27
+ ## Independent provider results
28
+
29
+ | Provider | Official/account result | Decision |
30
+ |---|---|---|
31
+ | DeepSeek | `/models` HTTP 200, two IDs: `deepseek-flash`, `deepseek-v4-pro`. Latest changelog release September 10, V4.1 Flash. Native low/high/max documented. Flash omission/low/high/xhigh/max and Pro low each HTTP 200. | Correct low serialization; retain all limits and prices. |
32
+ | Meta | Official Standard API confirms Muse 1.3 and retained 1.2 on `https://api.meta.ai/v1/responses`; 1.3 adds native max. No configured Meta credential. | Correct omission from official schema/reasoning docs and actual offline adapter tests; do not claim authenticated inference. |
33
+ | Google | Native model listing HTTP 200 (62 records); 3.8 Flash and 3.5 Flash-Lite confirmed, each 1,048,576 input / 65,536 output. 3.8 minimal smoke HTTP 200. | Omit deprecated temperature; preserve thinking-level mapping. New September 15 Live audio models require a Live adapter, not generateContent chat registration. September 17 Antigravity preview is an agent/Interactions API, not a drop-in model. |
34
+ | OpenAI | Native `/models` HTTP 200. Astra Responses omission/xhigh/max HTTP 200; none HTTP 400 explicitly reports supported low/medium/high/xhigh/max. | Existing native Responses max support is correct. Do not import another application's Chat Completions limitations. No catalog change. |
35
+ | Anthropic | Native `/v1/models` HTTP 200 (10 records) confirms Fable 5.1, Opus 5 and Sonnet 5. Fable adaptive thinking + output_config.effort=max HTTP 200. | Native max was already preserved by the adapter; no change. Account-gated models are not inferred from announcements. |
36
+ | xAI | Native `/v1/models` HTTP 200 (34 records) includes Grok 4.6. Omitted effort and xhigh Chat Completions requests HTTP 200. | xhigh was already supported. Keep exactly 200K owner context cap, 64K output, existing aliases. |
37
+
38
+ The model listing counts are account/time-specific, not availability promises.
39
+ No API list had an unconsumed pagination indicator. No 429 occurred and no
40
+ credential was printed or persisted in evidence. Minimal inference probes used
41
+ synthetic "Reply with OK only" prompts and at most 64 output tokens. Successful
42
+ HTTP parameter acceptance is not a long-output, all-modality or full tool-loop
43
+ inference certification. Offline existing tests separately cover streaming,
44
+ reasoning/signature/encrypted replay, tool calls and errors.
45
+
46
+ ## Metadata and compatibility
47
+
48
+ - DeepSeek Flash is V4.1 Flash (text/image), Pro is V4 Pro 0813 (text).
49
+ Both have 1,048,576 context / 393,216 maximum output. Native default is
50
+ thinking enabled/high. Nonthinking default max output is 8K; thinking high is
51
+ 64K and max is 128K. Temperature has no effect in thinking mode and remains
52
+ omitted. The adapter sends `max_tokens`, not `max_completion_tokens`, and
53
+ does not invent developer-role, store, or strict-tool support.
54
+ - Meta Standard 1.3/1.2 retain 1,048,576 context / 131,072 output and text/image
55
+ input in the harness. Models always reason. Omission lets the model choose
56
+ effort; none/off are invalid. Max is exclusive to Standard 1.3. Standard
57
+ pricing remains $1.25 input / $4.25 output / $0.15 cached input per million.
58
+ - Google current Flash/Lite use native thinking levels rather than a synthetic
59
+ budget. 3.8 MINIMAL remains mapped to LOW by the previously verified adapter;
60
+ 3.5 Lite supports MINIMAL. Omission stays omitted. Native API supports more
61
+ media than the pinned text/image agent interface; this patch does not claim
62
+ Live audio, Interactions, image/video output or unrelated tool adapters.
63
+ - OpenAI native context is 1,050,000 with 128K max output; conservative existing
64
+ harness budgets remain unchanged. Astra cannot disable reasoning. The native
65
+ Responses adapter already passes max exactly. The global `/think max` spelling
66
+ still normalizes to harness xhigh; adding a separate UI enum is not this patch.
67
+ - Modern Claude native effort levels include distinct xhigh and max. Adaptive
68
+ thinking omits manual budgets and sampling. Fable/Opus retain 300K owner caps;
69
+ Sonnet's conservative existing budget remains untouched.
70
+ - Grok 4.6 documents 500K native context / 64K output; the harness deliberately
71
+ retains 200K. Its reasoning_effort xhigh is transmitted exactly. Prices remain
72
+ $2 input / $6 output / $0.50 cached input per million; native long-prompt
73
+ multipliers are not represented by the basic catalog cost estimator.
74
+
75
+ Prices/limits not changed in this release remain at their last verified values;
76
+ there is no blanket claim of re-pricing every legacy registry entry. DeepSeek
77
+ current pricing and Meta Standard pricing were independently rechecked. No price,
78
+ context or output-limit diff was required on the changed entries.
79
+
80
+ ## First-party sources checked
81
+
82
+ - DeepSeek: [catalog/pricing](https://api-docs.deepseek.com/quick_start/pricing/),
83
+ [thinking](https://api-docs.deepseek.com/guides/thinking_mode/),
84
+ [complete Chat Completions reference](https://api-docs.deepseek.com/api/create-chat-completion/),
85
+ [changelog](https://api-docs.deepseek.com/updates).
86
+ - Meta: [models](https://ai.developer.meta.com/docs/models.md),
87
+ [reasoning](https://ai.developer.meta.com/docs/reasoning.md),
88
+ [Responses protocol](https://ai.developer.meta.com/docs/protocols/responses.md),
89
+ [create response](https://ai.developer.meta.com/docs/api-reference/responses/create-response.md),
90
+ [pricing](https://ai.developer.meta.com/docs/pricing-rate-limits.md).
91
+ HTML indexes intermittently returned HTTP 500; official `.md` pages succeeded.
92
+ - Google: [model catalog](https://ai.google.dev/gemini-api/docs/models),
93
+ [thinking](https://ai.google.dev/gemini-api/docs/thinking.md.txt),
94
+ [latest-model migration](https://ai.google.dev/gemini-api/docs/generate-content/latest-model),
95
+ [changelog](https://ai.google.dev/gemini-api/docs/changelog).
96
+ Some index requests redirected repeatedly; native model listing and direct
97
+ content fetch supplied the data. Locale variants do not alter exact API IDs.
98
+ - OpenAI: [catalog](https://developers.openai.com/api/docs/models/),
99
+ [Astra model reference](https://developers.openai.com/api/docs/models/gpt-6-astra.md).
100
+ - Anthropic: [overview](https://platform.claude.com/docs/en/models/overview),
101
+ [Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview),
102
+ [effort](https://platform.claude.com/docs/en/build-with-claude/effort).
103
+ - xAI: [catalog](https://docs.x.ai/developers/models),
104
+ [Grok 4.6](https://docs.x.ai/developers/models/grok-4.6).
105
+
106
+ ## Verification and recovery
107
+
108
+ Family 77 has exact preimage/postimage hashes, all-file preflight before writes,
109
+ read-only checks, alternate-root and dependency-only modes, and idempotent replay.
110
+ Family 67 validates its historical surface after reversing only the exact new
111
+ edits. Meta's Family 53 snapshot is synchronized. Payload tests exercise actual
112
+ installed adapters, not string-only fixtures. Existing alias contract tests and
113
+ all prior catalog/transport replay tests remain part of the release gates.
114
+
115
+ The release is a three-adapter patch on 2026.9.29. Roll back a standard package
116
+ installation to `johns-harness@2026.9.29`; focused installations must restore their
117
+ recorded pre-patch adapter files, not overwrite state or migrate directories.
118
+ A running gateway caches provider modules. Disk validation is not activation:
119
+ coordinate a safe restart after active work finishes, never interrupt jobs merely
120
+ to report a refreshed version.
@@ -1,4 +1,4 @@
1
- # Current provider catalog — checked 2026-09-13
1
+ # Current provider catalog — checked 2026-09-20
2
2
 
3
3
  Family 67 audits general-purpose agent models against first-party documentation and
4
4
  provider-authenticated model lists. It is a targeted registry/adapter refresh on the
@@ -31,7 +31,7 @@ Registry presence does not promise account entitlement.
31
31
 
32
32
  - **DeepSeek V4:** omitted effort leaves `thinking` and `reasoning_effort` absent,
33
33
  preserving the server's default enabled/high. Explicit low-level `off`/`none`
34
- sets `thinking: {type: "disabled"}`. Minimal/low/medium/high map to native `high`;
34
+ sets `thinking: {type: "disabled"}`. Minimal/low map to native `low`; medium/high map to native `high`;
35
35
  xhigh/max map to native `max`. Unknown efforts fail before HTTP. Disabled thinking
36
36
  must not carry `reasoning_effort`. In enabled/default mode temperature is omitted
37
37
  because the API ignores sampling controls. Use `max_tokens`, not
@@ -43,10 +43,12 @@ Registry presence does not promise account entitlement.
43
43
  reasoning replay, and tool-loop commentary `phase`. Always-reasoning;
44
44
  minimal/low/medium/high/xhigh, with native max only on Standard 1.3. The pinned
45
45
  high-level `/think max` spelling still normalizes to xhigh, not a new global enum.
46
+ Omitted effort stays absent and provider-controlled rather than forcing minimal.
46
47
  - **Gemini:** Flash-Lite now enters the native thinking-level branch instead of
47
48
  the legacy thinking-budget branch. It supports MINIMAL/LOW/MEDIUM/HIGH; omission
48
49
  preserves its native minimal default. 3.8 Flash still maps minimal to LOW because
49
50
  the live API rejects MINIMAL (Family 50). No synthetic thinkingBudget is sent.
51
+ Deprecated/ignored temperature is omitted on 3.8 Flash and 3.5 Flash-Lite.
50
52
  - **Claude:** current Fable 5/5.1, Opus 5 and Sonnet 5 preserve native xhigh rather
51
53
  than silently downgrading it to high or promoting it to max. Low-level max remains
52
54
  distinct. Minimal maps to low. Adaptive thinking has no manual token budget.
@@ -60,7 +62,7 @@ Registry presence does not promise account entitlement.
60
62
  public thinking enum still tops out at xhigh. Transport token ceilings are not
61
63
  raised merely because the provider advertises a larger native window.
62
64
  - **Grok:** existing native effort mapping and 200K owner cap retained; live low
63
- effort request accepted by `grok-4.6`.
65
+ and xhigh effort requests accepted by `grok-4.6`.
64
66
 
65
67
  The pinned high-level API represents `/think off` as an absent reasoning option on
66
68
  some routes. Consequently omission and explicit off cannot be distinguished there;
@@ -96,7 +98,8 @@ Subscription cost entries remain zero (subscription-billed), not API-price estim
96
98
 
97
99
  ## Evidence and sources
98
100
 
99
- Checked 2026-09-13; first-party direct pages unless noted. Credential-bearing API
101
+ Baseline sources checked 2026-09-13; six-provider re-audit 2026-09-20 is in
102
+ [the dated audit](PROVIDER-AUDIT-2026-09-20.md). First-party pages unless noted. Credential-bearing API
100
103
  requests were sent only to the relevant official host, with redirects disabled.
101
104
 
102
105
  - DeepSeek: <https://api-docs.deepseek.com/quick_start/pricing/>,
@@ -252,7 +252,8 @@ function createClient(model, apiKey, optionsHeaders) {
252
252
  function buildParams(model, context, options = {}) {
253
253
  const contents = convertMessages(model, context);
254
254
  const generationConfig = {};
255
- if (options.temperature !== undefined) {
255
+ // JOHNNESS_PATCH_PROVIDER_AUDIT_77: sampling is deprecated/ignored on these models.
256
+ if (options.temperature !== undefined && !["gemini-3.8-flash", "gemini-3.5-flash-lite"].includes(model.id)) {
256
257
  generationConfig.temperature = options.temperature;
257
258
  }
258
259
  if (options.maxTokens !== undefined) {
@@ -364,7 +364,10 @@ function buildParams(model, context, options) {
364
364
  } else {
365
365
  if (effort !== undefined) {
366
366
  params.thinking = { type: "enabled" };
367
- params.reasoning_effort = ["xhigh", "max"].includes(effort) ? "max" : "high";
367
+ // JOHNNESS_PATCH_PROVIDER_AUDIT_77: low is native on both V4 routes (2026-09-20).
368
+ // Preserve the harness's existing xhigh -> max compatibility mapping.
369
+ params.reasoning_effort = ["minimal", "low"].includes(effort) ? "low"
370
+ : ["xhigh", "max"].includes(effort) ? "max" : "high";
368
371
  }
369
372
  // Sampling controls have no effect in thinking mode; never imply they do.
370
373
  delete params.temperature;
@@ -176,11 +176,12 @@ function buildParams(model, context, options) {
176
176
  }
177
177
  // JOHNNESS_PATCH_MUSE_SPARK: always-reasoning, stateless encrypted replay.
178
178
  if (model.provider === "meta") {
179
- const effort = options?.reasoningEffort || "minimal";
180
- const allowed = ["minimal", "low", "medium", "high", "xhigh"];
179
+ // JOHNNESS_PATCH_PROVIDER_AUDIT_77: omission keeps Meta's model-selected effort.
180
+ const effort = options?.reasoningEffort;
181
+ const allowed = [undefined, "minimal", "low", "medium", "high", "xhigh"];
181
182
  if (model.id === "muse-spark-1.3") allowed.push("max");
182
183
  if (!allowed.includes(effort)) throw new Error("Muse Spark always reasons. Use minimal, low, medium, high or xhigh (max only on Standard 1.3); none/off is not supported.");
183
- params.reasoning = { effort, summary: options?.reasoningSummary || "auto" };
184
+ params.reasoning = { ...(effort === undefined ? {} : { effort }), summary: options?.reasoningSummary || "auto" };
184
185
  params.include = ["reasoning.encrypted_content"];
185
186
  params.store = false;
186
187
  // Avoid inventing Meta pricing multipliers from OpenAI service tiers.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "johns-harness",
3
- "version": "2026.9.28",
3
+ "version": "2026.9.30",
4
4
  "description": "John's Harness: a production agent harness that runs autonomous coding agents as one cooperative swarm.",
5
5
  "keywords": [
6
6
  "ai-agents",
@@ -252,7 +252,8 @@ function createClient(model, apiKey, optionsHeaders) {
252
252
  function buildParams(model, context, options = {}) {
253
253
  const contents = convertMessages(model, context);
254
254
  const generationConfig = {};
255
- if (options.temperature !== undefined) {
255
+ // JOHNNESS_PATCH_PROVIDER_AUDIT_77: sampling is deprecated/ignored on these models.
256
+ if (options.temperature !== undefined && !["gemini-3.8-flash", "gemini-3.5-flash-lite"].includes(model.id)) {
256
257
  generationConfig.temperature = options.temperature;
257
258
  }
258
259
  if (options.maxTokens !== undefined) {
@@ -364,7 +364,10 @@ function buildParams(model, context, options) {
364
364
  } else {
365
365
  if (effort !== undefined) {
366
366
  params.thinking = { type: "enabled" };
367
- params.reasoning_effort = ["xhigh", "max"].includes(effort) ? "max" : "high";
367
+ // JOHNNESS_PATCH_PROVIDER_AUDIT_77: low is native on both V4 routes (2026-09-20).
368
+ // Preserve the harness's existing xhigh -> max compatibility mapping.
369
+ params.reasoning_effort = ["minimal", "low"].includes(effort) ? "low"
370
+ : ["xhigh", "max"].includes(effort) ? "max" : "high";
368
371
  }
369
372
  // Sampling controls have no effect in thinking mode; never imply they do.
370
373
  delete params.temperature;
@@ -176,11 +176,12 @@ function buildParams(model, context, options) {
176
176
  }
177
177
  // JOHNNESS_PATCH_MUSE_SPARK: always-reasoning, stateless encrypted replay.
178
178
  if (model.provider === "meta") {
179
- const effort = options?.reasoningEffort || "minimal";
180
- const allowed = ["minimal", "low", "medium", "high", "xhigh"];
179
+ // JOHNNESS_PATCH_PROVIDER_AUDIT_77: omission keeps Meta's model-selected effort.
180
+ const effort = options?.reasoningEffort;
181
+ const allowed = [undefined, "minimal", "low", "medium", "high", "xhigh"];
181
182
  if (model.id === "muse-spark-1.3") allowed.push("max");
182
183
  if (!allowed.includes(effort)) throw new Error("Muse Spark always reasons. Use minimal, low, medium, high or xhigh (max only on Standard 1.3); none/off is not supported.");
183
- params.reasoning = { effort, summary: options?.reasoningSummary || "auto" };
184
+ params.reasoning = { ...(effort === undefined ? {} : { effort }), summary: options?.reasoningSummary || "auto" };
184
185
  params.include = ["reasoning.encrypted_content"];
185
186
  params.store = false;
186
187
  // Avoid inventing Meta pricing multipliers from OpenAI service tiers.