instar 1.3.792 → 1.3.794
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/dist/commands/server.d.ts.map +1 -1
- package/dist/commands/server.js +10 -0
- package/dist/commands/server.js.map +1 -1
- package/dist/core/IntelligenceRouter.d.ts +35 -0
- package/dist/core/IntelligenceRouter.d.ts.map +1 -1
- package/dist/core/IntelligenceRouter.js +118 -1
- package/dist/core/IntelligenceRouter.js.map +1 -1
- package/dist/core/ModelTierEscalation.d.ts +1 -1
- package/dist/core/ModelTierEscalation.d.ts.map +1 -1
- package/dist/core/ModelTierEscalation.js +7 -0
- package/dist/core/ModelTierEscalation.js.map +1 -1
- package/dist/core/PostUpdateMigrator.d.ts.map +1 -1
- package/dist/core/PostUpdateMigrator.js +16 -0
- package/dist/core/PostUpdateMigrator.js.map +1 -1
- package/dist/core/types.d.ts +22 -0
- package/dist/core/types.d.ts.map +1 -1
- package/dist/core/types.js.map +1 -1
- package/dist/scaffold/templates.d.ts.map +1 -1
- package/dist/scaffold/templates.js +1 -0
- package/dist/scaffold/templates.js.map +1 -1
- package/dist/server/routes.d.ts.map +1 -1
- package/dist/server/routes.js +4 -1
- package/dist/server/routes.js.map +1 -1
- package/package.json +1 -1
- package/scripts/model-registry-freshness.manifest.json +6 -3
- package/src/data/builtin-manifest.json +63 -63
- package/src/scaffold/templates.ts +1 -0
- package/upgrades/1.3.793.md +80 -0
- package/upgrades/1.3.794.md +19 -0
- package/upgrades/side-effects/codex-gpt56-allowlist.md +60 -0
- package/upgrades/side-effects/nongating-failure-swap.md +54 -0
|
@@ -0,0 +1,80 @@
|
|
|
1
|
+
# Upgrade Guide — vNEXT
|
|
2
|
+
|
|
3
|
+
<!-- assembled-by: assemble-next-md -->
|
|
4
|
+
<!-- bump: patch -->
|
|
5
|
+
|
|
6
|
+
## What Changed
|
|
7
|
+
|
|
8
|
+
The `IntelligenceRouter` failure-swap now covers NON-gating internal calls, fixing a class where a
|
|
9
|
+
non-gating background component hard-errored instead of trying a healthy fallback door.
|
|
10
|
+
|
|
11
|
+
Background: internal components run off Claude by default (Codex → Pi → Gemini → Claude). When a
|
|
12
|
+
GATING call's primary provider fails at runtime it already swaps down that chain. NON-gating calls
|
|
13
|
+
did not — they re-threw straight to the caller's heuristic. In production, `TopicIntentExtractor`
|
|
14
|
+
(non-gating, routed to codex-cli/gpt-5.4-mini) showed a 28% error rate (122/428 over 7 days), every
|
|
15
|
+
error row zero-usage — i.e. the codex `exec` invocation itself failed/timed out/returned empty. The
|
|
16
|
+
GATING calls beside it errored ~1.5% precisely because they swap.
|
|
17
|
+
|
|
18
|
+
The fix extends the swap to non-gating calls with a TIGHTER bound than gating calls get:
|
|
19
|
+
- It fires ONLY on an INVOCATION-level primary failure (the primary threw AND produced zero tokens).
|
|
20
|
+
A content/parse error that carried tokens does NOT swap (the caller fail-opens that, per
|
|
21
|
+
provider-fallback §6.4). The router composes an `onUsage` capture onto the primary attempt to
|
|
22
|
+
observe token production; a provider that never surfaces usage (gemini) is treated as
|
|
23
|
+
invocation-level (the conservative, error-reducing direction).
|
|
24
|
+
- At most `maxAttempts` (default 1) steps down the active `failureSwap` tail, each target
|
|
25
|
+
circuit-checked and bounded by the existing `intelligence.swapAttemptTimeoutMs` per-attempt cap
|
|
26
|
+
(also flowed through as the provider's `timeoutMs`).
|
|
27
|
+
- NEVER onto `claude-code` or the default framework — the provider-fallback §6.2 invariant that
|
|
28
|
+
non-gating background traffic must never herd onto the last-resort Claude tail. If the only
|
|
29
|
+
remaining tail entry is claude-code, the call re-throws to its heuristic (today's behavior).
|
|
30
|
+
- Metrics honesty is automatic: each provider's own CircuitBreaking wrapper records its own
|
|
31
|
+
feature_metrics row (the failed codex primary keeps its zero-usage error row; the pi swap records
|
|
32
|
+
pi's usage/model). `usageCoverage` is unaffected.
|
|
33
|
+
|
|
34
|
+
New config: `intelligence.nonGatingFailureSwap: { enabled?: boolean; maxAttempts?: number }`,
|
|
35
|
+
inline-defaulted at the router construction site (`enabled` default true; the codexExecJson /
|
|
36
|
+
swapAttemptTimeoutMs precedent — no ConfigDefaults/migrateConfig entry). Set `enabled: false` to
|
|
37
|
+
restore the old hard-error behavior. Gating-call behavior, deferrable behavior,
|
|
38
|
+
`sessions.componentFrameworks` semantics, the spawn-cap funnel, and nature-routing are all
|
|
39
|
+
untouched. On a Claude-only agent (no off-Claude CLI) the whole thing is a no-op.
|
|
40
|
+
|
|
41
|
+
## What to Tell Your User
|
|
42
|
+
|
|
43
|
+
Some of my quick background helpers — like the one that works out what a conversation topic is
|
|
44
|
+
about — used to just fail whenever the tool they run on had a brief hiccup, even though a healthy
|
|
45
|
+
backup tool was sitting right there. In real numbers, one of them was failing more than a quarter
|
|
46
|
+
of the time for exactly this reason. Now, when a background helper's tool fails to run, I quietly
|
|
47
|
+
try one backup tool before giving up, so far fewer of these little checks fall over. It never adds
|
|
48
|
+
a noticeable wait, it never pushes that background chatter onto your main Claude account, and it
|
|
49
|
+
only kicks in when the tool genuinely failed to run — not when it answered. You don't have to do
|
|
50
|
+
anything; it's on by default and it just makes me more reliable. If you ever want the old behavior
|
|
51
|
+
back, I can switch it off for you.
|
|
52
|
+
|
|
53
|
+
## Summary of New Capabilities
|
|
54
|
+
|
|
55
|
+
- Non-gating internal calls now get a bounded, herd-safe provider swap on an invocation-level
|
|
56
|
+
failure, instead of hard-erroring — sharply reducing user-visible errors on components like the
|
|
57
|
+
topic classifier.
|
|
58
|
+
- New off-switch: `intelligence.nonGatingFailureSwap.enabled: false` restores the old hard-error
|
|
59
|
+
behavior; `maxAttempts` tunes how many tail steps a non-gating call may take (default 1).
|
|
60
|
+
- Proactive trigger for the agent: "why did my background classifier's error rate drop / does a
|
|
61
|
+
non-gating call fall back too?" → this bounded swap.
|
|
62
|
+
|
|
63
|
+
## Evidence
|
|
64
|
+
|
|
65
|
+
Reproduction (production, 2026-07-09, from `GET /metrics/features` 7d + the overnight routing
|
|
66
|
+
investigation `/.instar/plans/overnight-routing-error-investigation.md`): `TopicIntentExtractor`
|
|
67
|
+
`byModel` (codex-cli/gpt-5.4-mini) = 434 calls, 123 errors, `errorRowsWithUsage: 0`, `fired: 0`,
|
|
68
|
+
311 noop — a 28% error rate. The zero-usage on every error row is the tell: these are
|
|
69
|
+
invocation-level codex-exec failures (no tokens produced), not rate-limits (codex ~3% used) and not
|
|
70
|
+
parse errors (those carry tokens). The GATING `MessagingToneGate` errored at 1.5% and
|
|
71
|
+
`CoherenceReviewer` at 2.6% because they ride the failure-swap tail; the non-gating call did not.
|
|
72
|
+
|
|
73
|
+
Before: a non-gating primary invocation failure re-throws immediately → the 28% user-visible error
|
|
74
|
+
rate. After: the same failure swaps once onto the next active off-Claude framework (pi-cli, ~1.5%),
|
|
75
|
+
so the call succeeds instead of erroring. Verified by driving the exact router path end-to-end:
|
|
76
|
+
`tests/unit/nongating-failure-swap.test.ts` (13 — invocation-failure→swap, content-error→no-swap,
|
|
77
|
+
disabled/absent→old behavior, target-down→re-throw original, herd-safety never-onto-Claude while
|
|
78
|
+
gating still swaps, maxAttempts bound, tier preserved, slow-target abandoned at the cap),
|
|
79
|
+
`tests/integration/nongating-failure-swap-routing.test.ts` (3), and
|
|
80
|
+
`tests/e2e/nongating-failure-swap-lifecycle.test.ts` (2, real AgentServer init path, default-ON).
|
|
@@ -0,0 +1,19 @@
|
|
|
1
|
+
# Upgrade Guide — vNEXT
|
|
2
|
+
|
|
3
|
+
<!-- assembled-by: assemble-next-md -->
|
|
4
|
+
<!-- bump: patch -->
|
|
5
|
+
|
|
6
|
+
## What Changed
|
|
7
|
+
|
|
8
|
+
The GPT-5.6 model family (`gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-5.6-luna`) went GA on the Codex subscription and is now on instar's accepted-model allowlists for the `codex-cli` framework. Two mirrored lists gained the three ids: `KNOWN_CODEX_MODEL_IDS` (the model-tier-escalation resolver's closed enum, `src/core/ModelTierEscalation.ts`) and `CODEX_MODELS_SUBSCRIPTION` (the `POST /sessions/spawn` model validator, `src/server/routes.ts`). The topic-profile validator reads the same enum, so it accepts the new ids automatically. The fail-closed design is unchanged — an id outside the closed enum is still rejected; we only widened the set of recognized ids.
|
|
9
|
+
|
|
10
|
+
The Doorway/Model Knowledge Registry (`scripts/model-registry-freshness.manifest.json`) also learned the three models as recognized, reachable entries with pricing, so `GET /doorways` knows they exist. They are deliberately NOT promoted to the frontier pin — the capable tier still resolves to `gpt-5.5`; a benchmark-driven promotion is a tracked follow-up.
|
|
11
|
+
|
|
12
|
+
## What to Tell Your User
|
|
13
|
+
|
|
14
|
+
If you run Codex sessions, you can now pin a topic or spawn a session on the new GPT-5.6 models — sol (the flagship), terra (mid-size), or luna (small and cheap) — and instar accepts them instead of rejecting them as unknown. This needs the Codex CLI at version 0.144.0 or newer; older CLIs reject GPT-5.6 with a "requires a newer version" error, so update the Codex CLI if you hit that. If your setup already escalates heavy work onto the GPT-5.6 flagship, it becomes live with this update. The Pro variants are intentionally not accepted yet (they're likely plan-gated and pricier — a future follow-up), and for now instar still picks GPT-5.5 for its heavy tier until the new models earn that spot through benchmarks.
|
|
15
|
+
|
|
16
|
+
## Summary of New Capabilities
|
|
17
|
+
|
|
18
|
+
- `codex-cli` sessions and topic profiles accept `gpt-5.6-sol`, `gpt-5.6-terra`, and `gpt-5.6-luna` (model-tier escalation, `/sessions/spawn`, and topic-profile pins).
|
|
19
|
+
- `GET /doorways` reports the three GPT-5.6 models (with pricing) as recognized, reachable Codex models.
|
|
@@ -0,0 +1,60 @@
|
|
|
1
|
+
# Side-Effects Review — Add the GPT-5.6 family to the codex model allowlists
|
|
2
|
+
|
|
3
|
+
**Change:** additive model-id allowlist extension for the `codex-cli` framework. **Tier 1** (small, low-risk, fail-closed design unchanged). **Parent principle:** Structure > Willpower (the closed model-id enumeration is a code-enforced gate; this widens the allowed set, it does not weaken the gate).
|
|
4
|
+
|
|
5
|
+
**Files changed:**
|
|
6
|
+
- `src/core/ModelTierEscalation.ts` — `KNOWN_CODEX_MODEL_IDS` gains `gpt-5.6-sol`, `gpt-5.6-terra`, `gpt-5.6-luna`.
|
|
7
|
+
- `src/server/routes.ts` — `CODEX_MODELS_SUBSCRIPTION` (the `POST /sessions/spawn` codex model validator) gains the same three ids, keeping the two lists identical (the comment pins them as mirrors).
|
|
8
|
+
- `scripts/model-registry-freshness.manifest.json` — the Doorway/Model Knowledge Registry's `codex-cli.topModels` gains the three ids as recognized entries (`frontier: false`, pricing populated) + the door note / `$flaggedStaleNote` prose corrected to reflect the 2026-07-09 GA (the earlier "gpt-5.6-sol is preview-only/unreachable" line was now false).
|
|
9
|
+
- `tests/unit/modelTierEscalation-resolver.test.ts`, `tests/unit/route-validation-edge.test.ts`, `tests/unit/topicProfileValidation.test.ts` — allowlist + resolver + spawn + validation coverage.
|
|
10
|
+
- `docs/specs/codex-gpt56-allowlist.eli16.md`, `upgrades/next/codex-gpt56-allowlist.md` — the ELI16 + release fragment.
|
|
11
|
+
|
|
12
|
+
## What changed
|
|
13
|
+
|
|
14
|
+
The GPT-5.6 family (sol/terra/luna) went GA on the codex subscription on 2026-07-09 and was live-verified working (codex CLI >= 0.144.0 required; older CLIs 400 with "requires a newer version"). instar fails closed on model ids outside its closed per-framework enums, so these ids were rejected everywhere. This change adds the three ids to the two mirrored acceptance lists (`KNOWN_CODEX_MODEL_IDS` + `CODEX_MODELS_SUBSCRIPTION`). The `-pro` variants are deliberately excluded (plan-gated + pricier — tracked follow-up).
|
|
15
|
+
|
|
16
|
+
## Lists touched vs deliberately left
|
|
17
|
+
|
|
18
|
+
**Touched (acceptance / validation lists):**
|
|
19
|
+
- `KNOWN_CODEX_MODEL_IDS` (`src/core/ModelTierEscalation.ts`) — the escalation resolver's closed enum, and (via `KNOWN_MODEL_IDS`) the source for `topicProfileValidation.validateModelId`. One edit flows to both surfaces.
|
|
20
|
+
- `CODEX_MODELS_SUBSCRIPTION` (`src/server/routes.ts`) — the `/sessions/spawn` codex-cli model validator. Kept byte-identical in membership to the enum above (they are documented mirrors).
|
|
21
|
+
- `scripts/model-registry-freshness.manifest.json` → `doors.codex-cli.topModels` — the Doorway/Model Knowledge registry, added the three ids with `frontier: false` + pricing so `GET /doorways` knows they exist and their cost.
|
|
22
|
+
|
|
23
|
+
**Deliberately LEFT (routing/selection — a model earns a lane via benchmarks, not by GA date):**
|
|
24
|
+
- `src/providers/adapters/openai-codex/models.ts` `TIER_MODEL` (capable → gpt-5.5, fast/balanced → gpt-5.4-mini) — model CHOICES per tier, not an acceptance list.
|
|
25
|
+
- `src/core/frameworkSessionLaunch.ts` tier→model resolution — same reason.
|
|
26
|
+
- `src/data/llmBenchCoverage.ts` (`ROUTING_LABEL_TO_MODEL_ID`, bench coverage, nature-routing chains) — the routing chains that pick models for lanes; models earn these via benchmarks.
|
|
27
|
+
- The doorway manifest's `codex-capable-tier` PIN and the `frontier: true` flag — left at `gpt-5.5`. Promoting `gpt-5.6-sol` to the capable pin / frontier is a benchmark-driven follow-up, not this acceptance PR.
|
|
28
|
+
|
|
29
|
+
## Blast radius
|
|
30
|
+
|
|
31
|
+
- **Purely additive.** No id removed, no existing id's behavior changed. A caller not asking for a GPT-5.6 id is entirely unaffected.
|
|
32
|
+
- **Fail-closed preserved.** A well-shaped id outside the enum is still rejected (`id-not-in-closed-enum`); a made-up `gpt-9.9-fake` still 400s at spawn and resolves to null in escalation — both covered by new tests.
|
|
33
|
+
- **No new route, no new config default, no schema change.** No dark-gate line shift (no `enabled:` line added to `ConfigDefaults.ts`).
|
|
34
|
+
|
|
35
|
+
## Risk + mitigation
|
|
36
|
+
|
|
37
|
+
- **Risk:** the two mirror lists drift apart. **Mitigation:** both edited in the same commit with a cross-referencing comment on each; new tests assert acceptance through both the resolver and the spawn route.
|
|
38
|
+
- **Risk:** the doorway manifest edit trips the CI-gating model-registry-freshness lint. **Mitigation:** the new entries are `frontier: false`, so the codex-cli derived frontier set stays `['gpt-5.5']` and the `codex-capable-tier` pin (still `gpt-5.5`) remains a member — the drift tooth is unaffected. Staleness is unchanged (`lastReviewedAt` untouched, well within the 45-day window). Verified: `node scripts/lint-model-registry-freshness.mjs` → PASS.
|
|
39
|
+
- **Risk:** a user on an old codex CLI selects a GPT-5.6 id and gets an opaque failure. **Mitigation:** the ELI16 + release fragment both call out the codex CLI >= 0.144.0 requirement explicitly.
|
|
40
|
+
|
|
41
|
+
## Framework generality
|
|
42
|
+
|
|
43
|
+
The change is scoped to the `codex-cli` framework's own acceptance surfaces and routes through the existing per-framework `KNOWN_MODEL_IDS` / spawn-validator abstraction — it does not touch the session-launch/inject abstraction and makes no Claude-specific assumption. Other frameworks (claude-code, gemini-cli, pi-cli) are untouched.
|
|
44
|
+
|
|
45
|
+
## Migration parity
|
|
46
|
+
|
|
47
|
+
- No agent-installed file changes (no `.claude/settings.json`, no `.instar/config.json` default, no CLAUDE.md template section, no hook/skill). The allowlists are shipped code read at runtime, so existing agents pick up the new ids on the normal server/dist update — no `PostUpdateMigrator` entry needed. The operator's already-set `models.tierEscalation.frameworks.codex-cli.escalated = "gpt-5.6-sol"` becomes live purely by deploying this code.
|
|
48
|
+
|
|
49
|
+
## Tests
|
|
50
|
+
|
|
51
|
+
- `tests/unit/modelTierEscalation-resolver.test.ts` — codex-cli escalated tier accepts each of gpt-5.6-sol/terra/luna; a made-up id fails closed (`id-not-in-closed-enum`); the `-pro` variants are asserted absent.
|
|
52
|
+
- `tests/unit/route-validation-edge.test.ts` — `POST /sessions/spawn` with framework `codex-cli` accepts the three GPT-5.6 ids (not 400) and rejects `gpt-9.9-fake` (400, error names "model").
|
|
53
|
+
- `tests/unit/topicProfileValidation.test.ts` — `validateModelId(..., 'codex-cli')` returns null for the three ids (the KNOWN_MODEL_IDS mirror path).
|
|
54
|
+
- Adjacent suites re-run green: `model-registry-freshness`, `codex-model-tier-resolution`, `frameworkSessionLaunch`, `model-tier-swap-route`, `model-tier-escalation-lifecycle`. `npx tsc --noEmit` clean.
|
|
55
|
+
|
|
56
|
+
## Follow-ups
|
|
57
|
+
|
|
58
|
+
- Add the `-pro` GPT-5.6 variants once their plan-gating + pricing are confirmed.
|
|
59
|
+
- Benchmark GPT-5.6-sol and, if it wins the lane, promote it to the codex capable tier pin (`TIER_MODEL.capable` + the manifest `frontier: true` + `lastReviewedAt` bump) and into the routing chains — a routing decision, not an acceptance change.
|
|
60
|
+
- Populate real per-token pricing across the doorway registry if/when the metered doorway-scan scopes consume it.
|
|
@@ -0,0 +1,54 @@
|
|
|
1
|
+
# Side-Effects Review — Non-Gating Failure-Swap (bounded provider swap for non-gating internal calls)
|
|
2
|
+
|
|
3
|
+
**Spec:** docs/specs/nongating-failure-swap.md (Tier-1 bug fix — bounded extension of the CONVERGED + approved `docs/specs/provider-fallback-default-policy.md`). **Parent principle:** No Silent Degradation to Brittle Fallback.
|
|
4
|
+
**Ships ON by default** (`intelligence.nonGatingFailureSwap.enabled`, inline-defaulted `?? true` at the router construction site — no persisted config block). No-op on a Claude-only agent (no off-Claude tail) and on any router constructed without the field (e.g. unit tests).
|
|
5
|
+
**Files:** src/core/IntelligenceRouter.ts, src/core/types.ts, src/commands/server.ts, src/scaffold/templates.ts, src/core/PostUpdateMigrator.ts, docs/specs/nongating-failure-swap.md (new), docs/specs/nongating-failure-swap.eli16.md (new), upgrades/next/nongating-failure-swap.md (new), upgrades/side-effects/nongating-failure-swap.md (new), tests/unit/nongating-failure-swap.test.ts (new), tests/unit/PostUpdateMigrator-nonGatingFailureSwap.test.ts (new), tests/integration/nongating-failure-swap-routing.test.ts (new), tests/e2e/nongating-failure-swap-lifecycle.test.ts (new)
|
|
6
|
+
|
|
7
|
+
## What changed
|
|
8
|
+
|
|
9
|
+
1. **IntelligenceRouter.ts — `IntelligenceRouterOptions`:** new optional `nonGatingFailureSwap?: { enabled: boolean; maxAttempts?: number }`. Absent ⇒ feature OFF (byte-identical legacy — a non-gating primary failure re-throws straight to the caller's heuristic).
|
|
10
|
+
2. **IntelligenceRouter.ts — `evaluate()`:** after the existing `swapPositions`/`gatingDeadlineAt` computation, compute `nonGatingSwapEligible = !gating && !deferrable && !enforced && nonGatingFailureSwap.enabled === true && cfg.failureSwap.length > 0`. On the eligible path ONLY, compose an `onUsage` capture onto the primary attempt (`primaryEvalOptions`) so `primaryProducedTokens` records whether the primary produced any tokens; gating/deferrable/enforced calls use `evalOptions` verbatim (byte-identical). Inside the existing `if (swapPositions.length === 0)` branch, BEFORE the deferrable-queue + heuristic-fallthrough, if `nonGatingSwapEligible && !primaryProducedTokens` call the new `tryNonGatingSwap(...)`; on success return its result (before any heuristic-fallthrough tracking).
|
|
11
|
+
3. **IntelligenceRouter.ts — new `tryNonGatingSwap()`:** attempts at most `maxAttempts` (default 1) steps down `cfg.failureSwap`, FILTERING OUT `claude-code`, the default framework, and the just-failed primary. Each target is `resolveProvider`-checked (binary-missing/circuit-open → skipped/caught) and bounded by the SAME `resolveSwapCap` + `withSwapTimeout` machinery the gating loop uses (the cap also flows through as the provider's `timeoutMs`). Emits `onDegrade` (`nongating-failure-swap:` on success, `nongating-swap-attempt-timeout:` on a cap fire) + `onResolved` on success. Returns `{ ok }`; on `{ ok:false }` the caller falls through to its existing heuristic (`onHeuristicFallthrough` + `throw err`).
|
|
12
|
+
4. **types.ts:** new `intelligence.nonGatingFailureSwap?: { enabled?: boolean; maxAttempts?: number }` config field, documented as inline-defaulted (codexExecJson/swapAttemptTimeoutMs precedent — deliberately NOT in ConfigDefaults/migrateConfig).
|
|
13
|
+
5. **server.ts (router construction):** wire `nonGatingFailureSwap: { enabled: config.intelligence?.nonGatingFailureSwap?.enabled ?? true, maxAttempts: config.intelligence?.nonGatingFailureSwap?.maxAttempts }` — the default-ON expression.
|
|
14
|
+
6. **templates.ts + PostUpdateMigrator.ts:** a bullet under Per-Component Framework Routing (new agents) + an idempotent content-sniffed `migrateClaudeMd` corrective subsection (existing agents), marker `non-gating internal calls also get a bounded`.
|
|
15
|
+
|
|
16
|
+
## Blast radius
|
|
17
|
+
|
|
18
|
+
- **Gating / deferrable / nature-enforced paths are untouched.** `nonGatingSwapEligible` is false for all of them, so `primaryEvalOptions === evalOptions` (no capture) and the new branch is never entered. The gating swap loop, the deferrable backoff/queue rungs, and the enforced-nature selection are byte-identical.
|
|
19
|
+
- **No new HTTP route, no new provider, no new spawn.** The non-gating swap reuses the existing per-framework providers (already built at boot via `buildProvider`) and the existing per-attempt cap machinery. `tryNonGatingSwap` never builds a new provider or spawns beyond what a normal swap attempt does.
|
|
20
|
+
- **Bounded blast on the swap itself:** at most `maxAttempts` (default 1) steps, each circuit-checked, each capped by `swapAttemptTimeoutMs` (default 5s). Worst-case added latency on a non-gating failure = `maxAttempts × cap`.
|
|
21
|
+
- **Off-Claude only.** `claude-code` and the default framework are FILTERED OUT of non-gating targets, so this can never push non-gating background traffic onto the last-resort Claude tail (the §6.2 herd invariant). On a Claude-only agent `cfg` is undefined / the tail is empty → strict no-op.
|
|
22
|
+
|
|
23
|
+
## Risk + mitigation
|
|
24
|
+
|
|
25
|
+
- **Risk:** reintroduces the §6.2 herd (non-gating traffic floods a fallback under a broad rate-limit). **Mitigation:** the non-gating swap is STRICTLY more conservative than the gating swap — one step (default), circuit-checked (a target whose breaker is open throws fast → skipped), and NEVER onto Claude. Under a genuine rate-limit the target's own breaker damps repeat attempts. Proven by the herd-safety lens test (`never onto claude-code … but a GATING call does`) and the maxAttempts-bound test.
|
|
26
|
+
- **Risk:** swapping on a content/parse error double-spends tokens on a request that already burned some. **Mitigation:** the swap fires ONLY when the primary produced ZERO tokens (`primaryProducedTokens` false). A token-carrying failure is NOT swapped — the caller fail-opens it (§6.4). Proven by the `content/parse error that CARRIED tokens → NO swap` test.
|
|
27
|
+
- **Risk:** a slow fallback adds latency to a high-volume noop path. **Mitigation:** the per-attempt cap (`swapAttemptTimeoutMs`, default 5s) abandons a slow target via `withSwapTimeout` (the shipped crash-safe Promise.race form; timer cleared on settle). Proven by the `SLOW target abandoned at the cap` test.
|
|
28
|
+
- **Risk:** the `onUsage` capture interferes with the primary's own metrics/usage recording. **Mitigation:** the capture COMPOSES with the caller's onUsage (`callerOnUsage?.(u)`) and is downstream of the CircuitBreaking wrapper's own capture — additive, no clobber. Metrics honesty is automatic (each provider's wrapper records its own row keyed by serving framework/model); `usageCoverage` is unaffected.
|
|
29
|
+
- **Risk:** an error in the swap helper breaks the LLM call path. **Mitigation:** every path in `tryNonGatingSwap` ends at either a returned result or `{ ok:false }` → the caller's existing `throw err` (heuristic). It never introduces a new fail-closed and never swallows silently — the catch emits `onDegrade` on a cap fire and `continue`s (the same non-silent resilience pattern as the gating loop; not counted by the no-silent-fallbacks ratchet).
|
|
30
|
+
|
|
31
|
+
## Migration parity
|
|
32
|
+
|
|
33
|
+
- **Config:** no `migrateConfig` needed — the knob is inline-defaulted at the construction site (`?? true`), so existing agents pick up the default-ON behavior purely from the new code shipping (the codexExecJson/swapAttemptTimeoutMs precedent). Absence ⇒ enabled default.
|
|
34
|
+
- **CLAUDE.md:** `generateClaudeMd` gains the bullet (new agents); `migrateClaudeMd` appends an idempotent content-sniffed corrective subsection (existing agents), marker `non-gating internal calls also get a bounded`. Covered by `tests/unit/PostUpdateMigrator-nonGatingFailureSwap.test.ts` (add-when-absent, idempotent, preserves content, skips when missing) + a template-emits-it assertion.
|
|
35
|
+
|
|
36
|
+
## Dark-gate line-map
|
|
37
|
+
|
|
38
|
+
- UNCHANGED. `nonGatingFailureSwap` is inline-defaulted in `src/commands/server.ts` (`?? true`) and declared as an optional type in `types.ts`; it is NOT an `enabled:` line in `ConfigDefaults.ts`. The dark-gate attributor reads `ConfigDefaults.ts` only and matches `enabled:` lines, so no line shifted. Verified: `tests/unit/lint-dev-agent-dark-gate.test.ts` → green in the run batch.
|
|
39
|
+
|
|
40
|
+
## Rollback
|
|
41
|
+
|
|
42
|
+
- Set `intelligence.nonGatingFailureSwap.enabled: false` → non-gating failures re-throw to the heuristic with no swap (today's behavior), no restart-to-rewire needed (config is read live in `resolveConfig`; the field is read at construction, so a restart is needed only if the operator wants to change it after boot — same posture as `swapAttemptTimeoutMs`). To fully revert: remove the `nonGatingFailureSwap` option + `tryNonGatingSwap` + the eligibility/capture block in `evaluate()` + the server wiring + the type + the CLAUDE.md bullet/migration. Additive throughout.
|
|
43
|
+
|
|
44
|
+
## Tests
|
|
45
|
+
|
|
46
|
+
- `tests/unit/nongating-failure-swap.test.ts` (13) — the core behavior + both sides of every decision boundary: invocation-failure → one swap; content-error-with-usage → no swap; disabled + absent → old behavior; target down/circuit-open → skip + re-throw ORIGINAL error; gemini-primary (no usage) → conservative swap; herd-safety (never onto claude-code/default while GATING still does); gating unchanged (full-tail swap); maxAttempts=1 vs 2; model tier preserved; per-attempt cap passthrough; slow target abandoned at the cap.
|
|
47
|
+
- `tests/integration/nongating-failure-swap-routing.test.ts` (3) — a production-shaped router (computed default + the knob) SWAPS on a non-gating invocation failure; `GET /intelligence/routing` is unchanged (resolution, not swap); `{ enabled:false }` hard-errors.
|
|
48
|
+
- `tests/e2e/nongating-failure-swap-lifecycle.test.ts` (2) — real AgentServer init path: the intelligence-routing route is alive AND the wired router performs the swap via the SHIPPED default expression (config unset ⇒ enabled:true), proving the feature is alive + ON, not dark.
|
|
49
|
+
- `tests/unit/PostUpdateMigrator-nonGatingFailureSwap.test.ts` (5) — the migrateClaudeMd corrective (add/idempotent/preserve/skip) + template-emits-it.
|
|
50
|
+
- Regression: `no-silent-fallbacks`, `lint-dev-agent-dark-gate`, `provider-fallback-swap-timeout`, `per-target-swap-timeout`, `internalFrameworkDefault`, `intelligence-router`, `nature-routing-resolver`, `degradation-ladder`, `opus-claude-cli-gating-guardrail`, `provider-fallback-default-routing`, `intelligence-routing-routes/lifecycle` all green. tsc clean.
|
|
51
|
+
|
|
52
|
+
## Agent awareness
|
|
53
|
+
|
|
54
|
+
- A "Non-gating calls also get a bounded swap now" bullet extends the Per-Component Framework Routing section in `generateClaudeMd`, and an idempotent `migrateClaudeMd` corrective subsection reaches existing agents. Proactive trigger documented: "why did my background classifier's error rate drop / does a non-gating call fall back too?".
|