prism-mcp-server 20.20.0 → 20.21.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -32,8 +32,11 @@ A paid subscription adds cloud sync, higher model tiers, and team features throu
32
32
  undercounts inline — a number you can check, not marketing.
33
33
  - **Route-output enforcement** — route mode returns only well-formed calls to
34
34
  tools the host actually advertised. Standard and higher plans can add
35
- authenticated deterministic correction; `route_guard: "local"` keeps the
36
- prompt and draft entirely on-device.
35
+ authenticated deterministic correction; `route_guard: "local"` disables that
36
+ correction only. `cloud_fallback: false` forbids cloud inference fallback and
37
+ `verify: false` (with no `evidence`) disables the grounding verifier. With all
38
+ three off, no request carries your prompt, draft or evidence; the per-call
39
+ entitlement check and telemetry still contact the portal and carry neither.
37
40
  - **One setup for every agent** — `prism connect` configures Claude Code,
38
41
  Claude Desktop, Cursor, Gemini CLI, and Codex while preserving unrelated
39
42
  settings.
@@ -1034,8 +1037,10 @@ The free tier runs entirely on your machine. Paid tiers add cloud sync through t
1034
1037
  redaction (SSNs, dates of birth, medical record numbers, phone numbers, emails,
1035
1038
  and clinical identifiers are stripped before storage). Cloud inference and
1036
1039
  route correction send the request over TLS for processing and do not store it
1037
- as Prism memory; use `route_guard: "local"` or the **local tier** for a full
1038
- air-gap. **Enterprise** includes a HIPAA Business Associate Agreement.
1040
+ as Prism memory. The **local tier** (no Synalux key) is the air-gap;
1041
+ `route_guard: "local"` only skips the route correction, and `cloud_fallback`
1042
+ and `verify` govern the other two channels. **Enterprise** includes a HIPAA
1043
+ Business Associate Agreement.
1039
1044
 
1040
1045
  ---
1041
1046
 
@@ -1050,8 +1055,12 @@ before it reaches the host. Malformed or unadvertised calls become `NO_TOOL`.
1050
1055
  With `route_guard: "auto"` (the default), Standard and higher plans also send
1051
1056
  a well-formed draft for one of Prism's seven trained tools—or an unadvertised
1052
1057
  draft that may need correction—to Synalux for authenticated deterministic
1053
- correction. Advertised custom host tools remain local. Set
1054
- `route_guard: "local"` for a fully on-device route path.
1058
+ correction. Advertised custom host tools remain local. `route_guard: "local"`
1059
+ disables that correction only: cloud inference fallback is governed by
1060
+ `cloud_fallback`, and the grounding verifier is a separate channel with its
1061
+ own switch (`verify`, on by default when `evidence` is given). With all three
1062
+ off, no request carries your prompt, draft or evidence; the entitlement check
1063
+ and telemetry still contact the portal and carry neither.
1055
1064
 
1056
1065
  | Model | Ollama tag | Size | Vision | Routing accuracy¹ | Role | Automatic routing tier |
1057
1066
  |---|---|---|---|---|---|---|
@@ -1152,12 +1161,19 @@ check, the local 9B passed 2/3 tasks; the local 27B and Gemini 3.6 Flash each
1152
1161
  passed 3/3. This is a self-published regression signal, not an independent
1153
1162
  leaderboard or a claim of broad model equivalence.
1154
1163
 
1155
- ### Cloud Escalation (`cloud_fallback: true`)
1164
+ ### Cloud Escalation (`cloud_fallback`)
1156
1165
 
1157
1166
  Prism always tries an eligible local model first. If the quality gate detects
1158
- an empty, truncated, think-only, or looping response, paid tiers can retry the
1159
- request through Gemini 3.6 Flash. Free-tier routing stays local and reports the
1160
- quality-gate outcome without making a cloud call.
1167
+ an empty, truncated, think-only, or looping response, or the safety screen
1168
+ finds the request uncertain or reserved, a paid plan escalates the request
1169
+ through Gemini 3.6 Flash. Free-tier routing stays local and reports the
1170
+ outcome without making a cloud call.
1171
+
1172
+ The flag follows the plan. Leave it unset and your plan decides: paid plans
1173
+ escalate, free plans never do. Pass `false` to forbid cloud inference fallback
1174
+ for a call, which the clinical delegation rules do for drafting; pass `true`
1175
+ to ask for it, which still requires a plan with cloud. A request that carries
1176
+ an image is never escalated; screenshots stay on this device.
1161
1177
 
1162
1178
  ---
1163
1179
 
@@ -1361,8 +1377,8 @@ prism_infer({
1361
1377
  })
1362
1378
  // → "n = countActiveUsers(data)" (local 9b, $0)
1363
1379
 
1364
- // A turn the on-device screen finds uncertain when read alone is not served
1365
- // locally: it goes to Synalux cloud on a paid plan, or is refused with
1380
+ // A turn the on-device screen finds uncertain, alone or in context, is not
1381
+ // served locally: it goes to Synalux cloud on a paid plan, or is refused with
1366
1382
  // cloud_fallback: false. The result names the reason (layer1_uncertain) so
1367
1383
  // the host can decide what to do with the thread.
1368
1384
  prism_infer({
@@ -9,7 +9,8 @@ export const LOCAL_FIRST_POLICY_LINES = [
9
9
  "## Prism local-first orchestration",
10
10
  "For bounded, verifiable delegated work, use Prism's local worker before any host-native or background subagent.",
11
11
  "When session_task_route is available, call it first. If it returns target `claw`, call `prism_infer` with the",
12
- "recommended arguments, `cloud_fallback: false`, and the `project` and `conversation_id` from this session when known.",
12
+ "recommended arguments and the `project` and `conversation_id` from this session when known. Leave `cloud_fallback`",
13
+ "unset so the plan decides it; pass `false` to forbid cloud inference fallback.",
13
14
  "If the route says `needs_history`, or the task follows up an earlier `prism_infer` answer, pass the accepted prior",
14
15
  "turns as `messages` (paid plans); never send a follow-up without them, and never re-send a turn you rejected.",
15
16
  "Prism loads that project's dashboard-configured quick, standard, or deep memory and selects a RAM-safe local model.",
@@ -42,8 +42,8 @@ const LEDGER_UNAVAILABLE_ERROR = "Inference metrics ledger is unavailable";
42
42
  const INSERT_METRIC_SQL = `INSERT OR IGNORE INTO infer_metrics
43
43
  (ts, caller, mode, backend, model, used_cloud, gate_outcome,
44
44
  refusal_reason, prompt_tokens, completion_tokens, latency_ms, ram_free_mb,
45
- source_event_id)
46
- VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)`;
45
+ source_event_id, history_turns, refusal_layer)
46
+ VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)`;
47
47
  function closeClient(context) {
48
48
  const activeClient = client;
49
49
  client = null;
@@ -83,17 +83,25 @@ function ensureTable() {
83
83
  completion_tokens INTEGER,
84
84
  latency_ms INTEGER,
85
85
  ram_free_mb INTEGER,
86
- source_event_id TEXT
86
+ source_event_id TEXT,
87
+ history_turns INTEGER,
88
+ refusal_layer TEXT
87
89
  )`);
88
90
  // Existing ledgers predate external panel-spool ingestion. SQLite has
89
91
  // no ADD COLUMN IF NOT EXISTS, so use the repository's established
90
92
  // idempotent migration pattern and reject only unexpected failures.
91
- try {
92
- await client.execute(`ALTER TABLE infer_metrics ADD COLUMN source_event_id TEXT`);
93
- }
94
- catch (e) {
95
- if (!(e instanceof Error) || !e.message.includes("duplicate column name"))
96
- throw e;
93
+ for (const column of [
94
+ "source_event_id TEXT",
95
+ "history_turns INTEGER",
96
+ "refusal_layer TEXT",
97
+ ]) {
98
+ try {
99
+ await client.execute(`ALTER TABLE infer_metrics ADD COLUMN ${column}`);
100
+ }
101
+ catch (e) {
102
+ if (!(e instanceof Error) || !e.message.includes("duplicate column name"))
103
+ throw e;
104
+ }
97
105
  }
98
106
  await client.execute(`CREATE INDEX IF NOT EXISTS idx_infer_metrics_ts ON infer_metrics (ts)`);
99
107
  await client.execute(`CREATE UNIQUE INDEX IF NOT EXISTS idx_infer_metrics_source_event
@@ -149,6 +157,7 @@ function metricArgs(row) {
149
157
  row.refusal_reason ?? null, row.prompt_tokens ?? null,
150
158
  row.completion_tokens ?? null, row.latency_ms ?? null,
151
159
  row.ram_free_mb ?? null, row.source_event_id ?? null,
160
+ row.history_turns ?? null, row.refusal_layer ?? null,
152
161
  ];
153
162
  }
154
163
  /** Aggregate all persisted rows (optionally since a timestamp). */
@@ -9,7 +9,7 @@
9
9
  * 1. Probe Ollama, list tags
10
10
  * 2. Pick largest viable local tier (pickLocalModel)
11
11
  * 3. Call /api/generate locally — return on success
12
- * 4. On local fail, if cloud_fallback=true:
12
+ * 4. On local fail, if the plan allows cloud (explicit cloud_fallback:false forbids it):
13
13
  * - exchange synalux_sk_ → JWT (cached)
14
14
  * - POST synalux portal /api/v1/prism/inference
15
15
  * - portal serves Gemini 3.6 Flash according to the user's tier
@@ -259,6 +259,11 @@ async function classifyHistoryWindow(l1fn, window, ollamaUrl, model, budget) {
259
259
  budget.tripped = true;
260
260
  return "UNCERTAIN";
261
261
  }
262
+ // deterministic:false is deliberate for a JOINT window and must stay so:
263
+ // the co-occurrence rules see words from different turns as one clause
264
+ // (the 2026-09-16 field refusal's joined window fires them; every turn
265
+ // alone is clean). The rules run per turn in proximity slices upstream;
266
+ // here only the model reads, and only to RAISE.
262
267
  const verdict = await l1fn(window, ollamaUrl, model, undefined, undefined, { deterministic: false });
263
268
  if (budget) {
264
269
  budget.consecutiveErrors = verdict === "ERROR" ? budget.consecutiveErrors + 1 : 0;
@@ -389,7 +394,7 @@ export const PRISM_INFER_TOOL = {
389
394
  "the caller's `task_complexity`, then validates loaded memory size, model context, " +
390
395
  "entitlements, installed models, and free RAM at call time. " +
391
396
  "Falls through to the Synalux portal Gemini 3.6 Flash cloud fallback " +
392
- "only when local is unviable AND `cloud_fallback=true`. " +
397
+ "only when local is unviable or refused and the plan allows cloud; `cloud_fallback: false` forbids it. " +
393
398
  "When `project` is provided, loads the dashboard-configured quick/standard/deep handoff and bounded history " +
394
399
  "as untrusted historical context for a memory-aware local worker. " +
395
400
  "Use this for code generation, summarisation, classification, or any synth task you would " +
@@ -474,8 +479,7 @@ export const PRISM_INFER_TOOL = {
474
479
  },
475
480
  cloud_fallback: {
476
481
  type: "boolean",
477
- description: "Fall through to the Synalux portal cascade on local failure. Default false: saving tokens is the point.",
478
- default: false,
482
+ description: "Synalux portal cascade when local is unviable or refused. Omitted: the plan decides; false forbids it.",
479
483
  },
480
484
  timeout_ms: {
481
485
  type: "number",
@@ -529,7 +533,7 @@ export const PRISM_INFER_TOOL = {
529
533
  type: "string",
530
534
  enum: ["auto", "local"],
531
535
  description: "'auto' (default): local advertised-tool contract plus, on paid plans, the private " +
532
- "Synalux deterministic route correction. 'local': prompt and draft stay on-device.",
536
+ "Synalux deterministic route correction. 'local': skips that correction only.",
533
537
  default: "auto",
534
538
  },
535
539
  think: {
@@ -1051,8 +1055,9 @@ export class ReservedRefusalError extends Error {
1051
1055
  const what = category ? `category="${category}"` : "matched the semantic classifier";
1052
1056
  const remedy = cloudWasAllowed
1053
1057
  ? "Cloud escalation was permitted and did not produce an answer; see attempts."
1054
- : "Reserved content is never answered by a local model. Pass cloud_fallback: true "
1055
- + "to escalate to a stronger model, or answer it in the host thread instead.";
1058
+ : "Reserved content is never answered by a local model. This call had no cloud: either it "
1059
+ + "passed cloud_fallback: false, or the plan has none. Pass cloud_fallback: true (or omit "
1060
+ + "it, on a paid plan) to escalate to a stronger model, or answer it in the host thread instead.";
1056
1061
  super(`prism_infer: Layer 1 verdict=${verdict}, ${what} — reserved content refused. `
1057
1062
  + `${remedy} attempts=${JSON.stringify(attempts)}`);
1058
1063
  this.attempts = attempts;
@@ -1060,7 +1065,7 @@ export class ReservedRefusalError extends Error {
1060
1065
  this.name = "ReservedRefusalError";
1061
1066
  }
1062
1067
  }
1063
- function makeReservedRefusal(verdict, attempts, category = null, cloudWasAllowed = false) {
1068
+ function makeReservedRefusal(verdict, attempts, category = null, cloudWasAllowed = false, ledger = {}) {
1064
1069
  // Ledger the refusal (fire-and-forget). No prompt content is persisted —
1065
1070
  // same HIPAA posture as the safety_gate exclusion. gate_outcome mirrors
1066
1071
  // the §5.2 report-mode row so refusal queries see both modes.
@@ -1068,6 +1073,8 @@ function makeReservedRefusal(verdict, attempts, category = null, cloudWasAllowed
1068
1073
  backend: "refused", model: null, used_cloud: false,
1069
1074
  gate_outcome: "refused",
1070
1075
  refusal_reason: "layer1_reserved",
1076
+ history_turns: ledger.history_turns,
1077
+ refusal_layer: ledger.refusal_layer,
1071
1078
  });
1072
1079
  return new ReservedRefusalError(verdict, attempts, category, cloudWasAllowed);
1073
1080
  }
@@ -1400,10 +1407,25 @@ export async function runInfer(args, deps) {
1400
1407
  // Retained for the log line, which describes the request rather than a
1401
1408
  // specific backend.
1402
1409
  const maxTokens = cloudMaxTokens;
1403
- // Cloud fallback only for paid plans
1410
+ // Cloud fallback is the PLAN's to give. An omitted flag means "whatever my
1411
+ // plan entitles me to": paid plans escalate, free plans do not. Explicit
1412
+ // false still forbids cloud INFERENCE fallback — the clinical delegation
1413
+ // rules and token-saving callers depend on that; the route guard and the
1414
+ // grounding verifier keep their own switches (route_guard, verify) — and
1415
+ // explicit true still needs a
1416
+ // plan with cloud. Until 2026-09-16 an omitted flag meant "no cloud", so a
1417
+ // paid, portal-ruled entitlement sat unused and an UNCERTAIN verdict
1418
+ // dead-ended instead of escalating; measured in production the day
1419
+ // multi-turn shipped, on a host that simply did not pass the argument.
1404
1420
  // let, not const: the reserved-image branch pins this off mid-call so no
1405
1421
  // later escalation path can carry even the prompt text off-device.
1406
- let allowCloud = args.cloud_fallback === true && ent.features.cloud_fallback;
1422
+ // The default is image-aware: cloud can never serve an image request
1423
+ // (screenshots stay on this device), so defaulting it ON for one would only
1424
+ // convert a gate-failed-but-usable local answer into a hard failure — an
1425
+ // explicit request still behaves as before and is refused at the point of
1426
+ // use with its attempt named.
1427
+ const planDefault = ent.features.cloud_fallback && !(args.images?.length);
1428
+ let allowCloud = (args.cloud_fallback ?? planDefault) && ent.features.cloud_fallback;
1407
1429
  // Verification only for paid plans (free users skip L3 grounding)
1408
1430
  const canVerify = ent.features.grounding_verifier;
1409
1431
  // The portal entitlement is authoritative. A paid plan alone must not
@@ -1438,6 +1460,20 @@ export async function runInfer(args, deps) {
1438
1460
  multi_turn: multiTurnPolicy(ent),
1439
1461
  history_turns: args.messages?.length ?? 0,
1440
1462
  };
1463
+ // Which screen layer decided the call; ledgered on a refusal. Bookkeeping
1464
+ // only: raise() is worseLayer1Verdict with a label and the assignment stays
1465
+ // at the call site, so it changes no outcome. Declared here because
1466
+ // refusedResult() can run before the Layer 1 block. Without it, "which
1467
+ // layer refused this" needs a transcript replay — exactly what a benign
1468
+ // production refusal cost on 2026-09-16.
1469
+ let l1Layer = null;
1470
+ const ledgerMeta = () => ({ history_turns: args.messages?.length ?? 0, refusal_layer: l1Layer ?? undefined });
1471
+ const raise = (cur, next, source) => {
1472
+ const merged = worseLayer1Verdict(cur, next);
1473
+ if (merged !== cur)
1474
+ l1Layer = source;
1475
+ return merged;
1476
+ };
1441
1477
  const refusedResult = (reason) => ({
1442
1478
  output: "",
1443
1479
  backend: "refused",
@@ -1448,6 +1484,7 @@ export async function runInfer(args, deps) {
1448
1484
  attempts,
1449
1485
  ...entMeta,
1450
1486
  gate_outcome: { status: "refused", reason, served_anyway: false },
1487
+ refusal_layer: l1Layer ?? undefined,
1451
1488
  });
1452
1489
  debugLog(`[prism_infer] plan=${ent.plan} ceiling=${effectiveCeiling} max_tokens=${maxTokens} ` +
1453
1490
  `cloud=${allowCloud} verify=${canVerify} route_guard=${canUsePrivateRouteGuard}`);
@@ -1593,6 +1630,8 @@ export async function runInfer(args, deps) {
1593
1630
  if (!args.messages?.length) {
1594
1631
  // Single turn: the exact call it always was.
1595
1632
  l1 = await l1fn(args.prompt, deps.ollamaUrl, l1Model, undefined, resolvedImages);
1633
+ if (l1 !== "OBVIOUS_NOT_RESERVED")
1634
+ l1Layer = "prompt";
1596
1635
  }
1597
1636
  else {
1598
1637
  l1 = "OBVIOUS_NOT_RESERVED";
@@ -1621,7 +1660,7 @@ export async function runInfer(args, deps) {
1621
1660
  for (const slice of windowsOf(turn.content, DETERMINISTIC_FLOOR_WINDOW_CHARS, DETERMINISTIC_FLOOR_WINDOW_OVERLAP)) {
1622
1661
  const det = classifyDeterministicLayer1(slice, { operational: isUser });
1623
1662
  if (det)
1624
- l1 = worseLayer1Verdict(l1, det);
1663
+ l1 = raise(l1, det, "rules");
1625
1664
  }
1626
1665
  }
1627
1666
  // 2. Semantic floor, per TURN in isolation; every verdict read
@@ -1651,10 +1690,10 @@ export async function runInfer(args, deps) {
1651
1690
  continue;
1652
1691
  const alone = await classifyHistoryWindow(l1fn, window, deps.ollamaUrl, l1Model, budget);
1653
1692
  if (alone === "OBVIOUS_RESERVED") {
1654
- l1 = "OBVIOUS_RESERVED";
1693
+ l1 = raise(l1, "OBVIOUS_RESERVED", "isolated");
1655
1694
  break history;
1656
1695
  }
1657
- l1 = worseLayer1Verdict(l1, alone);
1696
+ l1 = raise(l1, alone, "isolated");
1658
1697
  }
1659
1698
  }
1660
1699
  // The current prompt is a request: its deterministic floor runs
@@ -1667,7 +1706,7 @@ export async function runInfer(args, deps) {
1667
1706
  for (const slice of windowsOf(args.prompt, DETERMINISTIC_FLOOR_WINDOW_CHARS, DETERMINISTIC_FLOOR_WINDOW_OVERLAP)) {
1668
1707
  const promptDet = classifyDeterministicLayer1(slice);
1669
1708
  if (promptDet)
1670
- l1 = worseLayer1Verdict(l1, promptDet);
1709
+ l1 = raise(l1, promptDet, "rules");
1671
1710
  if (promptDet !== "OBVIOUS_NOT_RESERVED")
1672
1711
  promptRoutine = false;
1673
1712
  }
@@ -1681,7 +1720,7 @@ export async function runInfer(args, deps) {
1681
1720
  // (review round 19: skipping it there bypassed that floor).
1682
1721
  const promptFastPath = promptRoutine && args.prompt.length <= MAX_CLASSIFIER_PROMPT_LENGTH && (resolvedImages?.length ?? 0) === 0;
1683
1722
  if (l1 !== "OBVIOUS_RESERVED" && !promptFastPath) {
1684
- l1 = worseLayer1Verdict(l1, await l1fn(args.prompt, deps.ollamaUrl, l1Model, undefined, resolvedImages, { deterministic: false }));
1723
+ l1 = raise(l1, await l1fn(args.prompt, deps.ollamaUrl, l1Model, undefined, resolvedImages, { deterministic: false }), "prompt");
1685
1724
  }
1686
1725
  // 3. Context, raise only: one window per turn and one for the
1687
1726
  // prompt (see contextWindows), cached like any window. Skipped
@@ -1697,7 +1736,7 @@ export async function runInfer(args, deps) {
1697
1736
  for (const window of contextWindows(args)) {
1698
1737
  if (!window.trim())
1699
1738
  continue;
1700
- l1 = worseLayer1Verdict(l1, await classifyHistoryWindow(l1fn, window, deps.ollamaUrl, l1Model, budget));
1739
+ l1 = raise(l1, await classifyHistoryWindow(l1fn, window, deps.ollamaUrl, l1Model, budget), "context");
1701
1740
  if (l1 === "UNCERTAIN" || l1 === "OBVIOUS_RESERVED")
1702
1741
  break;
1703
1742
  }
@@ -1705,7 +1744,7 @@ export async function runInfer(args, deps) {
1705
1744
  // A budget or breaker trip raises to UNCERTAIN whatever the cache
1706
1745
  // held (text: cloud or refused; with an image: local only).
1707
1746
  if (budget.tripped)
1708
- l1 = worseLayer1Verdict(l1, "UNCERTAIN");
1747
+ l1 = raise(l1, "UNCERTAIN", "budget");
1709
1748
  if (budget.calls > LAYER1_SCREEN_CALL_BUDGET) {
1710
1749
  attempts.push({ tier: "layer1", reason: `layer1_screen_over_budget:${LAYER1_SCREEN_CALL_BUDGET}` });
1711
1750
  }
@@ -1753,7 +1792,7 @@ export async function runInfer(args, deps) {
1753
1792
  attempts.push({ tier: "synalux", reason: "reserved_escalation_refused_images_stay_local" });
1754
1793
  if (wantReport)
1755
1794
  return refusedResult("layer1_reserved");
1756
- throw makeReservedRefusal(l1, attempts, reservedCat, true);
1795
+ throw makeReservedRefusal(l1, attempts, reservedCat, true, ledgerMeta());
1757
1796
  }
1758
1797
  if (allowCloud) {
1759
1798
  const cloudTimeout = args.timeout_ms ?? 90_000;
@@ -1769,7 +1808,7 @@ export async function runInfer(args, deps) {
1769
1808
  attempts.push({ tier: "synalux", reason: `reserved_weak_backend:${cloud.backend}` });
1770
1809
  if (wantReport)
1771
1810
  return refusedResult("layer1_reserved");
1772
- throw makeReservedRefusal(l1, attempts, reservedCat, true);
1811
+ throw makeReservedRefusal(l1, attempts, reservedCat, true, ledgerMeta());
1773
1812
  }
1774
1813
  return await applyVerification(cloud.output, gatedArgs, deps, {
1775
1814
  backend: cloud.backend ?? "synalux",
@@ -1787,7 +1826,7 @@ export async function runInfer(args, deps) {
1787
1826
  }
1788
1827
  if (wantReport)
1789
1828
  return refusedResult("layer1_reserved");
1790
- throw makeReservedRefusal(l1, attempts, reservedCat, allowCloud);
1829
+ throw makeReservedRefusal(l1, attempts, reservedCat, allowCloud, ledgerMeta());
1791
1830
  }
1792
1831
  if (l1 === "UNCERTAIN_LENGTH") {
1793
1832
  // §5.3: prompt too long to classify in full, but the full-text
@@ -1818,7 +1857,7 @@ export async function runInfer(args, deps) {
1818
1857
  attempts.push({ tier: "synalux", reason: "error_escalation_refused_images_stay_local" });
1819
1858
  if (wantReport)
1820
1859
  return refusedResult("layer1_error");
1821
- throw makeReservedRefusal(l1, attempts);
1860
+ throw makeReservedRefusal(l1, attempts, null, false, ledgerMeta());
1822
1861
  }
1823
1862
  if (allowCloud) {
1824
1863
  const cloudTimeout = args.timeout_ms ?? 90_000;
@@ -1842,6 +1881,7 @@ export async function runInfer(args, deps) {
1842
1881
  debugLog(`[prism_infer] keyword backstop verdict=${backstop}`);
1843
1882
  attempts.push({ tier: "keyword_backstop", reason: `backstop_${backstop.toLowerCase()}` });
1844
1883
  if (backstop === "OBVIOUS_RESERVED") {
1884
+ l1Layer = "backstop"; // the regex net refused, whatever raised the verdict before it
1845
1885
  if (wantReport)
1846
1886
  return refusedResult("keyword_backstop_reserved");
1847
1887
  // Serve-mode backstop refusal previously wrote NO ledger row —
@@ -1850,6 +1890,8 @@ export async function runInfer(args, deps) {
1850
1890
  backend: "refused", model: null, used_cloud: false,
1851
1891
  gate_outcome: "refused",
1852
1892
  refusal_reason: "keyword_backstop_reserved",
1893
+ history_turns: args.messages?.length ?? 0,
1894
+ refusal_layer: "backstop",
1853
1895
  });
1854
1896
  throw new Error(`prism_infer: classifier failed + keyword backstop caught reserved content. attempts=${JSON.stringify(attempts)}`);
1855
1897
  }
@@ -347,7 +347,10 @@ function buildRecommendedArgs(args, complexityScore) {
347
347
  ...(args.project ? { project: args.project } : {}),
348
348
  mode: "code",
349
349
  task_complexity: complexityScore,
350
- cloud_fallback: false,
350
+ // cloud_fallback is deliberately absent: prism_infer resolves it from the
351
+ // plan's entitlements, which the router does not read. Pinning it false
352
+ // here made a paid plan's escalation unreachable for any host that copied
353
+ // these arguments verbatim (2026-09-16).
351
354
  escalation: "report",
352
355
  };
353
356
  }
@@ -139,6 +139,8 @@ export function recordInference(result) {
139
139
  completion_tokens: result.completion_tokens,
140
140
  latency_ms: result.latency_ms,
141
141
  ram_free_mb: result.ram_free_mb,
142
+ history_turns: result.history_turns,
143
+ refusal_layer: result.refusal_layer,
142
144
  });
143
145
  // §5.2: refused results (escalation:"report") get a ledger row above but
144
146
  // must NOT touch the session accumulators — no model ran and nothing was
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "prism-mcp-server",
3
- "version": "20.20.0",
3
+ "version": "20.21.0",
4
4
  "mcpName": "io.github.dcostenco/prism-coder",
5
5
  "description": "Persistent session memory for AI coding agents that never leaves your machine — including the on-device model that reasons over it. Restores your prior decisions, open TODOs, and changed files across sessions; adds associative recall of related past work, semantic drift detection, and local inference. Local-first by default. Works with Claude Code, Cursor, and Codex.",
6
6
  "module": "index.ts",