prism-mcp-server 20.20.0 → 20.21.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md
CHANGED
|
@@ -32,8 +32,11 @@ A paid subscription adds cloud sync, higher model tiers, and team features throu
|
|
|
32
32
|
undercounts inline — a number you can check, not marketing.
|
|
33
33
|
- **Route-output enforcement** — route mode returns only well-formed calls to
|
|
34
34
|
tools the host actually advertised. Standard and higher plans can add
|
|
35
|
-
authenticated deterministic correction; `route_guard: "local"`
|
|
36
|
-
|
|
35
|
+
authenticated deterministic correction; `route_guard: "local"` disables that
|
|
36
|
+
correction only. `cloud_fallback: false` forbids cloud inference fallback and
|
|
37
|
+
`verify: false` (with no `evidence`) disables the grounding verifier. With all
|
|
38
|
+
three off, no request carries your prompt, draft or evidence; the per-call
|
|
39
|
+
entitlement check and telemetry still contact the portal and carry neither.
|
|
37
40
|
- **One setup for every agent** — `prism connect` configures Claude Code,
|
|
38
41
|
Claude Desktop, Cursor, Gemini CLI, and Codex while preserving unrelated
|
|
39
42
|
settings.
|
|
@@ -1034,8 +1037,10 @@ The free tier runs entirely on your machine. Paid tiers add cloud sync through t
|
|
|
1034
1037
|
redaction (SSNs, dates of birth, medical record numbers, phone numbers, emails,
|
|
1035
1038
|
and clinical identifiers are stripped before storage). Cloud inference and
|
|
1036
1039
|
route correction send the request over TLS for processing and do not store it
|
|
1037
|
-
as Prism memory
|
|
1038
|
-
|
|
1040
|
+
as Prism memory. The **local tier** (no Synalux key) is the air-gap;
|
|
1041
|
+
`route_guard: "local"` only skips the route correction, and `cloud_fallback`
|
|
1042
|
+
and `verify` govern the other two channels. **Enterprise** includes a HIPAA
|
|
1043
|
+
Business Associate Agreement.
|
|
1039
1044
|
|
|
1040
1045
|
---
|
|
1041
1046
|
|
|
@@ -1050,8 +1055,12 @@ before it reaches the host. Malformed or unadvertised calls become `NO_TOOL`.
|
|
|
1050
1055
|
With `route_guard: "auto"` (the default), Standard and higher plans also send
|
|
1051
1056
|
a well-formed draft for one of Prism's seven trained tools—or an unadvertised
|
|
1052
1057
|
draft that may need correction—to Synalux for authenticated deterministic
|
|
1053
|
-
correction. Advertised custom host tools remain local.
|
|
1054
|
-
|
|
1058
|
+
correction. Advertised custom host tools remain local. `route_guard: "local"`
|
|
1059
|
+
disables that correction only: cloud inference fallback is governed by
|
|
1060
|
+
`cloud_fallback`, and the grounding verifier is a separate channel with its
|
|
1061
|
+
own switch (`verify`, on by default when `evidence` is given). With all three
|
|
1062
|
+
off, no request carries your prompt, draft or evidence; the entitlement check
|
|
1063
|
+
and telemetry still contact the portal and carry neither.
|
|
1055
1064
|
|
|
1056
1065
|
| Model | Ollama tag | Size | Vision | Routing accuracy¹ | Role | Automatic routing tier |
|
|
1057
1066
|
|---|---|---|---|---|---|---|
|
|
@@ -1152,12 +1161,19 @@ check, the local 9B passed 2/3 tasks; the local 27B and Gemini 3.6 Flash each
|
|
|
1152
1161
|
passed 3/3. This is a self-published regression signal, not an independent
|
|
1153
1162
|
leaderboard or a claim of broad model equivalence.
|
|
1154
1163
|
|
|
1155
|
-
### Cloud Escalation (`cloud_fallback
|
|
1164
|
+
### Cloud Escalation (`cloud_fallback`)
|
|
1156
1165
|
|
|
1157
1166
|
Prism always tries an eligible local model first. If the quality gate detects
|
|
1158
|
-
an empty, truncated, think-only, or looping response,
|
|
1159
|
-
|
|
1160
|
-
|
|
1167
|
+
an empty, truncated, think-only, or looping response, or the safety screen
|
|
1168
|
+
finds the request uncertain or reserved, a paid plan escalates the request
|
|
1169
|
+
through Gemini 3.6 Flash. Free-tier routing stays local and reports the
|
|
1170
|
+
outcome without making a cloud call.
|
|
1171
|
+
|
|
1172
|
+
The flag follows the plan. Leave it unset and your plan decides: paid plans
|
|
1173
|
+
escalate, free plans never do. Pass `false` to forbid cloud inference fallback
|
|
1174
|
+
for a call, which the clinical delegation rules do for drafting; pass `true`
|
|
1175
|
+
to ask for it, which still requires a plan with cloud. A request that carries
|
|
1176
|
+
an image is never escalated; screenshots stay on this device.
|
|
1161
1177
|
|
|
1162
1178
|
---
|
|
1163
1179
|
|
|
@@ -1361,8 +1377,8 @@ prism_infer({
|
|
|
1361
1377
|
})
|
|
1362
1378
|
// → "n = countActiveUsers(data)" (local 9b, $0)
|
|
1363
1379
|
|
|
1364
|
-
// A turn the on-device screen finds uncertain
|
|
1365
|
-
// locally: it goes to Synalux cloud on a paid plan, or is refused with
|
|
1380
|
+
// A turn the on-device screen finds uncertain, alone or in context, is not
|
|
1381
|
+
// served locally: it goes to Synalux cloud on a paid plan, or is refused with
|
|
1366
1382
|
// cloud_fallback: false. The result names the reason (layer1_uncertain) so
|
|
1367
1383
|
// the host can decide what to do with the thread.
|
|
1368
1384
|
prism_infer({
|
package/dist/localFirstPolicy.js
CHANGED
|
@@ -9,7 +9,8 @@ export const LOCAL_FIRST_POLICY_LINES = [
|
|
|
9
9
|
"## Prism local-first orchestration",
|
|
10
10
|
"For bounded, verifiable delegated work, use Prism's local worker before any host-native or background subagent.",
|
|
11
11
|
"When session_task_route is available, call it first. If it returns target `claw`, call `prism_infer` with the",
|
|
12
|
-
"recommended arguments
|
|
12
|
+
"recommended arguments and the `project` and `conversation_id` from this session when known. Leave `cloud_fallback`",
|
|
13
|
+
"unset so the plan decides it; pass `false` to forbid cloud inference fallback.",
|
|
13
14
|
"If the route says `needs_history`, or the task follows up an earlier `prism_infer` answer, pass the accepted prior",
|
|
14
15
|
"turns as `messages` (paid plans); never send a follow-up without them, and never re-send a turn you rejected.",
|
|
15
16
|
"Prism loads that project's dashboard-configured quick, standard, or deep memory and selects a RAM-safe local model.",
|
|
@@ -42,8 +42,8 @@ const LEDGER_UNAVAILABLE_ERROR = "Inference metrics ledger is unavailable";
|
|
|
42
42
|
const INSERT_METRIC_SQL = `INSERT OR IGNORE INTO infer_metrics
|
|
43
43
|
(ts, caller, mode, backend, model, used_cloud, gate_outcome,
|
|
44
44
|
refusal_reason, prompt_tokens, completion_tokens, latency_ms, ram_free_mb,
|
|
45
|
-
source_event_id)
|
|
46
|
-
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)`;
|
|
45
|
+
source_event_id, history_turns, refusal_layer)
|
|
46
|
+
VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)`;
|
|
47
47
|
function closeClient(context) {
|
|
48
48
|
const activeClient = client;
|
|
49
49
|
client = null;
|
|
@@ -83,17 +83,25 @@ function ensureTable() {
|
|
|
83
83
|
completion_tokens INTEGER,
|
|
84
84
|
latency_ms INTEGER,
|
|
85
85
|
ram_free_mb INTEGER,
|
|
86
|
-
source_event_id TEXT
|
|
86
|
+
source_event_id TEXT,
|
|
87
|
+
history_turns INTEGER,
|
|
88
|
+
refusal_layer TEXT
|
|
87
89
|
)`);
|
|
88
90
|
// Existing ledgers predate external panel-spool ingestion. SQLite has
|
|
89
91
|
// no ADD COLUMN IF NOT EXISTS, so use the repository's established
|
|
90
92
|
// idempotent migration pattern and reject only unexpected failures.
|
|
91
|
-
|
|
92
|
-
|
|
93
|
-
|
|
94
|
-
|
|
95
|
-
|
|
96
|
-
|
|
93
|
+
for (const column of [
|
|
94
|
+
"source_event_id TEXT",
|
|
95
|
+
"history_turns INTEGER",
|
|
96
|
+
"refusal_layer TEXT",
|
|
97
|
+
]) {
|
|
98
|
+
try {
|
|
99
|
+
await client.execute(`ALTER TABLE infer_metrics ADD COLUMN ${column}`);
|
|
100
|
+
}
|
|
101
|
+
catch (e) {
|
|
102
|
+
if (!(e instanceof Error) || !e.message.includes("duplicate column name"))
|
|
103
|
+
throw e;
|
|
104
|
+
}
|
|
97
105
|
}
|
|
98
106
|
await client.execute(`CREATE INDEX IF NOT EXISTS idx_infer_metrics_ts ON infer_metrics (ts)`);
|
|
99
107
|
await client.execute(`CREATE UNIQUE INDEX IF NOT EXISTS idx_infer_metrics_source_event
|
|
@@ -149,6 +157,7 @@ function metricArgs(row) {
|
|
|
149
157
|
row.refusal_reason ?? null, row.prompt_tokens ?? null,
|
|
150
158
|
row.completion_tokens ?? null, row.latency_ms ?? null,
|
|
151
159
|
row.ram_free_mb ?? null, row.source_event_id ?? null,
|
|
160
|
+
row.history_turns ?? null, row.refusal_layer ?? null,
|
|
152
161
|
];
|
|
153
162
|
}
|
|
154
163
|
/** Aggregate all persisted rows (optionally since a timestamp). */
|
|
@@ -9,7 +9,7 @@
|
|
|
9
9
|
* 1. Probe Ollama, list tags
|
|
10
10
|
* 2. Pick largest viable local tier (pickLocalModel)
|
|
11
11
|
* 3. Call /api/generate locally — return on success
|
|
12
|
-
* 4. On local fail, if cloud_fallback
|
|
12
|
+
* 4. On local fail, if the plan allows cloud (explicit cloud_fallback:false forbids it):
|
|
13
13
|
* - exchange synalux_sk_ → JWT (cached)
|
|
14
14
|
* - POST synalux portal /api/v1/prism/inference
|
|
15
15
|
* - portal serves Gemini 3.6 Flash according to the user's tier
|
|
@@ -259,6 +259,11 @@ async function classifyHistoryWindow(l1fn, window, ollamaUrl, model, budget) {
|
|
|
259
259
|
budget.tripped = true;
|
|
260
260
|
return "UNCERTAIN";
|
|
261
261
|
}
|
|
262
|
+
// deterministic:false is deliberate for a JOINT window and must stay so:
|
|
263
|
+
// the co-occurrence rules see words from different turns as one clause
|
|
264
|
+
// (the 2026-09-16 field refusal's joined window fires them; every turn
|
|
265
|
+
// alone is clean). The rules run per turn in proximity slices upstream;
|
|
266
|
+
// here only the model reads, and only to RAISE.
|
|
262
267
|
const verdict = await l1fn(window, ollamaUrl, model, undefined, undefined, { deterministic: false });
|
|
263
268
|
if (budget) {
|
|
264
269
|
budget.consecutiveErrors = verdict === "ERROR" ? budget.consecutiveErrors + 1 : 0;
|
|
@@ -389,7 +394,7 @@ export const PRISM_INFER_TOOL = {
|
|
|
389
394
|
"the caller's `task_complexity`, then validates loaded memory size, model context, " +
|
|
390
395
|
"entitlements, installed models, and free RAM at call time. " +
|
|
391
396
|
"Falls through to the Synalux portal Gemini 3.6 Flash cloud fallback " +
|
|
392
|
-
"only when local is unviable
|
|
397
|
+
"only when local is unviable or refused and the plan allows cloud; `cloud_fallback: false` forbids it. " +
|
|
393
398
|
"When `project` is provided, loads the dashboard-configured quick/standard/deep handoff and bounded history " +
|
|
394
399
|
"as untrusted historical context for a memory-aware local worker. " +
|
|
395
400
|
"Use this for code generation, summarisation, classification, or any synth task you would " +
|
|
@@ -474,8 +479,7 @@ export const PRISM_INFER_TOOL = {
|
|
|
474
479
|
},
|
|
475
480
|
cloud_fallback: {
|
|
476
481
|
type: "boolean",
|
|
477
|
-
description: "
|
|
478
|
-
default: false,
|
|
482
|
+
description: "Synalux portal cascade when local is unviable or refused. Omitted: the plan decides; false forbids it.",
|
|
479
483
|
},
|
|
480
484
|
timeout_ms: {
|
|
481
485
|
type: "number",
|
|
@@ -529,7 +533,7 @@ export const PRISM_INFER_TOOL = {
|
|
|
529
533
|
type: "string",
|
|
530
534
|
enum: ["auto", "local"],
|
|
531
535
|
description: "'auto' (default): local advertised-tool contract plus, on paid plans, the private " +
|
|
532
|
-
"Synalux deterministic route correction. 'local':
|
|
536
|
+
"Synalux deterministic route correction. 'local': skips that correction only.",
|
|
533
537
|
default: "auto",
|
|
534
538
|
},
|
|
535
539
|
think: {
|
|
@@ -1051,8 +1055,9 @@ export class ReservedRefusalError extends Error {
|
|
|
1051
1055
|
const what = category ? `category="${category}"` : "matched the semantic classifier";
|
|
1052
1056
|
const remedy = cloudWasAllowed
|
|
1053
1057
|
? "Cloud escalation was permitted and did not produce an answer; see attempts."
|
|
1054
|
-
: "Reserved content is never answered by a local model.
|
|
1055
|
-
+ "
|
|
1058
|
+
: "Reserved content is never answered by a local model. This call had no cloud: either it "
|
|
1059
|
+
+ "passed cloud_fallback: false, or the plan has none. Pass cloud_fallback: true (or omit "
|
|
1060
|
+
+ "it, on a paid plan) to escalate to a stronger model, or answer it in the host thread instead.";
|
|
1056
1061
|
super(`prism_infer: Layer 1 verdict=${verdict}, ${what} — reserved content refused. `
|
|
1057
1062
|
+ `${remedy} attempts=${JSON.stringify(attempts)}`);
|
|
1058
1063
|
this.attempts = attempts;
|
|
@@ -1060,7 +1065,7 @@ export class ReservedRefusalError extends Error {
|
|
|
1060
1065
|
this.name = "ReservedRefusalError";
|
|
1061
1066
|
}
|
|
1062
1067
|
}
|
|
1063
|
-
function makeReservedRefusal(verdict, attempts, category = null, cloudWasAllowed = false) {
|
|
1068
|
+
function makeReservedRefusal(verdict, attempts, category = null, cloudWasAllowed = false, ledger = {}) {
|
|
1064
1069
|
// Ledger the refusal (fire-and-forget). No prompt content is persisted —
|
|
1065
1070
|
// same HIPAA posture as the safety_gate exclusion. gate_outcome mirrors
|
|
1066
1071
|
// the §5.2 report-mode row so refusal queries see both modes.
|
|
@@ -1068,6 +1073,8 @@ function makeReservedRefusal(verdict, attempts, category = null, cloudWasAllowed
|
|
|
1068
1073
|
backend: "refused", model: null, used_cloud: false,
|
|
1069
1074
|
gate_outcome: "refused",
|
|
1070
1075
|
refusal_reason: "layer1_reserved",
|
|
1076
|
+
history_turns: ledger.history_turns,
|
|
1077
|
+
refusal_layer: ledger.refusal_layer,
|
|
1071
1078
|
});
|
|
1072
1079
|
return new ReservedRefusalError(verdict, attempts, category, cloudWasAllowed);
|
|
1073
1080
|
}
|
|
@@ -1400,10 +1407,25 @@ export async function runInfer(args, deps) {
|
|
|
1400
1407
|
// Retained for the log line, which describes the request rather than a
|
|
1401
1408
|
// specific backend.
|
|
1402
1409
|
const maxTokens = cloudMaxTokens;
|
|
1403
|
-
// Cloud fallback
|
|
1410
|
+
// Cloud fallback is the PLAN's to give. An omitted flag means "whatever my
|
|
1411
|
+
// plan entitles me to": paid plans escalate, free plans do not. Explicit
|
|
1412
|
+
// false still forbids cloud INFERENCE fallback — the clinical delegation
|
|
1413
|
+
// rules and token-saving callers depend on that; the route guard and the
|
|
1414
|
+
// grounding verifier keep their own switches (route_guard, verify) — and
|
|
1415
|
+
// explicit true still needs a
|
|
1416
|
+
// plan with cloud. Until 2026-09-16 an omitted flag meant "no cloud", so a
|
|
1417
|
+
// paid, portal-ruled entitlement sat unused and an UNCERTAIN verdict
|
|
1418
|
+
// dead-ended instead of escalating; measured in production the day
|
|
1419
|
+
// multi-turn shipped, on a host that simply did not pass the argument.
|
|
1404
1420
|
// let, not const: the reserved-image branch pins this off mid-call so no
|
|
1405
1421
|
// later escalation path can carry even the prompt text off-device.
|
|
1406
|
-
|
|
1422
|
+
// The default is image-aware: cloud can never serve an image request
|
|
1423
|
+
// (screenshots stay on this device), so defaulting it ON for one would only
|
|
1424
|
+
// convert a gate-failed-but-usable local answer into a hard failure — an
|
|
1425
|
+
// explicit request still behaves as before and is refused at the point of
|
|
1426
|
+
// use with its attempt named.
|
|
1427
|
+
const planDefault = ent.features.cloud_fallback && !(args.images?.length);
|
|
1428
|
+
let allowCloud = (args.cloud_fallback ?? planDefault) && ent.features.cloud_fallback;
|
|
1407
1429
|
// Verification only for paid plans (free users skip L3 grounding)
|
|
1408
1430
|
const canVerify = ent.features.grounding_verifier;
|
|
1409
1431
|
// The portal entitlement is authoritative. A paid plan alone must not
|
|
@@ -1438,6 +1460,20 @@ export async function runInfer(args, deps) {
|
|
|
1438
1460
|
multi_turn: multiTurnPolicy(ent),
|
|
1439
1461
|
history_turns: args.messages?.length ?? 0,
|
|
1440
1462
|
};
|
|
1463
|
+
// Which screen layer decided the call; ledgered on a refusal. Bookkeeping
|
|
1464
|
+
// only: raise() is worseLayer1Verdict with a label and the assignment stays
|
|
1465
|
+
// at the call site, so it changes no outcome. Declared here because
|
|
1466
|
+
// refusedResult() can run before the Layer 1 block. Without it, "which
|
|
1467
|
+
// layer refused this" needs a transcript replay — exactly what a benign
|
|
1468
|
+
// production refusal cost on 2026-09-16.
|
|
1469
|
+
let l1Layer = null;
|
|
1470
|
+
const ledgerMeta = () => ({ history_turns: args.messages?.length ?? 0, refusal_layer: l1Layer ?? undefined });
|
|
1471
|
+
const raise = (cur, next, source) => {
|
|
1472
|
+
const merged = worseLayer1Verdict(cur, next);
|
|
1473
|
+
if (merged !== cur)
|
|
1474
|
+
l1Layer = source;
|
|
1475
|
+
return merged;
|
|
1476
|
+
};
|
|
1441
1477
|
const refusedResult = (reason) => ({
|
|
1442
1478
|
output: "",
|
|
1443
1479
|
backend: "refused",
|
|
@@ -1448,6 +1484,7 @@ export async function runInfer(args, deps) {
|
|
|
1448
1484
|
attempts,
|
|
1449
1485
|
...entMeta,
|
|
1450
1486
|
gate_outcome: { status: "refused", reason, served_anyway: false },
|
|
1487
|
+
refusal_layer: l1Layer ?? undefined,
|
|
1451
1488
|
});
|
|
1452
1489
|
debugLog(`[prism_infer] plan=${ent.plan} ceiling=${effectiveCeiling} max_tokens=${maxTokens} ` +
|
|
1453
1490
|
`cloud=${allowCloud} verify=${canVerify} route_guard=${canUsePrivateRouteGuard}`);
|
|
@@ -1593,6 +1630,8 @@ export async function runInfer(args, deps) {
|
|
|
1593
1630
|
if (!args.messages?.length) {
|
|
1594
1631
|
// Single turn: the exact call it always was.
|
|
1595
1632
|
l1 = await l1fn(args.prompt, deps.ollamaUrl, l1Model, undefined, resolvedImages);
|
|
1633
|
+
if (l1 !== "OBVIOUS_NOT_RESERVED")
|
|
1634
|
+
l1Layer = "prompt";
|
|
1596
1635
|
}
|
|
1597
1636
|
else {
|
|
1598
1637
|
l1 = "OBVIOUS_NOT_RESERVED";
|
|
@@ -1621,7 +1660,7 @@ export async function runInfer(args, deps) {
|
|
|
1621
1660
|
for (const slice of windowsOf(turn.content, DETERMINISTIC_FLOOR_WINDOW_CHARS, DETERMINISTIC_FLOOR_WINDOW_OVERLAP)) {
|
|
1622
1661
|
const det = classifyDeterministicLayer1(slice, { operational: isUser });
|
|
1623
1662
|
if (det)
|
|
1624
|
-
l1 =
|
|
1663
|
+
l1 = raise(l1, det, "rules");
|
|
1625
1664
|
}
|
|
1626
1665
|
}
|
|
1627
1666
|
// 2. Semantic floor, per TURN in isolation; every verdict read
|
|
@@ -1651,10 +1690,10 @@ export async function runInfer(args, deps) {
|
|
|
1651
1690
|
continue;
|
|
1652
1691
|
const alone = await classifyHistoryWindow(l1fn, window, deps.ollamaUrl, l1Model, budget);
|
|
1653
1692
|
if (alone === "OBVIOUS_RESERVED") {
|
|
1654
|
-
l1 = "OBVIOUS_RESERVED";
|
|
1693
|
+
l1 = raise(l1, "OBVIOUS_RESERVED", "isolated");
|
|
1655
1694
|
break history;
|
|
1656
1695
|
}
|
|
1657
|
-
l1 =
|
|
1696
|
+
l1 = raise(l1, alone, "isolated");
|
|
1658
1697
|
}
|
|
1659
1698
|
}
|
|
1660
1699
|
// The current prompt is a request: its deterministic floor runs
|
|
@@ -1667,7 +1706,7 @@ export async function runInfer(args, deps) {
|
|
|
1667
1706
|
for (const slice of windowsOf(args.prompt, DETERMINISTIC_FLOOR_WINDOW_CHARS, DETERMINISTIC_FLOOR_WINDOW_OVERLAP)) {
|
|
1668
1707
|
const promptDet = classifyDeterministicLayer1(slice);
|
|
1669
1708
|
if (promptDet)
|
|
1670
|
-
l1 =
|
|
1709
|
+
l1 = raise(l1, promptDet, "rules");
|
|
1671
1710
|
if (promptDet !== "OBVIOUS_NOT_RESERVED")
|
|
1672
1711
|
promptRoutine = false;
|
|
1673
1712
|
}
|
|
@@ -1681,7 +1720,7 @@ export async function runInfer(args, deps) {
|
|
|
1681
1720
|
// (review round 19: skipping it there bypassed that floor).
|
|
1682
1721
|
const promptFastPath = promptRoutine && args.prompt.length <= MAX_CLASSIFIER_PROMPT_LENGTH && (resolvedImages?.length ?? 0) === 0;
|
|
1683
1722
|
if (l1 !== "OBVIOUS_RESERVED" && !promptFastPath) {
|
|
1684
|
-
l1 =
|
|
1723
|
+
l1 = raise(l1, await l1fn(args.prompt, deps.ollamaUrl, l1Model, undefined, resolvedImages, { deterministic: false }), "prompt");
|
|
1685
1724
|
}
|
|
1686
1725
|
// 3. Context, raise only: one window per turn and one for the
|
|
1687
1726
|
// prompt (see contextWindows), cached like any window. Skipped
|
|
@@ -1697,7 +1736,7 @@ export async function runInfer(args, deps) {
|
|
|
1697
1736
|
for (const window of contextWindows(args)) {
|
|
1698
1737
|
if (!window.trim())
|
|
1699
1738
|
continue;
|
|
1700
|
-
l1 =
|
|
1739
|
+
l1 = raise(l1, await classifyHistoryWindow(l1fn, window, deps.ollamaUrl, l1Model, budget), "context");
|
|
1701
1740
|
if (l1 === "UNCERTAIN" || l1 === "OBVIOUS_RESERVED")
|
|
1702
1741
|
break;
|
|
1703
1742
|
}
|
|
@@ -1705,7 +1744,7 @@ export async function runInfer(args, deps) {
|
|
|
1705
1744
|
// A budget or breaker trip raises to UNCERTAIN whatever the cache
|
|
1706
1745
|
// held (text: cloud or refused; with an image: local only).
|
|
1707
1746
|
if (budget.tripped)
|
|
1708
|
-
l1 =
|
|
1747
|
+
l1 = raise(l1, "UNCERTAIN", "budget");
|
|
1709
1748
|
if (budget.calls > LAYER1_SCREEN_CALL_BUDGET) {
|
|
1710
1749
|
attempts.push({ tier: "layer1", reason: `layer1_screen_over_budget:${LAYER1_SCREEN_CALL_BUDGET}` });
|
|
1711
1750
|
}
|
|
@@ -1753,7 +1792,7 @@ export async function runInfer(args, deps) {
|
|
|
1753
1792
|
attempts.push({ tier: "synalux", reason: "reserved_escalation_refused_images_stay_local" });
|
|
1754
1793
|
if (wantReport)
|
|
1755
1794
|
return refusedResult("layer1_reserved");
|
|
1756
|
-
throw makeReservedRefusal(l1, attempts, reservedCat, true);
|
|
1795
|
+
throw makeReservedRefusal(l1, attempts, reservedCat, true, ledgerMeta());
|
|
1757
1796
|
}
|
|
1758
1797
|
if (allowCloud) {
|
|
1759
1798
|
const cloudTimeout = args.timeout_ms ?? 90_000;
|
|
@@ -1769,7 +1808,7 @@ export async function runInfer(args, deps) {
|
|
|
1769
1808
|
attempts.push({ tier: "synalux", reason: `reserved_weak_backend:${cloud.backend}` });
|
|
1770
1809
|
if (wantReport)
|
|
1771
1810
|
return refusedResult("layer1_reserved");
|
|
1772
|
-
throw makeReservedRefusal(l1, attempts, reservedCat, true);
|
|
1811
|
+
throw makeReservedRefusal(l1, attempts, reservedCat, true, ledgerMeta());
|
|
1773
1812
|
}
|
|
1774
1813
|
return await applyVerification(cloud.output, gatedArgs, deps, {
|
|
1775
1814
|
backend: cloud.backend ?? "synalux",
|
|
@@ -1787,7 +1826,7 @@ export async function runInfer(args, deps) {
|
|
|
1787
1826
|
}
|
|
1788
1827
|
if (wantReport)
|
|
1789
1828
|
return refusedResult("layer1_reserved");
|
|
1790
|
-
throw makeReservedRefusal(l1, attempts, reservedCat, allowCloud);
|
|
1829
|
+
throw makeReservedRefusal(l1, attempts, reservedCat, allowCloud, ledgerMeta());
|
|
1791
1830
|
}
|
|
1792
1831
|
if (l1 === "UNCERTAIN_LENGTH") {
|
|
1793
1832
|
// §5.3: prompt too long to classify in full, but the full-text
|
|
@@ -1818,7 +1857,7 @@ export async function runInfer(args, deps) {
|
|
|
1818
1857
|
attempts.push({ tier: "synalux", reason: "error_escalation_refused_images_stay_local" });
|
|
1819
1858
|
if (wantReport)
|
|
1820
1859
|
return refusedResult("layer1_error");
|
|
1821
|
-
throw makeReservedRefusal(l1, attempts);
|
|
1860
|
+
throw makeReservedRefusal(l1, attempts, null, false, ledgerMeta());
|
|
1822
1861
|
}
|
|
1823
1862
|
if (allowCloud) {
|
|
1824
1863
|
const cloudTimeout = args.timeout_ms ?? 90_000;
|
|
@@ -1842,6 +1881,7 @@ export async function runInfer(args, deps) {
|
|
|
1842
1881
|
debugLog(`[prism_infer] keyword backstop verdict=${backstop}`);
|
|
1843
1882
|
attempts.push({ tier: "keyword_backstop", reason: `backstop_${backstop.toLowerCase()}` });
|
|
1844
1883
|
if (backstop === "OBVIOUS_RESERVED") {
|
|
1884
|
+
l1Layer = "backstop"; // the regex net refused, whatever raised the verdict before it
|
|
1845
1885
|
if (wantReport)
|
|
1846
1886
|
return refusedResult("keyword_backstop_reserved");
|
|
1847
1887
|
// Serve-mode backstop refusal previously wrote NO ledger row —
|
|
@@ -1850,6 +1890,8 @@ export async function runInfer(args, deps) {
|
|
|
1850
1890
|
backend: "refused", model: null, used_cloud: false,
|
|
1851
1891
|
gate_outcome: "refused",
|
|
1852
1892
|
refusal_reason: "keyword_backstop_reserved",
|
|
1893
|
+
history_turns: args.messages?.length ?? 0,
|
|
1894
|
+
refusal_layer: "backstop",
|
|
1853
1895
|
});
|
|
1854
1896
|
throw new Error(`prism_infer: classifier failed + keyword backstop caught reserved content. attempts=${JSON.stringify(attempts)}`);
|
|
1855
1897
|
}
|
|
@@ -347,7 +347,10 @@ function buildRecommendedArgs(args, complexityScore) {
|
|
|
347
347
|
...(args.project ? { project: args.project } : {}),
|
|
348
348
|
mode: "code",
|
|
349
349
|
task_complexity: complexityScore,
|
|
350
|
-
cloud_fallback:
|
|
350
|
+
// cloud_fallback is deliberately absent: prism_infer resolves it from the
|
|
351
|
+
// plan's entitlements, which the router does not read. Pinning it false
|
|
352
|
+
// here made a paid plan's escalation unreachable for any host that copied
|
|
353
|
+
// these arguments verbatim (2026-09-16).
|
|
351
354
|
escalation: "report",
|
|
352
355
|
};
|
|
353
356
|
}
|
|
@@ -139,6 +139,8 @@ export function recordInference(result) {
|
|
|
139
139
|
completion_tokens: result.completion_tokens,
|
|
140
140
|
latency_ms: result.latency_ms,
|
|
141
141
|
ram_free_mb: result.ram_free_mb,
|
|
142
|
+
history_turns: result.history_turns,
|
|
143
|
+
refusal_layer: result.refusal_layer,
|
|
142
144
|
});
|
|
143
145
|
// §5.2: refused results (escalation:"report") get a ledger row above but
|
|
144
146
|
// must NOT touch the session accumulators — no model ran and nothing was
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "prism-mcp-server",
|
|
3
|
-
"version": "20.
|
|
3
|
+
"version": "20.21.0",
|
|
4
4
|
"mcpName": "io.github.dcostenco/prism-coder",
|
|
5
5
|
"description": "Persistent session memory for AI coding agents that never leaves your machine — including the on-device model that reasons over it. Restores your prior decisions, open TODOs, and changed files across sessions; adds associative recall of related past work, semantic drift detection, and local inference. Local-first by default. Works with Claude Code, Cursor, and Codex.",
|
|
6
6
|
"module": "index.ts",
|