prism-mcp-server 20.19.0 → 20.21.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -32,8 +32,11 @@ A paid subscription adds cloud sync, higher model tiers, and team features throu
32
32
  undercounts inline — a number you can check, not marketing.
33
33
  - **Route-output enforcement** — route mode returns only well-formed calls to
34
34
  tools the host actually advertised. Standard and higher plans can add
35
- authenticated deterministic correction; `route_guard: "local"` keeps the
36
- prompt and draft entirely on-device.
35
+ authenticated deterministic correction; `route_guard: "local"` disables that
36
+ correction only. `cloud_fallback: false` forbids cloud inference fallback and
37
+ `verify: false` (with no `evidence`) disables the grounding verifier. With all
38
+ three off, no request carries your prompt, draft or evidence; the per-call
39
+ entitlement check and telemetry still contact the portal and carry neither.
37
40
  - **One setup for every agent** — `prism connect` configures Claude Code,
38
41
  Claude Desktop, Cursor, Gemini CLI, and Codex while preserving unrelated
39
42
  settings.
@@ -46,7 +49,8 @@ A paid subscription adds cloud sync, higher model tiers, and team features throu
46
49
  small ones on top: mid-session prompt routing, and a post-compaction
47
50
  re-injection of the protected-floor digest.
48
51
  - **Safe escalation and observability** — inference outcomes are explicit,
49
- reserved content remains fail-closed, and local/cloud usage is recorded for
52
+ reserved text remains fail-closed (clinical images are processed locally,
53
+ never sent to the cloud), and local/cloud usage is recorded for
50
54
  review.
51
55
 
52
56
  ## Get started
@@ -781,7 +785,7 @@ Qwen 3.5 models (9B/27B) with thinking enabled could burn all tokens on `<think>
781
785
  ## What's New in v20.0.3
782
786
 
783
787
  ### Layer 1 Cold-Model Resilience
784
- The reserved-category classifier now retries once with a longer timeout on cold-model failure, then falls back to a deterministic keyword backstop before refusing. Over-length prompts (>4K chars) are classified as UNCERTAIN before reaching the classifier — prompt padding can no longer force the ERROR branch. This eliminates the cold-start refusal problem without weakening the safety gate.
788
+ (As shipped in an earlier release; the current contract is the header of `src/utils/layer1.ts`.) The reserved-category classifier retries once with a longer timeout on cold-model failure, then falls back to a deterministic keyword backstop; keyword-clean text is served locally. Over-length prompts (>4K chars) get the full-text keyword floor plus a head+middle+tail excerpt read and a distinct UNCERTAIN_LENGTH marker — prompt padding cannot force the ERROR branch. This eliminates the cold-start refusal problem without weakening the safety gate.
785
789
 
786
790
  ### Keyword Backstop for Reserved Content
787
791
  When the LLM classifier fails (timeout, injection, resource pressure), a deterministic regex floor catches reserved vocabulary (restraint, seclusion, self-harm, suicide, overdose, crisis de-escalation, etc.) including inflected and verb forms. Blocks prompt-padding and classifier-injection attacks on the ERROR path.
@@ -1033,8 +1037,10 @@ The free tier runs entirely on your machine. Paid tiers add cloud sync through t
1033
1037
  redaction (SSNs, dates of birth, medical record numbers, phone numbers, emails,
1034
1038
  and clinical identifiers are stripped before storage). Cloud inference and
1035
1039
  route correction send the request over TLS for processing and do not store it
1036
- as Prism memory; use `route_guard: "local"` or the **local tier** for a full
1037
- air-gap. **Enterprise** includes a HIPAA Business Associate Agreement.
1040
+ as Prism memory. The **local tier** (no Synalux key) is the air-gap;
1041
+ `route_guard: "local"` only skips the route correction, and `cloud_fallback`
1042
+ and `verify` govern the other two channels. **Enterprise** includes a HIPAA
1043
+ Business Associate Agreement.
1038
1044
 
1039
1045
  ---
1040
1046
 
@@ -1049,8 +1055,12 @@ before it reaches the host. Malformed or unadvertised calls become `NO_TOOL`.
1049
1055
  With `route_guard: "auto"` (the default), Standard and higher plans also send
1050
1056
  a well-formed draft for one of Prism's seven trained tools—or an unadvertised
1051
1057
  draft that may need correction—to Synalux for authenticated deterministic
1052
- correction. Advertised custom host tools remain local. Set
1053
- `route_guard: "local"` for a fully on-device route path.
1058
+ correction. Advertised custom host tools remain local. `route_guard: "local"`
1059
+ disables that correction only: cloud inference fallback is governed by
1060
+ `cloud_fallback`, and the grounding verifier is a separate channel with its
1061
+ own switch (`verify`, on by default when `evidence` is given). With all three
1062
+ off, no request carries your prompt, draft or evidence; the entitlement check
1063
+ and telemetry still contact the portal and carry neither.
1054
1064
 
1055
1065
  | Model | Ollama tag | Size | Vision | Routing accuracy¹ | Role | Automatic routing tier |
1056
1066
  |---|---|---|---|---|---|---|
@@ -1151,12 +1161,19 @@ check, the local 9B passed 2/3 tasks; the local 27B and Gemini 3.6 Flash each
1151
1161
  passed 3/3. This is a self-published regression signal, not an independent
1152
1162
  leaderboard or a claim of broad model equivalence.
1153
1163
 
1154
- ### Cloud Escalation (`cloud_fallback: true`)
1164
+ ### Cloud Escalation (`cloud_fallback`)
1155
1165
 
1156
1166
  Prism always tries an eligible local model first. If the quality gate detects
1157
- an empty, truncated, think-only, or looping response, paid tiers can retry the
1158
- request through Gemini 3.6 Flash. Free-tier routing stays local and reports the
1159
- quality-gate outcome without making a cloud call.
1167
+ an empty, truncated, think-only, or looping response, or the safety screen
1168
+ finds the request uncertain or reserved, a paid plan escalates the request
1169
+ through Gemini 3.6 Flash. Free-tier routing stays local and reports the
1170
+ outcome without making a cloud call.
1171
+
1172
+ The flag follows the plan. Leave it unset and your plan decides: paid plans
1173
+ escalate, free plans never do. Pass `false` to forbid cloud inference fallback
1174
+ for a call, which the clinical delegation rules do for drafting; pass `true`
1175
+ to ask for it, which still requires a plan with cloud. A request that carries
1176
+ an image is never escalated; screenshots stay on this device.
1160
1177
 
1161
1178
  ---
1162
1179
 
@@ -1218,6 +1235,7 @@ All on-device models are free to run locally via Ollama on every tier. A subscri
1218
1235
  | Cloud search | -- | ✅ | ✅ | ✅ |
1219
1236
  | Max output tokens | 512 | 1,024 | 2,048 | 4,096 |
1220
1237
  | Cloud fallback | -- | Gemini 3.6 Flash | Gemini 3.6 Flash | Gemini 3.6 Flash (priority) |
1238
+ | Multi-turn `prism_infer` (conversation carried across calls) | -- | 12 turns / 32k chars | 20 turns / 64k chars | 30 turns / 96k chars |
1221
1239
  | Grounding verifier (fact-check AI output) | -- | ✅ | ✅ | ✅ |
1222
1240
  | Memory sync (cloud) | -- | ✅ | ✅ | ✅ |
1223
1241
  | Knowledge / session memory | limited | unlimited | unlimited | unlimited |
@@ -1270,8 +1288,115 @@ prism_infer({
1270
1288
  })
1271
1289
  // → 27B generates code locally ($0), with thinking for quality
1272
1290
  // → If quality gate fails + paid tier → auto-escalate to Gemini 3.6 Flash
1291
+
1292
+ // Follow-ups carry the conversation (paid plans). The host curates the turns;
1293
+ // Prism bounds them to your plan's caps (user+assistant text only), safety-
1294
+ // screens every turn (each alone, then each in context; a short routine
1295
+ // request skips its own model read), counts them against
1296
+ // the tier's context, and never stores them. A free plan or a host with no
1297
+ // portal is refused: multi_turn_not_in_plan.
1298
+ prism_infer({
1299
+ messages: [
1300
+ { role: "user", content: "Write a binary search in Python" },
1301
+ { role: "assistant", content: "<the accepted answer>" },
1302
+ ],
1303
+ prompt: "Now make it return the insertion point when the value is absent",
1304
+ mode: "code",
1305
+ })
1273
1306
  ```
1274
1307
 
1308
+ #### Multi-turn: why it matters, and why it is a paid feature
1309
+
1310
+ A single `prism_infer` call has no memory. The host asks a question, the local
1311
+ model answers, and the next call starts from nothing: "now add a timeout to it"
1312
+ means nothing to a model that never saw "it". Without history the host either
1313
+ restates the whole context in every prompt (tokens, and the answer drifts) or
1314
+ gives up on local delegation and does the follow-up itself in the cloud. With
1315
+ `messages`, the host hands back the turns it accepted and the local model
1316
+ continues the thread at $0, which is what makes local delegation useful for
1317
+ real work instead of one-shot snippets.
1318
+
1319
+ It is paid because it cannot run without Synalux behind it:
1320
+
1321
+ - **The caps are served by the portal, per plan.** How many turns and how many
1322
+ characters a conversation may carry is an entitlement your plan returns,
1323
+ not a number the client decides. No portal, no policy, and the call is
1324
+ refused as `multi_turn_not_in_plan`.
1325
+ - **Ambiguous turns go to Synalux cloud.** Every turn is safety-screened on
1326
+ device, alone and in context. A turn the screen finds uncertain is never
1327
+ served by the local model; it goes to the cloud when the plan allows it,
1328
+ and is refused otherwise. Free has no cloud, so the only honest answer for
1329
+ free is no history at all.
1330
+ - **Nothing is stored.** Prism bounds, screens, counts and forwards the turns
1331
+ the host sends, and keeps none of them.
1332
+
1333
+ <details>
1334
+ <summary>Without multi-turn (free plan, or no `messages`): every call starts cold</summary>
1335
+
1336
+ ```typescript
1337
+ // Call 1
1338
+ prism_infer({ prompt: "My project codename is Nightjar. Reply OK.", mode: "chat" })
1339
+ // → "OK" (local 9b, $0)
1340
+
1341
+ // Call 2 — the model never saw call 1
1342
+ prism_infer({ prompt: "What is my codename? One word.", mode: "chat" })
1343
+ // → "I don't have that information." (local 9b, correct and useless)
1344
+
1345
+ // Call 3 — a coding follow-up with no thread
1346
+ prism_infer({ prompt: "Now add a timeout parameter to it.", mode: "code" })
1347
+ // → guesses what "it" is, or asks (the host has to redo the work)
1348
+
1349
+ // A free plan that sends messages anyway:
1350
+ prism_infer({ messages: [/* … */], prompt: "…" })
1351
+ // → refused: multi_turn_not_in_plan (no portal entitlement, no cloud)
1352
+ ```
1353
+ </details>
1354
+
1355
+ <details>
1356
+ <summary>With multi-turn (paid plan): the thread continues locally, and the screen decides per turn</summary>
1357
+
1358
+ ```typescript
1359
+ // The host keeps the turns it accepted and passes them back.
1360
+ prism_infer({
1361
+ messages: [
1362
+ { role: "user", content: "My project codename is Nightjar. Reply OK." },
1363
+ { role: "assistant", content: "OK" },
1364
+ ],
1365
+ prompt: "What is my codename? One word.",
1366
+ mode: "chat",
1367
+ })
1368
+ // → "Nightjar" (local 9b, $0; history_turns: 2)
1369
+
1370
+ prism_infer({
1371
+ messages: [
1372
+ { role: "user", content: "We named the helper countActiveUsers(data). Confirm." },
1373
+ { role: "assistant", content: "Confirmed." },
1374
+ ],
1375
+ prompt: "Write the one-line call that stores its result in n.",
1376
+ mode: "code",
1377
+ })
1378
+ // → "n = countActiveUsers(data)" (local 9b, $0)
1379
+
1380
+ // A turn the on-device screen finds uncertain, alone or in context, is not
1381
+ // served locally: it goes to Synalux cloud on a paid plan, or is refused with
1382
+ // cloud_fallback: false. The result names the reason (layer1_uncertain) so
1383
+ // the host can decide what to do with the thread.
1384
+ prism_infer({
1385
+ messages: [
1386
+ { role: "user", content: "The ticket for this bug is SYN-4471. Acknowledge." },
1387
+ { role: "assistant", content: "Acknowledged." },
1388
+ ],
1389
+ prompt: "Which ticket is this bug filed under?",
1390
+ cloud_fallback: true,
1391
+ })
1392
+ // → "SYN-4471" (Gemini 3.6 Flash; used_cloud: true)
1393
+ ```
1394
+
1395
+ Measured on the real 9b through the real handler with cloud off: 7 of 12
1396
+ benign follow-ups are served locally, the rest refuse and name the reason;
1397
+ every injection variant the reviewers built refuses.
1398
+ </details>
1399
+
1275
1400
  | Mode | Think | Model | Use case |
1276
1401
  |------|-------|-------|----------|
1277
1402
  | `route` | Off (fast) — except a tier that reasons better, e.g. 9B | 9B default | MCP tool routing |
package/dist/cli.js CHANGED
@@ -1273,8 +1273,9 @@ program
1273
1273
  .option('--dry-run', 'Print what would be pulled/re-aliased without executing')
1274
1274
  .action(async (options) => {
1275
1275
  const { runOllamaConverge } = await import('./modelConvergeRunner.js');
1276
+ const { convergeFailed } = await import('./utils/modelConverge.js');
1276
1277
  const outcomes = await runOllamaConverge({ dryRun: options.dryRun === true });
1277
- if (outcomes.every(o => o.action === 'failed'))
1278
+ if (convergeFailed(outcomes))
1278
1279
  process.exitCode = 1;
1279
1280
  });
1280
1281
  // ─── prism register-models ────────────────────────────────────
@@ -9,7 +9,10 @@ export const LOCAL_FIRST_POLICY_LINES = [
9
9
  "## Prism local-first orchestration",
10
10
  "For bounded, verifiable delegated work, use Prism's local worker before any host-native or background subagent.",
11
11
  "When session_task_route is available, call it first. If it returns target `claw`, call `prism_infer` with the",
12
- "recommended arguments, `cloud_fallback: false`, and the `project` and `conversation_id` from this session when known.",
12
+ "recommended arguments and the `project` and `conversation_id` from this session when known. Leave `cloud_fallback`",
13
+ "unset so the plan decides it; pass `false` to forbid cloud inference fallback.",
14
+ "If the route says `needs_history`, or the task follows up an earlier `prism_infer` answer, pass the accepted prior",
15
+ "turns as `messages` (paid plans); never send a follow-up without them, and never re-send a turn you rejected.",
13
16
  "Prism loads that project's dashboard-configured quick, standard, or deep memory and selects a RAM-safe local model.",
14
17
  "Never create host-native or background subagents for routine work, never fan out, and never nest agents.",
15
18
  "If local inference is unavailable, refused, degraded, or the task requires host tools or reserved judgment, continue",
@@ -81,6 +81,20 @@ export async function runOllamaConverge(opts = {}) {
81
81
  const data = (await res.json());
82
82
  return (data.models ?? []).map((m) => ({ name: m.name, digest: m.digest }));
83
83
  },
84
+ tagFacts: async (name) => {
85
+ const res = await fetch(`${OLLAMA_URL}/api/show`, {
86
+ method: "POST", headers: { "Content-Type": "application/json" },
87
+ body: JSON.stringify({ model: name }), signal: AbortSignal.timeout(5_000),
88
+ });
89
+ if (!res.ok)
90
+ return null;
91
+ const data = (await res.json());
92
+ const from = /^FROM\s+(\S+)/m.exec(data.modelfile ?? "")?.[1];
93
+ if (!from)
94
+ return null;
95
+ const pin = /^\s*num_ctx\s+(\d+)\s*$/m.exec(data.parameters ?? "");
96
+ return { weightsBlob: from, pinnedNumCtx: pin ? Number(pin[1]) : null };
97
+ },
84
98
  pull: (ref) => runOllama(["pull", ref], true),
85
99
  copy: (from, to) => runOllama(["cp", from, to], false),
86
100
  log: (line) => console.log(` ${line}`),
@@ -42,8 +42,8 @@ const LEDGER_UNAVAILABLE_ERROR = "Inference metrics ledger is unavailable";
42
42
  const INSERT_METRIC_SQL = `INSERT OR IGNORE INTO infer_metrics
43
43
  (ts, caller, mode, backend, model, used_cloud, gate_outcome,
44
44
  refusal_reason, prompt_tokens, completion_tokens, latency_ms, ram_free_mb,
45
- source_event_id)
46
- VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)`;
45
+ source_event_id, history_turns, refusal_layer)
46
+ VALUES (?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?, ?)`;
47
47
  function closeClient(context) {
48
48
  const activeClient = client;
49
49
  client = null;
@@ -83,17 +83,25 @@ function ensureTable() {
83
83
  completion_tokens INTEGER,
84
84
  latency_ms INTEGER,
85
85
  ram_free_mb INTEGER,
86
- source_event_id TEXT
86
+ source_event_id TEXT,
87
+ history_turns INTEGER,
88
+ refusal_layer TEXT
87
89
  )`);
88
90
  // Existing ledgers predate external panel-spool ingestion. SQLite has
89
91
  // no ADD COLUMN IF NOT EXISTS, so use the repository's established
90
92
  // idempotent migration pattern and reject only unexpected failures.
91
- try {
92
- await client.execute(`ALTER TABLE infer_metrics ADD COLUMN source_event_id TEXT`);
93
- }
94
- catch (e) {
95
- if (!(e instanceof Error) || !e.message.includes("duplicate column name"))
96
- throw e;
93
+ for (const column of [
94
+ "source_event_id TEXT",
95
+ "history_turns INTEGER",
96
+ "refusal_layer TEXT",
97
+ ]) {
98
+ try {
99
+ await client.execute(`ALTER TABLE infer_metrics ADD COLUMN ${column}`);
100
+ }
101
+ catch (e) {
102
+ if (!(e instanceof Error) || !e.message.includes("duplicate column name"))
103
+ throw e;
104
+ }
97
105
  }
98
106
  await client.execute(`CREATE INDEX IF NOT EXISTS idx_infer_metrics_ts ON infer_metrics (ts)`);
99
107
  await client.execute(`CREATE UNIQUE INDEX IF NOT EXISTS idx_infer_metrics_source_event
@@ -149,6 +157,7 @@ function metricArgs(row) {
149
157
  row.refusal_reason ?? null, row.prompt_tokens ?? null,
150
158
  row.completion_tokens ?? null, row.latency_ms ?? null,
151
159
  row.ram_free_mb ?? null, row.source_event_id ?? null,
160
+ row.history_turns ?? null, row.refusal_layer ?? null,
152
161
  ];
153
162
  }
154
163
  /** Aggregate all persisted rows (optionally since a timestamp). */
@@ -25,7 +25,7 @@ import { buildVaultDirectory } from "../utils/vaultExporter.js";
25
25
  * ═══════════════════════════════════════════════════════════════════
26
26
  */
27
27
  import { debugLog } from "../utils/logger.js";
28
- import { FREE_ENTITLEMENTS } from "../utils/entitlements.js";
28
+ import { FREE_ENTITLEMENTS, peekEntitlements, multiTurnPolicy } from "../utils/entitlements.js";
29
29
  import { getStorage, activeStorageBackend } from "../storage/index.js";
30
30
  import { toKeywordArray } from "../utils/keywordExtractor.js";
31
31
  import { getLLMProvider } from "../utils/llm/factory.js";
@@ -550,6 +550,7 @@ async function buildNativeSystemReadyBlock(snapshot, depth) {
550
550
  `> - 🧠 **Context depth:** ${depth}\n` +
551
551
  `> - 🔄 **Skill sync:** ${SKILL_SYNC_STATUS_LABELS[snapshot.syncStatus]} · native materialization incomplete${conflictSuffix}` +
552
552
  conflictWarning +
553
+ localWorkerLine() +
553
554
  freeTierUpgradeLine(snapshot.tier));
554
555
  }
555
556
  if (snapshot.source === "tier-fallback") {
@@ -560,6 +561,7 @@ async function buildNativeSystemReadyBlock(snapshot, depth) {
560
561
  `> - 🧠 **Context depth:** ${depth}\n` +
561
562
  `> - 🔄 **Skill sync:** ${SKILL_SYNC_STATUS_LABELS[snapshot.syncStatus]} · no committed manifest${conflictSuffix}` +
562
563
  conflictWarning +
564
+ localWorkerLine() +
563
565
  freeTierUpgradeLine(snapshot.tier));
564
566
  }
565
567
  return wrap(`> **Prism System Ready**\n>` + undeliveredWarning + `\n` +
@@ -572,6 +574,7 @@ async function buildNativeSystemReadyBlock(snapshot, depth) {
572
574
  `> - 🧠 **Context depth:** ${depth}\n` +
573
575
  `> - 🔄 **Skill sync:** ${SKILL_SYNC_STATUS_LABELS[snapshot.syncStatus]} · committed manifest${conflictSuffix}` +
574
576
  conflictWarning +
577
+ localWorkerLine() +
575
578
  freeTierUpgradeLine(snapshot.tier));
576
579
  }
577
580
  /**
@@ -580,6 +583,22 @@ async function buildNativeSystemReadyBlock(snapshot, depth) {
580
583
  * queryMemoryNaturalHandler); the startup path — the only guaranteed
581
584
  * impression — referenced it zero times.
582
585
  */
586
+ /**
587
+ * One startup line so the host knows, before its first delegation, whether
588
+ * the local worker takes conversation history and how much. Reads the
589
+ * entitlements CACHE only — never a portal fetch on the startup path; when
590
+ * cold, the first prism_infer result carries the same policy.
591
+ */
592
+ function localWorkerLine() {
593
+ const ent = peekEntitlements();
594
+ if (!ent)
595
+ return "";
596
+ const p = multiTurnPolicy(ent);
597
+ return p.enabled
598
+ ? `\n> - 🧵 **Local worker multi-turn:** on — up to ${p.max_turns} turns / ` +
599
+ `${p.max_chars.toLocaleString("en-US")} chars per prism_infer call; pass accepted prior turns as \`messages\``
600
+ : `\n> - 🧵 **Local worker multi-turn:** off on the ${ent.plan} plan — a prism_infer follow-up is answered without context`;
601
+ }
583
602
  function freeTierUpgradeLine(tier) {
584
603
  if (tier !== "free")
585
604
  return "";