prism-mcp-server 20.19.0 → 20.20.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -46,7 +46,8 @@ A paid subscription adds cloud sync, higher model tiers, and team features throu
46
46
  small ones on top: mid-session prompt routing, and a post-compaction
47
47
  re-injection of the protected-floor digest.
48
48
  - **Safe escalation and observability** — inference outcomes are explicit,
49
- reserved content remains fail-closed, and local/cloud usage is recorded for
49
+ reserved text remains fail-closed (clinical images are processed locally,
50
+ never sent to the cloud), and local/cloud usage is recorded for
50
51
  review.
51
52
 
52
53
  ## Get started
@@ -781,7 +782,7 @@ Qwen 3.5 models (9B/27B) with thinking enabled could burn all tokens on `<think>
781
782
  ## What's New in v20.0.3
782
783
 
783
784
  ### Layer 1 Cold-Model Resilience
784
- The reserved-category classifier now retries once with a longer timeout on cold-model failure, then falls back to a deterministic keyword backstop before refusing. Over-length prompts (>4K chars) are classified as UNCERTAIN before reaching the classifier — prompt padding can no longer force the ERROR branch. This eliminates the cold-start refusal problem without weakening the safety gate.
785
+ (As shipped in an earlier release; the current contract is the header of `src/utils/layer1.ts`.) The reserved-category classifier retries once with a longer timeout on cold-model failure, then falls back to a deterministic keyword backstop; keyword-clean text is served locally. Over-length prompts (>4K chars) get the full-text keyword floor plus a head+middle+tail excerpt read and a distinct UNCERTAIN_LENGTH marker — prompt padding cannot force the ERROR branch. This eliminates the cold-start refusal problem without weakening the safety gate.
785
786
 
786
787
  ### Keyword Backstop for Reserved Content
787
788
  When the LLM classifier fails (timeout, injection, resource pressure), a deterministic regex floor catches reserved vocabulary (restraint, seclusion, self-harm, suicide, overdose, crisis de-escalation, etc.) including inflected and verb forms. Blocks prompt-padding and classifier-injection attacks on the ERROR path.
@@ -1218,6 +1219,7 @@ All on-device models are free to run locally via Ollama on every tier. A subscri
1218
1219
  | Cloud search | -- | ✅ | ✅ | ✅ |
1219
1220
  | Max output tokens | 512 | 1,024 | 2,048 | 4,096 |
1220
1221
  | Cloud fallback | -- | Gemini 3.6 Flash | Gemini 3.6 Flash | Gemini 3.6 Flash (priority) |
1222
+ | Multi-turn `prism_infer` (conversation carried across calls) | -- | 12 turns / 32k chars | 20 turns / 64k chars | 30 turns / 96k chars |
1221
1223
  | Grounding verifier (fact-check AI output) | -- | ✅ | ✅ | ✅ |
1222
1224
  | Memory sync (cloud) | -- | ✅ | ✅ | ✅ |
1223
1225
  | Knowledge / session memory | limited | unlimited | unlimited | unlimited |
@@ -1270,8 +1272,115 @@ prism_infer({
1270
1272
  })
1271
1273
  // → 27B generates code locally ($0), with thinking for quality
1272
1274
  // → If quality gate fails + paid tier → auto-escalate to Gemini 3.6 Flash
1275
+
1276
+ // Follow-ups carry the conversation (paid plans). The host curates the turns;
1277
+ // Prism bounds them to your plan's caps (user+assistant text only), safety-
1278
+ // screens every turn (each alone, then each in context; a short routine
1279
+ // request skips its own model read), counts them against
1280
+ // the tier's context, and never stores them. A free plan or a host with no
1281
+ // portal is refused: multi_turn_not_in_plan.
1282
+ prism_infer({
1283
+ messages: [
1284
+ { role: "user", content: "Write a binary search in Python" },
1285
+ { role: "assistant", content: "<the accepted answer>" },
1286
+ ],
1287
+ prompt: "Now make it return the insertion point when the value is absent",
1288
+ mode: "code",
1289
+ })
1290
+ ```
1291
+
1292
+ #### Multi-turn: why it matters, and why it is a paid feature
1293
+
1294
+ A single `prism_infer` call has no memory. The host asks a question, the local
1295
+ model answers, and the next call starts from nothing: "now add a timeout to it"
1296
+ means nothing to a model that never saw "it". Without history the host either
1297
+ restates the whole context in every prompt (tokens, and the answer drifts) or
1298
+ gives up on local delegation and does the follow-up itself in the cloud. With
1299
+ `messages`, the host hands back the turns it accepted and the local model
1300
+ continues the thread at $0, which is what makes local delegation useful for
1301
+ real work instead of one-shot snippets.
1302
+
1303
+ It is paid because it cannot run without Synalux behind it:
1304
+
1305
+ - **The caps are served by the portal, per plan.** How many turns and how many
1306
+ characters a conversation may carry is an entitlement your plan returns,
1307
+ not a number the client decides. No portal, no policy, and the call is
1308
+ refused as `multi_turn_not_in_plan`.
1309
+ - **Ambiguous turns go to Synalux cloud.** Every turn is safety-screened on
1310
+ device, alone and in context. A turn the screen finds uncertain is never
1311
+ served by the local model; it goes to the cloud when the plan allows it,
1312
+ and is refused otherwise. Free has no cloud, so the only honest answer for
1313
+ free is no history at all.
1314
+ - **Nothing is stored.** Prism bounds, screens, counts and forwards the turns
1315
+ the host sends, and keeps none of them.
1316
+
1317
+ <details>
1318
+ <summary>Without multi-turn (free plan, or no `messages`): every call starts cold</summary>
1319
+
1320
+ ```typescript
1321
+ // Call 1
1322
+ prism_infer({ prompt: "My project codename is Nightjar. Reply OK.", mode: "chat" })
1323
+ // → "OK" (local 9b, $0)
1324
+
1325
+ // Call 2 — the model never saw call 1
1326
+ prism_infer({ prompt: "What is my codename? One word.", mode: "chat" })
1327
+ // → "I don't have that information." (local 9b, correct and useless)
1328
+
1329
+ // Call 3 — a coding follow-up with no thread
1330
+ prism_infer({ prompt: "Now add a timeout parameter to it.", mode: "code" })
1331
+ // → guesses what "it" is, or asks (the host has to redo the work)
1332
+
1333
+ // A free plan that sends messages anyway:
1334
+ prism_infer({ messages: [/* … */], prompt: "…" })
1335
+ // → refused: multi_turn_not_in_plan (no portal entitlement, no cloud)
1336
+ ```
1337
+ </details>
1338
+
1339
+ <details>
1340
+ <summary>With multi-turn (paid plan): the thread continues locally, and the screen decides per turn</summary>
1341
+
1342
+ ```typescript
1343
+ // The host keeps the turns it accepted and passes them back.
1344
+ prism_infer({
1345
+ messages: [
1346
+ { role: "user", content: "My project codename is Nightjar. Reply OK." },
1347
+ { role: "assistant", content: "OK" },
1348
+ ],
1349
+ prompt: "What is my codename? One word.",
1350
+ mode: "chat",
1351
+ })
1352
+ // → "Nightjar" (local 9b, $0; history_turns: 2)
1353
+
1354
+ prism_infer({
1355
+ messages: [
1356
+ { role: "user", content: "We named the helper countActiveUsers(data). Confirm." },
1357
+ { role: "assistant", content: "Confirmed." },
1358
+ ],
1359
+ prompt: "Write the one-line call that stores its result in n.",
1360
+ mode: "code",
1361
+ })
1362
+ // → "n = countActiveUsers(data)" (local 9b, $0)
1363
+
1364
+ // A turn the on-device screen finds uncertain when read alone is not served
1365
+ // locally: it goes to Synalux cloud on a paid plan, or is refused with
1366
+ // cloud_fallback: false. The result names the reason (layer1_uncertain) so
1367
+ // the host can decide what to do with the thread.
1368
+ prism_infer({
1369
+ messages: [
1370
+ { role: "user", content: "The ticket for this bug is SYN-4471. Acknowledge." },
1371
+ { role: "assistant", content: "Acknowledged." },
1372
+ ],
1373
+ prompt: "Which ticket is this bug filed under?",
1374
+ cloud_fallback: true,
1375
+ })
1376
+ // → "SYN-4471" (Gemini 3.6 Flash; used_cloud: true)
1273
1377
  ```
1274
1378
 
1379
+ Measured on the real 9b through the real handler with cloud off: 7 of 12
1380
+ benign follow-ups are served locally, the rest refuse and name the reason;
1381
+ every injection variant the reviewers built refuses.
1382
+ </details>
1383
+
1275
1384
  | Mode | Think | Model | Use case |
1276
1385
  |------|-------|-------|----------|
1277
1386
  | `route` | Off (fast) — except a tier that reasons better, e.g. 9B | 9B default | MCP tool routing |
package/dist/cli.js CHANGED
@@ -1273,8 +1273,9 @@ program
1273
1273
  .option('--dry-run', 'Print what would be pulled/re-aliased without executing')
1274
1274
  .action(async (options) => {
1275
1275
  const { runOllamaConverge } = await import('./modelConvergeRunner.js');
1276
+ const { convergeFailed } = await import('./utils/modelConverge.js');
1276
1277
  const outcomes = await runOllamaConverge({ dryRun: options.dryRun === true });
1277
- if (outcomes.every(o => o.action === 'failed'))
1278
+ if (convergeFailed(outcomes))
1278
1279
  process.exitCode = 1;
1279
1280
  });
1280
1281
  // ─── prism register-models ────────────────────────────────────
@@ -10,6 +10,8 @@ export const LOCAL_FIRST_POLICY_LINES = [
10
10
  "For bounded, verifiable delegated work, use Prism's local worker before any host-native or background subagent.",
11
11
  "When session_task_route is available, call it first. If it returns target `claw`, call `prism_infer` with the",
12
12
  "recommended arguments, `cloud_fallback: false`, and the `project` and `conversation_id` from this session when known.",
13
+ "If the route says `needs_history`, or the task follows up an earlier `prism_infer` answer, pass the accepted prior",
14
+ "turns as `messages` (paid plans); never send a follow-up without them, and never re-send a turn you rejected.",
13
15
  "Prism loads that project's dashboard-configured quick, standard, or deep memory and selects a RAM-safe local model.",
14
16
  "Never create host-native or background subagents for routine work, never fan out, and never nest agents.",
15
17
  "If local inference is unavailable, refused, degraded, or the task requires host tools or reserved judgment, continue",
@@ -81,6 +81,20 @@ export async function runOllamaConverge(opts = {}) {
81
81
  const data = (await res.json());
82
82
  return (data.models ?? []).map((m) => ({ name: m.name, digest: m.digest }));
83
83
  },
84
+ tagFacts: async (name) => {
85
+ const res = await fetch(`${OLLAMA_URL}/api/show`, {
86
+ method: "POST", headers: { "Content-Type": "application/json" },
87
+ body: JSON.stringify({ model: name }), signal: AbortSignal.timeout(5_000),
88
+ });
89
+ if (!res.ok)
90
+ return null;
91
+ const data = (await res.json());
92
+ const from = /^FROM\s+(\S+)/m.exec(data.modelfile ?? "")?.[1];
93
+ if (!from)
94
+ return null;
95
+ const pin = /^\s*num_ctx\s+(\d+)\s*$/m.exec(data.parameters ?? "");
96
+ return { weightsBlob: from, pinnedNumCtx: pin ? Number(pin[1]) : null };
97
+ },
84
98
  pull: (ref) => runOllama(["pull", ref], true),
85
99
  copy: (from, to) => runOllama(["cp", from, to], false),
86
100
  log: (line) => console.log(` ${line}`),
@@ -25,7 +25,7 @@ import { buildVaultDirectory } from "../utils/vaultExporter.js";
25
25
  * ═══════════════════════════════════════════════════════════════════
26
26
  */
27
27
  import { debugLog } from "../utils/logger.js";
28
- import { FREE_ENTITLEMENTS } from "../utils/entitlements.js";
28
+ import { FREE_ENTITLEMENTS, peekEntitlements, multiTurnPolicy } from "../utils/entitlements.js";
29
29
  import { getStorage, activeStorageBackend } from "../storage/index.js";
30
30
  import { toKeywordArray } from "../utils/keywordExtractor.js";
31
31
  import { getLLMProvider } from "../utils/llm/factory.js";
@@ -550,6 +550,7 @@ async function buildNativeSystemReadyBlock(snapshot, depth) {
550
550
  `> - 🧠 **Context depth:** ${depth}\n` +
551
551
  `> - 🔄 **Skill sync:** ${SKILL_SYNC_STATUS_LABELS[snapshot.syncStatus]} · native materialization incomplete${conflictSuffix}` +
552
552
  conflictWarning +
553
+ localWorkerLine() +
553
554
  freeTierUpgradeLine(snapshot.tier));
554
555
  }
555
556
  if (snapshot.source === "tier-fallback") {
@@ -560,6 +561,7 @@ async function buildNativeSystemReadyBlock(snapshot, depth) {
560
561
  `> - 🧠 **Context depth:** ${depth}\n` +
561
562
  `> - 🔄 **Skill sync:** ${SKILL_SYNC_STATUS_LABELS[snapshot.syncStatus]} · no committed manifest${conflictSuffix}` +
562
563
  conflictWarning +
564
+ localWorkerLine() +
563
565
  freeTierUpgradeLine(snapshot.tier));
564
566
  }
565
567
  return wrap(`> **Prism System Ready**\n>` + undeliveredWarning + `\n` +
@@ -572,6 +574,7 @@ async function buildNativeSystemReadyBlock(snapshot, depth) {
572
574
  `> - 🧠 **Context depth:** ${depth}\n` +
573
575
  `> - 🔄 **Skill sync:** ${SKILL_SYNC_STATUS_LABELS[snapshot.syncStatus]} · committed manifest${conflictSuffix}` +
574
576
  conflictWarning +
577
+ localWorkerLine() +
575
578
  freeTierUpgradeLine(snapshot.tier));
576
579
  }
577
580
  /**
@@ -580,6 +583,22 @@ async function buildNativeSystemReadyBlock(snapshot, depth) {
580
583
  * queryMemoryNaturalHandler); the startup path — the only guaranteed
581
584
  * impression — referenced it zero times.
582
585
  */
586
+ /**
587
+ * One startup line so the host knows, before its first delegation, whether
588
+ * the local worker takes conversation history and how much. Reads the
589
+ * entitlements CACHE only — never a portal fetch on the startup path; when
590
+ * cold, the first prism_infer result carries the same policy.
591
+ */
592
+ function localWorkerLine() {
593
+ const ent = peekEntitlements();
594
+ if (!ent)
595
+ return "";
596
+ const p = multiTurnPolicy(ent);
597
+ return p.enabled
598
+ ? `\n> - 🧵 **Local worker multi-turn:** on — up to ${p.max_turns} turns / ` +
599
+ `${p.max_chars.toLocaleString("en-US")} chars per prism_infer call; pass accepted prior turns as \`messages\``
600
+ : `\n> - 🧵 **Local worker multi-turn:** off on the ${ent.plan} plan — a prism_infer follow-up is answered without context`;
601
+ }
583
602
  function freeTierUpgradeLine(tier) {
584
603
  if (tier !== "free")
585
604
  return "";