prism-mcp-server 20.19.0 → 20.20.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +111 -2
- package/dist/cli.js +2 -1
- package/dist/localFirstPolicy.js +2 -0
- package/dist/modelConvergeRunner.js +14 -0
- package/dist/tools/ledgerHandlers.js +20 -1
- package/dist/tools/prismInferHandler.js +578 -82
- package/dist/tools/taskRouterHandler.js +47 -0
- package/dist/utils/entitlements.js +32 -0
- package/dist/utils/layer1.js +38 -16
- package/dist/utils/modelConverge.js +47 -2
- package/dist/utils/safetyGate.js +11 -1
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -46,7 +46,8 @@ A paid subscription adds cloud sync, higher model tiers, and team features throu
|
|
|
46
46
|
small ones on top: mid-session prompt routing, and a post-compaction
|
|
47
47
|
re-injection of the protected-floor digest.
|
|
48
48
|
- **Safe escalation and observability** — inference outcomes are explicit,
|
|
49
|
-
reserved
|
|
49
|
+
reserved text remains fail-closed (clinical images are processed locally,
|
|
50
|
+
never sent to the cloud), and local/cloud usage is recorded for
|
|
50
51
|
review.
|
|
51
52
|
|
|
52
53
|
## Get started
|
|
@@ -781,7 +782,7 @@ Qwen 3.5 models (9B/27B) with thinking enabled could burn all tokens on `<think>
|
|
|
781
782
|
## What's New in v20.0.3
|
|
782
783
|
|
|
783
784
|
### Layer 1 Cold-Model Resilience
|
|
784
|
-
The reserved-category classifier
|
|
785
|
+
(As shipped in an earlier release; the current contract is the header of `src/utils/layer1.ts`.) The reserved-category classifier retries once with a longer timeout on cold-model failure, then falls back to a deterministic keyword backstop; keyword-clean text is served locally. Over-length prompts (>4K chars) get the full-text keyword floor plus a head+middle+tail excerpt read and a distinct UNCERTAIN_LENGTH marker — prompt padding cannot force the ERROR branch. This eliminates the cold-start refusal problem without weakening the safety gate.
|
|
785
786
|
|
|
786
787
|
### Keyword Backstop for Reserved Content
|
|
787
788
|
When the LLM classifier fails (timeout, injection, resource pressure), a deterministic regex floor catches reserved vocabulary (restraint, seclusion, self-harm, suicide, overdose, crisis de-escalation, etc.) including inflected and verb forms. Blocks prompt-padding and classifier-injection attacks on the ERROR path.
|
|
@@ -1218,6 +1219,7 @@ All on-device models are free to run locally via Ollama on every tier. A subscri
|
|
|
1218
1219
|
| Cloud search | -- | ✅ | ✅ | ✅ |
|
|
1219
1220
|
| Max output tokens | 512 | 1,024 | 2,048 | 4,096 |
|
|
1220
1221
|
| Cloud fallback | -- | Gemini 3.6 Flash | Gemini 3.6 Flash | Gemini 3.6 Flash (priority) |
|
|
1222
|
+
| Multi-turn `prism_infer` (conversation carried across calls) | -- | 12 turns / 32k chars | 20 turns / 64k chars | 30 turns / 96k chars |
|
|
1221
1223
|
| Grounding verifier (fact-check AI output) | -- | ✅ | ✅ | ✅ |
|
|
1222
1224
|
| Memory sync (cloud) | -- | ✅ | ✅ | ✅ |
|
|
1223
1225
|
| Knowledge / session memory | limited | unlimited | unlimited | unlimited |
|
|
@@ -1270,8 +1272,115 @@ prism_infer({
|
|
|
1270
1272
|
})
|
|
1271
1273
|
// → 27B generates code locally ($0), with thinking for quality
|
|
1272
1274
|
// → If quality gate fails + paid tier → auto-escalate to Gemini 3.6 Flash
|
|
1275
|
+
|
|
1276
|
+
// Follow-ups carry the conversation (paid plans). The host curates the turns;
|
|
1277
|
+
// Prism bounds them to your plan's caps (user+assistant text only), safety-
|
|
1278
|
+
// screens every turn (each alone, then each in context; a short routine
|
|
1279
|
+
// request skips its own model read), counts them against
|
|
1280
|
+
// the tier's context, and never stores them. A free plan or a host with no
|
|
1281
|
+
// portal is refused: multi_turn_not_in_plan.
|
|
1282
|
+
prism_infer({
|
|
1283
|
+
messages: [
|
|
1284
|
+
{ role: "user", content: "Write a binary search in Python" },
|
|
1285
|
+
{ role: "assistant", content: "<the accepted answer>" },
|
|
1286
|
+
],
|
|
1287
|
+
prompt: "Now make it return the insertion point when the value is absent",
|
|
1288
|
+
mode: "code",
|
|
1289
|
+
})
|
|
1290
|
+
```
|
|
1291
|
+
|
|
1292
|
+
#### Multi-turn: why it matters, and why it is a paid feature
|
|
1293
|
+
|
|
1294
|
+
A single `prism_infer` call has no memory. The host asks a question, the local
|
|
1295
|
+
model answers, and the next call starts from nothing: "now add a timeout to it"
|
|
1296
|
+
means nothing to a model that never saw "it". Without history the host either
|
|
1297
|
+
restates the whole context in every prompt (tokens, and the answer drifts) or
|
|
1298
|
+
gives up on local delegation and does the follow-up itself in the cloud. With
|
|
1299
|
+
`messages`, the host hands back the turns it accepted and the local model
|
|
1300
|
+
continues the thread at $0, which is what makes local delegation useful for
|
|
1301
|
+
real work instead of one-shot snippets.
|
|
1302
|
+
|
|
1303
|
+
It is paid because it cannot run without Synalux behind it:
|
|
1304
|
+
|
|
1305
|
+
- **The caps are served by the portal, per plan.** How many turns and how many
|
|
1306
|
+
characters a conversation may carry is an entitlement your plan returns,
|
|
1307
|
+
not a number the client decides. No portal, no policy, and the call is
|
|
1308
|
+
refused as `multi_turn_not_in_plan`.
|
|
1309
|
+
- **Ambiguous turns go to Synalux cloud.** Every turn is safety-screened on
|
|
1310
|
+
device, alone and in context. A turn the screen finds uncertain is never
|
|
1311
|
+
served by the local model; it goes to the cloud when the plan allows it,
|
|
1312
|
+
and is refused otherwise. Free has no cloud, so the only honest answer for
|
|
1313
|
+
free is no history at all.
|
|
1314
|
+
- **Nothing is stored.** Prism bounds, screens, counts and forwards the turns
|
|
1315
|
+
the host sends, and keeps none of them.
|
|
1316
|
+
|
|
1317
|
+
<details>
|
|
1318
|
+
<summary>Without multi-turn (free plan, or no `messages`): every call starts cold</summary>
|
|
1319
|
+
|
|
1320
|
+
```typescript
|
|
1321
|
+
// Call 1
|
|
1322
|
+
prism_infer({ prompt: "My project codename is Nightjar. Reply OK.", mode: "chat" })
|
|
1323
|
+
// → "OK" (local 9b, $0)
|
|
1324
|
+
|
|
1325
|
+
// Call 2 — the model never saw call 1
|
|
1326
|
+
prism_infer({ prompt: "What is my codename? One word.", mode: "chat" })
|
|
1327
|
+
// → "I don't have that information." (local 9b, correct and useless)
|
|
1328
|
+
|
|
1329
|
+
// Call 3 — a coding follow-up with no thread
|
|
1330
|
+
prism_infer({ prompt: "Now add a timeout parameter to it.", mode: "code" })
|
|
1331
|
+
// → guesses what "it" is, or asks (the host has to redo the work)
|
|
1332
|
+
|
|
1333
|
+
// A free plan that sends messages anyway:
|
|
1334
|
+
prism_infer({ messages: [/* … */], prompt: "…" })
|
|
1335
|
+
// → refused: multi_turn_not_in_plan (no portal entitlement, no cloud)
|
|
1336
|
+
```
|
|
1337
|
+
</details>
|
|
1338
|
+
|
|
1339
|
+
<details>
|
|
1340
|
+
<summary>With multi-turn (paid plan): the thread continues locally, and the screen decides per turn</summary>
|
|
1341
|
+
|
|
1342
|
+
```typescript
|
|
1343
|
+
// The host keeps the turns it accepted and passes them back.
|
|
1344
|
+
prism_infer({
|
|
1345
|
+
messages: [
|
|
1346
|
+
{ role: "user", content: "My project codename is Nightjar. Reply OK." },
|
|
1347
|
+
{ role: "assistant", content: "OK" },
|
|
1348
|
+
],
|
|
1349
|
+
prompt: "What is my codename? One word.",
|
|
1350
|
+
mode: "chat",
|
|
1351
|
+
})
|
|
1352
|
+
// → "Nightjar" (local 9b, $0; history_turns: 2)
|
|
1353
|
+
|
|
1354
|
+
prism_infer({
|
|
1355
|
+
messages: [
|
|
1356
|
+
{ role: "user", content: "We named the helper countActiveUsers(data). Confirm." },
|
|
1357
|
+
{ role: "assistant", content: "Confirmed." },
|
|
1358
|
+
],
|
|
1359
|
+
prompt: "Write the one-line call that stores its result in n.",
|
|
1360
|
+
mode: "code",
|
|
1361
|
+
})
|
|
1362
|
+
// → "n = countActiveUsers(data)" (local 9b, $0)
|
|
1363
|
+
|
|
1364
|
+
// A turn the on-device screen finds uncertain when read alone is not served
|
|
1365
|
+
// locally: it goes to Synalux cloud on a paid plan, or is refused with
|
|
1366
|
+
// cloud_fallback: false. The result names the reason (layer1_uncertain) so
|
|
1367
|
+
// the host can decide what to do with the thread.
|
|
1368
|
+
prism_infer({
|
|
1369
|
+
messages: [
|
|
1370
|
+
{ role: "user", content: "The ticket for this bug is SYN-4471. Acknowledge." },
|
|
1371
|
+
{ role: "assistant", content: "Acknowledged." },
|
|
1372
|
+
],
|
|
1373
|
+
prompt: "Which ticket is this bug filed under?",
|
|
1374
|
+
cloud_fallback: true,
|
|
1375
|
+
})
|
|
1376
|
+
// → "SYN-4471" (Gemini 3.6 Flash; used_cloud: true)
|
|
1273
1377
|
```
|
|
1274
1378
|
|
|
1379
|
+
Measured on the real 9b through the real handler with cloud off: 7 of 12
|
|
1380
|
+
benign follow-ups are served locally, the rest refuse and name the reason;
|
|
1381
|
+
every injection variant the reviewers built refuses.
|
|
1382
|
+
</details>
|
|
1383
|
+
|
|
1275
1384
|
| Mode | Think | Model | Use case |
|
|
1276
1385
|
|------|-------|-------|----------|
|
|
1277
1386
|
| `route` | Off (fast) — except a tier that reasons better, e.g. 9B | 9B default | MCP tool routing |
|
package/dist/cli.js
CHANGED
|
@@ -1273,8 +1273,9 @@ program
|
|
|
1273
1273
|
.option('--dry-run', 'Print what would be pulled/re-aliased without executing')
|
|
1274
1274
|
.action(async (options) => {
|
|
1275
1275
|
const { runOllamaConverge } = await import('./modelConvergeRunner.js');
|
|
1276
|
+
const { convergeFailed } = await import('./utils/modelConverge.js');
|
|
1276
1277
|
const outcomes = await runOllamaConverge({ dryRun: options.dryRun === true });
|
|
1277
|
-
if (outcomes
|
|
1278
|
+
if (convergeFailed(outcomes))
|
|
1278
1279
|
process.exitCode = 1;
|
|
1279
1280
|
});
|
|
1280
1281
|
// ─── prism register-models ────────────────────────────────────
|
package/dist/localFirstPolicy.js
CHANGED
|
@@ -10,6 +10,8 @@ export const LOCAL_FIRST_POLICY_LINES = [
|
|
|
10
10
|
"For bounded, verifiable delegated work, use Prism's local worker before any host-native or background subagent.",
|
|
11
11
|
"When session_task_route is available, call it first. If it returns target `claw`, call `prism_infer` with the",
|
|
12
12
|
"recommended arguments, `cloud_fallback: false`, and the `project` and `conversation_id` from this session when known.",
|
|
13
|
+
"If the route says `needs_history`, or the task follows up an earlier `prism_infer` answer, pass the accepted prior",
|
|
14
|
+
"turns as `messages` (paid plans); never send a follow-up without them, and never re-send a turn you rejected.",
|
|
13
15
|
"Prism loads that project's dashboard-configured quick, standard, or deep memory and selects a RAM-safe local model.",
|
|
14
16
|
"Never create host-native or background subagents for routine work, never fan out, and never nest agents.",
|
|
15
17
|
"If local inference is unavailable, refused, degraded, or the task requires host tools or reserved judgment, continue",
|
|
@@ -81,6 +81,20 @@ export async function runOllamaConverge(opts = {}) {
|
|
|
81
81
|
const data = (await res.json());
|
|
82
82
|
return (data.models ?? []).map((m) => ({ name: m.name, digest: m.digest }));
|
|
83
83
|
},
|
|
84
|
+
tagFacts: async (name) => {
|
|
85
|
+
const res = await fetch(`${OLLAMA_URL}/api/show`, {
|
|
86
|
+
method: "POST", headers: { "Content-Type": "application/json" },
|
|
87
|
+
body: JSON.stringify({ model: name }), signal: AbortSignal.timeout(5_000),
|
|
88
|
+
});
|
|
89
|
+
if (!res.ok)
|
|
90
|
+
return null;
|
|
91
|
+
const data = (await res.json());
|
|
92
|
+
const from = /^FROM\s+(\S+)/m.exec(data.modelfile ?? "")?.[1];
|
|
93
|
+
if (!from)
|
|
94
|
+
return null;
|
|
95
|
+
const pin = /^\s*num_ctx\s+(\d+)\s*$/m.exec(data.parameters ?? "");
|
|
96
|
+
return { weightsBlob: from, pinnedNumCtx: pin ? Number(pin[1]) : null };
|
|
97
|
+
},
|
|
84
98
|
pull: (ref) => runOllama(["pull", ref], true),
|
|
85
99
|
copy: (from, to) => runOllama(["cp", from, to], false),
|
|
86
100
|
log: (line) => console.log(` ${line}`),
|
|
@@ -25,7 +25,7 @@ import { buildVaultDirectory } from "../utils/vaultExporter.js";
|
|
|
25
25
|
* ═══════════════════════════════════════════════════════════════════
|
|
26
26
|
*/
|
|
27
27
|
import { debugLog } from "../utils/logger.js";
|
|
28
|
-
import { FREE_ENTITLEMENTS } from "../utils/entitlements.js";
|
|
28
|
+
import { FREE_ENTITLEMENTS, peekEntitlements, multiTurnPolicy } from "../utils/entitlements.js";
|
|
29
29
|
import { getStorage, activeStorageBackend } from "../storage/index.js";
|
|
30
30
|
import { toKeywordArray } from "../utils/keywordExtractor.js";
|
|
31
31
|
import { getLLMProvider } from "../utils/llm/factory.js";
|
|
@@ -550,6 +550,7 @@ async function buildNativeSystemReadyBlock(snapshot, depth) {
|
|
|
550
550
|
`> - 🧠 **Context depth:** ${depth}\n` +
|
|
551
551
|
`> - 🔄 **Skill sync:** ${SKILL_SYNC_STATUS_LABELS[snapshot.syncStatus]} · native materialization incomplete${conflictSuffix}` +
|
|
552
552
|
conflictWarning +
|
|
553
|
+
localWorkerLine() +
|
|
553
554
|
freeTierUpgradeLine(snapshot.tier));
|
|
554
555
|
}
|
|
555
556
|
if (snapshot.source === "tier-fallback") {
|
|
@@ -560,6 +561,7 @@ async function buildNativeSystemReadyBlock(snapshot, depth) {
|
|
|
560
561
|
`> - 🧠 **Context depth:** ${depth}\n` +
|
|
561
562
|
`> - 🔄 **Skill sync:** ${SKILL_SYNC_STATUS_LABELS[snapshot.syncStatus]} · no committed manifest${conflictSuffix}` +
|
|
562
563
|
conflictWarning +
|
|
564
|
+
localWorkerLine() +
|
|
563
565
|
freeTierUpgradeLine(snapshot.tier));
|
|
564
566
|
}
|
|
565
567
|
return wrap(`> **Prism System Ready**\n>` + undeliveredWarning + `\n` +
|
|
@@ -572,6 +574,7 @@ async function buildNativeSystemReadyBlock(snapshot, depth) {
|
|
|
572
574
|
`> - 🧠 **Context depth:** ${depth}\n` +
|
|
573
575
|
`> - 🔄 **Skill sync:** ${SKILL_SYNC_STATUS_LABELS[snapshot.syncStatus]} · committed manifest${conflictSuffix}` +
|
|
574
576
|
conflictWarning +
|
|
577
|
+
localWorkerLine() +
|
|
575
578
|
freeTierUpgradeLine(snapshot.tier));
|
|
576
579
|
}
|
|
577
580
|
/**
|
|
@@ -580,6 +583,22 @@ async function buildNativeSystemReadyBlock(snapshot, depth) {
|
|
|
580
583
|
* queryMemoryNaturalHandler); the startup path — the only guaranteed
|
|
581
584
|
* impression — referenced it zero times.
|
|
582
585
|
*/
|
|
586
|
+
/**
|
|
587
|
+
* One startup line so the host knows, before its first delegation, whether
|
|
588
|
+
* the local worker takes conversation history and how much. Reads the
|
|
589
|
+
* entitlements CACHE only — never a portal fetch on the startup path; when
|
|
590
|
+
* cold, the first prism_infer result carries the same policy.
|
|
591
|
+
*/
|
|
592
|
+
function localWorkerLine() {
|
|
593
|
+
const ent = peekEntitlements();
|
|
594
|
+
if (!ent)
|
|
595
|
+
return "";
|
|
596
|
+
const p = multiTurnPolicy(ent);
|
|
597
|
+
return p.enabled
|
|
598
|
+
? `\n> - 🧵 **Local worker multi-turn:** on — up to ${p.max_turns} turns / ` +
|
|
599
|
+
`${p.max_chars.toLocaleString("en-US")} chars per prism_infer call; pass accepted prior turns as \`messages\``
|
|
600
|
+
: `\n> - 🧵 **Local worker multi-turn:** off on the ${ent.plan} plan — a prism_infer follow-up is answered without context`;
|
|
601
|
+
}
|
|
583
602
|
function freeTierUpgradeLine(tier) {
|
|
584
603
|
if (tier !== "free")
|
|
585
604
|
return "";
|