prism-mcp-server 20.21.14 → 20.21.15
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md
CHANGED
|
@@ -158,6 +158,21 @@ or by re-enabling after each run.
|
|
|
158
158
|
<details>
|
|
159
159
|
<summary>Release history (optional)</summary>
|
|
160
160
|
|
|
161
|
+
## What's New in v20.21.15
|
|
162
|
+
|
|
163
|
+
### Follow-ups: see what runs locally, and a stricter screen when the classifier is down
|
|
164
|
+
|
|
165
|
+
- `local_savings` (and `prism savings`) shows a line for follow-ups, calls that
|
|
166
|
+
carried your conversation. It covers how many the local model answered (and
|
|
167
|
+
how many of those the 9b answered), how many the on-device screen refused and
|
|
168
|
+
at which stage, how many your plan's multi-turn limits refused, and how many
|
|
169
|
+
went to the cloud.
|
|
170
|
+
- A short chat answer ("16", "Yes") counts as an answer. Before, the quality
|
|
171
|
+
gate treated it as empty, and a paid plan re-asked the cloud.
|
|
172
|
+
- When the on-device classifier fails every read on a follow-up, the follow-up
|
|
173
|
+
is no longer answered on a keyword check alone. It goes to the cloud if your
|
|
174
|
+
plan allows it, and is refused otherwise.
|
|
175
|
+
|
|
161
176
|
## What's New in v20.21.14
|
|
162
177
|
|
|
163
178
|
### Skills load when they help, and you can see why they loaded
|
|
@@ -1445,11 +1460,11 @@ It is paid because it cannot run without Synalux behind it:
|
|
|
1445
1460
|
```typescript
|
|
1446
1461
|
// Call 1
|
|
1447
1462
|
prism_infer({ prompt: "My project codename is Nightjar. Reply OK.", mode: "chat" })
|
|
1448
|
-
// → "OK" (local
|
|
1463
|
+
// → "OK" (local model, $0)
|
|
1449
1464
|
|
|
1450
1465
|
// Call 2 — the model never saw call 1
|
|
1451
1466
|
prism_infer({ prompt: "What is my codename? One word.", mode: "chat" })
|
|
1452
|
-
// → "I don't have that information." (local
|
|
1467
|
+
// → "I don't have that information." (local model, correct and useless)
|
|
1453
1468
|
|
|
1454
1469
|
// Call 3 — a coding follow-up with no thread
|
|
1455
1470
|
prism_infer({ prompt: "Now add a timeout parameter to it.", mode: "code" })
|
|
@@ -1474,7 +1489,7 @@ prism_infer({
|
|
|
1474
1489
|
prompt: "What is my codename? One word.",
|
|
1475
1490
|
mode: "chat",
|
|
1476
1491
|
})
|
|
1477
|
-
// → "Nightjar" (local
|
|
1492
|
+
// → "Nightjar" (local model, $0; history_turns: 2)
|
|
1478
1493
|
|
|
1479
1494
|
prism_infer({
|
|
1480
1495
|
messages: [
|
|
@@ -1484,7 +1499,7 @@ prism_infer({
|
|
|
1484
1499
|
prompt: "Write the one-line call that stores its result in n.",
|
|
1485
1500
|
mode: "code",
|
|
1486
1501
|
})
|
|
1487
|
-
// → "n = countActiveUsers(data)" (local
|
|
1502
|
+
// → "n = countActiveUsers(data)" (local model, $0)
|
|
1488
1503
|
|
|
1489
1504
|
// A turn the on-device screen finds uncertain, alone or in context, is not
|
|
1490
1505
|
// served locally: it goes to Synalux cloud on a paid plan, or is refused with
|
|
@@ -1501,9 +1516,14 @@ prism_infer({
|
|
|
1501
1516
|
// → "SYN-4471" (Gemini 3.6 Flash; used_cloud: true)
|
|
1502
1517
|
```
|
|
1503
1518
|
|
|
1504
|
-
Measured on the real
|
|
1505
|
-
benign follow-
|
|
1506
|
-
|
|
1519
|
+
Measured on 20.21.14 through the real handler, on the local 9b with cloud off:
|
|
1520
|
+
- Every benign follow-up that reached the local 9b with its conversation
|
|
1521
|
+
attached was answered correctly. Without the conversation, most answers were
|
|
1522
|
+
invented.
|
|
1523
|
+
- The on-device screen still refuses too many benign follow-ups: 11 of 26 in
|
|
1524
|
+
our tests. Each refusal names its reason. Cutting these false refusals is
|
|
1525
|
+
the current work.
|
|
1526
|
+
- Every reserved multi-turn probe was refused before any generation (11 of 11).
|
|
1507
1527
|
</details>
|
|
1508
1528
|
|
|
1509
1529
|
| Mode | Think | Model | Use case |
|
|
@@ -1556,6 +1576,10 @@ host, or run `prism savings` from a terminal — `--period all|month|week|sessio
|
|
|
1556
1576
|
prism-coder:9b: 41 call(s), ~505K tokens
|
|
1557
1577
|
prism-coder:4b: 12 call(s), ~4.8K tokens
|
|
1558
1578
|
|
|
1579
|
+
Follow-ups with your conversation:
|
|
1580
|
+
12 answered locally (12 by the 9b) · 2 refused by the on-device screen · 0 sent to cloud
|
|
1581
|
+
Refusals by stage: follow-up alone 1 · turns together 1
|
|
1582
|
+
|
|
1559
1583
|
Counts tokens a local model handled instead of your cloud model. On the token
|
|
1560
1584
|
axis, the token count is measured — a floor, with known undercounts listed
|
|
1561
1585
|
when present. On the displacement axis, prism cannot observe the call your
|
|
@@ -1577,6 +1601,21 @@ the same tokens; and most users are on flat plans where a currency figure means
|
|
|
1577
1601
|
nothing at all. Tokens are the one unit prism measured itself. If you know your
|
|
1578
1602
|
own effective rate, multiply — the split is printed for exactly that reason.
|
|
1579
1603
|
|
|
1604
|
+
**Follow-ups.** Calls that carried your conversation (`messages`) get their
|
|
1605
|
+
own line:
|
|
1606
|
+
- how many the local model answered, and how many of those the 9b answered;
|
|
1607
|
+
- how many the on-device screen refused, and at which stage (the follow-up read
|
|
1608
|
+
alone, an earlier turn read alone, or the turns read together);
|
|
1609
|
+
- how many went to the cloud.
|
|
1610
|
+
|
|
1611
|
+
A refused follow-up went back to your host instead of being answered locally,
|
|
1612
|
+
so this line shows how much of your follow-up work local serving actually
|
|
1613
|
+
took. Refusals because your plan does not include multi-turn history, or the
|
|
1614
|
+
history is over its cap, are counted on their own line. They are recorded only
|
|
1615
|
+
when the host asks for a report (`escalation: "report"`); otherwise they fail
|
|
1616
|
+
before anything is recorded. The line appears for the `week`, `month`, `all`
|
|
1617
|
+
and `--days` views, which read the durable ledger.
|
|
1618
|
+
|
|
1580
1619
|
Refused calls are excluded, the VS Code panel-playground share is disclosed
|
|
1581
1620
|
separately, and the known sources of undercount are listed inline rather than
|
|
1582
1621
|
left implicit — so the durable (`week`/`month`/`all`/`--days`) headline is a
|
|
@@ -284,6 +284,32 @@ export async function queryLocalSavings(sinceTs) {
|
|
|
284
284
|
GROUP BY COALESCE(model, backend)`,
|
|
285
285
|
args: whereArgs,
|
|
286
286
|
});
|
|
287
|
+
const followWhere = `WHERE history_turns > 0${sinceTs != null ? " AND ts >= ?" : ""}`;
|
|
288
|
+
const PLAN_REFUSAL = `COALESCE(refusal_reason, '') IN ('multi_turn_not_in_plan', 'history_over_plan_cap')`;
|
|
289
|
+
// A 9b model: '9b' right after a non-digit, so a future '19b' is not counted.
|
|
290
|
+
const NINE_B = `(LOWER(COALESCE(model, backend)) GLOB '*[^0-9]9b*' OR LOWER(COALESCE(model, backend)) GLOB '9b*')`;
|
|
291
|
+
const follow = await client.execute({
|
|
292
|
+
sql: `SELECT
|
|
293
|
+
SUM(CASE WHEN ${SERVED_LOCAL} THEN 1 ELSE 0 END) AS served,
|
|
294
|
+
SUM(CASE WHEN ${SERVED_LOCAL} AND ${NINE_B} THEN 1 ELSE 0 END) AS served_9b,
|
|
295
|
+
SUM(CASE WHEN used_cloud = 1 THEN 1 ELSE 0 END) AS cloud,
|
|
296
|
+
SUM(CASE WHEN used_cloud = 0 AND NOT (${SERVED_LOCAL}) AND NOT (${PLAN_REFUSAL}) THEN 1 ELSE 0 END) AS refused,
|
|
297
|
+
SUM(CASE WHEN used_cloud = 0 AND NOT (${SERVED_LOCAL}) AND ${PLAN_REFUSAL} THEN 1 ELSE 0 END) AS refused_by_plan
|
|
298
|
+
FROM infer_metrics ${followWhere}`,
|
|
299
|
+
args: whereArgs,
|
|
300
|
+
});
|
|
301
|
+
const followLayers = await client.execute({
|
|
302
|
+
sql: `SELECT COALESCE(refusal_layer, 'unrecorded') AS layer, COUNT(*) AS n
|
|
303
|
+
FROM infer_metrics ${followWhere} AND used_cloud = 0 AND NOT (${SERVED_LOCAL})
|
|
304
|
+
AND NOT (${PLAN_REFUSAL})
|
|
305
|
+
GROUP BY COALESCE(refusal_layer, 'unrecorded')`,
|
|
306
|
+
args: whereArgs,
|
|
307
|
+
});
|
|
308
|
+
const f = follow.rows[0];
|
|
309
|
+
const refused_by_layer = {};
|
|
310
|
+
for (const row of followLayers.rows) {
|
|
311
|
+
refused_by_layer[String(row.layer)] = Number(row.n ?? 0);
|
|
312
|
+
}
|
|
287
313
|
const r = agg.rows[0];
|
|
288
314
|
const by_model = {};
|
|
289
315
|
for (const row of byM.rows) {
|
|
@@ -310,6 +336,14 @@ export async function queryLocalSavings(sinceTs) {
|
|
|
310
336
|
first_ts: r.first_ts == null ? null : Number(r.first_ts),
|
|
311
337
|
last_ts: r.last_ts == null ? null : Number(r.last_ts),
|
|
312
338
|
by_model,
|
|
339
|
+
followups: {
|
|
340
|
+
served_local: Number(f.served ?? 0),
|
|
341
|
+
served_local_9b: Number(f.served_9b ?? 0),
|
|
342
|
+
refused: Number(f.refused ?? 0),
|
|
343
|
+
refused_by_plan: Number(f.refused_by_plan ?? 0),
|
|
344
|
+
cloud: Number(f.cloud ?? 0),
|
|
345
|
+
refused_by_layer,
|
|
346
|
+
},
|
|
313
347
|
};
|
|
314
348
|
}
|
|
315
349
|
catch (e) {
|
|
@@ -1744,6 +1744,17 @@ export async function runInfer(args, deps) {
|
|
|
1744
1744
|
break;
|
|
1745
1745
|
}
|
|
1746
1746
|
}
|
|
1747
|
+
// A classifier that answered none of this request's window reads
|
|
1748
|
+
// read none of them: the keyword net must not become their sole
|
|
1749
|
+
// guard (review round 16). The consecutive-ERROR breaker catches
|
|
1750
|
+
// this only from its third read, and a follow-up that re-sends the
|
|
1751
|
+
// same turns leaves one or two uncached windows (measured
|
|
1752
|
+
// 2026-09-25: a dead classifier, two failed reads, served locally
|
|
1753
|
+
// on the keyword net).
|
|
1754
|
+
if (budget.calls > 0 && budget.consecutiveErrors === budget.calls && !budget.tripped) {
|
|
1755
|
+
budget.tripped = true;
|
|
1756
|
+
attempts.push({ tier: "layer1", reason: "layer1_screen_all_reads_failed" });
|
|
1757
|
+
}
|
|
1747
1758
|
// A budget or breaker trip raises to UNCERTAIN whatever the cache
|
|
1748
1759
|
// held (text: cloud or refused; with an image: local only).
|
|
1749
1760
|
if (budget.tripped)
|
|
@@ -129,6 +129,44 @@ function basisLine(s) {
|
|
|
129
129
|
"your host would have made, so whether all of it would have hit the cloud is an assumption. " +
|
|
130
130
|
`Read it as: at most this much displacement, of ${volumeWord} this token volume.`;
|
|
131
131
|
}
|
|
132
|
+
/** Plain names for the on-device screen's stages, as recorded in refusal_layer. */
|
|
133
|
+
const FOLLOWUP_STAGE_NAMES = {
|
|
134
|
+
isolated: "earlier turn alone",
|
|
135
|
+
prompt: "follow-up alone",
|
|
136
|
+
context: "turns together",
|
|
137
|
+
rules: "rules",
|
|
138
|
+
backstop: "keyword net",
|
|
139
|
+
budget: "screening limit",
|
|
140
|
+
};
|
|
141
|
+
/**
|
|
142
|
+
* Follow-ups: calls that carried the conversation. Shown even when none was
|
|
143
|
+
* served, because refusals are the part a user needs to see: a follow-up the
|
|
144
|
+
* screen refused went back to the host instead of being answered locally.
|
|
145
|
+
*/
|
|
146
|
+
export function followupLines(f) {
|
|
147
|
+
if (!f)
|
|
148
|
+
return [];
|
|
149
|
+
const total = f.served_local + f.refused + f.refused_by_plan + f.cloud;
|
|
150
|
+
if (total === 0)
|
|
151
|
+
return [];
|
|
152
|
+
const nineB = f.served_local > 0 ? ` (${fmt(f.served_local_9b)} by the 9b)` : "";
|
|
153
|
+
const lines = [
|
|
154
|
+
"",
|
|
155
|
+
" Follow-ups with your conversation:",
|
|
156
|
+
` ${fmt(f.served_local)} answered locally${nineB} · ${fmt(f.refused)} refused by the on-device screen · ${fmt(f.cloud)} sent to cloud`,
|
|
157
|
+
];
|
|
158
|
+
if (f.refused_by_plan > 0) {
|
|
159
|
+
lines.push(` ${fmt(f.refused_by_plan)} refused by your plan's multi-turn limits`);
|
|
160
|
+
}
|
|
161
|
+
// Most refusals first; ties in a fixed stage order, so the line never depends on SQL row order.
|
|
162
|
+
const order = (k) => { const i = Object.keys(FOLLOWUP_STAGE_NAMES).indexOf(k); return i < 0 ? 99 : i; };
|
|
163
|
+
const stages = Object.entries(f.refused_by_layer).filter(([, n]) => n > 0)
|
|
164
|
+
.sort((a, b) => b[1] - a[1] || order(a[0]) - order(b[0]) || a[0].localeCompare(b[0]));
|
|
165
|
+
if (stages.length > 0) {
|
|
166
|
+
lines.push(` Refusals by stage: ${stages.map(([k, n]) => `${FOLLOWUP_STAGE_NAMES[k] ?? k} ${fmt(n)}`).join(" · ")}`);
|
|
167
|
+
}
|
|
168
|
+
return lines;
|
|
169
|
+
}
|
|
132
170
|
export function renderSavings(s, period, customDays) {
|
|
133
171
|
const label = customDays !== undefined && Number.isFinite(customDays) && customDays > 0
|
|
134
172
|
? `LAST ${Math.floor(customDays)} DAYS`
|
|
@@ -144,6 +182,7 @@ export function renderSavings(s, period, customDays) {
|
|
|
144
182
|
lines.push(period === "session"
|
|
145
183
|
? " Delegate work with prism_infer, or use session_task_route to pick targets automatically."
|
|
146
184
|
: " Once prism starts serving locally, displaced token volume shows up here.");
|
|
185
|
+
lines.push(...followupLines(s.followups));
|
|
147
186
|
return { text: lines.join("\n"), data: { ...s, period } };
|
|
148
187
|
}
|
|
149
188
|
const totalRouted = s.local_calls + s.cloud_calls;
|
|
@@ -162,6 +201,7 @@ export function renderSavings(s, period, customDays) {
|
|
|
162
201
|
lines.push(` ${name}: ${fmt(m.calls)} call(s), ~${abbreviate(t)} tokens`);
|
|
163
202
|
}
|
|
164
203
|
}
|
|
204
|
+
lines.push(...followupLines(s.followups));
|
|
165
205
|
lines.push("");
|
|
166
206
|
lines.push(` ${basisLine(s)}`);
|
|
167
207
|
const caveats = caveatsFor(s);
|
|
@@ -10,7 +10,7 @@ export const TOOL_CALL_BLEED_RE = /<\|tool_call\|>|<\|tool_call_end\|>/;
|
|
|
10
10
|
* @param stripped Response AFTER think-stripping (use stripThink first)
|
|
11
11
|
* @param thinkOnly True if the response was only <think> blocks with no answer
|
|
12
12
|
* @param finishReason Ollama's finish_reason if available (e.g. "length" = truncated)
|
|
13
|
-
* @param mode Inference mode — "route"
|
|
13
|
+
* @param mode Inference mode — "route": empty only when blank; "chat": empty only with no letter or digit; "code"/unset: 4 chars or fewer
|
|
14
14
|
*/
|
|
15
15
|
export function passesQualityGate(stripped, thinkOnly, finishReason, mode) {
|
|
16
16
|
// Signal 1: Think-only — model reasoned but produced no answer (check before empty)
|
|
@@ -19,9 +19,17 @@ export function passesQualityGate(stripped, thinkOnly, finishReason, mode) {
|
|
|
19
19
|
}
|
|
20
20
|
// Signal 2: Mode-aware empty floor.
|
|
21
21
|
// Route legitimately returns 1–4 char labels ("P1", "YES", "CO4", "FIXED").
|
|
22
|
-
//
|
|
23
|
-
|
|
24
|
-
|
|
22
|
+
// Chat answers can be one short value: a follow-up asking "what is x times
|
|
23
|
+
// 6?" is correctly answered "42". Measured 2026-09-24: under the old <5
|
|
24
|
+
// floor, correct chat answers "16" and "36" failed here and, on a paid
|
|
25
|
+
// plan, were thrown away and re-asked of the cloud. So chat is empty only
|
|
26
|
+
// when it has no letter or digit at all.
|
|
27
|
+
// Code keeps <5: a 1–4 char code answer ("Hi", "DONE") is not an answer.
|
|
28
|
+
const trimmed = stripped.trim();
|
|
29
|
+
const empty = mode === "route" ? trimmed.length === 0
|
|
30
|
+
: mode === "chat" ? !/[\p{L}\p{N}]/u.test(trimmed)
|
|
31
|
+
: trimmed.length <= 4;
|
|
32
|
+
if (empty) {
|
|
25
33
|
return { pass: false, reason: "empty_response" };
|
|
26
34
|
}
|
|
27
35
|
// Signal 3: Hard truncation — Ollama reports finish_reason="length"
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "prism-mcp-server",
|
|
3
|
-
"version": "20.21.
|
|
3
|
+
"version": "20.21.15",
|
|
4
4
|
"mcpName": "io.github.dcostenco/prism-coder",
|
|
5
5
|
"description": "Persistent session memory for AI coding agents that never leaves your machine — including the on-device model that reasons over it. Restores your prior decisions, open TODOs, and changed files across sessions; adds associative recall of related past work, semantic drift detection, and local inference. Local-first by default. Works with Claude Code, Cursor, and Codex.",
|
|
6
6
|
"module": "index.ts",
|