prism-mcp-server 20.21.14 → 20.21.15

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -158,6 +158,21 @@ or by re-enabling after each run.
158
158
  <details>
159
159
  <summary>Release history (optional)</summary>
160
160
 
161
+ ## What's New in v20.21.15
162
+
163
+ ### Follow-ups: see what runs locally, and a stricter screen when the classifier is down
164
+
165
+ - `local_savings` (and `prism savings`) shows a line for follow-ups, calls that
166
+ carried your conversation. It covers how many the local model answered (and
167
+ how many of those the 9b answered), how many the on-device screen refused and
168
+ at which stage, how many your plan's multi-turn limits refused, and how many
169
+ went to the cloud.
170
+ - A short chat answer ("16", "Yes") counts as an answer. Before, the quality
171
+ gate treated it as empty, and a paid plan re-asked the cloud.
172
+ - When the on-device classifier fails every read on a follow-up, the follow-up
173
+ is no longer answered on a keyword check alone. It goes to the cloud if your
174
+ plan allows it, and is refused otherwise.
175
+
161
176
  ## What's New in v20.21.14
162
177
 
163
178
  ### Skills load when they help, and you can see why they loaded
@@ -1445,11 +1460,11 @@ It is paid because it cannot run without Synalux behind it:
1445
1460
  ```typescript
1446
1461
  // Call 1
1447
1462
  prism_infer({ prompt: "My project codename is Nightjar. Reply OK.", mode: "chat" })
1448
- // → "OK" (local 9b, $0)
1463
+ // → "OK" (local model, $0)
1449
1464
 
1450
1465
  // Call 2 — the model never saw call 1
1451
1466
  prism_infer({ prompt: "What is my codename? One word.", mode: "chat" })
1452
- // → "I don't have that information." (local 9b, correct and useless)
1467
+ // → "I don't have that information." (local model, correct and useless)
1453
1468
 
1454
1469
  // Call 3 — a coding follow-up with no thread
1455
1470
  prism_infer({ prompt: "Now add a timeout parameter to it.", mode: "code" })
@@ -1474,7 +1489,7 @@ prism_infer({
1474
1489
  prompt: "What is my codename? One word.",
1475
1490
  mode: "chat",
1476
1491
  })
1477
- // → "Nightjar" (local 9b, $0; history_turns: 2)
1492
+ // → "Nightjar" (local model, $0; history_turns: 2)
1478
1493
 
1479
1494
  prism_infer({
1480
1495
  messages: [
@@ -1484,7 +1499,7 @@ prism_infer({
1484
1499
  prompt: "Write the one-line call that stores its result in n.",
1485
1500
  mode: "code",
1486
1501
  })
1487
- // → "n = countActiveUsers(data)" (local 9b, $0)
1502
+ // → "n = countActiveUsers(data)" (local model, $0)
1488
1503
 
1489
1504
  // A turn the on-device screen finds uncertain, alone or in context, is not
1490
1505
  // served locally: it goes to Synalux cloud on a paid plan, or is refused with
@@ -1501,9 +1516,14 @@ prism_infer({
1501
1516
  // → "SYN-4471" (Gemini 3.6 Flash; used_cloud: true)
1502
1517
  ```
1503
1518
 
1504
- Measured on the real 9b through the real handler with cloud off: 7 of 12
1505
- benign follow-ups are served locally, the rest refuse and name the reason;
1506
- every injection variant the reviewers built refuses.
1519
+ Measured on 20.21.14 through the real handler, on the local 9b with cloud off:
1520
+ - Every benign follow-up that reached the local 9b with its conversation
1521
+ attached was answered correctly. Without the conversation, most answers were
1522
+ invented.
1523
+ - The on-device screen still refuses too many benign follow-ups: 11 of 26 in
1524
+ our tests. Each refusal names its reason. Cutting these false refusals is
1525
+ the current work.
1526
+ - Every reserved multi-turn probe was refused before any generation (11 of 11).
1507
1527
  </details>
1508
1528
 
1509
1529
  | Mode | Think | Model | Use case |
@@ -1556,6 +1576,10 @@ host, or run `prism savings` from a terminal — `--period all|month|week|sessio
1556
1576
  prism-coder:9b: 41 call(s), ~505K tokens
1557
1577
  prism-coder:4b: 12 call(s), ~4.8K tokens
1558
1578
 
1579
+ Follow-ups with your conversation:
1580
+ 12 answered locally (12 by the 9b) · 2 refused by the on-device screen · 0 sent to cloud
1581
+ Refusals by stage: follow-up alone 1 · turns together 1
1582
+
1559
1583
  Counts tokens a local model handled instead of your cloud model. On the token
1560
1584
  axis, the token count is measured — a floor, with known undercounts listed
1561
1585
  when present. On the displacement axis, prism cannot observe the call your
@@ -1577,6 +1601,21 @@ the same tokens; and most users are on flat plans where a currency figure means
1577
1601
  nothing at all. Tokens are the one unit prism measured itself. If you know your
1578
1602
  own effective rate, multiply — the split is printed for exactly that reason.
1579
1603
 
1604
+ **Follow-ups.** Calls that carried your conversation (`messages`) get their
1605
+ own line:
1606
+ - how many the local model answered, and how many of those the 9b answered;
1607
+ - how many the on-device screen refused, and at which stage (the follow-up read
1608
+ alone, an earlier turn read alone, or the turns read together);
1609
+ - how many went to the cloud.
1610
+
1611
+ A refused follow-up went back to your host instead of being answered locally,
1612
+ so this line shows how much of your follow-up work local serving actually
1613
+ took. Refusals because your plan does not include multi-turn history, or the
1614
+ history is over its cap, are counted on their own line. They are recorded only
1615
+ when the host asks for a report (`escalation: "report"`); otherwise they fail
1616
+ before anything is recorded. The line appears for the `week`, `month`, `all`
1617
+ and `--days` views, which read the durable ledger.
1618
+
1580
1619
  Refused calls are excluded, the VS Code panel-playground share is disclosed
1581
1620
  separately, and the known sources of undercount are listed inline rather than
1582
1621
  left implicit — so the durable (`week`/`month`/`all`/`--days`) headline is a
@@ -284,6 +284,32 @@ export async function queryLocalSavings(sinceTs) {
284
284
  GROUP BY COALESCE(model, backend)`,
285
285
  args: whereArgs,
286
286
  });
287
+ const followWhere = `WHERE history_turns > 0${sinceTs != null ? " AND ts >= ?" : ""}`;
288
+ const PLAN_REFUSAL = `COALESCE(refusal_reason, '') IN ('multi_turn_not_in_plan', 'history_over_plan_cap')`;
289
+ // A 9b model: '9b' right after a non-digit, so a future '19b' is not counted.
290
+ const NINE_B = `(LOWER(COALESCE(model, backend)) GLOB '*[^0-9]9b*' OR LOWER(COALESCE(model, backend)) GLOB '9b*')`;
291
+ const follow = await client.execute({
292
+ sql: `SELECT
293
+ SUM(CASE WHEN ${SERVED_LOCAL} THEN 1 ELSE 0 END) AS served,
294
+ SUM(CASE WHEN ${SERVED_LOCAL} AND ${NINE_B} THEN 1 ELSE 0 END) AS served_9b,
295
+ SUM(CASE WHEN used_cloud = 1 THEN 1 ELSE 0 END) AS cloud,
296
+ SUM(CASE WHEN used_cloud = 0 AND NOT (${SERVED_LOCAL}) AND NOT (${PLAN_REFUSAL}) THEN 1 ELSE 0 END) AS refused,
297
+ SUM(CASE WHEN used_cloud = 0 AND NOT (${SERVED_LOCAL}) AND ${PLAN_REFUSAL} THEN 1 ELSE 0 END) AS refused_by_plan
298
+ FROM infer_metrics ${followWhere}`,
299
+ args: whereArgs,
300
+ });
301
+ const followLayers = await client.execute({
302
+ sql: `SELECT COALESCE(refusal_layer, 'unrecorded') AS layer, COUNT(*) AS n
303
+ FROM infer_metrics ${followWhere} AND used_cloud = 0 AND NOT (${SERVED_LOCAL})
304
+ AND NOT (${PLAN_REFUSAL})
305
+ GROUP BY COALESCE(refusal_layer, 'unrecorded')`,
306
+ args: whereArgs,
307
+ });
308
+ const f = follow.rows[0];
309
+ const refused_by_layer = {};
310
+ for (const row of followLayers.rows) {
311
+ refused_by_layer[String(row.layer)] = Number(row.n ?? 0);
312
+ }
287
313
  const r = agg.rows[0];
288
314
  const by_model = {};
289
315
  for (const row of byM.rows) {
@@ -310,6 +336,14 @@ export async function queryLocalSavings(sinceTs) {
310
336
  first_ts: r.first_ts == null ? null : Number(r.first_ts),
311
337
  last_ts: r.last_ts == null ? null : Number(r.last_ts),
312
338
  by_model,
339
+ followups: {
340
+ served_local: Number(f.served ?? 0),
341
+ served_local_9b: Number(f.served_9b ?? 0),
342
+ refused: Number(f.refused ?? 0),
343
+ refused_by_plan: Number(f.refused_by_plan ?? 0),
344
+ cloud: Number(f.cloud ?? 0),
345
+ refused_by_layer,
346
+ },
313
347
  };
314
348
  }
315
349
  catch (e) {
@@ -1744,6 +1744,17 @@ export async function runInfer(args, deps) {
1744
1744
  break;
1745
1745
  }
1746
1746
  }
1747
+ // A classifier that answered none of this request's window reads
1748
+ // read none of them: the keyword net must not become their sole
1749
+ // guard (review round 16). The consecutive-ERROR breaker catches
1750
+ // this only from its third read, and a follow-up that re-sends the
1751
+ // same turns leaves one or two uncached windows (measured
1752
+ // 2026-09-25: a dead classifier, two failed reads, served locally
1753
+ // on the keyword net).
1754
+ if (budget.calls > 0 && budget.consecutiveErrors === budget.calls && !budget.tripped) {
1755
+ budget.tripped = true;
1756
+ attempts.push({ tier: "layer1", reason: "layer1_screen_all_reads_failed" });
1757
+ }
1747
1758
  // A budget or breaker trip raises to UNCERTAIN whatever the cache
1748
1759
  // held (text: cloud or refused; with an image: local only).
1749
1760
  if (budget.tripped)
@@ -129,6 +129,44 @@ function basisLine(s) {
129
129
  "your host would have made, so whether all of it would have hit the cloud is an assumption. " +
130
130
  `Read it as: at most this much displacement, of ${volumeWord} this token volume.`;
131
131
  }
132
+ /** Plain names for the on-device screen's stages, as recorded in refusal_layer. */
133
+ const FOLLOWUP_STAGE_NAMES = {
134
+ isolated: "earlier turn alone",
135
+ prompt: "follow-up alone",
136
+ context: "turns together",
137
+ rules: "rules",
138
+ backstop: "keyword net",
139
+ budget: "screening limit",
140
+ };
141
+ /**
142
+ * Follow-ups: calls that carried the conversation. Shown even when none was
143
+ * served, because refusals are the part a user needs to see: a follow-up the
144
+ * screen refused went back to the host instead of being answered locally.
145
+ */
146
+ export function followupLines(f) {
147
+ if (!f)
148
+ return [];
149
+ const total = f.served_local + f.refused + f.refused_by_plan + f.cloud;
150
+ if (total === 0)
151
+ return [];
152
+ const nineB = f.served_local > 0 ? ` (${fmt(f.served_local_9b)} by the 9b)` : "";
153
+ const lines = [
154
+ "",
155
+ " Follow-ups with your conversation:",
156
+ ` ${fmt(f.served_local)} answered locally${nineB} · ${fmt(f.refused)} refused by the on-device screen · ${fmt(f.cloud)} sent to cloud`,
157
+ ];
158
+ if (f.refused_by_plan > 0) {
159
+ lines.push(` ${fmt(f.refused_by_plan)} refused by your plan's multi-turn limits`);
160
+ }
161
+ // Most refusals first; ties in a fixed stage order, so the line never depends on SQL row order.
162
+ const order = (k) => { const i = Object.keys(FOLLOWUP_STAGE_NAMES).indexOf(k); return i < 0 ? 99 : i; };
163
+ const stages = Object.entries(f.refused_by_layer).filter(([, n]) => n > 0)
164
+ .sort((a, b) => b[1] - a[1] || order(a[0]) - order(b[0]) || a[0].localeCompare(b[0]));
165
+ if (stages.length > 0) {
166
+ lines.push(` Refusals by stage: ${stages.map(([k, n]) => `${FOLLOWUP_STAGE_NAMES[k] ?? k} ${fmt(n)}`).join(" · ")}`);
167
+ }
168
+ return lines;
169
+ }
132
170
  export function renderSavings(s, period, customDays) {
133
171
  const label = customDays !== undefined && Number.isFinite(customDays) && customDays > 0
134
172
  ? `LAST ${Math.floor(customDays)} DAYS`
@@ -144,6 +182,7 @@ export function renderSavings(s, period, customDays) {
144
182
  lines.push(period === "session"
145
183
  ? " Delegate work with prism_infer, or use session_task_route to pick targets automatically."
146
184
  : " Once prism starts serving locally, displaced token volume shows up here.");
185
+ lines.push(...followupLines(s.followups));
147
186
  return { text: lines.join("\n"), data: { ...s, period } };
148
187
  }
149
188
  const totalRouted = s.local_calls + s.cloud_calls;
@@ -162,6 +201,7 @@ export function renderSavings(s, period, customDays) {
162
201
  lines.push(` ${name}: ${fmt(m.calls)} call(s), ~${abbreviate(t)} tokens`);
163
202
  }
164
203
  }
204
+ lines.push(...followupLines(s.followups));
165
205
  lines.push("");
166
206
  lines.push(` ${basisLine(s)}`);
167
207
  const caveats = caveatsFor(s);
@@ -10,7 +10,7 @@ export const TOOL_CALL_BLEED_RE = /<\|tool_call\|>|<\|tool_call_end\|>/;
10
10
  * @param stripped Response AFTER think-stripping (use stripThink first)
11
11
  * @param thinkOnly True if the response was only <think> blocks with no answer
12
12
  * @param finishReason Ollama's finish_reason if available (e.g. "length" = truncated)
13
- * @param mode Inference mode — "route" uses length===0 floor; "code"/"chat" keep <5
13
+ * @param mode Inference mode — "route": empty only when blank; "chat": empty only with no letter or digit; "code"/unset: 4 chars or fewer
14
14
  */
15
15
  export function passesQualityGate(stripped, thinkOnly, finishReason, mode) {
16
16
  // Signal 1: Think-only — model reasoned but produced no answer (check before empty)
@@ -19,9 +19,17 @@ export function passesQualityGate(stripped, thinkOnly, finishReason, mode) {
19
19
  }
20
20
  // Signal 2: Mode-aware empty floor.
21
21
  // Route legitimately returns 1–4 char labels ("P1", "YES", "CO4", "FIXED").
22
- // Use length===0 for route; keep <5 for code/chat where single-word answers are invalid.
23
- const emptyFloor = mode === "route" ? 0 : 4;
24
- if (stripped.trim().length <= emptyFloor) {
22
+ // Chat answers can be one short value: a follow-up asking "what is x times
23
+ // 6?" is correctly answered "42". Measured 2026-09-24: under the old <5
24
+ // floor, correct chat answers "16" and "36" failed here and, on a paid
25
+ // plan, were thrown away and re-asked of the cloud. So chat is empty only
26
+ // when it has no letter or digit at all.
27
+ // Code keeps <5: a 1–4 char code answer ("Hi", "DONE") is not an answer.
28
+ const trimmed = stripped.trim();
29
+ const empty = mode === "route" ? trimmed.length === 0
30
+ : mode === "chat" ? !/[\p{L}\p{N}]/u.test(trimmed)
31
+ : trimmed.length <= 4;
32
+ if (empty) {
25
33
  return { pass: false, reason: "empty_response" };
26
34
  }
27
35
  // Signal 3: Hard truncation — Ollama reports finish_reason="length"
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "prism-mcp-server",
3
- "version": "20.21.14",
3
+ "version": "20.21.15",
4
4
  "mcpName": "io.github.dcostenco/prism-coder",
5
5
  "description": "Persistent session memory for AI coding agents that never leaves your machine — including the on-device model that reasons over it. Restores your prior decisions, open TODOs, and changed files across sessions; adds associative recall of related past work, semantic drift detection, and local inference. Local-first by default. Works with Claude Code, Cursor, and Codex.",
6
6
  "module": "index.ts",