formlab-mcp 0.6.40 → 0.6.42

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -16,6 +16,7 @@ Ask Claude (or another MCP client) things like:
16
16
 
17
17
  - _"Which of my formulations use silicone fluid at 5% or more?"_
18
18
  - _"Compare the composition of Formula FORM-001 and FORM-004 side by side."_
19
+ - _"Compare the test results of Coating A and Coating B — which parameters differ, and does either fail?"_
19
20
  - _"What parameters fail most often in Q2 testing?"_
20
21
  - _"Show me the DOE matrix at sample grain for all 'Anti-aging serum' family formulas."_
21
22
  - _"What's untested? Which of my approved formulas have no measurements yet?"_
@@ -45,9 +46,10 @@ Ask Claude (or another MCP client) things like:
45
46
  | `get_sample` | Full record + canonical variant + test reports + blend lineage |
46
47
  | `list_test_results` | Filtered list of test reports (by sample, parameter, **measured-value range** (`value_min`/`value_max`), date, or lab) |
47
48
  | `get_test_result` | Full report: every measurement's value, spec, any per-measurement run **conditions**, and the **resolved instrument** (`instrumentSource`: row / method / run). Complex types carry a `representative` (time-series → final, distribution → D50); a time series adds `aggregated` (per-timestamp mean ± sd) and `pointReplicates` (repeat readings at one `t`). Report-level `seriesReplicates` lists parameters measured 2+ times as separate rows — point vs series replicates are distinct. Each measurement also gives its **effective spec** (`effectiveSpec`, `specSource`: override / template / parameter / none) and `passFail` — panel and Test Method specs apply exactly as in the app, and every pass/fail count across the tools uses the same rule |
48
- | `get_doe_matrix` | Pivot matrix (CSV by default) — rows × ingredients × parameters. A parameter measured under two or more **run conditions** (storage 25 °C vs 40 °C) is one column per condition (`pH`, `pH · 40 °C`; the bare name is the base arm), never averaged together — same columns as the app's DOE Matrix |
49
+ | `get_doe_matrix` | Pivot matrix (CSV by default) — rows × ingredients × parameters. A parameter measured under two or more **run conditions** (storage 25 °C vs 40 °C) is one column per condition (`pH`, `pH · 40 °C`; the bare name is the base arm), never averaged together — same columns, in the same order (grouped by measurement type, the test panel's order within a group), as the app's DOE Matrix |
49
50
  | `find_failures` | Parameters ranked by failure **rate** (failing ÷ all readings; replicates in one report count once), `lowN` flag under 5 readings, plus formulations ranked by share of failing reports — same numbers as the app's Test Analytics → Failures |
50
51
  | `get_coverage_matrix` | TEST coverage by formulation, batch or sample (`grain`) × parameter. Each cell is the **latest** reading by test date (same-day reports averaged, any fail → fail) with its date and how many reports measured it; one column per **run condition** when a parameter was measured under two or more — same rollup as the app's Test Analytics → Coverage |
52
+ | `compare_test_results` | **Test results** side by side across 2+ formulas, batches, samples or reports — the app's Compare Results table. Parameters grouped by measurement type (Numbers, Curves, Distributions, Ranked scales, Text results) and, within a group, in the test panel's order (`order: method`, default) or `az`. One cell per item × parameter: `aggregation` `latest` (default; one value per test date, replicates averaged ± sd), `mean` (± sd between dates), `median` (min–max) or `trend` (change from the first date). Curves compare at their final point, distributions at D50; ranked scales and text show the latest reading (with its rank); a parameter measured under 2+ run conditions is split into `arms`, never averaged. `differs` flags a >5% spread across items, as the app highlights |
51
53
  | `compare_batches` | Reproducibility of 2+ runs (ideally of one formula): each run's **yield %**, actual produced mass, cost/kg; per-ingredient **drift** across the runs vs the formula's proposed wt-%; and the biggest **outlier** run. `basis`: `wt_percent` (default) or `amount` |
52
54
  | `get_batch_pivot` | Production analytics — aggregate batches by `group_by` (formula / status / month / prepared_by / project) × `metric` (count, avg_yield_pct, total_produced_kg, avg_cost_per_batch, total_samples, total_tests), with share + total for additive metrics. Optional `batch_uids` scope + `status` filter |
53
55
  | `get_project_pivot` | Portfolio analytics — aggregate projects by `group_by` (status / phase / priority / business_unit / site / customer / lead, or `cf:<custom field>`) × `metric` (count, total_formulas / batches / samples, total_cost, avg_formulas / batches / cost_per_batch), rolling child activity + cost up by any project attribute. Optional `project_uids` scope |
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "formlab-mcp",
3
- "version": "0.6.40",
3
+ "version": "0.6.42",
4
4
  "mcpName": "io.github.juliu1980/formlab-mcp",
5
5
  "description": "Read-only Model Context Protocol server for FormLab \u2014 lets Claude (and other MCP clients) read and analyze your FormLab data, from a local export file OR your live cloud workspace.",
6
6
  "type": "module",
package/server.json CHANGED
@@ -2,7 +2,7 @@
2
2
  "$schema": "https://static.modelcontextprotocol.io/schemas/2025-12-11/server.schema.json",
3
3
  "name": "io.github.juliu1980/formlab-mcp",
4
4
  "description": "Read-only MCP for FormLab — let Claude query your formulation lab. Free reads a local export; Pro connects to your live cloud workspace with a dedicated read-only token.",
5
- "version": "0.6.40",
5
+ "version": "0.6.42",
6
6
  "websiteUrl": "https://formvix.com/mcp",
7
7
  "repository": {
8
8
  "url": "https://github.com/juliu1980/FormLab",
@@ -13,7 +13,7 @@
13
13
  {
14
14
  "registryType": "npm",
15
15
  "identifier": "formlab-mcp",
16
- "version": "0.6.40",
16
+ "version": "0.6.42",
17
17
  "runtimeHint": "npx",
18
18
  "transport": {
19
19
  "type": "stdio"
@@ -164,7 +164,7 @@ const list_test_results = {
164
164
  const found = (t.measurements || []).some(m => {
165
165
  if (args.parameter && !norm(m.parameter).includes(pn)) return false;
166
166
  if (hasRange) {
167
- const v = _mcpRepNumber(m); // a curve's final value, not its point count (13a)
167
+ const v = _mcpRepNumber(m); // a curve's final value, not its point count (12v)
168
168
  if (v == null) return false;
169
169
  if (args.value_min != null && v < args.value_min) return false;
170
170
  if (args.value_max != null && v > args.value_max) return false;
@@ -226,7 +226,7 @@ const get_test_result = {
226
226
  const get_doe_matrix = {
227
227
  definition: {
228
228
  name: 'get_doe_matrix',
229
- description: 'Build a Design-of-Experiments style pivot matrix: rows = formulations (or samples / batches / tests), columns = composition wt-% per ingredient + aggregated test-parameter values. Same shape as the FormLab DOE Matrix view. A parameter measured under two or more RUN CONDITIONS (e.g. storage 25 °C vs 40 °C) is one column per condition — "pH · 40 °C" — and the bare parameter column is the base arm (readings with no condition, else the condition with the most readings, on a tie the one declared as the default on the Test Method / Test Panel); they are different responses, never averaged together. Returns CSV when format=csv (default) for compact LLM consumption, or structured JSON when format=json.',
229
+ description: 'Build a Design-of-Experiments style pivot matrix: rows = formulations (or samples / batches / tests), columns = composition wt-% per ingredient + aggregated test-parameter values. Same shape as the FormLab DOE Matrix view. A parameter measured under two or more RUN CONDITIONS (e.g. storage 25 °C vs 40 °C) is one column per condition — "pH · 40 °C" — and the bare parameter column is the base arm (readings with no condition, else the condition with the most readings, on a tie the one declared as the default on the Test Method / Test Panel); they are different responses, never averaged together. Parameter columns come in the same order as the app: grouped by measurement type (numbers, curves, distributions, ranked scales, text) and in the test panel order within a group. Returns CSV when format=csv (default) for compact LLM consumption, or structured JSON when format=json.',
230
230
  inputSchema: {
231
231
  type: 'object',
232
232
  properties: {
@@ -340,8 +340,18 @@ const get_doe_matrix = {
340
340
  const armNames = _mcpArmNames(rows.flatMap(r => r.measurements
341
341
  .filter(m => m && m.parameter && parameters.has(m.parameter))
342
342
  .map(m => ({ parameter: m.parameter, m }))));
343
+ // Parameter columns in the app's order (13h): grouped by measurement type
344
+ // (numbers, curves, distributions, ranked scales, text), the test panel's
345
+ // order within a group — the same order as compare_test_results.
346
+ const testOf = new Map();
347
+ (db.testResults || []).forEach(t => (t.measurements || []).forEach(x => testOf.set(x, t)));
348
+ const ordered = _mcpParamGroupsOf(rows.flatMap(r => r.measurements.filter(m => m && m.parameter && parameters.has(m.parameter)).map(m => ({ m, t: testOf.get(m) || null }))), 'method', db)
349
+ .flatMap(g => g.parameters);
350
+ const byTrim = new Map([...parameters].map(k => [String(k).trim(), k]));
351
+ const paramOrder = ordered.map(p => byTrim.get(p)).filter(k => k != null);
352
+ [...parameters].forEach(k => { if (!paramOrder.includes(k)) paramOrder.push(k); });
343
353
  const paramCols = [];
344
- [...parameters].forEach(p => {
354
+ paramOrder.forEach(p => {
345
355
  const arms = armNames.get(p);
346
356
  if (!arms) { paramCols.push({ id: p, cond: null, label: p }); return; }
347
357
  arms.forEach((label, sig) => paramCols.push({ id: p, cond: sig, label }));
@@ -352,7 +362,7 @@ const get_doe_matrix = {
352
362
  const cellParamValue = (r, col) => {
353
363
  const ms = r.measurements.filter(m => m.parameter === col.id && (col.cond == null || _mcpCondSig(m) === col.cond));
354
364
  if (!ms.length) return null;
355
- // The app's DOE Matrix numbers (13a): a curve's final value, a
365
+ // The app's DOE Matrix numbers (12v): a curve's final value, a
356
366
  // distribution's D50, an ordinal's 1-based rank; text → its first reading.
357
367
  const nums = ms.map(m => _mcpRepNumber(m, { ordinal: 'rank' })).filter(v => v != null);
358
368
  if (!nums.length) return ms[0].value ?? null;
@@ -407,7 +417,7 @@ const get_doe_matrix = {
407
417
  // so a parameter measured 3× on one occasion counts once in failure rates and
408
418
  // coverage, exactly as Test Analytics does.
409
419
  const FAIL_MIN_N = 5; // app: _TA_FAIL_MIN_N — rates on fewer readings are "low n"
410
- // One number per measurement — the app's flMeasNum (js/utils.js, 13a), kept
420
+ // One number per measurement — the app's flMeasNum (js/utils.js, 12v), kept
411
421
  // in step. Never parseFloat(m.value): a saved time series / distribution has
412
422
  // "12 points" / "8 bins" there, and an ordinal label "6B" parses to 6.
413
423
  // scalar → value; time_series → opts.ts ('final' default | 'initial' | 'peak');
@@ -449,7 +459,7 @@ function _mcpRepNumber(m, opts) {
449
459
  return isFinite(v) ? v : null;
450
460
  }
451
461
  // A ranked-scale (ordinal) or text (categorical) parameter's trend, as the
452
- // app's Stability view draws it (FormLab 13b): one point per test date × run
462
+ // app's Stability view draws it (FormLab 12w): one point per test date × run
453
463
  // condition — an ordinal date is the MEDIAN rank of its readings (rank =
454
464
  // 1-based position on the reading's own scale), a text date its value (or
455
465
  // "mixed: a / b") — failing when any reading fails; one series per condition.
@@ -774,7 +784,7 @@ function _evalPassFail(m, templateId, db) {
774
784
  if (!crit.requiredValue) return 'no-criteria';
775
785
  return raw.toLowerCase() === String(crit.requiredValue).toLowerCase() ? 'pass' : 'fail';
776
786
  }
777
- // As the app (13d): one representative value against the range — a curve's
787
+ // As the app (12y): one representative value against the range — a curve's
778
788
  // FINAL point, a distribution's statistic named by the spec (crit.stat:
779
789
  // "D90"), D50 when unnamed. It used to require every point / bin in range.
780
790
  if (type === 'time_series' || type === 'distribution') {
@@ -1232,7 +1242,7 @@ const get_stability = {
1232
1242
  if (!paramLc) return { error: 'parameter is required.' };
1233
1243
  const points = [];
1234
1244
  const skipped = { ordinal: 0, categorical: 0 };
1235
- const graded = []; // ordinal / text readings, for the step / strip trend (13b)
1245
+ const graded = []; // ordinal / text readings, for the step / strip trend (12w)
1236
1246
  let specMin = null, specMax = null, unit = '';
1237
1247
  // Each reading keeps the spec it was judged by (reading → panel row →
1238
1248
  // method). Mirrors the app: one band only when they all agree; 2+ distinct
@@ -1245,7 +1255,7 @@ const get_stability = {
1245
1255
  if (!Number.isFinite(date.getTime())) return;
1246
1256
  (t.measurements || []).forEach(m => {
1247
1257
  if ((m.parameter || '').trim().toLowerCase() !== paramLc) return;
1248
- // As the app's Stability chart (13a): a curve → its final value (not
1258
+ // As the app's Stability chart (12v): a curve → its final value (not
1249
1259
  // "12 points" → 12), and ordinal / text readings left out and counted.
1250
1260
  const _ty = _mcpMeasType(m);
1251
1261
  if (_ty === 'ordinal' || _ty === 'categorical') {
@@ -1375,8 +1385,225 @@ const get_stability = {
1375
1385
  };
1376
1386
 
1377
1387
  export { _mcpRepNumber, _mcpMeasType, _evalPassFail, _resolveSpec, _shelfLife, _shelfLifeAssess, _mcpCondTags, _mcpArmNames, _mcpCondDefaultScore };
1388
+ // ============================================================
1389
+ // compare_test_results — the app's Compare Results table (js/views/testing-
1390
+ // compare.js), 13g. Items are formulas, batches, samples or test reports; each
1391
+ // gathers the reports behind it (a formula: every report on its samples). The
1392
+ // parameters are grouped by measurement type — numbers, curves, distributions,
1393
+ // ranked scales, text — and, within a group, in the test panel's order
1394
+ // ('method', stable) or by name ('az'), the same rule as the app (utils.js
1395
+ // flParamGroups). One cell per item × parameter, rolled up as the app does:
1396
+ // one value per test DATE (replicates on a date averaged), a curve as its final
1397
+ // point and a distribution as its D50 (flMeasNum), a parameter measured under
1398
+ // two or more run conditions split into one arm per condition (never averaged
1399
+ // across them), a ranked scale or text result shown as its latest reading.
1400
+ // ============================================================
1401
+ const _MCP_PARAM_GROUPS = [
1402
+ ['scalar', 'Numbers'], ['time_series', 'Curves'], ['distribution', 'Distributions'],
1403
+ ['ordinal', 'Ranked scales'], ['categorical', 'Text results'],
1404
+ ];
1405
+ const _mcpCmp = (a, b) => String(a).localeCompare(String(b), undefined, { numeric: true, sensitivity: 'base' }) || (a < b ? -1 : a > b ? 1 : 0);
1406
+ function _mcpParamGroups(items, sort, db) {
1407
+ const pairs = [];
1408
+ (items || []).forEach(it => (it.tests || []).forEach(t => (t.measurements || []).forEach(m => pairs.push({ m, t }))));
1409
+ return _mcpParamGroupsOf(pairs, sort, db);
1410
+ }
1411
+ // From [{ m, t }] pairs (t may be null) — get_doe_matrix holds flattened measurements (13h).
1412
+ function _mcpParamGroupsOf(pairs, sort, db) {
1413
+ const sortAz = sort === 'az';
1414
+ const rank = new Map(_MCP_PARAM_GROUPS.map(([t], i) => [t, i]));
1415
+ const best = new Map(), panels = new Map(), onPanel = new Map();
1416
+ (pairs || []).forEach(({ m, t }) => {
1417
+ const p = String((m && m.parameter) || '').trim();
1418
+ if (!p) return;
1419
+ const k = rank.get(_mcpMeasType(m)) ?? 0;
1420
+ if (!best.has(p) || k < best.get(p)) best.set(p, k);
1421
+ if (sortAz || !t || !t.templateId) return;
1422
+ if (!panels.has(t.templateId)) panels.set(t.templateId, (db.templates || []).find(x => x && x.id === t.templateId) || null);
1423
+ const panel = panels.get(t.templateId);
1424
+ const params = (panel && panel.parameters) || [];
1425
+ const i = params.findIndex(x => x && ((m.parameterId && x.parameterId === m.parameterId) || String(x.parameter || x.name || '').trim() === p));
1426
+ if (i < 0) return;
1427
+ const at = [String(panel.name || ''), i];
1428
+ const cur = onPanel.get(p);
1429
+ if (!cur || _mcpCmp(at[0], cur[0]) < 0 || (at[0] === cur[0] && i < cur[1])) onPanel.set(p, at);
1430
+ });
1431
+ const byMethod = (a, b) => {
1432
+ const x = onPanel.get(a), y = onPanel.get(b);
1433
+ if (x && y) return _mcpCmp(x[0], y[0]) || (x[1] - y[1]) || _mcpCmp(a, b);
1434
+ if (x || y) return x ? -1 : 1;
1435
+ return _mcpCmp(a, b);
1436
+ };
1437
+ const names = [...best.keys()];
1438
+ return _MCP_PARAM_GROUPS.map(([type, label], i) => ({ type, label,
1439
+ parameters: names.filter(p => best.get(p) === i).sort(sortAz ? _mcpCmp : byMethod) }))
1440
+ .filter(g => g.parameters.length);
1441
+ }
1442
+
1443
+ const _tcDate = (t) => t.testDate ? new Date(t.testDate).getTime() : (t.createdAt ? new Date(t.createdAt).getTime() : 0);
1444
+ const _tcIso = (ms) => { const d = new Date(ms); return Number.isFinite(d.getTime()) && ms ? d.toISOString().slice(0, 10) : ''; };
1445
+
1446
+ // One item × parameter. Mirrors _tcCellRollup / _perfParamSummaries / _tcSeriesCell.
1447
+ function _mcpCompareCell(item, parameter, agg, db) {
1448
+ const hits = [];
1449
+ (item.tests || []).forEach(t => (t.measurements || []).forEach(m => { if (String(m.parameter || '').trim() === parameter) hits.push({ m, t }); }));
1450
+ if (!hits.length) return null;
1451
+ hits.sort((a, b) => _tcDate(a.t) - _tcDate(b.t));
1452
+ const type = _mcpMeasType(hits[0].m);
1453
+ const arms = new Map();
1454
+ hits.forEach(({ m, t }) => {
1455
+ const v = _mcpRepNumber(m);
1456
+ if (v == null || !isFinite(v)) return;
1457
+ const sig = _mcpCondSig(m);
1458
+ if (!arms.has(sig)) arms.set(sig, { sig, m, t, rows: [] });
1459
+ arms.get(sig).rows.push({ m, t, date: _tcDate(t), value: v, pf: _evalPassFail(m, t.templateId, db) });
1460
+ });
1461
+ // Nothing numeric (a ranked scale, a text result, a distribution with no D50):
1462
+ // the latest reading, as typed.
1463
+ if (!arms.size) {
1464
+ const { m, t } = hits[hits.length - 1];
1465
+ const out = { type, value: m.value ?? null, unit: m.unit || '', pf: _evalPassFail(m, t.templateId, db), date: t.testDate || '', readings: hits.length };
1466
+ if (type === 'ordinal' && Array.isArray(m.scale)) { const r = _mcpRepNumber(m, { ordinal: 'rank' }); if (r != null) { out.rank = r; out.scaleLength = m.scale.length; } }
1467
+ if (type === 'distribution' && Array.isArray(m.bins)) { out.value = null; out.bins = m.bins.map(b => ({ label: b.label, value: b.value })); out.note = 'no D50 bin, so no single value'; }
1468
+ if (hits.length > 1 && (agg === 'mean' || agg === 'median' || agg === 'trend')) out.note = (out.note ? out.note + '; ' : '') + 'not a number — the latest reading is shown';
1469
+ return out;
1470
+ }
1471
+ const series = [...arms.values()].map(a => {
1472
+ const byDay = new Map();
1473
+ a.rows.forEach(r => { if (!byDay.has(r.date)) byDay.set(r.date, []); byDay.get(r.date).push(r); });
1474
+ const points = [...byDay.keys()].sort((x, y) => x - y).map(d => {
1475
+ const g = byDay.get(d); const vals = g.map(r => r.value);
1476
+ const mean = vals.reduce((x, y) => x + y, 0) / vals.length;
1477
+ const sd = vals.length >= 2 ? Math.sqrt(vals.reduce((x, v) => x + (v - mean) ** 2, 0) / (vals.length - 1)) : 0;
1478
+ const pf = g.some(r => r.pf === 'fail') ? 'fail' : g.some(r => r.pf === 'pass') ? 'pass' : g[0].pf;
1479
+ return { date: d, value: mean, sd, pf, replicates: g.length, m: g[0].m, t: g[0].t };
1480
+ });
1481
+ return { sig: a.sig, m: a.m, t: points[points.length - 1].t, latestM: points[points.length - 1].m, points, readings: a.rows.length };
1482
+ });
1483
+ // Primary first: the longest record, then the declared default condition, then the most recent, then by name.
1484
+ series.sort((a, b) => (b.points.length - a.points.length)
1485
+ || (_mcpCondDefaultScore(b.latestM, b.t, db) - _mcpCondDefaultScore(a.latestM, a.t, db))
1486
+ || (b.points[b.points.length - 1].date - a.points[a.points.length - 1].date) || _mcpCmp(a.sig, b.sig));
1487
+ const basis = type === 'time_series' ? 'final' : type === 'distribution' ? 'D50' : undefined;
1488
+ const one = (s) => {
1489
+ const lat = s.points[s.points.length - 1]; const vals = s.points.map(p => p.value);
1490
+ const typed = (lat.replicates === 1 && isFinite(+String(lat.m.value ?? '').trim()) && String(lat.m.value ?? '').trim() !== '' && type === 'scalar') ? +String(lat.m.value).trim() : _round(lat.value, 6);
1491
+ const latest = { date: _tcIso(lat.date), value: typed, pf: lat.pf, replicates: lat.replicates, ...(lat.replicates > 1 ? { sd: _round(lat.sd, 6) } : {}) };
1492
+ const c = { value: latest.value, unit: s.m.unit || '', ...(basis ? { basis } : {}), n: s.points.length, readings: s.readings, latest };
1493
+ if (agg === 'mean' && vals.length > 1) {
1494
+ const mean = vals.reduce((x, y) => x + y, 0) / vals.length;
1495
+ c.value = _round(mean, 6); c.sd = _round(Math.sqrt(vals.reduce((x, v) => x + (v - mean) ** 2, 0) / (vals.length - 1)), 6);
1496
+ } else if (agg === 'median' && vals.length > 1) {
1497
+ const so = vals.slice().sort((x, y) => x - y), k = so.length;
1498
+ c.value = _round(k % 2 ? so[(k - 1) / 2] : (so[k / 2 - 1] + so[k / 2]) / 2, 6); c.min = _round(so[0], 6); c.max = _round(so[k - 1], 6);
1499
+ } else if (agg === 'trend') {
1500
+ const first = vals[0], last = vals[vals.length - 1];
1501
+ const pct = vals.length < 2 ? 0 : first !== 0 ? (last - first) / Math.abs(first) * 100 : (last === first ? 0 : null);
1502
+ c.trend = { first: _round(first, 6), last: _round(last, 6), changePct: pct == null ? null : _round(pct, 2),
1503
+ direction: pct == null ? (last > first ? 'up' : 'down') : Math.abs(pct) < 1 ? 'flat' : (pct > 0 ? 'up' : 'down') };
1504
+ }
1505
+ return c;
1506
+ };
1507
+ const cells = series.map(one);
1508
+ if (series.length === 1) return { type, ...cells[0], ...(series[0].sig ? { condition: _mcpCondLabel(series[0].m) } : {}) };
1509
+ const tags = _mcpCondTags(series.map(s => s.sig ? _mcpCondLabel(s.m) : ''));
1510
+ return { type, ...cells[0], splitByCondition: true,
1511
+ arms: series.map((s, i) => ({ condition: s.sig ? _mcpCondLabel(s.m) : null, tag: tags[i], ...cells[i] })),
1512
+ note: 'measured under more than one run condition: one arm per condition, never averaged; the top-level value is the first arm (the longest record)' };
1513
+ }
1514
+
1515
+ // The app's difference highlight (_tcKeysDiffer): numbers differ when their
1516
+ // spread is more than 5% of the largest magnitude; text on a case-insensitive match.
1517
+ function _mcpCellsDiffer(cells) {
1518
+ const live = cells.filter(c => c && (c.value != null || c.trend));
1519
+ if (live.length < 2) return false;
1520
+ const dirs = live.map(c => (c.trend && c.trend.direction) || '');
1521
+ if (dirs.some(d => d !== dirs[0])) return true;
1522
+ const nums = live.map(c => (typeof c.value === 'number' ? c.value : NaN));
1523
+ if (nums.every(Number.isFinite)) {
1524
+ const lo = Math.min(...nums), hi = Math.max(...nums);
1525
+ return hi !== lo && (hi - lo) / Math.max(Math.abs(lo), Math.abs(hi)) > 0.05;
1526
+ }
1527
+ const low = live.map(c => String(c.value ?? '').trim().toLowerCase());
1528
+ return low.some(k => k !== low[0]);
1529
+ }
1530
+
1531
+ const compare_test_results = {
1532
+ definition: {
1533
+ name: 'compare_test_results',
1534
+ description: 'Compare TEST RESULTS side by side across 2+ formulas, batches, samples or test reports — the app\'s Compare Results table. Each item gathers the reports behind it (a formula: every report on its samples; a batch: reports on samples drawn from it). Parameters come grouped by measurement type (Numbers, Curves, Distributions, Ranked scales, Text results) and, within a group, in the test panel\'s order (order "method", default) or A–Z ("az"). One cell per item × parameter: aggregation "latest" (default) = the most recent test date (replicates on one date averaged, ± sd), "mean" = mean of the per-date values ± sd between dates, "median" = median with min–max, "trend" = latest with the change from the first date. A curve is compared at its final point and a distribution at its D50; a ranked scale or text result shows its latest reading (with its rank on the scale). A parameter measured under 2+ run conditions (e.g. storage 25 °C vs 40 °C) is split into arms, never averaged across conditions. `differs` flags a parameter whose values spread by more than 5% across items, as the app highlights. For composition differences use compare_formulations; for production runs, compare_batches.',
1535
+ inputSchema: {
1536
+ type: 'object',
1537
+ properties: {
1538
+ mode: { type: 'string', enum: ['formulations', 'batches', 'samples', 'tests'], description: 'What the ids are. Default formulations.' },
1539
+ ids: { type: 'array', items: { type: 'string' }, description: '2–20 ids or UIDs (FORM-…, BATCH-…, SAMP-…, TEST-…); a formula may also be given by its exact name.' },
1540
+ aggregation: { type: 'string', enum: ['latest', 'mean', 'median', 'trend'], description: 'How each cell rolls up several test dates. Default latest. Ignored in tests mode (one report per item).' },
1541
+ order: { type: 'string', enum: ['method', 'az'], description: 'Parameter order within each type group: the test panel\'s order (default) or A–Z.' },
1542
+ parameter_filter: { type: 'string', description: 'Only parameters whose name contains this (case-insensitive).' },
1543
+ },
1544
+ required: ['ids'],
1545
+ },
1546
+ },
1547
+ handler: async (args) => {
1548
+ const { db } = getStore();
1549
+ const mode = ['formulations', 'batches', 'samples', 'tests'].includes(args.mode) ? args.mode : 'formulations';
1550
+ const agg = mode === 'tests' ? 'latest' : (['latest', 'mean', 'median', 'trend'].includes(args.aggregation) ? args.aggregation : 'latest');
1551
+ const order = args.order === 'az' ? 'az' : 'method';
1552
+ if (!Array.isArray(args.ids) || args.ids.length < 2) return { error: 'Pass at least 2 ids.' };
1553
+ if (args.ids.length > 20) return { error: 'Compare at most 20 items at once.' };
1554
+ const coll = { formulations: 'formulations', batches: 'batches', samples: 'samples', tests: 'testResults' }[mode];
1555
+ const lc = (s) => String(s || '').trim().toLowerCase();
1556
+ const resolve = (id) => resolveById(coll, id) || (mode === 'formulations' ? (db.formulations || []).find(f => f && lc(f.name) === lc(id)) : null);
1557
+ const recs = []; const missing = [];
1558
+ args.ids.forEach(id => { const r = resolve(id); if (!r) missing.push(id); else if (!recs.includes(r)) recs.push(r); });
1559
+ if (recs.length < 2) return { error: `Fewer than 2 of the ids resolved to ${mode}.`, missing };
1560
+ const samplesById = new Map((db.samples || []).map(s => [s.id, s]));
1561
+ const formOf = (s) => s && s.formulationId ? (db.formulations || []).find(f => f.id === s.formulationId) : null;
1562
+ const pf = lc(args.parameter_filter);
1563
+ const narrow = (t) => pf ? { ...t, measurements: (t.measurements || []).filter(m => lc(m.parameter).includes(pf)) } : t;
1564
+ const items = recs.map(r => {
1565
+ const tests = (mode === 'tests' ? [r] : (db.testResults || []).filter(t => {
1566
+ const s = samplesById.get(t.sampleId);
1567
+ return mode === 'samples' ? t.sampleId === r.id : mode === 'batches' ? !!(s && s.batchId === r.id) : !!(s && s.formulationId === r.id);
1568
+ })).map(narrow);
1569
+ const days = tests.map(t => t.testDate).filter(Boolean).sort();
1570
+ const smp = mode === 'tests' ? samplesById.get(r.sampleId) : mode === 'samples' ? r : null;
1571
+ const form = mode === 'formulations' ? r : mode === 'batches' ? (db.formulations || []).find(f => f.id === r.formulationId) : formOf(smp);
1572
+ return { rec: r, tests, head: {
1573
+ id: r.id, uid: r.uid || null,
1574
+ name: mode === 'formulations' ? (r.name || '') : mode === 'tests' ? (r.reportName || r.uid || '') : (r.uid || r.id),
1575
+ ...(mode !== 'formulations' && form ? { formula: form.uid || form.name } : {}),
1576
+ tests: tests.length, samples: new Set(tests.map(t => t.sampleId).filter(Boolean)).size,
1577
+ dateRange: days.length ? (days[0] === days[days.length - 1] ? days[0] : `${days[0]} … ${days[days.length - 1]}`) : null,
1578
+ } };
1579
+ });
1580
+ const keyOf = (it) => it.head.uid || it.head.id;
1581
+ const groups = _mcpParamGroups(items, order, db).map(g => ({
1582
+ type: g.type, label: g.label,
1583
+ parameters: g.parameters.map(p => {
1584
+ const cells = items.map(it => _mcpCompareCell(it, p, agg, db));
1585
+ const unit = (cells.find(c => c && c.unit) || {}).unit || '';
1586
+ return { parameter: p, ...(unit ? { unit } : {}), differs: _mcpCellsDiffer(cells),
1587
+ missingFrom: items.filter((_, i) => !cells[i]).map(keyOf),
1588
+ values: Object.fromEntries(items.map((it, i) => [keyOf(it), cells[i]])) };
1589
+ }),
1590
+ }));
1591
+ const paramCount = groups.reduce((n, g) => n + g.parameters.length, 0);
1592
+ return {
1593
+ mode, aggregation: agg, order,
1594
+ items: items.map(it => it.head),
1595
+ ...(missing.length ? { missing } : {}),
1596
+ parameterCount: paramCount,
1597
+ differingCount: groups.reduce((n, g) => n + g.parameters.filter(p => p.differs).length, 0),
1598
+ groups,
1599
+ note: 'values are keyed by each item\'s uid (or id). n = test dates behind a value; readings = every reading. A pass/fail (pf) is the latest date\'s verdict: a date fails when any of its readings fails.',
1600
+ };
1601
+ },
1602
+ };
1603
+
1378
1604
  export const tools = {
1379
1605
  list_test_results,
1606
+ compare_test_results,
1380
1607
  get_test_result,
1381
1608
  get_doe_matrix,
1382
1609
  find_failures,
@@ -346,7 +346,7 @@ function _roleCandidates(forms, roleMaps) {
346
346
  return out.sort((a, b) => b.score - a.score);
347
347
  }
348
348
  // A formula's value for a measured parameter, as the app's
349
- // _asmFormulaValueForCol (13a): readings of the primary run condition (the one
349
+ // _asmFormulaValueForCol (12v): readings of the primary run condition (the one
350
350
  // with the most readings), each as the DOE Matrix reads it (a curve's final
351
351
  // value, a distribution's D50, an ordinal's rank), then their mean.
352
352
  function _paramValue(f, param, db) {