ruvnet-brain 2.5.2 β†’ 2.9.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -4,7 +4,7 @@
4
4
 
5
5
  # 🧠 RuvNet Brain
6
6
 
7
- ### 🧠 RuvNet Brain β€” [![RuvNet Brain version 2.5.2 β€” updated 2026-07-12 21:11 EDT](https://img.shields.io/badge/version_2.5.2-updated_2026--07--12_21:11_EDT-1E90FF?style=for-the-badge&labelColor=0757BA)](https://github.com/stuinfla/ruvnet-brain/blob/main/plugin/.claude-plugin/plugin.json)
7
+ ### 🧠 RuvNet Brain β€” [![RuvNet Brain version 2.9.0 β€” updated 2026-07-14 05:59 EDT](https://img.shields.io/badge/version_2.9.0-updated_2026--07--14_05:59_EDT-1E90FF?style=for-the-badge&labelColor=0757BA)](https://github.com/stuinfla/ruvnet-brain/blob/main/plugin/.claude-plugin/plugin.json)
8
8
 
9
9
  **A portable, source-grounded brain over Reuven Cohen's (rUv's) RuvNet stack β€” delivered as a Claude Code plugin that makes Claude _use_ the stack instead of fighting it.**
10
10
 
@@ -77,7 +77,7 @@ So 2.5.1 makes it a **wall, not advice**: a `PreToolUse` gate that **blocks any
77
77
 
78
78
  | | v1 (0.x–1.x) | v2.0 |
79
79
  |---|---|---|
80
- | **Corpus** | 24 repos built | **32 repos** built (of 197 live ruvnet repos), each verified by a live retrieval query |
80
+ | **Corpus** | 24 repos built | **36 repos** built (of 197 live ruvnet repos), each verified by a live retrieval query |
81
81
  | **Depth** (flagship `ruvector`) | 18,491 passages Β· **0** full source bodies | **28,018 passages Β· 2,996 full bodies** β€” depth also restored to `agent-harness-generator` (8,896/715), `ruview` (7,434/765), `open-claude-code` (195/69) |
82
82
  | **Corpus QA gate** | none | every store must prove *embeds correctly + reads correctly* β€” vector count == passage count, depth floors, a 3-passage self-retrieval round-trip per store β€” **72/72 store-variants PASS**, wired fail-closed into the nightly publish |
83
83
  | **Retrieval eval** | 12 frozen questions | **120 frozen, hash-pinned questions** across 5 strata; promotion gated on Wilson lower bounds, fail-closed β€” it blocked a real release this morning, which is the feature working |
@@ -212,7 +212,7 @@ Plus: the **β€œtake the wheel” behavioral pipeline** (below), a **4-level beha
212
212
 
213
213
  ## How it works
214
214
 
215
- The expensive work happens **once, at build time**: every covered repo is deep-walked (whole files, full function bodies, plus a symbol index), embedded into **two** vector variants (MiniLM-384 for edge/portability, bge-768 for depth) stored on-disk in **RVF / HNSW**, and distilled into a concepts + capability layer of per-repo primers and cards. That's **129,037 source chunks**. At **query time**, `search_ruvnet` searches every repo's store at once, pools the hits, and runs them through **one cross-encoder rerank** on a common scale β€” so the truly relevant file wins regardless of which repo it lives in β€” then returns whole source files, each labeled by repo and path.
215
+ The expensive work happens **once, at build time**: every covered repo is deep-walked (whole files, full function bodies, plus a symbol index), embedded into **two** vector variants (MiniLM-384 for edge/portability, bge-768 for depth) stored on-disk in **RVF / HNSW**, and distilled into a concepts + capability layer of per-repo primers and cards. That's **129,685 source chunks**. At **query time**, `search_ruvnet` searches every repo's store at once, pools the hits, and runs them through **one cross-encoder rerank** on a common scale β€” so the truly relevant file wins regardless of which repo it lives in β€” then returns whole source files, each labeled by repo and path.
216
216
 
217
217
  ![RuvNet Brain architecture pipeline](assets/diagrams/architecture-pipeline.svg)
218
218
 
@@ -246,7 +246,7 @@ The brain answers **both** kinds of questions. **Name the repo or ask something
246
246
 
247
247
  ## What it covers
248
248
 
249
- 32 of rUv's repos in the [ruvnet](https://github.com/ruvnet) org β€” the reusable **building blocks** you'd actually compose into a system β€” each deep-walked and embedded in both variants. The core blocks below also carry symbol indexes and capability cards (the 8 newest repos are findable by name; their capability cards are coming).
249
+ 36 of rUv's repos in the [ruvnet](https://github.com/ruvnet) org β€” the reusable **building blocks** you'd actually compose into a system β€” each deep-walked and embedded in both variants. The core blocks below also carry symbol indexes and capability cards (the 8 newest repos are findable by name; their capability cards are coming).
250
250
 
251
251
  ![The RuvNet stack the brain covers](primer/assets/diagrams/ruvnet-stack.svg)
252
252
 
@@ -322,7 +322,7 @@ node forge-ask-all.mjs --dir . --q "How does RuVector implement HNSW vector sear
322
322
 
323
323
  This project versions in the open (see the live badge up top for the exact plugin version; the downloadable knowledge bundle is a separate track) β€” we don't claim β€œdone,” β€œcomplete,” or β€œzero hallucinations.” Where it stands:
324
324
 
325
- - βœ… **The grounding brain is real and proven** β€” 32 repos, 129,037 chunks, dual embeddings, cross-encoder rerank, plugin (MCP tool + enforcement hook + skill), all re-runnable.
325
+ - βœ… **The grounding brain is real and proven** β€” 36 repos, 129,685 chunks, dual embeddings, cross-encoder rerank, plugin (MCP tool + enforcement hook + skill), all re-runnable.
326
326
  - βœ… **Code-level depth** β€” the code-rich repos are indexed to full function bodies; β€œhow is it implemented?” returns the implementation. Verified in the shipped bundle (clean-room 3/3).
327
327
  - βœ… **Routing holds** β€” named 47/48, described 26/28, scenario 7/8; behavioral L1–L4 all pass; private stores fenced out of the public bundle (zero-leak verified).
328
328
  - ⚠️ **Two routing residuals** (above) β€” surfaced, not hidden.
package/package.json CHANGED
@@ -1,7 +1,7 @@
1
1
  {
2
2
  "name": "ruvnet-brain",
3
- "version": "2.5.2",
4
- "description": "One-command installer for RuvNet Brain \u2014 a portable, source-grounded brain over rUv's RuvNet building blocks, delivered as a Claude Code plugin so Claude uses the stack instead of fighting it.",
3
+ "version": "2.9.0",
4
+ "description": "One-command installer for RuvNet Brain β€” a portable, source-grounded brain over rUv's RuvNet building blocks, delivered as a Claude Code plugin so Claude uses the stack instead of fighting it.",
5
5
  "type": "module",
6
6
  "bin": {
7
7
  "ruvnet-brain": "bin/install.mjs"
@@ -147,6 +147,49 @@ export const CHECKS = [
147
147
  };
148
148
  },
149
149
  },
150
+ {
151
+ id: 'does-memory-actually-recall',
152
+ asked: '"Is AgentDB actually storing AND recalling, or is it just small bits?"',
153
+ why: 'I reported AgentDB as broken THREE times. It was never broken. I called `memory search` positionally when the CLI wants -q, then broke my own canary test with a bad grep, and reported MY test defects as PRODUCT defects. This check does the round trip so I can never again mistake my own sloppiness for a broken tool.',
154
+ run: () => {
155
+ // MEMORY_DB override exists so this check can be TESTED AGAINST A BROKEN STATE. A check nobody
156
+ // has watched fail is not a check. (My first two attempts to test it were themselves broken β€”
157
+ // which is the entire reason this check exists.)
158
+ const db = process.env.MEMORY_DB || path.join(ROOT, '.swarm', 'memory.db');
159
+ if (!fs.existsSync(db)) return { ok: false, detail: 'no memory.db β€” memory is not initialised' };
160
+
161
+ // 1. Is it COLLECTING? (rows, embedded, current)
162
+ const q = (sql) => (spawnSync('sqlite3', [db, sql], { encoding: 'utf8' }).stdout || '').trim();
163
+ const total = Number(q('SELECT COUNT(*) FROM memory_entries;'));
164
+ const embedded = Number(q("SELECT SUM(embedding IS NOT NULL AND LENGTH(embedding)>50) FROM memory_entries;"));
165
+ const ageH = (Date.now() - Number(q('SELECT MAX(created_at) FROM memory_entries;'))) / 3600_000;
166
+
167
+ // 2. Does it SURVIVE compaction / session end? (the whole point)
168
+ const pre = Number(q("SELECT COUNT(*) FROM memory_entries WHERE key LIKE 'session-precompact%';"));
169
+ const end = Number(q("SELECT COUNT(*) FROM memory_entries WHERE key LIKE 'session-sessionend%';"));
170
+
171
+ // 3. Does it RECALL? A real semantic round-trip through the REAL CLI β€” not a DB peek.
172
+ const r = spawnSync('npx', ['ruflo@latest', 'memory', 'search', '-q', 'nightly job supervision heartbeat receipts'], {
173
+ cwd: ROOT, encoding: 'utf8', timeout: 120_000,
174
+ });
175
+ const recalled = /Found\s+\d+\s+results/.test(r.stdout || '') && !/Found\s+0\s+results/.test(r.stdout || '');
176
+
177
+ const problems = [];
178
+ if (total < 10) problems.push(`only ${total} entries stored`);
179
+ if (embedded / Math.max(total, 1) < 0.9) problems.push(`only ${embedded}/${total} entries are embedded β€” semantic recall will be partial`);
180
+ if (ageH > 48) problems.push(`newest entry is ${ageH.toFixed(0)}h old β€” capture may have stopped`);
181
+ if (pre === 0) problems.push('ZERO PreCompact snapshots β€” context will NOT survive a compaction');
182
+ if (end === 0) problems.push('ZERO SessionEnd snapshots β€” context will NOT survive a session restart');
183
+ if (!recalled) problems.push('semantic search returned NOTHING β€” recall is dead');
184
+
185
+ return {
186
+ ok: problems.length === 0,
187
+ detail: problems.length
188
+ ? problems.join('; ')
189
+ : `${total} entries (${embedded} embedded), ${pre} PreCompact + ${end} SessionEnd snapshots, semantic recall returns results`,
190
+ };
191
+ },
192
+ },
150
193
  {
151
194
  id: 'is-ci-actually-green',
152
195
  asked: '"Why am I getting failure notifications from GitHub?"',
@@ -43,22 +43,29 @@ export function loadReceipts(file) {
43
43
  }
44
44
 
45
45
  const fmt$ = (n) => `$${n < 0.01 ? n.toFixed(5) : n.toFixed(4)}`;
46
+ // Measured wall-clock, human-readable. '-' when the receipt has no duration β€” never invented.
47
+ const fmtT = (ms) => {
48
+ if (typeof ms !== 'number' || !(ms > 0)) return '-';
49
+ if (ms < 60_000) return `${(ms / 1000).toFixed(1)}s`;
50
+ return `${Math.floor(ms / 60_000)}m${String(Math.round((ms % 60_000) / 1000)).padStart(2, '0')}s`;
51
+ };
46
52
 
47
53
  export function formatTable(rows) {
48
54
  if (!rows.length) return 'No routing receipts yet.\nRoute something cheap first: node scripts/route-cheap.mjs --task "<text>"';
49
55
 
50
56
  // `channel` + `instead of` (2026-07-13): subagent receipts arrived with a per-row baseline β€” the
51
57
  // model that agent WOULD have inherited β€” so a single global "frontier" column would misreport them.
52
- const header = ['date', 'channel', 'task class', 'model used', 'instead of', 'est. cost', 'est. baseline', 'saved'];
58
+ const header = ['date', 'channel', 'task class', 'model used', 'instead of', 'est. cost', 'est. baseline', 'saved', 'time'];
53
59
  const body = rows.map((r) => [
54
60
  (r.ts || '').replace('T', ' ').slice(0, 16),
55
- r.source === 'claude-subagent' ? 'subagent' : 'openrouter',
61
+ r.source === 'claude-subagent' ? 'subagent' : r.source === 'calibration' ? 'calibrate' : 'openrouter',
56
62
  r.task_class || '?',
57
63
  r.model,
58
64
  r.frontier_ref || 'claude-opus-4.8',
59
65
  fmt$(r.est_cost ?? 0),
60
66
  fmt$(r.est_frontier_cost ?? 0),
61
67
  fmt$(r.saved),
68
+ fmtT(r.duration_ms), // measured, never estimated β€” '-' when absent
62
69
  ]);
63
70
  const widths = header.map((h, i) => Math.max(h.length, ...body.map((row) => row[i].length)));
64
71
  const line = (cells) => cells.map((cell, i) => cell.padEnd(widths[i])).join(' ');
@@ -67,19 +74,58 @@ export function formatTable(rows) {
67
74
  const totalFrontier = rows.reduce((s, r) => s + (r.est_frontier_cost || 0), 0);
68
75
  const totalSaved = rows.reduce((s, r) => s + r.saved, 0);
69
76
  const ratio = totalCost > 0 ? (totalFrontier / totalCost).toFixed(1) : '?';
77
+ // Lead with the PERCENTAGE β€” "$1.83" reads as pocket change; "68% cheaper" is the actual
78
+ // message (Stuart, 2026-07-13). Dollars stay for auditability; percent carries the story.
79
+ const pct = totalFrontier > 0 ? Math.round((totalSaved / totalFrontier) * 100) : 0;
70
80
 
71
81
  // Baselines now vary per row; name them all rather than picking one and implying it covers everything.
72
82
  const baselines = [...new Set(rows.map((r) => r.frontier_ref || 'claude-opus-4.8'))].join(', ');
83
+ const timedTotal = rows.reduce((s, r) => s + (typeof r.duration_ms === 'number' && r.duration_ms > 0 ? r.duration_ms : 0), 0);
84
+ const timedRows = rows.filter((r) => typeof r.duration_ms === 'number' && r.duration_ms > 0).length;
85
+ // TIME as a percentage, like cost β€” "20 seconds" is trivia; "~40% faster" is the message
86
+ // (Stuart, 2026-07-13: cheaper AND faster = fundamentally more efficient). Only rows carrying a
87
+ // MEASURED baseline_duration_ms (written by the calibration harness) enter this comparison.
88
+ const paired = rows.filter((r) => r.duration_ms > 0 && r.baseline_duration_ms > 0);
89
+ const pairedRouted = paired.reduce((s, r) => s + r.duration_ms, 0);
90
+ const pairedBase = paired.reduce((s, r) => s + r.baseline_duration_ms, 0);
91
+ const fasterPct = pairedBase > 0 ? Math.round((1 - pairedRouted / pairedBase) * 100) : null;
92
+ // Say what the measurement actually says β€” FASTER, SLOWER, or parity. The 2026-07-13
93
+ // calibration measured cheap tiers at speed PARITY on micro-tasks (startup dominates);
94
+ // a card that only knows how to say "faster" would have lied.
95
+ const speedBadge = fasterPct === null ? ''
96
+ : fasterPct >= 5 ? ` · ⚑ ~${fasterPct}% FASTER (measured, n=${paired.length})`
97
+ : fasterPct <= -5 ? ` · ⏱ ~${-fasterPct}% slower on routed tier (measured, n=${paired.length})`
98
+ : ` · ⚑ speed parity (measured, n=${paired.length})`;
73
99
  const subagents = rows.filter((r) => r.source === 'claude-subagent').length;
74
100
 
101
+ // The card leads with the PERCENTAGE and a spent-vs-unrouted bar β€” "$1.83" reads as pocket
102
+ // change; "68% cheaper", drawn, is the message (Stuart, 2026-07-13). The dollar table stays
103
+ // below for auditability; every number still traces to a receipt row.
104
+ const BAR = 30;
105
+ const withBar = totalFrontier > 0 ? Math.min(BAR, Math.max(1, Math.round(BAR * (totalCost / totalFrontier)))) : BAR;
106
+ const drawBar = (n) => 'β–ˆ'.repeat(n) + 'β–‘'.repeat(BAR - n);
107
+ const rule = '─'.repeat(70);
75
108
  return [
109
+ rule,
110
+ ` πŸ’° SAVED ~${pct}%${speedBadge} Β· ~${ratio}Γ— cheaper Β· ${rows.length} routed task(s) (${subagents} subagent, ${rows.length - subagents} openrouter/calibration)`,
111
+ '',
112
+ ` without routing ${drawBar(BAR)} ${fmt$(totalFrontier)}`,
113
+ ` with MetaHarness ${drawBar(withBar)} ${fmt$(totalCost)} β†’ ~${fmt$(totalSaved)} kept`,
114
+ rule,
76
115
  line(header),
77
116
  line(widths.map((w) => '-'.repeat(w))),
78
117
  ...body.map(line),
79
118
  '',
80
- `${rows.length} routed task(s) (${subagents} subagent, ${rows.length - subagents} openrouter) Β· est. spent ${fmt$(totalCost)} vs ${fmt$(totalFrontier)} unrouted (${baselines}) Β· saved ~${fmt$(totalSaved)} (~${ratio}x cheaper)`,
119
+ `Baselines: ${baselines} β€” the model each task would have run on if it had not been routed.`,
81
120
  'Pricing is live-verified. Token counts are measured OR estimated per row (each row records which, in token_source).',
82
- 'Baseline = the model the task would have run on if it had not been routed.',
121
+ // Time honesty: durations are MEASURED wall-clock of the routed run. A "time saved vs baseline"
122
+ // number requires a measured per-tier speed baseline (the calibration batch) β€” until that exists
123
+ // we state the gap instead of inventing the comparison.
124
+ fasterPct !== null
125
+ ? `Time, measured: routed ${fmtT(pairedRouted)} vs baseline ${fmtT(pairedBase)} on ${paired.length} calibrated task(s) β†’ ${fasterPct >= 5 ? `~${fasterPct}% faster` : fasterPct <= -5 ? `~${-fasterPct}% slower (startup-dominated micro-tasks β€” speed wins come from parallel fan-out and long generations, measured from real dispatches)` : 'speed parity'}. Uncalibrated rows show routed time only.`
126
+ : timedTotal > 0
127
+ ? `Measured time on routed models: ${fmtT(timedTotal)} across ${timedRows} timed task(s). Time-SAVED vs baseline: not yet measured β€” appears after the per-tier calibration run.`
128
+ : 'No measured durations in these receipts yet. Time-saved reporting activates after the per-tier calibration run.',
83
129
  ].join('\n');
84
130
  }
85
131
 
@@ -82,7 +82,9 @@ const fmt$ = (n) => `$${n < 0.01 ? n.toFixed(5) : n.toFixed(4)}`;
82
82
 
83
83
  export function receiptLine(model, costs) {
84
84
  const ref = costs.ref && costs.ref !== FRONTIER.name ? costs.ref : 'frontier';
85
- return `\x1b[2m⚑ MetaHarness: routed to ${model} (est. ${fmt$(costs.cost)} vs ${fmt$(costs.frontier)} ${ref} β€” saved ~${fmt$(costs.saved)})\x1b[0m`;
85
+ // Percentage leads β€” "saved ~$0.005" reads as noise, "~97% cheaper" is the message (Stuart, 2026-07-13).
86
+ const pct = costs.frontier > 0 ? Math.round((costs.saved / costs.frontier) * 100) : 0;
87
+ return `\x1b[2m⚑ MetaHarness: routed to ${model} β€” ~${pct}% cheaper (est. ${fmt$(costs.cost)} vs ${fmt$(costs.frontier)} ${ref}, saved ~${fmt$(costs.saved)})\x1b[0m`;
86
88
  }
87
89
 
88
90
  // Load OPENROUTER_API_KEY from ruvnet-brain/.env if not already in env. Value never printed/logged.