ruvnet-brain 3.4.6-dev β†’ 3.4.9-dev

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -4,7 +4,7 @@
4
4
 
5
5
  # 🧠 RuvNet Brain
6
6
 
7
- ### 🧠 RuvNet Brain β€” [![RuvNet Brain version 3.4.6-dev β€” updated 2026-07-17 09:40 EDT](https://img.shields.io/badge/version_3.4.6--dev-updated_2026--07--17_09:40_EDT-1E90FF?style=for-the-badge&labelColor=0757BA)](https://github.com/stuinfla/ruvnet-brain/blob/main/plugin/.claude-plugin/plugin.json)
7
+ ### 🧠 RuvNet Brain β€” [![RuvNet Brain version 3.4.9-dev β€” updated 2026-07-18 02:41 EDT](https://img.shields.io/badge/version_3.4.9--dev-updated_2026--07--18_02:41_EDT-1E90FF?style=for-the-badge&labelColor=0757BA)](https://github.com/stuinfla/ruvnet-brain/blob/main/plugin/.claude-plugin/plugin.json)
8
8
 
9
9
  **A portable, source-grounded brain over Reuven Cohen's (rUv's) RuvNet stack β€” delivered as a Claude Code plugin that makes Claude _use_ the stack instead of fighting it.**
10
10
 
@@ -14,7 +14,7 @@
14
14
  [![explainer](https://img.shields.io/badge/β–Ά%20see%20it%20live-isovision.ai%2Fruvnet--brain-e8a13a?style=flat-square)](https://isovision.ai/ruvnet-brain/)
15
15
  [![license](https://img.shields.io/badge/license-MIT-8ecae6?style=flat-square)](LICENSE)
16
16
  [![grounded](https://img.shields.io/badge/answers-cited%20rUv%20source-333?style=flat-square)](#testing--proof)
17
- [![coverage](https://img.shields.io/badge/coverage-10%25%20of%20ALL%20source%20Β·%20honest-b58900?style=flat-square)](#testing--proof)
17
+ [![coverage](https://img.shields.io/badge/coverage-14%25%20of%20ALL%20source%20Β·%20honest-b58900?style=flat-square)](#testing--proof)
18
18
 
19
19
  > **Three independent things version separately here β€” by design, not drift. Every number below is live (read straight from its real source, never hand-typed), so none of them can go stale:**
20
20
  > - **`plugin`** (badge above) β€” the Claude Code plugin itself: SKILL.md, the grounding hooks, the MCP server. Read live from [`plugin/.claude-plugin/plugin.json`](plugin/.claude-plugin/plugin.json). Updates often β€” this is where behavior fixes land.
@@ -45,7 +45,7 @@
45
45
  **Shipped 2026-07-17.** 3.3 made every card lead with a point. Then Stuart looked at the two pages meant to *teach* the stack and scored them 55/100: the graphics were mediocre, the one page that should show how the pieces fit was a wall of text, and two diagrams were literally unreadable. 3.4 is the fix β€” the harness's invisible work, finally drawn, and drawn so you can actually read it.
46
46
 
47
47
  - **The animation that shows the whole argument** β€” one prompt, sent two ways. Plain Claude Code: a bare wire to one expensive model, no grounding, no memory, no gates. The same model **wrapped in the harness**: it grounds the prompt, routes it to the cheapest model that can do the job, hands back what your project already decided, **inspects the write and can refuse it**, and checks it on the way out. The left lane finishes first β€” and that's the problem.
48
- - **MetaHarness, as a picture** β€” the old card was rectangles with words in them. It's now the thesis at a glance: the model a **frozen** cyan crystal that never moves, seven policy surfaces evolving in a warm orbit around it, the kept branch merging back and the pruned one dying mid-air. *Freeze the model, evolve the harness.* Every number traces to an accepted ADR (28.5% cheaper at 98.1% bar-compliance β€” ADR-073/075/076).
48
+ - **MetaHarness, as a picture** β€” the old card was rectangles with words in them. It's now the thesis at a glance: the model a **frozen** cyan crystal that never moves, seven policy surfaces evolving in a warm orbit around it, the kept branch merging back and the pruned one dying mid-air. *Freeze the model, evolve the harness.* Per agentic-flow's ADR-076: 28.5% cheaper at 98.1% bar-compliance.
49
49
  - **Diagrams you can actually read** β€” two of them shipped with labels rendering at **8px and 3px**: present, un-clipped, and invisible. An SVG sized in one coordinate system, crushed into a narrower column, silently shrinks its own text and no error ever fires. There's now a gate for exactly that (`scripts/check-legibility.mjs`) β€” it measures the *effective* pixel size in the live page, proven to catch the known-bad case before it was trusted. Every diagram label now clears 12px on a phone.
50
50
  - **The tips page has a door** β€” its only link wore the same style as a status readout beside it, and on a phone was hidden entirely β€” the page was **unreachable under 640px**. There's now an unmissable button where it belongs.
51
51
  - **Real generated imagery** β€” three commissioned stills built from the console's own palette, so they belong to the page instead of sitting on top of it.
@@ -125,7 +125,7 @@ So 2.5.1 makes it a **wall, not advice**: a `PreToolUse` gate that **blocks any
125
125
  </details>
126
126
 
127
127
  <details>
128
- <summary><b>Earlier &#8212; what 2.0 proved</b> &#183; the release where the brain stopped taking its own word for anything: 32 verified repos, a 120-question fail-closed eval gate, ~90% cheaper per-turn injection, and an 8-dimension evidence-backed scorecard (55 &#8594; 83 in two days). <i>Expand for the receipts.</i></summary>
128
+ <summary><b>Earlier &#8212; what 2.0 proved</b> &#183; the release where the brain stopped taking its own word for anything: 57 verified repos, a 120-question fail-closed eval gate, ~90% cheaper per-turn injection, and an 8-dimension evidence-backed scorecard (55 &#8594; 83 in two days). <i>Expand for the receipts.</i></summary>
129
129
 
130
130
  ### 2.0 β€” the receipts
131
131
 
@@ -133,7 +133,7 @@ So 2.5.1 makes it a **wall, not advice**: a `PreToolUse` gate that **blocks any
133
133
 
134
134
  | | v1 (0.x–1.x) | v2.0 |
135
135
  |---|---|---|
136
- | **Corpus** | 24 repos built | **36 repos** built (of 197 live ruvnet repos), each verified by a live retrieval query |
136
+ | **Corpus** | 24 repos built | **57 repos** built (of 248 in the org), each verified by a live retrieval query |
137
137
  | **Depth** (flagship `ruvector`) | 18,491 passages Β· **0** full source bodies | **28,018 passages Β· 2,996 full bodies** β€” depth also restored to `agent-harness-generator` (8,896/715), `ruview` (7,434/765), `open-claude-code` (195/69) |
138
138
  | **Corpus QA gate** | none | every store must prove *embeds correctly + reads correctly* β€” vector count == passage count, depth floors, a 3-passage self-retrieval round-trip per store β€” **72/72 store-variants PASS**, wired fail-closed into the nightly publish |
139
139
  | **Retrieval eval** | 12 frozen questions | **120 frozen, hash-pinned questions** across 5 strata; promotion gated on Wilson lower bounds, fail-closed β€” it blocked a real release this morning, which is the feature working |
@@ -151,7 +151,7 @@ The depth jump wasn't tuning β€” it was two pipeline root-causes fixed for good:
151
151
  | Dimension | v1 (2026-07-09) | v2.0 (2026-07-10) | Ξ” | What moved it |
152
152
  |---|---:|---:|---:|---|
153
153
  | End-user experience | 54 | 83 | +29 | One-command install now offers nightly self-updates (default yes); the page lives on isovision.ai; publishing renews itself |
154
- | Knowledge corpus | 71 | 88 | +17 | 24β†’32 verified repos; a 72/72 embeds-and-reads QA gate; full source depth restored β€” flagship went 0β†’2,996 source bodies |
154
+ | Knowledge corpus | 71 | 88 | +17 | 24β†’57 verified repos; a 72/72 embeds-and-reads QA gate; full source depth restored β€” flagship went 0β†’2,996 source bodies |
155
155
  | Effectiveness (eval-proven retrieval) | 58 | 88 | +30 | 120-question Wilson-bound gate β€” it blocked a bad release, then passed the fix *above* the old baseline |
156
156
  | Acting like rUv | 38 | 72 | +34 | Memory layer root-caused and fixed with proofs; 6 exact patches queued upstream; real multi-agent swarm operations |
157
157
  | Developer smarter | 62 | 84 | +22 | `/brain-score` runs this same scorecard on any repo; honest tool announcements; per-answer source receipts |
@@ -271,7 +271,7 @@ Plus: the **β€œtake the wheel” behavioral pipeline** (below), a **4-level beha
271
271
 
272
272
  ## How it works
273
273
 
274
- The expensive work happens **once, at build time**: every covered repo is deep-walked (whole files, full function bodies, plus a symbol index), embedded into **two** vector variants (MiniLM-384 for edge/portability, bge-768 for depth) stored on-disk in **RVF / HNSW**, and distilled into a concepts + capability layer of per-repo primers and cards. That's **132,131 source chunks**. At **query time**, `search_ruvnet` searches every repo's store at once, pools the hits, and runs them through **one cross-encoder rerank** on a common scale β€” so the truly relevant file wins regardless of which repo it lives in β€” then returns whole source files, each labeled by repo and path.
274
+ The expensive work happens **once, at build time**: every covered repo is deep-walked (whole files, full function bodies, plus a symbol index), embedded into **two** vector variants (MiniLM-384 for edge/portability, bge-768 for depth) stored on-disk in **RVF / HNSW**, and distilled into a concepts + capability layer of per-repo primers and cards. That's **132,135 source chunks**. At **query time**, `search_ruvnet` searches every repo's store at once, pools the hits, and runs them through **one cross-encoder rerank** on a common scale β€” so the truly relevant file wins regardless of which repo it lives in β€” then returns whole source files, each labeled by repo and path.
275
275
 
276
276
  ![RuvNet Brain architecture pipeline](assets/diagrams/architecture-pipeline.svg)
277
277
 
@@ -305,7 +305,7 @@ The brain answers **both** kinds of questions. **Name the repo or ask something
305
305
 
306
306
  ## What it covers
307
307
 
308
- 36 of rUv's repos in the [ruvnet](https://github.com/ruvnet) org β€” the reusable **building blocks** you'd actually compose into a system β€” each deep-walked and embedded in both variants. The core blocks below also carry symbol indexes and capability cards (the 8 newest repos are findable by name; their capability cards are coming).
308
+ 57 of rUv's repos in the [ruvnet](https://github.com/ruvnet) org β€” the reusable **building blocks** you'd actually compose into a system β€” each deep-walked and embedded in both variants. The core blocks below also carry symbol indexes and capability cards (the 8 newest repos are findable by name; their capability cards are coming).
309
309
 
310
310
  ![The RuvNet stack the brain covers](primer/assets/diagrams/ruvnet-stack.svg)
311
311
 
@@ -344,7 +344,7 @@ node plugin/test/run-tests.mjs # full plugin QA over real JSO
344
344
  | **L1–L4 behavioral harness** | **all pass** | route Β· deep-recall (returns _code_) Β· implement (cites the API) Β· orchestrate (the hook drives the full pipeline) |
345
345
  | **Plugin QA** | **26 / 26** | manifests, hook firing, MCP `initialize`/`tools/list`, capability battery |
346
346
  | **Clean-room install** | **3 / 3** | download the published bundle fresh β†’ unzip β†’ query β†’ grounded, cited answers |
347
- | **Unit tests** | **257 passing** Β· 10% of ALL source covered | `npm run test:cov` β€” the floor fails CI if it slips. 10% is the honest number over every shipped file; the previous "75%" measured a hand-picked 8-file subset |
347
+ | **Unit tests** | **548 passing, 169 todo** Β· 14% of ALL source covered | `npm run test:cov` regenerates both β€” the coverage floor fails CI if it slips (`claims:verify` re-derives the %, it is not a hand-typed badge). 14% is the honest number over every shipped file; the previous "75%" measured a hand-picked 8-file subset |
348
348
  | **Grounding proof** | `npx ruvnet-brain --doctor` | asks a real question, then checks the cited path really exists in the on-disk store; a citation that doesn't resolve is reported as **NOT grounded** |
349
349
  | **Held-out eval** | **grounded 100/100** Β· routed 63/80 | `npm run eval` β€” 120 frozen, hash-pinned questions across 5 strata, never used for tuning, graded on ground truth, never by a model |
350
350
 
@@ -381,7 +381,7 @@ node forge-ask-all.mjs --dir . --q "How does RuVector implement HNSW vector sear
381
381
 
382
382
  This project versions in the open (see the live badge up top for the exact plugin version; the downloadable knowledge bundle is a separate track) β€” we don't claim β€œdone,” β€œcomplete,” or β€œzero hallucinations.” Where it stands:
383
383
 
384
- - βœ… **The grounding brain is real and proven** β€” 36 repos, 132,131 chunks, dual embeddings, cross-encoder rerank, plugin (MCP tool + enforcement hook + skill), all re-runnable.
384
+ - βœ… **The grounding brain is real and proven** β€” 57 repos, 132,135 chunks, dual embeddings, cross-encoder rerank, plugin (MCP tool + enforcement hook + skill), all re-runnable.
385
385
  - βœ… **Code-level depth** β€” the code-rich repos are indexed to full function bodies; β€œhow is it implemented?” returns the implementation. Verified in the shipped bundle (clean-room 3/3).
386
386
  - βœ… **Routing holds** β€” named 47/48, described 26/28, scenario 7/8; behavioral L1–L4 all pass; private stores fenced out of the public bundle (zero-leak verified).
387
387
  - ⚠️ **Two routing residuals** (above) β€” surfaced, not hidden.
@@ -1,7 +1,7 @@
1
1
  {
2
- "_why": "THE REGISTRY OF WHAT MUST BE RUNNING. Created 2026-07-13 after a failure that must never repeat: com.ruvnet.brain-nightly's launchd trigger had NEVER fired, and nothing noticed, because 'is it running?' was only ever answered by looking at a job's own exit code β€” and launchd reports exit 0 for a job that has never run, which is indistinguishable from success. Silence was being read as health.",
2
+ "_why": "THE REGISTRY OF WHAT MUST BE RUNNING. Created 2026-07-13 after a failure that must never repeat: com.ruvnet.brain-nightly's launchd trigger had NEVER fired, and nothing noticed, because 'is it running?' was only ever answered by looking at a job's own exit code \u2014 and launchd reports exit 0 for a job that has never run, which is indistinguishable from success. Silence was being read as health.",
3
3
  "_how_it_works": "scripts/job-heartbeat.sh wraps every job and writes start/end/exit receipts that a dying job cannot forge or skip (trap-protected). scripts/nightly-watchdog.mjs compares THIS registry against reality: a job listed here that is not loaded, or loaded but has no fresh receipt, or has a receipt with a non-zero exit, is a VIOLATION and gongs the phone. Absence of evidence is failure, never 'probably fine'.",
4
- "_adding_a_job": "Add it here FIRST, then wrap its plist command in job-heartbeat.sh. A job not in this registry is unwatched by definition β€” that is the whole point of a registry rather than per-job good intentions.",
4
+ "_adding_a_job": "Add it here FIRST, then wrap its plist command in job-heartbeat.sh. A job not in this registry is unwatched by definition \u2014 that is the whole point of a registry rather than per-job good intentions.",
5
5
  "heartbeatDir": "~/.cache/ruvnet-brain/heartbeats",
6
6
  "jobs": [
7
7
  {
@@ -41,15 +41,15 @@
41
41
  "schedule": "daily 03:33",
42
42
  "maxAgeHours": 26,
43
43
  "required": true,
44
- "_was_blind": "THE PUREST CASE OF THE BUG. On a healthy day this script returns early and writes ZERO BYTES anywhere β€” its log and state file sat frozen 2+ days stale while it fired on schedule every night. 'Ran and did nothing (healthy)' was literally indistinguishable from 'never ran', ~76 days out of every 90. The heartbeat wrapper fixes this for free: the RECEIPT is written by the wrapper, so the job no longer has to remember to report."
44
+ "_was_blind": "THE PUREST CASE OF THE BUG. On a healthy day this script returns early and writes ZERO BYTES anywhere \u2014 its log and state file sat frozen 2+ days stale while it fired on schedule every night. 'Ran and did nothing (healthy)' was literally indistinguishable from 'never ran', ~76 days out of every 90. The heartbeat wrapper fixes this for free: the RECEIPT is written by the wrapper, so the job no longer has to remember to report."
45
45
  },
46
46
  {
47
47
  "label": "com.stuartkerr.api-spend-watchdog",
48
- "what": "Hourly guard against runaway API spend and agent bursts β€” the one that actually pages the phone",
48
+ "what": "Hourly guard against runaway API spend and agent bursts \u2014 the one that actually pages the phone",
49
49
  "schedule": "hourly",
50
50
  "maxAgeHours": 3,
51
51
  "required": true,
52
- "_note": "It has been pushing '4 scheduled jobs failing silently' every hour, correctly β€” but off launchctl's exit-code field, which CANNOT see a job that never ran. It alerts; it just can't see the hole. Now it is itself watched: nothing supervised the supervisor."
52
+ "_note": "It has been pushing '4 scheduled jobs failing silently' every hour, correctly \u2014 but off launchctl's exit-code field, which CANNOT see a job that never ran. It alerts; it just can't see the hole. Now it is itself watched: nothing supervised the supervisor."
53
53
  },
54
54
  {
55
55
  "label": "com.stuartkerr.ruflo-autoupdate",
@@ -57,14 +57,7 @@
57
57
  "schedule": "daily 03:30",
58
58
  "maxAgeHours": 26,
59
59
  "required": true,
60
- "_note": "This is THE job that keeps rUv's stack fresh on this machine. It was unwatched until 2026-07-13 β€” the single most important update job here had no supervision at all."
61
- },
62
- {
63
- "label": "io.ruv.auto-subscribe",
64
- "what": "Hourly: installs RuvNet package updates the moment they publish",
65
- "schedule": "hourly",
66
- "maxAgeHours": 3,
67
- "required": true
60
+ "_note": "This is THE job that keeps rUv's stack fresh on this machine. It was unwatched until 2026-07-13 \u2014 the single most important update job here had no supervision at all."
68
61
  },
69
62
  {
70
63
  "label": "com.cognitum.ruvector-autoupdate",
@@ -72,7 +65,7 @@
72
65
  "schedule": "every 6h",
73
66
  "maxAgeHours": 8,
74
67
  "required": true,
75
- "_note": "Its git stash pop has silently failed once (orphaned stash@{0} from 2026-06-28 β€” it swallowed local edits and nothing said so). The heartbeat now catches a non-zero exit; the orphaned stash still needs a human decision."
68
+ "_note": "Its git stash pop has silently failed once (orphaned stash@{0} from 2026-06-28 \u2014 it swallowed local edits and nothing said so). The heartbeat now catches a non-zero exit; the orphaned stash still needs a human decision."
76
69
  },
77
70
  {
78
71
  "label": "com.stuartkerr.clear-claude-tmp",
@@ -80,25 +73,25 @@
80
73
  "schedule": "every 3h",
81
74
  "maxAgeHours": 5,
82
75
  "required": true,
83
- "_was_lying": "Its log line had the date BAKED IN at plist-write time β€” all 43 entries since 2026-04-06 were byte-identical. It worked; its log was a lie. Now a real script with a real $(date) and a real delete count."
76
+ "_was_lying": "Its log line had the date BAKED IN at plist-write time \u2014 all 43 entries since 2026-04-06 were byte-identical. It worked; its log was a lie. Now a real script with a real $(date) and a real delete count."
84
77
  },
85
78
  {
86
79
  "label": "com.ruvnet.nightly-watchdog",
87
- "what": "Watches all of the above. Listed here so that IT is watched too β€” a watchdog nobody watches is the same blind spot one level up",
80
+ "what": "Watches all of the above. Listed here so that IT is watched too \u2014 a watchdog nobody watches is the same blind spot one level up",
88
81
  "schedule": "daily 09:00",
89
82
  "maxAgeHours": 26,
90
83
  "required": true
91
84
  },
92
85
  {
93
86
  "label": "com.ruvnet.issue-watch",
94
- "what": "Hourly GitHub-issues SLA watcher (stuinfla/ruvnet-brain) β€” pages ntfy when an open issue sits >4h with no comment from the repo owner",
87
+ "what": "Hourly GitHub-issues SLA watcher (stuinfla/ruvnet-brain) \u2014 pages ntfy when an open issue sits >4h with no comment from the repo owner",
95
88
  "schedule": "hourly",
96
89
  "maxAgeHours": 3,
97
90
  "required": true
98
91
  },
99
92
  {
100
93
  "label": "com.ruvnet.issue-fix",
101
- "what": "Every 10 min: auto-fixes newly opened GitHub issues (stuinfla/ruvnet-brain) β€” bounded headless `claude -p` per issue in a disposable git worktree, pushes an issue-fix/<N> branch + comment or posts an honest triage comment, never touches main, never closes an issue",
94
+ "what": "Every 10 min: auto-fixes newly opened GitHub issues (stuinfla/ruvnet-brain) \u2014 bounded headless `claude -p` per issue in a disposable git worktree, pushes an issue-fix/<N> branch + comment or posts an honest triage comment, never touches main, never closes an issue",
102
95
  "schedule": "every 10 min",
103
96
  "maxAgeHours": 1,
104
97
  "required": true
@@ -109,5 +102,12 @@
109
102
  "maxAgeHours": 26,
110
103
  "what": "nightly bounded live flywheel over routing policy (candidate+receipt only, never live routing)"
111
104
  }
105
+ ],
106
+ "_retired": [
107
+ {
108
+ "label": "io.ruv.auto-subscribe",
109
+ "retired": "2026-07-14",
110
+ "why": "archived as cruft (archive-cruft-2026-07-14/); hourly RuvNet package auto-install is superseded by com.stuartkerr.ruflo-autoupdate (daily) + com.cognitum.ruvector-autoupdate. Removed from watch registry 2026-07-18 so nightly-watchdog stops flagging a deliberately-killed job."
111
+ }
112
112
  ]
113
113
  }
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "ruvnet-brain",
3
- "version": "3.4.6-dev",
3
+ "version": "3.4.9-dev",
4
4
  "description": "One-command installer for RuvNet Brain β€” a portable, source-grounded brain over rUv's RuvNet building blocks, delivered as a Claude Code plugin so Claude uses the stack instead of fighting it.",
5
5
  "type": "module",
6
6
  "bin": {
@@ -94,6 +94,18 @@ export function applyProfile(candidates, profile) {
94
94
  }));
95
95
  }
96
96
 
97
+ // Honest provenance of the catalog the engine is actually using, so no surface can pass the
98
+ // built-in stub off as a real personal catalog (trust rule: never present a fallback as the thing).
99
+ // Returns 'catalog' when a real ~/.claude/model-router/catalog.json is present + valid, else
100
+ // 'built-in-fallback'. Same check loadCatalog() uses β€” kept in lockstep.
101
+ export function catalogSource() {
102
+ try {
103
+ const j = JSON.parse(fs.readFileSync(CATALOG_PATH, 'utf8'));
104
+ if (Array.isArray(j.candidates) && j.candidates.length) return 'catalog';
105
+ } catch { /* fall through */ }
106
+ return 'built-in-fallback';
107
+ }
108
+
97
109
  export function loadCatalog() {
98
110
  try {
99
111
  const j = JSON.parse(fs.readFileSync(CATALOG_PATH, 'utf8'));