ruvnet-brain 3.9.84-dev β†’ 3.9.129-dev

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -4,7 +4,7 @@
4
4
 
5
5
  # 🧠 RuvNet Brain
6
6
 
7
- ### 🧠 RuvNet Brain β€” [![RuvNet Brain version 3.9.84-dev β€” updated 2026-07-24 16:09 EDT](https://img.shields.io/badge/version_3.9.84--dev-updated_2026--07--24_16:09_EDT-1E90FF?style=for-the-badge&labelColor=0757BA)](https://github.com/stuinfla/ruvnet-brain/blob/main/plugin/.claude-plugin/plugin.json)
7
+ ### 🧠 RuvNet Brain β€” [![RuvNet Brain version 3.9.129-dev β€” updated 2026-07-24 16:09 EDT](https://img.shields.io/badge/version_3.9.129--dev-updated_2026--07--24_16:09_EDT-1E90FF?style=for-the-badge&labelColor=0757BA)](https://github.com/stuinfla/ruvnet-brain/blob/main/plugin/.claude-plugin/plugin.json)
8
8
 
9
9
  **A portable, source-grounded brain over Reuven Cohen's (rUv's) RuvNet stack β€” delivered as a Claude Code plugin that makes Claude _use_ the stack instead of fighting it.**
10
10
 
@@ -30,9 +30,9 @@
30
30
  [![explainer](https://img.shields.io/badge/β–Ά%20see%20it%20live-isovision.ai%2Fruvnet--brain-e8a13a?style=flat-square)](https://isovision.ai/ruvnet-brain/)
31
31
  [![license](https://img.shields.io/badge/license-MIT-8ecae6?style=flat-square)](LICENSE)
32
32
  [![grounded](https://img.shields.io/badge/answers-cited%20rUv%20source-333?style=flat-square)](#testing--proof)
33
- [![coverage](https://img.shields.io/badge/coverage-26%25%20of%20ALL%20source%20Β·%20honest-b58900?style=flat-square)](#testing--proof)
33
+ [![coverage](https://img.shields.io/badge/coverage-32%25%20of%20ALL%20source%20Β·%20honest-b58900?style=flat-square)](#testing--proof)
34
34
 
35
- > **Three independent things version separately here β€” by design, not drift. Headline claims are regenerated and checked by the claims ledger (`scripts/claims-verify.mjs`); other numbers below are hand-stamped and dated:**
35
+ > **One Brain generation everywhere.** npm, the GitHub tag/release, bundle manifests, source metadata, and checksum-bound RVF generations must share the same product version. Headline claims are regenerated and checked by the claims ledger (`scripts/claims-verify.mjs`); other numbers below are hand-stamped and dated:
36
36
  > - **`plugin`** (badge above) β€” the Claude Code plugin itself: SKILL.md, the grounding hooks, the MCP server. Read live from [`plugin/.claude-plugin/plugin.json`](plugin/.claude-plugin/plugin.json). Updates often β€” this is where behavior fixes land.
37
37
  > - **`installer (npm)`** (badge above) β€” the `npx ruvnet-brain` setup script. Read live from the [npm registry](https://www.npmjs.com/package/ruvnet-brain). Only moves when the installer script itself changes β€” rare.
38
38
  > - **Brain Release** (the downloadable knowledge bundle, linked from the "download" badge above) β€” always resolves to [`releases/latest`](https://github.com/stuinfla/ruvnet-brain/releases/latest) (the nightly publishes fresh bundles as the corpus grows). Only moves when the underlying knowledge base is rebuilt β€” separate again from the two above.
@@ -270,7 +270,7 @@ So 2.5.1 makes it a **wall, not advice**: a `PreToolUse` gate that **blocks any
270
270
  |---|---|---|
271
271
  | **Corpus** | 24 repos built | **69 repos** built (of 248 in the org), each verified by a live retrieval query |
272
272
  | **Depth** (flagship `ruvector`) | 18,491 passages Β· **0** full source bodies | **28,018 passages Β· 2,996 full bodies** β€” depth also restored to `agent-harness-generator` (8,896/715), `ruview` (7,434/765), `open-claude-code` (195/69) |
273
- | **Corpus QA gate** | none | every store must prove *embeds correctly + reads correctly* β€” vector count == passage count, depth floors, a 3-passage self-retrieval round-trip per store β€” **72/72 store-variants PASS**, wired fail-closed into the nightly publish |
273
+ | **Corpus QA gate** | none | every canonical store must prove *embeds correctly + reads correctly* β€” vector count == passage count, depth floors, and a deterministic self-retrieval round trip β€” wired fail-closed into nightly promotion |
274
274
  | **Retrieval eval** | 12 frozen questions | **120 frozen, hash-pinned questions** across 5 strata; promotion gated on Wilson lower bounds, fail-closed β€” it blocked a real release this morning, which is the feature working |
275
275
  | **Token cost** | 6,183 bytes injected per hook turn Β· zero self-measurement | **684 bytes (~90% cut)** with an eval-PASS proving zero quality loss Β· a live token meter measuring real bytes/tokens per prompt class Β· ~27% faster repeat queries via the KB cache |
276
276
  | **rUv's gists** | not indexed | **444 gists** indexed with per-chunk freshness/provenance banners, refreshed nightly with cost-disciplined skip |
@@ -286,7 +286,7 @@ The depth jump wasn't tuning β€” it was two pipeline root-causes fixed for good:
286
286
  | Dimension | v1 (2026-07-09) | v2.0 (2026-07-10) | Ξ” | What moved it |
287
287
  |---|---:|---:|---:|---|
288
288
  | End-user experience | 54 | 83 | +29 | One-command install now offers nightly self-updates (default yes); the page lives on isovision.ai; publishing renews itself |
289
- | Knowledge corpus | 71 | 88 | +17 | 24β†’69 verified repos; a 72/72 embeds-and-reads QA gate; full source depth restored β€” flagship went 0β†’2,996 source bodies |
289
+ | Knowledge corpus | 71 | 88 | +17 | 24β†’69 verified repos; an embeds-and-reads QA gate; full source depth restored β€” flagship went 0β†’2,996 source bodies |
290
290
  | Effectiveness (eval-proven retrieval) | 58 | 88 | +30 | 120-question Wilson-bound gate β€” it blocked a bad release, then passed the fix *above* the old baseline |
291
291
  | Acting like rUv | 38 | 72 | +34 | Memory layer root-caused and fixed with proofs; 6 exact patches queued upstream; real multi-agent swarm operations |
292
292
  | Developer smarter | 62 | 84 | +22 | `/brain-score` runs this same scorecard on any repo; honest tool announcements; per-answer source receipts |
@@ -311,7 +311,7 @@ But **Claude was trained on _classical_ software development.** Point it at rUv'
311
311
 
312
312
  > **RuvNet Brain is the missing instruction manual.** It reads rUv's real source, hands Claude the _answer key_, and removes Claude's permission to make things up about the stack. Install it once, aim it at any repo, and a newcomer can build ~9 months ahead β€” without being rUv.
313
313
 
314
- The novelty is **structural grounding, not plain retrieval.** Plain RAG only decides what to _add_ to context. This ships a `UserPromptSubmit` hook that injects a grounding directive on **every** RuvNet-relevant turn β€” the harness consumes that stdout structurally, so the directive is _always present_, not a decline-able suggestion. It's a **strong, always-on nudge** β€” Claude is pointed at the real source and told to ground before asserting on every relevant turn β€” not a hard block on the model's output. **RAG decides what to add; this makes grounding the default the model has to actively argue its way out of.**
314
+ The novelty is **structural grounding, not plain retrieval.** Plain RAG only decides what to _add_ to context. This ships a `UserPromptSubmit` hook that injects a grounding directive on **every** RuvNet-relevant turn β€” the harness consumes that stdout structurally, so the directive is _always present_, not a decline-able suggestion. It's a **strong, always-on nudge** β€” Claude is pointed at the real source and told to ground before asserting on every relevant turn β€” not a hard block on the model's output. (That describes the *grounding* hook. Separately, a few `PreToolUse` gates **can** block a call β€” opt-in (model-router profile), repo-scoped, or path-scoped to the brain's own consent switch β€” and all fail open; [SECURITY.md](SECURITY.md#what-runs-automatically-and-when) lists every hook and whether it can block you.) **RAG decides what to add; this makes grounding the default the model has to actively argue its way out of.**
315
315
 
316
316
  ---
317
317
 
@@ -389,7 +389,7 @@ You install once. After that, three mechanisms keep you on the current brain wit
389
389
  `🧠 RuvNet Brain jumped in Β· guidance only, no source read Β· v3.4.18-dev`
390
390
  An unearned citation is worse than no citation, so the line may only name a path the tools genuinely returned β€” and on a prompt where nothing fires, it stays silent rather than manufacture a receipt. The version shown is the one **actually loaded in memory** for this session; if a newer one is staged awaiting a restart, the line says so plainly (`… vX staged, restart to load`). So you never have to wonder whether the brain is on, which version is acting, or whether an answer was grounded or guessed.
391
391
 
392
- - **Nightly publish β†’ `releases/latest` chain** (`scripts/self-update.mjs --publish`, run by the `deploy/com.ruvnet.brain-nightly.plist` LaunchAgent at 03:15). The nightly rebuilds only the repos whose upstream changed, and **if anything was rebuilt** it bumps the product version, cuts a GitHub Release, and advances [`releases/latest`](https://github.com/stuinfla/ruvnet-brain/releases/latest). Plugin and knowledge bundle move under **one** version number, so the heartbeat above picks up both automatically. (The LaunchAgent is not auto-installed β€” enabling a system scheduler needs explicit owner approval.)
392
+ - **Nightly publish β†’ `releases/latest` chain** (`scripts/self-update.mjs --publish`, run by the `deploy/com.ruvnet.brain-nightly.plist` LaunchAgent at 03:15). The nightly rebuilds only the repos whose upstream changed, and **if anything was rebuilt** it bumps the product version, cuts a GitHub Release, and advances [`releases/latest`](https://github.com/stuinfla/ruvnet-brain/releases/latest). Plugin and knowledge bundle move under **one** version number, so the heartbeat above picks up both automatically. The exact author-vs-end-user schedules, incremental algorithm, failure behavior, and hosting recommendation are documented in [Nightly refresh and publish](docs/NIGHTLY-REFRESH.md). (The LaunchAgent is not auto-installed β€” enabling a system scheduler needs explicit owner approval.)
393
393
 
394
394
  ---
395
395
 
@@ -411,12 +411,12 @@ Plus: the **β€œtake the wheel” behavioral pipeline** (below), a **4-level beha
411
411
 
412
412
  ## How it works
413
413
 
414
- The expensive work happens **once, at build time**: every covered repo is deep-walked (whole files, full function bodies, plus a symbol index), embedded into **two** vector variants (MiniLM-384 for edge/portability, bge-768 for depth) stored on-disk in **RVF / HNSW**, and distilled into a concepts + capability layer of per-repo primers and cards. That's **149,930 source chunks**. At **query time**, `search_ruvnet` searches every repo's store at once, pools the hits, and runs them through **one cross-encoder rerank** on a common scale β€” so the truly relevant file wins regardless of which repo it lives in β€” then returns whole source files, each labeled by repo and path.
414
+ The expensive work happens **once per changed source chunk, at build time**: every covered repo is deep-walked (whole files, full function bodies, plus a symbol index), embedded into one canonical computer-class **bge-768** store in **RVF / HNSW**, and distilled into a concepts + capability layer of per-repo primers and cards. Stable content-addressed chunk IDs let the nightly keep unchanged vectors, delete departed chunks, and embed only additions. At **query time**, `search_ruvnet` searches every repo's store at once, pools the hits, and runs them through **one cross-encoder rerank** on a common scale β€” so the truly relevant file wins regardless of which repo it lives in β€” then returns whole source files, each labeled by repo and path.
415
415
 
416
416
  ![RuvNet Brain architecture pipeline](assets/diagrams/architecture-pipeline.svg)
417
417
 
418
418
  - **Per-repo RVF stores** β€” each repo gets its own HNSW graph; the tool queries across them and normalizes.
419
- - **Dual embeddings + cross-encoder** β€” MiniLM-384 for edge/portability, bge-768 for depth; a single cross-encoder rerank puts every candidate on one comparable scale. (The cross-encoder β€” a third model reading query+passage together β€” is the real quality lever, not a fusion of the two embedders.)
419
+ - **One canonical embedding + cross-encoder** β€” bge-base-en-v1.5 at 768 dimensions for the computer-class Brain; a single cross-encoder rerank puts every candidate on one comparable scale. A frozen bake-off, not dimension alone, governs any future model change.
420
420
  - **Concepts / capability layer** β€” per-repo primers and capability cards let the model ground _capability_ claims and route a described need to the right repo, not just do file lookups.
421
421
 
422
422
  ---
@@ -445,7 +445,7 @@ The brain answers **both** kinds of questions. **Name the repo or ask something
445
445
 
446
446
  ## What it covers
447
447
 
448
- 69 of rUv's repos in the [ruvnet](https://github.com/ruvnet) org β€” the reusable **building blocks** you'd actually compose into a system β€” each deep-walked and embedded in both variants. The core blocks below also carry symbol indexes and capability cards (the 8 newest repos are findable by name; their capability cards are coming).
448
+ 69 of rUv's repos in the [ruvnet](https://github.com/ruvnet) org β€” the reusable **building blocks** you'd actually compose into a system β€” each deep-walked into its own canonical RVF segment. The core blocks below also carry symbol indexes and capability cards (the 8 newest repos are findable by name; their capability cards are coming).
449
449
 
450
450
  ![The RuvNet stack the brain covers](primer/assets/diagrams/ruvnet-stack.svg)
451
451
 
@@ -481,10 +481,11 @@ node plugin/test/run-tests.mjs # full plugin QA over real JSO
481
481
  | **Named / specific routing** | **47 / 48 (98%)** | name the tool β†’ right repo |
482
482
  | **Described-need routing** | **26 / 28 (93%)** | describe the need, no name β†’ right repo (was 33% before capability cards) |
483
483
  | **Context-scenario routing** | **7 / 8 (88%)** | full-scenario prompts route correctly |
484
- | **L1–L4 behavioral harness** | **all pass** | route Β· deep-recall (returns _code_) Β· implement (cites the API) Β· orchestrate (the hook drives the full pipeline) |
485
- | **Plugin QA** | **26 / 26** | manifests, hook firing, MCP `initialize`/`tools/list`, capability battery |
484
+ | **L1–L3 behavioral harness** | **all pass** | route Β· deep-recall (returns _code_) Β· implement (cites the API) β€” each graded on a retrieval outcome |
485
+ | **L4 "orchestrate"** | **downgraded β€” measures speech, not obedience** | L4 asserts the hook's own injected prose contains required words (`must: ['take the wheel','SPARC','swarm',…]`). That proves **the brain spoke**. It cannot fail when the advice is read and ignored β€” which is the failure this product exists to prevent. Counterfactual replay against a brain-off control (ADR-058 Β§D4) is what will earn this row back |
486
+ | **Plugin QA** | **60 / 60** | manifests, hook firing, MCP `initialize`/`tools/list`, capability battery |
486
487
  | **Clean-room install** | **3 / 3** | download the published bundle fresh β†’ unzip β†’ query β†’ grounded, cited answers |
487
- | **Unit tests** | **548 passing, 169 todo** Β· 26% of ALL source covered | `npm run test:cov` regenerates both β€” the coverage floor fails CI if it slips (`claims:verify` re-derives the %, it is not a hand-typed badge). 26% is the honest number over every shipped file; the previous "75%" measured a hand-picked 8-file subset |
488
+ | **Unit tests** | **2,327 passing, 169 todo** Β· 32% of ALL source covered | `npm run test:cov` regenerates both β€” the coverage floor fails CI if it slips (`claims:verify` re-derives the %, it is not a hand-typed badge). 32% is the honest number over every shipped file; the previous "75%" measured a hand-picked 8-file subset |
488
489
  | **Grounding proof** | `npx ruvnet-brain --doctor` | asks a real question, then checks the cited path really exists in the on-disk store; a citation that doesn't resolve is reported as **NOT grounded** |
489
490
  | **Held-out eval** | **grounded 100/100** Β· routed 63/80 | `npm run eval` β€” 120 frozen, hash-pinned questions across 5 strata, never used for tuning, graded on ground truth, never by a model |
490
491
 
@@ -521,11 +522,12 @@ node forge-ask-all.mjs --dir . --q "How does RuVector implement HNSW vector sear
521
522
 
522
523
  This project versions in the open (see the live badge up top for the exact plugin version; the downloadable knowledge bundle is a separate track) β€” we don't claim β€œdone,” β€œcomplete,” or β€œzero hallucinations.” Where it stands:
523
524
 
524
- - βœ… **The grounding brain is real and proven** β€” 54 public stores Β· 149,930 public source chunks (57 built stores incl. private), dual embeddings, cross-encoder rerank, plugin (MCP tool + enforcement hook + skill), all re-runnable.
525
+ - βœ… **The grounding brain is real and proven** β€” 62 public stores Β· 150,163 public source chunks (69 built stores incl. private), dual embeddings, cross-encoder rerank, plugin (MCP tool + enforcement hook + skill), all re-runnable.
525
526
  - βœ… **Code-level depth** β€” the code-rich repos are indexed to full function bodies; β€œhow is it implemented?” returns the implementation. Verified in the shipped bundle (clean-room 3/3).
526
- - βœ… **Routing holds** β€” named 47/48, described 26/28, scenario 7/8; behavioral L1–L4 all pass; private stores fenced out of the public bundle (zero-leak verified).
527
+ - βœ… **Routing holds** β€” named 47/48, described 26/28, scenario 7/8; behavioral L1–L3 all pass (**L4 downgraded β€” it measures that the brain spoke, not that anything listened**); private stores fenced out of the public bundle (zero-leak verified).
527
528
  - ⚠️ **Two routing residuals** (above) β€” surfaced, not hidden.
528
529
  - βœ… **Published on npm** β€” `npx ruvnet-brain` (short form); `npx github:stuinfla/ruvnet-brain` always tracks the latest commit if you want it even fresher.
530
+ - βœ… **Every automatic hook is documented** β€” [SECURITY.md](SECURITY.md#what-runs-automatically-and-when) lists each one, what it reads, and whether it can block a turn. Most are advisory; the gates that *can* block are opt-in (model-router profile) or scoped to this repo, and all fail open.
529
531
  - ⏳ **The fully-autonomous engineering loop** ([ADR-0008](docs/adr/)) β€” the behavioral hook injects the loop CONTRACT (assess β†’ SPARC β†’ ADR/DDD β†’ QA β†’ score); the fully-autonomous loop is ADR-0008's open work.
530
532
 
531
533
  ---