ruvnet-brain 3.4.18-dev β 3.4.20-dev
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +13 -13
- package/package.json +1 -1
- package/scripts/nightly-watchdog.mjs +20 -4
package/README.md
CHANGED
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# π§ RuvNet Brain
|
|
6
6
|
|
|
7
|
-
### π§ RuvNet Brain β [](https://github.com/stuinfla/ruvnet-brain/blob/main/plugin/.claude-plugin/plugin.json)
|
|
8
8
|
|
|
9
9
|
**A portable, source-grounded brain over Reuven Cohen's (rUv's) RuvNet stack β delivered as a Claude Code plugin that makes Claude _use_ the stack instead of fighting it.**
|
|
10
10
|
|
|
@@ -16,7 +16,7 @@
|
|
|
16
16
|
[](#testing--proof)
|
|
17
17
|
[](#testing--proof)
|
|
18
18
|
|
|
19
|
-
> **Three independent things version separately here β by design, not drift.
|
|
19
|
+
> **Three independent things version separately here β by design, not drift. Headline claims are regenerated and checked by the claims ledger (`scripts/claims-verify.mjs`); other numbers below are hand-stamped and dated:**
|
|
20
20
|
> - **`plugin`** (badge above) β the Claude Code plugin itself: SKILL.md, the grounding hooks, the MCP server. Read live from [`plugin/.claude-plugin/plugin.json`](plugin/.claude-plugin/plugin.json). Updates often β this is where behavior fixes land.
|
|
21
21
|
> - **`installer (npm)`** (badge above) β the `npx ruvnet-brain` setup script. Read live from the [npm registry](https://www.npmjs.com/package/ruvnet-brain). Only moves when the installer script itself changes β rare.
|
|
22
22
|
> - **Brain Release** (the downloadable knowledge bundle, linked from the "download" badge above) β always resolves to [`releases/latest`](https://github.com/stuinfla/ruvnet-brain/releases/latest) (the nightly publishes fresh bundles as the corpus grows). Only moves when the underlying knowledge base is rebuilt β separate again from the two above.
|
|
@@ -125,7 +125,7 @@ So 2.5.1 makes it a **wall, not advice**: a `PreToolUse` gate that **blocks any
|
|
|
125
125
|
</details>
|
|
126
126
|
|
|
127
127
|
<details>
|
|
128
|
-
<summary><b>Earlier — what 2.0 proved</b> · the release where the brain stopped taking its own word for anything:
|
|
128
|
+
<summary><b>Earlier — what 2.0 proved</b> · the release where the brain stopped taking its own word for anything: 69 verified repos, a 120-question fail-closed eval gate, ~90% cheaper per-turn injection, and an 8-dimension evidence-backed scorecard (55 → 83 in two days). <i>Expand for the receipts.</i></summary>
|
|
129
129
|
|
|
130
130
|
### 2.0 β the receipts
|
|
131
131
|
|
|
@@ -133,12 +133,12 @@ So 2.5.1 makes it a **wall, not advice**: a `PreToolUse` gate that **blocks any
|
|
|
133
133
|
|
|
134
134
|
| | v1 (0.xβ1.x) | v2.0 |
|
|
135
135
|
|---|---|---|
|
|
136
|
-
| **Corpus** | 24 repos built | **
|
|
136
|
+
| **Corpus** | 24 repos built | **69 repos** built (of 248 in the org), each verified by a live retrieval query |
|
|
137
137
|
| **Depth** (flagship `ruvector`) | 18,491 passages Β· **0** full source bodies | **28,018 passages Β· 2,996 full bodies** β depth also restored to `agent-harness-generator` (8,896/715), `ruview` (7,434/765), `open-claude-code` (195/69) |
|
|
138
138
|
| **Corpus QA gate** | none | every store must prove *embeds correctly + reads correctly* β vector count == passage count, depth floors, a 3-passage self-retrieval round-trip per store β **72/72 store-variants PASS**, wired fail-closed into the nightly publish |
|
|
139
139
|
| **Retrieval eval** | 12 frozen questions | **120 frozen, hash-pinned questions** across 5 strata; promotion gated on Wilson lower bounds, fail-closed β it blocked a real release this morning, which is the feature working |
|
|
140
140
|
| **Token cost** | 6,183 bytes injected per hook turn Β· zero self-measurement | **684 bytes (~90% cut)** with an eval-PASS proving zero quality loss Β· a live token meter measuring real bytes/tokens per prompt class Β· ~27% faster repeat queries via the KB cache |
|
|
141
|
-
| **rUv's gists** | not indexed | **
|
|
141
|
+
| **rUv's gists** | not indexed | **444 gists** indexed with per-chunk freshness/provenance banners, refreshed nightly with cost-disciplined skip |
|
|
142
142
|
| **Reliability** | claims were prose | **claims ledger** (6 marketing claims mechanically re-verified in CI) Β· integration tests in CI incl. Linux Β· a Windows CI job Β· an honest coverage denominator (all source files) |
|
|
143
143
|
| **Autonomy** | hooks asked questions to an empty room β the #1 real-user complaint | **`/loop` contract** β checkpoint / resume / done-criteria; hooks detect autonomous mode and stop asking β 10 mutation-verified tests |
|
|
144
144
|
| **Publishing** | manual npm token Β· a nightly release PATH bug | **self-renewing npm token** (launchd daemon, proven end-to-end) Β· PATH bug root-caused and cured |
|
|
@@ -151,7 +151,7 @@ The depth jump wasn't tuning β it was two pipeline root-causes fixed for good:
|
|
|
151
151
|
| Dimension | v1 (2026-07-09) | v2.0 (2026-07-10) | Ξ | What moved it |
|
|
152
152
|
|---|---:|---:|---:|---|
|
|
153
153
|
| End-user experience | 54 | 83 | +29 | One-command install now offers nightly self-updates (default yes); the page lives on isovision.ai; publishing renews itself |
|
|
154
|
-
| Knowledge corpus | 71 | 88 | +17 | 24β
|
|
154
|
+
| Knowledge corpus | 71 | 88 | +17 | 24β69 verified repos; a 72/72 embeds-and-reads QA gate; full source depth restored β flagship went 0β2,996 source bodies |
|
|
155
155
|
| Effectiveness (eval-proven retrieval) | 58 | 88 | +30 | 120-question Wilson-bound gate β it blocked a bad release, then passed the fix *above* the old baseline |
|
|
156
156
|
| Acting like rUv | 38 | 72 | +34 | Memory layer root-caused and fixed with proofs; 6 exact patches queued upstream; real multi-agent swarm operations |
|
|
157
157
|
| Developer smarter | 62 | 84 | +22 | `/brain-score` runs this same scorecard on any repo; honest tool announcements; per-answer source receipts |
|
|
@@ -245,8 +245,8 @@ You install once. After that, three mechanisms keep you on the current brain wit
|
|
|
245
245
|
- **Consent-gated auto-update heartbeat** (the `SessionStart` hook, `plugin/scripts/session-start.sh`). The **first** time the plugin runs on a machine it asks you **once** whether it may keep itself updated in the background β a security-conscious opt-in, because self-update can change the model's own instructions. Your answer is remembered (`~/.cache/ruvnet-brain/.auto-update-pref`) and never asked again. On each session start it does a rate-limited (~15 min) 3s-capped check of the live GitHub `plugin.json`. If a newer plugin version exists **and** you opted in, it downloads it in the background through Claude Code's own trusted marketplace path β but the new version is **staged, not active**: Claude Code only loads plugins at process start, so **this session keeps running the version it started with** until you restart (`claude --continue` brings your conversation right back on the new version). If you declined, it just tells you the command to run. The knowledge bundle is handled more conservatively β **detect + notify only**, never auto-applied, because the bundle isn't cryptographically signed yet and applying it would overwrite executable tool files (SEC-0010 #6).
|
|
246
246
|
|
|
247
247
|
- **Grounding receipt line** (the `UserPromptSubmit` gate, `plugin/scripts/ground-ruvnet.sh`). When the brain engages on a prompt, the answer ends with one dim line stating what it actually did β either it read rUv's real source and **names the file**, or it says plainly that it didn't:
|
|
248
|
-
`π§ RuvNet Brain jumped in Β· cited agentic-flow/docs/adr/ADR-076-reposition-agentic-flow-as-agentic-meta-harness.md Β· v3.
|
|
249
|
-
`π§ RuvNet Brain jumped in Β· guidance only, no source read Β· v3.
|
|
248
|
+
`π§ RuvNet Brain jumped in Β· cited agentic-flow/docs/adr/ADR-076-reposition-agentic-flow-as-agentic-meta-harness.md Β· v3.4.18-dev`
|
|
249
|
+
`π§ RuvNet Brain jumped in Β· guidance only, no source read Β· v3.4.18-dev`
|
|
250
250
|
An unearned citation is worse than no citation, so the line may only name a path the tools genuinely returned β and on a prompt where nothing fires, it stays silent rather than manufacture a receipt. The version shown is the one **actually loaded in memory** for this session; if a newer one is staged awaiting a restart, the line says so plainly (`β¦ vX staged, restart to load`). So you never have to wonder whether the brain is on, which version is acting, or whether an answer was grounded or guessed.
|
|
251
251
|
|
|
252
252
|
- **Nightly publish β `releases/latest` chain** (`scripts/self-update.mjs --publish`, run by the `deploy/com.ruvnet.brain-nightly.plist` LaunchAgent at 03:15). The nightly rebuilds only the repos whose upstream changed, and **if anything was rebuilt** it bumps the product version, cuts a GitHub Release, and advances [`releases/latest`](https://github.com/stuinfla/ruvnet-brain/releases/latest). Plugin and knowledge bundle move under **one** version number, so the heartbeat above picks up both automatically. (The LaunchAgent is not auto-installed β enabling a system scheduler needs explicit owner approval.)
|
|
@@ -271,7 +271,7 @@ Plus: the **βtake the wheelβ behavioral pipeline** (below), a **4-level beha
|
|
|
271
271
|
|
|
272
272
|
## How it works
|
|
273
273
|
|
|
274
|
-
The expensive work happens **once, at build time**: every covered repo is deep-walked (whole files, full function bodies, plus a symbol index), embedded into **two** vector variants (MiniLM-384 for edge/portability, bge-768 for depth) stored on-disk in **RVF / HNSW**, and distilled into a concepts + capability layer of per-repo primers and cards. That's **
|
|
274
|
+
The expensive work happens **once, at build time**: every covered repo is deep-walked (whole files, full function bodies, plus a symbol index), embedded into **two** vector variants (MiniLM-384 for edge/portability, bge-768 for depth) stored on-disk in **RVF / HNSW**, and distilled into a concepts + capability layer of per-repo primers and cards. That's **149,664 source chunks**. At **query time**, `search_ruvnet` searches every repo's store at once, pools the hits, and runs them through **one cross-encoder rerank** on a common scale β so the truly relevant file wins regardless of which repo it lives in β then returns whole source files, each labeled by repo and path.
|
|
275
275
|
|
|
276
276
|

|
|
277
277
|
|
|
@@ -305,7 +305,7 @@ The brain answers **both** kinds of questions. **Name the repo or ask something
|
|
|
305
305
|
|
|
306
306
|
## What it covers
|
|
307
307
|
|
|
308
|
-
|
|
308
|
+
69 of rUv's repos in the [ruvnet](https://github.com/ruvnet) org β the reusable **building blocks** you'd actually compose into a system β each deep-walked and embedded in both variants. The core blocks below also carry symbol indexes and capability cards (the 8 newest repos are findable by name; their capability cards are coming).
|
|
309
309
|
|
|
310
310
|

|
|
311
311
|
|
|
@@ -350,7 +350,7 @@ node plugin/test/run-tests.mjs # full plugin QA over real JSO
|
|
|
350
350
|
|
|
351
351
|
<sub>The suite also carries **169 `it.todo` stubs** β a written backlog, each naming an untested behavior and what it would take to cover. They are deliberately **not** counted as tests: a stub proves nothing, and a number that flatters is worse than no number.</sub>
|
|
352
352
|
|
|
353
|
-
Two honest residuals, not hidden: one described question (_βroute to cheaper models to cut costβ_)
|
|
353
|
+
Two honest residuals, not hidden: one described question (_βroute to cheaper models to cut costβ_) routes to `open-claude-code` instead of `agentic-flow`; one _βmethodologyβ_ question routes to `agent-harness-generator` (in `PROOF.md`) or `cognitum-cogs` (in `HELIX-DEMO-NOHELIX.md`) instead of `sparc`/`ruflo`. Proof reports land in [`PROOF.md`](PROOF.md), [`DESCRIBED-PROOF.md`](DESCRIBED-PROOF.md), and [`HELIX-DEMO-NOHELIX.md`](HELIX-DEMO-NOHELIX.md).
|
|
354
354
|
|
|
355
355
|
### The eval flywheel
|
|
356
356
|
|
|
@@ -381,12 +381,12 @@ node forge-ask-all.mjs --dir . --q "How does RuVector implement HNSW vector sear
|
|
|
381
381
|
|
|
382
382
|
This project versions in the open (see the live badge up top for the exact plugin version; the downloadable knowledge bundle is a separate track) β we don't claim βdone,β βcomplete,β or βzero hallucinations.β Where it stands:
|
|
383
383
|
|
|
384
|
-
- β
**The grounding brain is real and proven** β
|
|
384
|
+
- β
**The grounding brain is real and proven** β 54 public stores Β· 149,664 public source chunks (57 built stores incl. private), dual embeddings, cross-encoder rerank, plugin (MCP tool + enforcement hook + skill), all re-runnable.
|
|
385
385
|
- β
**Code-level depth** β the code-rich repos are indexed to full function bodies; βhow is it implemented?β returns the implementation. Verified in the shipped bundle (clean-room 3/3).
|
|
386
386
|
- β
**Routing holds** β named 47/48, described 26/28, scenario 7/8; behavioral L1βL4 all pass; private stores fenced out of the public bundle (zero-leak verified).
|
|
387
387
|
- β οΈ **Two routing residuals** (above) β surfaced, not hidden.
|
|
388
388
|
- β
**Published on npm** β `npx ruvnet-brain` (short form); `npx github:stuinfla/ruvnet-brain` always tracks the latest commit if you want it even fresher.
|
|
389
|
-
- β³ **The fully-autonomous engineering loop** ([ADR-0008](docs/adr/)) β the behavioral hook
|
|
389
|
+
- β³ **The fully-autonomous engineering loop** ([ADR-0008](docs/adr/)) β the behavioral hook injects the loop CONTRACT (assess β SPARC β ADR/DDD β QA β score); the fully-autonomous loop is ADR-0008's open work.
|
|
390
390
|
|
|
391
391
|
---
|
|
392
392
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "ruvnet-brain",
|
|
3
|
-
"version": "3.4.
|
|
3
|
+
"version": "3.4.20-dev",
|
|
4
4
|
"description": "One-command installer for RuvNet Brain β a portable, source-grounded brain over rUv's RuvNet building blocks, delivered as a Claude Code plugin so Claude uses the stack instead of fighting it.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
|
@@ -36,7 +36,7 @@
|
|
|
36
36
|
import fs from 'node:fs';
|
|
37
37
|
import path from 'node:path';
|
|
38
38
|
import os from 'node:os';
|
|
39
|
-
import { spawnSync } from 'node:child_process';
|
|
39
|
+
import { spawnSync, execFileSync } from 'node:child_process';
|
|
40
40
|
import { fileURLToPath, pathToFileURL } from 'node:url';
|
|
41
41
|
|
|
42
42
|
const ROOT = path.resolve(path.dirname(fileURLToPath(import.meta.url)), '..');
|
|
@@ -79,10 +79,26 @@ export function judge(job, hb, loaded, now) {
|
|
|
79
79
|
}
|
|
80
80
|
const ageHours = (now - stamp) / HOUR;
|
|
81
81
|
|
|
82
|
-
// Started and never finished
|
|
83
|
-
//
|
|
82
|
+
// Started and never finished? DERIVE it from process liveness, not wall-clock alone (2026-07-19):
|
|
83
|
+
// a full corpus rebuild legitimately runs 12h+ (five changed repos = a long day), and the old
|
|
84
|
+
// ">6h running β FAILING" rule false-alarmed on exactly that β a 13.5h nightly whose worker was
|
|
85
|
+
// verifiably at 70% CPU got paged as "hung". The receipt carries the wrapper's pid: if that pid is
|
|
86
|
+
// STILL ALIVE (and is genuinely our wrapper, not a recycled pid), the job is a long run in
|
|
87
|
+
// progress β OK, stated as such. If the pid is GONE while the receipt still says "running", THAT
|
|
88
|
+
// is the real started-and-never-finished (SIGKILL/power loss skipped the trap) β FAILING.
|
|
84
89
|
if (hb.state === 'running' && ageHours > 6) {
|
|
85
|
-
|
|
90
|
+
let alive = false;
|
|
91
|
+
if (hb.pid) {
|
|
92
|
+
try {
|
|
93
|
+
process.kill(hb.pid, 0); // signal 0 = existence check, sends nothing
|
|
94
|
+
const cmd = execFileSync('ps', ['-o', 'command=', '-p', String(hb.pid)], { encoding: 'utf8' });
|
|
95
|
+
alive = cmd.includes('job-heartbeat.sh') && cmd.includes(job.label); // guard against pid recycling
|
|
96
|
+
} catch { alive = false; }
|
|
97
|
+
}
|
|
98
|
+
if (alive) {
|
|
99
|
+
return { state: OK, ageHours, detail: `LONG RUN in progress β running ${ageHours.toFixed(1)}h, wrapper pid ${hb.pid} verified alive (full rebuilds legitimately take 12h+)` };
|
|
100
|
+
}
|
|
101
|
+
return { state: FAILING, ageHours, detail: `started ${ageHours.toFixed(1)}h ago and NEVER FINISHED β receipt says "running" but pid ${hb.pid ?? '?'} is GONE (killed/power loss skipped the trap)` };
|
|
86
102
|
}
|
|
87
103
|
if (ageHours > job.maxAgeHours) {
|
|
88
104
|
return { state: STALE, ageHours, detail: `last ran ${ageHours.toFixed(1)}h ago β its schedule allows ${job.maxAgeHours}h. It stopped.` };
|