ruvnet-brain 2.4.3 β 2.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +22 -12
- package/bin/install.mjs +5 -1
- package/config/ruvnet-autoupdate.sh.snapshot +109 -0
- package/config/scheduled-jobs.json +93 -0
- package/package.json +9 -2
- package/scripts/dispatch-receipt.mjs +112 -0
- package/scripts/metaharness-receipts.mjs +95 -0
- package/scripts/model-router-engine.mjs +3 -1
- package/scripts/route-cheap.mjs +27 -6
package/README.md
CHANGED
|
@@ -4,7 +4,7 @@
|
|
|
4
4
|
|
|
5
5
|
# π§ RuvNet Brain
|
|
6
6
|
|
|
7
|
-
### π§ RuvNet Brain β [](https://github.com/stuinfla/ruvnet-brain/blob/main/plugin/.claude-plugin/plugin.json)
|
|
8
8
|
|
|
9
9
|
**A portable, source-grounded brain over Reuven Cohen's (rUv's) RuvNet stack β delivered as a Claude Code plugin that makes Claude _use_ the stack instead of fighting it.**
|
|
10
10
|
|
|
@@ -36,17 +36,27 @@
|
|
|
36
36
|
|
|
37
37
|
---
|
|
38
38
|
|
|
39
|
-
## What's new in 2.
|
|
39
|
+
## What's new in 2.5 β it uses rUv's real tools, and every job proves it ran
|
|
40
40
|
|
|
41
|
-
**Shipped 2026-07-
|
|
41
|
+
**Shipped 2026-07-13. Two hard lessons, both fixed at the root.**
|
|
42
42
|
|
|
43
|
-
|
|
44
|
-
|
|
45
|
-
|
|
46
|
-
|
|
47
|
-
|
|
48
|
-
|
|
49
|
-
|
|
43
|
+
### 1. It stopped faking rUv's tools β and now a CI gate makes faking impossible
|
|
44
|
+
|
|
45
|
+
The 2.4 router was **hand-rolled**: our own heuristic with a placeholder policy, presented as "the MetaHarness router." rUv had already shipped the real thing. 2.5 replaces it with **[`@metaharness/router`](https://www.npmjs.com/package/@metaharness/router)** β his actual learned cost-optimal router (k-NN over labelled embeddings, the productized DRACO Phase-2 finding, ADR-040/043).
|
|
46
|
+
|
|
47
|
+
Our code now does **one honest job**: a price transform. A model your subscription already covers becomes `costPerMTok = 0`, and rUv's router does the rest natively β a $0 model that clears the quality bar simply *is* the cheapest sufficient candidate. Cost-optimal routing and "already paid for" compose; they never competed.
|
|
48
|
+
|
|
49
|
+
> **`npm run substitution:check`** β a new CI gate that fails the build if any code implements a capability rUv already ships *and* wears his name without either using the real package or openly disclosing the hand-roll. **Verified to fire on the exact commit where we got this wrong.** You may hand-roll. You may never hand-roll *silently*.
|
|
50
|
+
|
|
51
|
+
### 2. Every scheduled job must PROVE it ran β silence is no longer health
|
|
52
|
+
|
|
53
|
+
`launchctl` reports **exit 0 for a job that has never run** β byte-identical to success. So "check the exit code" cannot tell triumph from total absence, and a nightly job sat unfired for its entire life while every surface read green.
|
|
54
|
+
|
|
55
|
+
- **A registry** (`config/scheduled-jobs.json`) of what *must* run β so unloaded, deleted, or never-fired is a **violation**, not a silence.
|
|
56
|
+
- **A wrapper** every job runs through β start receipt, end receipt, real exit code, urgent phone push on failure. Break-tested through all four death modes (including SIGKILL, which no trap can catch β the watchdog catches that one).
|
|
57
|
+
- **A watchdog** that treats *absence of evidence as failure*, and which is **in its own registry** β because a supervisor nobody supervises just moves the blind spot up one level.
|
|
58
|
+
|
|
59
|
+
Result: **11 jobs supervised, every one producing a fresh successful receipt.** One had been *totally blind* β writing zero bytes on a healthy day, so "ran fine" and "never ran" were indistinguishable. Cured without changing a line of its logic.
|
|
50
60
|
|
|
51
61
|

|
|
52
62
|
|
|
@@ -194,7 +204,7 @@ Plus: the **βtake the wheelβ behavioral pipeline** (below), a **4-level beha
|
|
|
194
204
|
|
|
195
205
|
## How it works
|
|
196
206
|
|
|
197
|
-
The expensive work happens **once, at build time**: every covered repo is deep-walked (whole files, full function bodies, plus a symbol index), embedded into **two** vector variants (MiniLM-384 for edge/portability, bge-768 for depth) stored on-disk in **RVF / HNSW**, and distilled into a concepts + capability layer of per-repo primers and cards. That's **129,
|
|
207
|
+
The expensive work happens **once, at build time**: every covered repo is deep-walked (whole files, full function bodies, plus a symbol index), embedded into **two** vector variants (MiniLM-384 for edge/portability, bge-768 for depth) stored on-disk in **RVF / HNSW**, and distilled into a concepts + capability layer of per-repo primers and cards. That's **129,034 source chunks**. At **query time**, `search_ruvnet` searches every repo's store at once, pools the hits, and runs them through **one cross-encoder rerank** on a common scale β so the truly relevant file wins regardless of which repo it lives in β then returns whole source files, each labeled by repo and path.
|
|
198
208
|
|
|
199
209
|

|
|
200
210
|
|
|
@@ -304,7 +314,7 @@ node forge-ask-all.mjs --dir . --q "How does RuVector implement HNSW vector sear
|
|
|
304
314
|
|
|
305
315
|
This project versions in the open (see the live badge up top for the exact plugin version; the downloadable knowledge bundle is a separate track) β we don't claim βdone,β βcomplete,β or βzero hallucinations.β Where it stands:
|
|
306
316
|
|
|
307
|
-
- β
**The grounding brain is real and proven** β 32 repos, 129,
|
|
317
|
+
- β
**The grounding brain is real and proven** β 32 repos, 129,034 chunks, dual embeddings, cross-encoder rerank, plugin (MCP tool + enforcement hook + skill), all re-runnable.
|
|
308
318
|
- β
**Code-level depth** β the code-rich repos are indexed to full function bodies; βhow is it implemented?β returns the implementation. Verified in the shipped bundle (clean-room 3/3).
|
|
309
319
|
- β
**Routing holds** β named 47/48, described 26/28, scenario 7/8; behavioral L1βL4 all pass; private stores fenced out of the public bundle (zero-leak verified).
|
|
310
320
|
- β οΈ **Two routing residuals** (above) β surfaced, not hidden.
|
package/bin/install.mjs
CHANGED
|
@@ -1151,7 +1151,11 @@ export async function offerRouterProfile() {
|
|
|
1151
1151
|
if (fs.existsSync(s) && !fs.existsSync(d)) { fs.copyFileSync(s, d); ok(`installed ${dst} (edit freely β goldie keeps prices fresh where scheduled)`); }
|
|
1152
1152
|
}
|
|
1153
1153
|
let copied = 0;
|
|
1154
|
-
|
|
1154
|
+
// dispatch-receipt + metaharness-receipts added 2026-07-13: without the LOGGER, subagent routing is
|
|
1155
|
+
// invisible; without the VIEWER, the user has no scoreboard to hold it to. Shipping one without the
|
|
1156
|
+
// other is how a router ends up "working" with three test pings in its log and nobody the wiser.
|
|
1157
|
+
// (dispatch-receipt.mjs relative-imports route-cheap.mjs β they land in the same bin/ dir, so it resolves.)
|
|
1158
|
+
for (const t of ['model-router-engine.mjs', 'model-router-setup.mjs', 'model-router-status.mjs', 'model-router-outcome.mjs', 'route-cheap.mjs', 'dispatch-receipt.mjs', 'metaharness-receipts.mjs', 'codex-routed.sh']) {
|
|
1155
1159
|
const s = path.join(pkgRoot, 'scripts', t);
|
|
1156
1160
|
if (fs.existsSync(s)) { fs.copyFileSync(s, path.join(routerDir, 'bin', t)); copied++; }
|
|
1157
1161
|
}
|
|
@@ -0,0 +1,109 @@
|
|
|
1
|
+
#!/bin/bash
|
|
2
|
+
# Ruvnet Ecosystem Auto-Update
|
|
3
|
+
# Updated: 2026-05-07
|
|
4
|
+
#
|
|
5
|
+
# Runs every night at 03:30 via ~/Library/LaunchAgents/com.stuartkerr.ruflo-autoupdate.plist
|
|
6
|
+
# Discovers all currently-installed Ruvnet ecosystem packages and refreshes them.
|
|
7
|
+
#
|
|
8
|
+
# Tag policy:
|
|
9
|
+
# - ruflo, @claude-flow/cli β @alpha (Stuart runs alpha track)
|
|
10
|
+
# - agentdb β RETIRED 2026-05-30: AgentDB is bundled inside
|
|
11
|
+
# @claude-flow/memory (Ruflo); no standalone install.
|
|
12
|
+
# - everything else (@ruvector/*, ruvector, @claude-flow/aidefence,
|
|
13
|
+
# agent-browser, agentic-flow, flow-nexus)
|
|
14
|
+
# β @latest
|
|
15
|
+
|
|
16
|
+
set -uo pipefail
|
|
17
|
+
|
|
18
|
+
LOG="$HOME/Library/Logs/ruflo-autoupdate.log"
|
|
19
|
+
PATH="$HOME/.npm-global/bin:/usr/local/bin:/opt/homebrew/bin:/usr/bin:/bin"
|
|
20
|
+
export PATH
|
|
21
|
+
|
|
22
|
+
ts() { date -u +"[%Y-%m-%dT%H:%M:%SZ]"; }
|
|
23
|
+
|
|
24
|
+
echo "$(ts) ββββ Ruvnet ecosystem update starting ββββ" | tee -a "$LOG"
|
|
25
|
+
|
|
26
|
+
# Packages known to fail on this platform β skipped from auto-update.
|
|
27
|
+
# @ruvector/edge-net pulls wrtc which needs a darwin-arm64 prebuilt binary
|
|
28
|
+
# from S3 that returns 404. Bug in upstream wrtc package, not actionable here.
|
|
29
|
+
SKIP_PACKAGES="@ruvector/edge-net"
|
|
30
|
+
|
|
31
|
+
is_skipped() {
|
|
32
|
+
local pkg="$1"
|
|
33
|
+
for sk in $SKIP_PACKAGES; do
|
|
34
|
+
[ "$pkg" = "$sk" ] && return 0
|
|
35
|
+
done
|
|
36
|
+
return 1
|
|
37
|
+
}
|
|
38
|
+
|
|
39
|
+
# Discover all currently-installed packages of interest
|
|
40
|
+
# Discover all currently-installed RuvNet packages.
|
|
41
|
+
# 2026-07-13: the old filter listed only 7 names/prefixes, so agentic-qe, @metaharness/*, qudag,
|
|
42
|
+
# ruv-swarm, ruvbot, ruvi, ruvector-extensions and agentic-robotics were installed but NEVER updated
|
|
43
|
+
# β silently frozen at whatever version they landed on. Live drift found that day: agentic-qe
|
|
44
|
+
# 3.11.5 vs 3.12.0, ruvbot 0.3.1 vs 0.3.2, ruvector-extensions 0.1.1 vs 0.1.2.
|
|
45
|
+
# Principle: if it is an installed RuvNet package, it gets kept current. No arbitrary allowlist.
|
|
46
|
+
# (@marketing/ai-swarms is a LOCAL npm link β never npm-install over it; it is excluded by name.)
|
|
47
|
+
PACKAGES=$(npm ls -g --depth=0 --json 2>/dev/null | jq -r '
|
|
48
|
+
.dependencies // {} | keys[] | select(
|
|
49
|
+
startswith("@ruvector/") or
|
|
50
|
+
startswith("@claude-flow/") or
|
|
51
|
+
startswith("@metaharness/") or
|
|
52
|
+
startswith("@agentic-robotics/") or
|
|
53
|
+
. == "ruflo" or
|
|
54
|
+
. == "ruvector" or
|
|
55
|
+
. == "ruvector-extensions" or
|
|
56
|
+
. == "agent-browser" or
|
|
57
|
+
. == "agentic-flow" or
|
|
58
|
+
. == "agentic-qe" or
|
|
59
|
+
. == "agentic-robotics" or
|
|
60
|
+
. == "flow-nexus" or
|
|
61
|
+
. == "qudag" or
|
|
62
|
+
. == "ruv-swarm" or
|
|
63
|
+
. == "ruvbot" or
|
|
64
|
+
. == "ruvi"
|
|
65
|
+
) | select(startswith("@marketing/") | not)
|
|
66
|
+
')
|
|
67
|
+
|
|
68
|
+
if [ -z "$PACKAGES" ]; then
|
|
69
|
+
echo "$(ts) ERROR: no Ruvnet packages found via 'npm ls -g'" | tee -a "$LOG"
|
|
70
|
+
exit 1
|
|
71
|
+
fi
|
|
72
|
+
|
|
73
|
+
# Build install args with the right tag per package
|
|
74
|
+
INSTALL_ARGS=""
|
|
75
|
+
SKIPPED=""
|
|
76
|
+
ALPHA_PACKAGES="ruflo @claude-flow/cli"
|
|
77
|
+
for pkg in $PACKAGES; do
|
|
78
|
+
if is_skipped "$pkg"; then
|
|
79
|
+
SKIPPED="$SKIPPED $pkg"
|
|
80
|
+
continue
|
|
81
|
+
fi
|
|
82
|
+
TAG="@latest"
|
|
83
|
+
for alpha_pkg in $ALPHA_PACKAGES; do
|
|
84
|
+
if [ "$pkg" = "$alpha_pkg" ]; then
|
|
85
|
+
TAG="@alpha"
|
|
86
|
+
break
|
|
87
|
+
fi
|
|
88
|
+
done
|
|
89
|
+
INSTALL_ARGS="$INSTALL_ARGS ${pkg}${TAG}"
|
|
90
|
+
done
|
|
91
|
+
|
|
92
|
+
if [ -n "$SKIPPED" ]; then
|
|
93
|
+
echo "$(ts) Skipping known-broken packages:$SKIPPED" | tee -a "$LOG"
|
|
94
|
+
fi
|
|
95
|
+
|
|
96
|
+
echo "$(ts) Updating: $INSTALL_ARGS" | tee -a "$LOG"
|
|
97
|
+
|
|
98
|
+
# Run the install β keep going even if one package fails
|
|
99
|
+
npm install -g $INSTALL_ARGS 2>&1 | tee -a "$LOG"
|
|
100
|
+
|
|
101
|
+
EXIT_CODE=$?
|
|
102
|
+
if [ $EXIT_CODE -eq 0 ]; then
|
|
103
|
+
echo "$(ts) β auto-update finished cleanly" | tee -a "$LOG"
|
|
104
|
+
else
|
|
105
|
+
echo "$(ts) β auto-update finished with exit code $EXIT_CODE β check log for failed packages" | tee -a "$LOG"
|
|
106
|
+
fi
|
|
107
|
+
|
|
108
|
+
echo "" | tee -a "$LOG"
|
|
109
|
+
exit $EXIT_CODE
|
|
@@ -0,0 +1,93 @@
|
|
|
1
|
+
{
|
|
2
|
+
"_why": "THE REGISTRY OF WHAT MUST BE RUNNING. Created 2026-07-13 after a failure that must never repeat: com.ruvnet.brain-nightly's launchd trigger had NEVER fired, and nothing noticed, because 'is it running?' was only ever answered by looking at a job's own exit code β and launchd reports exit 0 for a job that has never run, which is indistinguishable from success. Silence was being read as health.",
|
|
3
|
+
"_how_it_works": "scripts/job-heartbeat.sh wraps every job and writes start/end/exit receipts that a dying job cannot forge or skip (trap-protected). scripts/nightly-watchdog.mjs compares THIS registry against reality: a job listed here that is not loaded, or loaded but has no fresh receipt, or has a receipt with a non-zero exit, is a VIOLATION and gongs the phone. Absence of evidence is failure, never 'probably fine'.",
|
|
4
|
+
"_adding_a_job": "Add it here FIRST, then wrap its plist command in job-heartbeat.sh. A job not in this registry is unwatched by definition β that is the whole point of a registry rather than per-job good intentions.",
|
|
5
|
+
"heartbeatDir": "~/.cache/ruvnet-brain/heartbeats",
|
|
6
|
+
"jobs": [
|
|
7
|
+
{
|
|
8
|
+
"label": "com.ruvnet.brain-nightly",
|
|
9
|
+
"what": "Refreshes every RuvNet repo, rebuilds the brain, publishes the GitHub Release",
|
|
10
|
+
"schedule": "daily 03:15",
|
|
11
|
+
"maxAgeHours": 26,
|
|
12
|
+
"required": true,
|
|
13
|
+
"legacyLog": "logs/nightly.log"
|
|
14
|
+
},
|
|
15
|
+
{
|
|
16
|
+
"label": "com.ruvnet.brain-gists",
|
|
17
|
+
"what": "Pulls rUv's new gists into the brain and re-embeds when they change",
|
|
18
|
+
"schedule": "daily 21:47",
|
|
19
|
+
"maxAgeHours": 26,
|
|
20
|
+
"required": true,
|
|
21
|
+
"legacyLog": "logs/gists-nightly.log"
|
|
22
|
+
},
|
|
23
|
+
{
|
|
24
|
+
"label": "com.ruvnet.goldie-weekly",
|
|
25
|
+
"what": "Re-researches the cheapest/best models and refreshes the router catalog",
|
|
26
|
+
"schedule": "Mondays 07:30",
|
|
27
|
+
"maxAgeHours": 192,
|
|
28
|
+
"required": true,
|
|
29
|
+
"legacyLog": "logs/goldie.log"
|
|
30
|
+
},
|
|
31
|
+
{
|
|
32
|
+
"label": "com.ruvnet.model-refresh",
|
|
33
|
+
"what": "Refreshes the live model list (~/.claude/models.json) that the router prices against",
|
|
34
|
+
"schedule": "daily 09:00",
|
|
35
|
+
"maxAgeHours": 26,
|
|
36
|
+
"required": true
|
|
37
|
+
},
|
|
38
|
+
{
|
|
39
|
+
"label": "com.ruvnet.npm-token-renew",
|
|
40
|
+
"what": "Renews the npm publish token before it expires (silently lets publishing die if it stops)",
|
|
41
|
+
"schedule": "daily 03:33",
|
|
42
|
+
"maxAgeHours": 26,
|
|
43
|
+
"required": true,
|
|
44
|
+
"_was_blind": "THE PUREST CASE OF THE BUG. On a healthy day this script returns early and writes ZERO BYTES anywhere β its log and state file sat frozen 2+ days stale while it fired on schedule every night. 'Ran and did nothing (healthy)' was literally indistinguishable from 'never ran', ~76 days out of every 90. The heartbeat wrapper fixes this for free: the RECEIPT is written by the wrapper, so the job no longer has to remember to report."
|
|
45
|
+
},
|
|
46
|
+
{
|
|
47
|
+
"label": "com.stuartkerr.api-spend-watchdog",
|
|
48
|
+
"what": "Hourly guard against runaway API spend and agent bursts β the one that actually pages the phone",
|
|
49
|
+
"schedule": "hourly",
|
|
50
|
+
"maxAgeHours": 3,
|
|
51
|
+
"required": true,
|
|
52
|
+
"_note": "It has been pushing '4 scheduled jobs failing silently' every hour, correctly β but off launchctl's exit-code field, which CANNOT see a job that never ran. It alerts; it just can't see the hole. Now it is itself watched: nothing supervised the supervisor."
|
|
53
|
+
},
|
|
54
|
+
{
|
|
55
|
+
"label": "com.stuartkerr.ruflo-autoupdate",
|
|
56
|
+
"what": "Keeps the whole RuvNet/Ruflo npm stack current (npm i -g ~26 packages on @latest/@alpha)",
|
|
57
|
+
"schedule": "daily 03:30",
|
|
58
|
+
"maxAgeHours": 26,
|
|
59
|
+
"required": true,
|
|
60
|
+
"_note": "This is THE job that keeps rUv's stack fresh on this machine. It was unwatched until 2026-07-13 β the single most important update job here had no supervision at all."
|
|
61
|
+
},
|
|
62
|
+
{
|
|
63
|
+
"label": "io.ruv.auto-subscribe",
|
|
64
|
+
"what": "Hourly: installs RuvNet package updates the moment they publish",
|
|
65
|
+
"schedule": "hourly",
|
|
66
|
+
"maxAgeHours": 3,
|
|
67
|
+
"required": true
|
|
68
|
+
},
|
|
69
|
+
{
|
|
70
|
+
"label": "com.cognitum.ruvector-autoupdate",
|
|
71
|
+
"what": "Every 6h: git pull of ~/RuVector_Clean (the RuVector source the Rust crates build from)",
|
|
72
|
+
"schedule": "every 6h",
|
|
73
|
+
"maxAgeHours": 8,
|
|
74
|
+
"required": true,
|
|
75
|
+
"_note": "Its git stash pop has silently failed once (orphaned stash@{0} from 2026-06-28 β it swallowed local edits and nothing said so). The heartbeat now catches a non-zero exit; the orphaned stash still needs a human decision."
|
|
76
|
+
},
|
|
77
|
+
{
|
|
78
|
+
"label": "com.stuartkerr.clear-claude-tmp",
|
|
79
|
+
"what": "Every 3h: purges Claude Code task-output temp files older than 2 days (the dir hits ~800MB)",
|
|
80
|
+
"schedule": "every 3h",
|
|
81
|
+
"maxAgeHours": 5,
|
|
82
|
+
"required": true,
|
|
83
|
+
"_was_lying": "Its log line had the date BAKED IN at plist-write time β all 43 entries since 2026-04-06 were byte-identical. It worked; its log was a lie. Now a real script with a real $(date) and a real delete count."
|
|
84
|
+
},
|
|
85
|
+
{
|
|
86
|
+
"label": "com.ruvnet.nightly-watchdog",
|
|
87
|
+
"what": "Watches all of the above. Listed here so that IT is watched too β a watchdog nobody watches is the same blind spot one level up",
|
|
88
|
+
"schedule": "daily 09:00",
|
|
89
|
+
"maxAgeHours": 26,
|
|
90
|
+
"required": true
|
|
91
|
+
}
|
|
92
|
+
]
|
|
93
|
+
}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "ruvnet-brain",
|
|
3
|
-
"version": "2.
|
|
3
|
+
"version": "2.5.0",
|
|
4
4
|
"description": "One-command installer for RuvNet Brain \u2014 a portable, source-grounded brain over rUv's RuvNet building blocks, delivered as a Claude Code plugin so Claude uses the stack instead of fighting it.",
|
|
5
5
|
"type": "module",
|
|
6
6
|
"bin": {
|
|
@@ -16,6 +16,7 @@
|
|
|
16
16
|
"test:all": "npm run test:unit && npm run test:integration && npm test",
|
|
17
17
|
"metaharness:receipts": "node scripts/metaharness-receipts.mjs",
|
|
18
18
|
"route:cheap": "node scripts/route-cheap.mjs",
|
|
19
|
+
"route:receipt": "node scripts/dispatch-receipt.mjs",
|
|
19
20
|
"metaharness:fix": "node scripts/fix-metaharness-memretrieve.mjs --apply",
|
|
20
21
|
"metaharness:check": "node scripts/fix-metaharness-memretrieve.mjs --check",
|
|
21
22
|
"eval": "node scripts/eval-brain.mjs",
|
|
@@ -25,7 +26,8 @@
|
|
|
25
26
|
"embed:check": "node scripts/embed-verifier.mjs --check",
|
|
26
27
|
"gists:index": "node scripts/ingest-gists.mjs --index-only",
|
|
27
28
|
"gists:sync": "node scripts/ingest-gists.mjs && node kb/forge-big.mjs both --dir kb --name ruv-gists",
|
|
28
|
-
"test:integration": "vitest run tests/integration"
|
|
29
|
+
"test:integration": "vitest run tests/integration",
|
|
30
|
+
"substitution:check": "node scripts/no-silent-substitution.mjs"
|
|
29
31
|
},
|
|
30
32
|
"files": [
|
|
31
33
|
"bin/install.mjs",
|
|
@@ -37,6 +39,8 @@
|
|
|
37
39
|
"scripts/model-router-status.mjs",
|
|
38
40
|
"scripts/model-router-outcome.mjs",
|
|
39
41
|
"scripts/route-cheap.mjs",
|
|
42
|
+
"scripts/dispatch-receipt.mjs",
|
|
43
|
+
"scripts/metaharness-receipts.mjs",
|
|
40
44
|
"scripts/codex-routed.sh"
|
|
41
45
|
],
|
|
42
46
|
"engines": {
|
|
@@ -67,5 +71,8 @@
|
|
|
67
71
|
"@metaharness/darwin": "~0.8.0",
|
|
68
72
|
"@vitest/coverage-v8": "^4.1.10",
|
|
69
73
|
"vitest": "^4.1.10"
|
|
74
|
+
},
|
|
75
|
+
"dependencies": {
|
|
76
|
+
"@metaharness/router": "^0.3.2"
|
|
70
77
|
}
|
|
71
78
|
}
|
|
@@ -0,0 +1,112 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
// scripts/dispatch-receipt.mjs β make SUBAGENT routing visible.
|
|
3
|
+
//
|
|
4
|
+
// WHY THIS EXISTS (2026-07-13, Stuart: "the meta harness isn't really doing anything for me").
|
|
5
|
+
// He was right, and the receipts log proved it: 3 entries, all of them test pings, $0.018 saved in
|
|
6
|
+
// the router's entire life. Two separate faults hid behind that number:
|
|
7
|
+
// 1. BEHAVIOR β mechanical work was being done inline in the main loop instead of dispatched.
|
|
8
|
+
// 2. MEASUREMENT β the biggest lever wrote no receipt AT ALL. A Claude Code subagent INHERITS the
|
|
9
|
+
// main-loop model unless explicitly overridden, so a fan-out of five agents on a Fable session
|
|
10
|
+
// is five Fable agents. Overriding them to haiku is the single largest saving available β and
|
|
11
|
+
// route-cheap.mjs only logged OpenRouter calls, so that saving was invisible even when it happened.
|
|
12
|
+
// This closes #2 so #1 becomes auditable: the receipts file either grows with real work, or the
|
|
13
|
+
// hard rule in SKILL.md is being ignored and the log says so.
|
|
14
|
+
//
|
|
15
|
+
// The baseline is the INHERITED model, not "frontier" β that is the honest counterfactual. An
|
|
16
|
+
// un-overridden subagent genuinely would have run on the session's model.
|
|
17
|
+
//
|
|
18
|
+
// Usage (call it right after the Agent/Task returns, with the REAL text sizes):
|
|
19
|
+
// node scripts/dispatch-receipt.mjs --model claude-haiku-4.5 --inherited claude-fable-5 \
|
|
20
|
+
// --task "sweep tests/ for machine-state deps" --in-chars 2400 --out-chars 9100
|
|
21
|
+
//
|
|
22
|
+
// Costs are estimates (chars/4 tokens x live-verified $/Mtok) and are labeled "est." everywhere.
|
|
23
|
+
// Unknown model or unknown baseline β NO receipt, non-zero exit. Never invent a savings number.
|
|
24
|
+
|
|
25
|
+
import fs from 'node:fs';
|
|
26
|
+
import path from 'node:path';
|
|
27
|
+
import { pathToFileURL } from 'node:url';
|
|
28
|
+
import { CLAUDE_TIERS, PRICING, estTokens, estimateCosts, receiptLine, receiptsPath, priceOf } from './route-cheap.mjs';
|
|
29
|
+
|
|
30
|
+
// Default input share when only a MEASURED TOTAL is known. A subagent's tokens are dominated by input
|
|
31
|
+
// (it re-reads files and tool output on every turn); its final report is small. 0.9 is an assumption,
|
|
32
|
+
// not a measurement, so it is NAMED in the receipt's token_source and overridable with --split.
|
|
33
|
+
const DEFAULT_INPUT_SHARE = 0.9;
|
|
34
|
+
|
|
35
|
+
export function parseArgs(argv) {
|
|
36
|
+
const args = { model: 'claude-haiku-4.5', inherited: 'claude-opus-4.8', class: 'mechanical' };
|
|
37
|
+
for (let i = 0; i < argv.length; i++) {
|
|
38
|
+
const k = argv[i];
|
|
39
|
+
if (['--model', '--inherited', '--task', '--class', '--in-chars', '--out-chars', '--label', '--total-tokens', '--split'].includes(k)) {
|
|
40
|
+
args[k.slice(2).replace(/-([a-z])/g, (_, c) => c.toUpperCase())] = argv[++i];
|
|
41
|
+
}
|
|
42
|
+
}
|
|
43
|
+
return args;
|
|
44
|
+
}
|
|
45
|
+
|
|
46
|
+
/**
|
|
47
|
+
* Token sizing, in descending order of honesty:
|
|
48
|
+
* 1. --total-tokens N β the harness REPORTED this agent's real usage. Split by --split (named).
|
|
49
|
+
* 2. --in-chars/--out-chars β chars/4 estimate of the text we actually sent/received.
|
|
50
|
+
* 3. the task text alone β weakest; use only when nothing else is known.
|
|
51
|
+
* Why this matters: a subagent that reads 40 files burns ~20x the tokens its prompt+report suggest.
|
|
52
|
+
* Sizing it from the prompt alone would UNDERSTATE the saving by that factor and make real routing
|
|
53
|
+
* look pointless β the same "it isn't doing anything" trap, just with the error flipped.
|
|
54
|
+
*/
|
|
55
|
+
export function sizeTokens(args) {
|
|
56
|
+
const total = Number(args.totalTokens);
|
|
57
|
+
if (total > 0) {
|
|
58
|
+
const share = Number(args.split) > 0 && Number(args.split) < 1 ? Number(args.split) : DEFAULT_INPUT_SHARE;
|
|
59
|
+
return {
|
|
60
|
+
inTok: Math.round(total * share),
|
|
61
|
+
outTok: Math.round(total * (1 - share)),
|
|
62
|
+
source: `measured total ${total} tok, assumed ${Math.round(share * 100)}/${Math.round((1 - share) * 100)} in/out split`,
|
|
63
|
+
};
|
|
64
|
+
}
|
|
65
|
+
if (Number(args.inChars) || Number(args.outChars)) {
|
|
66
|
+
return {
|
|
67
|
+
inTok: Number(args.inChars) ? estTokens('x'.repeat(Number(args.inChars))) : 0,
|
|
68
|
+
outTok: Number(args.outChars) ? estTokens('x'.repeat(Number(args.outChars))) : 0,
|
|
69
|
+
source: 'chars/4 est',
|
|
70
|
+
};
|
|
71
|
+
}
|
|
72
|
+
return { inTok: estTokens(args.task || ''), outTok: 0, source: 'chars/4 est (prompt only β likely an undercount)' };
|
|
73
|
+
}
|
|
74
|
+
|
|
75
|
+
/** Build the receipt row. Returns null when either model is unpriced β caller must not fabricate. */
|
|
76
|
+
export function buildReceipt(args, now) {
|
|
77
|
+
if (!priceOf(args.model) || !priceOf(args.inherited)) return null;
|
|
78
|
+
const { inTok, outTok, source } = sizeTokens(args);
|
|
79
|
+
const costs = estimateCosts(args.model, inTok, outTok, args.inherited);
|
|
80
|
+
if (!costs) return null;
|
|
81
|
+
return {
|
|
82
|
+
ts: now,
|
|
83
|
+
task_class: args.class,
|
|
84
|
+
task: (args.task || args.label || '(unlabeled subagent dispatch)').slice(0, 200),
|
|
85
|
+
model: args.model,
|
|
86
|
+
inherited: args.inherited,
|
|
87
|
+
est_in_tokens: inTok,
|
|
88
|
+
est_out_tokens: outTok,
|
|
89
|
+
est_cost: costs.cost,
|
|
90
|
+
est_frontier_cost: costs.frontier, // same schema key as route-cheap receipts; here = inherited-model cost
|
|
91
|
+
saved: costs.saved,
|
|
92
|
+
frontier_ref: args.inherited,
|
|
93
|
+
token_source: source,
|
|
94
|
+
source: 'claude-subagent',
|
|
95
|
+
};
|
|
96
|
+
}
|
|
97
|
+
|
|
98
|
+
function main() {
|
|
99
|
+
const args = parseArgs(process.argv.slice(2));
|
|
100
|
+
const known = [...Object.keys(CLAUDE_TIERS), ...Object.keys(PRICING)].join(', ');
|
|
101
|
+
const receipt = buildReceipt(args, new Date().toISOString());
|
|
102
|
+
if (!receipt) {
|
|
103
|
+
console.error(`dispatch-receipt: unpriced model ("${args.model}") or baseline ("${args.inherited}") β refusing to invent savings. Known: ${known}`);
|
|
104
|
+
process.exit(2);
|
|
105
|
+
}
|
|
106
|
+
const file = receiptsPath();
|
|
107
|
+
fs.mkdirSync(path.dirname(file), { recursive: true });
|
|
108
|
+
fs.appendFileSync(file, JSON.stringify(receipt) + '\n');
|
|
109
|
+
console.log(receiptLine(receipt.model, { cost: receipt.est_cost, frontier: receipt.est_frontier_cost, saved: receipt.saved, ref: receipt.inherited }));
|
|
110
|
+
}
|
|
111
|
+
|
|
112
|
+
if (process.argv[1] && import.meta.url === pathToFileURL(process.argv[1]).href) main();
|
|
@@ -0,0 +1,95 @@
|
|
|
1
|
+
#!/usr/bin/env node
|
|
2
|
+
// scripts/metaharness-receipts.mjs β plain-language routing receipts: what MetaHarness cheap-routing
|
|
3
|
+
// actually did and what it saved. Reads the REAL log written by scripts/route-cheap.mjs at
|
|
4
|
+
// ~/.claude/metaharness/routing-receipts.jsonl (override: METAHARNESS_RECEIPTS env, used by tests).
|
|
5
|
+
// No data β says so plainly. Never invents numbers (all costs are estimates from verified
|
|
6
|
+
// OpenRouter pricing + chars/4 token estimates, and are labeled "est.").
|
|
7
|
+
//
|
|
8
|
+
// Usage: node scripts/metaharness-receipts.mjs
|
|
9
|
+
|
|
10
|
+
import fs from 'node:fs';
|
|
11
|
+
import path from 'node:path';
|
|
12
|
+
import os from 'node:os';
|
|
13
|
+
import { pathToFileURL } from 'node:url';
|
|
14
|
+
|
|
15
|
+
export function receiptsPath() {
|
|
16
|
+
return (
|
|
17
|
+
process.env.METAHARNESS_RECEIPTS ||
|
|
18
|
+
path.join(os.homedir(), '.claude', 'metaharness', 'routing-receipts.jsonl')
|
|
19
|
+
);
|
|
20
|
+
}
|
|
21
|
+
|
|
22
|
+
// Parse the JSONL log. Corrupt lines are skipped (counted), never guessed at.
|
|
23
|
+
export function loadReceipts(file) {
|
|
24
|
+
let raw;
|
|
25
|
+
try {
|
|
26
|
+
raw = fs.readFileSync(file, 'utf8');
|
|
27
|
+
} catch {
|
|
28
|
+
return { rows: [], skipped: 0 };
|
|
29
|
+
}
|
|
30
|
+
const rows = [];
|
|
31
|
+
let skipped = 0;
|
|
32
|
+
for (const line of raw.split('\n')) {
|
|
33
|
+
if (!line.trim()) continue;
|
|
34
|
+
try {
|
|
35
|
+
const r = JSON.parse(line);
|
|
36
|
+
if (typeof r.saved === 'number' && r.model) rows.push(r);
|
|
37
|
+
else skipped++;
|
|
38
|
+
} catch {
|
|
39
|
+
skipped++;
|
|
40
|
+
}
|
|
41
|
+
}
|
|
42
|
+
return { rows, skipped };
|
|
43
|
+
}
|
|
44
|
+
|
|
45
|
+
const fmt$ = (n) => `$${n < 0.01 ? n.toFixed(5) : n.toFixed(4)}`;
|
|
46
|
+
|
|
47
|
+
export function formatTable(rows) {
|
|
48
|
+
if (!rows.length) return 'No routing receipts yet.\nRoute something cheap first: node scripts/route-cheap.mjs --task "<text>"';
|
|
49
|
+
|
|
50
|
+
// `channel` + `instead of` (2026-07-13): subagent receipts arrived with a per-row baseline β the
|
|
51
|
+
// model that agent WOULD have inherited β so a single global "frontier" column would misreport them.
|
|
52
|
+
const header = ['date', 'channel', 'task class', 'model used', 'instead of', 'est. cost', 'est. baseline', 'saved'];
|
|
53
|
+
const body = rows.map((r) => [
|
|
54
|
+
(r.ts || '').replace('T', ' ').slice(0, 16),
|
|
55
|
+
r.source === 'claude-subagent' ? 'subagent' : 'openrouter',
|
|
56
|
+
r.task_class || '?',
|
|
57
|
+
r.model,
|
|
58
|
+
r.frontier_ref || 'claude-opus-4.8',
|
|
59
|
+
fmt$(r.est_cost ?? 0),
|
|
60
|
+
fmt$(r.est_frontier_cost ?? 0),
|
|
61
|
+
fmt$(r.saved),
|
|
62
|
+
]);
|
|
63
|
+
const widths = header.map((h, i) => Math.max(h.length, ...body.map((row) => row[i].length)));
|
|
64
|
+
const line = (cells) => cells.map((cell, i) => cell.padEnd(widths[i])).join(' ');
|
|
65
|
+
|
|
66
|
+
const totalCost = rows.reduce((s, r) => s + (r.est_cost || 0), 0);
|
|
67
|
+
const totalFrontier = rows.reduce((s, r) => s + (r.est_frontier_cost || 0), 0);
|
|
68
|
+
const totalSaved = rows.reduce((s, r) => s + r.saved, 0);
|
|
69
|
+
const ratio = totalCost > 0 ? (totalFrontier / totalCost).toFixed(1) : '?';
|
|
70
|
+
|
|
71
|
+
// Baselines now vary per row; name them all rather than picking one and implying it covers everything.
|
|
72
|
+
const baselines = [...new Set(rows.map((r) => r.frontier_ref || 'claude-opus-4.8'))].join(', ');
|
|
73
|
+
const subagents = rows.filter((r) => r.source === 'claude-subagent').length;
|
|
74
|
+
|
|
75
|
+
return [
|
|
76
|
+
line(header),
|
|
77
|
+
line(widths.map((w) => '-'.repeat(w))),
|
|
78
|
+
...body.map(line),
|
|
79
|
+
'',
|
|
80
|
+
`${rows.length} routed task(s) (${subagents} subagent, ${rows.length - subagents} openrouter) Β· est. spent ${fmt$(totalCost)} vs ${fmt$(totalFrontier)} unrouted (${baselines}) Β· saved ~${fmt$(totalSaved)} (~${ratio}x cheaper)`,
|
|
81
|
+
'Pricing is live-verified. Token counts are measured OR estimated per row (each row records which, in token_source).',
|
|
82
|
+
'Baseline = the model the task would have run on if it had not been routed.',
|
|
83
|
+
].join('\n');
|
|
84
|
+
}
|
|
85
|
+
|
|
86
|
+
function main() {
|
|
87
|
+
const file = receiptsPath();
|
|
88
|
+
const { rows, skipped } = loadReceipts(file);
|
|
89
|
+
console.log(`MetaHarness routing receipts β ${file}`);
|
|
90
|
+
console.log('');
|
|
91
|
+
console.log(formatTable(rows));
|
|
92
|
+
if (skipped) console.log(`(${skipped} corrupt line(s) skipped)`);
|
|
93
|
+
}
|
|
94
|
+
|
|
95
|
+
if (process.argv[1] && import.meta.url === pathToFileURL(process.argv[1]).href) main();
|
|
@@ -39,7 +39,9 @@ import { estTokens } from './route-cheap.mjs'; // reuse the verified char/4 esti
|
|
|
39
39
|
|
|
40
40
|
const __dirname = path.dirname(fileURLToPath(import.meta.url));
|
|
41
41
|
export const CONFIG_DIR = path.join(os.homedir(), '.claude', 'model-router');
|
|
42
|
-
|
|
42
|
+
// Overridable for hermetic tests + CI (runners have no ~/.claude): the 2026-07-12 CI redness was
|
|
43
|
+
// exactly this β tests that silently depended on one developer's machine state.
|
|
44
|
+
const CATALOG_PATH = process.env.MODEL_ROUTER_CATALOG || path.join(CONFIG_DIR, 'catalog.json');
|
|
43
45
|
const POLICY_USER = path.join(CONFIG_DIR, 'policy.mjs');
|
|
44
46
|
const POLICY_DEFAULT = path.join(CONFIG_DIR, 'policy.default.mjs');
|
|
45
47
|
const DECISIONS_LOG =
|
package/scripts/route-cheap.mjs
CHANGED
|
@@ -40,6 +40,22 @@ export const PRICING = {
|
|
|
40
40
|
};
|
|
41
41
|
export const FRONTIER = { name: 'claude-opus-4.8', in: 5.0, out: 25.0 };
|
|
42
42
|
|
|
43
|
+
// Claude tiers β $/Mtok, verified live from the OpenRouter /models API 2026-07-13.
|
|
44
|
+
// These are NOT routed through here (Claude Code's own Agent/Task tool spawns them). They are priced
|
|
45
|
+
// so a SUBAGENT dispatch can get a real receipt: until 2026-07-13 the single biggest routing lever β
|
|
46
|
+
// every subagent inherits the main-loop model unless told otherwise β wrote no receipt at all, so the
|
|
47
|
+
// whole router looked unused. It WAS unused; it was also unmeasurable. Both had to be fixed.
|
|
48
|
+
// The spread is the whole argument: fable-5 costs 10x haiku-4.5 for identical mechanical work.
|
|
49
|
+
export const CLAUDE_TIERS = {
|
|
50
|
+
'claude-haiku-4.5': { in: 1.0, out: 5.0 },
|
|
51
|
+
'claude-sonnet-5': { in: 2.0, out: 10.0 },
|
|
52
|
+
'claude-opus-4.8': { in: 5.0, out: 25.0 },
|
|
53
|
+
'claude-fable-5': { in: 10.0, out: 50.0 },
|
|
54
|
+
};
|
|
55
|
+
|
|
56
|
+
/** Price lookup across both tables. Unknown model β null (never invent a savings number). */
|
|
57
|
+
export const priceOf = (model) => PRICING[model] || CLAUDE_TIERS[model] || null;
|
|
58
|
+
|
|
43
59
|
export function receiptsPath() {
|
|
44
60
|
return (
|
|
45
61
|
process.env.METAHARNESS_RECEIPTS ||
|
|
@@ -50,18 +66,23 @@ export function receiptsPath() {
|
|
|
50
66
|
// Honest token estimate: ~4 chars/token. Labeled "est." everywhere β never presented as measured.
|
|
51
67
|
export const estTokens = (text) => Math.max(1, Math.ceil((text || '').length / 4));
|
|
52
68
|
|
|
53
|
-
|
|
54
|
-
|
|
55
|
-
|
|
69
|
+
// `ref` = what this task WOULD have run on. For a cheap OpenRouter call that's the frontier default;
|
|
70
|
+
// for a subagent it's the model the agent would have INHERITED (i.e. the main-loop model), which is
|
|
71
|
+
// the only honest baseline β an un-overridden subagent really does run on whatever the session is on.
|
|
72
|
+
export function estimateCosts(model, inTokens, outTokens, ref = FRONTIER.name) {
|
|
73
|
+
const p = priceOf(model);
|
|
74
|
+
const r = priceOf(ref);
|
|
75
|
+
if (!p || !r) return null; // unknown model β no receipt rather than an invented number
|
|
56
76
|
const cost = (inTokens * p.in + outTokens * p.out) / 1e6;
|
|
57
|
-
const frontier = (inTokens *
|
|
58
|
-
return { cost, frontier, saved: frontier - cost };
|
|
77
|
+
const frontier = (inTokens * r.in + outTokens * r.out) / 1e6;
|
|
78
|
+
return { cost, frontier, saved: frontier - cost, ref };
|
|
59
79
|
}
|
|
60
80
|
|
|
61
81
|
const fmt$ = (n) => `$${n < 0.01 ? n.toFixed(5) : n.toFixed(4)}`;
|
|
62
82
|
|
|
63
83
|
export function receiptLine(model, costs) {
|
|
64
|
-
|
|
84
|
+
const ref = costs.ref && costs.ref !== FRONTIER.name ? costs.ref : 'frontier';
|
|
85
|
+
return `\x1b[2mβ‘ MetaHarness: routed to ${model} (est. ${fmt$(costs.cost)} vs ${fmt$(costs.frontier)} ${ref} β saved ~${fmt$(costs.saved)})\x1b[0m`;
|
|
65
86
|
}
|
|
66
87
|
|
|
67
88
|
// Load OPENROUTER_API_KEY from ruvnet-brain/.env if not already in env. Value never printed/logged.
|