nomarmy 0.1.0-alpha.2 → 0.1.0-alpha.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +33 -29
- package/lib/admission.mjs +17 -0
- package/lib/army.mjs +13 -3
- package/lib/coordinator-instructions.mjs +1 -0
- package/lib/health.mjs +14 -2
- package/lib/statusline.mjs +32 -6
- package/lib/usage-limits.mjs +126 -0
- package/mcp/server.mjs +7 -2
- package/package.json +1 -1
package/README.md
CHANGED
|
@@ -13,22 +13,24 @@
|
|
|
13
13
|
## TL;DR
|
|
14
14
|
|
|
15
15
|
1. **Have** Git, Node 20+ and [Podman](https://podman.io) (on macOS: `brew install podman && podman machine init && podman machine start`).
|
|
16
|
-
2. **Install** (
|
|
16
|
+
2. **Install** (OpenClaw and the sandbox, and registers nomArmy with Claude Code):
|
|
17
17
|
```bash
|
|
18
18
|
git clone https://github.com/rayson-tech/nomarmy.git && cd nomarmy
|
|
19
|
-
|
|
20
|
-
|
|
19
|
+
npm install && npm link
|
|
20
|
+
nomarmy setup --hosted
|
|
21
|
+
./install.sh --profile hosted
|
|
21
22
|
```
|
|
22
|
-
|
|
23
|
-
|
|
24
|
-
|
|
25
|
-
|
|
23
|
+
3. **Add your workers:** `nomarmy agents add` (an API key, or your ChatGPT or Muse Code subscription), then `nomarmy army init --agent <name>` to put every role on it. `nomarmy doctor` checks the lot.
|
|
24
|
+
4. **Set up your repo:** in the project, run `nomarmy init`. It proposes a `.nomarmy.yml` with your test command.
|
|
25
|
+
5. **Use it:** restart Claude Code in that project and ask it to use nomArmy for one small bug that has a test. When that works, try `/feature <what you want built>`.
|
|
26
|
+
|
|
27
|
+
**Have a GPU or a Mac with plenty of memory?** Workers can also run free on a local model: `./install.sh --profile macbook-pro` (or `nvidia-linux`, `cpu-linux`, `dgx-spark`) builds llama.cpp and starts it; see [Install](#install). A team GPU server works too: [a shared model server](#a-shared-model-server). Codex or Cursor as the coordinator: `nomarmy connect codex cursor`.
|
|
26
28
|
|
|
27
29
|
Stuck? `nomarmy doctor` checks the machine, and `nomarmy health` checks everything nomArmy runs on.
|
|
28
30
|
|
|
29
31
|
## What it is
|
|
30
32
|
|
|
31
|
-
Your coding assistant (Claude Code, Codex or Cursor) stays in charge as the **General**: it decides what gets built and whether the result is acceptable. The work goes to **noms**, workers that implement, test and repair in their own git worktree and sandbox, on
|
|
33
|
+
Your coding assistant (Claude Code, Codex or Cursor) stays in charge as the **General**: it decides what gets built and whether the result is acceptable. The work goes to **noms**, workers that implement, test and repair in their own git worktree and sandbox, on an API key, your own ChatGPT or Muse Code subscription, or a local model. nomArmy owns everything in between: worktrees, git, sandboxes, verification, and the evidence that decides whether work is accepted.
|
|
32
34
|
|
|
33
35
|
**What you get is work you don't have to take on faith**, not cheaper work. Delegating costs the General tokens too: briefing and reviewing. On small, already-diagnosed tickets we measured 4 to 8 times more of the General's tokens than fixing the bug directly, and break-even at roughly 150 lines of context a fix needs to read ([the measurements](docs/experiments/2026-09-20-model-bakeoff-and-economics.md)). It pays off on bigger tickets, on parallel work, and anywhere you'd otherwise have to trust an agent's say-so.
|
|
34
36
|
|
|
@@ -48,22 +50,38 @@ Around that core: **agents** say where a job can run, the **army** says which ro
|
|
|
48
50
|
|
|
49
51
|
## Install
|
|
50
52
|
|
|
51
|
-
|
|
|
53
|
+
| Setup | Guide |
|
|
52
54
|
|---|---|
|
|
53
|
-
|
|
|
54
|
-
|
|
|
55
|
+
| API keys and subscriptions, no local model (most people) | [Hosted workers only](#hosted-workers-only) |
|
|
56
|
+
| A local model on macOS (Apple Silicon) | [macOS](#macos-apple-silicon) |
|
|
57
|
+
| A local model on Linux, with or without an NVIDIA GPU | [Linux](#linux) |
|
|
55
58
|
| Windows | [Windows](#windows) |
|
|
56
59
|
| NVIDIA DGX Spark | [DGX Spark](#dgx-spark) |
|
|
57
|
-
| No local model: API keys and subscriptions only | [Hosted workers only](#hosted-workers-only) |
|
|
58
60
|
| A shared GPU server (or a tunnel to one) | [A shared model server](#a-shared-model-server) |
|
|
59
61
|
| No GPU, with Bedrock | [Cloud (Bedrock)](#cloud-bedrock) |
|
|
60
62
|
|
|
61
63
|
Every platform needs Git and Podman. `nomarmy doctor` checks the host and prints a fix for anything missing.
|
|
62
64
|
|
|
63
|
-
`install.sh` builds llama.cpp
|
|
65
|
+
`install.sh` builds llama.cpp when you run a local model, installs and configures [OpenClaw](https://github.com/openclaw/openclaw) (the host-side broker every model call goes through), builds the sandbox image, and registers the MCP server if Claude Code is installed. `nomarmy connect` (run by `install.sh`, or by hand for Codex and Cursor) also installs the `/feature` command, Claude Code's status line and, on macOS, nomArmy's notifier. The coordinator gets nomArmy's instructions from the MCP server itself, so there's nothing to copy into your projects.
|
|
64
66
|
|
|
65
67
|
**From npm:** `npm install -g nomarmy@alpha` gives you the `nomarmy` command; `nomarmy setup` then picks a profile and model and prints the `install.sh` command to run (`nomarmy setup --hosted` or `--llama-url <server>` without a local model). Installing from a clone, as in the TL;DR, is the most tested path.
|
|
66
68
|
|
|
69
|
+
### Hosted workers only
|
|
70
|
+
|
|
71
|
+
No GPU and no local model: every job runs on an API key or a subscription (ChatGPT, Muse Code) you add as an agent. Git worktrees, the sandbox and verification still run on your machine, so you still need Git, Node and Podman.
|
|
72
|
+
|
|
73
|
+
```bash
|
|
74
|
+
git clone https://github.com/rayson-tech/nomarmy.git && cd nomarmy
|
|
75
|
+
npm install && npm link # or: npm install -g nomarmy@alpha
|
|
76
|
+
nomarmy setup --hosted # records that this install has no local model
|
|
77
|
+
./install.sh --profile hosted # OpenClaw, the sandbox, and the Claude Code registration; no llama.cpp
|
|
78
|
+
nomarmy agents add # an API key or a subscription login
|
|
79
|
+
nomarmy army init --agent <name> # every role on that agent (add --model <model> to pick one)
|
|
80
|
+
nomarmy doctor
|
|
81
|
+
```
|
|
82
|
+
|
|
83
|
+
A hosted install refuses a job that names no role or agent, rather than falling back to a local model that isn't there. `nomarmy health` warns about any role still on `local`. `e2e.sh` tests the local model, so it has nothing to do here; `nomarmy army assign` tests each role's route instead.
|
|
84
|
+
|
|
67
85
|
### macOS (Apple Silicon)
|
|
68
86
|
|
|
69
87
|
```bash
|
|
@@ -122,22 +140,6 @@ chmod +x install.sh e2e.sh scripts/*.sh
|
|
|
122
140
|
|
|
123
141
|
Moving from a Mac install? Don't copy a Mac binary or model cache over: clone fresh and let `install.sh` build llama.cpp for CUDA on that machine.
|
|
124
142
|
|
|
125
|
-
### Hosted workers only
|
|
126
|
-
|
|
127
|
-
No GPU and no local model: every job runs on an API key or a subscription (ChatGPT, Muse Code) you add as an agent. Git worktrees, the sandbox and verification still run on your machine, so you still need Git, Node and Podman.
|
|
128
|
-
|
|
129
|
-
```bash
|
|
130
|
-
git clone https://github.com/rayson-tech/nomarmy.git && cd nomarmy
|
|
131
|
-
npm install && npm link # or: npm install -g nomarmy@alpha
|
|
132
|
-
nomarmy setup --hosted # records that this install has no local model
|
|
133
|
-
./install.sh --profile hosted # OpenClaw, the sandbox, and the Claude Code registration; no llama.cpp
|
|
134
|
-
nomarmy agents add # an API key or a subscription login
|
|
135
|
-
nomarmy army init --agent <name> # every role on that agent (add --model <model> to pick one)
|
|
136
|
-
nomarmy doctor
|
|
137
|
-
```
|
|
138
|
-
|
|
139
|
-
A hosted install refuses a job that names no role or agent, rather than falling back to a local model that isn't there. `nomarmy health` warns about any role still on `local`. `e2e.sh` tests the local model, so it has nothing to do here; `nomarmy army assign` tests each role's route instead.
|
|
140
|
-
|
|
141
143
|
### A shared model server
|
|
142
144
|
|
|
143
145
|
A team GPU box (a DGX, a workstation) runs one llama-server; everyone else points nomArmy at it. That works through an SSH tunnel too (`ssh -L 8080:localhost:8080 gpu-box`, then `http://127.0.0.1:8080`).
|
|
@@ -208,6 +210,8 @@ Changes apply to the next job with no restart. The exception is a **new** api ag
|
|
|
208
210
|
|
|
209
211
|
**Your plan decides which models run.** A model can be listed and still refused: on a ChatGPT plan, the Codex route runs gpt-6-astra and the gpt-5.6 models but refuses gpt-6-sol and gpt-6-luna. `army assign` and `agents update --probe` test the exact route a job takes, so they catch this before a job does.
|
|
210
212
|
|
|
213
|
+
**Usage limits.** nomArmy reads how much of a subscription's limit is used where the vendor reports it: Codex in every job's session log, Claude in what Claude Code passes to the status line (Pro and Max plans). It shows in `army`, `local_worker_capacity` and the status line (`⚠ openai 85% wk`), and `nomarmy health` warns from 80%. At the limit, a job on that agent is held rather than left to fail: the General asks you, and resubmits with `confirm_over_limit: true` if you say go. Muse, Grok and API keys don't report their limits; a job that hits one in a `/feature` run pauses that agent for the run.
|
|
214
|
+
|
|
211
215
|
**Vendor terms and platform risk.** Every model call goes through [OpenClaw](https://github.com/openclaw/openclaw), and subscriptions are reached through each vendor's own CLI or login. We've read the terms that apply (see above), but using a personal subscription through a harness is exactly the kind of use vendors tighten, and a change in a vendor's terms or in OpenClaw can stop a subscription agent from working. Local models and API keys don't carry that risk. Plan on subscriptions as a convenience, not the only way your roles can run.
|
|
212
216
|
|
|
213
217
|
**Picking an agent.** Build work goes to a sandboxed agent: `local`, an api key, Codex or Muse. `local` for a bounded change against a written spec with a test; your code never leaves your machine. An api or subscription agent when the work needs more than the local model, knowing it sends code to that vendor. That's a decision about where your source travels, separate from the trust boundary, which is the same for every agent. The General itself when the answer isn't known yet.
|
package/lib/admission.mjs
CHANGED
|
@@ -8,6 +8,7 @@ import { readJson } from "./openclaw-run.mjs";
|
|
|
8
8
|
import { readOpenClawTranscriptTail } from "./transcript.mjs";
|
|
9
9
|
import { readClaudeSessionTranscript } from "./claude-transcript.mjs";
|
|
10
10
|
import { notify } from "./notify.mjs";
|
|
11
|
+
import { readCodexRateLimits, recordUsageSnapshot, readUsageSnapshots, usageStatus } from "./usage-limits.mjs";
|
|
11
12
|
import { recentModelRefusal } from "./health.mjs";
|
|
12
13
|
import { writeLease, removeLease, liveLeases, liveSlots, acquireSlot } from "./slots.mjs";
|
|
13
14
|
import { loadRun, runTotals, runAdmissionProblems, recordRunJob, detectUsageLimit } from "./runs.mjs";
|
|
@@ -140,6 +141,11 @@ export function createJobRuntime(deps) {
|
|
|
140
141
|
* watching hears about it from any coordinator without polling.
|
|
141
142
|
*/
|
|
142
143
|
function notifyJobFinished(entry, result, error) {
|
|
144
|
+
try {
|
|
145
|
+
const provider = entry.agent ? agentProviderId(agentsConfig().agents[entry.agent]) : null;
|
|
146
|
+
const snapshot = provider ? readCodexRateLimits(path.join(jobsRoot, entry.jobId)) : null;
|
|
147
|
+
if (snapshot) recordUsageSnapshot(stateRoot, provider, snapshot);
|
|
148
|
+
} catch { /* Usage telemetry must never affect the job result. */ }
|
|
143
149
|
if (!entry.lane) return; // only tracked jobs, never internal helpers
|
|
144
150
|
const m = result?.manifest ?? {};
|
|
145
151
|
const outcome = error ? "failed" : String(m.outcome ?? (result?.ok ? "done" : "finished")).toLowerCase().replace(/_/g, " ");
|
|
@@ -161,6 +167,7 @@ export function createJobRuntime(deps) {
|
|
|
161
167
|
admission,
|
|
162
168
|
memory: deps.budgetState.hardwareSnapshot?.memory ?? null,
|
|
163
169
|
running: [...activeJobs.values()].filter(j => !j.settled).map(j => ({ jobId: j.jobId, workerId: j.workerId, mode: j.mode, lane: j.lane, startedAt: j.startedAt, phase: readJson(path.join(jobsRoot, j.jobId, "status.json"))?.phase ?? "starting" })),
|
|
170
|
+
usageLimits: Object.fromEntries(Object.entries(readUsageSnapshots(stateRoot)).map(([provider, snapshot]) => [provider, usageStatus(snapshot)])),
|
|
164
171
|
maxWorkers: currentMaxWorkers(),
|
|
165
172
|
remote: { running: runningCount("remote"), maxWorkers: currentMaxPoolWorkers(), note: "api and subscription agents; each agent's own max_concurrent also applies" }
|
|
166
173
|
};
|
|
@@ -169,6 +176,16 @@ export function createJobRuntime(deps) {
|
|
|
169
176
|
await deps.budgetState.refresh();
|
|
170
177
|
if (jobs.some((j) => jobLane(j) === "remote")) await modelCatalogReady();
|
|
171
178
|
const problems = [];
|
|
179
|
+
const snapshots = readUsageSnapshots(stateRoot);
|
|
180
|
+
jobs.forEach((j, i) => {
|
|
181
|
+
if (!j.agentName || j.confirm_over_limit === true) return;
|
|
182
|
+
let provider;
|
|
183
|
+
try { provider = agentProviderId(agentsConfig().agents[j.agentName]); } catch { return; }
|
|
184
|
+
const snapshot = snapshots[provider];
|
|
185
|
+
if (!snapshot) return;
|
|
186
|
+
const status = usageStatus(snapshot);
|
|
187
|
+
if (status.level === "over") problems.push(`${jobs.length > 1 ? `job ${i + 1}: ` : ""}agent "${j.agentName}" is held at its usage limit: ${status.text} (reading ${status.ageMinutes} minutes old). Ask the operator before resubmitting with confirm_over_limit: true, or send the job to another agent.`);
|
|
188
|
+
});
|
|
172
189
|
// A pool-routed job is checked against that pool's OWN (model-dependent)
|
|
173
190
|
// budget, not the local-derived global one -- see budgetsForPool. Which
|
|
174
191
|
// specific entry pickProvider will land on isn't known yet at admission
|
package/lib/army.mjs
CHANGED
|
@@ -30,6 +30,7 @@ import path from "node:path";
|
|
|
30
30
|
import YAML from "yaml";
|
|
31
31
|
import { z } from "zod";
|
|
32
32
|
import { agentRunsToolsOnHost } from "./dispatch-schema.mjs";
|
|
33
|
+
import { usageStatus } from "./usage-limits.mjs";
|
|
33
34
|
|
|
34
35
|
export const GLOBAL_CONFIG_FILENAME = "config.yml";
|
|
35
36
|
export const PROJECT_CONFIG_FILENAMES = Object.freeze([".nomarmy.yml", ".nomarmy.yaml"]);
|
|
@@ -361,11 +362,20 @@ export function generalOverlap(army, agents = {}) {
|
|
|
361
362
|
* description, phase, suggested mode, agent, any problem or overlap with
|
|
362
363
|
* the General, and which layer set each field.
|
|
363
364
|
*/
|
|
364
|
-
export function describeArmy(loaded, { agents = {}, describeAgent = null } = {}) {
|
|
365
|
+
export function describeArmy(loaded, { agents = {}, describeAgent = null, usageSnapshots = null, agentProviderId = null, now = Date.now() } = {}) {
|
|
365
366
|
const army = loaded.army;
|
|
366
367
|
const problems = armyTargetProblems(army, agents);
|
|
367
368
|
const overlap = generalOverlap(army, agents);
|
|
368
369
|
const runsOn = (name) => (name && agents[name] && describeAgent ? describeAgent(agents[name]) : null);
|
|
370
|
+
const usage = (name) => {
|
|
371
|
+
if (!name || !agents[name] || !usageSnapshots || !agentProviderId) return null;
|
|
372
|
+
let provider;
|
|
373
|
+
try { provider = agentProviderId(agents[name]); } catch { return null; }
|
|
374
|
+
const snapshot = usageSnapshots[provider];
|
|
375
|
+
if (!snapshot) return null;
|
|
376
|
+
const { level, text } = usageStatus(snapshot, now);
|
|
377
|
+
return { level, text };
|
|
378
|
+
};
|
|
369
379
|
// A subscription whose owner isn't the General's own is worth a word,
|
|
370
380
|
// not a refusal: it may be the same person's other account (a real Senti
|
|
371
381
|
// General skipped a role's agent for exactly this, unsure whose it was).
|
|
@@ -378,7 +388,7 @@ export function describeArmy(loaded, { agents = {}, describeAgent = null } = {})
|
|
|
378
388
|
const roles = Object.fromEntries(Object.entries(army.roles).map(([name, role]) => [name, {
|
|
379
389
|
description: role.description ?? null, phase: role.phase ?? null, mode: role.mode ?? null,
|
|
380
390
|
agent: role.agent ?? null, model: role.model ?? agents[role.agent]?.model ?? null, modelIsAuto: role.model === "auto",
|
|
381
|
-
agentRunsOn: runsOn(role.agent),
|
|
391
|
+
agentRunsOn: runsOn(role.agent), usage: usage(role.agent),
|
|
382
392
|
// Its agent's own tools run on this machine (dispatch-schema.mjs): an
|
|
383
393
|
// implement role there is refused per job unless the agent allows it.
|
|
384
394
|
hostTools: agentRunsToolsOnHost(agents[role.agent]) ? { allowed: agents[role.agent].allow_host_tools === true, implementRole: (role.mode ?? "implement") === "implement" } : null,
|
|
@@ -388,7 +398,7 @@ export function describeArmy(loaded, { agents = {}, describeAgent = null } = {})
|
|
|
388
398
|
? "not defined -- run `nomarmy army general <agent>` so nomArmy can flag roles that share the General's model or usage"
|
|
389
399
|
: !Object.prototype.hasOwnProperty.call(agents, army.general) ? `agent "${army.general}" is not defined in your agents.yml` : null;
|
|
390
400
|
return {
|
|
391
|
-
general: { ...GENERAL, agent: army.general, agentRunsOn: runsOn(army.general), problem: generalProblem, setBy: loaded.sources.general },
|
|
401
|
+
general: { ...GENERAL, agent: army.general, agentRunsOn: runsOn(army.general), usage: usage(army.general), problem: generalProblem, setBy: loaded.sources.general },
|
|
392
402
|
workflow: army.workflow,
|
|
393
403
|
runLimits: army.runLimits ?? {},
|
|
394
404
|
roles,
|
|
@@ -11,6 +11,7 @@ export const COORDINATOR_INSTRUCTIONS = `nomArmy runs bounded engineering jobs (
|
|
|
11
11
|
Before dispatching:
|
|
12
12
|
- Answer where-is / who-calls / grep questions with repo_evidence (deterministic, [path:line] on every hit). Use mode: scout only for read-only research that would otherwise pull many files into your own context.
|
|
13
13
|
- Call army to see this repo's roles and which agent each runs on; dispatch by army_role when a role fits. A job on a subscription agent needs on_behalf_of set to that agent's owner.
|
|
14
|
+
- Agents' usage limits show in army and local_worker_capacity; a job on an agent at its limit is held. Ask the operator before resubmitting with confirm_over_limit: true, or move the job to another agent. Never set it on your own.
|
|
14
15
|
- A Claude subscription agent (claude-cli) runs its tools on this machine, outside the sandbox: use it for scout and review work. nomArmy refuses implement jobs on it unless the operator set allow_host_tools; send build work to a sandboxed agent.
|
|
15
16
|
- Brief outcomes, not edits: a task, explicit acceptance criteria, and the tests that prove it. Put facts you've already resolved in evidence.
|
|
16
17
|
- Prefer local_worker_start + local_worker_status for anything longer than a few minutes. For a whole feature, use /feature (run_start keeps a run's jobs, spend and hours bounded).
|
package/lib/health.mjs
CHANGED
|
@@ -22,6 +22,7 @@ import path from "node:path";
|
|
|
22
22
|
import { executionMode } from "./execution.mjs";
|
|
23
23
|
import { modelRejection } from "./openclaw-errors.mjs";
|
|
24
24
|
import { providerConfigured, readOpenclawConfig } from "./openclaw-config.mjs";
|
|
25
|
+
import { readUsageSnapshots, usageStatus } from "./usage-limits.mjs";
|
|
25
26
|
|
|
26
27
|
const DAY = 86400000;
|
|
27
28
|
|
|
@@ -183,12 +184,23 @@ export function leftoverIssues({ retainedWorktrees = 0, jobsBytes = 0, staleRunn
|
|
|
183
184
|
* Run every check. `env` supplies what each needs, with real defaults;
|
|
184
185
|
* tests pass their own.
|
|
185
186
|
*/
|
|
186
|
-
export async function runHealthChecks({ now = Date.now(), mode = "local", openclawCmd = process.env.NOMARMY_OPENCLAW_CMD || "openclaw", run = runBounded, armySummary = null, agentsError = null, jobsRoot = null, pidAlive = () => true, agents = null, openclawConfig = null, vendors = {}, modelsInUse = null, autoPruned = null } = {}) {
|
|
187
|
+
export async function runHealthChecks({ now = Date.now(), mode = "local", openclawCmd = process.env.NOMARMY_OPENCLAW_CMD || "openclaw", run = runBounded, armySummary = null, agentsError = null, jobsRoot = null, pidAlive = () => true, agents = null, openclawConfig = null, vendors = {}, modelsInUse = null, autoPruned = null, usageSnapshots = null } = {}) {
|
|
187
188
|
const issues = [];
|
|
188
189
|
if (autoPruned?.freedBytes) issues.push({ id: `auto-prune:${new Date(now).toISOString()}`, severity: "info",
|
|
189
190
|
title: `Freed ${(autoPruned.freedBytes / 1024 ** 3).toFixed(2)} GB: ${[autoPruned.pruned ? `runtime data of ${autoPruned.pruned} finished job${autoPruned.pruned === 1 ? "" : "s"} older than ${autoPruned.olderThanHours}h` : null, autoPruned.scratchCleared ? `OpenClaw scratch files of ${autoPruned.scratchCleared} more` : null].filter(Boolean).join(", ")}`,
|
|
190
191
|
detail: "Automatic; each job's record and report are kept. NOMARMY_AUTO_PRUNE_HOURS sets the age (0 turns it off).", fix: null, short: null });
|
|
191
192
|
if (agents) issues.push(...providerConfigIssues({ agents, openclawConfig, vendors }));
|
|
193
|
+
if (usageSnapshots) {
|
|
194
|
+
for (const [provider, snapshot] of Object.entries(usageSnapshots)) {
|
|
195
|
+
const status = usageStatus(snapshot, now);
|
|
196
|
+
if (status.level === "ok") continue;
|
|
197
|
+
const short = `${provider} ${status.short}`;
|
|
198
|
+
issues.push({ id: `usage:${provider}:${status.level}`, severity: "warn", title: status.level === "over" ? `${provider} is at its usage limit` : `${provider} usage is ${status.short}`,
|
|
199
|
+
detail: `${status.text} (reading ${status.ageMinutes} minutes old).`,
|
|
200
|
+
fix: status.level === "over" ? "wait for the reset, or move its roles with nomarmy army assign" : "plan remaining work or move its roles with nomarmy army assign",
|
|
201
|
+
short });
|
|
202
|
+
}
|
|
203
|
+
}
|
|
192
204
|
const [auth, version, latest, plugins] = await Promise.all([
|
|
193
205
|
run(openclawCmd, ["models", "auth", "list", "--json"]),
|
|
194
206
|
run(openclawCmd, ["--version"]),
|
|
@@ -278,7 +290,7 @@ export async function checkAndRecordHealth({ projectDir, stateRoot, configDir, n
|
|
|
278
290
|
if (ageMs !== null) { try { autoPruned = { ...pruneJobRuntime({ stateRoot, olderThanMs: ageMs, now }), olderThanHours: ageMs / 3600000 }; } catch { /* best-effort */ } }
|
|
279
291
|
const mode = executionMode(env).mode;
|
|
280
292
|
const result = await runHealthChecks({ now, mode, armySummary, agentsError, jobsRoot: path.join(stateRoot, "jobs"), pidAlive,
|
|
281
|
-
agents, openclawConfig: readOpenclawConfig(), vendors: SUBSCRIPTION_VENDORS, modelsInUse, autoPruned });
|
|
293
|
+
agents, openclawConfig: readOpenclawConfig(), vendors: SUBSCRIPTION_VENDORS, modelsInUse, autoPruned, usageSnapshots: readUsageSnapshots(stateRoot) });
|
|
282
294
|
const toNotify = recordHealth(path.join(stateRoot, "health.json"), result, { now });
|
|
283
295
|
return { result, toNotify };
|
|
284
296
|
}
|
package/lib/statusline.mjs
CHANGED
|
@@ -27,6 +27,7 @@ import fs from "node:fs";
|
|
|
27
27
|
import os from "node:os";
|
|
28
28
|
import path from "node:path";
|
|
29
29
|
import { fileURLToPath } from "node:url";
|
|
30
|
+
import { mergeUsageSnapshot, normalizeClaudeRateLimits, readUsageSnapshots, recordUsageSnapshot, usageStatus } from "./usage-limits.mjs";
|
|
30
31
|
|
|
31
32
|
function readJson(file) { try { return JSON.parse(fs.readFileSync(file, "utf8")); } catch { return null; } }
|
|
32
33
|
function pidAlive(pid) {
|
|
@@ -47,6 +48,17 @@ const minutes = (ms) => (ms < 60000 ? `${Math.max(1, Math.round(ms / 1000))}s` :
|
|
|
47
48
|
*/
|
|
48
49
|
export function statusLineText({ session = {}, stateRoot, now = Date.now(), maxLength = Number(process.env.NOMARMY_STATUSLINE_MAX) || 90 } = {}) {
|
|
49
50
|
const root = stateRoot ?? (process.env.NOMARMY_AGENT_STATE || path.join(os.homedir(), ".local", "share", "nomarmy-local-agents"));
|
|
51
|
+
// Every Claude Code window redraws this line with its own last-seen
|
|
52
|
+
// limits, idle ones included, so a reading is merged rather than trusted
|
|
53
|
+
// (mergeUsageSnapshot), and written only when it moves the figure.
|
|
54
|
+
try {
|
|
55
|
+
const snapshot = normalizeClaudeRateLimits(session.rate_limits);
|
|
56
|
+
if (snapshot) {
|
|
57
|
+
const previous = readUsageSnapshots(root)["claude-cli"];
|
|
58
|
+
const merged = mergeUsageSnapshot(previous, { ...snapshot, observedAt: now }, now);
|
|
59
|
+
if (merged !== previous) recordUsageSnapshot(root, "claude-cli", merged);
|
|
60
|
+
}
|
|
61
|
+
} catch { /* usage capture must never break the status line */ }
|
|
50
62
|
// workspace.project_dir is where Claude Code was launched -- the repo that
|
|
51
63
|
// session's nomArmy server works in; current_dir can be a subdirectory.
|
|
52
64
|
const cwd = session.workspace?.project_dir ?? session.workspace?.current_dir ?? session.cwd ?? process.cwd();
|
|
@@ -99,23 +111,37 @@ export function statusLineText({ session = {}, stateRoot, now = Date.now(), maxL
|
|
|
99
111
|
try {
|
|
100
112
|
const health = readJson(path.join(root, "health.json"));
|
|
101
113
|
if (health && now - Date.parse(health.checkedAt) < 2 * 86400000) {
|
|
102
|
-
|
|
114
|
+
// Usage warnings get their own part below, so they aren't repeated here.
|
|
115
|
+
const top = (health.issues ?? []).find((i) => i.severity !== "info" && i.short && !String(i.id).startsWith("usage:"));
|
|
103
116
|
if (top) healthPart = ` │ ⚠ ${top.short}`;
|
|
104
117
|
}
|
|
105
118
|
} catch { /* no health yet */ }
|
|
106
119
|
runPart += healthPart;
|
|
107
|
-
|
|
120
|
+
let usageParts = [];
|
|
121
|
+
try {
|
|
122
|
+
const warnings = Object.entries(readUsageSnapshots(root)).map(([provider, snapshot]) => {
|
|
123
|
+
const status = usageStatus(snapshot, now);
|
|
124
|
+
if (status.level === "ok") return null;
|
|
125
|
+
return { level: status.level, text: ` │ ${status.level === "over" ? "⛔" : "⚠"} ${provider} ${status.short}` };
|
|
126
|
+
}).filter(Boolean).sort((a, b) => (a.level === "over" ? 0 : 1) - (b.level === "over" ? 0 : 1));
|
|
127
|
+
usageParts = warnings.map((warning) => warning.text);
|
|
128
|
+
} catch { /* usage display is best-effort */ }
|
|
129
|
+
// As many jobs as fit, then "+N". Only after every job is collapsed may
|
|
130
|
+
// lower-priority usage warnings be omitted to honor the hard cap.
|
|
108
131
|
const prefix = `${head ? `${head} │ ` : ""}🍪 `;
|
|
109
132
|
const count = jobs.length > 1 ? `${jobs.length}: ` : "";
|
|
110
|
-
let shown = jobs.length, army;
|
|
133
|
+
let shown = jobs.length, army, suffix;
|
|
111
134
|
for (;;) {
|
|
112
135
|
const rest = jobs.length - shown;
|
|
113
136
|
army = (!jobs.length ? "idle" : `${count}${jobs.slice(0, shown).join(" · ")}${rest ? `${shown ? " " : ""}+${rest}` : ""}`)
|
|
114
137
|
+ (elsewhere ? ` · ${elsewhere} in other repo${elsewhere === 1 ? "" : "s"}` : "");
|
|
115
|
-
|
|
116
|
-
|
|
138
|
+
suffix = runPart + usageParts.join("");
|
|
139
|
+
if ([...`${prefix}${army}${suffix}`].length <= maxLength) break;
|
|
140
|
+
if (shown > 0) { shown--; continue; }
|
|
141
|
+
if (usageParts.length) { usageParts.pop(); continue; }
|
|
142
|
+
break;
|
|
117
143
|
}
|
|
118
|
-
return `${prefix}${army}${
|
|
144
|
+
return `${prefix}${army}${suffix}`;
|
|
119
145
|
}
|
|
120
146
|
|
|
121
147
|
const isMain = (() => { try { return path.resolve(process.argv[1] ?? "") === fileURLToPath(import.meta.url); } catch { return false; } })();
|
|
@@ -0,0 +1,126 @@
|
|
|
1
|
+
import fs from "node:fs";
|
|
2
|
+
import path from "node:path";
|
|
3
|
+
import { randomUUID } from "node:crypto";
|
|
4
|
+
|
|
5
|
+
// How much of each subscription's usage limit is used, as its vendor reports
|
|
6
|
+
// it: Codex writes rate_limits into every job's session log, and Claude Code
|
|
7
|
+
// passes them to the status line. Kept per OpenClaw provider id (one host
|
|
8
|
+
// login per provider) in <stateRoot>/usage-limits.json. Admission holds a job
|
|
9
|
+
// on an agent that's over its limit until the General confirms.
|
|
10
|
+
|
|
11
|
+
const object = value => value !== null && typeof value === "object" && !Array.isArray(value);
|
|
12
|
+
const finite = value => typeof value === "number" && Number.isFinite(value);
|
|
13
|
+
const windowName = minutes => ({ 300: "5h", 10080: "week", 1440: "day" })[minutes] ?? `${minutes}m`;
|
|
14
|
+
function window(value, name, minutes, percentKey) {
|
|
15
|
+
if (!object(value) || !finite(value[percentKey]) || value[percentKey] < 0) return null;
|
|
16
|
+
return { name, usedPercent: value[percentKey], windowMinutes: minutes,
|
|
17
|
+
resetsAt: finite(value.resets_at) ? value.resets_at * 1000 : null };
|
|
18
|
+
}
|
|
19
|
+
|
|
20
|
+
/** Read the last rate-limit event from the newest rollout (by modification time). */
|
|
21
|
+
export function readCodexRateLimits(jobDir) {
|
|
22
|
+
try {
|
|
23
|
+
const files = [];
|
|
24
|
+
function walk(dir) {
|
|
25
|
+
for (const entry of fs.readdirSync(dir, { withFileTypes: true })) {
|
|
26
|
+
const file = path.join(dir, entry.name);
|
|
27
|
+
if (entry.isDirectory()) walk(file);
|
|
28
|
+
else if (entry.isFile() && /^rollout-.*\.jsonl$/.test(entry.name)) files.push({ file, mtime: fs.statSync(file).mtimeMs });
|
|
29
|
+
}
|
|
30
|
+
}
|
|
31
|
+
const agents = path.join(jobDir, "runtime/state/agents");
|
|
32
|
+
for (const agent of fs.readdirSync(agents, { withFileTypes: true })) {
|
|
33
|
+
if (!agent.isDirectory()) continue;
|
|
34
|
+
const sessions = path.join(agents, agent.name, "agent/codex-home/sessions");
|
|
35
|
+
if (fs.existsSync(sessions)) walk(sessions);
|
|
36
|
+
}
|
|
37
|
+
files.sort((a, b) => b.mtime - a.mtime || b.file.localeCompare(a.file));
|
|
38
|
+
if (!files.length) return null;
|
|
39
|
+
let last = null;
|
|
40
|
+
for (const line of fs.readFileSync(files[0].file, "utf8").split(/\r?\n/)) {
|
|
41
|
+
try {
|
|
42
|
+
const event = JSON.parse(line);
|
|
43
|
+
if (event.type === "event_msg" && event.payload?.type === "token_count" && object(event.payload.rate_limits)) last = event;
|
|
44
|
+
} catch { /* A partial trailing write is not an event. */ }
|
|
45
|
+
}
|
|
46
|
+
if (!last) return null;
|
|
47
|
+
const limits = last.payload.rate_limits;
|
|
48
|
+
const windows = [limits.primary, limits.secondary].map(value => {
|
|
49
|
+
const minutes = finite(value?.window_minutes) ? value.window_minutes : null;
|
|
50
|
+
return window(value, minutes === null ? "unknown" : windowName(minutes), minutes, "used_percent");
|
|
51
|
+
}).filter(Boolean);
|
|
52
|
+
const limitReached = limits.rate_limit_reached_type != null || Boolean(limits.spend_control_reached);
|
|
53
|
+
if (!windows.length && !limitReached) return null;
|
|
54
|
+
const timestamp = Date.parse(last.timestamp);
|
|
55
|
+
return { source: "codex", plan: typeof limits.plan_type === "string" ? limits.plan_type : null,
|
|
56
|
+
limitReached, observedAt: Number.isFinite(timestamp) ? timestamp : files[0].mtime, windows };
|
|
57
|
+
} catch { return null; }
|
|
58
|
+
}
|
|
59
|
+
|
|
60
|
+
export function normalizeClaudeRateLimits(rateLimits) {
|
|
61
|
+
if (!object(rateLimits)) return null;
|
|
62
|
+
const windows = [["five_hour", "5h", 300], ["seven_day", "week", 10080], ["spend_limit", "spend", null]]
|
|
63
|
+
.map(([key, name, minutes]) => window(rateLimits[key], name, minutes, "used_percentage")).filter(Boolean);
|
|
64
|
+
// Claude supplies percentages, not an independent provider-wide reached flag.
|
|
65
|
+
return windows.length ? { source: "claude", plan: null, limitReached: false, observedAt: Date.now(), windows } : null;
|
|
66
|
+
}
|
|
67
|
+
|
|
68
|
+
/**
|
|
69
|
+
* Combine a stored snapshot with a new reading that may be stale. Every
|
|
70
|
+
* Claude Code window passes the status line the limits from its own last
|
|
71
|
+
* response, so an idle window keeps reporting an old, lower figure. Usage
|
|
72
|
+
* only rises within a window, so per window: a later reset wins, and at the
|
|
73
|
+
* same reset the higher figure wins. A stored window the reading lacks is
|
|
74
|
+
* kept while it's live. Returns the stored snapshot itself when nothing
|
|
75
|
+
* changed, so callers can skip the write.
|
|
76
|
+
*/
|
|
77
|
+
export function mergeUsageSnapshot(previous, incoming, now = Date.now()) {
|
|
78
|
+
if (!previous) return incoming;
|
|
79
|
+
const key = (w) => [w.resetsAt ?? 0, w.usedPercent];
|
|
80
|
+
const newer = (a, b) => { const [ra, pa] = key(a), [rb, pb] = key(b); return ra > rb || (ra === rb && pa > pb); };
|
|
81
|
+
const byName = new Map(previous.windows.filter((w) => w.resetsAt === null || w.resetsAt > now).map((w) => [w.name, w]));
|
|
82
|
+
let changed = byName.size !== previous.windows.length;
|
|
83
|
+
for (const w of incoming.windows) {
|
|
84
|
+
const stored = byName.get(w.name);
|
|
85
|
+
if (!stored || newer(w, stored)) { byName.set(w.name, w); changed = true; }
|
|
86
|
+
}
|
|
87
|
+
const limitReached = incoming.limitReached || previous.limitReached;
|
|
88
|
+
if (!changed && limitReached === previous.limitReached) return previous;
|
|
89
|
+
return { ...incoming, limitReached, observedAt: now, windows: [...byName.values()] };
|
|
90
|
+
}
|
|
91
|
+
|
|
92
|
+
export function readUsageSnapshots(stateRoot) {
|
|
93
|
+
try {
|
|
94
|
+
const data = JSON.parse(fs.readFileSync(path.join(stateRoot, "usage-limits.json"), "utf8"));
|
|
95
|
+
if (!object(data)) return {};
|
|
96
|
+
return Object.fromEntries(Object.entries(data).filter(([, s]) => object(s) && ["codex", "claude"].includes(s.source)
|
|
97
|
+
&& finite(s.observedAt) && typeof s.limitReached === "boolean" && (s.plan === null || typeof s.plan === "string")
|
|
98
|
+
&& Array.isArray(s.windows) && s.windows.every(w => object(w) && typeof w.name === "string" && finite(w.usedPercent)
|
|
99
|
+
&& (w.windowMinutes === null || finite(w.windowMinutes)) && (w.resetsAt === null || finite(w.resetsAt)))));
|
|
100
|
+
} catch { return {}; }
|
|
101
|
+
}
|
|
102
|
+
|
|
103
|
+
/** Save a provider's snapshot, unless the one on file is newer (jobs finish out of order). */
|
|
104
|
+
export function recordUsageSnapshot(stateRoot, provider, snapshot) {
|
|
105
|
+
const current = readUsageSnapshots(stateRoot);
|
|
106
|
+
if (current[provider]?.observedAt > snapshot.observedAt) return;
|
|
107
|
+
const snapshots = { ...current, [provider]: snapshot };
|
|
108
|
+
fs.mkdirSync(stateRoot, { recursive: true });
|
|
109
|
+
const file = path.join(stateRoot, "usage-limits.json"), tmp = `${file}.${randomUUID()}.tmp`;
|
|
110
|
+
try {
|
|
111
|
+
fs.writeFileSync(tmp, JSON.stringify(snapshots, null, 2));
|
|
112
|
+
fs.renameSync(tmp, file);
|
|
113
|
+
} finally { fs.rmSync(tmp, { force: true }); }
|
|
114
|
+
}
|
|
115
|
+
|
|
116
|
+
export function usageStatus(snapshot, now = Date.now()) {
|
|
117
|
+
const live = snapshot.windows.filter(w => w.resetsAt === null || w.resetsAt > now).sort((a, b) => b.usedPercent - a.usedPercent);
|
|
118
|
+
const reached = snapshot.limitReached && (snapshot.windows.length === 0 || live.length > 0);
|
|
119
|
+
const highest = live[0];
|
|
120
|
+
const level = reached || highest?.usedPercent >= 100 ? "over" : highest?.usedPercent >= 80 ? "high" : "ok";
|
|
121
|
+
const text = live.length ? live.map(w => `${w.usedPercent}% of ${w.name}, resets ${w.resetsAt === null ? "unknown" : new Date(w.resetsAt).toLocaleString("en-US", { weekday: "short", hour: "2-digit", minute: "2-digit", hour12: false })}`).join("; ") : reached ? "limit reached, reset unknown" : "no live usage windows";
|
|
122
|
+
// A short label for tight spaces (status line, health): "85% wk".
|
|
123
|
+
const short = highest ? `${highest.usedPercent}% ${highest.name === "week" ? "wk" : highest.name}` : reached ? "limit" : null;
|
|
124
|
+
return { level, text, short, resetsAt: level === "over" ? highest?.resetsAt ?? null : null,
|
|
125
|
+
ageMinutes: Math.max(0, Math.floor((now - snapshot.observedAt) / 60000)) };
|
|
126
|
+
}
|
package/mcp/server.mjs
CHANGED
|
@@ -32,6 +32,7 @@ import { liveLeases } from "../lib/slots.mjs";
|
|
|
32
32
|
import { createRun, loadRun, runTotals, finishRun, resolveRunLimits, describeLoweredLimits } from "../lib/runs.mjs";
|
|
33
33
|
import { agentDispatchFields, resolveAgentModel, agentProviderId, describeAgent } from "../lib/agents.mjs";
|
|
34
34
|
import { OUTCOMES, COORDINATOR_STATUS_BY_OUTCOME } from "../lib/outcomes.mjs";
|
|
35
|
+
import { readUsageSnapshots, usageStatus } from "../lib/usage-limits.mjs";
|
|
35
36
|
import { createBuildMetrics, resolveOutcome, finalText, workerMetadata, usageMetrics, policyAdmissionProblems, applyRefactorContract, applyVerificationPolicy, resolveVerifyRegression } from "../lib/outcome.mjs";
|
|
36
37
|
import { compactJobRecord, formatResult, formatUnion, testChangeBanner, regressionCheckBanner, decomposeOverlapBanner } from "../lib/job-format.mjs";
|
|
37
38
|
|
|
@@ -282,6 +283,7 @@ export const jobSchema = z.object({
|
|
|
282
283
|
report: z.enum(["brief", "standard", "full"]).optional().describe("How much the worker may report back, capped by its agent's tier: brief (today's local-sized report), standard (the default), full (the frontier ceiling: about 2k tokens for implement, 4k for a scout). The report lands in your own context and is re-read every later turn, so ask for full only when the job's findings are the point (a broad review). No effect on the local model, whose caps are calibrated."),
|
|
283
284
|
commit_subject: z.string().max(200).optional().describe("implement: the subject line of the commit nomArmy makes on the worker branch, e.g. \"Keep held-back tables in the list_tables cache\". Defaults to the task's first sentence; the body is the worker's NOTE, and the job id is a trailer."),
|
|
284
285
|
army_role: z.string().regex(/^[a-z][a-z0-9-]{0,63}$/).optional().describe("Dispatch by army role (e.g. \"sr-dev\", \"security-analyst\"): nomArmy runs it on the agent this repo assigns to that role and puts the role's description at the top of the brief. Call the `army` tool first to see this repo's roles. Mutually exclusive with agent. Add on_behalf_of in case the role's agent is a subscription; it's ignored otherwise."),
|
|
286
|
+
confirm_over_limit: z.boolean().optional().describe("Override a reached usage limit: the General must ask the operator before resubmitting with confirm_over_limit: true, or send the job to another agent. nomArmy never sets it itself."),
|
|
285
287
|
on_behalf_of: z.string().min(1).max(254).optional().describe("Required when the job's agent is a subscription: must exactly match that agent's owner in agents.yml, or nomArmy refuses the job. A self-reported attestation, not an independently verified identity check -- nomArmy has no caller-identity boundary today, so what this guarantees is explicit, auditable intent and hard refusal on mismatch or omission, not cryptographic proof of who issued the call. Ignored for a local or api agent."),
|
|
286
288
|
evidence: z.string().max(maxEvidenceChars,
|
|
287
289
|
`Evidence exceeds the ${maxEvidenceChars}-character budget. This is for facts already resolved (e.g. with repo_evidence), not more description of the task -- if it needs more than this, resolve less per job or put the pointer (a path and line range) here instead of the material itself.`
|
|
@@ -474,14 +476,17 @@ server.tool("run_finish", "Close a /feature run as complete or stopped, with a o
|
|
|
474
476
|
server.tool("army", "Who you, the General, are and who you call for what in this repository: your fixed charter and the agent you're defined as, the army's workflow, then each role's description, phase (build, review, acceptance), suggested mode, and the agent it runs on, with which config layer set each value (global, project .nomarmy.yml, local .nomarmy.local.yml). Flags roles with no usable agent, and roles that share your model or subscription (not an independent review). Dispatch a role with `army_role`, or an agent directly with `agent`. Read-only, re-read on every call.", {}, async () => {
|
|
475
477
|
try {
|
|
476
478
|
const agents = agentsConfig().agents;
|
|
477
|
-
const
|
|
479
|
+
const usageSnapshots = readUsageSnapshots(stateRoot);
|
|
480
|
+
const summary = describeArmy(currentArmy(), { agents, describeAgent, usageSnapshots, agentProviderId });
|
|
478
481
|
// Each agent's models, from OpenClaw's catalog, so the General can pick
|
|
479
482
|
// one for a role set to "auto". The catalog can lag a brand-new model.
|
|
480
483
|
const catalog = await modelCatalogReady();
|
|
481
484
|
summary.agents = Object.fromEntries(Object.entries(agents).map(([name, agent]) => {
|
|
482
485
|
const provider = agentProviderId(agent);
|
|
483
486
|
const models = provider && catalog ? [...catalog.keys()].filter((k) => k.startsWith(`${provider}/`)).map((k) => k.slice(provider.length + 1)) : [];
|
|
484
|
-
|
|
487
|
+
const snapshot = usageSnapshots[provider];
|
|
488
|
+
const usage = snapshot ? (() => { const { level, text } = usageStatus(snapshot); return { level, text }; })() : null;
|
|
489
|
+
return [name, { runsOn: describeAgent(agent), defaultModel: agent.model ?? null, models, usage }];
|
|
485
490
|
}));
|
|
486
491
|
// A pinned model missing from the catalog isn't necessarily wrong:
|
|
487
492
|
// `army assign` proves an unlisted model with a real test call, and the
|
package/package.json
CHANGED
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
"description": "A harness for AI coding workers whose claims are never trusted: your coding assistant stays in charge while workers implement and test in sandboxes, on local models, API keys or your own subscriptions.",
|
|
4
4
|
"author": "Rayson Technologies",
|
|
5
5
|
"license": "Apache-2.0",
|
|
6
|
-
"version": "0.1.0-alpha.
|
|
6
|
+
"version": "0.1.0-alpha.3",
|
|
7
7
|
"private": false,
|
|
8
8
|
"type": "module",
|
|
9
9
|
"engines": {
|