nomarmy 0.1.0-alpha.2 → 0.1.0-alpha.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -13,22 +13,24 @@
13
13
  ## TL;DR
14
14
 
15
15
  1. **Have** Git, Node 20+ and [Podman](https://podman.io) (on macOS: `brew install podman && podman machine init && podman machine start`).
16
- 2. **Install** (builds the local model server, OpenClaw and the sandbox, and registers nomArmy with Claude Code):
16
+ 2. **Install** (OpenClaw and the sandbox, and registers nomArmy with Claude Code):
17
17
  ```bash
18
18
  git clone https://github.com/rayson-tech/nomarmy.git && cd nomarmy
19
- ./install.sh --profile macbook-pro # or nvidia-linux, cpu-linux, dgx-spark, bedrock
20
- ./e2e.sh --profile macbook-pro # should end with: === E2E PASS ===
19
+ npm install && npm link
20
+ nomarmy setup --hosted
21
+ ./install.sh --profile hosted
21
22
  ```
22
- No local model? See [hosted workers only](#hosted-workers-only) or [a shared model server](#a-shared-model-server).
23
- 3. **Set up your repo:** in the project, run `nomarmy init`. It proposes a `.nomarmy.yml` with your test command.
24
- 4. **Use it:** restart Claude Code in that project and ask it to use nomArmy for one small bug that has a test. When that works, try `/feature <what you want built>`.
25
- 5. **Optional:** add hosted workers with `nomarmy agents add`, give roles to them with `nomarmy army init`, or connect Codex or Cursor with `nomarmy connect codex cursor`.
23
+ 3. **Add your workers:** `nomarmy agents add` (an API key, or your ChatGPT or Muse Code subscription), then `nomarmy army init --agent <name>` to put every role on it. `nomarmy doctor` checks the lot.
24
+ 4. **Set up your repo:** in the project, run `nomarmy init`. It proposes a `.nomarmy.yml` with your test command.
25
+ 5. **Use it:** restart Claude Code in that project and ask it to use nomArmy for one small bug that has a test. When that works, try `/feature <what you want built>`.
26
+
27
+ **Have a GPU or a Mac with plenty of memory?** Workers can also run free on a local model: `./install.sh --profile macbook-pro` (or `nvidia-linux`, `cpu-linux`, `dgx-spark`) builds llama.cpp and starts it; see [Install](#install). A team GPU server works too: [a shared model server](#a-shared-model-server). Codex or Cursor as the coordinator: `nomarmy connect codex cursor`.
26
28
 
27
29
  Stuck? `nomarmy doctor` checks the machine, and `nomarmy health` checks everything nomArmy runs on.
28
30
 
29
31
  ## What it is
30
32
 
31
- Your coding assistant (Claude Code, Codex or Cursor) stays in charge as the **General**: it decides what gets built and whether the result is acceptable. The work goes to **noms**, workers that implement, test and repair in their own git worktree and sandbox, on a local model, an API key, or your own ChatGPT or Muse Code subscription. nomArmy owns everything in between: worktrees, git, sandboxes, verification, and the evidence that decides whether work is accepted.
33
+ Your coding assistant (Claude Code, Codex or Cursor) stays in charge as the **General**: it decides what gets built and whether the result is acceptable. The work goes to **noms**, workers that implement, test and repair in their own git worktree and sandbox, on an API key, your own ChatGPT or Muse Code subscription, or a local model. nomArmy owns everything in between: worktrees, git, sandboxes, verification, and the evidence that decides whether work is accepted.
32
34
 
33
35
  **What you get is work you don't have to take on faith**, not cheaper work. Delegating costs the General tokens too: briefing and reviewing. On small, already-diagnosed tickets we measured 4 to 8 times more of the General's tokens than fixing the bug directly, and break-even at roughly 150 lines of context a fix needs to read ([the measurements](docs/experiments/2026-09-20-model-bakeoff-and-economics.md)). It pays off on bigger tickets, on parallel work, and anywhere you'd otherwise have to trust an agent's say-so.
34
36
 
@@ -48,22 +50,38 @@ Around that core: **agents** say where a job can run, the **army** says which ro
48
50
 
49
51
  ## Install
50
52
 
51
- | Platform | Guide |
53
+ | Setup | Guide |
52
54
  |---|---|
53
- | macOS (Apple Silicon) | [macOS](#macos-apple-silicon) |
54
- | Linux, with or without an NVIDIA GPU | [Linux](#linux) |
55
+ | API keys and subscriptions, no local model (most people) | [Hosted workers only](#hosted-workers-only) |
56
+ | A local model on macOS (Apple Silicon) | [macOS](#macos-apple-silicon) |
57
+ | A local model on Linux, with or without an NVIDIA GPU | [Linux](#linux) |
55
58
  | Windows | [Windows](#windows) |
56
59
  | NVIDIA DGX Spark | [DGX Spark](#dgx-spark) |
57
- | No local model: API keys and subscriptions only | [Hosted workers only](#hosted-workers-only) |
58
60
  | A shared GPU server (or a tunnel to one) | [A shared model server](#a-shared-model-server) |
59
61
  | No GPU, with Bedrock | [Cloud (Bedrock)](#cloud-bedrock) |
60
62
 
61
63
  Every platform needs Git and Podman. `nomarmy doctor` checks the host and prints a fix for anything missing.
62
64
 
63
- `install.sh` builds llama.cpp for local inference, installs and configures [OpenClaw](https://github.com/openclaw/openclaw) (the host-side broker every model call goes through), builds the sandbox image, and registers the MCP server if Claude Code is installed. `nomarmy connect` (run by `install.sh`, or by hand for Codex and Cursor) also installs the `/feature` command, Claude Code's status line and, on macOS, nomArmy's notifier. The coordinator gets nomArmy's instructions from the MCP server itself, so there's nothing to copy into your projects.
65
+ `install.sh` builds llama.cpp when you run a local model, installs and configures [OpenClaw](https://github.com/openclaw/openclaw) (the host-side broker every model call goes through), builds the sandbox image, and registers the MCP server if Claude Code is installed. `nomarmy connect` (run by `install.sh`, or by hand for Codex and Cursor) also installs the `/feature` command, Claude Code's status line and, on macOS, nomArmy's notifier. The coordinator gets nomArmy's instructions from the MCP server itself, so there's nothing to copy into your projects.
64
66
 
65
67
  **From npm:** `npm install -g nomarmy@alpha` gives you the `nomarmy` command; `nomarmy setup` then picks a profile and model and prints the `install.sh` command to run (`nomarmy setup --hosted` or `--llama-url <server>` without a local model). Installing from a clone, as in the TL;DR, is the most tested path.
66
68
 
69
+ ### Hosted workers only
70
+
71
+ No GPU and no local model: every job runs on an API key or a subscription (ChatGPT, Muse Code) you add as an agent. Git worktrees, the sandbox and verification still run on your machine, so you still need Git, Node and Podman.
72
+
73
+ ```bash
74
+ git clone https://github.com/rayson-tech/nomarmy.git && cd nomarmy
75
+ npm install && npm link # or: npm install -g nomarmy@alpha
76
+ nomarmy setup --hosted # records that this install has no local model
77
+ ./install.sh --profile hosted # OpenClaw, the sandbox, and the Claude Code registration; no llama.cpp
78
+ nomarmy agents add # an API key or a subscription login
79
+ nomarmy army init --agent <name> # every role on that agent (add --model <model> to pick one)
80
+ nomarmy doctor
81
+ ```
82
+
83
+ A hosted install refuses a job that names no role or agent, rather than falling back to a local model that isn't there. `nomarmy health` warns about any role still on `local`. `e2e.sh` tests the local model, so it has nothing to do here; `nomarmy army assign` tests each role's route instead.
84
+
67
85
  ### macOS (Apple Silicon)
68
86
 
69
87
  ```bash
@@ -122,22 +140,6 @@ chmod +x install.sh e2e.sh scripts/*.sh
122
140
 
123
141
  Moving from a Mac install? Don't copy a Mac binary or model cache over: clone fresh and let `install.sh` build llama.cpp for CUDA on that machine.
124
142
 
125
- ### Hosted workers only
126
-
127
- No GPU and no local model: every job runs on an API key or a subscription (ChatGPT, Muse Code) you add as an agent. Git worktrees, the sandbox and verification still run on your machine, so you still need Git, Node and Podman.
128
-
129
- ```bash
130
- git clone https://github.com/rayson-tech/nomarmy.git && cd nomarmy
131
- npm install && npm link # or: npm install -g nomarmy@alpha
132
- nomarmy setup --hosted # records that this install has no local model
133
- ./install.sh --profile hosted # OpenClaw, the sandbox, and the Claude Code registration; no llama.cpp
134
- nomarmy agents add # an API key or a subscription login
135
- nomarmy army init --agent <name> # every role on that agent (add --model <model> to pick one)
136
- nomarmy doctor
137
- ```
138
-
139
- A hosted install refuses a job that names no role or agent, rather than falling back to a local model that isn't there. `nomarmy health` warns about any role still on `local`. `e2e.sh` tests the local model, so it has nothing to do here; `nomarmy army assign` tests each role's route instead.
140
-
141
143
  ### A shared model server
142
144
 
143
145
  A team GPU box (a DGX, a workstation) runs one llama-server; everyone else points nomArmy at it. That works through an SSH tunnel too (`ssh -L 8080:localhost:8080 gpu-box`, then `http://127.0.0.1:8080`).
@@ -208,6 +210,8 @@ Changes apply to the next job with no restart. The exception is a **new** api ag
208
210
 
209
211
  **Your plan decides which models run.** A model can be listed and still refused: on a ChatGPT plan, the Codex route runs gpt-6-astra and the gpt-5.6 models but refuses gpt-6-sol and gpt-6-luna. `army assign` and `agents update --probe` test the exact route a job takes, so they catch this before a job does.
210
212
 
213
+ **Usage limits.** nomArmy reads how much of a subscription's limit is used where the vendor reports it: Codex in every job's session log, Claude in what Claude Code passes to the status line (Pro and Max plans). It shows in `army`, `local_worker_capacity` and the status line (`⚠ openai 85% wk`), and `nomarmy health` warns from 80%. At the limit, a job on that agent is held rather than left to fail: the General asks you, and resubmits with `confirm_over_limit: true` if you say go. Muse, Grok and API keys don't report their limits; a job that hits one in a `/feature` run pauses that agent for the run.
214
+
211
215
  **Vendor terms and platform risk.** Every model call goes through [OpenClaw](https://github.com/openclaw/openclaw), and subscriptions are reached through each vendor's own CLI or login. We've read the terms that apply (see above), but using a personal subscription through a harness is exactly the kind of use vendors tighten, and a change in a vendor's terms or in OpenClaw can stop a subscription agent from working. Local models and API keys don't carry that risk. Plan on subscriptions as a convenience, not the only way your roles can run.
212
216
 
213
217
  **Picking an agent.** Build work goes to a sandboxed agent: `local`, an api key, Codex or Muse. `local` for a bounded change against a written spec with a test; your code never leaves your machine. An api or subscription agent when the work needs more than the local model, knowing it sends code to that vendor. That's a decision about where your source travels, separate from the trust boundary, which is the same for every agent. The General itself when the answer isn't known yet.
package/lib/admission.mjs CHANGED
@@ -8,6 +8,7 @@ import { readJson } from "./openclaw-run.mjs";
8
8
  import { readOpenClawTranscriptTail } from "./transcript.mjs";
9
9
  import { readClaudeSessionTranscript } from "./claude-transcript.mjs";
10
10
  import { notify } from "./notify.mjs";
11
+ import { readCodexRateLimits, recordUsageSnapshot, readUsageSnapshots, usageStatus } from "./usage-limits.mjs";
11
12
  import { recentModelRefusal } from "./health.mjs";
12
13
  import { writeLease, removeLease, liveLeases, liveSlots, acquireSlot } from "./slots.mjs";
13
14
  import { loadRun, runTotals, runAdmissionProblems, recordRunJob, detectUsageLimit } from "./runs.mjs";
@@ -140,6 +141,11 @@ export function createJobRuntime(deps) {
140
141
  * watching hears about it from any coordinator without polling.
141
142
  */
142
143
  function notifyJobFinished(entry, result, error) {
144
+ try {
145
+ const provider = entry.agent ? agentProviderId(agentsConfig().agents[entry.agent]) : null;
146
+ const snapshot = provider ? readCodexRateLimits(path.join(jobsRoot, entry.jobId)) : null;
147
+ if (snapshot) recordUsageSnapshot(stateRoot, provider, snapshot);
148
+ } catch { /* Usage telemetry must never affect the job result. */ }
143
149
  if (!entry.lane) return; // only tracked jobs, never internal helpers
144
150
  const m = result?.manifest ?? {};
145
151
  const outcome = error ? "failed" : String(m.outcome ?? (result?.ok ? "done" : "finished")).toLowerCase().replace(/_/g, " ");
@@ -161,6 +167,7 @@ export function createJobRuntime(deps) {
161
167
  admission,
162
168
  memory: deps.budgetState.hardwareSnapshot?.memory ?? null,
163
169
  running: [...activeJobs.values()].filter(j => !j.settled).map(j => ({ jobId: j.jobId, workerId: j.workerId, mode: j.mode, lane: j.lane, startedAt: j.startedAt, phase: readJson(path.join(jobsRoot, j.jobId, "status.json"))?.phase ?? "starting" })),
170
+ usageLimits: Object.fromEntries(Object.entries(readUsageSnapshots(stateRoot)).map(([provider, snapshot]) => [provider, usageStatus(snapshot)])),
164
171
  maxWorkers: currentMaxWorkers(),
165
172
  remote: { running: runningCount("remote"), maxWorkers: currentMaxPoolWorkers(), note: "api and subscription agents; each agent's own max_concurrent also applies" }
166
173
  };
@@ -169,6 +176,16 @@ export function createJobRuntime(deps) {
169
176
  await deps.budgetState.refresh();
170
177
  if (jobs.some((j) => jobLane(j) === "remote")) await modelCatalogReady();
171
178
  const problems = [];
179
+ const snapshots = readUsageSnapshots(stateRoot);
180
+ jobs.forEach((j, i) => {
181
+ if (!j.agentName || j.confirm_over_limit === true) return;
182
+ let provider;
183
+ try { provider = agentProviderId(agentsConfig().agents[j.agentName]); } catch { return; }
184
+ const snapshot = snapshots[provider];
185
+ if (!snapshot) return;
186
+ const status = usageStatus(snapshot);
187
+ if (status.level === "over") problems.push(`${jobs.length > 1 ? `job ${i + 1}: ` : ""}agent "${j.agentName}" is held at its usage limit: ${status.text} (reading ${status.ageMinutes} minutes old). Ask the operator before resubmitting with confirm_over_limit: true, or send the job to another agent.`);
188
+ });
172
189
  // A pool-routed job is checked against that pool's OWN (model-dependent)
173
190
  // budget, not the local-derived global one -- see budgetsForPool. Which
174
191
  // specific entry pickProvider will land on isn't known yet at admission
package/lib/army.mjs CHANGED
@@ -30,6 +30,7 @@ import path from "node:path";
30
30
  import YAML from "yaml";
31
31
  import { z } from "zod";
32
32
  import { agentRunsToolsOnHost } from "./dispatch-schema.mjs";
33
+ import { usageStatus } from "./usage-limits.mjs";
33
34
 
34
35
  export const GLOBAL_CONFIG_FILENAME = "config.yml";
35
36
  export const PROJECT_CONFIG_FILENAMES = Object.freeze([".nomarmy.yml", ".nomarmy.yaml"]);
@@ -361,11 +362,20 @@ export function generalOverlap(army, agents = {}) {
361
362
  * description, phase, suggested mode, agent, any problem or overlap with
362
363
  * the General, and which layer set each field.
363
364
  */
364
- export function describeArmy(loaded, { agents = {}, describeAgent = null } = {}) {
365
+ export function describeArmy(loaded, { agents = {}, describeAgent = null, usageSnapshots = null, agentProviderId = null, now = Date.now() } = {}) {
365
366
  const army = loaded.army;
366
367
  const problems = armyTargetProblems(army, agents);
367
368
  const overlap = generalOverlap(army, agents);
368
369
  const runsOn = (name) => (name && agents[name] && describeAgent ? describeAgent(agents[name]) : null);
370
+ const usage = (name) => {
371
+ if (!name || !agents[name] || !usageSnapshots || !agentProviderId) return null;
372
+ let provider;
373
+ try { provider = agentProviderId(agents[name]); } catch { return null; }
374
+ const snapshot = usageSnapshots[provider];
375
+ if (!snapshot) return null;
376
+ const { level, text } = usageStatus(snapshot, now);
377
+ return { level, text };
378
+ };
369
379
  // A subscription whose owner isn't the General's own is worth a word,
370
380
  // not a refusal: it may be the same person's other account (a real Senti
371
381
  // General skipped a role's agent for exactly this, unsure whose it was).
@@ -378,7 +388,7 @@ export function describeArmy(loaded, { agents = {}, describeAgent = null } = {})
378
388
  const roles = Object.fromEntries(Object.entries(army.roles).map(([name, role]) => [name, {
379
389
  description: role.description ?? null, phase: role.phase ?? null, mode: role.mode ?? null,
380
390
  agent: role.agent ?? null, model: role.model ?? agents[role.agent]?.model ?? null, modelIsAuto: role.model === "auto",
381
- agentRunsOn: runsOn(role.agent),
391
+ agentRunsOn: runsOn(role.agent), usage: usage(role.agent),
382
392
  // Its agent's own tools run on this machine (dispatch-schema.mjs): an
383
393
  // implement role there is refused per job unless the agent allows it.
384
394
  hostTools: agentRunsToolsOnHost(agents[role.agent]) ? { allowed: agents[role.agent].allow_host_tools === true, implementRole: (role.mode ?? "implement") === "implement" } : null,
@@ -388,7 +398,7 @@ export function describeArmy(loaded, { agents = {}, describeAgent = null } = {})
388
398
  ? "not defined -- run `nomarmy army general <agent>` so nomArmy can flag roles that share the General's model or usage"
389
399
  : !Object.prototype.hasOwnProperty.call(agents, army.general) ? `agent "${army.general}" is not defined in your agents.yml` : null;
390
400
  return {
391
- general: { ...GENERAL, agent: army.general, agentRunsOn: runsOn(army.general), problem: generalProblem, setBy: loaded.sources.general },
401
+ general: { ...GENERAL, agent: army.general, agentRunsOn: runsOn(army.general), usage: usage(army.general), problem: generalProblem, setBy: loaded.sources.general },
392
402
  workflow: army.workflow,
393
403
  runLimits: army.runLimits ?? {},
394
404
  roles,
@@ -11,6 +11,7 @@ export const COORDINATOR_INSTRUCTIONS = `nomArmy runs bounded engineering jobs (
11
11
  Before dispatching:
12
12
  - Answer where-is / who-calls / grep questions with repo_evidence (deterministic, [path:line] on every hit). Use mode: scout only for read-only research that would otherwise pull many files into your own context.
13
13
  - Call army to see this repo's roles and which agent each runs on; dispatch by army_role when a role fits. A job on a subscription agent needs on_behalf_of set to that agent's owner.
14
+ - Agents' usage limits show in army and local_worker_capacity; a job on an agent at its limit is held. Ask the operator before resubmitting with confirm_over_limit: true, or move the job to another agent. Never set it on your own.
14
15
  - A Claude subscription agent (claude-cli) runs its tools on this machine, outside the sandbox: use it for scout and review work. nomArmy refuses implement jobs on it unless the operator set allow_host_tools; send build work to a sandboxed agent.
15
16
  - Brief outcomes, not edits: a task, explicit acceptance criteria, and the tests that prove it. Put facts you've already resolved in evidence.
16
17
  - Prefer local_worker_start + local_worker_status for anything longer than a few minutes. For a whole feature, use /feature (run_start keeps a run's jobs, spend and hours bounded).
package/lib/health.mjs CHANGED
@@ -22,6 +22,7 @@ import path from "node:path";
22
22
  import { executionMode } from "./execution.mjs";
23
23
  import { modelRejection } from "./openclaw-errors.mjs";
24
24
  import { providerConfigured, readOpenclawConfig } from "./openclaw-config.mjs";
25
+ import { readUsageSnapshots, usageStatus } from "./usage-limits.mjs";
25
26
 
26
27
  const DAY = 86400000;
27
28
 
@@ -183,12 +184,23 @@ export function leftoverIssues({ retainedWorktrees = 0, jobsBytes = 0, staleRunn
183
184
  * Run every check. `env` supplies what each needs, with real defaults;
184
185
  * tests pass their own.
185
186
  */
186
- export async function runHealthChecks({ now = Date.now(), mode = "local", openclawCmd = process.env.NOMARMY_OPENCLAW_CMD || "openclaw", run = runBounded, armySummary = null, agentsError = null, jobsRoot = null, pidAlive = () => true, agents = null, openclawConfig = null, vendors = {}, modelsInUse = null, autoPruned = null } = {}) {
187
+ export async function runHealthChecks({ now = Date.now(), mode = "local", openclawCmd = process.env.NOMARMY_OPENCLAW_CMD || "openclaw", run = runBounded, armySummary = null, agentsError = null, jobsRoot = null, pidAlive = () => true, agents = null, openclawConfig = null, vendors = {}, modelsInUse = null, autoPruned = null, usageSnapshots = null } = {}) {
187
188
  const issues = [];
188
189
  if (autoPruned?.freedBytes) issues.push({ id: `auto-prune:${new Date(now).toISOString()}`, severity: "info",
189
190
  title: `Freed ${(autoPruned.freedBytes / 1024 ** 3).toFixed(2)} GB: ${[autoPruned.pruned ? `runtime data of ${autoPruned.pruned} finished job${autoPruned.pruned === 1 ? "" : "s"} older than ${autoPruned.olderThanHours}h` : null, autoPruned.scratchCleared ? `OpenClaw scratch files of ${autoPruned.scratchCleared} more` : null].filter(Boolean).join(", ")}`,
190
191
  detail: "Automatic; each job's record and report are kept. NOMARMY_AUTO_PRUNE_HOURS sets the age (0 turns it off).", fix: null, short: null });
191
192
  if (agents) issues.push(...providerConfigIssues({ agents, openclawConfig, vendors }));
193
+ if (usageSnapshots) {
194
+ for (const [provider, snapshot] of Object.entries(usageSnapshots)) {
195
+ const status = usageStatus(snapshot, now);
196
+ if (status.level === "ok") continue;
197
+ const short = `${provider} ${status.short}`;
198
+ issues.push({ id: `usage:${provider}:${status.level}`, severity: "warn", title: status.level === "over" ? `${provider} is at its usage limit` : `${provider} usage is ${status.short}`,
199
+ detail: `${status.text} (reading ${status.ageMinutes} minutes old).`,
200
+ fix: status.level === "over" ? "wait for the reset, or move its roles with nomarmy army assign" : "plan remaining work or move its roles with nomarmy army assign",
201
+ short });
202
+ }
203
+ }
192
204
  const [auth, version, latest, plugins] = await Promise.all([
193
205
  run(openclawCmd, ["models", "auth", "list", "--json"]),
194
206
  run(openclawCmd, ["--version"]),
@@ -278,7 +290,7 @@ export async function checkAndRecordHealth({ projectDir, stateRoot, configDir, n
278
290
  if (ageMs !== null) { try { autoPruned = { ...pruneJobRuntime({ stateRoot, olderThanMs: ageMs, now }), olderThanHours: ageMs / 3600000 }; } catch { /* best-effort */ } }
279
291
  const mode = executionMode(env).mode;
280
292
  const result = await runHealthChecks({ now, mode, armySummary, agentsError, jobsRoot: path.join(stateRoot, "jobs"), pidAlive,
281
- agents, openclawConfig: readOpenclawConfig(), vendors: SUBSCRIPTION_VENDORS, modelsInUse, autoPruned });
293
+ agents, openclawConfig: readOpenclawConfig(), vendors: SUBSCRIPTION_VENDORS, modelsInUse, autoPruned, usageSnapshots: readUsageSnapshots(stateRoot) });
282
294
  const toNotify = recordHealth(path.join(stateRoot, "health.json"), result, { now });
283
295
  return { result, toNotify };
284
296
  }
@@ -27,6 +27,7 @@ import fs from "node:fs";
27
27
  import os from "node:os";
28
28
  import path from "node:path";
29
29
  import { fileURLToPath } from "node:url";
30
+ import { mergeUsageSnapshot, normalizeClaudeRateLimits, readUsageSnapshots, recordUsageSnapshot, usageStatus } from "./usage-limits.mjs";
30
31
 
31
32
  function readJson(file) { try { return JSON.parse(fs.readFileSync(file, "utf8")); } catch { return null; } }
32
33
  function pidAlive(pid) {
@@ -47,6 +48,17 @@ const minutes = (ms) => (ms < 60000 ? `${Math.max(1, Math.round(ms / 1000))}s` :
47
48
  */
48
49
  export function statusLineText({ session = {}, stateRoot, now = Date.now(), maxLength = Number(process.env.NOMARMY_STATUSLINE_MAX) || 90 } = {}) {
49
50
  const root = stateRoot ?? (process.env.NOMARMY_AGENT_STATE || path.join(os.homedir(), ".local", "share", "nomarmy-local-agents"));
51
+ // Every Claude Code window redraws this line with its own last-seen
52
+ // limits, idle ones included, so a reading is merged rather than trusted
53
+ // (mergeUsageSnapshot), and written only when it moves the figure.
54
+ try {
55
+ const snapshot = normalizeClaudeRateLimits(session.rate_limits);
56
+ if (snapshot) {
57
+ const previous = readUsageSnapshots(root)["claude-cli"];
58
+ const merged = mergeUsageSnapshot(previous, { ...snapshot, observedAt: now }, now);
59
+ if (merged !== previous) recordUsageSnapshot(root, "claude-cli", merged);
60
+ }
61
+ } catch { /* usage capture must never break the status line */ }
50
62
  // workspace.project_dir is where Claude Code was launched -- the repo that
51
63
  // session's nomArmy server works in; current_dir can be a subdirectory.
52
64
  const cwd = session.workspace?.project_dir ?? session.workspace?.current_dir ?? session.cwd ?? process.cwd();
@@ -99,23 +111,37 @@ export function statusLineText({ session = {}, stateRoot, now = Date.now(), maxL
99
111
  try {
100
112
  const health = readJson(path.join(root, "health.json"));
101
113
  if (health && now - Date.parse(health.checkedAt) < 2 * 86400000) {
102
- const top = (health.issues ?? []).find((i) => i.severity !== "info" && i.short);
114
+ // Usage warnings get their own part below, so they aren't repeated here.
115
+ const top = (health.issues ?? []).find((i) => i.severity !== "info" && i.short && !String(i.id).startsWith("usage:"));
103
116
  if (top) healthPart = ` │ ⚠ ${top.short}`;
104
117
  }
105
118
  } catch { /* no health yet */ }
106
119
  runPart += healthPart;
107
- // As many jobs as fit, then "+N": the run summary is never what gets cut.
120
+ let usageParts = [];
121
+ try {
122
+ const warnings = Object.entries(readUsageSnapshots(root)).map(([provider, snapshot]) => {
123
+ const status = usageStatus(snapshot, now);
124
+ if (status.level === "ok") return null;
125
+ return { level: status.level, text: ` │ ${status.level === "over" ? "⛔" : "⚠"} ${provider} ${status.short}` };
126
+ }).filter(Boolean).sort((a, b) => (a.level === "over" ? 0 : 1) - (b.level === "over" ? 0 : 1));
127
+ usageParts = warnings.map((warning) => warning.text);
128
+ } catch { /* usage display is best-effort */ }
129
+ // As many jobs as fit, then "+N". Only after every job is collapsed may
130
+ // lower-priority usage warnings be omitted to honor the hard cap.
108
131
  const prefix = `${head ? `${head} │ ` : ""}🍪 `;
109
132
  const count = jobs.length > 1 ? `${jobs.length}: ` : "";
110
- let shown = jobs.length, army;
133
+ let shown = jobs.length, army, suffix;
111
134
  for (;;) {
112
135
  const rest = jobs.length - shown;
113
136
  army = (!jobs.length ? "idle" : `${count}${jobs.slice(0, shown).join(" · ")}${rest ? `${shown ? " " : ""}+${rest}` : ""}`)
114
137
  + (elsewhere ? ` · ${elsewhere} in other repo${elsewhere === 1 ? "" : "s"}` : "");
115
- if (shown === 0 || [...`${prefix}${army}${runPart}`].length <= maxLength) break;
116
- shown--;
138
+ suffix = runPart + usageParts.join("");
139
+ if ([...`${prefix}${army}${suffix}`].length <= maxLength) break;
140
+ if (shown > 0) { shown--; continue; }
141
+ if (usageParts.length) { usageParts.pop(); continue; }
142
+ break;
117
143
  }
118
- return `${prefix}${army}${runPart}`;
144
+ return `${prefix}${army}${suffix}`;
119
145
  }
120
146
 
121
147
  const isMain = (() => { try { return path.resolve(process.argv[1] ?? "") === fileURLToPath(import.meta.url); } catch { return false; } })();
@@ -0,0 +1,126 @@
1
+ import fs from "node:fs";
2
+ import path from "node:path";
3
+ import { randomUUID } from "node:crypto";
4
+
5
+ // How much of each subscription's usage limit is used, as its vendor reports
6
+ // it: Codex writes rate_limits into every job's session log, and Claude Code
7
+ // passes them to the status line. Kept per OpenClaw provider id (one host
8
+ // login per provider) in <stateRoot>/usage-limits.json. Admission holds a job
9
+ // on an agent that's over its limit until the General confirms.
10
+
11
+ const object = value => value !== null && typeof value === "object" && !Array.isArray(value);
12
+ const finite = value => typeof value === "number" && Number.isFinite(value);
13
+ const windowName = minutes => ({ 300: "5h", 10080: "week", 1440: "day" })[minutes] ?? `${minutes}m`;
14
+ function window(value, name, minutes, percentKey) {
15
+ if (!object(value) || !finite(value[percentKey]) || value[percentKey] < 0) return null;
16
+ return { name, usedPercent: value[percentKey], windowMinutes: minutes,
17
+ resetsAt: finite(value.resets_at) ? value.resets_at * 1000 : null };
18
+ }
19
+
20
+ /** Read the last rate-limit event from the newest rollout (by modification time). */
21
+ export function readCodexRateLimits(jobDir) {
22
+ try {
23
+ const files = [];
24
+ function walk(dir) {
25
+ for (const entry of fs.readdirSync(dir, { withFileTypes: true })) {
26
+ const file = path.join(dir, entry.name);
27
+ if (entry.isDirectory()) walk(file);
28
+ else if (entry.isFile() && /^rollout-.*\.jsonl$/.test(entry.name)) files.push({ file, mtime: fs.statSync(file).mtimeMs });
29
+ }
30
+ }
31
+ const agents = path.join(jobDir, "runtime/state/agents");
32
+ for (const agent of fs.readdirSync(agents, { withFileTypes: true })) {
33
+ if (!agent.isDirectory()) continue;
34
+ const sessions = path.join(agents, agent.name, "agent/codex-home/sessions");
35
+ if (fs.existsSync(sessions)) walk(sessions);
36
+ }
37
+ files.sort((a, b) => b.mtime - a.mtime || b.file.localeCompare(a.file));
38
+ if (!files.length) return null;
39
+ let last = null;
40
+ for (const line of fs.readFileSync(files[0].file, "utf8").split(/\r?\n/)) {
41
+ try {
42
+ const event = JSON.parse(line);
43
+ if (event.type === "event_msg" && event.payload?.type === "token_count" && object(event.payload.rate_limits)) last = event;
44
+ } catch { /* A partial trailing write is not an event. */ }
45
+ }
46
+ if (!last) return null;
47
+ const limits = last.payload.rate_limits;
48
+ const windows = [limits.primary, limits.secondary].map(value => {
49
+ const minutes = finite(value?.window_minutes) ? value.window_minutes : null;
50
+ return window(value, minutes === null ? "unknown" : windowName(minutes), minutes, "used_percent");
51
+ }).filter(Boolean);
52
+ const limitReached = limits.rate_limit_reached_type != null || Boolean(limits.spend_control_reached);
53
+ if (!windows.length && !limitReached) return null;
54
+ const timestamp = Date.parse(last.timestamp);
55
+ return { source: "codex", plan: typeof limits.plan_type === "string" ? limits.plan_type : null,
56
+ limitReached, observedAt: Number.isFinite(timestamp) ? timestamp : files[0].mtime, windows };
57
+ } catch { return null; }
58
+ }
59
+
60
+ export function normalizeClaudeRateLimits(rateLimits) {
61
+ if (!object(rateLimits)) return null;
62
+ const windows = [["five_hour", "5h", 300], ["seven_day", "week", 10080], ["spend_limit", "spend", null]]
63
+ .map(([key, name, minutes]) => window(rateLimits[key], name, minutes, "used_percentage")).filter(Boolean);
64
+ // Claude supplies percentages, not an independent provider-wide reached flag.
65
+ return windows.length ? { source: "claude", plan: null, limitReached: false, observedAt: Date.now(), windows } : null;
66
+ }
67
+
68
+ /**
69
+ * Combine a stored snapshot with a new reading that may be stale. Every
70
+ * Claude Code window passes the status line the limits from its own last
71
+ * response, so an idle window keeps reporting an old, lower figure. Usage
72
+ * only rises within a window, so per window: a later reset wins, and at the
73
+ * same reset the higher figure wins. A stored window the reading lacks is
74
+ * kept while it's live. Returns the stored snapshot itself when nothing
75
+ * changed, so callers can skip the write.
76
+ */
77
+ export function mergeUsageSnapshot(previous, incoming, now = Date.now()) {
78
+ if (!previous) return incoming;
79
+ const key = (w) => [w.resetsAt ?? 0, w.usedPercent];
80
+ const newer = (a, b) => { const [ra, pa] = key(a), [rb, pb] = key(b); return ra > rb || (ra === rb && pa > pb); };
81
+ const byName = new Map(previous.windows.filter((w) => w.resetsAt === null || w.resetsAt > now).map((w) => [w.name, w]));
82
+ let changed = byName.size !== previous.windows.length;
83
+ for (const w of incoming.windows) {
84
+ const stored = byName.get(w.name);
85
+ if (!stored || newer(w, stored)) { byName.set(w.name, w); changed = true; }
86
+ }
87
+ const limitReached = incoming.limitReached || previous.limitReached;
88
+ if (!changed && limitReached === previous.limitReached) return previous;
89
+ return { ...incoming, limitReached, observedAt: now, windows: [...byName.values()] };
90
+ }
91
+
92
+ export function readUsageSnapshots(stateRoot) {
93
+ try {
94
+ const data = JSON.parse(fs.readFileSync(path.join(stateRoot, "usage-limits.json"), "utf8"));
95
+ if (!object(data)) return {};
96
+ return Object.fromEntries(Object.entries(data).filter(([, s]) => object(s) && ["codex", "claude"].includes(s.source)
97
+ && finite(s.observedAt) && typeof s.limitReached === "boolean" && (s.plan === null || typeof s.plan === "string")
98
+ && Array.isArray(s.windows) && s.windows.every(w => object(w) && typeof w.name === "string" && finite(w.usedPercent)
99
+ && (w.windowMinutes === null || finite(w.windowMinutes)) && (w.resetsAt === null || finite(w.resetsAt)))));
100
+ } catch { return {}; }
101
+ }
102
+
103
+ /** Save a provider's snapshot, unless the one on file is newer (jobs finish out of order). */
104
+ export function recordUsageSnapshot(stateRoot, provider, snapshot) {
105
+ const current = readUsageSnapshots(stateRoot);
106
+ if (current[provider]?.observedAt > snapshot.observedAt) return;
107
+ const snapshots = { ...current, [provider]: snapshot };
108
+ fs.mkdirSync(stateRoot, { recursive: true });
109
+ const file = path.join(stateRoot, "usage-limits.json"), tmp = `${file}.${randomUUID()}.tmp`;
110
+ try {
111
+ fs.writeFileSync(tmp, JSON.stringify(snapshots, null, 2));
112
+ fs.renameSync(tmp, file);
113
+ } finally { fs.rmSync(tmp, { force: true }); }
114
+ }
115
+
116
+ export function usageStatus(snapshot, now = Date.now()) {
117
+ const live = snapshot.windows.filter(w => w.resetsAt === null || w.resetsAt > now).sort((a, b) => b.usedPercent - a.usedPercent);
118
+ const reached = snapshot.limitReached && (snapshot.windows.length === 0 || live.length > 0);
119
+ const highest = live[0];
120
+ const level = reached || highest?.usedPercent >= 100 ? "over" : highest?.usedPercent >= 80 ? "high" : "ok";
121
+ const text = live.length ? live.map(w => `${w.usedPercent}% of ${w.name}, resets ${w.resetsAt === null ? "unknown" : new Date(w.resetsAt).toLocaleString("en-US", { weekday: "short", hour: "2-digit", minute: "2-digit", hour12: false })}`).join("; ") : reached ? "limit reached, reset unknown" : "no live usage windows";
122
+ // A short label for tight spaces (status line, health): "85% wk".
123
+ const short = highest ? `${highest.usedPercent}% ${highest.name === "week" ? "wk" : highest.name}` : reached ? "limit" : null;
124
+ return { level, text, short, resetsAt: level === "over" ? highest?.resetsAt ?? null : null,
125
+ ageMinutes: Math.max(0, Math.floor((now - snapshot.observedAt) / 60000)) };
126
+ }
package/mcp/server.mjs CHANGED
@@ -32,6 +32,7 @@ import { liveLeases } from "../lib/slots.mjs";
32
32
  import { createRun, loadRun, runTotals, finishRun, resolveRunLimits, describeLoweredLimits } from "../lib/runs.mjs";
33
33
  import { agentDispatchFields, resolveAgentModel, agentProviderId, describeAgent } from "../lib/agents.mjs";
34
34
  import { OUTCOMES, COORDINATOR_STATUS_BY_OUTCOME } from "../lib/outcomes.mjs";
35
+ import { readUsageSnapshots, usageStatus } from "../lib/usage-limits.mjs";
35
36
  import { createBuildMetrics, resolveOutcome, finalText, workerMetadata, usageMetrics, policyAdmissionProblems, applyRefactorContract, applyVerificationPolicy, resolveVerifyRegression } from "../lib/outcome.mjs";
36
37
  import { compactJobRecord, formatResult, formatUnion, testChangeBanner, regressionCheckBanner, decomposeOverlapBanner } from "../lib/job-format.mjs";
37
38
 
@@ -282,6 +283,7 @@ export const jobSchema = z.object({
282
283
  report: z.enum(["brief", "standard", "full"]).optional().describe("How much the worker may report back, capped by its agent's tier: brief (today's local-sized report), standard (the default), full (the frontier ceiling: about 2k tokens for implement, 4k for a scout). The report lands in your own context and is re-read every later turn, so ask for full only when the job's findings are the point (a broad review). No effect on the local model, whose caps are calibrated."),
283
284
  commit_subject: z.string().max(200).optional().describe("implement: the subject line of the commit nomArmy makes on the worker branch, e.g. \"Keep held-back tables in the list_tables cache\". Defaults to the task's first sentence; the body is the worker's NOTE, and the job id is a trailer."),
284
285
  army_role: z.string().regex(/^[a-z][a-z0-9-]{0,63}$/).optional().describe("Dispatch by army role (e.g. \"sr-dev\", \"security-analyst\"): nomArmy runs it on the agent this repo assigns to that role and puts the role's description at the top of the brief. Call the `army` tool first to see this repo's roles. Mutually exclusive with agent. Add on_behalf_of in case the role's agent is a subscription; it's ignored otherwise."),
286
+ confirm_over_limit: z.boolean().optional().describe("Override a reached usage limit: the General must ask the operator before resubmitting with confirm_over_limit: true, or send the job to another agent. nomArmy never sets it itself."),
285
287
  on_behalf_of: z.string().min(1).max(254).optional().describe("Required when the job's agent is a subscription: must exactly match that agent's owner in agents.yml, or nomArmy refuses the job. A self-reported attestation, not an independently verified identity check -- nomArmy has no caller-identity boundary today, so what this guarantees is explicit, auditable intent and hard refusal on mismatch or omission, not cryptographic proof of who issued the call. Ignored for a local or api agent."),
286
288
  evidence: z.string().max(maxEvidenceChars,
287
289
  `Evidence exceeds the ${maxEvidenceChars}-character budget. This is for facts already resolved (e.g. with repo_evidence), not more description of the task -- if it needs more than this, resolve less per job or put the pointer (a path and line range) here instead of the material itself.`
@@ -474,14 +476,17 @@ server.tool("run_finish", "Close a /feature run as complete or stopped, with a o
474
476
  server.tool("army", "Who you, the General, are and who you call for what in this repository: your fixed charter and the agent you're defined as, the army's workflow, then each role's description, phase (build, review, acceptance), suggested mode, and the agent it runs on, with which config layer set each value (global, project .nomarmy.yml, local .nomarmy.local.yml). Flags roles with no usable agent, and roles that share your model or subscription (not an independent review). Dispatch a role with `army_role`, or an agent directly with `agent`. Read-only, re-read on every call.", {}, async () => {
475
477
  try {
476
478
  const agents = agentsConfig().agents;
477
- const summary = describeArmy(currentArmy(), { agents, describeAgent });
479
+ const usageSnapshots = readUsageSnapshots(stateRoot);
480
+ const summary = describeArmy(currentArmy(), { agents, describeAgent, usageSnapshots, agentProviderId });
478
481
  // Each agent's models, from OpenClaw's catalog, so the General can pick
479
482
  // one for a role set to "auto". The catalog can lag a brand-new model.
480
483
  const catalog = await modelCatalogReady();
481
484
  summary.agents = Object.fromEntries(Object.entries(agents).map(([name, agent]) => {
482
485
  const provider = agentProviderId(agent);
483
486
  const models = provider && catalog ? [...catalog.keys()].filter((k) => k.startsWith(`${provider}/`)).map((k) => k.slice(provider.length + 1)) : [];
484
- return [name, { runsOn: describeAgent(agent), defaultModel: agent.model ?? null, models }];
487
+ const snapshot = usageSnapshots[provider];
488
+ const usage = snapshot ? (() => { const { level, text } = usageStatus(snapshot); return { level, text }; })() : null;
489
+ return [name, { runsOn: describeAgent(agent), defaultModel: agent.model ?? null, models, usage }];
485
490
  }));
486
491
  // A pinned model missing from the catalog isn't necessarily wrong:
487
492
  // `army assign` proves an unlisted model with a real test call, and the
package/package.json CHANGED
@@ -3,7 +3,7 @@
3
3
  "description": "A harness for AI coding workers whose claims are never trusted: your coding assistant stays in charge while workers implement and test in sandboxes, on local models, API keys or your own subscriptions.",
4
4
  "author": "Rayson Technologies",
5
5
  "license": "Apache-2.0",
6
- "version": "0.1.0-alpha.2",
6
+ "version": "0.1.0-alpha.3",
7
7
  "private": false,
8
8
  "type": "module",
9
9
  "engines": {