nomarmy 0.1.0-alpha.0 → 0.1.0-alpha.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -8,9 +8,28 @@
8
8
 
9
9
  <p align="center"><em>Tiny coders, big appetites for bounded tickets.</em> 🍪</p>
10
10
 
11
- **An agent harness where a worker's claims are never trusted, and the environment your tests need is declared, disposable and reproducible.**
11
+ **Your coding assistant plans; sandboxed workers build; nothing counts until nomArmy has checked it.**
12
12
 
13
- Your coding assistant (Claude Code, Codex or Cursor) stays in charge as the **General**: it decides what gets built and whether the result is acceptable. The work goes to **noms**, workers that implement, test and repair in their own git worktree and sandbox. A nom can run on a local model at no token cost, on an API key, or on your own Claude, ChatGPT or Muse Code subscription. nomArmy owns everything in between: worktrees, git, sandboxes, verification, and the evidence that decides whether work is accepted.
13
+ ## TL;DR
14
+
15
+ 1. **Have** Git, Node 20+ and [Podman](https://podman.io) (on macOS: `brew install podman && podman machine init && podman machine start`).
16
+ 2. **Install** (builds the local model server, OpenClaw and the sandbox, and registers nomArmy with Claude Code):
17
+ ```bash
18
+ git clone https://github.com/rayson-tech/nomarmy.git && cd nomarmy
19
+ ./install.sh --profile macbook-pro # or nvidia-linux, cpu-linux, dgx-spark, bedrock
20
+ ./e2e.sh --profile macbook-pro # should end with: === E2E PASS ===
21
+ ```
22
+ 3. **Set up your repo:** in the project, run `nomarmy init`. It proposes a `.nomarmy.yml` with your test command.
23
+ 4. **Use it:** restart Claude Code in that project and ask it to use nomArmy for one small bug that has a test. When that works, try `/feature <what you want built>`.
24
+ 5. **Optional:** add hosted workers with `nomarmy agents add`, give roles to them with `nomarmy army init`, or connect Codex or Cursor with `nomarmy connect codex cursor`.
25
+
26
+ Stuck? `nomarmy doctor` checks the machine, and `nomarmy health` checks everything nomArmy runs on.
27
+
28
+ ## What it is
29
+
30
+ Your coding assistant (Claude Code, Codex or Cursor) stays in charge as the **General**: it decides what gets built and whether the result is acceptable. The work goes to **noms**, workers that implement, test and repair in their own git worktree and sandbox, on a local model, an API key, or your own ChatGPT or Muse Code subscription. nomArmy owns everything in between: worktrees, git, sandboxes, verification, and the evidence that decides whether work is accepted.
31
+
32
+ **What you get is work you don't have to take on faith**, not cheaper work. Delegating costs the General tokens too: briefing and reviewing. On small, already-diagnosed tickets we measured 4 to 8 times more of the General's tokens than fixing the bug directly, and break-even at roughly 150 lines of context a fix needs to read ([the measurements](docs/experiments/2026-09-20-model-bakeoff-and-economics.md)). It pays off on bigger tickets, on parallel work, and anywhere you'd otherwise have to trust an agent's say-so.
14
33
 
15
34
  Developed and maintained by Rayson Technologies. This is an alpha (`0.1.0-alpha`).
16
35
 
@@ -22,35 +41,10 @@ Developed and maintained by Rayson Technologies. This is an alpha (`0.1.0-alpha`
22
41
  4. nomArmy treats that report as a claim. It reads the real diff from git, runs your verification profile itself in a fresh sandbox, reverts the production change to check the tests actually fail without it, and scans for secrets.
23
42
  5. Only then does it commit, on the worker's own branch. It never merges into yours: reviewing and integrating stay with the General, and with you.
24
43
 
25
- A malformed report isn't automatically a failure: if the repository changed, nomArmy verifies independently and may recover the work. Failing verification stays failed, unconditionally.
44
+ A malformed report isn't automatically a failure: if the repository changed, nomArmy verifies independently and may recover the work. Failing verification stays failed, unconditionally. And the checks aren't the General's to waive: a repo's `.nomarmy.yml` policy (on by default for new repos) makes verification and the revert check mandatory for every job.
26
45
 
27
46
  Around that core: **agents** say where a job can run, the **army** says which role runs on which agent, and **`/feature`** runs a whole feature end to end, from plan through build, review and acceptance, handing you a branch to merge.
28
47
 
29
- ## Quick start
30
-
31
- You need Git, Node 20+, and [Podman](https://podman.io).
32
-
33
- ```bash
34
- npm install -g nomarmy@alpha
35
- nomarmy doctor # what's missing on this machine, and how to fix it
36
- nomarmy setup # pick a profile and model; prints the install.sh command to run
37
- nomarmy connect claude # or codex, cursor: registers nomArmy with your coordinator
38
- ```
39
-
40
- Installing from a clone is the most tested path, and it's what `install.sh` expects: see [Install](#install). `install.sh` builds llama.cpp for local inference, installs and configures [OpenClaw](https://github.com/openclaw/openclaw) (the host-side broker every model call goes through), builds the sandbox image and registers the MCP server.
41
-
42
- `nomarmy connect` also installs the `/feature` command, Claude Code's status line and (on macOS) nomArmy's notifier. The coordinator gets nomArmy's instructions from the MCP server itself, so there's nothing to copy into your projects.
43
-
44
- Then, in any git repository:
45
-
46
- ```bash
47
- nomarmy init # propose a .nomarmy.yml with your test command
48
- nomarmy agents add # optional: an API key or a subscription
49
- nomarmy army init # optional: the default roster of roles
50
- ```
51
-
52
- Ask your coordinator to delegate one small, well-tested ticket before anything bigger. When that works, try `/feature <what you want built>`.
53
-
54
48
  ## Install
55
49
 
56
50
  | Platform | Guide |
@@ -63,6 +57,10 @@ Ask your coordinator to delegate one small, well-tested ticket before anything b
63
57
 
64
58
  Every platform needs Git and Podman. `nomarmy doctor` checks the host and prints a fix for anything missing.
65
59
 
60
+ `install.sh` builds llama.cpp for local inference, installs and configures [OpenClaw](https://github.com/openclaw/openclaw) (the host-side broker every model call goes through), builds the sandbox image, and registers the MCP server if Claude Code is installed. `nomarmy connect` (run by `install.sh`, or by hand for Codex and Cursor) also installs the `/feature` command, Claude Code's status line and, on macOS, nomArmy's notifier. The coordinator gets nomArmy's instructions from the MCP server itself, so there's nothing to copy into your projects.
61
+
62
+ **From npm:** `npm install -g nomarmy@alpha` gives you the `nomarmy` command; `nomarmy setup` then picks a profile and model and prints the `install.sh` command to run. Installing from a clone, as in the TL;DR, is the most tested path.
63
+
66
64
  ### macOS (Apple Silicon)
67
65
 
68
66
  ```bash
@@ -179,6 +177,8 @@ Changes apply to the next job with no restart. The exception is a **new** api ag
179
177
 
180
178
  **Your plan decides which models run.** A model can be listed and still refused: on a ChatGPT plan, the Codex route runs gpt-6-astra and the gpt-5.6 models but refuses gpt-6-sol and gpt-6-luna. `army assign` and `agents update --probe` test the exact route a job takes, so they catch this before a job does.
181
179
 
180
+ **Vendor terms and platform risk.** Every model call goes through [OpenClaw](https://github.com/openclaw/openclaw), and subscriptions are reached through each vendor's own CLI or login. We've read the terms that apply (see above), but using a personal subscription through a harness is exactly the kind of use vendors tighten, and a change in a vendor's terms or in OpenClaw can stop a subscription agent from working. Local models and API keys don't carry that risk. Plan on subscriptions as a convenience, not the only way your roles can run.
181
+
182
182
  **Picking an agent.** Build work goes to a sandboxed agent: `local`, an api key, Codex or Muse. `local` for a bounded change against a written spec with a test; your code never leaves your machine. An api or subscription agent when the work needs more than the local model, knowing it sends code to that vendor. That's a decision about where your source travels, separate from the trust boundary, which is the same for every agent. The General itself when the answer isn't known yet.
183
183
 
184
184
  ## The army: who does what
@@ -270,6 +270,18 @@ verification:
270
270
 
271
271
  `nomarmy validate` checks the file against the schema; `nomarmy scan --check` diffs it against what the repo actually contains.
272
272
 
273
+ **Policy: what no job can skip.**
274
+
275
+ ```yaml
276
+ policy:
277
+ require_verification: true # every implement job needs a verification profile; only passing work commits
278
+ require_regression_check: true # verify_regression can't be switched off per job
279
+ ```
280
+
281
+ `nomarmy init` proposes both for new repos. Without them, a job with no verification profile still commits (flagged for review, not blocked), and the General decides per job whether to run the revert check. With them, those are the repo's rules, not the General's judgment calls, and since nomArmy reads this file only from your checkout, neither the General nor a worker can relax it.
282
+
283
+ **Refactors.** Reverting a behavior-preserving change restores code that works, so the revert check can't prove anything about it. A job can declare `refactor: true` instead: nomArmy then commits it only if verification passes **and no test file was added, changed or deleted**. The existing tests passing unchanged is the evidence. A change that alters behavior has to alter tests to show it, so it can't pass as a refactor.
284
+
273
285
  **Add a check for what unit tests can't see.** A module left out of a deploy bundle passes every unit test and crashes at deploy. When `nomarmy init` sees a bundle or packaging step (Lambda asset scripts, SAM, Serverless, CDK), it suggests a profile that runs it and then imports each entry point from the built bundle.
274
286
 
275
287
  ### Languages and dependencies
package/e2e.sh CHANGED
@@ -56,18 +56,35 @@ fi
56
56
  # bind-mounting paths under the host home directory, and the cleanup step
57
57
  # below bind-mounts this directory into a container.
58
58
  TMP="$(mktemp -d "$HOME/.nomarmy-e2e.XXXXXX")"
59
+ # OpenClaw's own state for this run, also under $HOME, the way nomArmy's
60
+ # real dispatch keeps it (--state-dir). Left unset, OpenClaw puts its
61
+ # working files in the system temp folder, which the Podman sandbox can't
62
+ # mount on macOS, and it uses the operator's main state, whose memory index
63
+ # holds their past sessions.
64
+ STATE="$(mktemp -d "$HOME/.nomarmy-e2e-state.XXXXXX")"
59
65
 
60
66
  cleanup() {
61
67
  local exit_code=$?
62
68
 
63
- if [[ -n "${TMP:-}" && -d "$TMP" ]]; then
69
+ # OpenClaw leaves this run's sandbox container running; its name carries
70
+ # the hash recorded under the state dir (as nomArmy's job cleanup does).
71
+ if [[ -n "${STATE:-}" && -d "$STATE/state/sandbox/skills-workspaces" ]] && command -v podman >/dev/null 2>&1; then
72
+ for ws in "$STATE"/state/sandbox/skills-workspaces/workspace-*; do
73
+ [[ -d "$ws" ]] || continue
74
+ podman ps -a --filter "name=${ws##*/workspace-}" --format '{{.Names}}' 2>/dev/null \
75
+ | xargs -r podman rm -f -v >/dev/null 2>&1 || true
76
+ done
77
+ fi
78
+
79
+ for dir in "${TMP:-}" "${STATE:-}"; do
80
+ if [[ -n "$dir" && -d "$dir" ]]; then
64
81
 
65
82
  # OpenClaw's Podman sandbox may create files that the host user
66
83
  # cannot delete directly. Use a disposable container to clean
67
84
  # the temporary workspace first.
68
85
  if command -v podman >/dev/null 2>&1 && podman info >/dev/null 2>&1; then
69
86
  podman run --rm \
70
- -v "$TMP:/cleanup" \
87
+ -v "$dir:/cleanup" \
71
88
  alpine:3.20 \
72
89
  sh -c '
73
90
  find /cleanup -mindepth 1 -maxdepth 1 -exec rm -rf -- {} + \
@@ -76,13 +93,14 @@ cleanup() {
76
93
  >/dev/null 2>&1 || true
77
94
  fi
78
95
 
79
- rm -rf "$TMP" 2>/dev/null || true
96
+ rm -rf "$dir" 2>/dev/null || true
80
97
 
81
- if [[ -d "$TMP" ]]; then
98
+ if [[ -d "$dir" ]]; then
82
99
  echo "WARN: E2E temporary directory could not be completely removed:"
83
- echo " $TMP"
100
+ echo " $dir"
84
101
  fi
85
102
  fi
103
+ done
86
104
 
87
105
  exit "$exit_code"
88
106
  }
@@ -126,15 +144,29 @@ PROMPT='Fix the bug so npm test passes. Work only in the workspace. Run npm test
126
144
 
127
145
  OUT="$TMP/openclaw.json"
128
146
 
147
+ # The same privacy settings every nomArmy job gets (lib/openclaw-run.mjs
148
+ # withJobPrivacy): OpenClaw's memory search and session-memory hook off, so
149
+ # nothing is indexed or sent for embedding.
150
+ CONFIG="$STATE/openclaw.job.json"
151
+ mkdir -p "$STATE/state"
152
+ node --input-type=module -e '
153
+ import fs from "node:fs";
154
+ import { readOpenclawConfig } from "'"$ROOT"'/lib/openclaw-config.mjs";
155
+ import { withJobPrivacy } from "'"$ROOT"'/lib/openclaw-run.mjs";
156
+ fs.writeFileSync(process.argv[1], JSON.stringify(withJobPrivacy(readOpenclawConfig() ?? {})), { mode: 0o600 });
157
+ ' "$CONFIG"
158
+
129
159
  openclaw agent exec "$PROMPT" \
130
160
  --model "$NOMARMY_WORKER_PROVIDER/$NOMARMY_WORKER_MODEL" \
131
161
  --cwd "$TMP" \
162
+ --state-dir "$STATE/state" \
163
+ --config "$CONFIG" \
132
164
  --code-mode direct \
133
165
  --local-model-lean \
134
166
  --thinking off \
135
167
  --timeout 600 \
136
168
  --json \
137
- > "$OUT"
169
+ > "$OUT" 2> "$STATE/openclaw.stderr.log"
138
170
 
139
171
  # Independent verification. We do not trust the worker's claim
140
172
  # that its implementation is correct.
@@ -0,0 +1,370 @@
1
+ import fs from "node:fs";
2
+ import path from "node:path";
3
+ import { clampInt } from "./budget-state.mjs";
4
+ import { checkBrief, assessAdmission, describeBudgets } from "./budget.mjs";
5
+ import { parseStatusPorcelainZ, isRuntimeJunk } from "./git-record.mjs";
6
+ import { readJson } from "./openclaw-run.mjs";
7
+ import { readOpenClawTranscriptTail } from "./transcript.mjs";
8
+ import { readClaudeSessionTranscript } from "./claude-transcript.mjs";
9
+ import { notify } from "./notify.mjs";
10
+ import { recentModelRefusal } from "./health.mjs";
11
+ import { writeLease, removeLease, liveLeases, liveSlots, acquireSlot } from "./slots.mjs";
12
+ import { loadRun, runTotals, runAdmissionProblems, recordRunJob, detectUsageLimit } from "./runs.mjs";
13
+ import { agentProviderId, hostToolsImplementProblem } from "./agents.mjs";
14
+ import { policyAdmissionProblems } from "./outcome.mjs";
15
+
16
+ // `lane` is "local" (the local model on llama-server) or "remote" (an api
17
+ // or subscription agent: the inference runs at the vendor). The local-slot
18
+ // admission check must only ever count the local lane. A subscription job
19
+ // used to land in "local" (the lane was decided by `pool` alone), so a
20
+ // Claude or Codex job took llama-server's only slot and blocked local work
21
+ // it never competed with -- reported from a real Senti run.
22
+ export function jobLane(job) {
23
+ return job.pool || job.subscription_worker ? "remote" : "local";
24
+ }
25
+
26
+ // A static, operator-declared ceiling on how many remote jobs (api and
27
+ // subscription agents) may run at once, independent of and additive to
28
+ // currentMaxWorkers()'s local ceiling. Each still runs a sandbox and a
29
+ // worktree on this machine, which is what this bounds; each agent's own
30
+ // max_concurrent bounds its vendor. The env name predates agents.yml
31
+ // (remote jobs were all "pool" jobs then) -- exactly the "more real concurrency, not just diversity"
32
+ // benefit of spreading load across providers with their own separate rate
33
+ // limits. Not rate-limit-aware (see config/providers.yml.example); read
34
+ // fresh each call, matching currentMaxWorkers()'s own env-read pattern.
35
+ export function currentMaxPoolWorkers() {
36
+ return clampInt(process.env.NOMARMY_MAX_POOL_WORKERS, 1, 32, 4);
37
+ }
38
+
39
+ // Pure partition of a batch's ORIGINAL indices by lane -- pulled out of
40
+ // local_workers' handler so this specific invariant (every job lands in
41
+ // exactly one lane, indices preserved) is directly testable without also
42
+ // exercising the full async dispatch/mapLimit machinery around it. This is
43
+ // the exact split that used to not exist at all: every job in a batch
44
+ // shared one `parallel` slot count derived only from the local ceiling,
45
+ // which let an all-pool batch ignore NOMARMY_MAX_POOL_WORKERS entirely.
46
+ export function splitJobsByLane(jobs) {
47
+ const localIndices = [], remoteIndices = [];
48
+ jobs.forEach((j, i) => (jobLane(j) === "remote" ? remoteIndices : localIndices).push(i));
49
+ return { localIndices, remoteIndices };
50
+ }
51
+
52
+ export function toolText(text, isError = false) { return { content: [{ type: "text", text }], isError }; }
53
+
54
+ // The capacity snapshot only when a problem is about capacity: a
55
+ // model_not_found or bad-field refusal came with ~60 lines of local-model
56
+ // capacity JSON that had nothing to do with it (a Senti review).
57
+ export function refusalText(problems, snapshot) {
58
+ const aboutCapacity = problems.some((p) => /capacity|memory|context|slot|MAX_(POOL_)?WORKERS|max_concurrent/i.test(p));
59
+ return `REFUSED - nothing was started.\n${problems.map(p => `- ${p}`).join("\n")}${aboutCapacity ? `\n\nCapacity right now:\n${JSON.stringify(snapshot(), null, 2)}` : ""}`;
60
+ }
61
+
62
+ export function createJobRuntime(deps) {
63
+ const { projectDir, stateRoot, jobsRoot, runsRoot, leasesRoot, slotsRoot, run, currentMaxWorkers, slug, agentsConfig, modelCatalogReady, budgetsForJob, resolveSubscriptionSelection, executeJob, subscriptionJobFieldProblems, repoPolicy, jobArgs } = deps;
64
+
65
+ // Staggers concurrent job starts by `slot * staggerMs` before each runner
66
+ // begins pulling work. Verified root cause: two OpenClaw sandbox containers
67
+ // created in the same instant reliably hit a podman/crun race ("crun: mount
68
+ // `devpts` to `dev/pts`: Invalid argument"), even with ample host and VM
69
+ // memory free -- reproduced twice, unrelated to memory pressure. A short
70
+ // stagger between concurrent `podman create`/`run` invocations gives crun's
71
+ // container-creation critical section enough separation to not collide.
72
+ //
73
+ // That original fix/measurement was only verified at 2-way concurrency.
74
+ // Re-verified at 4-way (this session): the same race still fired with the
75
+ // stagger active -- one job failed on this exact error within 5.2s of a
76
+ // 4-job concurrent dispatch. 1500ms of separation between ADJACENT slot
77
+ // starts is not consistently enough once 4 containers are all competing for
78
+ // the same crun critical section under real system load, not 2. Raised to
79
+ // 3000ms as a direct response to that reproduction; RETRY_TRANSIENT_SANDBOX_ERRORS
80
+ // below is the second, more robust layer -- no fixed stagger value can be
81
+ // proven sufficient for every load condition, only likely-sufficient.
82
+ const WORKER_START_STAGGER_MS = Number.parseInt(process.env.NOMARMY_WORKER_START_STAGGER_MS ?? "", 10) || 3000;
83
+
84
+ // ---------------------------------------------------------------------------
85
+ // Job registry and admission. Every job, blocking or backgrounded, is tracked
86
+ // here so capacity counts all of them. Admission re-reads the budget (a
87
+ // restarted llama-server or changed profile is picked up) and refuses under
88
+ // memory pressure rather than shrinking the brief and hoping.
89
+ // ---------------------------------------------------------------------------
90
+ const activeJobs = new Map();
91
+ // Counted across every session on this machine, not just this server's own
92
+ // jobs: each coordinator session runs its own server, and per-process
93
+ // counts let six sessions each run their "one" local job at once. Idle
94
+ // sessions hold no leases and count for nothing.
95
+ function runningCount(lane = null) {
96
+ return liveLeases(leasesRoot, lane ? { lane } : {}).length;
97
+ }
98
+
99
+ /** An api or subscription agent's max_concurrent (1 for a subscription, 2 for api by default); null for local. */
100
+ function agentMaxConcurrent(agentName) {
101
+ try {
102
+ const agent = agentsConfig().agents[agentName];
103
+ return agent && agent.kind !== "local" ? agent.max_concurrent ?? (agent.kind === "subscription" ? 1 : 2) : null;
104
+ } catch { return null; }
105
+ }
106
+
107
+ /**
108
+ * Run a job holding one of its agent's max_concurrent slots, machine-wide
109
+ * (lib/slots.mjs), so `max_concurrent: 1` on a subscription means one job
110
+ * on it across every session -- per-session counting never enforced that,
111
+ * and for subscriptions the count was never checked at all. `waitMs` lets a
112
+ * batch queue for a slot instead of failing.
113
+ */
114
+ function withAgentSlot(args, jobId, fn, { waitMs = 0 } = {}) {
115
+ const max = args.agentName ? agentMaxConcurrent(args.agentName) : null;
116
+ if (!max) return fn();
117
+ return (async () => {
118
+ const slot = await acquireSlot(slotsRoot, args.agentName, max, { jobId, waitMs });
119
+ if (!slot) throw new Error(`agent "${args.agentName}" is at its max_concurrent (${max}) across every nomArmy session on this machine; try again when one of its jobs finishes`);
120
+ try { return await fn(); } finally { slot.release(); }
121
+ })();
122
+ }
123
+ function track(jobId, meta, promise) {
124
+ const entry = { ...meta, jobId, startedAt: new Date().toISOString(), settled: false, result: null, error: null, promise: null };
125
+ // A machine-wide lease for as long as the job runs, so every session's
126
+ // admission counts it (runningCount); released however the job ends.
127
+ // `repo` lets each session's status line show its own repo's jobs.
128
+ if (meta.lane) writeLease(leasesRoot, jobId, { lane: meta.lane, agent: meta.agent ?? null, runId: meta.runId ?? null, role: meta.role ?? null, model: meta.model ?? null, repo: projectDir });
129
+ const release = () => removeLease(leasesRoot, jobId);
130
+ entry.promise = promise.then(
131
+ r => { entry.settled = true; entry.result = r; release(); notifyJobFinished(entry, r, null); return r; },
132
+ e => { entry.settled = true; entry.error = e; release(); notifyJobFinished(entry, null, e); throw e; });
133
+ entry.promise.catch(() => {});
134
+ activeJobs.set(jobId, entry);
135
+ return entry;
136
+ }
137
+ /**
138
+ * A desktop notification when a job ends (lib/notify.mjs), so the person
139
+ * watching hears about it from any coordinator without polling.
140
+ */
141
+ function notifyJobFinished(entry, result, error) {
142
+ if (!entry.lane) return; // only tracked jobs, never internal helpers
143
+ const m = result?.manifest ?? {};
144
+ const outcome = error ? "failed" : String(m.outcome ?? (result?.ok ? "done" : "finished")).toLowerCase().replace(/_/g, " ");
145
+ const who = entry.agent ? `${entry.agent}${entry.model ? `/${entry.model}` : ""}` : "local model";
146
+ const took = Math.round((Date.now() - Date.parse(entry.startedAt)) / 60000);
147
+ const ok = !error && (result?.ok || m.coordinatorStatus === "complete");
148
+ notify(`nomArmy: ${entry.role ?? entry.mode ?? "job"} ${ok ? "done" : outcome}`, `${entry.workerId ?? entry.jobId} on ${who}: ${outcome} after ${took}m. ${ok ? "Ready for the General's review." : "Needs a look."}`);
149
+ }
150
+ function capacitySnapshot() {
151
+ const admission = assessAdmission({ hardware: deps.budgetState.hardwareSnapshot, runningJobs: runningCount("local"), slots: deps.budgetState.contextInfo.slots, maxWorkers: currentMaxWorkers() });
152
+ return {
153
+ // The local model's budget. An api or subscription job's scales with
154
+ // its own model; local_worker_start reports that job's.
155
+ budgets: { ...deps.budgetState.budgets, describe: describeBudgets(deps.budgetState.budgets) },
156
+ context: deps.budgetState.contextInfo,
157
+ admission,
158
+ memory: deps.budgetState.hardwareSnapshot?.memory ?? null,
159
+ running: [...activeJobs.values()].filter(j => !j.settled).map(j => ({ jobId: j.jobId, workerId: j.workerId, mode: j.mode, lane: j.lane, startedAt: j.startedAt, phase: readJson(path.join(jobsRoot, j.jobId, "status.json"))?.phase ?? "starting" })),
160
+ maxWorkers: currentMaxWorkers(),
161
+ remote: { running: runningCount("remote"), maxWorkers: currentMaxPoolWorkers(), note: "api and subscription agents; each agent's own max_concurrent also applies" }
162
+ };
163
+ }
164
+ async function admit(jobs) {
165
+ await deps.budgetState.refresh();
166
+ if (jobs.some((j) => jobLane(j) === "remote")) await modelCatalogReady();
167
+ const problems = [];
168
+ // A pool-routed job is checked against that pool's OWN (model-dependent)
169
+ // budget, not the local-derived global one -- see budgetsForPool. Which
170
+ // specific entry pickProvider will land on isn't known yet at admission
171
+ // time, so this is the conservative minimum across the pool's currently
172
+ // available entries, not any one entry's precise number. A
173
+ // subscription_worker job budgets against that one named entry directly
174
+ // (see budgetsForSubscriptionWorker) -- there's no "which entry" unknown
175
+ // the way a weighted pool has, since the name given IS the entry.
176
+ jobs.forEach((j, i) => {
177
+ const jobBudgets = budgetsForJob(j);
178
+ for (const p of checkBrief(j, jobBudgets)) problems.push(jobs.length > 1 ? `job ${i + 1}: ${p}` : p);
179
+ });
180
+ // verify_regression re-runs `verification`; with no profile set there is
181
+ // nothing to re-run. Refuse before starting anything, matching every other
182
+ // admission check here, rather than silently no-op at runtime.
183
+ jobs.forEach((j, i) => {
184
+ if (j.verify_regression && !j.verification) {
185
+ problems.push(`${jobs.length > 1 ? `job ${i + 1}: ` : ""}verify_regression requires a verification profile; there is nothing to run twice without one`);
186
+ }
187
+ });
188
+ // subscription_worker/on_behalf_of: the owner-match attestation refusal
189
+ // happens here, before a container is ever provisioned -- matching how a
190
+ // bad `pool` name is already caught before dispatch, not mid-flight. Only
191
+ // attempted once the plain field-presence problems above are already
192
+ // clean, so a missing on_behalf_of is never reported twice in two
193
+ // different shapes.
194
+ jobs.forEach((j, i) => {
195
+ const fieldProblems = subscriptionJobFieldProblems(j);
196
+ for (const p of fieldProblems) problems.push(jobs.length > 1 ? `job ${i + 1}: ${p}` : p);
197
+ if (fieldProblems.length === 0 && j.on_behalf_of) {
198
+ try {
199
+ if (j.subscription_worker) resolveSubscriptionSelection(j.subscription_worker, j.on_behalf_of, j.reasoning, { model: j.model });
200
+ } catch (error) { problems.push(jobs.length > 1 ? `job ${i + 1}: ${error.message}` : error.message); }
201
+ }
202
+ });
203
+ // The repo's own policy: verification required, revert check required.
204
+ const policy = repoPolicy();
205
+ jobs.forEach((j, i) => { for (const p of policyAdmissionProblems(j, policy)) problems.push(`${jobs.length > 1 ? `job ${i + 1}: ` : ""}${p}`); });
206
+ // An implement job on an agent whose own tools run on this machine (the
207
+ // Claude CLI) isn't bounded by the sandbox, so it's refused unless that
208
+ // agent says allow_host_tools (lib/agents.mjs). Scouts and reviews still run.
209
+ jobs.forEach((j, i) => {
210
+ if (!j.agentName || (j.mode ?? "implement") !== "implement") return;
211
+ let problem = null;
212
+ try { problem = hostToolsImplementProblem(j.agentName, agentsConfig().agents[j.agentName]); } catch { return; }
213
+ if (problem) problems.push(`${jobs.length > 1 ? `job ${i + 1}: ` : ""}${problem}`);
214
+ });
215
+ // A model its vendor refused on a job today, with nothing working on it
216
+ // since, isn't sent another job (lib/health.mjs recentModelRefusal).
217
+ jobs.forEach((j, i) => {
218
+ if (!j.agentName || !j.model) return;
219
+ let provider = null;
220
+ try { provider = agentProviderId(agentsConfig().agents[j.agentName]); } catch { return; }
221
+ if (!provider) return;
222
+ const refusal = recentModelRefusal(stateRoot, `${provider}/${j.model}`);
223
+ if (refusal) problems.push(`${jobs.length > 1 ? `job ${i + 1}: ` : ""}model_not_found: ${provider}/${j.model} was refused on an earlier job today and hasn't worked since, so this job wasn't sent. Use another model (the job's \`model\`, or \`nomarmy army assign\`); \`nomarmy army assign <role> ${j.agentName} ${j.model}\` re-tests it, and a passing test clears this.`);
224
+ });
225
+ // An agent's max_concurrent, machine-wide. Batch jobs on the same agent
226
+ // queue for its slot at launch instead (withAgentSlot's waitMs).
227
+ if (jobs.length === 1) {
228
+ const [j] = jobs;
229
+ const max = j.agentName ? agentMaxConcurrent(j.agentName) : null;
230
+ const held = max ? liveSlots(slotsRoot, j.agentName) : 0;
231
+ if (max && held >= max) problems.push(`not admitted (capacity): agent "${j.agentName}" already has ${held} job(s) running across this machine's nomArmy sessions, at its max_concurrent of ${max}`);
232
+ }
233
+ // A job in a /feature run: the run's own limits and paused agents.
234
+ jobs.forEach((j, i) => {
235
+ if (!j.run_id) return;
236
+ try {
237
+ const run = loadRun(runsRoot, j.run_id);
238
+ if (run.repo !== projectDir) problems.push(`${jobs.length > 1 ? `job ${i + 1}: ` : ""}run "${run.id}" belongs to ${run.repo}, not this repository`);
239
+ const running = liveLeases(leasesRoot, { runId: run.id }).length + jobs.slice(0, i).filter((o) => o.run_id === run.id).length;
240
+ for (const p of runAdmissionProblems(run, { agentName: j.agentName ?? "local", running })) problems.push(jobs.length > 1 ? `job ${i + 1}: ${p}` : p);
241
+ } catch (error) { problems.push(jobs.length > 1 ? `job ${i + 1}: ${error.message}` : error.message); }
242
+ });
243
+ // Slot capacity only concerns local jobs: a remote job's inference runs
244
+ // at its vendor and never competes for llama-server's slots. Free memory
245
+ // still applies to every job (each one runs a local sandbox), so a
246
+ // remote-only batch is checked for memory alone. Remote jobs have their
247
+ // own, additive ceiling (currentMaxPoolWorkers).
248
+ const anyLocal = jobs.some((j) => jobLane(j) === "local");
249
+ const admission = anyLocal
250
+ ? assessAdmission({ hardware: deps.budgetState.hardwareSnapshot, runningJobs: runningCount("local"), slots: deps.budgetState.contextInfo.slots, maxWorkers: currentMaxWorkers() })
251
+ : assessAdmission({ hardware: deps.budgetState.hardwareSnapshot, runningJobs: 0, slots: null, maxWorkers: Infinity });
252
+ if (!admission.admit) problems.push(...admission.reasons.map(r => `not admitted (${admission.level}): ${r}`));
253
+ if (jobs.some((j) => jobLane(j) === "remote")) {
254
+ const remoteCeiling = currentMaxPoolWorkers(), runningRemote = runningCount("remote");
255
+ if (runningRemote >= remoteCeiling) {
256
+ problems.push(`not admitted (capacity): ${runningRemote} remote job(s) (api or subscription agents) already running, at NOMARMY_MAX_POOL_WORKERS=${remoteCeiling}`);
257
+ }
258
+ }
259
+ return { problems, admission };
260
+ }
261
+
262
+ function refusal(problems) {
263
+ return toolText(refusalText(problems, capacitySnapshot), true);
264
+ }
265
+ /** A run's totals and warnings, for a tool response. */
266
+ function runBrief(runId) {
267
+ try {
268
+ const run = loadRun(runsRoot, runId);
269
+ const totals = runTotals(run);
270
+ return { id: run.id, status: run.status, limits: run.limits, used: totals.used, warnings: totals.warnings };
271
+ } catch (error) { return { id: runId, error: error.message }; }
272
+ }
273
+
274
+ /**
275
+ * Record a finished job into its run. A usage-limit message is looked for
276
+ * only in error text (OpenClaw's failure envelope, and the error lines of
277
+ * a thrown run), never in the worker's report or tool output, where "rate
278
+ * limit" may just be the code under review.
279
+ */
280
+ function recordJobInRun(args, jobId, result, error = null) {
281
+ if (!args.run_id) return;
282
+ const kind = args.pool ? "api" : args.subscription_worker ? "subscription" : "local";
283
+ const m = result?.manifest ?? {};
284
+ const errorLines = [m.worker?.error, error?.message,
285
+ ...String(m.workerError ?? "").split(/\r?\n/).filter((l) => /error|limit|429/i.test(l))].filter(Boolean).join("\n");
286
+ const usageLimit = kind === "local" ? null : detectUsageLimit(errorLines);
287
+ try {
288
+ const before = runTotals(loadRun(runsRoot, args.run_id)).warnings;
289
+ const updated = recordRunJob(runsRoot, args.run_id, {
290
+ jobId, agent: args.agentName ?? "local", kind, model: args.model ?? null, role: args.armyRole ?? null, mode: args.mode,
291
+ outcome: m.outcome ?? (error ? "ERROR" : null), costUsd: m.metrics?.worker_cost_usd ?? null,
292
+ tokens: m.metrics?.worker_tokens_total ?? null, usageLimit,
293
+ });
294
+ // A limit crossed or an agent paused by this job is worth interrupting for.
295
+ const fresh = runTotals(updated).warnings.filter((w) => !before.includes(w) && /OVER|paused/.test(w));
296
+ if (fresh.length) notify(`nomArmy run ${updated.name}: stopped short`, fresh.join("; "));
297
+ } catch (recordError) {
298
+ fs.appendFileSync(path.join(jobsRoot, jobId, "coordinator.log"), `${new Date().toISOString()} could not record into run ${args.run_id}: ${recordError.message}\n`);
299
+ }
300
+ }
301
+ function trackInRun(args, entry) {
302
+ if (args.run_id) entry.promise.then((r) => recordJobInRun(args, entry.jobId, r), (e) => recordJobInRun(args, entry.jobId, null, e));
303
+ return entry;
304
+ }
305
+ function launch(args) {
306
+ const workerId = args.worker_id || null;
307
+ const jobId = slug(workerId || (args.mode === "scout" ? "scout" : "worker"));
308
+ return trackInRun(args, track(jobId, { mode: args.mode, workerId: workerId || jobId, lane: jobLane(args), agent: args.agentName ?? null, runId: args.run_id ?? null, role: args.armyRole ?? null, model: args.model ?? null },
309
+ withAgentSlot(args, jobId, () => executeJob({ ...jobArgs(args, workerId), jobId }))));
310
+ }
311
+ // Best-effort progress signal for a job still mid-run: a plain "phase: worker,
312
+ // elapsed: Ns" told a caller nothing about whether the worker was still
313
+ // reading or already editing, short of running `git status` on the worktree
314
+ // by hand. Both lookups here are read-only and disposable -- a job's worktree
315
+ // mid-write or a transcript sqlite file mid-append can legitimately fail to
316
+ // read, and that must never fail the status call, only omit the field.
317
+ async function liveProgress(jobDir) {
318
+ const out = {};
319
+ try {
320
+ const worktree = path.join(jobDir, "worktree");
321
+ if (fs.existsSync(worktree)) {
322
+ // --untracked-files=normal, not all: "all" descends into every
323
+ // untracked directory (a virtualenv, a cache) a job creates. Bounded:
324
+ // a live progress read must never hold anything up.
325
+ const statusOut = (await run("git", ["status", "--porcelain=v1", "-z", "--untracked-files=normal"], { cwd: worktree, trim: false, timeoutMs: 10000 })).stdout;
326
+ // Same runtime-junk filter as collectGitRecord/makeIdleDiffTick: .npm/
327
+ // etc. is the sandbox's own churn, not the worker's progress, and
328
+ // counting it made a job that had made zero real edits report
329
+ // filesChangedLive: 1 anyway.
330
+ out.filesChangedLive = parseStatusPorcelainZ(statusOut).map(e => e.file).filter(f => !isRuntimeJunk(f)).length;
331
+ }
332
+ } catch { /* worktree not ready yet, or mutated mid-read; omit */ }
333
+ try {
334
+ const stateDir = path.join(jobDir, "runtime", "state");
335
+ const transcript = await readOpenClawTranscriptTail(stateDir, { limit: 6 });
336
+ if (transcript.available) {
337
+ const last = transcript.toolCalls.at(-1);
338
+ if (last) out.lastTool = { tool: last.tool, target: last.path ?? last.command ?? null };
339
+ }
340
+ // A claude-cli worker's tools only appear in Claude Code's own session
341
+ // transcript, not OpenClaw's.
342
+ if (!out.lastTool) {
343
+ const startedMs = Date.parse(readJson(path.join(jobDir, "status.json"))?.startedAt ?? "") || 0;
344
+ const claude = readClaudeSessionTranscript(path.join(jobDir, "worktree"), { sinceMs: startedMs, tailBytes: 262144 });
345
+ const last = claude.available ? claude.toolCalls.at(-1) : null;
346
+ if (last) { out.lastTool = { tool: last.tool, target: last.path ?? last.command ?? null }; out.toolCallsLive = claude.toolCalls.length; }
347
+ }
348
+ } catch { /* transcript not created yet, or locked mid-write; omit */ }
349
+ return out;
350
+ }
351
+
352
+ async function summarize(entry, files, jobDir = null) {
353
+ const status = files.status, meta = files.meta ?? files.failure;
354
+ const elapsedSeconds = status?.startedAt ? Math.round((Date.now() - Date.parse(status.startedAt)) / 1000) : entry ? Math.round((Date.now() - Date.parse(entry.startedAt)) / 1000) : null;
355
+ const out = { jobId: entry?.jobId ?? status?.jobId ?? meta?.jobId ?? null, workerId: entry?.workerId ?? status?.workerId ?? meta?.workerId ?? null,
356
+ mode: entry?.mode ?? status?.mode ?? meta?.mode ?? null, state: null, phase: status?.phase ?? "starting", elapsedSeconds,
357
+ timeoutSeconds: status?.timeoutSeconds ?? null, coordinatorStatus: meta?.coordinatorStatus ?? null, outcome: meta?.outcome ?? null,
358
+ reviewRequired: meta?.reviewRequired ?? null, issues: (meta?.issues ?? []).slice(0, 6), worktree: meta?.worktree ?? null, branch: meta?.branch ?? null,
359
+ commit: meta?.commit?.sha ?? null, scout: meta?.scout ? { supported: meta.scout.supported, unsupported: meta.scout.unsupported } : null };
360
+ if (entry && !entry.settled) out.state = "running";
361
+ else if (entry?.error) { out.state = "failed"; out.error = String(entry.error.message ?? entry.error).split("\n")[0]; }
362
+ else if (meta) out.state = "finished";
363
+ else if (status?.state === "running") { out.state = status.serverPid === process.pid ? "running" : "orphaned"; if (out.state === "orphaned") out.error = `the MCP server that ran this job (pid ${status.serverPid}) is gone; outcome unknown, see the job directory logs`; }
364
+ else out.state = "unknown";
365
+ if (out.state === "running" && jobDir) Object.assign(out, await liveProgress(jobDir));
366
+ return out;
367
+ }
368
+
369
+ return { WORKER_START_STAGGER_MS, activeJobs, runningCount, agentMaxConcurrent, withAgentSlot, track, notifyJobFinished, capacitySnapshot, admit, refusal, runBrief, recordJobInRun, trackInRun, launch, liveProgress, summarize };
370
+ }