nomarmy 0.1.0-alpha.0 → 0.1.0-alpha.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +40 -28
- package/e2e.sh +38 -6
- package/lib/admission.mjs +370 -0
- package/lib/agent-config.mjs +85 -0
- package/lib/budget-state.mjs +48 -0
- package/lib/diff-checks.mjs +418 -0
- package/lib/execute.mjs +728 -0
- package/lib/git-record.mjs +189 -0
- package/lib/job-budgets.mjs +104 -0
- package/lib/job-format.mjs +118 -0
- package/lib/openclaw-run.mjs +658 -0
- package/lib/outcome.mjs +237 -0
- package/lib/outcomes.mjs +27 -0
- package/lib/process.mjs +131 -0
- package/lib/propose.mjs +5 -0
- package/lib/report.mjs +104 -0
- package/lib/schema.mjs +20 -0
- package/lib/selection.mjs +109 -0
- package/lib/server-context.mjs +13 -0
- package/lib/verification-flow.mjs +316 -0
- package/lib/worker-prompt.mjs +78 -0
- package/mcp/server.mjs +138 -3531
- package/package.json +1 -1
- package/policies/coder.md +1 -1
package/README.md
CHANGED
|
@@ -8,9 +8,28 @@
|
|
|
8
8
|
|
|
9
9
|
<p align="center"><em>Tiny coders, big appetites for bounded tickets.</em> 🍪</p>
|
|
10
10
|
|
|
11
|
-
**
|
|
11
|
+
**Your coding assistant plans; sandboxed workers build; nothing counts until nomArmy has checked it.**
|
|
12
12
|
|
|
13
|
-
|
|
13
|
+
## TL;DR
|
|
14
|
+
|
|
15
|
+
1. **Have** Git, Node 20+ and [Podman](https://podman.io) (on macOS: `brew install podman && podman machine init && podman machine start`).
|
|
16
|
+
2. **Install** (builds the local model server, OpenClaw and the sandbox, and registers nomArmy with Claude Code):
|
|
17
|
+
```bash
|
|
18
|
+
git clone https://github.com/rayson-tech/nomarmy.git && cd nomarmy
|
|
19
|
+
./install.sh --profile macbook-pro # or nvidia-linux, cpu-linux, dgx-spark, bedrock
|
|
20
|
+
./e2e.sh --profile macbook-pro # should end with: === E2E PASS ===
|
|
21
|
+
```
|
|
22
|
+
3. **Set up your repo:** in the project, run `nomarmy init`. It proposes a `.nomarmy.yml` with your test command.
|
|
23
|
+
4. **Use it:** restart Claude Code in that project and ask it to use nomArmy for one small bug that has a test. When that works, try `/feature <what you want built>`.
|
|
24
|
+
5. **Optional:** add hosted workers with `nomarmy agents add`, give roles to them with `nomarmy army init`, or connect Codex or Cursor with `nomarmy connect codex cursor`.
|
|
25
|
+
|
|
26
|
+
Stuck? `nomarmy doctor` checks the machine, and `nomarmy health` checks everything nomArmy runs on.
|
|
27
|
+
|
|
28
|
+
## What it is
|
|
29
|
+
|
|
30
|
+
Your coding assistant (Claude Code, Codex or Cursor) stays in charge as the **General**: it decides what gets built and whether the result is acceptable. The work goes to **noms**, workers that implement, test and repair in their own git worktree and sandbox, on a local model, an API key, or your own ChatGPT or Muse Code subscription. nomArmy owns everything in between: worktrees, git, sandboxes, verification, and the evidence that decides whether work is accepted.
|
|
31
|
+
|
|
32
|
+
**What you get is work you don't have to take on faith**, not cheaper work. Delegating costs the General tokens too: briefing and reviewing. On small, already-diagnosed tickets we measured 4 to 8 times more of the General's tokens than fixing the bug directly, and break-even at roughly 150 lines of context a fix needs to read ([the measurements](docs/experiments/2026-09-20-model-bakeoff-and-economics.md)). It pays off on bigger tickets, on parallel work, and anywhere you'd otherwise have to trust an agent's say-so.
|
|
14
33
|
|
|
15
34
|
Developed and maintained by Rayson Technologies. This is an alpha (`0.1.0-alpha`).
|
|
16
35
|
|
|
@@ -22,35 +41,10 @@ Developed and maintained by Rayson Technologies. This is an alpha (`0.1.0-alpha`
|
|
|
22
41
|
4. nomArmy treats that report as a claim. It reads the real diff from git, runs your verification profile itself in a fresh sandbox, reverts the production change to check the tests actually fail without it, and scans for secrets.
|
|
23
42
|
5. Only then does it commit, on the worker's own branch. It never merges into yours: reviewing and integrating stay with the General, and with you.
|
|
24
43
|
|
|
25
|
-
A malformed report isn't automatically a failure: if the repository changed, nomArmy verifies independently and may recover the work. Failing verification stays failed, unconditionally.
|
|
44
|
+
A malformed report isn't automatically a failure: if the repository changed, nomArmy verifies independently and may recover the work. Failing verification stays failed, unconditionally. And the checks aren't the General's to waive: a repo's `.nomarmy.yml` policy (on by default for new repos) makes verification and the revert check mandatory for every job.
|
|
26
45
|
|
|
27
46
|
Around that core: **agents** say where a job can run, the **army** says which role runs on which agent, and **`/feature`** runs a whole feature end to end, from plan through build, review and acceptance, handing you a branch to merge.
|
|
28
47
|
|
|
29
|
-
## Quick start
|
|
30
|
-
|
|
31
|
-
You need Git, Node 20+, and [Podman](https://podman.io).
|
|
32
|
-
|
|
33
|
-
```bash
|
|
34
|
-
npm install -g nomarmy@alpha
|
|
35
|
-
nomarmy doctor # what's missing on this machine, and how to fix it
|
|
36
|
-
nomarmy setup # pick a profile and model; prints the install.sh command to run
|
|
37
|
-
nomarmy connect claude # or codex, cursor: registers nomArmy with your coordinator
|
|
38
|
-
```
|
|
39
|
-
|
|
40
|
-
Installing from a clone is the most tested path, and it's what `install.sh` expects: see [Install](#install). `install.sh` builds llama.cpp for local inference, installs and configures [OpenClaw](https://github.com/openclaw/openclaw) (the host-side broker every model call goes through), builds the sandbox image and registers the MCP server.
|
|
41
|
-
|
|
42
|
-
`nomarmy connect` also installs the `/feature` command, Claude Code's status line and (on macOS) nomArmy's notifier. The coordinator gets nomArmy's instructions from the MCP server itself, so there's nothing to copy into your projects.
|
|
43
|
-
|
|
44
|
-
Then, in any git repository:
|
|
45
|
-
|
|
46
|
-
```bash
|
|
47
|
-
nomarmy init # propose a .nomarmy.yml with your test command
|
|
48
|
-
nomarmy agents add # optional: an API key or a subscription
|
|
49
|
-
nomarmy army init # optional: the default roster of roles
|
|
50
|
-
```
|
|
51
|
-
|
|
52
|
-
Ask your coordinator to delegate one small, well-tested ticket before anything bigger. When that works, try `/feature <what you want built>`.
|
|
53
|
-
|
|
54
48
|
## Install
|
|
55
49
|
|
|
56
50
|
| Platform | Guide |
|
|
@@ -63,6 +57,10 @@ Ask your coordinator to delegate one small, well-tested ticket before anything b
|
|
|
63
57
|
|
|
64
58
|
Every platform needs Git and Podman. `nomarmy doctor` checks the host and prints a fix for anything missing.
|
|
65
59
|
|
|
60
|
+
`install.sh` builds llama.cpp for local inference, installs and configures [OpenClaw](https://github.com/openclaw/openclaw) (the host-side broker every model call goes through), builds the sandbox image, and registers the MCP server if Claude Code is installed. `nomarmy connect` (run by `install.sh`, or by hand for Codex and Cursor) also installs the `/feature` command, Claude Code's status line and, on macOS, nomArmy's notifier. The coordinator gets nomArmy's instructions from the MCP server itself, so there's nothing to copy into your projects.
|
|
61
|
+
|
|
62
|
+
**From npm:** `npm install -g nomarmy@alpha` gives you the `nomarmy` command; `nomarmy setup` then picks a profile and model and prints the `install.sh` command to run. Installing from a clone, as in the TL;DR, is the most tested path.
|
|
63
|
+
|
|
66
64
|
### macOS (Apple Silicon)
|
|
67
65
|
|
|
68
66
|
```bash
|
|
@@ -179,6 +177,8 @@ Changes apply to the next job with no restart. The exception is a **new** api ag
|
|
|
179
177
|
|
|
180
178
|
**Your plan decides which models run.** A model can be listed and still refused: on a ChatGPT plan, the Codex route runs gpt-6-astra and the gpt-5.6 models but refuses gpt-6-sol and gpt-6-luna. `army assign` and `agents update --probe` test the exact route a job takes, so they catch this before a job does.
|
|
181
179
|
|
|
180
|
+
**Vendor terms and platform risk.** Every model call goes through [OpenClaw](https://github.com/openclaw/openclaw), and subscriptions are reached through each vendor's own CLI or login. We've read the terms that apply (see above), but using a personal subscription through a harness is exactly the kind of use vendors tighten, and a change in a vendor's terms or in OpenClaw can stop a subscription agent from working. Local models and API keys don't carry that risk. Plan on subscriptions as a convenience, not the only way your roles can run.
|
|
181
|
+
|
|
182
182
|
**Picking an agent.** Build work goes to a sandboxed agent: `local`, an api key, Codex or Muse. `local` for a bounded change against a written spec with a test; your code never leaves your machine. An api or subscription agent when the work needs more than the local model, knowing it sends code to that vendor. That's a decision about where your source travels, separate from the trust boundary, which is the same for every agent. The General itself when the answer isn't known yet.
|
|
183
183
|
|
|
184
184
|
## The army: who does what
|
|
@@ -270,6 +270,18 @@ verification:
|
|
|
270
270
|
|
|
271
271
|
`nomarmy validate` checks the file against the schema; `nomarmy scan --check` diffs it against what the repo actually contains.
|
|
272
272
|
|
|
273
|
+
**Policy: what no job can skip.**
|
|
274
|
+
|
|
275
|
+
```yaml
|
|
276
|
+
policy:
|
|
277
|
+
require_verification: true # every implement job needs a verification profile; only passing work commits
|
|
278
|
+
require_regression_check: true # verify_regression can't be switched off per job
|
|
279
|
+
```
|
|
280
|
+
|
|
281
|
+
`nomarmy init` proposes both for new repos. Without them, a job with no verification profile still commits (flagged for review, not blocked), and the General decides per job whether to run the revert check. With them, those are the repo's rules, not the General's judgment calls, and since nomArmy reads this file only from your checkout, neither the General nor a worker can relax it.
|
|
282
|
+
|
|
283
|
+
**Refactors.** Reverting a behavior-preserving change restores code that works, so the revert check can't prove anything about it. A job can declare `refactor: true` instead: nomArmy then commits it only if verification passes **and no test file was added, changed or deleted**. The existing tests passing unchanged is the evidence. A change that alters behavior has to alter tests to show it, so it can't pass as a refactor.
|
|
284
|
+
|
|
273
285
|
**Add a check for what unit tests can't see.** A module left out of a deploy bundle passes every unit test and crashes at deploy. When `nomarmy init` sees a bundle or packaging step (Lambda asset scripts, SAM, Serverless, CDK), it suggests a profile that runs it and then imports each entry point from the built bundle.
|
|
274
286
|
|
|
275
287
|
### Languages and dependencies
|
package/e2e.sh
CHANGED
|
@@ -56,18 +56,35 @@ fi
|
|
|
56
56
|
# bind-mounting paths under the host home directory, and the cleanup step
|
|
57
57
|
# below bind-mounts this directory into a container.
|
|
58
58
|
TMP="$(mktemp -d "$HOME/.nomarmy-e2e.XXXXXX")"
|
|
59
|
+
# OpenClaw's own state for this run, also under $HOME, the way nomArmy's
|
|
60
|
+
# real dispatch keeps it (--state-dir). Left unset, OpenClaw puts its
|
|
61
|
+
# working files in the system temp folder, which the Podman sandbox can't
|
|
62
|
+
# mount on macOS, and it uses the operator's main state, whose memory index
|
|
63
|
+
# holds their past sessions.
|
|
64
|
+
STATE="$(mktemp -d "$HOME/.nomarmy-e2e-state.XXXXXX")"
|
|
59
65
|
|
|
60
66
|
cleanup() {
|
|
61
67
|
local exit_code=$?
|
|
62
68
|
|
|
63
|
-
|
|
69
|
+
# OpenClaw leaves this run's sandbox container running; its name carries
|
|
70
|
+
# the hash recorded under the state dir (as nomArmy's job cleanup does).
|
|
71
|
+
if [[ -n "${STATE:-}" && -d "$STATE/state/sandbox/skills-workspaces" ]] && command -v podman >/dev/null 2>&1; then
|
|
72
|
+
for ws in "$STATE"/state/sandbox/skills-workspaces/workspace-*; do
|
|
73
|
+
[[ -d "$ws" ]] || continue
|
|
74
|
+
podman ps -a --filter "name=${ws##*/workspace-}" --format '{{.Names}}' 2>/dev/null \
|
|
75
|
+
| xargs -r podman rm -f -v >/dev/null 2>&1 || true
|
|
76
|
+
done
|
|
77
|
+
fi
|
|
78
|
+
|
|
79
|
+
for dir in "${TMP:-}" "${STATE:-}"; do
|
|
80
|
+
if [[ -n "$dir" && -d "$dir" ]]; then
|
|
64
81
|
|
|
65
82
|
# OpenClaw's Podman sandbox may create files that the host user
|
|
66
83
|
# cannot delete directly. Use a disposable container to clean
|
|
67
84
|
# the temporary workspace first.
|
|
68
85
|
if command -v podman >/dev/null 2>&1 && podman info >/dev/null 2>&1; then
|
|
69
86
|
podman run --rm \
|
|
70
|
-
-v "$
|
|
87
|
+
-v "$dir:/cleanup" \
|
|
71
88
|
alpine:3.20 \
|
|
72
89
|
sh -c '
|
|
73
90
|
find /cleanup -mindepth 1 -maxdepth 1 -exec rm -rf -- {} + \
|
|
@@ -76,13 +93,14 @@ cleanup() {
|
|
|
76
93
|
>/dev/null 2>&1 || true
|
|
77
94
|
fi
|
|
78
95
|
|
|
79
|
-
rm -rf "$
|
|
96
|
+
rm -rf "$dir" 2>/dev/null || true
|
|
80
97
|
|
|
81
|
-
if [[ -d "$
|
|
98
|
+
if [[ -d "$dir" ]]; then
|
|
82
99
|
echo "WARN: E2E temporary directory could not be completely removed:"
|
|
83
|
-
echo " $
|
|
100
|
+
echo " $dir"
|
|
84
101
|
fi
|
|
85
102
|
fi
|
|
103
|
+
done
|
|
86
104
|
|
|
87
105
|
exit "$exit_code"
|
|
88
106
|
}
|
|
@@ -126,15 +144,29 @@ PROMPT='Fix the bug so npm test passes. Work only in the workspace. Run npm test
|
|
|
126
144
|
|
|
127
145
|
OUT="$TMP/openclaw.json"
|
|
128
146
|
|
|
147
|
+
# The same privacy settings every nomArmy job gets (lib/openclaw-run.mjs
|
|
148
|
+
# withJobPrivacy): OpenClaw's memory search and session-memory hook off, so
|
|
149
|
+
# nothing is indexed or sent for embedding.
|
|
150
|
+
CONFIG="$STATE/openclaw.job.json"
|
|
151
|
+
mkdir -p "$STATE/state"
|
|
152
|
+
node --input-type=module -e '
|
|
153
|
+
import fs from "node:fs";
|
|
154
|
+
import { readOpenclawConfig } from "'"$ROOT"'/lib/openclaw-config.mjs";
|
|
155
|
+
import { withJobPrivacy } from "'"$ROOT"'/lib/openclaw-run.mjs";
|
|
156
|
+
fs.writeFileSync(process.argv[1], JSON.stringify(withJobPrivacy(readOpenclawConfig() ?? {})), { mode: 0o600 });
|
|
157
|
+
' "$CONFIG"
|
|
158
|
+
|
|
129
159
|
openclaw agent exec "$PROMPT" \
|
|
130
160
|
--model "$NOMARMY_WORKER_PROVIDER/$NOMARMY_WORKER_MODEL" \
|
|
131
161
|
--cwd "$TMP" \
|
|
162
|
+
--state-dir "$STATE/state" \
|
|
163
|
+
--config "$CONFIG" \
|
|
132
164
|
--code-mode direct \
|
|
133
165
|
--local-model-lean \
|
|
134
166
|
--thinking off \
|
|
135
167
|
--timeout 600 \
|
|
136
168
|
--json \
|
|
137
|
-
> "$OUT"
|
|
169
|
+
> "$OUT" 2> "$STATE/openclaw.stderr.log"
|
|
138
170
|
|
|
139
171
|
# Independent verification. We do not trust the worker's claim
|
|
140
172
|
# that its implementation is correct.
|
|
@@ -0,0 +1,370 @@
|
|
|
1
|
+
import fs from "node:fs";
|
|
2
|
+
import path from "node:path";
|
|
3
|
+
import { clampInt } from "./budget-state.mjs";
|
|
4
|
+
import { checkBrief, assessAdmission, describeBudgets } from "./budget.mjs";
|
|
5
|
+
import { parseStatusPorcelainZ, isRuntimeJunk } from "./git-record.mjs";
|
|
6
|
+
import { readJson } from "./openclaw-run.mjs";
|
|
7
|
+
import { readOpenClawTranscriptTail } from "./transcript.mjs";
|
|
8
|
+
import { readClaudeSessionTranscript } from "./claude-transcript.mjs";
|
|
9
|
+
import { notify } from "./notify.mjs";
|
|
10
|
+
import { recentModelRefusal } from "./health.mjs";
|
|
11
|
+
import { writeLease, removeLease, liveLeases, liveSlots, acquireSlot } from "./slots.mjs";
|
|
12
|
+
import { loadRun, runTotals, runAdmissionProblems, recordRunJob, detectUsageLimit } from "./runs.mjs";
|
|
13
|
+
import { agentProviderId, hostToolsImplementProblem } from "./agents.mjs";
|
|
14
|
+
import { policyAdmissionProblems } from "./outcome.mjs";
|
|
15
|
+
|
|
16
|
+
// `lane` is "local" (the local model on llama-server) or "remote" (an api
|
|
17
|
+
// or subscription agent: the inference runs at the vendor). The local-slot
|
|
18
|
+
// admission check must only ever count the local lane. A subscription job
|
|
19
|
+
// used to land in "local" (the lane was decided by `pool` alone), so a
|
|
20
|
+
// Claude or Codex job took llama-server's only slot and blocked local work
|
|
21
|
+
// it never competed with -- reported from a real Senti run.
|
|
22
|
+
export function jobLane(job) {
|
|
23
|
+
return job.pool || job.subscription_worker ? "remote" : "local";
|
|
24
|
+
}
|
|
25
|
+
|
|
26
|
+
// A static, operator-declared ceiling on how many remote jobs (api and
|
|
27
|
+
// subscription agents) may run at once, independent of and additive to
|
|
28
|
+
// currentMaxWorkers()'s local ceiling. Each still runs a sandbox and a
|
|
29
|
+
// worktree on this machine, which is what this bounds; each agent's own
|
|
30
|
+
// max_concurrent bounds its vendor. The env name predates agents.yml
|
|
31
|
+
// (remote jobs were all "pool" jobs then) -- exactly the "more real concurrency, not just diversity"
|
|
32
|
+
// benefit of spreading load across providers with their own separate rate
|
|
33
|
+
// limits. Not rate-limit-aware (see config/providers.yml.example); read
|
|
34
|
+
// fresh each call, matching currentMaxWorkers()'s own env-read pattern.
|
|
35
|
+
export function currentMaxPoolWorkers() {
|
|
36
|
+
return clampInt(process.env.NOMARMY_MAX_POOL_WORKERS, 1, 32, 4);
|
|
37
|
+
}
|
|
38
|
+
|
|
39
|
+
// Pure partition of a batch's ORIGINAL indices by lane -- pulled out of
|
|
40
|
+
// local_workers' handler so this specific invariant (every job lands in
|
|
41
|
+
// exactly one lane, indices preserved) is directly testable without also
|
|
42
|
+
// exercising the full async dispatch/mapLimit machinery around it. This is
|
|
43
|
+
// the exact split that used to not exist at all: every job in a batch
|
|
44
|
+
// shared one `parallel` slot count derived only from the local ceiling,
|
|
45
|
+
// which let an all-pool batch ignore NOMARMY_MAX_POOL_WORKERS entirely.
|
|
46
|
+
export function splitJobsByLane(jobs) {
|
|
47
|
+
const localIndices = [], remoteIndices = [];
|
|
48
|
+
jobs.forEach((j, i) => (jobLane(j) === "remote" ? remoteIndices : localIndices).push(i));
|
|
49
|
+
return { localIndices, remoteIndices };
|
|
50
|
+
}
|
|
51
|
+
|
|
52
|
+
export function toolText(text, isError = false) { return { content: [{ type: "text", text }], isError }; }
|
|
53
|
+
|
|
54
|
+
// The capacity snapshot only when a problem is about capacity: a
|
|
55
|
+
// model_not_found or bad-field refusal came with ~60 lines of local-model
|
|
56
|
+
// capacity JSON that had nothing to do with it (a Senti review).
|
|
57
|
+
export function refusalText(problems, snapshot) {
|
|
58
|
+
const aboutCapacity = problems.some((p) => /capacity|memory|context|slot|MAX_(POOL_)?WORKERS|max_concurrent/i.test(p));
|
|
59
|
+
return `REFUSED - nothing was started.\n${problems.map(p => `- ${p}`).join("\n")}${aboutCapacity ? `\n\nCapacity right now:\n${JSON.stringify(snapshot(), null, 2)}` : ""}`;
|
|
60
|
+
}
|
|
61
|
+
|
|
62
|
+
export function createJobRuntime(deps) {
|
|
63
|
+
const { projectDir, stateRoot, jobsRoot, runsRoot, leasesRoot, slotsRoot, run, currentMaxWorkers, slug, agentsConfig, modelCatalogReady, budgetsForJob, resolveSubscriptionSelection, executeJob, subscriptionJobFieldProblems, repoPolicy, jobArgs } = deps;
|
|
64
|
+
|
|
65
|
+
// Staggers concurrent job starts by `slot * staggerMs` before each runner
|
|
66
|
+
// begins pulling work. Verified root cause: two OpenClaw sandbox containers
|
|
67
|
+
// created in the same instant reliably hit a podman/crun race ("crun: mount
|
|
68
|
+
// `devpts` to `dev/pts`: Invalid argument"), even with ample host and VM
|
|
69
|
+
// memory free -- reproduced twice, unrelated to memory pressure. A short
|
|
70
|
+
// stagger between concurrent `podman create`/`run` invocations gives crun's
|
|
71
|
+
// container-creation critical section enough separation to not collide.
|
|
72
|
+
//
|
|
73
|
+
// That original fix/measurement was only verified at 2-way concurrency.
|
|
74
|
+
// Re-verified at 4-way (this session): the same race still fired with the
|
|
75
|
+
// stagger active -- one job failed on this exact error within 5.2s of a
|
|
76
|
+
// 4-job concurrent dispatch. 1500ms of separation between ADJACENT slot
|
|
77
|
+
// starts is not consistently enough once 4 containers are all competing for
|
|
78
|
+
// the same crun critical section under real system load, not 2. Raised to
|
|
79
|
+
// 3000ms as a direct response to that reproduction; RETRY_TRANSIENT_SANDBOX_ERRORS
|
|
80
|
+
// below is the second, more robust layer -- no fixed stagger value can be
|
|
81
|
+
// proven sufficient for every load condition, only likely-sufficient.
|
|
82
|
+
const WORKER_START_STAGGER_MS = Number.parseInt(process.env.NOMARMY_WORKER_START_STAGGER_MS ?? "", 10) || 3000;
|
|
83
|
+
|
|
84
|
+
// ---------------------------------------------------------------------------
|
|
85
|
+
// Job registry and admission. Every job, blocking or backgrounded, is tracked
|
|
86
|
+
// here so capacity counts all of them. Admission re-reads the budget (a
|
|
87
|
+
// restarted llama-server or changed profile is picked up) and refuses under
|
|
88
|
+
// memory pressure rather than shrinking the brief and hoping.
|
|
89
|
+
// ---------------------------------------------------------------------------
|
|
90
|
+
const activeJobs = new Map();
|
|
91
|
+
// Counted across every session on this machine, not just this server's own
|
|
92
|
+
// jobs: each coordinator session runs its own server, and per-process
|
|
93
|
+
// counts let six sessions each run their "one" local job at once. Idle
|
|
94
|
+
// sessions hold no leases and count for nothing.
|
|
95
|
+
function runningCount(lane = null) {
|
|
96
|
+
return liveLeases(leasesRoot, lane ? { lane } : {}).length;
|
|
97
|
+
}
|
|
98
|
+
|
|
99
|
+
/** An api or subscription agent's max_concurrent (1 for a subscription, 2 for api by default); null for local. */
|
|
100
|
+
function agentMaxConcurrent(agentName) {
|
|
101
|
+
try {
|
|
102
|
+
const agent = agentsConfig().agents[agentName];
|
|
103
|
+
return agent && agent.kind !== "local" ? agent.max_concurrent ?? (agent.kind === "subscription" ? 1 : 2) : null;
|
|
104
|
+
} catch { return null; }
|
|
105
|
+
}
|
|
106
|
+
|
|
107
|
+
/**
|
|
108
|
+
* Run a job holding one of its agent's max_concurrent slots, machine-wide
|
|
109
|
+
* (lib/slots.mjs), so `max_concurrent: 1` on a subscription means one job
|
|
110
|
+
* on it across every session -- per-session counting never enforced that,
|
|
111
|
+
* and for subscriptions the count was never checked at all. `waitMs` lets a
|
|
112
|
+
* batch queue for a slot instead of failing.
|
|
113
|
+
*/
|
|
114
|
+
function withAgentSlot(args, jobId, fn, { waitMs = 0 } = {}) {
|
|
115
|
+
const max = args.agentName ? agentMaxConcurrent(args.agentName) : null;
|
|
116
|
+
if (!max) return fn();
|
|
117
|
+
return (async () => {
|
|
118
|
+
const slot = await acquireSlot(slotsRoot, args.agentName, max, { jobId, waitMs });
|
|
119
|
+
if (!slot) throw new Error(`agent "${args.agentName}" is at its max_concurrent (${max}) across every nomArmy session on this machine; try again when one of its jobs finishes`);
|
|
120
|
+
try { return await fn(); } finally { slot.release(); }
|
|
121
|
+
})();
|
|
122
|
+
}
|
|
123
|
+
function track(jobId, meta, promise) {
|
|
124
|
+
const entry = { ...meta, jobId, startedAt: new Date().toISOString(), settled: false, result: null, error: null, promise: null };
|
|
125
|
+
// A machine-wide lease for as long as the job runs, so every session's
|
|
126
|
+
// admission counts it (runningCount); released however the job ends.
|
|
127
|
+
// `repo` lets each session's status line show its own repo's jobs.
|
|
128
|
+
if (meta.lane) writeLease(leasesRoot, jobId, { lane: meta.lane, agent: meta.agent ?? null, runId: meta.runId ?? null, role: meta.role ?? null, model: meta.model ?? null, repo: projectDir });
|
|
129
|
+
const release = () => removeLease(leasesRoot, jobId);
|
|
130
|
+
entry.promise = promise.then(
|
|
131
|
+
r => { entry.settled = true; entry.result = r; release(); notifyJobFinished(entry, r, null); return r; },
|
|
132
|
+
e => { entry.settled = true; entry.error = e; release(); notifyJobFinished(entry, null, e); throw e; });
|
|
133
|
+
entry.promise.catch(() => {});
|
|
134
|
+
activeJobs.set(jobId, entry);
|
|
135
|
+
return entry;
|
|
136
|
+
}
|
|
137
|
+
/**
|
|
138
|
+
* A desktop notification when a job ends (lib/notify.mjs), so the person
|
|
139
|
+
* watching hears about it from any coordinator without polling.
|
|
140
|
+
*/
|
|
141
|
+
function notifyJobFinished(entry, result, error) {
|
|
142
|
+
if (!entry.lane) return; // only tracked jobs, never internal helpers
|
|
143
|
+
const m = result?.manifest ?? {};
|
|
144
|
+
const outcome = error ? "failed" : String(m.outcome ?? (result?.ok ? "done" : "finished")).toLowerCase().replace(/_/g, " ");
|
|
145
|
+
const who = entry.agent ? `${entry.agent}${entry.model ? `/${entry.model}` : ""}` : "local model";
|
|
146
|
+
const took = Math.round((Date.now() - Date.parse(entry.startedAt)) / 60000);
|
|
147
|
+
const ok = !error && (result?.ok || m.coordinatorStatus === "complete");
|
|
148
|
+
notify(`nomArmy: ${entry.role ?? entry.mode ?? "job"} ${ok ? "done" : outcome}`, `${entry.workerId ?? entry.jobId} on ${who}: ${outcome} after ${took}m. ${ok ? "Ready for the General's review." : "Needs a look."}`);
|
|
149
|
+
}
|
|
150
|
+
function capacitySnapshot() {
|
|
151
|
+
const admission = assessAdmission({ hardware: deps.budgetState.hardwareSnapshot, runningJobs: runningCount("local"), slots: deps.budgetState.contextInfo.slots, maxWorkers: currentMaxWorkers() });
|
|
152
|
+
return {
|
|
153
|
+
// The local model's budget. An api or subscription job's scales with
|
|
154
|
+
// its own model; local_worker_start reports that job's.
|
|
155
|
+
budgets: { ...deps.budgetState.budgets, describe: describeBudgets(deps.budgetState.budgets) },
|
|
156
|
+
context: deps.budgetState.contextInfo,
|
|
157
|
+
admission,
|
|
158
|
+
memory: deps.budgetState.hardwareSnapshot?.memory ?? null,
|
|
159
|
+
running: [...activeJobs.values()].filter(j => !j.settled).map(j => ({ jobId: j.jobId, workerId: j.workerId, mode: j.mode, lane: j.lane, startedAt: j.startedAt, phase: readJson(path.join(jobsRoot, j.jobId, "status.json"))?.phase ?? "starting" })),
|
|
160
|
+
maxWorkers: currentMaxWorkers(),
|
|
161
|
+
remote: { running: runningCount("remote"), maxWorkers: currentMaxPoolWorkers(), note: "api and subscription agents; each agent's own max_concurrent also applies" }
|
|
162
|
+
};
|
|
163
|
+
}
|
|
164
|
+
async function admit(jobs) {
|
|
165
|
+
await deps.budgetState.refresh();
|
|
166
|
+
if (jobs.some((j) => jobLane(j) === "remote")) await modelCatalogReady();
|
|
167
|
+
const problems = [];
|
|
168
|
+
// A pool-routed job is checked against that pool's OWN (model-dependent)
|
|
169
|
+
// budget, not the local-derived global one -- see budgetsForPool. Which
|
|
170
|
+
// specific entry pickProvider will land on isn't known yet at admission
|
|
171
|
+
// time, so this is the conservative minimum across the pool's currently
|
|
172
|
+
// available entries, not any one entry's precise number. A
|
|
173
|
+
// subscription_worker job budgets against that one named entry directly
|
|
174
|
+
// (see budgetsForSubscriptionWorker) -- there's no "which entry" unknown
|
|
175
|
+
// the way a weighted pool has, since the name given IS the entry.
|
|
176
|
+
jobs.forEach((j, i) => {
|
|
177
|
+
const jobBudgets = budgetsForJob(j);
|
|
178
|
+
for (const p of checkBrief(j, jobBudgets)) problems.push(jobs.length > 1 ? `job ${i + 1}: ${p}` : p);
|
|
179
|
+
});
|
|
180
|
+
// verify_regression re-runs `verification`; with no profile set there is
|
|
181
|
+
// nothing to re-run. Refuse before starting anything, matching every other
|
|
182
|
+
// admission check here, rather than silently no-op at runtime.
|
|
183
|
+
jobs.forEach((j, i) => {
|
|
184
|
+
if (j.verify_regression && !j.verification) {
|
|
185
|
+
problems.push(`${jobs.length > 1 ? `job ${i + 1}: ` : ""}verify_regression requires a verification profile; there is nothing to run twice without one`);
|
|
186
|
+
}
|
|
187
|
+
});
|
|
188
|
+
// subscription_worker/on_behalf_of: the owner-match attestation refusal
|
|
189
|
+
// happens here, before a container is ever provisioned -- matching how a
|
|
190
|
+
// bad `pool` name is already caught before dispatch, not mid-flight. Only
|
|
191
|
+
// attempted once the plain field-presence problems above are already
|
|
192
|
+
// clean, so a missing on_behalf_of is never reported twice in two
|
|
193
|
+
// different shapes.
|
|
194
|
+
jobs.forEach((j, i) => {
|
|
195
|
+
const fieldProblems = subscriptionJobFieldProblems(j);
|
|
196
|
+
for (const p of fieldProblems) problems.push(jobs.length > 1 ? `job ${i + 1}: ${p}` : p);
|
|
197
|
+
if (fieldProblems.length === 0 && j.on_behalf_of) {
|
|
198
|
+
try {
|
|
199
|
+
if (j.subscription_worker) resolveSubscriptionSelection(j.subscription_worker, j.on_behalf_of, j.reasoning, { model: j.model });
|
|
200
|
+
} catch (error) { problems.push(jobs.length > 1 ? `job ${i + 1}: ${error.message}` : error.message); }
|
|
201
|
+
}
|
|
202
|
+
});
|
|
203
|
+
// The repo's own policy: verification required, revert check required.
|
|
204
|
+
const policy = repoPolicy();
|
|
205
|
+
jobs.forEach((j, i) => { for (const p of policyAdmissionProblems(j, policy)) problems.push(`${jobs.length > 1 ? `job ${i + 1}: ` : ""}${p}`); });
|
|
206
|
+
// An implement job on an agent whose own tools run on this machine (the
|
|
207
|
+
// Claude CLI) isn't bounded by the sandbox, so it's refused unless that
|
|
208
|
+
// agent says allow_host_tools (lib/agents.mjs). Scouts and reviews still run.
|
|
209
|
+
jobs.forEach((j, i) => {
|
|
210
|
+
if (!j.agentName || (j.mode ?? "implement") !== "implement") return;
|
|
211
|
+
let problem = null;
|
|
212
|
+
try { problem = hostToolsImplementProblem(j.agentName, agentsConfig().agents[j.agentName]); } catch { return; }
|
|
213
|
+
if (problem) problems.push(`${jobs.length > 1 ? `job ${i + 1}: ` : ""}${problem}`);
|
|
214
|
+
});
|
|
215
|
+
// A model its vendor refused on a job today, with nothing working on it
|
|
216
|
+
// since, isn't sent another job (lib/health.mjs recentModelRefusal).
|
|
217
|
+
jobs.forEach((j, i) => {
|
|
218
|
+
if (!j.agentName || !j.model) return;
|
|
219
|
+
let provider = null;
|
|
220
|
+
try { provider = agentProviderId(agentsConfig().agents[j.agentName]); } catch { return; }
|
|
221
|
+
if (!provider) return;
|
|
222
|
+
const refusal = recentModelRefusal(stateRoot, `${provider}/${j.model}`);
|
|
223
|
+
if (refusal) problems.push(`${jobs.length > 1 ? `job ${i + 1}: ` : ""}model_not_found: ${provider}/${j.model} was refused on an earlier job today and hasn't worked since, so this job wasn't sent. Use another model (the job's \`model\`, or \`nomarmy army assign\`); \`nomarmy army assign <role> ${j.agentName} ${j.model}\` re-tests it, and a passing test clears this.`);
|
|
224
|
+
});
|
|
225
|
+
// An agent's max_concurrent, machine-wide. Batch jobs on the same agent
|
|
226
|
+
// queue for its slot at launch instead (withAgentSlot's waitMs).
|
|
227
|
+
if (jobs.length === 1) {
|
|
228
|
+
const [j] = jobs;
|
|
229
|
+
const max = j.agentName ? agentMaxConcurrent(j.agentName) : null;
|
|
230
|
+
const held = max ? liveSlots(slotsRoot, j.agentName) : 0;
|
|
231
|
+
if (max && held >= max) problems.push(`not admitted (capacity): agent "${j.agentName}" already has ${held} job(s) running across this machine's nomArmy sessions, at its max_concurrent of ${max}`);
|
|
232
|
+
}
|
|
233
|
+
// A job in a /feature run: the run's own limits and paused agents.
|
|
234
|
+
jobs.forEach((j, i) => {
|
|
235
|
+
if (!j.run_id) return;
|
|
236
|
+
try {
|
|
237
|
+
const run = loadRun(runsRoot, j.run_id);
|
|
238
|
+
if (run.repo !== projectDir) problems.push(`${jobs.length > 1 ? `job ${i + 1}: ` : ""}run "${run.id}" belongs to ${run.repo}, not this repository`);
|
|
239
|
+
const running = liveLeases(leasesRoot, { runId: run.id }).length + jobs.slice(0, i).filter((o) => o.run_id === run.id).length;
|
|
240
|
+
for (const p of runAdmissionProblems(run, { agentName: j.agentName ?? "local", running })) problems.push(jobs.length > 1 ? `job ${i + 1}: ${p}` : p);
|
|
241
|
+
} catch (error) { problems.push(jobs.length > 1 ? `job ${i + 1}: ${error.message}` : error.message); }
|
|
242
|
+
});
|
|
243
|
+
// Slot capacity only concerns local jobs: a remote job's inference runs
|
|
244
|
+
// at its vendor and never competes for llama-server's slots. Free memory
|
|
245
|
+
// still applies to every job (each one runs a local sandbox), so a
|
|
246
|
+
// remote-only batch is checked for memory alone. Remote jobs have their
|
|
247
|
+
// own, additive ceiling (currentMaxPoolWorkers).
|
|
248
|
+
const anyLocal = jobs.some((j) => jobLane(j) === "local");
|
|
249
|
+
const admission = anyLocal
|
|
250
|
+
? assessAdmission({ hardware: deps.budgetState.hardwareSnapshot, runningJobs: runningCount("local"), slots: deps.budgetState.contextInfo.slots, maxWorkers: currentMaxWorkers() })
|
|
251
|
+
: assessAdmission({ hardware: deps.budgetState.hardwareSnapshot, runningJobs: 0, slots: null, maxWorkers: Infinity });
|
|
252
|
+
if (!admission.admit) problems.push(...admission.reasons.map(r => `not admitted (${admission.level}): ${r}`));
|
|
253
|
+
if (jobs.some((j) => jobLane(j) === "remote")) {
|
|
254
|
+
const remoteCeiling = currentMaxPoolWorkers(), runningRemote = runningCount("remote");
|
|
255
|
+
if (runningRemote >= remoteCeiling) {
|
|
256
|
+
problems.push(`not admitted (capacity): ${runningRemote} remote job(s) (api or subscription agents) already running, at NOMARMY_MAX_POOL_WORKERS=${remoteCeiling}`);
|
|
257
|
+
}
|
|
258
|
+
}
|
|
259
|
+
return { problems, admission };
|
|
260
|
+
}
|
|
261
|
+
|
|
262
|
+
function refusal(problems) {
|
|
263
|
+
return toolText(refusalText(problems, capacitySnapshot), true);
|
|
264
|
+
}
|
|
265
|
+
/** A run's totals and warnings, for a tool response. */
|
|
266
|
+
function runBrief(runId) {
|
|
267
|
+
try {
|
|
268
|
+
const run = loadRun(runsRoot, runId);
|
|
269
|
+
const totals = runTotals(run);
|
|
270
|
+
return { id: run.id, status: run.status, limits: run.limits, used: totals.used, warnings: totals.warnings };
|
|
271
|
+
} catch (error) { return { id: runId, error: error.message }; }
|
|
272
|
+
}
|
|
273
|
+
|
|
274
|
+
/**
|
|
275
|
+
* Record a finished job into its run. A usage-limit message is looked for
|
|
276
|
+
* only in error text (OpenClaw's failure envelope, and the error lines of
|
|
277
|
+
* a thrown run), never in the worker's report or tool output, where "rate
|
|
278
|
+
* limit" may just be the code under review.
|
|
279
|
+
*/
|
|
280
|
+
function recordJobInRun(args, jobId, result, error = null) {
|
|
281
|
+
if (!args.run_id) return;
|
|
282
|
+
const kind = args.pool ? "api" : args.subscription_worker ? "subscription" : "local";
|
|
283
|
+
const m = result?.manifest ?? {};
|
|
284
|
+
const errorLines = [m.worker?.error, error?.message,
|
|
285
|
+
...String(m.workerError ?? "").split(/\r?\n/).filter((l) => /error|limit|429/i.test(l))].filter(Boolean).join("\n");
|
|
286
|
+
const usageLimit = kind === "local" ? null : detectUsageLimit(errorLines);
|
|
287
|
+
try {
|
|
288
|
+
const before = runTotals(loadRun(runsRoot, args.run_id)).warnings;
|
|
289
|
+
const updated = recordRunJob(runsRoot, args.run_id, {
|
|
290
|
+
jobId, agent: args.agentName ?? "local", kind, model: args.model ?? null, role: args.armyRole ?? null, mode: args.mode,
|
|
291
|
+
outcome: m.outcome ?? (error ? "ERROR" : null), costUsd: m.metrics?.worker_cost_usd ?? null,
|
|
292
|
+
tokens: m.metrics?.worker_tokens_total ?? null, usageLimit,
|
|
293
|
+
});
|
|
294
|
+
// A limit crossed or an agent paused by this job is worth interrupting for.
|
|
295
|
+
const fresh = runTotals(updated).warnings.filter((w) => !before.includes(w) && /OVER|paused/.test(w));
|
|
296
|
+
if (fresh.length) notify(`nomArmy run ${updated.name}: stopped short`, fresh.join("; "));
|
|
297
|
+
} catch (recordError) {
|
|
298
|
+
fs.appendFileSync(path.join(jobsRoot, jobId, "coordinator.log"), `${new Date().toISOString()} could not record into run ${args.run_id}: ${recordError.message}\n`);
|
|
299
|
+
}
|
|
300
|
+
}
|
|
301
|
+
function trackInRun(args, entry) {
|
|
302
|
+
if (args.run_id) entry.promise.then((r) => recordJobInRun(args, entry.jobId, r), (e) => recordJobInRun(args, entry.jobId, null, e));
|
|
303
|
+
return entry;
|
|
304
|
+
}
|
|
305
|
+
function launch(args) {
|
|
306
|
+
const workerId = args.worker_id || null;
|
|
307
|
+
const jobId = slug(workerId || (args.mode === "scout" ? "scout" : "worker"));
|
|
308
|
+
return trackInRun(args, track(jobId, { mode: args.mode, workerId: workerId || jobId, lane: jobLane(args), agent: args.agentName ?? null, runId: args.run_id ?? null, role: args.armyRole ?? null, model: args.model ?? null },
|
|
309
|
+
withAgentSlot(args, jobId, () => executeJob({ ...jobArgs(args, workerId), jobId }))));
|
|
310
|
+
}
|
|
311
|
+
// Best-effort progress signal for a job still mid-run: a plain "phase: worker,
|
|
312
|
+
// elapsed: Ns" told a caller nothing about whether the worker was still
|
|
313
|
+
// reading or already editing, short of running `git status` on the worktree
|
|
314
|
+
// by hand. Both lookups here are read-only and disposable -- a job's worktree
|
|
315
|
+
// mid-write or a transcript sqlite file mid-append can legitimately fail to
|
|
316
|
+
// read, and that must never fail the status call, only omit the field.
|
|
317
|
+
async function liveProgress(jobDir) {
|
|
318
|
+
const out = {};
|
|
319
|
+
try {
|
|
320
|
+
const worktree = path.join(jobDir, "worktree");
|
|
321
|
+
if (fs.existsSync(worktree)) {
|
|
322
|
+
// --untracked-files=normal, not all: "all" descends into every
|
|
323
|
+
// untracked directory (a virtualenv, a cache) a job creates. Bounded:
|
|
324
|
+
// a live progress read must never hold anything up.
|
|
325
|
+
const statusOut = (await run("git", ["status", "--porcelain=v1", "-z", "--untracked-files=normal"], { cwd: worktree, trim: false, timeoutMs: 10000 })).stdout;
|
|
326
|
+
// Same runtime-junk filter as collectGitRecord/makeIdleDiffTick: .npm/
|
|
327
|
+
// etc. is the sandbox's own churn, not the worker's progress, and
|
|
328
|
+
// counting it made a job that had made zero real edits report
|
|
329
|
+
// filesChangedLive: 1 anyway.
|
|
330
|
+
out.filesChangedLive = parseStatusPorcelainZ(statusOut).map(e => e.file).filter(f => !isRuntimeJunk(f)).length;
|
|
331
|
+
}
|
|
332
|
+
} catch { /* worktree not ready yet, or mutated mid-read; omit */ }
|
|
333
|
+
try {
|
|
334
|
+
const stateDir = path.join(jobDir, "runtime", "state");
|
|
335
|
+
const transcript = await readOpenClawTranscriptTail(stateDir, { limit: 6 });
|
|
336
|
+
if (transcript.available) {
|
|
337
|
+
const last = transcript.toolCalls.at(-1);
|
|
338
|
+
if (last) out.lastTool = { tool: last.tool, target: last.path ?? last.command ?? null };
|
|
339
|
+
}
|
|
340
|
+
// A claude-cli worker's tools only appear in Claude Code's own session
|
|
341
|
+
// transcript, not OpenClaw's.
|
|
342
|
+
if (!out.lastTool) {
|
|
343
|
+
const startedMs = Date.parse(readJson(path.join(jobDir, "status.json"))?.startedAt ?? "") || 0;
|
|
344
|
+
const claude = readClaudeSessionTranscript(path.join(jobDir, "worktree"), { sinceMs: startedMs, tailBytes: 262144 });
|
|
345
|
+
const last = claude.available ? claude.toolCalls.at(-1) : null;
|
|
346
|
+
if (last) { out.lastTool = { tool: last.tool, target: last.path ?? last.command ?? null }; out.toolCallsLive = claude.toolCalls.length; }
|
|
347
|
+
}
|
|
348
|
+
} catch { /* transcript not created yet, or locked mid-write; omit */ }
|
|
349
|
+
return out;
|
|
350
|
+
}
|
|
351
|
+
|
|
352
|
+
async function summarize(entry, files, jobDir = null) {
|
|
353
|
+
const status = files.status, meta = files.meta ?? files.failure;
|
|
354
|
+
const elapsedSeconds = status?.startedAt ? Math.round((Date.now() - Date.parse(status.startedAt)) / 1000) : entry ? Math.round((Date.now() - Date.parse(entry.startedAt)) / 1000) : null;
|
|
355
|
+
const out = { jobId: entry?.jobId ?? status?.jobId ?? meta?.jobId ?? null, workerId: entry?.workerId ?? status?.workerId ?? meta?.workerId ?? null,
|
|
356
|
+
mode: entry?.mode ?? status?.mode ?? meta?.mode ?? null, state: null, phase: status?.phase ?? "starting", elapsedSeconds,
|
|
357
|
+
timeoutSeconds: status?.timeoutSeconds ?? null, coordinatorStatus: meta?.coordinatorStatus ?? null, outcome: meta?.outcome ?? null,
|
|
358
|
+
reviewRequired: meta?.reviewRequired ?? null, issues: (meta?.issues ?? []).slice(0, 6), worktree: meta?.worktree ?? null, branch: meta?.branch ?? null,
|
|
359
|
+
commit: meta?.commit?.sha ?? null, scout: meta?.scout ? { supported: meta.scout.supported, unsupported: meta.scout.unsupported } : null };
|
|
360
|
+
if (entry && !entry.settled) out.state = "running";
|
|
361
|
+
else if (entry?.error) { out.state = "failed"; out.error = String(entry.error.message ?? entry.error).split("\n")[0]; }
|
|
362
|
+
else if (meta) out.state = "finished";
|
|
363
|
+
else if (status?.state === "running") { out.state = status.serverPid === process.pid ? "running" : "orphaned"; if (out.state === "orphaned") out.error = `the MCP server that ran this job (pid ${status.serverPid}) is gone; outcome unknown, see the job directory logs`; }
|
|
364
|
+
else out.state = "unknown";
|
|
365
|
+
if (out.state === "running" && jobDir) Object.assign(out, await liveProgress(jobDir));
|
|
366
|
+
return out;
|
|
367
|
+
}
|
|
368
|
+
|
|
369
|
+
return { WORKER_START_STAGGER_MS, activeJobs, runningCount, agentMaxConcurrent, withAgentSlot, track, notifyJobFinished, capacitySnapshot, admit, refusal, runBrief, recordJobInRun, trackInRun, launch, liveProgress, summarize };
|
|
370
|
+
}
|