nomarmy 0.1.0-alpha.6 → 0.1.0-alpha.8

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -6,20 +6,22 @@
6
6
  <a href="https://github.com/rayson-tech/nomarmy/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-Apache%202.0-blue.svg" alt="License: Apache 2.0"></a>
7
7
  </p>
8
8
 
9
- <p align="center"><em>Tiny coders, big appetites for bounded tickets.</em> 🍪</p>
9
+ <p align="center"><em>Every byte verified.</em> 🍪</p>
10
10
 
11
11
  **Your coding assistant plans; sandboxed workers build; nothing counts until nomArmy has checked it.**
12
12
 
13
+ AI coding workers are confident. Their "done, all tests pass" is a claim, not evidence. nomArmy lets your coding assistant (Claude Code, Codex or Cursor) hand work to workers called **noms**, then checks every change itself before anything is committed: the real diff, your tests run in a fresh sandbox, a check that those tests actually catch the change, and a secret scan.
14
+
13
15
  ## TL;DR
14
16
 
15
- 1. **Have** Git, Node 20+ and [Podman](https://podman.io) (on macOS: `brew install podman && podman machine init --memory 8192 && podman machine start`; Podman's 2 GiB default is too small for nomArmy's sandboxes).
17
+ 1. **Have** Git, Node 24.16+ (or 26.1+) and [Podman](https://podman.io). On macOS, give Podman 8 GiB: `brew install podman && podman machine init --memory 8192 && podman machine start`.
16
18
  2. **Install and set up:**
17
19
  ```bash
18
20
  npm install -g nomarmy@alpha
19
21
  cd your-project
20
22
  nomarmy setup
21
23
  ```
22
- `nomarmy setup` is the playbook. It shows a checklist and runs the next step each time you say yes:
24
+ `nomarmy setup` is a playbook. It shows a checklist and runs the next step each time you say yes:
23
25
  ```text
24
26
  ✓ Where models run: hosted
25
27
  ✓ Installed: OpenClaw 2026.9.6
@@ -29,34 +31,39 @@
29
31
  Check: verify the installation
30
32
  Run `nomarmy agents add` now? [Y/n]
31
33
  ```
32
- In order: pick where models run (API keys and subscriptions for most people), install OpenClaw and the sandbox, add your agents (an API key, or your ChatGPT or Muse Code subscription), put the roles on them, write this repo's `.nomarmy.yml`, then check it all. Stop anytime; `nomarmy setup` picks up where you left off.
33
-
34
- **Want every step spelled out?** [Example setup: Claude Code, Codex and an API key](https://github.com/rayson-tech/nomarmy/blob/main/docs/setup/example.md) walks through a complete setup, command by command.
34
+ Stop anytime; `nomarmy setup` picks up where you left off. **Want every step spelled out?** [Example setup: Claude Code, Codex and an API key](https://github.com/rayson-tech/nomarmy/blob/main/docs/setup/example.md) goes command by command.
35
35
  3. **Use it:** restart Claude Code in the project and ask it to use nomArmy for one small bug that has a test. When that works, try `/feature <what you want built>`.
36
36
 
37
- **Have a GPU or a Mac with plenty of memory?** Choose "a local model" in `nomarmy setup` and workers run on llama.cpp on your own machine: no per-token bill and your code stays home, but you pay in hardware, power and speed, and a model too big for your memory crawls. `nomarmy sizing` tells you what fits; see [Install](https://github.com/rayson-tech/nomarmy/blob/main/docs/install.md#install). A team GPU server works too: [a shared model server](https://github.com/rayson-tech/nomarmy/blob/main/docs/install.md#a-shared-model-server). Codex or Cursor as the coordinator: `nomarmy connect codex cursor`.
37
+ Stuck? `nomarmy doctor` checks the machine and `nomarmy health` checks everything nomArmy runs on. Upgrading later? `nomarmy update`.
38
38
 
39
- Stuck? `nomarmy doctor` checks the machine, and `nomarmy health` checks everything nomArmy runs on.
39
+ ## How every byte gets verified
40
40
 
41
- ## What it is
41
+ 1. Your coding assistant, the **General**, briefs a job: a task, acceptance criteria, and the tests that prove it.
42
+ 2. nomArmy creates a git worktree from your branch and runs the nom in a Podman sandbox with no network and no host credentials. (One exception, the Claude subscription: see [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md#security-posture).)
43
+ 3. The nom edits, runs tests, and ends with a four-line report: `STATUS`, `TESTS`, `NOT_DONE`, `NOTE`.
44
+ 4. nomArmy treats that report as a claim and checks the evidence itself:
45
+ - reads the real diff from git, not the nom's description of it
46
+ - runs your verification profile in a fresh sandbox
47
+ - reverts the production change and reruns the tests: a test that still passes proves nothing, so the job goes to review instead of being committed
48
+ - blocks on secrets, and flags tests made to pass (new skips, stubbed imports) and code nothing calls
49
+ 5. Only then does it commit, on the nom's own branch. It never merges into yours: reviewing and integrating stay with the General, and with you.
42
50
 
43
- Your coding assistant (Claude Code, Codex or Cursor) stays in charge as the **General**: it decides what gets built and whether the result is acceptable. The work goes to **noms**, workers that implement, test and repair in their own git worktree and sandbox, on an API key, your own ChatGPT or Muse Code subscription, or a local model. nomArmy owns everything in between: worktrees, git, sandboxes, verification, and the evidence that decides whether work is accepted.
51
+ Failing verification stays failed, unconditionally. A malformed report isn't automatically a failure: if the repository changed, nomArmy verifies independently and may recover the work. And the checks aren't the General's to waive: a repo's `.nomarmy.yml` policy (on by default for new repos) makes verification and the revert check mandatory for every job.
44
52
 
45
- **What you get is work you don't have to take on faith**, not cheaper work. Delegating costs the General tokens too: briefing and reviewing. On small, already-diagnosed tickets we measured 4 to 8 times more of the General's tokens than fixing the bug directly, and break-even at roughly 150 lines of context a fix needs to read ([the measurements](https://github.com/rayson-tech/nomarmy/blob/main/docs/experiments/2026-09-20-model-bakeoff-and-economics.md)). It pays off on bigger tickets, on parallel work, and anywhere you'd otherwise have to trust an agent's say-so.
53
+ **Checking without building** costs nothing: `mode: verify` runs a verification profile against any branch, with no worker and no model tokens.
46
54
 
47
- Developed and maintained by Rayson Technologies. This is an alpha (`0.1.0-alpha`).
55
+ ## Where the work runs
48
56
 
49
- ## How it works
57
+ - **Agents** say where a job can run: an API key, your own ChatGPT or Muse Code subscription, or a local model on llama.cpp.
58
+ - **The army** says which role runs on which agent: Sr and Jr devs build, a security analyst and a data architect review, a PM checks the plan, a PO accepts.
59
+ - **`/feature`** runs a whole feature end to end, from plan through build, review and acceptance, and hands you a branch to merge.
60
+ - **Harnesses** give each repo the right sandbox: Go, Rust, Python and Node (mixed repos too), Playwright browser tests, and fake services like a mock login server, all offline. [Adding one](https://github.com/rayson-tech/nomarmy/blob/main/CONTRIBUTING.md#adding-a-harness) never touches core code.
50
61
 
51
- 1. The General briefs a job: a task, acceptance criteria, the tests that prove it.
52
- 2. nomArmy creates a worktree from your branch and runs the worker in a Podman sandbox with no network and no host credentials. (One exception, the Claude subscription: see [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md#security-posture).)
53
- 3. The worker edits, runs tests, and ends with a four-line report: `STATUS`, `TESTS`, `NOT_DONE`, `NOTE`.
54
- 4. nomArmy treats that report as a claim. It reads the real diff from git, runs your verification profile itself in a fresh sandbox, reverts the production change to check the tests actually fail without it, and scans for secrets.
55
- 5. Only then does it commit, on the worker's own branch. It never merges into yours: reviewing and integrating stay with the General, and with you.
62
+ **What you get is work you don't have to take on faith**, not cheaper work. Delegating costs the General tokens too, for briefing and review: on small, already-diagnosed tickets we measured 4 to 8 times more of the General's tokens than fixing the bug directly, with break-even around 150 lines of context a fix needs to read ([the measurements](https://github.com/rayson-tech/nomarmy/blob/main/docs/experiments/2026-09-20-model-bakeoff-and-economics.md)). It pays off on bigger tickets, parallel work, and anywhere you'd otherwise trust an agent's say-so.
56
63
 
57
- A malformed report isn't automatically a failure: if the repository changed, nomArmy verifies independently and may recover the work. Failing verification stays failed, unconditionally. And the checks aren't the General's to waive: a repo's `.nomarmy.yml` policy (on by default for new repos) makes verification and the revert check mandatory for every job.
64
+ **Have a GPU or a Mac with plenty of memory?** Choose "a local model" in `nomarmy setup`: no per-token bill and your code stays home, but you pay in hardware, power and speed. `nomarmy sizing` tells you what fits. A [shared model server](https://github.com/rayson-tech/nomarmy/blob/main/docs/install.md#a-shared-model-server) works too.
58
65
 
59
- Around that core: **agents** say where a job can run, the **army** says which role runs on which agent, and **`/feature`** runs a whole feature end to end, from plan through build, review and acceptance, handing you a branch to merge.
66
+ Developed and maintained by Rayson Technologies. This is an alpha (`0.1.0-alpha`).
60
67
 
61
68
  ## Docs
62
69
 
@@ -66,7 +73,7 @@ Around that core: **agents** say where a job can run, the **army** says which ro
66
73
  | [Example setup](https://github.com/rayson-tech/nomarmy/blob/main/docs/setup/example.md) | Claude Code, Codex and an API key, command by command |
67
74
  | [Agents and the army](https://github.com/rayson-tech/nomarmy/blob/main/docs/agents-and-army.md) | Where a job can run, who does what, usage limits, picking an agent |
68
75
  | [`/feature` runs](https://github.com/rayson-tech/nomarmy/blob/main/docs/feature-runs.md) | A feature end to end, and watching what nomArmy is doing |
69
- | [Your repository](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md) | `.nomarmy.yml`, verification, languages and dependencies, what nomArmy checks |
76
+ | [Your repository](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md) | `.nomarmy.yml`, verification, dependencies, private registries, what nomArmy checks |
70
77
  | [Harnesses](https://github.com/rayson-tech/nomarmy/blob/main/docs/harnesses.md) | Ecosystem registry, detection, network levels, and requirements |
71
78
  | [Configuration](https://github.com/rayson-tech/nomarmy/blob/main/docs/configuration.md) | Settings, swapping the local model, sizing, admission |
72
79
  | [Reference](https://github.com/rayson-tech/nomarmy/blob/main/docs/reference.md) | Every CLI command and MCP tool |
@@ -75,23 +82,20 @@ Around that core: **agents** say where a job can run, the **army** says which ro
75
82
 
76
83
  ## Security
77
84
 
78
- A worker gets a writable git worktree inside a Podman sandbox and nothing else: no network, no host credentials, no Podman socket. Every model call is made by OpenClaw on your machine, never from inside the sandbox. **The exception is a Claude subscription**, whose tools run on your machine, so nomArmy refuses build jobs on it unless you allow it. Never hand a worker production credentials, deployment access or SSH keys. Details: [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md); to report a vulnerability, [SECURITY.md](https://github.com/rayson-tech/nomarmy/blob/main/SECURITY.md).
85
+ A nom gets a writable git worktree inside a Podman sandbox and nothing else: no network, no host credentials, no Podman socket. Every model call is made by OpenClaw on your machine, never from inside the sandbox. Verification can climb a network ladder one rung at a time (fake services on a private network, then an allowlist you approve for a test tenant), but noms never leave `network none`. **The exception is a Claude subscription**, whose tools run on your machine, so nomArmy refuses build jobs on it unless you allow it. Never hand a nom production credentials, deployment access or SSH keys. Details: [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md); to report a vulnerability, [SECURITY.md](https://github.com/rayson-tech/nomarmy/blob/main/SECURITY.md).
79
86
 
80
87
  ## Status
81
88
 
82
89
  | Capability | Status |
83
90
  |---|---|
84
- | Delegation core: worktrees, nomArmy-owned git, independent verification, kept failed worktrees | Working, end-to-end tested |
85
- | Local (llama.cpp) and Bedrock profiles | Working |
86
- | Scout and decompose modes | Unit and live tested |
87
- | `auto_union`, `verify_regression`, test-selection and unwired-code checks | Unit and live tested; the heuristics are review flags |
88
- | Secret scanning (secretlint, hard block) | Unit tested against the real dependency |
89
- | Agents: api keys | Live-verified with xAI; other providers built to OpenClaw's documented interface |
91
+ | Verification core: worktrees, nomArmy-owned git, independent verification, the revert check, kept failed worktrees | Working, end-to-end tested |
92
+ | `mode: verify`, secret scanning (secretlint, hard block), test-workaround and unwired-code checks | Unit and live tested; the heuristics are review flags |
93
+ | Scout and decompose modes, `auto_union` | Unit and live tested |
94
+ | Agents: API keys | Live-verified with xAI; other providers built to OpenClaw's documented interface |
90
95
  | Agents: subscriptions | ChatGPT (Codex) and Muse Code sandboxed and live-verified; Claude live-verified, but its tools run on the host (scout and review by default) |
96
+ | Local (llama.cpp) and Bedrock profiles | Working |
91
97
  | The army and `/feature` | Driven by a real Claude Code General across three runs, about 18 implement jobs |
92
- | Go, Rust, Python and Node repos, and mixed ones | Harness images live-verified: Go modules and Rust crates prefetched; npm, pnpm, yarn, bun and workspaces; pip, pyproject, uv and poetry |
93
- | Fake services beside the app (mock login server, mock APIs) | Live-verified on a private network with no route out (the `services` harness level) |
94
- | Browser tests (Playwright + Chromium) | Live-verified offline, with screenshots and traces kept as job evidence |
98
+ | Harnesses: Go, Rust, Python, Node and mixed repos; Playwright; fake services | Live-verified offline |
95
99
  | Private registries and a verification-only network allowlist | Live-verified; each passed an independent security review |
96
100
 
97
101
  What we've learned from real runs, including where delegating pays and where it doesn't, is in [docs/findings.md](https://github.com/rayson-tech/nomarmy/blob/main/docs/findings.md).
@@ -99,13 +103,13 @@ What we've learned from real runs, including where delegating pays and where it
99
103
  ### Known limitations
100
104
 
101
105
  - **A Claude subscription isn't sandboxed.** Its tools run on your machine, so implement jobs on it are refused unless you set `allow_host_tools: true`. See [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md#security-posture).
102
- - **A refused model costs one job.** When a vendor refuses a model at run time that OpenClaw lists (gpt-6-sol on a ChatGPT plan), the first job on it fails with `model_not_found`. After that nomArmy refuses to dispatch it until a job or test call on it works. `army assign` tests the job's route and catches this before any job.
103
- - **Claude subscription token counts** come from the Claude CLI's own session log, since OpenClaw sees only the final reply. Totals include cache reads and writes, which make up most of an agent's prompt; each part is also kept separately.
106
+ - **Your own compose services aren't started yet.** A verification profile that needs a real database from your compose file (`environment: basic` or higher) reports `not_run` rather than running without it (and the compose parser doesn't resolve YAML anchors). Fake services from harnesses, like the mock login server, do run.
107
+ - **Private registries don't cover Poetry or Yarn Berry** yet: Poetry can't guarantee a credentialed install runs no package code, and Yarn Berry doesn't read `.npmrc`. uv, pip wheels, npm, pnpm, Yarn Classic and bun work. See [Private registries](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md#private-registries).
108
+ - **A refused model costs one job.** When a vendor refuses a model at run time that OpenClaw lists (gpt-6-sol on a ChatGPT plan), the first job on it fails with `model_not_found`; after that nomArmy won't dispatch it until a job or test call on it works. `army assign` tests the route and catches this before any job.
109
+ - **Claude subscription token counts** come from the Claude CLI's own session log, since OpenClaw sees only the final reply; totals include cache reads and writes.
104
110
  - **Test-workaround detection is a flag, not a verdict**: a legitimate new skip still gets flagged.
105
111
  - **Deploy-time failures need your own check.** See [Add a check for what unit tests can't see](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md#nomarmyyml).
106
- - **Node private registries are not supported yet**: npm, pnpm, yarn and bun lockfiles and workspaces are supported, but the image build has no credentials for private registries (dependency plan step 8).
107
- - **Verification needing services** (a database, a mock server) reports `not_run` instead of running without them. The compose parser doesn't resolve YAML anchors.
108
- - **Same-host sandboxes**: the MCP server, OpenClaw and every job's sandbox run on the machine with the coordinator. Only the model can be elsewhere (an agent, or [a shared model server](https://github.com/rayson-tech/nomarmy/blob/main/docs/install.md#a-shared-model-server)).
112
+ - **Same-host sandboxes**: the MCP server, OpenClaw and every job's sandbox run on the machine with the coordinator. Only the model can be elsewhere.
109
113
 
110
114
  ## More
111
115
 
package/bin/nomarmy.mjs CHANGED
@@ -18,7 +18,8 @@ import { buildConfigProposal } from "../lib/propose.mjs";
18
18
  import { detectHardware } from "../lib/hardware.mjs";
19
19
  import { readGGUFMetadata, resolveModelPath, totalSplitBytes } from "../lib/gguf.mjs";
20
20
  import { recommend, customRecommendation, evaluateConfig, bytesPerKvElementForCacheTypes, MIN_CONTEXT_PER_NOM } from "../lib/sizing.mjs";
21
- import { connectClaude, connectCodex, connectCursor, cursorAlreadyConnected, deriveWorkerModelEnv } from "../lib/connect.mjs";
21
+ import { connectClaude, connectCodex, connectCursor, cursorAlreadyConnected, deriveWorkerModelEnv, defaultInstallDir } from "../lib/connect.mjs";
22
+ import { compareVersions, readPackageVersion, readInstallVersions } from "../lib/install-freshness.mjs";
22
23
  import { ID_RE, AUTH_ENV_NAME_RE, OPENCLAW_PROVIDER_ID_RE, openclawProviderId, isNativeProviderType } from "../lib/dispatch-schema.mjs";
23
24
  import { loadAgents, readAgentsFile, writeAgentsFile, agentsConfigPath, apiAgentAsPoolEntry, describeAgent as describeAgentLabel, agentRunsToolsOnHost, agentProviderId, AGENT_KINDS, API_PROVIDER_TYPES, RESERVED_AGENT_NAMES, BUILTIN_LOCAL_AGENT } from "../lib/agents.mjs";
24
25
  import { loadArmy, mergeArmy, describeArmy, readArmyFile, updateArmyInFile, assignRoleInFile, parseTargetSpec, armyLayerPath, globalConfigDir, DEFAULT_ARMY, ARMY_PHASES, LOCAL_CONFIG_FILENAME } from "../lib/army.mjs";
@@ -115,8 +116,9 @@ Usage: nomarmy <command> [options]
115
116
  registration (never done silently)
116
117
  --restart-inference with --json, also stop/start local
117
118
  inference (never done silently)
118
- update Pull the latest nomArmy code and re-sync the installed
119
- MCP copy (fast-forward only; refuses on local changes).
119
+ update Update nomArmy and reconnect your coordinators: installs
120
+ npm's latest alpha, or for a git checkout pulls (fast-forward
121
+ only; refuses on local changes). Then restart open sessions.
120
122
  agents <list|add|update|remove>
121
123
  Every account a job can run on, in one list:
122
124
  ~/.config/nomarmy/agents.yml (or NOMARMY_CONFIG_DIR).
@@ -1642,14 +1644,9 @@ function git(args) {
1642
1644
  */
1643
1645
  async function cmdUpdate() {
1644
1646
  const say = (s) => { if (!json) console.log(s); };
1645
- // Installed from npm: there's no checkout to pull. npm updates the
1646
- // package; connect resyncs the copy each coordinator runs.
1647
- if (!fs.existsSync(path.join(nomarmyRoot, ".git"))) {
1648
- const how = "npm install -g nomarmy@alpha && nomarmy connect";
1649
- if (json) return out({ error: "installed from npm, not a git checkout", fix: how });
1650
- console.log(`This nomArmy was installed from npm, so there's nothing to pull. Update with:\n ${how}`);
1651
- return;
1652
- }
1647
+ // Installed from npm: npm updates the package, then connect resyncs the
1648
+ // copy each coordinator runs.
1649
+ if (!fs.existsSync(path.join(nomarmyRoot, ".git"))) return updateFromNpm();
1653
1650
  const status = git(["status", "--porcelain"]);
1654
1651
  if (status) {
1655
1652
  if (json) { out({ error: "working tree is not clean; refusing to pull over local changes", status }); process.exit(1); }
@@ -1704,6 +1701,45 @@ async function cmdUpdate() {
1704
1701
  console.log(c.yellow("\nThe MCP server is a per-session child process: every open Claude Code / Codex / Cursor session needs a restart to pick this up, not just this one."));
1705
1702
  }
1706
1703
 
1704
+ // The coordinators nomArmy is registered with. Cursor has no CLI to probe,
1705
+ // so it counts when its own config already lists nomArmy.
1706
+ function connectedTargets() {
1707
+ return [commandExists("claude") && "claude", commandExists("codex") && "codex", cursorAlreadyConnected() && "cursor"].filter(Boolean);
1708
+ }
1709
+
1710
+ async function updateFromNpm() {
1711
+ const current = readPackageVersion(nomarmyRoot);
1712
+ let latest = null;
1713
+ try { latest = execFileSync("npm", ["view", "nomarmy", "dist-tags.alpha"], { encoding: "utf8", timeout: 20000 }).trim(); } catch { /* offline */ }
1714
+ if (!latest) {
1715
+ const fix = "npm install -g nomarmy@alpha && nomarmy connect claude";
1716
+ if (json) { out({ error: "could not read nomarmy's latest release from npm", fix }); process.exit(1); }
1717
+ console.log(c.red("Couldn't reach npm to find nomArmy's latest release.") + ` Update by hand:\n ${fix}`);
1718
+ process.exit(1);
1719
+ }
1720
+ const { copyVersion } = readInstallVersions(defaultInstallDir());
1721
+ const upgrade = compareVersions(current, latest) < 0;
1722
+ const staleCopy = !copyVersion || compareVersions(copyVersion, upgrade ? latest : current) < 0;
1723
+ if (!upgrade && !staleCopy) {
1724
+ if (json) return out({ updated: false, version: current, reason: "already up to date" });
1725
+ console.log(c.green(`✓ nomArmy ${current} is the latest, and your coordinators run it.`));
1726
+ return;
1727
+ }
1728
+ if (upgrade) {
1729
+ if (!json) console.log(c.bold(`🍪 Updating nomArmy ${current} → ${latest}\n`));
1730
+ execFileSync("npm", ["install", "-g", `nomarmy@${latest}`, "--no-audit", "--no-fund"], { stdio: json ? "ignore" : "inherit" });
1731
+ }
1732
+ // A child process, so the reconnect runs the code just installed rather
1733
+ // than the old code this process loaded.
1734
+ const targets = connectedTargets();
1735
+ if (targets.length) {
1736
+ if (!json) console.log(`\nReconnecting ${targets.join(", ")}...`);
1737
+ execFileSync(process.execPath, [path.join(nomarmyRoot, "bin", "nomarmy.mjs"), "connect", ...targets, ...(json ? ["--json"] : [])], { stdio: json ? "ignore" : "inherit" });
1738
+ }
1739
+ if (json) return out({ updated: upgrade, from: current, version: upgrade ? latest : current, resynced: targets });
1740
+ console.log(c.yellow("\nRestart every open Claude Code, Codex and Cursor session: each keeps the code it started with until then."));
1741
+ }
1742
+
1707
1743
  function commandExists(cmd) {
1708
1744
  try { execFileSync(process.platform === "win32" ? "where" : "which", [cmd], { stdio: "ignore" }); return true; }
1709
1745
  catch { return false; }
@@ -9,15 +9,17 @@ RUN apt-get update && apt-get install -y --no-install-recommends \
9
9
  # (verified live: apt installed 1.19 here while go.dev's stable was 1.27) --
10
10
  # same reasoning as using rustup instead of apt for Rust in Dockerfile.rust.
11
11
  # Pinned, not "latest", for a reproducible build; bump GO_VERSION by hand
12
- # periodically against https://go.dev/dl/.
12
+ # periodically against https://go.dev/dl/, with both SHA256 values from
13
+ # https://go.dev/dl/?mode=json.
13
14
  ENV GO_VERSION=1.27.1
14
15
  RUN ARCH="$(dpkg --print-architecture)" \
15
16
  && case "$ARCH" in \
16
- amd64) GOARCH=amd64 ;; \
17
- arm64) GOARCH=arm64 ;; \
17
+ amd64) GOARCH=amd64; SHA256=63d339f0da5ab53635a56f2490a7984dfe12dfcff22ad749f63edaf590168445 ;; \
18
+ arm64) GOARCH=arm64; SHA256=3450b45a3f9ee8568792736a5c5e70a1f2e9b36c35a8f74958c03e51d7d92bec ;; \
18
19
  *) echo "unsupported architecture for Go install: $ARCH" >&2; exit 1 ;; \
19
20
  esac \
20
21
  && curl -fsSL "https://go.dev/dl/go${GO_VERSION}.linux-${GOARCH}.tar.gz" -o /tmp/go.tgz \
22
+ && echo "${SHA256} /tmp/go.tgz" | sha256sum -c - \
21
23
  && tar -C /usr/local -xzf /tmp/go.tgz \
22
24
  && rm /tmp/go.tgz
23
25
  ENV PATH="/usr/local/go/bin:${PATH}"
@@ -14,6 +14,21 @@ USER node
14
14
  ENV RUSTUP_HOME=/home/node/.rustup
15
15
  ENV CARGO_HOME=/home/node/.cargo
16
16
  ENV PATH="${CARGO_HOME}/bin:${PATH}"
17
- RUN curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y --profile minimal --default-toolchain stable
17
+ # rustup-init is pinned by version and checksum instead of piping
18
+ # sh.rustup.rs into sh; rustup then verifies the toolchain it downloads.
19
+ # Bump RUSTUP_VERSION and both hashes together, from
20
+ # https://static.rust-lang.org/rustup/archive/<version>/<target>/rustup-init.sha256.
21
+ ENV RUSTUP_VERSION=1.29.1
22
+ RUN ARCH="$(dpkg --print-architecture)" \
23
+ && case "$ARCH" in \
24
+ amd64) TARGET=x86_64-unknown-linux-gnu; SHA256=dda7234360b7f578ca8b0ddcb80145646fa61a67c1720a5abc7051b35c9fcb71 ;; \
25
+ arm64) TARGET=aarch64-unknown-linux-gnu; SHA256=15f6e4ce9f583b929c996c91562bad6d4454f3281de858b02cdfdef615fac433 ;; \
26
+ *) echo "unsupported architecture for Rust install: $ARCH" >&2; exit 1 ;; \
27
+ esac \
28
+ && curl --proto '=https' --tlsv1.2 -fsSL "https://static.rust-lang.org/rustup/archive/${RUSTUP_VERSION}/${TARGET}/rustup-init" -o /tmp/rustup-init \
29
+ && echo "${SHA256} /tmp/rustup-init" | sha256sum -c - \
30
+ && chmod +x /tmp/rustup-init \
31
+ && /tmp/rustup-init -y --no-modify-path --profile minimal --default-toolchain stable \
32
+ && rm /tmp/rustup-init
18
33
  WORKDIR /workspace
19
34
  CMD ["sleep", "infinity"]
package/install.sh CHANGED
@@ -99,9 +99,23 @@ fi
99
99
  # on a fresh macOS install nothing has initialized that VM yet at this point.
100
100
  if nomarmy_is_cloud; then need aws || { echo 'ERROR: the AWS CLI is required for cloud profiles.'; exit 1; }; fi
101
101
  "$ROOT/scripts/install-llama-cpp.sh" "$NOMARMY_PROFILE"
102
+ # A pinned npm release, which npm checks against the registry's integrity
103
+ # hash, rather than piping a remote installer script into bash.
104
+ OPENCLAW_VERSION="${NOMARMY_OPENCLAW_VERSION:-2026.9.6}"
102
105
  if ! command -v openclaw >/dev/null 2>&1; then
103
- echo '==> Installing OpenClaw (non-interactive)'
104
- curl -fsSL https://openclaw.ai/install.sh | bash -s -- --no-onboard
106
+ if ! node -e 'const [a, b] = process.versions.node.split(".").map(Number); process.exit((a === 24 && b >= 16) || (a === 26 && b >= 1) || a > 26 ? 0 : 1)'; then
107
+ echo "ERROR: OpenClaw $OPENCLAW_VERSION needs Node 24.16+ or 26.1+; this is Node $(node -v)."
108
+ echo ' Upgrade Node and re-run this installer, or install OpenClaw yourself first:'
109
+ echo ' https://openclaw.ai (then re-run this installer)'
110
+ exit 1
111
+ fi
112
+ echo "==> Installing OpenClaw $OPENCLAW_VERSION from npm"
113
+ if [[ -w "$(npm prefix -g)" ]]; then
114
+ npm install -g "openclaw@$OPENCLAW_VERSION" --no-audit --no-fund
115
+ else
116
+ # No writable global prefix: install for this user, as OpenClaw's own installer does.
117
+ npm install -g --prefix "$HOME/.npm-global" "openclaw@$OPENCLAW_VERSION" --no-audit --no-fund
118
+ fi
105
119
  export PATH="$HOME/.local/bin:$HOME/.npm-global/bin:$PATH"
106
120
  fi
107
121
  need openclaw
package/lib/admission.mjs CHANGED
@@ -1,6 +1,7 @@
1
1
  import fs from "node:fs";
2
2
  import path from "node:path";
3
3
  import { executionMode } from "./execution.mjs";
4
+ import { jobLabel, jobElapsedSeconds } from "./job-format.mjs";
4
5
  import { loadConfig } from "./config.mjs";
5
6
  import { clampInt } from "./budget-state.mjs";
6
7
  import { checkBrief, assessAdmission, describeBudgets } from "./budget.mjs";
@@ -160,7 +161,7 @@ export function createJobRuntime(deps) {
160
161
  const who = entry.mode === "verify" ? "verification runner" : entry.agent ? `${entry.agent}${entry.model ? `/${entry.model}` : ""}` : "local model";
161
162
  const took = Math.round((Date.now() - Date.parse(entry.startedAt)) / 60000);
162
163
  const ok = !error && (result?.ok || m.coordinatorStatus === "complete");
163
- notify(`nomArmy: ${entry.role ?? entry.mode ?? "job"} ${ok ? "done" : outcome}`, `${entry.workerId ?? entry.jobId} on ${who}: ${outcome} after ${took}m. ${ok ? "Ready for the General's review." : "Needs a look."}`);
164
+ notify(`nomArmy: ${entry.role ?? entry.mode ?? "job"} ${ok ? "done" : outcome}`, `${entry.label ? `${entry.label} (${entry.workerId ?? entry.jobId})` : entry.workerId ?? entry.jobId} on ${who}: ${outcome} after ${took}m. ${ok ? "Ready for the General's review." : "Needs a look."}`);
164
165
  }
165
166
  function admissionHardware() {
166
167
  return executionMode(deps.env).managesModelServer ? deps.budgetState.hardwareSnapshot : null;
@@ -349,7 +350,7 @@ export function createJobRuntime(deps) {
349
350
  function launch(args) {
350
351
  const workerId = args.worker_id || null;
351
352
  const jobId = slug(workerId || (args.mode === "scout" ? "scout" : "worker"));
352
- return trackInRun(args, track(jobId, { mode: args.mode, workerId: workerId || jobId, lane: jobLane(args), agent: args.agentName ?? null, runId: args.run_id ?? null, role: args.armyRole ?? null, model: args.model ?? null },
353
+ return trackInRun(args, track(jobId, { mode: args.mode, workerId: workerId || jobId, lane: jobLane(args), agent: args.agentName ?? null, runId: args.run_id ?? null, role: args.armyRole ?? null, model: args.model ?? null, label: jobLabel(args) },
353
354
  withAgentSlot(args, jobId, () => executeJob({ ...jobArgs(args, workerId), jobId }))));
354
355
  }
355
356
  // Best-effort progress signal for a job still mid-run: a plain "phase: worker,
@@ -395,7 +396,7 @@ export function createJobRuntime(deps) {
395
396
 
396
397
  async function summarize(entry, files, jobDir = null) {
397
398
  const status = files.status, meta = files.meta ?? files.failure;
398
- const elapsedSeconds = status?.startedAt ? Math.round((Date.now() - Date.parse(status.startedAt)) / 1000) : entry ? Math.round((Date.now() - Date.parse(entry.startedAt)) / 1000) : null;
399
+ const elapsedSeconds = jobElapsedSeconds({ status, meta, entry });
399
400
  const out = { jobId: entry?.jobId ?? status?.jobId ?? meta?.jobId ?? null, workerId: entry?.workerId ?? status?.workerId ?? meta?.workerId ?? null,
400
401
  mode: entry?.mode ?? status?.mode ?? meta?.mode ?? null, state: null, phase: status?.phase ?? "starting", elapsedSeconds,
401
402
  timeoutSeconds: status?.timeoutSeconds ?? null, coordinatorStatus: meta?.coordinatorStatus ?? null, outcome: meta?.outcome ?? null,
package/lib/agents.mjs CHANGED
@@ -30,6 +30,7 @@ import fs from "node:fs";
30
30
  import path from "node:path";
31
31
  import YAML from "yaml";
32
32
  import { z } from "zod";
33
+ import { typeError, isMissingField, isUnknownDiscriminator } from "./zod-issues.mjs";
33
34
 
34
35
  import { providerEntrySchema, formatDispatchIssues, openclawProviderId, PROVIDER_TYPES, ID_RE, HOST_TOOL_PROVIDERS, agentRunsToolsOnHost } from "./dispatch-schema.mjs";
35
36
 
@@ -45,10 +46,10 @@ export const RESERVED_AGENT_NAMES = Object.freeze(["__proto__", "constructor", "
45
46
  export const BUILTIN_LOCAL_AGENT = Object.freeze({ kind: "local", slot: "coder" });
46
47
 
47
48
  const requiredString = () =>
48
- z.string({ required_error: "is required", invalid_type_error: "must be a string" })
49
+ z.string({ error: typeError("a string") })
49
50
  .refine((value) => value.trim().length > 0, { message: "must not be empty" });
50
51
  const thinkingSchema = z.union([z.boolean(), z.enum(["low", "medium", "high"])]).default(true);
51
- const positiveInt = () => z.number({ invalid_type_error: "must be a number" }).int("must be a whole number").positive("must be a positive number");
52
+ const positiveInt = () => z.number({ error: typeError("a number") }).int("must be a whole number").positive("must be a positive number");
52
53
 
53
54
  const localAgentSchema = z.object({
54
55
  kind: z.literal("local"),
@@ -62,7 +63,7 @@ const localAgentSchema = z.object({
62
63
  // providerEntrySchema in agentsFileSchema's superRefine, so they exist once.
63
64
  const apiAgentSchema = z.object({
64
65
  kind: z.literal("api"),
65
- provider: z.enum(API_PROVIDER_TYPES, { errorMap: () => ({ message: `must be one of ${API_PROVIDER_TYPES.join(", ")}` }) }),
66
+ provider: z.enum(API_PROVIDER_TYPES, { error: `must be one of ${API_PROVIDER_TYPES.join(", ")}` }),
66
67
  model: requiredString().optional(),
67
68
  auth_env: requiredString(),
68
69
  base_url: z.string().optional(),
@@ -95,7 +96,7 @@ export function hostToolsImplementProblem(name, agent) {
95
96
  }
96
97
 
97
98
  export const agentSchema = z.discriminatedUnion("kind", [localAgentSchema, apiAgentSchema, subscriptionAgentSchema], {
98
- errorMap: (issue, ctx) => (issue.code === "invalid_union_discriminator" ? { message: `kind must be one of ${AGENT_KINDS.join(", ")}` } : { message: ctx.defaultError }),
99
+ error: (issue) => (isUnknownDiscriminator(issue) ? `kind must be one of ${AGENT_KINDS.join(", ")}` : undefined),
99
100
  });
100
101
 
101
102
  /** An api agent as the pool entry the dispatch path understands. `model` may be absent (see resolveAgentModel). */
@@ -148,7 +149,7 @@ export function formatAgentIssues(error) {
148
149
  const lines = error.issues.map((issue) => {
149
150
  const where = issue.path.length ? issue.path.join(".") : "config";
150
151
  if (issue.code === "unrecognized_keys") return `${where}: unexpected field(s) ${issue.keys.map((k) => `"${k}"`).join(", ")}`;
151
- if (issue.code === "invalid_type" && issue.received === "undefined") return `${where}: is required`;
152
+ if (isMissingField(issue)) return `${where}: is required`;
152
153
  return `${where}: ${issue.message}`;
153
154
  });
154
155
  return [...new Set(lines)];
package/lib/army.mjs CHANGED
@@ -29,6 +29,7 @@ import os from "node:os";
29
29
  import path from "node:path";
30
30
  import YAML from "yaml";
31
31
  import { z } from "zod";
32
+ import { typeError } from "./zod-issues.mjs";
32
33
  import { agentRunsToolsOnHost } from "./dispatch-schema.mjs";
33
34
  import { usageStatus } from "./usage-limits.mjs";
34
35
 
@@ -116,7 +117,7 @@ const roleSchema = z.object({
116
117
  disabled: z.boolean().optional(),
117
118
  }).strict();
118
119
 
119
- const positiveNumber = () => z.number({ invalid_type_error: "must be a number" }).positive("must be positive");
120
+ const positiveNumber = () => z.number({ error: typeError("a number") }).positive("must be positive");
120
121
  // Limits on one /feature run (lib/runs.mjs). Personal, like the General:
121
122
  // global or local only, never a committed project file.
122
123
  const runLimitsSchema = z.object({
package/lib/connect.mjs CHANGED
@@ -28,6 +28,7 @@ import { execFileSync } from "node:child_process";
28
28
  import { loadAgents } from "./agents.mjs";
29
29
  import { globalConfigDir } from "./army.mjs";
30
30
  import { buildNotifierApp } from "./notifier-app.mjs";
31
+ import { recordCopySource } from "./install-freshness.mjs";
31
32
 
32
33
  export function defaultInstallDir() {
33
34
  return process.env.NOMARMY_AGENT_INSTALL_DIR || path.join(process.env.HOME ?? process.env.USERPROFILE ?? ".", ".local", "share", "nomarmy-local-worker");
@@ -142,6 +143,7 @@ export function installMcpCopy({ nomarmyRoot, installDir, run = defaultRun }) {
142
143
  fs.mkdirSync(path.join(installDir, "mcp"), { recursive: true });
143
144
  fs.copyFileSync(path.join(nomarmyRoot, "package.json"), path.join(installDir, "package.json"));
144
145
  fs.copyFileSync(path.join(nomarmyRoot, "mcp", "server.mjs"), path.join(installDir, "mcp", "server.mjs"));
146
+ recordCopySource(installDir, nomarmyRoot);
145
147
  fs.rmSync(path.join(installDir, "lib"), { recursive: true, force: true });
146
148
  fs.cpSync(path.join(nomarmyRoot, "lib"), path.join(installDir, "lib"), { recursive: true });
147
149
  // lib/sandbox-images.mjs resolves docker/ relative to its own location
@@ -1,3 +1,5 @@
1
+ import path from "node:path";
2
+
1
3
  // ---------------------------------------------------------------------------
2
4
  // Test-change classification (plan 16). One tunable constant, on purpose:
3
5
  // every heuristic about what counts as a test file lives here and nowhere else.
@@ -416,3 +418,91 @@ export function mergeUntrackedIntoNameStatus(nameStatus, untrackedFiles) {
416
418
  return [...(nameStatus ?? []), ...extra];
417
419
  }
418
420
 
421
+
422
+ // ---------------------------------------------------------------------------
423
+ // Verification inputs: .nomarmy.yml is read from the operator's checkout so a
424
+ // worker can't weaken its own checks, but the files its commands run come
425
+ // from the worker's worktree. Found live: a worker whose verification ran
426
+ // `node check.js` rewrote check.js to print PASS and exit 0, and the job
427
+ // committed as done. The revert check can't see it: reverting check.js along
428
+ // with the code makes verification fail, which reads as coverage.
429
+ //
430
+ // Blocking (the diff changes what a command runs): a changed non-test file a
431
+ // command names (`node check.js`, `bash scripts/verify.sh`), a changed
432
+ // Makefile or justfile under `make`/`just`, or a changed package.json script
433
+ // that `npm test`/`pnpm run x`/`yarn x`/`bun run x` calls. Test files named by
434
+ // a command are left to the test-change review, since adding cases is normal.
435
+ // Flagged only: test-runner configuration, which is often a legitimate edit.
436
+ // ---------------------------------------------------------------------------
437
+
438
+ const RUNNER_CONFIG_RE = /(^|\/)(conftest\.py|pytest\.ini|tox\.ini|setup\.cfg|\.coveragerc|(jest|vitest|vite|playwright|karma|cypress)\.config\.[cm]?[jt]s|\.mocharc(\.[a-z]+)?|phpunit\.xml(\.dist)?|\.rspec)$/;
439
+ const TASK_FILES = { make: ["Makefile", "makefile", "GNUmakefile"], just: ["justfile", "Justfile", ".justfile"] };
440
+
441
+ function shellWords(segment) {
442
+ return (segment.match(/"[^"]*"|'[^']*'|\S+/g) ?? []).map((w) => w.replace(/^["']|["']$/g, ""));
443
+ }
444
+
445
+ // Which package.json script a package-manager command runs, if any.
446
+ function packageScript(words) {
447
+ const [tool, first, second] = words;
448
+ if (!["npm", "pnpm", "yarn", "bun"].includes(tool) || !first) return null;
449
+ if (["run", "run-script"].includes(first)) return second && !second.startsWith("-") ? second : null;
450
+ if (tool === "npm") return ["test", "t", "tst"].includes(first) ? "test" : ["start", "stop", "restart"].includes(first) ? first : null;
451
+ if (tool === "bun") return null; // `bun test` is bun's own runner, not a script
452
+ return first.startsWith("-") || ["install", "add", "remove", "exec", "dlx", "x"].includes(first) ? null : first;
453
+ }
454
+
455
+ function scriptsOf(text) {
456
+ try { return JSON.parse(text ?? "")?.scripts ?? {}; } catch { return {}; }
457
+ }
458
+
459
+ function pytestSection(text) {
460
+ const match = /^\[tool\.pytest[^\]]*\]\s*$([\s\S]*?)(?=^\[|(?![\s\S]))/m.exec(String(text ?? ""));
461
+ return match ? match[1].trim() : null;
462
+ }
463
+
464
+ /**
465
+ * @param {{ commands: string[], changedFiles: string[], readBase: (file: string) => string|null|Promise<string|null>,
466
+ * readHead: (file: string) => string|null|Promise<string|null>, isTestPathFn?: (file: string) => boolean }} input
467
+ * @returns {Promise<{ blocked: {file: string, command: string, why: string}[], flagged: {file: string, why: string}[] } | null>}
468
+ */
469
+ export async function detectVerificationInputChanges({ commands = [], changedFiles = [], readBase = () => null, readHead = () => null, isTestPathFn = isTestPath } = {}) {
470
+ const changed = new Set(changedFiles.map((f) => path.posix.normalize(f)));
471
+ if (!changed.size) return null;
472
+ const blocked = [], flagged = [], seen = new Set();
473
+ const block = (file, command, why) => { if (!seen.has(file)) { seen.add(file); blocked.push({ file, command, why }); } };
474
+ for (const command of commands) {
475
+ let dir = "";
476
+ for (const segment of String(command ?? "").split(/\s*(?:&&|\|\||;|\|)\s*/)) {
477
+ const all = shellWords(segment);
478
+ let lead = 0;
479
+ while (lead < all.length && /^[A-Za-z_][A-Za-z0-9_]*=/.test(all[lead])) lead++; // FOO=1 cmd ...
480
+ const words = all.slice(lead);
481
+ if (!words.length) continue;
482
+ const at = (p) => path.posix.normalize(path.posix.join(dir, p));
483
+ if (words[0] === "cd" && words[1]) { dir = at(words[1]); continue; }
484
+ for (const word of words) {
485
+ if (word.startsWith("-") || word.includes("$")) continue;
486
+ const file = at(word);
487
+ if (changed.has(file) && !isTestPathFn(file)) block(file, command, `run by \`${command}\``);
488
+ }
489
+ for (const name of TASK_FILES[words[0]] ?? []) {
490
+ const file = at(name);
491
+ if (changed.has(file)) block(file, command, `read by \`${command}\``);
492
+ }
493
+ const script = packageScript(words);
494
+ const manifest = at("package.json");
495
+ if (script && changed.has(manifest)) {
496
+ const before = scriptsOf(await readBase(manifest)), after = scriptsOf(await readHead(manifest));
497
+ const touched = [script, `pre${script}`, `post${script}`].filter((s) => before[s] !== after[s]);
498
+ if (touched.length) block(manifest, command, `its script${touched.length > 1 ? "s" : ""} ${touched.map((s) => `"${s}"`).join(", ")} changed, and \`${command}\` runs it`);
499
+ }
500
+ }
501
+ }
502
+ for (const file of changed) {
503
+ if (seen.has(file)) continue;
504
+ if (RUNNER_CONFIG_RE.test(file)) flagged.push({ file, why: "test-runner configuration" });
505
+ else if (/(^|\/)pyproject\.toml$/.test(file) && pytestSection(await readBase(file)) !== pytestSection(await readHead(file))) flagged.push({ file, why: "its [tool.pytest] settings changed" });
506
+ }
507
+ return blocked.length || flagged.length ? { blocked, flagged } : null;
508
+ }
@@ -15,10 +15,11 @@
15
15
  // never in a file nomArmy writes or reads back.
16
16
 
17
17
  import { z } from "zod";
18
+ import { typeError, isMissingField, isUnknownDiscriminator } from "./zod-issues.mjs";
18
19
 
19
20
  const requiredString = () =>
20
21
  z
21
- .string({ required_error: "is required", invalid_type_error: "must be a string" })
22
+ .string({ error: typeError("a string") })
22
23
  .refine((value) => value.trim().length > 0, { message: "must not be empty" });
23
24
 
24
25
  // Exported (not just used internally) so bin/nomarmy.mjs's interactive
@@ -101,11 +102,11 @@ export function isNativeProviderType(type) {
101
102
  export const CUSTOM_ENDPOINT_PROVIDER_TYPES = Object.freeze(["bedrock", "azure-openai", "openai-compatible"]);
102
103
 
103
104
  const idSchema = z
104
- .string({ required_error: "is required", invalid_type_error: "must be a string" })
105
+ .string({ error: typeError("a string") })
105
106
  .regex(ID_RE, "must be 1-64 characters of letters, numbers, dot, underscore or hyphen");
106
107
 
107
108
  const weightSchema = z
108
- .number({ required_error: "is required", invalid_type_error: "must be a number" })
109
+ .number({ error: typeError("a number") })
109
110
  .positive("must be a positive number")
110
111
  .finite("must be a finite number");
111
112
 
@@ -113,13 +114,13 @@ const weightSchema = z
113
114
  // one entry at once -- a stand-in for real rate-limit-aware admission (see
114
115
  // README's dispatch-pool section for why that's out of scope for now).
115
116
  const maxConcurrentSchema = z
116
- .number({ invalid_type_error: "must be a number" })
117
+ .number({ error: typeError("a number") })
117
118
  .int("must be a whole number")
118
119
  .positive("must be a positive number")
119
120
  .default(2);
120
121
 
121
122
  const authEnvSchema = z
122
- .string({ required_error: "is required", invalid_type_error: "must be a string" })
123
+ .string({ error: typeError("a string") })
123
124
  .regex(
124
125
  AUTH_ENV_NAME_RE,
125
126
  "must be an environment variable NAME (uppercase letters, digits, underscores), never the credential itself",
@@ -127,7 +128,7 @@ const authEnvSchema = z
127
128
 
128
129
  const urlSchema = () =>
129
130
  z
130
- .string({ required_error: "is required", invalid_type_error: "must be a string" })
131
+ .string({ error: typeError("a string") })
131
132
  .url("must be a valid URL");
132
133
 
133
134
  // Whether/how the entry's model uses a thinking/reasoning mode. `true`
@@ -152,7 +153,7 @@ const thinkingSchema = z.union([z.boolean(), z.enum(["low", "medium", "high"])])
152
153
  // knows about yet (see lib/model-catalog.mjs), or an operator who wants to
153
154
  // be more conservative than the model's rated maximum.
154
155
  const contextWindowSchema = z
155
- .number({ invalid_type_error: "must be a number" })
156
+ .number({ error: typeError("a number") })
156
157
  .int("must be a whole number")
157
158
  .positive("must be a positive number")
158
159
  .optional();
@@ -189,7 +190,7 @@ const llamaCppEntrySchema = z
189
190
  // `agents add` installs it before registering the key.
190
191
  const genericOpenclawEntrySchema = hostedProviderSchema("openclaw", { requireBaseUrl: false }).extend({
191
192
  openclaw_provider: z
192
- .string({ required_error: "is required", invalid_type_error: "must be a string" })
193
+ .string({ error: typeError("a string") })
193
194
  .regex(OPENCLAW_PROVIDER_ID_RE, "must be an OpenClaw provider id (lowercase letters, digits, dot, underscore, hyphen)")
194
195
  .refine((id) => !PROVIDER_TYPES.includes(id) || NATIVE_PROVIDER_TYPES.includes(id), "names a nomArmy provider type with its own setup -- use that type instead"),
195
196
  plugin: z.string().regex(/^\S+$/, "must be one plugin install spec, e.g. clawhub:@openclaw/deepseek-provider").optional(),
@@ -205,7 +206,7 @@ export const providerEntrySchema = z.discriminatedUnion("provider", [
205
206
  export const poolsSchema = z.record(
206
207
  z.string().min(1, "pool name must not be empty"),
207
208
  z
208
- .array(providerEntrySchema, { invalid_type_error: "must be an array of provider entries" })
209
+ .array(providerEntrySchema, { error: typeError("an array of provider entries") })
209
210
  .min(1, "must list at least one provider entry"),
210
211
  );
211
212
 
@@ -250,11 +251,11 @@ export function formatDispatchIssues(error) {
250
251
  lines.push(`${where}: unexpected field(s) ${keys}`);
251
252
  continue;
252
253
  }
253
- if (issue.code === "invalid_union_discriminator") {
254
+ if (isUnknownDiscriminator(issue)) {
254
255
  lines.push(`${where}: must be one of ${PROVIDER_TYPES.join(", ")}`);
255
256
  continue;
256
257
  }
257
- if (issue.code === "invalid_type" && issue.received === "undefined") {
258
+ if (isMissingField(issue)) {
258
259
  lines.push(`${where}: is required`);
259
260
  continue;
260
261
  }
package/lib/execute.mjs CHANGED
@@ -17,7 +17,7 @@ import { describeRecoveryChanges, reportRecoveryPrompt } from "./worker-prompt.m
17
17
  import { parseWorkerReport } from "./report.mjs";
18
18
  import { OUTCOMES, COORDINATOR_STATUS_BY_OUTCOME } from "./outcomes.mjs";
19
19
  import { resolveOutcome, finalText, workerMetadata, applyRefactorContract, applyVerificationPolicy } from "./outcome.mjs";
20
- import { isTestPath, isDocumentationPath, detectScopedTestSelectionRisk, detectUnwiredNewDefinitions, detectMislabeledTestNames, detectPossibleSecrets } from "./diff-checks.mjs";
20
+ import { isTestPath, isDocumentationPath, detectScopedTestSelectionRisk, detectUnwiredNewDefinitions, detectMislabeledTestNames, detectPossibleSecrets, detectVerificationInputChanges } from "./diff-checks.mjs";
21
21
 
22
22
  // ---------------------------------------------------------------------------
23
23
  // Job status for polling. `status.json` is written at every phase transition
@@ -250,7 +250,7 @@ export function createExecutor(deps) {
250
250
  // which reads as evidence about work that never happened.
251
251
  independentVerification = normalizeVerification({ status: "not_run", basis: "not-applicable", reason: "the worker changed nothing, so there was none of its work to verify" }, verification ?? null);
252
252
  } else if (verificationFlow.verificationRunner || !reportValidation.valid) {
253
- independentVerification = await runIndependentVerification({ profile: verification ?? null, cwd, jobId, baseSha: base.sha, branch, mode, record: preCommit });
253
+ independentVerification = await runIndependentVerification({ profile: verification ?? null, cwd, jobId, baseSha: base.sha, branch, mode, record: preCommit, logFile: path.join(jobDir, "verification.log") });
254
254
  }
255
255
 
256
256
  // verify_regression: on by default whenever there's a verification
@@ -299,12 +299,19 @@ export function createExecutor(deps) {
299
299
  // Cheap, always-on, additive: never changes commitAllowed/commitBlockedReason
300
300
  // on its own (unlike the regression-check override above), only flags for
301
301
  // review -- see detectScopedTestSelectionRisk's own doc comment for why.
302
- let selectionRisk = null;
302
+ let selectionRisk = null, verificationInputs = null;
303
303
  if (mode === "implement" && verification) {
304
304
  try {
305
305
  const loaded = loadConfig(projectDir); // the operator's contract; see registerVerificationRunner's call
306
306
  const profileCommands = loaded.found ? (loaded.config?.verification?.[verification]?.commands ?? []) : [];
307
307
  selectionRisk = detectScopedTestSelectionRisk({ commands: profileCommands, testChanges: preCommit.testChanges });
308
+ // A diff that changes what those commands run (see
309
+ // detectVerificationInputChanges): applied below, after the others.
310
+ verificationInputs = await detectVerificationInputChanges({
311
+ commands: profileCommands, changedFiles: preCommit.changedFiles ?? [],
312
+ readBase: (file) => gitRaw(["show", `${base.sha}:${file}`], cwd).catch(() => null),
313
+ readHead: (file) => { try { return fs.readFileSync(path.join(cwd, file), "utf8"); } catch { return null; } },
314
+ });
308
315
  } catch { /* a config load failure here is the verification runner's own problem to report, not this check's */ }
309
316
  }
310
317
  const afterSelectionRisk = selectionRisk
@@ -399,7 +406,21 @@ export function createExecutor(deps) {
399
406
  commitBlockedReason: `possible secret detected: ${possibleSecrets.reason}`,
400
407
  reasons: [...afterHostInstalls.reasons, `POSSIBLE SECRET DETECTED: ${possibleSecrets.reason}`] }
401
408
  : afterHostInstalls;
402
- const finalOutcome = applyRefactorContract(applyVerificationPolicy(afterSecrets, independentVerification.status, repoPolicy()),
409
+ // A worker must not be judged by a check it rewrote: a changed script,
410
+ // Makefile or package.json script that a verification command runs
411
+ // blocks the commit; changed test-runner config only asks for review.
412
+ const blockedInputs = verificationInputs?.blocked ?? [], flaggedInputs = verificationInputs?.flagged ?? [];
413
+ const inputLine = blockedInputs.map((b) => `${b.file} (${b.why})`).join("; ");
414
+ const afterInputs = blockedInputs.length
415
+ ? { ...afterSecrets, outcome: afterSecrets.commitAllowed || afterSecrets.outcome === OUTCOMES.WORKER_DONE ? OUTCOMES.NEEDS_REVIEW : afterSecrets.outcome,
416
+ reviewRequired: true, commitAllowed: false,
417
+ commitBlockedReason: afterSecrets.commitAllowed ? `the diff changes what verification runs: ${inputLine}` : afterSecrets.commitBlockedReason,
418
+ reasons: [...afterSecrets.reasons, `VERIFICATION INPUT CHANGED: the diff changes what profile '${verification}' runs, so its result can't be trusted: ${inputLine}`] }
419
+ : afterSecrets;
420
+ const afterConfig = flaggedInputs.length
421
+ ? { ...afterInputs, reviewRequired: true, reasons: [...afterInputs.reasons, `TEST CONFIG CHANGED: ${flaggedInputs.map((c) => `${c.file} (${c.why})`).join("; ")}`] }
422
+ : afterInputs;
423
+ const finalOutcome = applyRefactorContract(applyVerificationPolicy(afterConfig, independentVerification.status, repoPolicy()),
403
424
  { refactor, verificationStatus: independentVerification.status, testChanges: preCommit.testChanges });
404
425
 
405
426
  progress("commit");
@@ -411,6 +432,7 @@ export function createExecutor(deps) {
411
432
  let coordinatorStatus = COORDINATOR_STATUS_BY_OUTCOME[finalOutcome.outcome] ?? "incomplete";
412
433
  const issues = [...finalOutcome.reasons, ...(preCommit.issues ?? [])];
413
434
  if (workerError) issues.push(`worker error: ${String(workerError).split("\n")[0]}`);
435
+ if ((result ?? attempted)?.salvaged) issues.push(`runner cleanup failed after the run (${(result ?? attempted).salvagedFrom}); the worker's report was recovered from the run's transcript`);
414
436
  if (workerFailed || workerTimedOut) { const restarted = vmRestartIssue(vmStartedBefore, deps.podmanVmStartedAt?.() ?? null); if (restarted) issues.unshift(restarted); }
415
437
  if (repositoryChanged && !commit.created) {
416
438
  if (coordinatorStatus === "complete") coordinatorStatus = "incomplete";
@@ -575,6 +597,7 @@ export function createExecutor(deps) {
575
597
  const worker = workerMetadata(result ?? attempted);
576
598
  const issues = [...outcome.reasons];
577
599
  if (workerError) issues.push(`scout error: ${String(workerError).split("\n")[0]}`);
600
+ if ((result ?? attempted)?.salvaged) issues.push(`runner cleanup failed after the run (${(result ?? attempted).salvagedFrom}); the scout's report was recovered from the run's transcript`);
578
601
  const failures = worker.toolSummary?.failures ?? 0; if (failures > 0) issues.push(`scout recorded ${failures} tool failure(s)`);
579
602
  if (dirty) issues.push(`snapshot changed: ${record.repoStatusFiles.join(", ")}`);
580
603
  if (reportRecoveryAttempted) {
@@ -706,6 +729,7 @@ export function createExecutor(deps) {
706
729
  const worker = workerMetadata(result ?? attempted);
707
730
  const issues = [...outcome.reasons];
708
731
  if (workerError) issues.push(`decompose error: ${String(workerError).split("\n")[0]}`);
732
+ if ((result ?? attempted)?.salvaged) issues.push(`runner cleanup failed after the run (${(result ?? attempted).salvagedFrom}); the decomposer's report was recovered from the run's transcript`);
709
733
  const failures = worker.toolSummary?.failures ?? 0; if (failures > 0) issues.push(`decomposer recorded ${failures} tool failure(s)`);
710
734
  if (dirty) issues.push(`snapshot changed: ${record.repoStatusFiles.join(", ")}`);
711
735
  if (overlaps.length) issues.push(`${overlaps.length} subtask pair(s) claim overlapping files; not safe to dispatch as independent jobs as proposed`);
package/lib/health.mjs CHANGED
@@ -11,6 +11,8 @@
11
11
  // army a role or the General pointing at something unusable
12
12
  // config agents.yml unloadable or unsafe
13
13
  // leftovers retained worktrees, job storage, stale "running" jobs
14
+ // nomarmy-update a newer nomArmy on npm; nomarmy-copy: coordinators run
15
+ // an older copy than the installed CLI (install-freshness.mjs)
14
16
  //
15
17
  // Each issue: { id, severity: "error"|"warn"|"info", title, detail, fix,
16
18
  // short }. `id` is stable across runs, so a notification goes out once per
@@ -23,6 +25,7 @@ import { executionMode } from "./execution.mjs";
23
25
  import { modelRejection } from "./openclaw-errors.mjs";
24
26
  import { providerConfigured, readOpenclawConfig } from "./openclaw-config.mjs";
25
27
  import { readUsageSnapshots, usageStatus } from "./usage-limits.mjs";
28
+ import { freshnessIssues, readInstallVersions } from "./install-freshness.mjs";
26
29
 
27
30
  const DAY = 86400000;
28
31
 
@@ -213,7 +216,7 @@ export function leftoverIssues({ retainedWorktrees = 0, jobsBytes = 0, staleRunn
213
216
  * Run every check. `env` supplies what each needs, with real defaults;
214
217
  * tests pass their own.
215
218
  */
216
- export async function runHealthChecks({ now = Date.now(), mode = "local", openclawCmd = process.env.NOMARMY_OPENCLAW_CMD || "openclaw", run = runBounded, armySummary = null, agentsError = null, jobsRoot = null, pidAlive = () => true, agents = null, openclawConfig = null, vendors = {}, modelsInUse = null, autoPruned = null, usageSnapshots = null } = {}) {
219
+ export async function runHealthChecks({ now = Date.now(), mode = "local", openclawCmd = process.env.NOMARMY_OPENCLAW_CMD || "openclaw", run = runBounded, armySummary = null, agentsError = null, jobsRoot = null, pidAlive = () => true, agents = null, openclawConfig = null, vendors = {}, modelsInUse = null, autoPruned = null, usageSnapshots = null, install = null } = {}) {
217
220
  const issues = [];
218
221
  if (autoPruned?.freedBytes) issues.push({ id: `auto-prune:${new Date(now).toISOString()}`, severity: "info",
219
222
  title: `Freed ${(autoPruned.freedBytes / 1024 ** 3).toFixed(2)} GB: ${[autoPruned.pruned ? `runtime data of ${autoPruned.pruned} finished job${autoPruned.pruned === 1 ? "" : "s"} older than ${autoPruned.olderThanHours}h` : null, autoPruned.scratchCleared ? `OpenClaw scratch files of ${autoPruned.scratchCleared} more` : null].filter(Boolean).join(", ")}`,
@@ -230,12 +233,14 @@ export async function runHealthChecks({ now = Date.now(), mode = "local", opencl
230
233
  short });
231
234
  }
232
235
  }
233
- const [auth, version, latest, plugins] = await Promise.all([
236
+ const [auth, version, latest, plugins, nomarmyLatest] = await Promise.all([
234
237
  run(openclawCmd, ["models", "auth", "list", "--json"]),
235
238
  run(openclawCmd, ["--version"]),
236
239
  run("npm", ["view", "openclaw", "version"], { timeoutMs: 15000 }),
237
240
  run(openclawCmd, ["plugins", "inspect", "codex"]),
241
+ install ? run("npm", ["view", "nomarmy", "dist-tags.alpha"], { timeoutMs: 15000 }) : null,
238
242
  ]);
243
+ if (install) issues.push(...freshnessIssues({ ...install, latestVersion: nomarmyLatest?.ok ? nomarmyLatest.stdout : null }));
239
244
  if (auth.ok) { try { issues.push(...loginExpiryIssues(JSON.parse(auth.stdout.slice(auth.stdout.indexOf("{"))), { now })); } catch { /* unparseable: skip */ } }
240
245
  const pluginVersion = /Version:\s*(\S+)/.exec(plugins.stdout ?? "")?.[1];
241
246
  if (version.ok) issues.push(...versionIssues({ installed: version.stdout, latest: latest.ok ? latest.stdout : null, plugins: pluginVersion ? [{ id: "codex", version: pluginVersion }] : [] }));
@@ -318,8 +323,10 @@ export async function checkAndRecordHealth({ projectDir, stateRoot, configDir, n
318
323
  let autoPruned = null;
319
324
  if (ageMs !== null) { try { autoPruned = { ...pruneJobRuntime({ stateRoot, olderThanMs: ageMs, now }), olderThanHours: ageMs / 3600000 }; } catch { /* best-effort */ } }
320
325
  const mode = executionMode(env).mode;
326
+ const { defaultInstallDir } = await import("./connect.mjs");
327
+ const install = readInstallVersions(defaultInstallDir());
321
328
  const result = await runHealthChecks({ now, mode, armySummary, agentsError, jobsRoot: path.join(stateRoot, "jobs"), pidAlive,
322
- agents, openclawConfig: readOpenclawConfig(), vendors: SUBSCRIPTION_VENDORS, modelsInUse, autoPruned, usageSnapshots: readUsageSnapshots(stateRoot) });
329
+ agents, openclawConfig: readOpenclawConfig(), vendors: SUBSCRIPTION_VENDORS, modelsInUse, autoPruned, usageSnapshots: readUsageSnapshots(stateRoot), install });
323
330
  const toNotify = recordHealth(path.join(stateRoot, "health.json"), result, { now });
324
331
  return { result, toNotify };
325
332
  }
@@ -0,0 +1,83 @@
1
+ // Whether the nomArmy a coordinator runs is current. Claude Code, Codex and
2
+ // Cursor run a copy of the server that `nomarmy connect` puts in the install
3
+ // dir, and each session keeps the code it started with. So an upgrade can
4
+ // stall at three points, each checked here:
5
+ //
6
+ // nomarmy-update npm's alpha release is newer than the installed CLI
7
+ // nomarmy-copy the CLI is newer than the copy coordinators run
8
+ // (`nomarmy connect` wasn't re-run)
9
+ // restart the copy on disk changed after this server started
10
+ // (the session wasn't restarted); per session, so it is
11
+ // reported by that session's server, not in health.json
12
+
13
+ import fs from "node:fs";
14
+ import path from "node:path";
15
+
16
+ /** Written into the install dir by `nomarmy connect`: where the copy came from. */
17
+ export const SOURCE_FILE = "source.json";
18
+
19
+ export function readPackageVersion(root) {
20
+ try { return JSON.parse(fs.readFileSync(path.join(root, "package.json"), "utf8")).version ?? null; }
21
+ catch { return null; }
22
+ }
23
+
24
+ export function recordCopySource(installDir, nomarmyRoot) {
25
+ fs.writeFileSync(path.join(installDir, SOURCE_FILE), JSON.stringify({ root: nomarmyRoot, version: readPackageVersion(nomarmyRoot) }, null, 2) + "\n");
26
+ }
27
+
28
+ /** The copy's version and, when connect recorded it, the version now at its source. */
29
+ export function readInstallVersions(installDir) {
30
+ let source = null;
31
+ try { source = JSON.parse(fs.readFileSync(path.join(installDir, SOURCE_FILE), "utf8")); } catch { /* connected before source.json existed */ }
32
+ return { copyVersion: readPackageVersion(installDir), sourceVersion: source?.root ? readPackageVersion(source.root) : null };
33
+ }
34
+
35
+ /** Semver order, prerelease included (0.1.0-alpha.7 < 0.1.0-alpha.10 < 0.1.0). */
36
+ export function compareVersions(a, b) {
37
+ const split = (v) => { const [main, pre] = String(v).trim().replace(/^v/, "").split("-", 2); return { main: main.split(".").map(Number), pre: pre ? pre.split(".") : null }; };
38
+ const x = split(a), y = split(b);
39
+ for (let i = 0; i < 3; i++) if ((x.main[i] || 0) !== (y.main[i] || 0)) return (x.main[i] || 0) < (y.main[i] || 0) ? -1 : 1;
40
+ if (!x.pre || !y.pre) return x.pre === y.pre ? 0 : x.pre ? -1 : 1;
41
+ for (let i = 0; i < Math.max(x.pre.length, y.pre.length); i++) {
42
+ const p = x.pre[i], q = y.pre[i];
43
+ if (p === undefined || q === undefined) return p === undefined ? -1 : 1;
44
+ if (p === q) continue;
45
+ const pn = /^\d+$/.test(p), qn = /^\d+$/.test(q);
46
+ if (pn && qn) return Number(p) < Number(q) ? -1 : 1;
47
+ if (pn !== qn) return pn ? -1 : 1;
48
+ return p < q ? -1 : 1;
49
+ }
50
+ return 0;
51
+ }
52
+
53
+ const valid = (v) => typeof v === "string" && /^v?\d+\.\d+\.\d+/.test(v.trim());
54
+
55
+ /** Health issues for a stale CLI or a stale installed copy. */
56
+ export function freshnessIssues({ copyVersion = null, sourceVersion = null, latestVersion = null }) {
57
+ const issues = [];
58
+ if (valid(latestVersion) && valid(sourceVersion) && compareVersions(sourceVersion, latestVersion) < 0) {
59
+ const latest = latestVersion.trim();
60
+ issues.push({ id: `nomarmy-update:${latest}`, severity: "info", title: `nomArmy ${latest} is out (you have ${sourceVersion})`,
61
+ detail: "Updating installs it, reconnects your coordinators and tells you which sessions to restart.",
62
+ fix: "nomarmy update", short: "nomarmy update" });
63
+ }
64
+ if (valid(sourceVersion) && valid(copyVersion) && compareVersions(copyVersion, sourceVersion) < 0) {
65
+ issues.push({ id: `nomarmy-copy:${sourceVersion}`, severity: "warn", title: `Your coordinators run nomArmy ${copyVersion}, but ${sourceVersion} is installed`,
66
+ detail: "Claude Code, Codex and Cursor run a copy of nomArmy that only `nomarmy connect` refreshes.",
67
+ fix: "nomarmy connect claude (and codex, cursor), then restart those sessions", short: "nomarmy reconnect" });
68
+ }
69
+ return issues;
70
+ }
71
+
72
+ /**
73
+ * For the running server: a notice when its copy on disk changed after it
74
+ * started, so this session still runs the old code. Null when current.
75
+ */
76
+ export function restartNotice({ serverFile, startedAtMs, runningVersion, stat = fs.statSync, readVersion = readPackageVersion }) {
77
+ let changedMs;
78
+ try { changedMs = stat(serverFile).mtimeMs; } catch { return null; }
79
+ if (!(changedMs > startedAtMs)) return null;
80
+ const onDisk = readVersion(path.join(path.dirname(serverFile), ".."));
81
+ const versions = onDisk && onDisk !== runningVersion ? ` (this session runs ${runningVersion}; ${onDisk} is installed)` : "";
82
+ return `nomArmy was updated after this session started${versions}. Restart this session to use the new version; until then it runs the old code.`;
83
+ }
@@ -124,3 +124,27 @@ export function formatUnion(union) {
124
124
  const artifacts = union.worktree ? `\n\nUnion artifacts: ${path.dirname(union.worktree)}\nWorktree retained for review: ${union.worktree}\nBranch retained for review: ${union.branch}` : "";
125
125
  return `${banner}--- UNION RECORD ---\n${JSON.stringify(union, null, 2)}${artifacts}`;
126
126
  }
127
+
128
+ /**
129
+ * What a person calls a job: its commit subject, else the task's first
130
+ * sentence, capped. Two jobs on one agent and model differ here even with no
131
+ * role, so a notification can tell them apart.
132
+ */
133
+ export function jobLabel(args) {
134
+ const text = String(args?.commit_subject || args?.task || "").split("\n")[0].split(/(?<=\.)\s/)[0].trim();
135
+ return text.length > 60 ? `${text.slice(0, 57)}...` : text || null;
136
+ }
137
+
138
+ /**
139
+ * Seconds a job has run: to its finish once it has one, not to whenever it's
140
+ * asked about. The record's total_elapsed covers the whole job; an implement
141
+ * job's finishedAt marks only the worker's end, before verification and commit.
142
+ */
143
+ export function jobElapsedSeconds({ status = null, meta = null, entry = null, now = Date.now() } = {}) {
144
+ const total = meta?.metrics?.total_elapsed;
145
+ if (Number.isFinite(total)) return Math.round(total / 1000);
146
+ const startedMs = Date.parse(status?.startedAt ?? entry?.startedAt ?? meta?.startedAt ?? "");
147
+ if (!Number.isFinite(startedMs)) return null;
148
+ const finishedMs = Date.parse((status?.state === "finished" ? status.updatedAt : null) ?? meta?.finishedAt ?? "");
149
+ return Math.round(((Number.isFinite(finishedMs) ? finishedMs : now) - startedMs) / 1000);
150
+ }
package/lib/outcome.mjs CHANGED
@@ -7,6 +7,16 @@ import { parseWorkerReport } from "./report.mjs";
7
7
  // failed. Recovery exists so that a mangled REPORT cannot destroy correct WORK.
8
8
  // It does not exist to launder a failure into a success.
9
9
  // ---------------------------------------------------------------------------
10
+ // A failed verification's own detail (the command, its exit code, the end of
11
+ // its output), capped so an issue line stays readable; the full output is in
12
+ // the job's verification.log.
13
+ const FAILURE_DETAIL_CHARS = 600;
14
+ function failureDetail(independentVerification) {
15
+ const detail = String(independentVerification?.detail ?? independentVerification?.reason ?? "").trim();
16
+ if (!detail) return "";
17
+ return `: ${detail.length > FAILURE_DETAIL_CHARS ? `${detail.slice(0, FAILURE_DETAIL_CHARS)}...` : detail}`;
18
+ }
19
+
10
20
  export function resolveOutcome({ report, repositoryChanged = false, independentVerification = null, regressionCheck = null, workerFailed = false, workerTimedOut = false, mode = "implement" }) {
11
21
  const verification = independentVerification?.status ?? "not_run";
12
22
  const parsed = report ?? parseWorkerReport("");
@@ -34,7 +44,7 @@ export function resolveOutcome({ report, repositoryChanged = false, independentV
34
44
  if (verification === "fail") {
35
45
  return { ...base, outcome: OUTCOMES.NEEDS_REVIEW, reviewRequired: true,
36
46
  commitBlockedReason: "independent verification failed despite a clean done/pass report",
37
- reasons: ["worker claimed done/pass but independent verification failed"] };
47
+ reasons: [`worker claimed done/pass but independent verification failed${failureDetail(independentVerification)}`] };
38
48
  }
39
49
  // verify_regression: reverting just the production files and re-running
40
50
  // the SAME verification profile still passed (or came back genuinely
@@ -92,7 +102,7 @@ export function resolveOutcome({ report, repositoryChanged = false, independentV
92
102
  if (verification === "fail") {
93
103
  return { ...recovery, outcome: OUTCOMES.WORKER_REPORT_INVALID,
94
104
  commitBlockedReason: "independent verification failed; recovery cannot promote a failure",
95
- reasons: [...recovery.reasons, "independent verification FAILED"] };
105
+ reasons: [...recovery.reasons, `independent verification FAILED${failureDetail(independentVerification)}`] };
96
106
  }
97
107
  if (verification === "pass") {
98
108
  // A leniently recovered `done` plus a passing independent check is the
package/lib/schema.mjs CHANGED
@@ -12,6 +12,7 @@
12
12
  // approval before the job runs. See `collectElevated`.
13
13
 
14
14
  import { z } from "zod";
15
+ import { typeError, isMissingField, isUnknownDiscriminator } from "./zod-issues.mjs";
15
16
  import { armySchema } from "./army.mjs";
16
17
 
17
18
  export const SERVICE_SOURCES = Object.freeze([
@@ -57,12 +58,12 @@ export const SERVICE_FIELDS = Object.freeze({
57
58
  // the dotted path, so repeating it reads as stutter.
58
59
  const requiredString = () =>
59
60
  z
60
- .string({ required_error: "is required", invalid_type_error: "must be a string" })
61
+ .string({ error: typeError("a string") })
61
62
  .refine((value) => value.trim().length > 0, { message: "must not be empty" });
62
63
 
63
64
  const enumOf = (values) =>
64
65
  z.enum(values, {
65
- errorMap: () => ({ message: `must be one of ${values.join(", ")}` }),
66
+ error: `must be one of ${values.join(", ")}`,
66
67
  });
67
68
 
68
69
  // A plain hostname: no scheme, no path, no port, no whitespace.
@@ -90,7 +91,7 @@ export function hostnameProblem(value) {
90
91
  }
91
92
 
92
93
  const hostnameSchema = z
93
- .string({ invalid_type_error: "must be a string hostname" })
94
+ .string({ error: typeError("a string hostname") })
94
95
  .superRefine((value, ctx) => {
95
96
  const problem = hostnameProblem(value);
96
97
  if (problem) ctx.addIssue({ code: z.ZodIssueCode.custom, message: problem });
@@ -162,7 +163,7 @@ export const serviceSchema = z.discriminatedUnion("source", [
162
163
  export const pythonEnvironmentSchema = z
163
164
  .object({
164
165
  requirements: z
165
- .array(requiredString(), { invalid_type_error: "must be an array of strings" })
166
+ .array(requiredString(), { error: typeError("an array of strings") })
166
167
  .min(1, "must list at least one requirements file"),
167
168
  })
168
169
  .strict();
@@ -211,10 +212,7 @@ export const verificationProfileSchema = z
211
212
  .object({
212
213
  environment: enumOf(ENVIRONMENT_LEVELS).default("none"),
213
214
  commands: z
214
- .array(requiredString(), {
215
- required_error: "is required",
216
- invalid_type_error: "must be an array of strings",
217
- })
215
+ .array(requiredString(), { error: typeError("an array of strings") })
218
216
  .min(1, "must list at least one command"),
219
217
  })
220
218
  .strict();
@@ -235,7 +233,7 @@ export const environmentRetentionSchema = z
235
233
  debug: enumOf(RETENTION_ACTIONS).default(DEFAULT_RETENTION.debug),
236
234
  })
237
235
  .strict()
238
- .default({});
236
+ .prefault({});
239
237
 
240
238
  // ---------------------------------------------------------------------------
241
239
  // root
@@ -295,11 +293,11 @@ export function formatIssues(error) {
295
293
  lines.push(`${where}: unexpected field(s) ${keys}`);
296
294
  continue;
297
295
  }
298
- if (issue.code === "invalid_union_discriminator") {
296
+ if (isUnknownDiscriminator(issue)) {
299
297
  lines.push(`${where}: must be one of ${SERVICE_SOURCES.join(", ")}`);
300
298
  continue;
301
299
  }
302
- if (issue.code === "invalid_type" && issue.received === "undefined") {
300
+ if (isMissingField(issue)) {
303
301
  lines.push(`${where}: is required`);
304
302
  continue;
305
303
  }
@@ -0,0 +1,15 @@
1
+ // Zod 4 helpers shared by nomArmy's config schemas and their issue formatters.
2
+
3
+ /**
4
+ * The `error` option for a primitive schema: "is required" when the field is
5
+ * absent, "must be <expected>" when it has the wrong type. Zod 4 replaced
6
+ * zod 3's required_error and invalid_type_error with this single function.
7
+ * @param {string} expected e.g. "a string"
8
+ */
9
+ export const typeError = (expected) => (issue) => (issue.input === undefined ? "is required" : `must be ${expected}`);
10
+
11
+ /** A field that is absent entirely, with no custom message of its own. */
12
+ export const isMissingField = (issue) => issue.code === "invalid_type" && / received undefined$/.test(issue.message);
13
+
14
+ /** A discriminated union whose discriminator matched none of its options. */
15
+ export const isUnknownDiscriminator = (issue) => issue.code === "invalid_union" && issue.note === "No matching discriminator";
package/mcp/server.mjs CHANGED
@@ -35,8 +35,9 @@ import { OUTCOMES, COORDINATOR_STATUS_BY_OUTCOME } from "../lib/outcomes.mjs";
35
35
  import { readUsageSnapshots, usageStatus } from "../lib/usage-limits.mjs";
36
36
  import { modelRefusals } from "../lib/health.mjs";
37
37
  import { podmanProblem, podmanVmStartedAt } from "../lib/podman-health.mjs";
38
+ import { restartNotice } from "../lib/install-freshness.mjs";
38
39
  import { createBuildMetrics, resolveOutcome, finalText, workerMetadata, usageMetrics, policyAdmissionProblems, applyRefactorContract, applyVerificationPolicy, resolveVerifyRegression } from "../lib/outcome.mjs";
39
- import { compactJobRecord, formatResult, formatUnion, testChangeBanner, regressionCheckBanner, decomposeOverlapBanner } from "../lib/job-format.mjs";
40
+ import { jobLabel, compactJobRecord, formatResult, formatUnion, testChangeBanner, regressionCheckBanner, decomposeOverlapBanner } from "../lib/job-format.mjs";
40
41
 
41
42
  export { run, mapLimit };
42
43
  export { readsMeasurable, measureReads };
@@ -57,6 +58,7 @@ export { TEST_PATH_PATTERNS, isTestPath, testPatternFor, classifyTestChanges, de
57
58
  // package.json to the same relative location next to the installed
58
59
  // mcp/server.mjs, so this resolves identically in a dev checkout or an
59
60
  // installed copy.
61
+ const SERVER_STARTED_MS = Date.now();
60
62
  const VERSION = JSON.parse(fs.readFileSync(path.join(path.dirname(fileURLToPath(import.meta.url)), "..", "package.json"), "utf8")).version;
61
63
  // Sent to every coordinator on connect, so no project needs a copied CLAUDE.md.
62
64
  const server = new McpServer({ name: "nomarmy-local-worker", version: VERSION }, { instructions: COORDINATOR_INSTRUCTIONS });
@@ -406,9 +408,20 @@ server.tool("local_worker_status", `Status of one job started by this server: ph
406
408
  if (full && files.meta) return toolText(JSON.stringify(files.meta, null, 2), summary.coordinatorStatus !== "complete");
407
409
  return toolText(JSON.stringify({ ...summary, jobDir, hint: entry?.result || files.meta ? "call again with full=true for the complete report" : null }, null, 2), summary.state === "orphaned" || summary.state === "failed");
408
410
  });
411
+ // Set when this session's copy of nomArmy changed on disk after it started
412
+ // (nomarmy connect or update ran): shown first in army and capacity, and
413
+ // notified once, since only a restart of this session picks it up.
414
+ let restartNotified = false;
415
+ function currentRestartNotice() {
416
+ const notice = restartNotice({ serverFile: fileURLToPath(import.meta.url), startedAtMs: SERVER_STARTED_MS, runningVersion: VERSION });
417
+ if (notice && !restartNotified) { restartNotified = true; try { notify("nomArmy: restart this session", notice); } catch { /* best-effort */ } }
418
+ return notice;
419
+ }
420
+ const withRestartNotice = (value) => { const notice = currentRestartNotice(); return notice ? { restartNeeded: notice, ...value } : value; };
421
+
409
422
  server.tool("local_worker_capacity", "What this host can take right now: context per nom and the brief/report budgets derived from it, memory pressure and whether another job would be admitted, and the jobs currently running. Read-only.", {}, async () => {
410
423
  await budgetState.refresh();
411
- return toolText(JSON.stringify(capacitySnapshot(), null, 2));
424
+ return toolText(JSON.stringify(withRestartNotice(capacitySnapshot()), null, 2));
412
425
  });
413
426
  // The only way to know what `verification`/`union_verification`/
414
427
  // `verify_regression` profile names are actually valid for this repo used to
@@ -515,7 +528,7 @@ server.tool("army", "Who you, the General, are and who you call for what in this
515
528
  role.modelNote = `${role.model} isn't in OpenClaw's catalog for ${role.agent}; \`army assign\` checked it with a real test call when it was set, and the catalog can lag new models. Use it as assigned; if a job reports "Unknown model", reassign.`;
516
529
  }
517
530
  }
518
- return toolText(JSON.stringify(summary, null, 2));
531
+ return toolText(JSON.stringify(withRestartNotice(summary), null, 2));
519
532
  } catch (error) {
520
533
  return toolText(error.message, true);
521
534
  }
@@ -569,7 +582,7 @@ server.tool("local_workers", "Run independent jobs (implement or scout) with bou
569
582
  // so it was invisible to both ceilings while it ran.
570
583
  // A batch job waits for its agent's slot (up to its own timeout) rather
571
584
  // than failing because an earlier job in the same batch holds it.
572
- return trackInRun(j, track(jobId, { mode: j.mode, workerId, lane: jobLane(j), agent: j.agentName ?? null, runId: j.run_id ?? null, role: j.armyRole ?? null, model: j.model ?? null },
585
+ return trackInRun(j, track(jobId, { mode: j.mode, workerId, lane: jobLane(j), agent: j.agentName ?? null, runId: j.run_id ?? null, role: j.armyRole ?? null, model: j.model ?? null, label: jobLabel(j) },
573
586
  withAgentSlot(j, jobId, () => executeJob({ ...jobArgs(effectiveJob, workerId), jobId }), { waitMs: (j.timeout_seconds ?? 600) * 1000 }))).promise;
574
587
  }, { staggerMs: WORKER_START_STAGGER_MS });
575
588
  indices.forEach((i, laneI) => { results[i] = laneResults[laneI]; });
package/package.json CHANGED
@@ -1,9 +1,9 @@
1
1
  {
2
2
  "name": "nomarmy",
3
- "description": "A harness for AI coding workers whose claims are never trusted: your coding assistant stays in charge while workers implement and test in sandboxes, on local models, API keys or your own subscriptions.",
3
+ "description": "Every byte verified: a harness for AI coding workers whose claims are never trusted. Your coding assistant stays in charge while workers implement and test in sandboxes, and nomArmy checks every change before it is committed.",
4
4
  "author": "Rayson Technologies",
5
5
  "license": "Apache-2.0",
6
- "version": "0.1.0-alpha.6",
6
+ "version": "0.1.0-alpha.8",
7
7
  "private": false,
8
8
  "type": "module",
9
9
  "engines": {
@@ -21,11 +21,11 @@
21
21
  "tag": "alpha"
22
22
  },
23
23
  "dependencies": {
24
- "@modelcontextprotocol/sdk": "^1.0.0",
24
+ "@modelcontextprotocol/sdk": "^1.30.1",
25
25
  "@secretlint/node": "^13.0.5",
26
26
  "@secretlint/secretlint-rule-preset-recommend": "^13.0.5",
27
27
  "yaml": "^2.5.0",
28
- "zod": "^3.24.0"
28
+ "zod": "^4.6.5"
29
29
  },
30
30
  "scripts": {
31
31
  "test": "node --test tests/*.test.mjs",