nomarmy 0.1.0-alpha.7 → 0.1.0-alpha.9

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -6,20 +6,22 @@
6
6
  <a href="https://github.com/rayson-tech/nomarmy/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-Apache%202.0-blue.svg" alt="License: Apache 2.0"></a>
7
7
  </p>
8
8
 
9
- <p align="center"><em>Tiny coders, big appetites for bounded tickets.</em> 🍪</p>
9
+ <p align="center"><em>Every byte verified.</em> 🍪</p>
10
10
 
11
11
  **Your coding assistant plans; sandboxed workers build; nothing counts until nomArmy has checked it.**
12
12
 
13
+ AI coding workers are confident. Their "done, all tests pass" is a claim, not evidence. nomArmy lets your coding assistant (Claude Code, Codex or Cursor) hand work to workers called **noms**, then checks every change itself before anything is committed: the real diff, your tests run in a fresh sandbox, a check that those tests actually catch the change, and a secret scan.
14
+
13
15
  ## TL;DR
14
16
 
15
- 1. **Have** Git, Node 24.16+ (or 26.1+; OpenClaw, which nomArmy installs, needs it) and [Podman](https://podman.io) (on macOS: `brew install podman && podman machine init --memory 8192 && podman machine start`; Podman's 2 GiB default is too small for nomArmy's sandboxes).
17
+ 1. **Have** Git, Node 24.16+ (or 26.1+) and [Podman](https://podman.io). On macOS, give Podman 8 GiB: `brew install podman && podman machine init --memory 8192 && podman machine start`.
16
18
  2. **Install and set up:**
17
19
  ```bash
18
20
  npm install -g nomarmy@alpha
19
21
  cd your-project
20
22
  nomarmy setup
21
23
  ```
22
- `nomarmy setup` is the playbook. It shows a checklist and runs the next step each time you say yes:
24
+ `nomarmy setup` is a playbook. It shows a checklist and runs the next step each time you say yes:
23
25
  ```text
24
26
  ✓ Where models run: hosted
25
27
  ✓ Installed: OpenClaw 2026.9.6
@@ -29,34 +31,39 @@
29
31
  Check: verify the installation
30
32
  Run `nomarmy agents add` now? [Y/n]
31
33
  ```
32
- In order: pick where models run (API keys and subscriptions for most people), install OpenClaw and the sandbox, add your agents (an API key, or your ChatGPT or Muse Code subscription), put the roles on them, write this repo's `.nomarmy.yml`, then check it all. Stop anytime; `nomarmy setup` picks up where you left off.
33
-
34
- **Want every step spelled out?** [Example setup: Claude Code, Codex and an API key](https://github.com/rayson-tech/nomarmy/blob/main/docs/setup/example.md) walks through a complete setup, command by command.
34
+ Stop anytime; `nomarmy setup` picks up where you left off. **Want every step spelled out?** [Example setup: Claude Code, Codex and an API key](https://github.com/rayson-tech/nomarmy/blob/main/docs/setup/example.md) goes command by command.
35
35
  3. **Use it:** restart Claude Code in the project and ask it to use nomArmy for one small bug that has a test. When that works, try `/feature <what you want built>`.
36
36
 
37
- **Have a GPU or a Mac with plenty of memory?** Choose "a local model" in `nomarmy setup` and workers run on llama.cpp on your own machine: no per-token bill and your code stays home, but you pay in hardware, power and speed, and a model too big for your memory crawls. `nomarmy sizing` tells you what fits; see [Install](https://github.com/rayson-tech/nomarmy/blob/main/docs/install.md#install). A team GPU server works too: [a shared model server](https://github.com/rayson-tech/nomarmy/blob/main/docs/install.md#a-shared-model-server). Codex or Cursor as the coordinator: `nomarmy connect codex cursor`.
37
+ Stuck? `nomarmy doctor` checks the machine and `nomarmy health` checks everything nomArmy runs on. Upgrading later? `nomarmy update`.
38
38
 
39
- Stuck? `nomarmy doctor` checks the machine, and `nomarmy health` checks everything nomArmy runs on.
39
+ ## How every byte gets verified
40
40
 
41
- ## What it is
41
+ 1. Your coding assistant, the **General**, briefs a job: a task, acceptance criteria, and the tests that prove it.
42
+ 2. nomArmy creates a git worktree from your branch and runs the nom in a Podman sandbox with no network and no host credentials. (One exception, the Claude subscription: see [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md#security-posture).)
43
+ 3. The nom edits, runs tests, and ends with a four-line report: `STATUS`, `TESTS`, `NOT_DONE`, `NOTE`.
44
+ 4. nomArmy treats that report as a claim and checks the evidence itself:
45
+ - reads the real diff from git, not the nom's description of it
46
+ - runs your verification profile in a fresh sandbox
47
+ - reverts the production change and reruns the tests: a test that still passes proves nothing, so the job goes to review instead of being committed
48
+ - blocks on secrets, and flags tests made to pass (new skips, stubbed imports) and code nothing calls
49
+ 5. Only then does it commit, on the nom's own branch. It never merges into yours: reviewing and integrating stay with the General, and with you.
42
50
 
43
- Your coding assistant (Claude Code, Codex or Cursor) stays in charge as the **General**: it decides what gets built and whether the result is acceptable. The work goes to **noms**, workers that implement, test and repair in their own git worktree and sandbox, on an API key, your own ChatGPT or Muse Code subscription, or a local model. nomArmy owns everything in between: worktrees, git, sandboxes, verification, and the evidence that decides whether work is accepted.
51
+ Failing verification stays failed, unconditionally. A malformed report isn't automatically a failure: if the repository changed, nomArmy verifies independently and may recover the work. And the checks aren't the General's to waive: a repo's `.nomarmy.yml` policy (on by default for new repos) makes verification and the revert check mandatory for every job.
44
52
 
45
- **What you get is work you don't have to take on faith**, not cheaper work. Delegating costs the General tokens too: briefing and reviewing. On small, already-diagnosed tickets we measured 4 to 8 times more of the General's tokens than fixing the bug directly, and break-even at roughly 150 lines of context a fix needs to read ([the measurements](https://github.com/rayson-tech/nomarmy/blob/main/docs/experiments/2026-09-20-model-bakeoff-and-economics.md)). It pays off on bigger tickets, on parallel work, and anywhere you'd otherwise have to trust an agent's say-so.
53
+ **Checking without building** costs nothing: `mode: verify` runs a verification profile against any branch, with no worker and no model tokens.
46
54
 
47
- Developed and maintained by Rayson Technologies. This is an alpha (`0.1.0-alpha`).
55
+ ## Where the work runs
48
56
 
49
- ## How it works
57
+ - **Agents** say where a job can run: an API key, your own ChatGPT or Muse Code subscription, or a local model on llama.cpp.
58
+ - **The army** says which role runs on which agent: Sr and Jr devs build, a security analyst and a data architect review, a PM checks the plan, a PO accepts.
59
+ - **`/feature`** runs a whole feature end to end, from plan through build, review and acceptance, and hands you a branch to merge.
60
+ - **Harnesses** give each repo the right sandbox: Go, Rust, Python and Node (mixed repos too), Playwright browser tests, and fake services like a mock login server, all offline. [Adding one](https://github.com/rayson-tech/nomarmy/blob/main/CONTRIBUTING.md#adding-a-harness) never touches core code.
50
61
 
51
- 1. The General briefs a job: a task, acceptance criteria, the tests that prove it.
52
- 2. nomArmy creates a worktree from your branch and runs the worker in a Podman sandbox with no network and no host credentials. (One exception, the Claude subscription: see [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md#security-posture).)
53
- 3. The worker edits, runs tests, and ends with a four-line report: `STATUS`, `TESTS`, `NOT_DONE`, `NOTE`.
54
- 4. nomArmy treats that report as a claim. It reads the real diff from git, runs your verification profile itself in a fresh sandbox, reverts the production change to check the tests actually fail without it, and scans for secrets.
55
- 5. Only then does it commit, on the worker's own branch. It never merges into yours: reviewing and integrating stay with the General, and with you.
62
+ **What you get is work you don't have to take on faith**, not cheaper work. Delegating costs the General tokens too, for briefing and review: on small, already-diagnosed tickets we measured 4 to 8 times more of the General's tokens than fixing the bug directly, with break-even around 150 lines of context a fix needs to read ([the measurements](https://github.com/rayson-tech/nomarmy/blob/main/docs/experiments/2026-09-20-model-bakeoff-and-economics.md)). It pays off on bigger tickets, parallel work, and anywhere you'd otherwise trust an agent's say-so.
56
63
 
57
- A malformed report isn't automatically a failure: if the repository changed, nomArmy verifies independently and may recover the work. Failing verification stays failed, unconditionally. And the checks aren't the General's to waive: a repo's `.nomarmy.yml` policy (on by default for new repos) makes verification and the revert check mandatory for every job.
64
+ **Have a GPU or a Mac with plenty of memory?** Choose "a local model" in `nomarmy setup`: no per-token bill and your code stays home, but you pay in hardware, power and speed. `nomarmy sizing` tells you what fits. A [shared model server](https://github.com/rayson-tech/nomarmy/blob/main/docs/install.md#a-shared-model-server) works too.
58
65
 
59
- Around that core: **agents** say where a job can run, the **army** says which role runs on which agent, and **`/feature`** runs a whole feature end to end, from plan through build, review and acceptance, handing you a branch to merge.
66
+ Developed and maintained by Rayson Technologies. This is an alpha (`0.1.0-alpha`).
60
67
 
61
68
  ## Docs
62
69
 
@@ -66,7 +73,7 @@ Around that core: **agents** say where a job can run, the **army** says which ro
66
73
  | [Example setup](https://github.com/rayson-tech/nomarmy/blob/main/docs/setup/example.md) | Claude Code, Codex and an API key, command by command |
67
74
  | [Agents and the army](https://github.com/rayson-tech/nomarmy/blob/main/docs/agents-and-army.md) | Where a job can run, who does what, usage limits, picking an agent |
68
75
  | [`/feature` runs](https://github.com/rayson-tech/nomarmy/blob/main/docs/feature-runs.md) | A feature end to end, and watching what nomArmy is doing |
69
- | [Your repository](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md) | `.nomarmy.yml`, verification, languages and dependencies, what nomArmy checks |
76
+ | [Your repository](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md) | `.nomarmy.yml`, verification, dependencies, private registries, what nomArmy checks |
70
77
  | [Harnesses](https://github.com/rayson-tech/nomarmy/blob/main/docs/harnesses.md) | Ecosystem registry, detection, network levels, and requirements |
71
78
  | [Configuration](https://github.com/rayson-tech/nomarmy/blob/main/docs/configuration.md) | Settings, swapping the local model, sizing, admission |
72
79
  | [Reference](https://github.com/rayson-tech/nomarmy/blob/main/docs/reference.md) | Every CLI command and MCP tool |
@@ -75,23 +82,20 @@ Around that core: **agents** say where a job can run, the **army** says which ro
75
82
 
76
83
  ## Security
77
84
 
78
- A worker gets a writable git worktree inside a Podman sandbox and nothing else: no network, no host credentials, no Podman socket. Every model call is made by OpenClaw on your machine, never from inside the sandbox. **The exception is a Claude subscription**, whose tools run on your machine, so nomArmy refuses build jobs on it unless you allow it. Never hand a worker production credentials, deployment access or SSH keys. Details: [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md); to report a vulnerability, [SECURITY.md](https://github.com/rayson-tech/nomarmy/blob/main/SECURITY.md).
85
+ A nom gets a writable git worktree inside a Podman sandbox and nothing else: no network, no host credentials, no Podman socket. Every model call is made by OpenClaw on your machine, never from inside the sandbox. Verification can climb a network ladder one rung at a time (fake services on a private network, then an allowlist you approve for a test tenant), but noms never leave `network none`. **The exception is a Claude subscription**, whose tools run on your machine, so nomArmy refuses build jobs on it unless you allow it. Never hand a nom production credentials, deployment access or SSH keys. Details: [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md); to report a vulnerability, [SECURITY.md](https://github.com/rayson-tech/nomarmy/blob/main/SECURITY.md).
79
86
 
80
87
  ## Status
81
88
 
82
89
  | Capability | Status |
83
90
  |---|---|
84
- | Delegation core: worktrees, nomArmy-owned git, independent verification, kept failed worktrees | Working, end-to-end tested |
85
- | Local (llama.cpp) and Bedrock profiles | Working |
86
- | Scout and decompose modes | Unit and live tested |
87
- | `auto_union`, `verify_regression`, test-selection and unwired-code checks | Unit and live tested; the heuristics are review flags |
88
- | Secret scanning (secretlint, hard block) | Unit tested against the real dependency |
89
- | Agents: api keys | Live-verified with xAI; other providers built to OpenClaw's documented interface |
91
+ | Verification core: worktrees, nomArmy-owned git, independent verification, the revert check, kept failed worktrees | Working, end-to-end tested |
92
+ | `mode: verify`, secret scanning (secretlint, hard block), test-workaround and unwired-code checks | Unit and live tested; the heuristics are review flags |
93
+ | Scout and decompose modes, `auto_union` | Unit and live tested |
94
+ | Agents: API keys | Live-verified with xAI; other providers built to OpenClaw's documented interface |
90
95
  | Agents: subscriptions | ChatGPT (Codex) and Muse Code sandboxed and live-verified; Claude live-verified, but its tools run on the host (scout and review by default) |
96
+ | Local (llama.cpp) and Bedrock profiles | Working |
91
97
  | The army and `/feature` | Driven by a real Claude Code General across three runs, about 18 implement jobs |
92
- | Go, Rust, Python and Node repos, and mixed ones | Harness images live-verified: Go modules and Rust crates prefetched; npm, pnpm, yarn, bun and workspaces; pip, pyproject, uv and poetry |
93
- | Fake services beside the app (mock login server, mock APIs) | Live-verified on a private network with no route out (the `services` harness level) |
94
- | Browser tests (Playwright + Chromium) | Live-verified offline, with screenshots and traces kept as job evidence |
98
+ | Harnesses: Go, Rust, Python, Node and mixed repos; Playwright; fake services | Live-verified offline |
95
99
  | Private registries and a verification-only network allowlist | Live-verified; each passed an independent security review |
96
100
 
97
101
  What we've learned from real runs, including where delegating pays and where it doesn't, is in [docs/findings.md](https://github.com/rayson-tech/nomarmy/blob/main/docs/findings.md).
@@ -99,13 +103,13 @@ What we've learned from real runs, including where delegating pays and where it
99
103
  ### Known limitations
100
104
 
101
105
  - **A Claude subscription isn't sandboxed.** Its tools run on your machine, so implement jobs on it are refused unless you set `allow_host_tools: true`. See [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md#security-posture).
102
- - **A refused model costs one job.** When a vendor refuses a model at run time that OpenClaw lists (gpt-6-sol on a ChatGPT plan), the first job on it fails with `model_not_found`. After that nomArmy refuses to dispatch it until a job or test call on it works. `army assign` tests the job's route and catches this before any job.
103
- - **Claude subscription token counts** come from the Claude CLI's own session log, since OpenClaw sees only the final reply. Totals include cache reads and writes, which make up most of an agent's prompt; each part is also kept separately.
106
+ - **Your own compose services aren't started yet.** A verification profile that needs a real database from your compose file (`environment: basic` or higher) reports `not_run` rather than running without it (and the compose parser doesn't resolve YAML anchors). Fake services from harnesses, like the mock login server, do run.
107
+ - **Private registries don't cover Poetry or Yarn Berry** yet: Poetry can't guarantee a credentialed install runs no package code, and Yarn Berry doesn't read `.npmrc`. uv, pip wheels, npm, pnpm, Yarn Classic and bun work. See [Private registries](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md#private-registries).
108
+ - **A refused model costs one job.** When a vendor refuses a model at run time that OpenClaw lists (gpt-6-sol on a ChatGPT plan), the first job on it fails with `model_not_found`; after that nomArmy won't dispatch it until a job or test call on it works. `army assign` tests the route and catches this before any job.
109
+ - **Claude subscription token counts** come from the Claude CLI's own session log, since OpenClaw sees only the final reply; totals include cache reads and writes.
104
110
  - **Test-workaround detection is a flag, not a verdict**: a legitimate new skip still gets flagged.
105
111
  - **Deploy-time failures need your own check.** See [Add a check for what unit tests can't see](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md#nomarmyyml).
106
- - **Node private registries are not supported yet**: npm, pnpm, yarn and bun lockfiles and workspaces are supported, but the image build has no credentials for private registries (dependency plan step 8).
107
- - **Verification needing services** (a database, a mock server) reports `not_run` instead of running without them. The compose parser doesn't resolve YAML anchors.
108
- - **Same-host sandboxes**: the MCP server, OpenClaw and every job's sandbox run on the machine with the coordinator. Only the model can be elsewhere (an agent, or [a shared model server](https://github.com/rayson-tech/nomarmy/blob/main/docs/install.md#a-shared-model-server)).
112
+ - **Same-host sandboxes**: the MCP server, OpenClaw and every job's sandbox run on the machine with the coordinator. Only the model can be elsewhere.
109
113
 
110
114
  ## More
111
115
 
package/bin/nomarmy.mjs CHANGED
@@ -19,7 +19,7 @@ import { detectHardware } from "../lib/hardware.mjs";
19
19
  import { readGGUFMetadata, resolveModelPath, totalSplitBytes } from "../lib/gguf.mjs";
20
20
  import { recommend, customRecommendation, evaluateConfig, bytesPerKvElementForCacheTypes, MIN_CONTEXT_PER_NOM } from "../lib/sizing.mjs";
21
21
  import { connectClaude, connectCodex, connectCursor, cursorAlreadyConnected, deriveWorkerModelEnv, defaultInstallDir } from "../lib/connect.mjs";
22
- import { compareVersions, readPackageVersion, readInstallVersions } from "../lib/install-freshness.mjs";
22
+ import { compareVersions, readPackageVersion, readInstallVersions, copyIsStale } from "../lib/install-freshness.mjs";
23
23
  import { ID_RE, AUTH_ENV_NAME_RE, OPENCLAW_PROVIDER_ID_RE, openclawProviderId, isNativeProviderType } from "../lib/dispatch-schema.mjs";
24
24
  import { loadAgents, readAgentsFile, writeAgentsFile, agentsConfigPath, apiAgentAsPoolEntry, describeAgent as describeAgentLabel, agentRunsToolsOnHost, agentProviderId, AGENT_KINDS, API_PROVIDER_TYPES, RESERVED_AGENT_NAMES, BUILTIN_LOCAL_AGENT } from "../lib/agents.mjs";
25
25
  import { loadArmy, mergeArmy, describeArmy, readArmyFile, updateArmyInFile, assignRoleInFile, parseTargetSpec, armyLayerPath, globalConfigDir, DEFAULT_ARMY, ARMY_PHASES, LOCAL_CONFIG_FILENAME } from "../lib/army.mjs";
@@ -543,14 +543,16 @@ function defaultLocalProfile() {
543
543
  return spawnSync("nvidia-smi", ["-L"], { stdio: "ignore", timeout: 5000 }).status === 0 ? "nvidia-linux" : "cpu-linux";
544
544
  }
545
545
 
546
- function setupChecklist() {
546
+ // Which profile this machine is set up for, the same way for `setup` and
547
+ // `install`: the one setup recorded, else the install marker's, else the
548
+ // execution mode's, else (a working local install) install.sh's default.
549
+ function setupProfileState() {
547
550
  const common = path.join(nomarmyRoot, "config", "common.env");
548
551
  const chosen = readEnvValue(common, "NOMARMY_SETUP_PROFILE");
549
552
  const probeCommand = (binary, args) => {
550
553
  const result = spawnSync(binary, args, { encoding: "utf8", timeout: 10000 });
551
554
  return result.status === 0 ? result.stdout.trim() : "";
552
555
  };
553
- const project = setupProjectDir();
554
556
  const profileFile = chosen ? path.join(nomarmyRoot, "config", "profiles", `${chosen}.env`) : null;
555
557
  const root = (process.env.NOMARMY_INSTALL_ROOT || (profileFile && readEnvValue(profileFile, "NOMARMY_INSTALL_ROOT")) || readEnvValue(common, "NOMARMY_INSTALL_ROOT") || "$HOME/.local/share/nomarmy-local-agents").replace(/\$HOME|\$\{HOME\}/g, os.homedir());
556
558
  let marker = null;
@@ -564,6 +566,12 @@ function setupChecklist() {
564
566
  const profile = chosen ?? marker?.profile
565
567
  ?? (["hosted", "remote", "bedrock"].includes(execution) ? execution : null)
566
568
  ?? (registered ? defaultLocalProfile() : null);
569
+ return { common, profile, marker, version, registered };
570
+ }
571
+
572
+ function setupChecklist() {
573
+ const { common, profile, marker, version, registered } = setupProfileState();
574
+ const project = setupProjectDir();
567
575
  return setupSteps({
568
576
  mode: () => ({ profile, host: readEnvValue(common, "NOMARMY_LLAMA_HOST"), port: readEnvValue(common, "NOMARMY_LLAMA_PORT") }),
569
577
  install: () => ({ marker, version, registered }),
@@ -590,8 +598,9 @@ function runSetupChild(args) {
590
598
  }
591
599
 
592
600
  function cmdInstall() {
593
- const profile = value("profile", readEnvValue(path.join(nomarmyRoot, "config", "common.env"), "NOMARMY_SETUP_PROFILE"));
601
+ const profile = value("profile", null) ?? setupProfileState().profile;
594
602
  if (!profile) throw new Error("Choose a profile first: nomarmy setup --choose (or install --profile <name>).");
603
+ if (!value("profile", null)) console.log(`Installing for profile ${profile} (from nomarmy setup; pass --profile to choose another).`);
595
604
  const result = spawnSync("bash", [path.join(nomarmyRoot, "install.sh"), "--profile", profile, ...(flag("no-claude") ? ["--no-claude"] : [])], { stdio: "inherit", cwd: nomarmyRoot });
596
605
  process.exitCode = result.status ?? 1;
597
606
  }
@@ -1660,8 +1669,18 @@ async function cmdUpdate() {
1660
1669
  const remote = git(["rev-parse", "@{u}"]);
1661
1670
  const base = git(["merge-base", "HEAD", "@{u}"]);
1662
1671
  if (local === remote) {
1663
- if (json) return out({ updated: false, reason: "already up to date" });
1664
- console.log(c.green("✓ Already up to date."));
1672
+ // Nothing to pull, but the copy coordinators run can still be behind
1673
+ // this checkout (commits made or pulled here without a reconnect).
1674
+ if (!copyIsStale(defaultInstallDir(), nomarmyRoot)) {
1675
+ if (json) return out({ updated: false, reason: "already up to date" });
1676
+ console.log(c.green("✓ Already up to date, and your coordinators run this checkout."));
1677
+ return;
1678
+ }
1679
+ say(c.bold("🍪 nomArmy update\n"));
1680
+ say("Nothing to pull, but your coordinators run an older copy of this checkout.");
1681
+ const resynced = reconnectCoordinators();
1682
+ if (json) return out({ updated: false, resynced, sha: local });
1683
+ console.log(c.yellow("\nRestart every open Claude Code, Codex and Cursor session: each keeps the code it started with until then."));
1665
1684
  return;
1666
1685
  }
1667
1686
  if (base !== local) {
@@ -1677,25 +1696,7 @@ async function cmdUpdate() {
1677
1696
  say("\nInstalling dependencies...");
1678
1697
  execFileSync("npm", ["install", "--omit=dev", "--no-audit", "--no-fund"], { cwd: nomarmyRoot, stdio: json ? "ignore" : "inherit" });
1679
1698
 
1680
- const resynced = [];
1681
- const runInherit = (cmd, args, opts = {}) => execFileSync(cmd, args, { stdio: json ? "ignore" : "inherit", ...opts });
1682
- if (commandExists("claude")) {
1683
- say("\nRe-syncing the Claude Code MCP install...");
1684
- connectClaude({ nomarmyRoot, run: runInherit });
1685
- resynced.push("claude");
1686
- }
1687
- if (commandExists("codex")) {
1688
- say("\nRe-syncing the Codex MCP install...");
1689
- connectCodex({ nomarmyRoot, run: runInherit });
1690
- resynced.push("codex");
1691
- }
1692
- // Cursor has no CLI/PATH binary to probe with commandExists -- "already
1693
- // connected" is read from its own config file instead.
1694
- if (cursorAlreadyConnected()) {
1695
- say("\nRe-syncing the Cursor MCP install...");
1696
- connectCursor({ nomarmyRoot, run: runInherit });
1697
- resynced.push("cursor");
1698
- }
1699
+ const resynced = reconnectCoordinators();
1699
1700
 
1700
1701
  if (json) return out({ updated: true, sha: git(["rev-parse", "HEAD"]), resynced });
1701
1702
  console.log(c.yellow("\nThe MCP server is a per-session child process: every open Claude Code / Codex / Cursor session needs a restart to pick this up, not just this one."));
@@ -1703,6 +1704,18 @@ async function cmdUpdate() {
1703
1704
 
1704
1705
  // The coordinators nomArmy is registered with. Cursor has no CLI to probe,
1705
1706
  // so it counts when its own config already lists nomArmy.
1707
+ // Reconnect every connected coordinator through a child process, so it runs
1708
+ // the code now on disk (just pulled or installed) rather than the old code
1709
+ // this process loaded. Returns the targets reconnected.
1710
+ function reconnectCoordinators() {
1711
+ const targets = connectedTargets();
1712
+ if (targets.length) {
1713
+ if (!json) console.log(`\nReconnecting ${targets.join(", ")}...`);
1714
+ execFileSync(process.execPath, [path.join(nomarmyRoot, "bin", "nomarmy.mjs"), "connect", ...targets, ...(json ? ["--json"] : [])], { stdio: json ? "ignore" : "inherit" });
1715
+ }
1716
+ return targets;
1717
+ }
1718
+
1706
1719
  function connectedTargets() {
1707
1720
  return [commandExists("claude") && "claude", commandExists("codex") && "codex", cursorAlreadyConnected() && "cursor"].filter(Boolean);
1708
1721
  }
@@ -1729,13 +1742,7 @@ async function updateFromNpm() {
1729
1742
  if (!json) console.log(c.bold(`🍪 Updating nomArmy ${current} → ${latest}\n`));
1730
1743
  execFileSync("npm", ["install", "-g", `nomarmy@${latest}`, "--no-audit", "--no-fund"], { stdio: json ? "ignore" : "inherit" });
1731
1744
  }
1732
- // A child process, so the reconnect runs the code just installed rather
1733
- // than the old code this process loaded.
1734
- const targets = connectedTargets();
1735
- if (targets.length) {
1736
- if (!json) console.log(`\nReconnecting ${targets.join(", ")}...`);
1737
- execFileSync(process.execPath, [path.join(nomarmyRoot, "bin", "nomarmy.mjs"), "connect", ...targets, ...(json ? ["--json"] : [])], { stdio: json ? "ignore" : "inherit" });
1738
- }
1745
+ const targets = reconnectCoordinators();
1739
1746
  if (json) return out({ updated: upgrade, from: current, version: upgrade ? latest : current, resynced: targets });
1740
1747
  console.log(c.yellow("\nRestart every open Claude Code, Codex and Cursor session: each keeps the code it started with until then."));
1741
1748
  }
package/lib/admission.mjs CHANGED
@@ -1,6 +1,7 @@
1
1
  import fs from "node:fs";
2
2
  import path from "node:path";
3
3
  import { executionMode } from "./execution.mjs";
4
+ import { jobLabel, jobElapsedSeconds } from "./job-format.mjs";
4
5
  import { loadConfig } from "./config.mjs";
5
6
  import { clampInt } from "./budget-state.mjs";
6
7
  import { checkBrief, assessAdmission, describeBudgets } from "./budget.mjs";
@@ -160,7 +161,7 @@ export function createJobRuntime(deps) {
160
161
  const who = entry.mode === "verify" ? "verification runner" : entry.agent ? `${entry.agent}${entry.model ? `/${entry.model}` : ""}` : "local model";
161
162
  const took = Math.round((Date.now() - Date.parse(entry.startedAt)) / 60000);
162
163
  const ok = !error && (result?.ok || m.coordinatorStatus === "complete");
163
- notify(`nomArmy: ${entry.role ?? entry.mode ?? "job"} ${ok ? "done" : outcome}`, `${entry.workerId ?? entry.jobId} on ${who}: ${outcome} after ${took}m. ${ok ? "Ready for the General's review." : "Needs a look."}`);
164
+ notify(`nomArmy: ${entry.role ?? entry.mode ?? "job"} ${ok ? "done" : outcome}`, `${entry.label ? `${entry.label} (${entry.workerId ?? entry.jobId})` : entry.workerId ?? entry.jobId} on ${who}: ${outcome} after ${took}m. ${ok ? "Ready for the General's review." : "Needs a look."}`);
164
165
  }
165
166
  function admissionHardware() {
166
167
  return executionMode(deps.env).managesModelServer ? deps.budgetState.hardwareSnapshot : null;
@@ -349,7 +350,7 @@ export function createJobRuntime(deps) {
349
350
  function launch(args) {
350
351
  const workerId = args.worker_id || null;
351
352
  const jobId = slug(workerId || (args.mode === "scout" ? "scout" : "worker"));
352
- return trackInRun(args, track(jobId, { mode: args.mode, workerId: workerId || jobId, lane: jobLane(args), agent: args.agentName ?? null, runId: args.run_id ?? null, role: args.armyRole ?? null, model: args.model ?? null },
353
+ return trackInRun(args, track(jobId, { mode: args.mode, workerId: workerId || jobId, lane: jobLane(args), agent: args.agentName ?? null, runId: args.run_id ?? null, role: args.armyRole ?? null, model: args.model ?? null, label: jobLabel(args) },
353
354
  withAgentSlot(args, jobId, () => executeJob({ ...jobArgs(args, workerId), jobId }))));
354
355
  }
355
356
  // Best-effort progress signal for a job still mid-run: a plain "phase: worker,
@@ -395,7 +396,7 @@ export function createJobRuntime(deps) {
395
396
 
396
397
  async function summarize(entry, files, jobDir = null) {
397
398
  const status = files.status, meta = files.meta ?? files.failure;
398
- const elapsedSeconds = status?.startedAt ? Math.round((Date.now() - Date.parse(status.startedAt)) / 1000) : entry ? Math.round((Date.now() - Date.parse(entry.startedAt)) / 1000) : null;
399
+ const elapsedSeconds = jobElapsedSeconds({ status, meta, entry });
399
400
  const out = { jobId: entry?.jobId ?? status?.jobId ?? meta?.jobId ?? null, workerId: entry?.workerId ?? status?.workerId ?? meta?.workerId ?? null,
400
401
  mode: entry?.mode ?? status?.mode ?? meta?.mode ?? null, state: null, phase: status?.phase ?? "starting", elapsedSeconds,
401
402
  timeoutSeconds: status?.timeoutSeconds ?? null, coordinatorStatus: meta?.coordinatorStatus ?? null, outcome: meta?.outcome ?? null,
package/lib/connect.mjs CHANGED
@@ -171,6 +171,14 @@ export function installMcpCopy({ nomarmyRoot, installDir, run = defaultRun }) {
171
171
  fs.rmSync(path.join(installDir, "config"), { recursive: true, force: true });
172
172
  fs.cpSync(configSrc, path.join(installDir, "config"), { recursive: true });
173
173
  }
174
+ // lib/harnesses.mjs finds its definitions at ../harnesses/, so the copy
175
+ // needs them too: without this, every connected session matched no
176
+ // harness and built its jobs in the plain base image, silently.
177
+ const harnessSrc = path.join(nomarmyRoot, "harnesses");
178
+ if (fs.existsSync(harnessSrc)) {
179
+ fs.rmSync(path.join(installDir, "harnesses"), { recursive: true, force: true });
180
+ fs.cpSync(harnessSrc, path.join(installDir, "harnesses"), { recursive: true });
181
+ }
174
182
  run("npm", ["install", "--omit=dev"], { cwd: installDir });
175
183
  run(process.execPath, ["--check", "mcp/server.mjs"], { cwd: installDir });
176
184
  }
@@ -1,3 +1,5 @@
1
+ import path from "node:path";
2
+
1
3
  // ---------------------------------------------------------------------------
2
4
  // Test-change classification (plan 16). One tunable constant, on purpose:
3
5
  // every heuristic about what counts as a test file lives here and nowhere else.
@@ -416,3 +418,91 @@ export function mergeUntrackedIntoNameStatus(nameStatus, untrackedFiles) {
416
418
  return [...(nameStatus ?? []), ...extra];
417
419
  }
418
420
 
421
+
422
+ // ---------------------------------------------------------------------------
423
+ // Verification inputs: .nomarmy.yml is read from the operator's checkout so a
424
+ // worker can't weaken its own checks, but the files its commands run come
425
+ // from the worker's worktree. Found live: a worker whose verification ran
426
+ // `node check.js` rewrote check.js to print PASS and exit 0, and the job
427
+ // committed as done. The revert check can't see it: reverting check.js along
428
+ // with the code makes verification fail, which reads as coverage.
429
+ //
430
+ // Blocking (the diff changes what a command runs): a changed non-test file a
431
+ // command names (`node check.js`, `bash scripts/verify.sh`), a changed
432
+ // Makefile or justfile under `make`/`just`, or a changed package.json script
433
+ // that `npm test`/`pnpm run x`/`yarn x`/`bun run x` calls. Test files named by
434
+ // a command are left to the test-change review, since adding cases is normal.
435
+ // Flagged only: test-runner configuration, which is often a legitimate edit.
436
+ // ---------------------------------------------------------------------------
437
+
438
+ const RUNNER_CONFIG_RE = /(^|\/)(conftest\.py|pytest\.ini|tox\.ini|setup\.cfg|\.coveragerc|(jest|vitest|vite|playwright|karma|cypress)\.config\.[cm]?[jt]s|\.mocharc(\.[a-z]+)?|phpunit\.xml(\.dist)?|\.rspec)$/;
439
+ const TASK_FILES = { make: ["Makefile", "makefile", "GNUmakefile"], just: ["justfile", "Justfile", ".justfile"] };
440
+
441
+ function shellWords(segment) {
442
+ return (segment.match(/"[^"]*"|'[^']*'|\S+/g) ?? []).map((w) => w.replace(/^["']|["']$/g, ""));
443
+ }
444
+
445
+ // Which package.json script a package-manager command runs, if any.
446
+ function packageScript(words) {
447
+ const [tool, first, second] = words;
448
+ if (!["npm", "pnpm", "yarn", "bun"].includes(tool) || !first) return null;
449
+ if (["run", "run-script"].includes(first)) return second && !second.startsWith("-") ? second : null;
450
+ if (tool === "npm") return ["test", "t", "tst"].includes(first) ? "test" : ["start", "stop", "restart"].includes(first) ? first : null;
451
+ if (tool === "bun") return null; // `bun test` is bun's own runner, not a script
452
+ return first.startsWith("-") || ["install", "add", "remove", "exec", "dlx", "x"].includes(first) ? null : first;
453
+ }
454
+
455
+ function scriptsOf(text) {
456
+ try { return JSON.parse(text ?? "")?.scripts ?? {}; } catch { return {}; }
457
+ }
458
+
459
+ function pytestSection(text) {
460
+ const match = /^\[tool\.pytest[^\]]*\]\s*$([\s\S]*?)(?=^\[|(?![\s\S]))/m.exec(String(text ?? ""));
461
+ return match ? match[1].trim() : null;
462
+ }
463
+
464
+ /**
465
+ * @param {{ commands: string[], changedFiles: string[], readBase: (file: string) => string|null|Promise<string|null>,
466
+ * readHead: (file: string) => string|null|Promise<string|null>, isTestPathFn?: (file: string) => boolean }} input
467
+ * @returns {Promise<{ blocked: {file: string, command: string, why: string}[], flagged: {file: string, why: string}[] } | null>}
468
+ */
469
+ export async function detectVerificationInputChanges({ commands = [], changedFiles = [], readBase = () => null, readHead = () => null, isTestPathFn = isTestPath } = {}) {
470
+ const changed = new Set(changedFiles.map((f) => path.posix.normalize(f)));
471
+ if (!changed.size) return null;
472
+ const blocked = [], flagged = [], seen = new Set();
473
+ const block = (file, command, why) => { if (!seen.has(file)) { seen.add(file); blocked.push({ file, command, why }); } };
474
+ for (const command of commands) {
475
+ let dir = "";
476
+ for (const segment of String(command ?? "").split(/\s*(?:&&|\|\||;|\|)\s*/)) {
477
+ const all = shellWords(segment);
478
+ let lead = 0;
479
+ while (lead < all.length && /^[A-Za-z_][A-Za-z0-9_]*=/.test(all[lead])) lead++; // FOO=1 cmd ...
480
+ const words = all.slice(lead);
481
+ if (!words.length) continue;
482
+ const at = (p) => path.posix.normalize(path.posix.join(dir, p));
483
+ if (words[0] === "cd" && words[1]) { dir = at(words[1]); continue; }
484
+ for (const word of words) {
485
+ if (word.startsWith("-") || word.includes("$")) continue;
486
+ const file = at(word);
487
+ if (changed.has(file) && !isTestPathFn(file)) block(file, command, `run by \`${command}\``);
488
+ }
489
+ for (const name of TASK_FILES[words[0]] ?? []) {
490
+ const file = at(name);
491
+ if (changed.has(file)) block(file, command, `read by \`${command}\``);
492
+ }
493
+ const script = packageScript(words);
494
+ const manifest = at("package.json");
495
+ if (script && changed.has(manifest)) {
496
+ const before = scriptsOf(await readBase(manifest)), after = scriptsOf(await readHead(manifest));
497
+ const touched = [script, `pre${script}`, `post${script}`].filter((s) => before[s] !== after[s]);
498
+ if (touched.length) block(manifest, command, `its script${touched.length > 1 ? "s" : ""} ${touched.map((s) => `"${s}"`).join(", ")} changed, and \`${command}\` runs it`);
499
+ }
500
+ }
501
+ }
502
+ for (const file of changed) {
503
+ if (seen.has(file)) continue;
504
+ if (RUNNER_CONFIG_RE.test(file)) flagged.push({ file, why: "test-runner configuration" });
505
+ else if (/(^|\/)pyproject\.toml$/.test(file) && pytestSection(await readBase(file)) !== pytestSection(await readHead(file))) flagged.push({ file, why: "its [tool.pytest] settings changed" });
506
+ }
507
+ return blocked.length || flagged.length ? { blocked, flagged } : null;
508
+ }
package/lib/execute.mjs CHANGED
@@ -17,7 +17,7 @@ import { describeRecoveryChanges, reportRecoveryPrompt } from "./worker-prompt.m
17
17
  import { parseWorkerReport } from "./report.mjs";
18
18
  import { OUTCOMES, COORDINATOR_STATUS_BY_OUTCOME } from "./outcomes.mjs";
19
19
  import { resolveOutcome, finalText, workerMetadata, applyRefactorContract, applyVerificationPolicy } from "./outcome.mjs";
20
- import { isTestPath, isDocumentationPath, detectScopedTestSelectionRisk, detectUnwiredNewDefinitions, detectMislabeledTestNames, detectPossibleSecrets } from "./diff-checks.mjs";
20
+ import { isTestPath, isDocumentationPath, detectScopedTestSelectionRisk, detectUnwiredNewDefinitions, detectMislabeledTestNames, detectPossibleSecrets, detectVerificationInputChanges } from "./diff-checks.mjs";
21
21
 
22
22
  // ---------------------------------------------------------------------------
23
23
  // Job status for polling. `status.json` is written at every phase transition
@@ -212,7 +212,7 @@ export function createExecutor(deps) {
212
212
  const recoveryResult = await runOpenClaw({
213
213
  task, acceptance, verification, mode, cwd, baseRef: base.ref, baseSha: base.sha,
214
214
  timeoutSeconds: timeBudget.reportReserveSeconds, runtimeDir, profile, reasoning, pool, subscriptionWorker, onBehalfOf, model, reportSize, jobDir, workerId: workerId || jobId,
215
- overridePrompt: reportRecoveryPrompt({ report: budgetState.budgets.report.implement, changes }), logSuffix: "-recovery",
215
+ overridePrompt: reportRecoveryPrompt({ report: budgetState.budgets.report.implement, changes, task }), logSuffix: "-recovery",
216
216
  });
217
217
  const recoveryText = finalText(recoveryResult);
218
218
  const recoveryValidation = parseWorkerReport(recoveryText);
@@ -250,7 +250,7 @@ export function createExecutor(deps) {
250
250
  // which reads as evidence about work that never happened.
251
251
  independentVerification = normalizeVerification({ status: "not_run", basis: "not-applicable", reason: "the worker changed nothing, so there was none of its work to verify" }, verification ?? null);
252
252
  } else if (verificationFlow.verificationRunner || !reportValidation.valid) {
253
- independentVerification = await runIndependentVerification({ profile: verification ?? null, cwd, jobId, baseSha: base.sha, branch, mode, record: preCommit });
253
+ independentVerification = await runIndependentVerification({ profile: verification ?? null, cwd, jobId, baseSha: base.sha, branch, mode, record: preCommit, logFile: path.join(jobDir, "verification.log") });
254
254
  }
255
255
 
256
256
  // verify_regression: on by default whenever there's a verification
@@ -299,12 +299,19 @@ export function createExecutor(deps) {
299
299
  // Cheap, always-on, additive: never changes commitAllowed/commitBlockedReason
300
300
  // on its own (unlike the regression-check override above), only flags for
301
301
  // review -- see detectScopedTestSelectionRisk's own doc comment for why.
302
- let selectionRisk = null;
302
+ let selectionRisk = null, verificationInputs = null;
303
303
  if (mode === "implement" && verification) {
304
304
  try {
305
305
  const loaded = loadConfig(projectDir); // the operator's contract; see registerVerificationRunner's call
306
306
  const profileCommands = loaded.found ? (loaded.config?.verification?.[verification]?.commands ?? []) : [];
307
307
  selectionRisk = detectScopedTestSelectionRisk({ commands: profileCommands, testChanges: preCommit.testChanges });
308
+ // A diff that changes what those commands run (see
309
+ // detectVerificationInputChanges): applied below, after the others.
310
+ verificationInputs = await detectVerificationInputChanges({
311
+ commands: profileCommands, changedFiles: preCommit.changedFiles ?? [],
312
+ readBase: (file) => gitRaw(["show", `${base.sha}:${file}`], cwd).catch(() => null),
313
+ readHead: (file) => { try { return fs.readFileSync(path.join(cwd, file), "utf8"); } catch { return null; } },
314
+ });
308
315
  } catch { /* a config load failure here is the verification runner's own problem to report, not this check's */ }
309
316
  }
310
317
  const afterSelectionRisk = selectionRisk
@@ -399,7 +406,21 @@ export function createExecutor(deps) {
399
406
  commitBlockedReason: `possible secret detected: ${possibleSecrets.reason}`,
400
407
  reasons: [...afterHostInstalls.reasons, `POSSIBLE SECRET DETECTED: ${possibleSecrets.reason}`] }
401
408
  : afterHostInstalls;
402
- const finalOutcome = applyRefactorContract(applyVerificationPolicy(afterSecrets, independentVerification.status, repoPolicy()),
409
+ // A worker must not be judged by a check it rewrote: a changed script,
410
+ // Makefile or package.json script that a verification command runs
411
+ // blocks the commit; changed test-runner config only asks for review.
412
+ const blockedInputs = verificationInputs?.blocked ?? [], flaggedInputs = verificationInputs?.flagged ?? [];
413
+ const inputLine = blockedInputs.map((b) => `${b.file} (${b.why})`).join("; ");
414
+ const afterInputs = blockedInputs.length
415
+ ? { ...afterSecrets, outcome: afterSecrets.commitAllowed || afterSecrets.outcome === OUTCOMES.WORKER_DONE ? OUTCOMES.NEEDS_REVIEW : afterSecrets.outcome,
416
+ reviewRequired: true, commitAllowed: false,
417
+ commitBlockedReason: afterSecrets.commitAllowed ? `the diff changes what verification runs: ${inputLine}` : afterSecrets.commitBlockedReason,
418
+ reasons: [...afterSecrets.reasons, `VERIFICATION INPUT CHANGED: the diff changes what profile '${verification}' runs, so its result can't be trusted: ${inputLine}`] }
419
+ : afterSecrets;
420
+ const afterConfig = flaggedInputs.length
421
+ ? { ...afterInputs, reviewRequired: true, reasons: [...afterInputs.reasons, `TEST CONFIG CHANGED: ${flaggedInputs.map((c) => `${c.file} (${c.why})`).join("; ")}`] }
422
+ : afterInputs;
423
+ const finalOutcome = applyRefactorContract(applyVerificationPolicy(afterConfig, independentVerification.status, repoPolicy()),
403
424
  { refactor, verificationStatus: independentVerification.status, testChanges: preCommit.testChanges });
404
425
 
405
426
  progress("commit");
@@ -411,6 +432,7 @@ export function createExecutor(deps) {
411
432
  let coordinatorStatus = COORDINATOR_STATUS_BY_OUTCOME[finalOutcome.outcome] ?? "incomplete";
412
433
  const issues = [...finalOutcome.reasons, ...(preCommit.issues ?? [])];
413
434
  if (workerError) issues.push(`worker error: ${String(workerError).split("\n")[0]}`);
435
+ if ((result ?? attempted)?.salvaged) issues.push(`runner cleanup failed after the run (${(result ?? attempted).salvagedFrom}); the worker's report was recovered from the run's transcript`);
414
436
  if (workerFailed || workerTimedOut) { const restarted = vmRestartIssue(vmStartedBefore, deps.podmanVmStartedAt?.() ?? null); if (restarted) issues.unshift(restarted); }
415
437
  if (repositoryChanged && !commit.created) {
416
438
  if (coordinatorStatus === "complete") coordinatorStatus = "incomplete";
@@ -546,7 +568,7 @@ export function createExecutor(deps) {
546
568
  task, acceptance, verification: null, mode, cwd: worktree, baseRef: base.ref, baseSha: base.sha,
547
569
  timeoutSeconds: remainingSeconds, runtimeDir, profile, reasoning, pool, subscriptionWorker, onBehalfOf, model, reportSize, jobDir, workerId: workerId || jobId,
548
570
  evidenceTool: evidencePlaced ? evidenceTool : null,
549
- overridePrompt: scoutReportRecoveryPrompt({ report: used.report.scout }), logSuffix: "-recovery",
571
+ overridePrompt: scoutReportRecoveryPrompt({ report: used.report.scout, question: task, acceptance }), logSuffix: "-recovery",
550
572
  });
551
573
  const recoveryReport = parseScoutReport(finalText(recoveryResult), (recoveryResult?.budgetsUsed ?? used).scout);
552
574
  if (!isScoutReportUnusable(recoveryReport)) {
@@ -575,6 +597,7 @@ export function createExecutor(deps) {
575
597
  const worker = workerMetadata(result ?? attempted);
576
598
  const issues = [...outcome.reasons];
577
599
  if (workerError) issues.push(`scout error: ${String(workerError).split("\n")[0]}`);
600
+ if ((result ?? attempted)?.salvaged) issues.push(`runner cleanup failed after the run (${(result ?? attempted).salvagedFrom}); the scout's report was recovered from the run's transcript`);
578
601
  const failures = worker.toolSummary?.failures ?? 0; if (failures > 0) issues.push(`scout recorded ${failures} tool failure(s)`);
579
602
  if (dirty) issues.push(`snapshot changed: ${record.repoStatusFiles.join(", ")}`);
580
603
  if (reportRecoveryAttempted) {
@@ -706,6 +729,7 @@ export function createExecutor(deps) {
706
729
  const worker = workerMetadata(result ?? attempted);
707
730
  const issues = [...outcome.reasons];
708
731
  if (workerError) issues.push(`decompose error: ${String(workerError).split("\n")[0]}`);
732
+ if ((result ?? attempted)?.salvaged) issues.push(`runner cleanup failed after the run (${(result ?? attempted).salvagedFrom}); the decomposer's report was recovered from the run's transcript`);
709
733
  const failures = worker.toolSummary?.failures ?? 0; if (failures > 0) issues.push(`decomposer recorded ${failures} tool failure(s)`);
710
734
  if (dirty) issues.push(`snapshot changed: ${record.repoStatusFiles.join(", ")}`);
711
735
  if (overlaps.length) issues.push(`${overlaps.length} subtask pair(s) claim overlapping files; not safe to dispatch as independent jobs as proposed`);
package/lib/health.mjs CHANGED
@@ -324,7 +324,9 @@ export async function checkAndRecordHealth({ projectDir, stateRoot, configDir, n
324
324
  if (ageMs !== null) { try { autoPruned = { ...pruneJobRuntime({ stateRoot, olderThanMs: ageMs, now }), olderThanHours: ageMs / 3600000 }; } catch { /* best-effort */ } }
325
325
  const mode = executionMode(env).mode;
326
326
  const { defaultInstallDir } = await import("./connect.mjs");
327
- const install = readInstallVersions(defaultInstallDir());
327
+ const installDir = defaultInstallDir();
328
+ const { loadHarnesses } = await import("./harnesses.mjs");
329
+ const install = { ...readInstallVersions(installDir), copyHarnesses: Object.keys(loadHarnesses(path.join(installDir, "harnesses")).harnesses).length };
328
330
  const result = await runHealthChecks({ now, mode, armySummary, agentsError, jobsRoot: path.join(stateRoot, "jobs"), pidAlive,
329
331
  agents, openclawConfig: readOpenclawConfig(), vendors: SUBSCRIPTION_VENDORS, modelsInUse, autoPruned, usageSnapshots: readUsageSnapshots(stateRoot), install });
330
332
  const toNotify = recordHealth(path.join(stateRoot, "health.json"), result, { now });
@@ -10,6 +10,7 @@
10
10
  // (the session wasn't restarted); per session, so it is
11
11
  // reported by that session's server, not in health.json
12
12
 
13
+ import { execFileSync } from "node:child_process";
13
14
  import fs from "node:fs";
14
15
  import path from "node:path";
15
16
 
@@ -21,8 +22,30 @@ export function readPackageVersion(root) {
21
22
  catch { return null; }
22
23
  }
23
24
 
25
+ /** The checkout's commit, for a git install; null for an npm install. */
26
+ export function readSourceCommit(root) {
27
+ if (!fs.existsSync(path.join(root, ".git"))) return null;
28
+ try { return execFileSync("git", ["rev-parse", "HEAD"], { cwd: root, encoding: "utf8", stdio: ["ignore", "pipe", "ignore"] }).trim() || null; }
29
+ catch { return null; }
30
+ }
31
+
24
32
  export function recordCopySource(installDir, nomarmyRoot) {
25
- fs.writeFileSync(path.join(installDir, SOURCE_FILE), JSON.stringify({ root: nomarmyRoot, version: readPackageVersion(nomarmyRoot) }, null, 2) + "\n");
33
+ const commit = readSourceCommit(nomarmyRoot);
34
+ fs.writeFileSync(path.join(installDir, SOURCE_FILE), JSON.stringify({ root: nomarmyRoot, version: readPackageVersion(nomarmyRoot), ...(commit ? { commit } : {}) }, null, 2) + "\n");
35
+ }
36
+
37
+ /**
38
+ * Whether the copy in installDir is behind the checkout or package at
39
+ * nomarmyRoot: an older version, a different commit (a git install moves on
40
+ * without a version bump), or no record of its source at all.
41
+ */
42
+ export function copyIsStale(installDir, nomarmyRoot) {
43
+ let source = null;
44
+ try { source = JSON.parse(fs.readFileSync(path.join(installDir, SOURCE_FILE), "utf8")); } catch { return true; }
45
+ const copyVersion = readPackageVersion(installDir), rootVersion = readPackageVersion(nomarmyRoot);
46
+ if (!copyVersion || !rootVersion || copyVersion !== rootVersion) return true;
47
+ const commit = readSourceCommit(nomarmyRoot);
48
+ return Boolean(commit && source.commit !== commit);
26
49
  }
27
50
 
28
51
  /** The copy's version and, when connect recorded it, the version now at its source. */
@@ -52,9 +75,14 @@ export function compareVersions(a, b) {
52
75
 
53
76
  const valid = (v) => typeof v === "string" && /^v?\d+\.\d+\.\d+/.test(v.trim());
54
77
 
55
- /** Health issues for a stale CLI or a stale installed copy. */
56
- export function freshnessIssues({ copyVersion = null, sourceVersion = null, latestVersion = null }) {
78
+ /** Health issues for a stale CLI, a stale installed copy, or a copy missing its harnesses. */
79
+ export function freshnessIssues({ copyVersion = null, sourceVersion = null, latestVersion = null, copyHarnesses = null }) {
57
80
  const issues = [];
81
+ if (valid(copyVersion) && copyHarnesses === 0) {
82
+ issues.push({ id: `nomarmy-harnesses:${copyVersion}`, severity: "error", title: "The nomArmy your coordinators run has no harnesses",
83
+ detail: "Jobs match no harness, so they run in the plain base image: no dependencies, fake services, browser tests or artifacts.",
84
+ fix: "nomarmy connect claude (and codex, cursor), then restart those sessions", short: "no harnesses" });
85
+ }
58
86
  if (valid(latestVersion) && valid(sourceVersion) && compareVersions(sourceVersion, latestVersion) < 0) {
59
87
  const latest = latestVersion.trim();
60
88
  issues.push({ id: `nomarmy-update:${latest}`, severity: "info", title: `nomArmy ${latest} is out (you have ${sourceVersion})`,
@@ -124,3 +124,27 @@ export function formatUnion(union) {
124
124
  const artifacts = union.worktree ? `\n\nUnion artifacts: ${path.dirname(union.worktree)}\nWorktree retained for review: ${union.worktree}\nBranch retained for review: ${union.branch}` : "";
125
125
  return `${banner}--- UNION RECORD ---\n${JSON.stringify(union, null, 2)}${artifacts}`;
126
126
  }
127
+
128
+ /**
129
+ * What a person calls a job: its commit subject, else the task's first
130
+ * sentence, capped. Two jobs on one agent and model differ here even with no
131
+ * role, so a notification can tell them apart.
132
+ */
133
+ export function jobLabel(args) {
134
+ const text = String(args?.commit_subject || args?.task || "").split("\n")[0].split(/(?<=\.)\s/)[0].trim();
135
+ return text.length > 60 ? `${text.slice(0, 57)}...` : text || null;
136
+ }
137
+
138
+ /**
139
+ * Seconds a job has run: to its finish once it has one, not to whenever it's
140
+ * asked about. The record's total_elapsed covers the whole job; an implement
141
+ * job's finishedAt marks only the worker's end, before verification and commit.
142
+ */
143
+ export function jobElapsedSeconds({ status = null, meta = null, entry = null, now = Date.now() } = {}) {
144
+ const total = meta?.metrics?.total_elapsed;
145
+ if (Number.isFinite(total)) return Math.round(total / 1000);
146
+ const startedMs = Date.parse(status?.startedAt ?? entry?.startedAt ?? meta?.startedAt ?? "");
147
+ if (!Number.isFinite(startedMs)) return null;
148
+ const finishedMs = Date.parse((status?.state === "finished" ? status.updatedAt : null) ?? meta?.finishedAt ?? "");
149
+ return Math.round(((Number.isFinite(finishedMs) ? finishedMs : now) - startedMs) / 1000);
150
+ }
package/lib/outcome.mjs CHANGED
@@ -7,6 +7,16 @@ import { parseWorkerReport } from "./report.mjs";
7
7
  // failed. Recovery exists so that a mangled REPORT cannot destroy correct WORK.
8
8
  // It does not exist to launder a failure into a success.
9
9
  // ---------------------------------------------------------------------------
10
+ // A failed verification's own detail (the command, its exit code, the end of
11
+ // its output), capped so an issue line stays readable; the full output is in
12
+ // the job's verification.log.
13
+ const FAILURE_DETAIL_CHARS = 600;
14
+ function failureDetail(independentVerification) {
15
+ const detail = String(independentVerification?.detail ?? independentVerification?.reason ?? "").trim();
16
+ if (!detail) return "";
17
+ return `: ${detail.length > FAILURE_DETAIL_CHARS ? `${detail.slice(0, FAILURE_DETAIL_CHARS)}...` : detail}`;
18
+ }
19
+
10
20
  export function resolveOutcome({ report, repositoryChanged = false, independentVerification = null, regressionCheck = null, workerFailed = false, workerTimedOut = false, mode = "implement" }) {
11
21
  const verification = independentVerification?.status ?? "not_run";
12
22
  const parsed = report ?? parseWorkerReport("");
@@ -34,7 +44,7 @@ export function resolveOutcome({ report, repositoryChanged = false, independentV
34
44
  if (verification === "fail") {
35
45
  return { ...base, outcome: OUTCOMES.NEEDS_REVIEW, reviewRequired: true,
36
46
  commitBlockedReason: "independent verification failed despite a clean done/pass report",
37
- reasons: ["worker claimed done/pass but independent verification failed"] };
47
+ reasons: [`worker claimed done/pass but independent verification failed${failureDetail(independentVerification)}`] };
38
48
  }
39
49
  // verify_regression: reverting just the production files and re-running
40
50
  // the SAME verification profile still passed (or came back genuinely
@@ -92,7 +102,7 @@ export function resolveOutcome({ report, repositoryChanged = false, independentV
92
102
  if (verification === "fail") {
93
103
  return { ...recovery, outcome: OUTCOMES.WORKER_REPORT_INVALID,
94
104
  commitBlockedReason: "independent verification failed; recovery cannot promote a failure",
95
- reasons: [...recovery.reasons, "independent verification FAILED"] };
105
+ reasons: [...recovery.reasons, `independent verification FAILED${failureDetail(independentVerification)}`] };
96
106
  }
97
107
  if (verification === "pass") {
98
108
  // A leniently recovered `done` plus a passing independent check is the
package/lib/scout.mjs CHANGED
@@ -435,8 +435,12 @@ export function isScoutReportUnusable(report) {
435
435
  * already found -- never asks it to look further, since a fresh read pass is
436
436
  * exactly the cost a scout exists to avoid paying twice.
437
437
  */
438
- export function scoutReportRecoveryPrompt({ report = { targetTokens: 600, hardCapTokens: 1024 } } = {}) {
439
- return `Your previous reply ended without the required SCOUT REPORT, or was cut off before reaching END.\n\nDo not repeat, redo, retry, or explore further. Do not call any tool. Based only on what you already found, reply with ONLY the report below, nothing before it, nothing after it:\n\nSCOUT REPORT\nQUESTION: <the question restated in one line>\nCONFIDENCE: high | medium | low\nFINDING: <one sentence> [src/example.js:10-24]\nNOT_FOUND: none | <what you looked for and could not find>\nEND\n\nIf you did not actually find anything worth a FINDING, say so under NOT_FOUND rather than inventing one. Target ${report.targetTokens} tokens; ${report.hardCapTokens} is the hard cap.`;
438
+ // The recovery call carries the question itself: a reply cut off mid-run can
439
+ // leave the resumed session without it, and a scout told only to "restate the
440
+ // question" then came back with an empty report (a live PM scout, twice).
441
+ export function scoutReportRecoveryPrompt({ report = { targetTokens: 600, hardCapTokens: 1024 }, question = null, acceptance = [] } = {}) {
442
+ const asked = question ? `\n\nThe question you were answering:\n${String(question).trim()}${acceptance?.length ? `\n\nA complete answer covers:\n${acceptance.map((a) => `- ${a}`).join("\n")}` : ""}` : "";
443
+ return `Your previous reply ended without the required SCOUT REPORT, or was cut off before reaching END.${asked}\n\nDo not repeat, redo, retry, or explore further. Do not call any tool. Based only on what you already found, reply with ONLY the report below, nothing before it, nothing after it:\n\nSCOUT REPORT\nQUESTION: <the question restated in one line>\nCONFIDENCE: high | medium | low\nFINDING: <one sentence> [src/example.js:10-24]\nNOT_FOUND: none | <what you looked for and could not find>\nEND\n\nIf you did not actually find anything worth a FINDING, say so under NOT_FOUND rather than inventing one. Target ${report.targetTokens} tokens; ${report.hardCapTokens} is the hard cap.`;
440
444
  }
441
445
 
442
446
  // ---------------------------------------------------------------------------
package/lib/verify.mjs CHANGED
@@ -135,10 +135,19 @@ function tailOf(result) {
135
135
  * command that never started is absence of evidence, so if nothing ever ran the
136
136
  * verdict is `not_run`, never `fail`.
137
137
  *
138
+ * A command guarded on nomArmy's changed-file variables (`if [ -n
139
+ * "$NOMARMY_CHANGED_TEST_FILES" ]; then ...; fi`) exits 0 without running
140
+ * anything when those are empty, as on an unchanged checkout in mode: verify.
141
+ * That's skipped, not passed: exit 0, no output at all, and every
142
+ * NOMARMY_CHANGED_* variable the command names empty. A command with a
143
+ * fallback (`${NOMARMY_CHANGED_TEST_FILES:-tests/}`) prints test output, so
144
+ * it counts as run. If every command skipped, nothing was verified: not_run.
145
+ *
138
146
  * @param {Array<object>} results
139
- * @returns {{ status: "pass"|"fail"|"not_run", detail: string }}
147
+ * @param {{ env?: Record<string, string> }} [context] the NOMARMY_CHANGED_* values the commands saw
148
+ * @returns {{ status: "pass"|"fail"|"not_run", detail: string, basis?: string }}
140
149
  */
141
- export function classifyResults(results) {
150
+ export function classifyResults(results, { env = {} } = {}) {
142
151
  const list = Array.isArray(results) ? results.filter(Boolean) : [];
143
152
  if (list.length === 0) {
144
153
  return { status: "not_run", detail: "no commands were executed" };
@@ -176,9 +185,28 @@ export function classifyResults(results) {
176
185
  }
177
186
  }
178
187
 
188
+ const skipped = list.map((result, index) => ({ index, why: guardedNoOp(result, env) })).filter((s) => s.why);
189
+ if (skipped.length === total) {
190
+ return { status: "not_run", basis: "nothing-changed", detail: `every command is scoped to changed files and nothing changed, so none ran (${[...new Set(skipped.map((s) => s.why))].join("; ")}); add an unscoped profile to run the full suite` };
191
+ }
192
+ if (skipped.length) {
193
+ const which = skipped.map((s) => `command ${s.index + 1} (\`${list[s.index].command ?? "?"}\`) did nothing: ${s.why}`).join("; ");
194
+ return { status: "pass", detail: `${total - skipped.length} of ${total} commands passed; ${skipped.length} skipped: ${which}` };
195
+ }
179
196
  return { status: "pass", detail: `${total} of ${total} commands passed` };
180
197
  }
181
198
 
199
+ const CHANGED_VAR_RE = /\$\{?(NOMARMY_CHANGED_[A-Z_]+)/g;
200
+
201
+ /** Why a zero-exit, silent command was a guarded no-op, or null if it ran. */
202
+ function guardedNoOp(result, env) {
203
+ if (result.exitCode !== 0 || result.timedOut || result.started === false) return null;
204
+ if (result.silent === false || (result.silent === undefined && `${result.stdout ?? ""}${result.stderr ?? ""}`.trim())) return null;
205
+ const names = [...new Set([...String(result.command ?? "").matchAll(CHANGED_VAR_RE)].map((m) => m[1]))];
206
+ if (!names.length || names.some((name) => String(env[name] ?? "").trim())) return null;
207
+ return `${names.join(" and ")} ${names.length > 1 ? "are" : "is"} empty`;
208
+ }
209
+
182
210
  // ---------------------------------------------------------------------------
183
211
  // pure: output capping
184
212
  // ---------------------------------------------------------------------------
@@ -853,6 +881,9 @@ export function createVerificationRunner(options = {}) {
853
881
  timeoutMs,
854
882
  stdout: stdout.text,
855
883
  stderr: stderr.text,
884
+ // Whether the command printed nothing at all, judged before any
885
+ // output is withheld, so a guarded no-op is still recognized.
886
+ silent: !String(raw?.stdout ?? "").trim() && !String(raw?.stderr ?? "").trim(),
856
887
  truncated: { stdout: stdout.dropped, stderr: stderr.dropped },
857
888
  };
858
889
  results.push(entry);
@@ -876,7 +907,7 @@ export function createVerificationRunner(options = {}) {
876
907
  }
877
908
  }
878
909
 
879
- const verdict = classifyResults(results);
910
+ const verdict = classifyResults(results, { env: verificationEnv });
880
911
  const executed = results.filter((r) => r.started).length;
881
912
  const dropped = results.reduce(
882
913
  (total, r) => total + (r.truncated?.stdout ?? 0) + (r.truncated?.stderr ?? 0),
@@ -888,7 +919,7 @@ export function createVerificationRunner(options = {}) {
888
919
  : "";
889
920
 
890
921
  if (verdict.status === "not_run") {
891
- return notRun(verdict.detail, "sandbox-unavailable", { ...networkInfo, output: renderOutput() });
922
+ return notRun(verdict.detail, verdict.basis ?? "sandbox-unavailable", { ...networkInfo, output: renderOutput() });
892
923
  }
893
924
 
894
925
  return {
@@ -69,10 +69,13 @@ export function describeRecoveryChanges(record) {
69
69
  return `${record.repoStatusFiles.length} file(s) differ from a clean checkout: ${record.repoStatusFiles.join(", ")}`;
70
70
  }
71
71
 
72
- export function reportRecoveryPrompt({ report = { targetTokens: 256, hardCapTokens: 512 }, changes = null } = {}) {
72
+ export function reportRecoveryPrompt({ report = { targetTokens: 256, hardCapTokens: 512 }, changes = null, task = null } = {}) {
73
+ // The objective itself, so STATUS is judged against it even when the
74
+ // resumed session lost it with the cut-off reply.
75
+ const objective = task ? `\nThe objective you were working on:\n${String(task).trim()}\n` : "";
73
76
  const changesLine = changes
74
77
  ? `\nThe repository (checked independently just now, not from your memory of this session) already shows: ${changes}. Trust this over any uncertainty about what you did or did not do.\n`
75
78
  : `\nThe repository (checked independently just now, not from your memory of this session) shows no changes at all.\n`;
76
- return `Your previous reply ended without the required final report, or was cut off before completing it.\n${changesLine}\nDo not repeat, redo, retry, or describe any action you already took. Do not call any tool. Reply with ONLY the four lines below, nothing before them, nothing after them:\n\nSTATUS: done | partial | blocked\nTESTS: pass | fail | not_run\nNOT_DONE: none | <brief>\nNOTE: <brief implementation or risk note>\n\nUse the exact field names above, including the underscore in NOT_DONE. Target ${report.targetTokens} tokens; ${report.hardCapTokens} is the hard cap. Base STATUS on the repository state above, not on what you recall attempting: if it shows the edit landed, you may report done; if it shows nothing relevant, report blocked or partial rather than guessing done.`;
79
+ return `Your previous reply ended without the required final report, or was cut off before completing it.\n${objective}${changesLine}\nDo not repeat, redo, retry, or describe any action you already took. Do not call any tool. Reply with ONLY the four lines below, nothing before them, nothing after them:\n\nSTATUS: done | partial | blocked\nTESTS: pass | fail | not_run\nNOT_DONE: none | <brief>\nNOTE: <brief implementation or risk note>\n\nUse the exact field names above, including the underscore in NOT_DONE. Target ${report.targetTokens} tokens; ${report.hardCapTokens} is the hard cap. Base STATUS on the repository state above, not on what you recall attempting: if it shows the edit landed, you may report done; if it shows nothing relevant, report blocked or partial rather than guessing done.`;
77
80
  }
78
81
 
package/mcp/server.mjs CHANGED
@@ -37,7 +37,7 @@ import { modelRefusals } from "../lib/health.mjs";
37
37
  import { podmanProblem, podmanVmStartedAt } from "../lib/podman-health.mjs";
38
38
  import { restartNotice } from "../lib/install-freshness.mjs";
39
39
  import { createBuildMetrics, resolveOutcome, finalText, workerMetadata, usageMetrics, policyAdmissionProblems, applyRefactorContract, applyVerificationPolicy, resolveVerifyRegression } from "../lib/outcome.mjs";
40
- import { compactJobRecord, formatResult, formatUnion, testChangeBanner, regressionCheckBanner, decomposeOverlapBanner } from "../lib/job-format.mjs";
40
+ import { jobLabel, compactJobRecord, formatResult, formatUnion, testChangeBanner, regressionCheckBanner, decomposeOverlapBanner } from "../lib/job-format.mjs";
41
41
 
42
42
  export { run, mapLimit };
43
43
  export { readsMeasurable, measureReads };
@@ -582,7 +582,7 @@ server.tool("local_workers", "Run independent jobs (implement or scout) with bou
582
582
  // so it was invisible to both ceilings while it ran.
583
583
  // A batch job waits for its agent's slot (up to its own timeout) rather
584
584
  // than failing because an earlier job in the same batch holds it.
585
- return trackInRun(j, track(jobId, { mode: j.mode, workerId, lane: jobLane(j), agent: j.agentName ?? null, runId: j.run_id ?? null, role: j.armyRole ?? null, model: j.model ?? null },
585
+ return trackInRun(j, track(jobId, { mode: j.mode, workerId, lane: jobLane(j), agent: j.agentName ?? null, runId: j.run_id ?? null, role: j.armyRole ?? null, model: j.model ?? null, label: jobLabel(j) },
586
586
  withAgentSlot(j, jobId, () => executeJob({ ...jobArgs(effectiveJob, workerId), jobId }), { waitMs: (j.timeout_seconds ?? 600) * 1000 }))).promise;
587
587
  }, { staggerMs: WORKER_START_STAGGER_MS });
588
588
  indices.forEach((i, laneI) => { results[i] = laneResults[laneI]; });
package/package.json CHANGED
@@ -1,9 +1,9 @@
1
1
  {
2
2
  "name": "nomarmy",
3
- "description": "A harness for AI coding workers whose claims are never trusted: your coding assistant stays in charge while workers implement and test in sandboxes, on local models, API keys or your own subscriptions.",
3
+ "description": "Every byte verified: a harness for AI coding workers whose claims are never trusted. Your coding assistant stays in charge while workers implement and test in sandboxes, and nomArmy checks every change before it is committed.",
4
4
  "author": "Rayson Technologies",
5
5
  "license": "Apache-2.0",
6
- "version": "0.1.0-alpha.7",
6
+ "version": "0.1.0-alpha.9",
7
7
  "private": false,
8
8
  "type": "module",
9
9
  "engines": {