nomarmy 0.1.0-alpha.7 → 0.1.0-alpha.8

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -6,20 +6,22 @@
6
6
  <a href="https://github.com/rayson-tech/nomarmy/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-Apache%202.0-blue.svg" alt="License: Apache 2.0"></a>
7
7
  </p>
8
8
 
9
- <p align="center"><em>Tiny coders, big appetites for bounded tickets.</em> 🍪</p>
9
+ <p align="center"><em>Every byte verified.</em> 🍪</p>
10
10
 
11
11
  **Your coding assistant plans; sandboxed workers build; nothing counts until nomArmy has checked it.**
12
12
 
13
+ AI coding workers are confident. Their "done, all tests pass" is a claim, not evidence. nomArmy lets your coding assistant (Claude Code, Codex or Cursor) hand work to workers called **noms**, then checks every change itself before anything is committed: the real diff, your tests run in a fresh sandbox, a check that those tests actually catch the change, and a secret scan.
14
+
13
15
  ## TL;DR
14
16
 
15
- 1. **Have** Git, Node 24.16+ (or 26.1+; OpenClaw, which nomArmy installs, needs it) and [Podman](https://podman.io) (on macOS: `brew install podman && podman machine init --memory 8192 && podman machine start`; Podman's 2 GiB default is too small for nomArmy's sandboxes).
17
+ 1. **Have** Git, Node 24.16+ (or 26.1+) and [Podman](https://podman.io). On macOS, give Podman 8 GiB: `brew install podman && podman machine init --memory 8192 && podman machine start`.
16
18
  2. **Install and set up:**
17
19
  ```bash
18
20
  npm install -g nomarmy@alpha
19
21
  cd your-project
20
22
  nomarmy setup
21
23
  ```
22
- `nomarmy setup` is the playbook. It shows a checklist and runs the next step each time you say yes:
24
+ `nomarmy setup` is a playbook. It shows a checklist and runs the next step each time you say yes:
23
25
  ```text
24
26
  ✓ Where models run: hosted
25
27
  ✓ Installed: OpenClaw 2026.9.6
@@ -29,34 +31,39 @@
29
31
  Check: verify the installation
30
32
  Run `nomarmy agents add` now? [Y/n]
31
33
  ```
32
- In order: pick where models run (API keys and subscriptions for most people), install OpenClaw and the sandbox, add your agents (an API key, or your ChatGPT or Muse Code subscription), put the roles on them, write this repo's `.nomarmy.yml`, then check it all. Stop anytime; `nomarmy setup` picks up where you left off.
33
-
34
- **Want every step spelled out?** [Example setup: Claude Code, Codex and an API key](https://github.com/rayson-tech/nomarmy/blob/main/docs/setup/example.md) walks through a complete setup, command by command.
34
+ Stop anytime; `nomarmy setup` picks up where you left off. **Want every step spelled out?** [Example setup: Claude Code, Codex and an API key](https://github.com/rayson-tech/nomarmy/blob/main/docs/setup/example.md) goes command by command.
35
35
  3. **Use it:** restart Claude Code in the project and ask it to use nomArmy for one small bug that has a test. When that works, try `/feature <what you want built>`.
36
36
 
37
- **Have a GPU or a Mac with plenty of memory?** Choose "a local model" in `nomarmy setup` and workers run on llama.cpp on your own machine: no per-token bill and your code stays home, but you pay in hardware, power and speed, and a model too big for your memory crawls. `nomarmy sizing` tells you what fits; see [Install](https://github.com/rayson-tech/nomarmy/blob/main/docs/install.md#install). A team GPU server works too: [a shared model server](https://github.com/rayson-tech/nomarmy/blob/main/docs/install.md#a-shared-model-server). Codex or Cursor as the coordinator: `nomarmy connect codex cursor`.
37
+ Stuck? `nomarmy doctor` checks the machine and `nomarmy health` checks everything nomArmy runs on. Upgrading later? `nomarmy update`.
38
38
 
39
- Stuck? `nomarmy doctor` checks the machine, and `nomarmy health` checks everything nomArmy runs on.
39
+ ## How every byte gets verified
40
40
 
41
- ## What it is
41
+ 1. Your coding assistant, the **General**, briefs a job: a task, acceptance criteria, and the tests that prove it.
42
+ 2. nomArmy creates a git worktree from your branch and runs the nom in a Podman sandbox with no network and no host credentials. (One exception, the Claude subscription: see [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md#security-posture).)
43
+ 3. The nom edits, runs tests, and ends with a four-line report: `STATUS`, `TESTS`, `NOT_DONE`, `NOTE`.
44
+ 4. nomArmy treats that report as a claim and checks the evidence itself:
45
+ - reads the real diff from git, not the nom's description of it
46
+ - runs your verification profile in a fresh sandbox
47
+ - reverts the production change and reruns the tests: a test that still passes proves nothing, so the job goes to review instead of being committed
48
+ - blocks on secrets, and flags tests made to pass (new skips, stubbed imports) and code nothing calls
49
+ 5. Only then does it commit, on the nom's own branch. It never merges into yours: reviewing and integrating stay with the General, and with you.
42
50
 
43
- Your coding assistant (Claude Code, Codex or Cursor) stays in charge as the **General**: it decides what gets built and whether the result is acceptable. The work goes to **noms**, workers that implement, test and repair in their own git worktree and sandbox, on an API key, your own ChatGPT or Muse Code subscription, or a local model. nomArmy owns everything in between: worktrees, git, sandboxes, verification, and the evidence that decides whether work is accepted.
51
+ Failing verification stays failed, unconditionally. A malformed report isn't automatically a failure: if the repository changed, nomArmy verifies independently and may recover the work. And the checks aren't the General's to waive: a repo's `.nomarmy.yml` policy (on by default for new repos) makes verification and the revert check mandatory for every job.
44
52
 
45
- **What you get is work you don't have to take on faith**, not cheaper work. Delegating costs the General tokens too: briefing and reviewing. On small, already-diagnosed tickets we measured 4 to 8 times more of the General's tokens than fixing the bug directly, and break-even at roughly 150 lines of context a fix needs to read ([the measurements](https://github.com/rayson-tech/nomarmy/blob/main/docs/experiments/2026-09-20-model-bakeoff-and-economics.md)). It pays off on bigger tickets, on parallel work, and anywhere you'd otherwise have to trust an agent's say-so.
53
+ **Checking without building** costs nothing: `mode: verify` runs a verification profile against any branch, with no worker and no model tokens.
46
54
 
47
- Developed and maintained by Rayson Technologies. This is an alpha (`0.1.0-alpha`).
55
+ ## Where the work runs
48
56
 
49
- ## How it works
57
+ - **Agents** say where a job can run: an API key, your own ChatGPT or Muse Code subscription, or a local model on llama.cpp.
58
+ - **The army** says which role runs on which agent: Sr and Jr devs build, a security analyst and a data architect review, a PM checks the plan, a PO accepts.
59
+ - **`/feature`** runs a whole feature end to end, from plan through build, review and acceptance, and hands you a branch to merge.
60
+ - **Harnesses** give each repo the right sandbox: Go, Rust, Python and Node (mixed repos too), Playwright browser tests, and fake services like a mock login server, all offline. [Adding one](https://github.com/rayson-tech/nomarmy/blob/main/CONTRIBUTING.md#adding-a-harness) never touches core code.
50
61
 
51
- 1. The General briefs a job: a task, acceptance criteria, the tests that prove it.
52
- 2. nomArmy creates a worktree from your branch and runs the worker in a Podman sandbox with no network and no host credentials. (One exception, the Claude subscription: see [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md#security-posture).)
53
- 3. The worker edits, runs tests, and ends with a four-line report: `STATUS`, `TESTS`, `NOT_DONE`, `NOTE`.
54
- 4. nomArmy treats that report as a claim. It reads the real diff from git, runs your verification profile itself in a fresh sandbox, reverts the production change to check the tests actually fail without it, and scans for secrets.
55
- 5. Only then does it commit, on the worker's own branch. It never merges into yours: reviewing and integrating stay with the General, and with you.
62
+ **What you get is work you don't have to take on faith**, not cheaper work. Delegating costs the General tokens too, for briefing and review: on small, already-diagnosed tickets we measured 4 to 8 times more of the General's tokens than fixing the bug directly, with break-even around 150 lines of context a fix needs to read ([the measurements](https://github.com/rayson-tech/nomarmy/blob/main/docs/experiments/2026-09-20-model-bakeoff-and-economics.md)). It pays off on bigger tickets, parallel work, and anywhere you'd otherwise trust an agent's say-so.
56
63
 
57
- A malformed report isn't automatically a failure: if the repository changed, nomArmy verifies independently and may recover the work. Failing verification stays failed, unconditionally. And the checks aren't the General's to waive: a repo's `.nomarmy.yml` policy (on by default for new repos) makes verification and the revert check mandatory for every job.
64
+ **Have a GPU or a Mac with plenty of memory?** Choose "a local model" in `nomarmy setup`: no per-token bill and your code stays home, but you pay in hardware, power and speed. `nomarmy sizing` tells you what fits. A [shared model server](https://github.com/rayson-tech/nomarmy/blob/main/docs/install.md#a-shared-model-server) works too.
58
65
 
59
- Around that core: **agents** say where a job can run, the **army** says which role runs on which agent, and **`/feature`** runs a whole feature end to end, from plan through build, review and acceptance, handing you a branch to merge.
66
+ Developed and maintained by Rayson Technologies. This is an alpha (`0.1.0-alpha`).
60
67
 
61
68
  ## Docs
62
69
 
@@ -66,7 +73,7 @@ Around that core: **agents** say where a job can run, the **army** says which ro
66
73
  | [Example setup](https://github.com/rayson-tech/nomarmy/blob/main/docs/setup/example.md) | Claude Code, Codex and an API key, command by command |
67
74
  | [Agents and the army](https://github.com/rayson-tech/nomarmy/blob/main/docs/agents-and-army.md) | Where a job can run, who does what, usage limits, picking an agent |
68
75
  | [`/feature` runs](https://github.com/rayson-tech/nomarmy/blob/main/docs/feature-runs.md) | A feature end to end, and watching what nomArmy is doing |
69
- | [Your repository](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md) | `.nomarmy.yml`, verification, languages and dependencies, what nomArmy checks |
76
+ | [Your repository](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md) | `.nomarmy.yml`, verification, dependencies, private registries, what nomArmy checks |
70
77
  | [Harnesses](https://github.com/rayson-tech/nomarmy/blob/main/docs/harnesses.md) | Ecosystem registry, detection, network levels, and requirements |
71
78
  | [Configuration](https://github.com/rayson-tech/nomarmy/blob/main/docs/configuration.md) | Settings, swapping the local model, sizing, admission |
72
79
  | [Reference](https://github.com/rayson-tech/nomarmy/blob/main/docs/reference.md) | Every CLI command and MCP tool |
@@ -75,23 +82,20 @@ Around that core: **agents** say where a job can run, the **army** says which ro
75
82
 
76
83
  ## Security
77
84
 
78
- A worker gets a writable git worktree inside a Podman sandbox and nothing else: no network, no host credentials, no Podman socket. Every model call is made by OpenClaw on your machine, never from inside the sandbox. **The exception is a Claude subscription**, whose tools run on your machine, so nomArmy refuses build jobs on it unless you allow it. Never hand a worker production credentials, deployment access or SSH keys. Details: [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md); to report a vulnerability, [SECURITY.md](https://github.com/rayson-tech/nomarmy/blob/main/SECURITY.md).
85
+ A nom gets a writable git worktree inside a Podman sandbox and nothing else: no network, no host credentials, no Podman socket. Every model call is made by OpenClaw on your machine, never from inside the sandbox. Verification can climb a network ladder one rung at a time (fake services on a private network, then an allowlist you approve for a test tenant), but noms never leave `network none`. **The exception is a Claude subscription**, whose tools run on your machine, so nomArmy refuses build jobs on it unless you allow it. Never hand a nom production credentials, deployment access or SSH keys. Details: [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md); to report a vulnerability, [SECURITY.md](https://github.com/rayson-tech/nomarmy/blob/main/SECURITY.md).
79
86
 
80
87
  ## Status
81
88
 
82
89
  | Capability | Status |
83
90
  |---|---|
84
- | Delegation core: worktrees, nomArmy-owned git, independent verification, kept failed worktrees | Working, end-to-end tested |
85
- | Local (llama.cpp) and Bedrock profiles | Working |
86
- | Scout and decompose modes | Unit and live tested |
87
- | `auto_union`, `verify_regression`, test-selection and unwired-code checks | Unit and live tested; the heuristics are review flags |
88
- | Secret scanning (secretlint, hard block) | Unit tested against the real dependency |
89
- | Agents: api keys | Live-verified with xAI; other providers built to OpenClaw's documented interface |
91
+ | Verification core: worktrees, nomArmy-owned git, independent verification, the revert check, kept failed worktrees | Working, end-to-end tested |
92
+ | `mode: verify`, secret scanning (secretlint, hard block), test-workaround and unwired-code checks | Unit and live tested; the heuristics are review flags |
93
+ | Scout and decompose modes, `auto_union` | Unit and live tested |
94
+ | Agents: API keys | Live-verified with xAI; other providers built to OpenClaw's documented interface |
90
95
  | Agents: subscriptions | ChatGPT (Codex) and Muse Code sandboxed and live-verified; Claude live-verified, but its tools run on the host (scout and review by default) |
96
+ | Local (llama.cpp) and Bedrock profiles | Working |
91
97
  | The army and `/feature` | Driven by a real Claude Code General across three runs, about 18 implement jobs |
92
- | Go, Rust, Python and Node repos, and mixed ones | Harness images live-verified: Go modules and Rust crates prefetched; npm, pnpm, yarn, bun and workspaces; pip, pyproject, uv and poetry |
93
- | Fake services beside the app (mock login server, mock APIs) | Live-verified on a private network with no route out (the `services` harness level) |
94
- | Browser tests (Playwright + Chromium) | Live-verified offline, with screenshots and traces kept as job evidence |
98
+ | Harnesses: Go, Rust, Python, Node and mixed repos; Playwright; fake services | Live-verified offline |
95
99
  | Private registries and a verification-only network allowlist | Live-verified; each passed an independent security review |
96
100
 
97
101
  What we've learned from real runs, including where delegating pays and where it doesn't, is in [docs/findings.md](https://github.com/rayson-tech/nomarmy/blob/main/docs/findings.md).
@@ -99,13 +103,13 @@ What we've learned from real runs, including where delegating pays and where it
99
103
  ### Known limitations
100
104
 
101
105
  - **A Claude subscription isn't sandboxed.** Its tools run on your machine, so implement jobs on it are refused unless you set `allow_host_tools: true`. See [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md#security-posture).
102
- - **A refused model costs one job.** When a vendor refuses a model at run time that OpenClaw lists (gpt-6-sol on a ChatGPT plan), the first job on it fails with `model_not_found`. After that nomArmy refuses to dispatch it until a job or test call on it works. `army assign` tests the job's route and catches this before any job.
103
- - **Claude subscription token counts** come from the Claude CLI's own session log, since OpenClaw sees only the final reply. Totals include cache reads and writes, which make up most of an agent's prompt; each part is also kept separately.
106
+ - **Your own compose services aren't started yet.** A verification profile that needs a real database from your compose file (`environment: basic` or higher) reports `not_run` rather than running without it (and the compose parser doesn't resolve YAML anchors). Fake services from harnesses, like the mock login server, do run.
107
+ - **Private registries don't cover Poetry or Yarn Berry** yet: Poetry can't guarantee a credentialed install runs no package code, and Yarn Berry doesn't read `.npmrc`. uv, pip wheels, npm, pnpm, Yarn Classic and bun work. See [Private registries](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md#private-registries).
108
+ - **A refused model costs one job.** When a vendor refuses a model at run time that OpenClaw lists (gpt-6-sol on a ChatGPT plan), the first job on it fails with `model_not_found`; after that nomArmy won't dispatch it until a job or test call on it works. `army assign` tests the route and catches this before any job.
109
+ - **Claude subscription token counts** come from the Claude CLI's own session log, since OpenClaw sees only the final reply; totals include cache reads and writes.
104
110
  - **Test-workaround detection is a flag, not a verdict**: a legitimate new skip still gets flagged.
105
111
  - **Deploy-time failures need your own check.** See [Add a check for what unit tests can't see](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md#nomarmyyml).
106
- - **Node private registries are not supported yet**: npm, pnpm, yarn and bun lockfiles and workspaces are supported, but the image build has no credentials for private registries (dependency plan step 8).
107
- - **Verification needing services** (a database, a mock server) reports `not_run` instead of running without them. The compose parser doesn't resolve YAML anchors.
108
- - **Same-host sandboxes**: the MCP server, OpenClaw and every job's sandbox run on the machine with the coordinator. Only the model can be elsewhere (an agent, or [a shared model server](https://github.com/rayson-tech/nomarmy/blob/main/docs/install.md#a-shared-model-server)).
112
+ - **Same-host sandboxes**: the MCP server, OpenClaw and every job's sandbox run on the machine with the coordinator. Only the model can be elsewhere.
109
113
 
110
114
  ## More
111
115
 
package/lib/admission.mjs CHANGED
@@ -1,6 +1,7 @@
1
1
  import fs from "node:fs";
2
2
  import path from "node:path";
3
3
  import { executionMode } from "./execution.mjs";
4
+ import { jobLabel, jobElapsedSeconds } from "./job-format.mjs";
4
5
  import { loadConfig } from "./config.mjs";
5
6
  import { clampInt } from "./budget-state.mjs";
6
7
  import { checkBrief, assessAdmission, describeBudgets } from "./budget.mjs";
@@ -160,7 +161,7 @@ export function createJobRuntime(deps) {
160
161
  const who = entry.mode === "verify" ? "verification runner" : entry.agent ? `${entry.agent}${entry.model ? `/${entry.model}` : ""}` : "local model";
161
162
  const took = Math.round((Date.now() - Date.parse(entry.startedAt)) / 60000);
162
163
  const ok = !error && (result?.ok || m.coordinatorStatus === "complete");
163
- notify(`nomArmy: ${entry.role ?? entry.mode ?? "job"} ${ok ? "done" : outcome}`, `${entry.workerId ?? entry.jobId} on ${who}: ${outcome} after ${took}m. ${ok ? "Ready for the General's review." : "Needs a look."}`);
164
+ notify(`nomArmy: ${entry.role ?? entry.mode ?? "job"} ${ok ? "done" : outcome}`, `${entry.label ? `${entry.label} (${entry.workerId ?? entry.jobId})` : entry.workerId ?? entry.jobId} on ${who}: ${outcome} after ${took}m. ${ok ? "Ready for the General's review." : "Needs a look."}`);
164
165
  }
165
166
  function admissionHardware() {
166
167
  return executionMode(deps.env).managesModelServer ? deps.budgetState.hardwareSnapshot : null;
@@ -349,7 +350,7 @@ export function createJobRuntime(deps) {
349
350
  function launch(args) {
350
351
  const workerId = args.worker_id || null;
351
352
  const jobId = slug(workerId || (args.mode === "scout" ? "scout" : "worker"));
352
- return trackInRun(args, track(jobId, { mode: args.mode, workerId: workerId || jobId, lane: jobLane(args), agent: args.agentName ?? null, runId: args.run_id ?? null, role: args.armyRole ?? null, model: args.model ?? null },
353
+ return trackInRun(args, track(jobId, { mode: args.mode, workerId: workerId || jobId, lane: jobLane(args), agent: args.agentName ?? null, runId: args.run_id ?? null, role: args.armyRole ?? null, model: args.model ?? null, label: jobLabel(args) },
353
354
  withAgentSlot(args, jobId, () => executeJob({ ...jobArgs(args, workerId), jobId }))));
354
355
  }
355
356
  // Best-effort progress signal for a job still mid-run: a plain "phase: worker,
@@ -395,7 +396,7 @@ export function createJobRuntime(deps) {
395
396
 
396
397
  async function summarize(entry, files, jobDir = null) {
397
398
  const status = files.status, meta = files.meta ?? files.failure;
398
- const elapsedSeconds = status?.startedAt ? Math.round((Date.now() - Date.parse(status.startedAt)) / 1000) : entry ? Math.round((Date.now() - Date.parse(entry.startedAt)) / 1000) : null;
399
+ const elapsedSeconds = jobElapsedSeconds({ status, meta, entry });
399
400
  const out = { jobId: entry?.jobId ?? status?.jobId ?? meta?.jobId ?? null, workerId: entry?.workerId ?? status?.workerId ?? meta?.workerId ?? null,
400
401
  mode: entry?.mode ?? status?.mode ?? meta?.mode ?? null, state: null, phase: status?.phase ?? "starting", elapsedSeconds,
401
402
  timeoutSeconds: status?.timeoutSeconds ?? null, coordinatorStatus: meta?.coordinatorStatus ?? null, outcome: meta?.outcome ?? null,
@@ -1,3 +1,5 @@
1
+ import path from "node:path";
2
+
1
3
  // ---------------------------------------------------------------------------
2
4
  // Test-change classification (plan 16). One tunable constant, on purpose:
3
5
  // every heuristic about what counts as a test file lives here and nowhere else.
@@ -416,3 +418,91 @@ export function mergeUntrackedIntoNameStatus(nameStatus, untrackedFiles) {
416
418
  return [...(nameStatus ?? []), ...extra];
417
419
  }
418
420
 
421
+
422
+ // ---------------------------------------------------------------------------
423
+ // Verification inputs: .nomarmy.yml is read from the operator's checkout so a
424
+ // worker can't weaken its own checks, but the files its commands run come
425
+ // from the worker's worktree. Found live: a worker whose verification ran
426
+ // `node check.js` rewrote check.js to print PASS and exit 0, and the job
427
+ // committed as done. The revert check can't see it: reverting check.js along
428
+ // with the code makes verification fail, which reads as coverage.
429
+ //
430
+ // Blocking (the diff changes what a command runs): a changed non-test file a
431
+ // command names (`node check.js`, `bash scripts/verify.sh`), a changed
432
+ // Makefile or justfile under `make`/`just`, or a changed package.json script
433
+ // that `npm test`/`pnpm run x`/`yarn x`/`bun run x` calls. Test files named by
434
+ // a command are left to the test-change review, since adding cases is normal.
435
+ // Flagged only: test-runner configuration, which is often a legitimate edit.
436
+ // ---------------------------------------------------------------------------
437
+
438
+ const RUNNER_CONFIG_RE = /(^|\/)(conftest\.py|pytest\.ini|tox\.ini|setup\.cfg|\.coveragerc|(jest|vitest|vite|playwright|karma|cypress)\.config\.[cm]?[jt]s|\.mocharc(\.[a-z]+)?|phpunit\.xml(\.dist)?|\.rspec)$/;
439
+ const TASK_FILES = { make: ["Makefile", "makefile", "GNUmakefile"], just: ["justfile", "Justfile", ".justfile"] };
440
+
441
+ function shellWords(segment) {
442
+ return (segment.match(/"[^"]*"|'[^']*'|\S+/g) ?? []).map((w) => w.replace(/^["']|["']$/g, ""));
443
+ }
444
+
445
+ // Which package.json script a package-manager command runs, if any.
446
+ function packageScript(words) {
447
+ const [tool, first, second] = words;
448
+ if (!["npm", "pnpm", "yarn", "bun"].includes(tool) || !first) return null;
449
+ if (["run", "run-script"].includes(first)) return second && !second.startsWith("-") ? second : null;
450
+ if (tool === "npm") return ["test", "t", "tst"].includes(first) ? "test" : ["start", "stop", "restart"].includes(first) ? first : null;
451
+ if (tool === "bun") return null; // `bun test` is bun's own runner, not a script
452
+ return first.startsWith("-") || ["install", "add", "remove", "exec", "dlx", "x"].includes(first) ? null : first;
453
+ }
454
+
455
+ function scriptsOf(text) {
456
+ try { return JSON.parse(text ?? "")?.scripts ?? {}; } catch { return {}; }
457
+ }
458
+
459
+ function pytestSection(text) {
460
+ const match = /^\[tool\.pytest[^\]]*\]\s*$([\s\S]*?)(?=^\[|(?![\s\S]))/m.exec(String(text ?? ""));
461
+ return match ? match[1].trim() : null;
462
+ }
463
+
464
+ /**
465
+ * @param {{ commands: string[], changedFiles: string[], readBase: (file: string) => string|null|Promise<string|null>,
466
+ * readHead: (file: string) => string|null|Promise<string|null>, isTestPathFn?: (file: string) => boolean }} input
467
+ * @returns {Promise<{ blocked: {file: string, command: string, why: string}[], flagged: {file: string, why: string}[] } | null>}
468
+ */
469
+ export async function detectVerificationInputChanges({ commands = [], changedFiles = [], readBase = () => null, readHead = () => null, isTestPathFn = isTestPath } = {}) {
470
+ const changed = new Set(changedFiles.map((f) => path.posix.normalize(f)));
471
+ if (!changed.size) return null;
472
+ const blocked = [], flagged = [], seen = new Set();
473
+ const block = (file, command, why) => { if (!seen.has(file)) { seen.add(file); blocked.push({ file, command, why }); } };
474
+ for (const command of commands) {
475
+ let dir = "";
476
+ for (const segment of String(command ?? "").split(/\s*(?:&&|\|\||;|\|)\s*/)) {
477
+ const all = shellWords(segment);
478
+ let lead = 0;
479
+ while (lead < all.length && /^[A-Za-z_][A-Za-z0-9_]*=/.test(all[lead])) lead++; // FOO=1 cmd ...
480
+ const words = all.slice(lead);
481
+ if (!words.length) continue;
482
+ const at = (p) => path.posix.normalize(path.posix.join(dir, p));
483
+ if (words[0] === "cd" && words[1]) { dir = at(words[1]); continue; }
484
+ for (const word of words) {
485
+ if (word.startsWith("-") || word.includes("$")) continue;
486
+ const file = at(word);
487
+ if (changed.has(file) && !isTestPathFn(file)) block(file, command, `run by \`${command}\``);
488
+ }
489
+ for (const name of TASK_FILES[words[0]] ?? []) {
490
+ const file = at(name);
491
+ if (changed.has(file)) block(file, command, `read by \`${command}\``);
492
+ }
493
+ const script = packageScript(words);
494
+ const manifest = at("package.json");
495
+ if (script && changed.has(manifest)) {
496
+ const before = scriptsOf(await readBase(manifest)), after = scriptsOf(await readHead(manifest));
497
+ const touched = [script, `pre${script}`, `post${script}`].filter((s) => before[s] !== after[s]);
498
+ if (touched.length) block(manifest, command, `its script${touched.length > 1 ? "s" : ""} ${touched.map((s) => `"${s}"`).join(", ")} changed, and \`${command}\` runs it`);
499
+ }
500
+ }
501
+ }
502
+ for (const file of changed) {
503
+ if (seen.has(file)) continue;
504
+ if (RUNNER_CONFIG_RE.test(file)) flagged.push({ file, why: "test-runner configuration" });
505
+ else if (/(^|\/)pyproject\.toml$/.test(file) && pytestSection(await readBase(file)) !== pytestSection(await readHead(file))) flagged.push({ file, why: "its [tool.pytest] settings changed" });
506
+ }
507
+ return blocked.length || flagged.length ? { blocked, flagged } : null;
508
+ }
package/lib/execute.mjs CHANGED
@@ -17,7 +17,7 @@ import { describeRecoveryChanges, reportRecoveryPrompt } from "./worker-prompt.m
17
17
  import { parseWorkerReport } from "./report.mjs";
18
18
  import { OUTCOMES, COORDINATOR_STATUS_BY_OUTCOME } from "./outcomes.mjs";
19
19
  import { resolveOutcome, finalText, workerMetadata, applyRefactorContract, applyVerificationPolicy } from "./outcome.mjs";
20
- import { isTestPath, isDocumentationPath, detectScopedTestSelectionRisk, detectUnwiredNewDefinitions, detectMislabeledTestNames, detectPossibleSecrets } from "./diff-checks.mjs";
20
+ import { isTestPath, isDocumentationPath, detectScopedTestSelectionRisk, detectUnwiredNewDefinitions, detectMislabeledTestNames, detectPossibleSecrets, detectVerificationInputChanges } from "./diff-checks.mjs";
21
21
 
22
22
  // ---------------------------------------------------------------------------
23
23
  // Job status for polling. `status.json` is written at every phase transition
@@ -250,7 +250,7 @@ export function createExecutor(deps) {
250
250
  // which reads as evidence about work that never happened.
251
251
  independentVerification = normalizeVerification({ status: "not_run", basis: "not-applicable", reason: "the worker changed nothing, so there was none of its work to verify" }, verification ?? null);
252
252
  } else if (verificationFlow.verificationRunner || !reportValidation.valid) {
253
- independentVerification = await runIndependentVerification({ profile: verification ?? null, cwd, jobId, baseSha: base.sha, branch, mode, record: preCommit });
253
+ independentVerification = await runIndependentVerification({ profile: verification ?? null, cwd, jobId, baseSha: base.sha, branch, mode, record: preCommit, logFile: path.join(jobDir, "verification.log") });
254
254
  }
255
255
 
256
256
  // verify_regression: on by default whenever there's a verification
@@ -299,12 +299,19 @@ export function createExecutor(deps) {
299
299
  // Cheap, always-on, additive: never changes commitAllowed/commitBlockedReason
300
300
  // on its own (unlike the regression-check override above), only flags for
301
301
  // review -- see detectScopedTestSelectionRisk's own doc comment for why.
302
- let selectionRisk = null;
302
+ let selectionRisk = null, verificationInputs = null;
303
303
  if (mode === "implement" && verification) {
304
304
  try {
305
305
  const loaded = loadConfig(projectDir); // the operator's contract; see registerVerificationRunner's call
306
306
  const profileCommands = loaded.found ? (loaded.config?.verification?.[verification]?.commands ?? []) : [];
307
307
  selectionRisk = detectScopedTestSelectionRisk({ commands: profileCommands, testChanges: preCommit.testChanges });
308
+ // A diff that changes what those commands run (see
309
+ // detectVerificationInputChanges): applied below, after the others.
310
+ verificationInputs = await detectVerificationInputChanges({
311
+ commands: profileCommands, changedFiles: preCommit.changedFiles ?? [],
312
+ readBase: (file) => gitRaw(["show", `${base.sha}:${file}`], cwd).catch(() => null),
313
+ readHead: (file) => { try { return fs.readFileSync(path.join(cwd, file), "utf8"); } catch { return null; } },
314
+ });
308
315
  } catch { /* a config load failure here is the verification runner's own problem to report, not this check's */ }
309
316
  }
310
317
  const afterSelectionRisk = selectionRisk
@@ -399,7 +406,21 @@ export function createExecutor(deps) {
399
406
  commitBlockedReason: `possible secret detected: ${possibleSecrets.reason}`,
400
407
  reasons: [...afterHostInstalls.reasons, `POSSIBLE SECRET DETECTED: ${possibleSecrets.reason}`] }
401
408
  : afterHostInstalls;
402
- const finalOutcome = applyRefactorContract(applyVerificationPolicy(afterSecrets, independentVerification.status, repoPolicy()),
409
+ // A worker must not be judged by a check it rewrote: a changed script,
410
+ // Makefile or package.json script that a verification command runs
411
+ // blocks the commit; changed test-runner config only asks for review.
412
+ const blockedInputs = verificationInputs?.blocked ?? [], flaggedInputs = verificationInputs?.flagged ?? [];
413
+ const inputLine = blockedInputs.map((b) => `${b.file} (${b.why})`).join("; ");
414
+ const afterInputs = blockedInputs.length
415
+ ? { ...afterSecrets, outcome: afterSecrets.commitAllowed || afterSecrets.outcome === OUTCOMES.WORKER_DONE ? OUTCOMES.NEEDS_REVIEW : afterSecrets.outcome,
416
+ reviewRequired: true, commitAllowed: false,
417
+ commitBlockedReason: afterSecrets.commitAllowed ? `the diff changes what verification runs: ${inputLine}` : afterSecrets.commitBlockedReason,
418
+ reasons: [...afterSecrets.reasons, `VERIFICATION INPUT CHANGED: the diff changes what profile '${verification}' runs, so its result can't be trusted: ${inputLine}`] }
419
+ : afterSecrets;
420
+ const afterConfig = flaggedInputs.length
421
+ ? { ...afterInputs, reviewRequired: true, reasons: [...afterInputs.reasons, `TEST CONFIG CHANGED: ${flaggedInputs.map((c) => `${c.file} (${c.why})`).join("; ")}`] }
422
+ : afterInputs;
423
+ const finalOutcome = applyRefactorContract(applyVerificationPolicy(afterConfig, independentVerification.status, repoPolicy()),
403
424
  { refactor, verificationStatus: independentVerification.status, testChanges: preCommit.testChanges });
404
425
 
405
426
  progress("commit");
@@ -411,6 +432,7 @@ export function createExecutor(deps) {
411
432
  let coordinatorStatus = COORDINATOR_STATUS_BY_OUTCOME[finalOutcome.outcome] ?? "incomplete";
412
433
  const issues = [...finalOutcome.reasons, ...(preCommit.issues ?? [])];
413
434
  if (workerError) issues.push(`worker error: ${String(workerError).split("\n")[0]}`);
435
+ if ((result ?? attempted)?.salvaged) issues.push(`runner cleanup failed after the run (${(result ?? attempted).salvagedFrom}); the worker's report was recovered from the run's transcript`);
414
436
  if (workerFailed || workerTimedOut) { const restarted = vmRestartIssue(vmStartedBefore, deps.podmanVmStartedAt?.() ?? null); if (restarted) issues.unshift(restarted); }
415
437
  if (repositoryChanged && !commit.created) {
416
438
  if (coordinatorStatus === "complete") coordinatorStatus = "incomplete";
@@ -575,6 +597,7 @@ export function createExecutor(deps) {
575
597
  const worker = workerMetadata(result ?? attempted);
576
598
  const issues = [...outcome.reasons];
577
599
  if (workerError) issues.push(`scout error: ${String(workerError).split("\n")[0]}`);
600
+ if ((result ?? attempted)?.salvaged) issues.push(`runner cleanup failed after the run (${(result ?? attempted).salvagedFrom}); the scout's report was recovered from the run's transcript`);
578
601
  const failures = worker.toolSummary?.failures ?? 0; if (failures > 0) issues.push(`scout recorded ${failures} tool failure(s)`);
579
602
  if (dirty) issues.push(`snapshot changed: ${record.repoStatusFiles.join(", ")}`);
580
603
  if (reportRecoveryAttempted) {
@@ -706,6 +729,7 @@ export function createExecutor(deps) {
706
729
  const worker = workerMetadata(result ?? attempted);
707
730
  const issues = [...outcome.reasons];
708
731
  if (workerError) issues.push(`decompose error: ${String(workerError).split("\n")[0]}`);
732
+ if ((result ?? attempted)?.salvaged) issues.push(`runner cleanup failed after the run (${(result ?? attempted).salvagedFrom}); the decomposer's report was recovered from the run's transcript`);
709
733
  const failures = worker.toolSummary?.failures ?? 0; if (failures > 0) issues.push(`decomposer recorded ${failures} tool failure(s)`);
710
734
  if (dirty) issues.push(`snapshot changed: ${record.repoStatusFiles.join(", ")}`);
711
735
  if (overlaps.length) issues.push(`${overlaps.length} subtask pair(s) claim overlapping files; not safe to dispatch as independent jobs as proposed`);
@@ -124,3 +124,27 @@ export function formatUnion(union) {
124
124
  const artifacts = union.worktree ? `\n\nUnion artifacts: ${path.dirname(union.worktree)}\nWorktree retained for review: ${union.worktree}\nBranch retained for review: ${union.branch}` : "";
125
125
  return `${banner}--- UNION RECORD ---\n${JSON.stringify(union, null, 2)}${artifacts}`;
126
126
  }
127
+
128
+ /**
129
+ * What a person calls a job: its commit subject, else the task's first
130
+ * sentence, capped. Two jobs on one agent and model differ here even with no
131
+ * role, so a notification can tell them apart.
132
+ */
133
+ export function jobLabel(args) {
134
+ const text = String(args?.commit_subject || args?.task || "").split("\n")[0].split(/(?<=\.)\s/)[0].trim();
135
+ return text.length > 60 ? `${text.slice(0, 57)}...` : text || null;
136
+ }
137
+
138
+ /**
139
+ * Seconds a job has run: to its finish once it has one, not to whenever it's
140
+ * asked about. The record's total_elapsed covers the whole job; an implement
141
+ * job's finishedAt marks only the worker's end, before verification and commit.
142
+ */
143
+ export function jobElapsedSeconds({ status = null, meta = null, entry = null, now = Date.now() } = {}) {
144
+ const total = meta?.metrics?.total_elapsed;
145
+ if (Number.isFinite(total)) return Math.round(total / 1000);
146
+ const startedMs = Date.parse(status?.startedAt ?? entry?.startedAt ?? meta?.startedAt ?? "");
147
+ if (!Number.isFinite(startedMs)) return null;
148
+ const finishedMs = Date.parse((status?.state === "finished" ? status.updatedAt : null) ?? meta?.finishedAt ?? "");
149
+ return Math.round(((Number.isFinite(finishedMs) ? finishedMs : now) - startedMs) / 1000);
150
+ }
package/lib/outcome.mjs CHANGED
@@ -7,6 +7,16 @@ import { parseWorkerReport } from "./report.mjs";
7
7
  // failed. Recovery exists so that a mangled REPORT cannot destroy correct WORK.
8
8
  // It does not exist to launder a failure into a success.
9
9
  // ---------------------------------------------------------------------------
10
+ // A failed verification's own detail (the command, its exit code, the end of
11
+ // its output), capped so an issue line stays readable; the full output is in
12
+ // the job's verification.log.
13
+ const FAILURE_DETAIL_CHARS = 600;
14
+ function failureDetail(independentVerification) {
15
+ const detail = String(independentVerification?.detail ?? independentVerification?.reason ?? "").trim();
16
+ if (!detail) return "";
17
+ return `: ${detail.length > FAILURE_DETAIL_CHARS ? `${detail.slice(0, FAILURE_DETAIL_CHARS)}...` : detail}`;
18
+ }
19
+
10
20
  export function resolveOutcome({ report, repositoryChanged = false, independentVerification = null, regressionCheck = null, workerFailed = false, workerTimedOut = false, mode = "implement" }) {
11
21
  const verification = independentVerification?.status ?? "not_run";
12
22
  const parsed = report ?? parseWorkerReport("");
@@ -34,7 +44,7 @@ export function resolveOutcome({ report, repositoryChanged = false, independentV
34
44
  if (verification === "fail") {
35
45
  return { ...base, outcome: OUTCOMES.NEEDS_REVIEW, reviewRequired: true,
36
46
  commitBlockedReason: "independent verification failed despite a clean done/pass report",
37
- reasons: ["worker claimed done/pass but independent verification failed"] };
47
+ reasons: [`worker claimed done/pass but independent verification failed${failureDetail(independentVerification)}`] };
38
48
  }
39
49
  // verify_regression: reverting just the production files and re-running
40
50
  // the SAME verification profile still passed (or came back genuinely
@@ -92,7 +102,7 @@ export function resolveOutcome({ report, repositoryChanged = false, independentV
92
102
  if (verification === "fail") {
93
103
  return { ...recovery, outcome: OUTCOMES.WORKER_REPORT_INVALID,
94
104
  commitBlockedReason: "independent verification failed; recovery cannot promote a failure",
95
- reasons: [...recovery.reasons, "independent verification FAILED"] };
105
+ reasons: [...recovery.reasons, `independent verification FAILED${failureDetail(independentVerification)}`] };
96
106
  }
97
107
  if (verification === "pass") {
98
108
  // A leniently recovered `done` plus a passing independent check is the
package/mcp/server.mjs CHANGED
@@ -37,7 +37,7 @@ import { modelRefusals } from "../lib/health.mjs";
37
37
  import { podmanProblem, podmanVmStartedAt } from "../lib/podman-health.mjs";
38
38
  import { restartNotice } from "../lib/install-freshness.mjs";
39
39
  import { createBuildMetrics, resolveOutcome, finalText, workerMetadata, usageMetrics, policyAdmissionProblems, applyRefactorContract, applyVerificationPolicy, resolveVerifyRegression } from "../lib/outcome.mjs";
40
- import { compactJobRecord, formatResult, formatUnion, testChangeBanner, regressionCheckBanner, decomposeOverlapBanner } from "../lib/job-format.mjs";
40
+ import { jobLabel, compactJobRecord, formatResult, formatUnion, testChangeBanner, regressionCheckBanner, decomposeOverlapBanner } from "../lib/job-format.mjs";
41
41
 
42
42
  export { run, mapLimit };
43
43
  export { readsMeasurable, measureReads };
@@ -582,7 +582,7 @@ server.tool("local_workers", "Run independent jobs (implement or scout) with bou
582
582
  // so it was invisible to both ceilings while it ran.
583
583
  // A batch job waits for its agent's slot (up to its own timeout) rather
584
584
  // than failing because an earlier job in the same batch holds it.
585
- return trackInRun(j, track(jobId, { mode: j.mode, workerId, lane: jobLane(j), agent: j.agentName ?? null, runId: j.run_id ?? null, role: j.armyRole ?? null, model: j.model ?? null },
585
+ return trackInRun(j, track(jobId, { mode: j.mode, workerId, lane: jobLane(j), agent: j.agentName ?? null, runId: j.run_id ?? null, role: j.armyRole ?? null, model: j.model ?? null, label: jobLabel(j) },
586
586
  withAgentSlot(j, jobId, () => executeJob({ ...jobArgs(effectiveJob, workerId), jobId }), { waitMs: (j.timeout_seconds ?? 600) * 1000 }))).promise;
587
587
  }, { staggerMs: WORKER_START_STAGGER_MS });
588
588
  indices.forEach((i, laneI) => { results[i] = laneResults[laneI]; });
package/package.json CHANGED
@@ -1,9 +1,9 @@
1
1
  {
2
2
  "name": "nomarmy",
3
- "description": "A harness for AI coding workers whose claims are never trusted: your coding assistant stays in charge while workers implement and test in sandboxes, on local models, API keys or your own subscriptions.",
3
+ "description": "Every byte verified: a harness for AI coding workers whose claims are never trusted. Your coding assistant stays in charge while workers implement and test in sandboxes, and nomArmy checks every change before it is committed.",
4
4
  "author": "Rayson Technologies",
5
5
  "license": "Apache-2.0",
6
- "version": "0.1.0-alpha.7",
6
+ "version": "0.1.0-alpha.8",
7
7
  "private": false,
8
8
  "type": "module",
9
9
  "engines": {