nomarmy 0.1.0-alpha.6 → 0.1.0-alpha.8
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +40 -36
- package/bin/nomarmy.mjs +47 -11
- package/docker/Dockerfile.go +5 -3
- package/docker/Dockerfile.rust +16 -1
- package/install.sh +16 -2
- package/lib/admission.mjs +4 -3
- package/lib/agents.mjs +6 -5
- package/lib/army.mjs +2 -1
- package/lib/connect.mjs +2 -0
- package/lib/diff-checks.mjs +90 -0
- package/lib/dispatch-schema.mjs +12 -11
- package/lib/execute.mjs +28 -4
- package/lib/health.mjs +10 -3
- package/lib/install-freshness.mjs +83 -0
- package/lib/job-format.mjs +24 -0
- package/lib/outcome.mjs +12 -2
- package/lib/schema.mjs +9 -11
- package/lib/zod-issues.mjs +15 -0
- package/mcp/server.mjs +17 -4
- package/package.json +4 -4
package/README.md
CHANGED
|
@@ -6,20 +6,22 @@
|
|
|
6
6
|
<a href="https://github.com/rayson-tech/nomarmy/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-Apache%202.0-blue.svg" alt="License: Apache 2.0"></a>
|
|
7
7
|
</p>
|
|
8
8
|
|
|
9
|
-
<p align="center"><em>
|
|
9
|
+
<p align="center"><em>Every byte verified.</em> 🍪</p>
|
|
10
10
|
|
|
11
11
|
**Your coding assistant plans; sandboxed workers build; nothing counts until nomArmy has checked it.**
|
|
12
12
|
|
|
13
|
+
AI coding workers are confident. Their "done, all tests pass" is a claim, not evidence. nomArmy lets your coding assistant (Claude Code, Codex or Cursor) hand work to workers called **noms**, then checks every change itself before anything is committed: the real diff, your tests run in a fresh sandbox, a check that those tests actually catch the change, and a secret scan.
|
|
14
|
+
|
|
13
15
|
## TL;DR
|
|
14
16
|
|
|
15
|
-
1. **Have** Git, Node
|
|
17
|
+
1. **Have** Git, Node 24.16+ (or 26.1+) and [Podman](https://podman.io). On macOS, give Podman 8 GiB: `brew install podman && podman machine init --memory 8192 && podman machine start`.
|
|
16
18
|
2. **Install and set up:**
|
|
17
19
|
```bash
|
|
18
20
|
npm install -g nomarmy@alpha
|
|
19
21
|
cd your-project
|
|
20
22
|
nomarmy setup
|
|
21
23
|
```
|
|
22
|
-
`nomarmy setup` is
|
|
24
|
+
`nomarmy setup` is a playbook. It shows a checklist and runs the next step each time you say yes:
|
|
23
25
|
```text
|
|
24
26
|
✓ Where models run: hosted
|
|
25
27
|
✓ Installed: OpenClaw 2026.9.6
|
|
@@ -29,34 +31,39 @@
|
|
|
29
31
|
Check: verify the installation
|
|
30
32
|
Run `nomarmy agents add` now? [Y/n]
|
|
31
33
|
```
|
|
32
|
-
|
|
33
|
-
|
|
34
|
-
**Want every step spelled out?** [Example setup: Claude Code, Codex and an API key](https://github.com/rayson-tech/nomarmy/blob/main/docs/setup/example.md) walks through a complete setup, command by command.
|
|
34
|
+
Stop anytime; `nomarmy setup` picks up where you left off. **Want every step spelled out?** [Example setup: Claude Code, Codex and an API key](https://github.com/rayson-tech/nomarmy/blob/main/docs/setup/example.md) goes command by command.
|
|
35
35
|
3. **Use it:** restart Claude Code in the project and ask it to use nomArmy for one small bug that has a test. When that works, try `/feature <what you want built>`.
|
|
36
36
|
|
|
37
|
-
|
|
37
|
+
Stuck? `nomarmy doctor` checks the machine and `nomarmy health` checks everything nomArmy runs on. Upgrading later? `nomarmy update`.
|
|
38
38
|
|
|
39
|
-
|
|
39
|
+
## How every byte gets verified
|
|
40
40
|
|
|
41
|
-
|
|
41
|
+
1. Your coding assistant, the **General**, briefs a job: a task, acceptance criteria, and the tests that prove it.
|
|
42
|
+
2. nomArmy creates a git worktree from your branch and runs the nom in a Podman sandbox with no network and no host credentials. (One exception, the Claude subscription: see [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md#security-posture).)
|
|
43
|
+
3. The nom edits, runs tests, and ends with a four-line report: `STATUS`, `TESTS`, `NOT_DONE`, `NOTE`.
|
|
44
|
+
4. nomArmy treats that report as a claim and checks the evidence itself:
|
|
45
|
+
- reads the real diff from git, not the nom's description of it
|
|
46
|
+
- runs your verification profile in a fresh sandbox
|
|
47
|
+
- reverts the production change and reruns the tests: a test that still passes proves nothing, so the job goes to review instead of being committed
|
|
48
|
+
- blocks on secrets, and flags tests made to pass (new skips, stubbed imports) and code nothing calls
|
|
49
|
+
5. Only then does it commit, on the nom's own branch. It never merges into yours: reviewing and integrating stay with the General, and with you.
|
|
42
50
|
|
|
43
|
-
|
|
51
|
+
Failing verification stays failed, unconditionally. A malformed report isn't automatically a failure: if the repository changed, nomArmy verifies independently and may recover the work. And the checks aren't the General's to waive: a repo's `.nomarmy.yml` policy (on by default for new repos) makes verification and the revert check mandatory for every job.
|
|
44
52
|
|
|
45
|
-
**
|
|
53
|
+
**Checking without building** costs nothing: `mode: verify` runs a verification profile against any branch, with no worker and no model tokens.
|
|
46
54
|
|
|
47
|
-
|
|
55
|
+
## Where the work runs
|
|
48
56
|
|
|
49
|
-
|
|
57
|
+
- **Agents** say where a job can run: an API key, your own ChatGPT or Muse Code subscription, or a local model on llama.cpp.
|
|
58
|
+
- **The army** says which role runs on which agent: Sr and Jr devs build, a security analyst and a data architect review, a PM checks the plan, a PO accepts.
|
|
59
|
+
- **`/feature`** runs a whole feature end to end, from plan through build, review and acceptance, and hands you a branch to merge.
|
|
60
|
+
- **Harnesses** give each repo the right sandbox: Go, Rust, Python and Node (mixed repos too), Playwright browser tests, and fake services like a mock login server, all offline. [Adding one](https://github.com/rayson-tech/nomarmy/blob/main/CONTRIBUTING.md#adding-a-harness) never touches core code.
|
|
50
61
|
|
|
51
|
-
|
|
52
|
-
2. nomArmy creates a worktree from your branch and runs the worker in a Podman sandbox with no network and no host credentials. (One exception, the Claude subscription: see [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md#security-posture).)
|
|
53
|
-
3. The worker edits, runs tests, and ends with a four-line report: `STATUS`, `TESTS`, `NOT_DONE`, `NOTE`.
|
|
54
|
-
4. nomArmy treats that report as a claim. It reads the real diff from git, runs your verification profile itself in a fresh sandbox, reverts the production change to check the tests actually fail without it, and scans for secrets.
|
|
55
|
-
5. Only then does it commit, on the worker's own branch. It never merges into yours: reviewing and integrating stay with the General, and with you.
|
|
62
|
+
**What you get is work you don't have to take on faith**, not cheaper work. Delegating costs the General tokens too, for briefing and review: on small, already-diagnosed tickets we measured 4 to 8 times more of the General's tokens than fixing the bug directly, with break-even around 150 lines of context a fix needs to read ([the measurements](https://github.com/rayson-tech/nomarmy/blob/main/docs/experiments/2026-09-20-model-bakeoff-and-economics.md)). It pays off on bigger tickets, parallel work, and anywhere you'd otherwise trust an agent's say-so.
|
|
56
63
|
|
|
57
|
-
|
|
64
|
+
**Have a GPU or a Mac with plenty of memory?** Choose "a local model" in `nomarmy setup`: no per-token bill and your code stays home, but you pay in hardware, power and speed. `nomarmy sizing` tells you what fits. A [shared model server](https://github.com/rayson-tech/nomarmy/blob/main/docs/install.md#a-shared-model-server) works too.
|
|
58
65
|
|
|
59
|
-
|
|
66
|
+
Developed and maintained by Rayson Technologies. This is an alpha (`0.1.0-alpha`).
|
|
60
67
|
|
|
61
68
|
## Docs
|
|
62
69
|
|
|
@@ -66,7 +73,7 @@ Around that core: **agents** say where a job can run, the **army** says which ro
|
|
|
66
73
|
| [Example setup](https://github.com/rayson-tech/nomarmy/blob/main/docs/setup/example.md) | Claude Code, Codex and an API key, command by command |
|
|
67
74
|
| [Agents and the army](https://github.com/rayson-tech/nomarmy/blob/main/docs/agents-and-army.md) | Where a job can run, who does what, usage limits, picking an agent |
|
|
68
75
|
| [`/feature` runs](https://github.com/rayson-tech/nomarmy/blob/main/docs/feature-runs.md) | A feature end to end, and watching what nomArmy is doing |
|
|
69
|
-
| [Your repository](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md) | `.nomarmy.yml`, verification,
|
|
76
|
+
| [Your repository](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md) | `.nomarmy.yml`, verification, dependencies, private registries, what nomArmy checks |
|
|
70
77
|
| [Harnesses](https://github.com/rayson-tech/nomarmy/blob/main/docs/harnesses.md) | Ecosystem registry, detection, network levels, and requirements |
|
|
71
78
|
| [Configuration](https://github.com/rayson-tech/nomarmy/blob/main/docs/configuration.md) | Settings, swapping the local model, sizing, admission |
|
|
72
79
|
| [Reference](https://github.com/rayson-tech/nomarmy/blob/main/docs/reference.md) | Every CLI command and MCP tool |
|
|
@@ -75,23 +82,20 @@ Around that core: **agents** say where a job can run, the **army** says which ro
|
|
|
75
82
|
|
|
76
83
|
## Security
|
|
77
84
|
|
|
78
|
-
A
|
|
85
|
+
A nom gets a writable git worktree inside a Podman sandbox and nothing else: no network, no host credentials, no Podman socket. Every model call is made by OpenClaw on your machine, never from inside the sandbox. Verification can climb a network ladder one rung at a time (fake services on a private network, then an allowlist you approve for a test tenant), but noms never leave `network none`. **The exception is a Claude subscription**, whose tools run on your machine, so nomArmy refuses build jobs on it unless you allow it. Never hand a nom production credentials, deployment access or SSH keys. Details: [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md); to report a vulnerability, [SECURITY.md](https://github.com/rayson-tech/nomarmy/blob/main/SECURITY.md).
|
|
79
86
|
|
|
80
87
|
## Status
|
|
81
88
|
|
|
82
89
|
| Capability | Status |
|
|
83
90
|
|---|---|
|
|
84
|
-
|
|
|
85
|
-
|
|
|
86
|
-
| Scout and decompose modes | Unit and live tested |
|
|
87
|
-
|
|
|
88
|
-
| Secret scanning (secretlint, hard block) | Unit tested against the real dependency |
|
|
89
|
-
| Agents: api keys | Live-verified with xAI; other providers built to OpenClaw's documented interface |
|
|
91
|
+
| Verification core: worktrees, nomArmy-owned git, independent verification, the revert check, kept failed worktrees | Working, end-to-end tested |
|
|
92
|
+
| `mode: verify`, secret scanning (secretlint, hard block), test-workaround and unwired-code checks | Unit and live tested; the heuristics are review flags |
|
|
93
|
+
| Scout and decompose modes, `auto_union` | Unit and live tested |
|
|
94
|
+
| Agents: API keys | Live-verified with xAI; other providers built to OpenClaw's documented interface |
|
|
90
95
|
| Agents: subscriptions | ChatGPT (Codex) and Muse Code sandboxed and live-verified; Claude live-verified, but its tools run on the host (scout and review by default) |
|
|
96
|
+
| Local (llama.cpp) and Bedrock profiles | Working |
|
|
91
97
|
| The army and `/feature` | Driven by a real Claude Code General across three runs, about 18 implement jobs |
|
|
92
|
-
| Go, Rust, Python
|
|
93
|
-
| Fake services beside the app (mock login server, mock APIs) | Live-verified on a private network with no route out (the `services` harness level) |
|
|
94
|
-
| Browser tests (Playwright + Chromium) | Live-verified offline, with screenshots and traces kept as job evidence |
|
|
98
|
+
| Harnesses: Go, Rust, Python, Node and mixed repos; Playwright; fake services | Live-verified offline |
|
|
95
99
|
| Private registries and a verification-only network allowlist | Live-verified; each passed an independent security review |
|
|
96
100
|
|
|
97
101
|
What we've learned from real runs, including where delegating pays and where it doesn't, is in [docs/findings.md](https://github.com/rayson-tech/nomarmy/blob/main/docs/findings.md).
|
|
@@ -99,13 +103,13 @@ What we've learned from real runs, including where delegating pays and where it
|
|
|
99
103
|
### Known limitations
|
|
100
104
|
|
|
101
105
|
- **A Claude subscription isn't sandboxed.** Its tools run on your machine, so implement jobs on it are refused unless you set `allow_host_tools: true`. See [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md#security-posture).
|
|
102
|
-
- **
|
|
103
|
-
- **
|
|
106
|
+
- **Your own compose services aren't started yet.** A verification profile that needs a real database from your compose file (`environment: basic` or higher) reports `not_run` rather than running without it (and the compose parser doesn't resolve YAML anchors). Fake services from harnesses, like the mock login server, do run.
|
|
107
|
+
- **Private registries don't cover Poetry or Yarn Berry** yet: Poetry can't guarantee a credentialed install runs no package code, and Yarn Berry doesn't read `.npmrc`. uv, pip wheels, npm, pnpm, Yarn Classic and bun work. See [Private registries](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md#private-registries).
|
|
108
|
+
- **A refused model costs one job.** When a vendor refuses a model at run time that OpenClaw lists (gpt-6-sol on a ChatGPT plan), the first job on it fails with `model_not_found`; after that nomArmy won't dispatch it until a job or test call on it works. `army assign` tests the route and catches this before any job.
|
|
109
|
+
- **Claude subscription token counts** come from the Claude CLI's own session log, since OpenClaw sees only the final reply; totals include cache reads and writes.
|
|
104
110
|
- **Test-workaround detection is a flag, not a verdict**: a legitimate new skip still gets flagged.
|
|
105
111
|
- **Deploy-time failures need your own check.** See [Add a check for what unit tests can't see](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md#nomarmyyml).
|
|
106
|
-
- **
|
|
107
|
-
- **Verification needing services** (a database, a mock server) reports `not_run` instead of running without them. The compose parser doesn't resolve YAML anchors.
|
|
108
|
-
- **Same-host sandboxes**: the MCP server, OpenClaw and every job's sandbox run on the machine with the coordinator. Only the model can be elsewhere (an agent, or [a shared model server](https://github.com/rayson-tech/nomarmy/blob/main/docs/install.md#a-shared-model-server)).
|
|
112
|
+
- **Same-host sandboxes**: the MCP server, OpenClaw and every job's sandbox run on the machine with the coordinator. Only the model can be elsewhere.
|
|
109
113
|
|
|
110
114
|
## More
|
|
111
115
|
|
package/bin/nomarmy.mjs
CHANGED
|
@@ -18,7 +18,8 @@ import { buildConfigProposal } from "../lib/propose.mjs";
|
|
|
18
18
|
import { detectHardware } from "../lib/hardware.mjs";
|
|
19
19
|
import { readGGUFMetadata, resolveModelPath, totalSplitBytes } from "../lib/gguf.mjs";
|
|
20
20
|
import { recommend, customRecommendation, evaluateConfig, bytesPerKvElementForCacheTypes, MIN_CONTEXT_PER_NOM } from "../lib/sizing.mjs";
|
|
21
|
-
import { connectClaude, connectCodex, connectCursor, cursorAlreadyConnected, deriveWorkerModelEnv } from "../lib/connect.mjs";
|
|
21
|
+
import { connectClaude, connectCodex, connectCursor, cursorAlreadyConnected, deriveWorkerModelEnv, defaultInstallDir } from "../lib/connect.mjs";
|
|
22
|
+
import { compareVersions, readPackageVersion, readInstallVersions } from "../lib/install-freshness.mjs";
|
|
22
23
|
import { ID_RE, AUTH_ENV_NAME_RE, OPENCLAW_PROVIDER_ID_RE, openclawProviderId, isNativeProviderType } from "../lib/dispatch-schema.mjs";
|
|
23
24
|
import { loadAgents, readAgentsFile, writeAgentsFile, agentsConfigPath, apiAgentAsPoolEntry, describeAgent as describeAgentLabel, agentRunsToolsOnHost, agentProviderId, AGENT_KINDS, API_PROVIDER_TYPES, RESERVED_AGENT_NAMES, BUILTIN_LOCAL_AGENT } from "../lib/agents.mjs";
|
|
24
25
|
import { loadArmy, mergeArmy, describeArmy, readArmyFile, updateArmyInFile, assignRoleInFile, parseTargetSpec, armyLayerPath, globalConfigDir, DEFAULT_ARMY, ARMY_PHASES, LOCAL_CONFIG_FILENAME } from "../lib/army.mjs";
|
|
@@ -115,8 +116,9 @@ Usage: nomarmy <command> [options]
|
|
|
115
116
|
registration (never done silently)
|
|
116
117
|
--restart-inference with --json, also stop/start local
|
|
117
118
|
inference (never done silently)
|
|
118
|
-
update
|
|
119
|
-
|
|
119
|
+
update Update nomArmy and reconnect your coordinators: installs
|
|
120
|
+
npm's latest alpha, or for a git checkout pulls (fast-forward
|
|
121
|
+
only; refuses on local changes). Then restart open sessions.
|
|
120
122
|
agents <list|add|update|remove>
|
|
121
123
|
Every account a job can run on, in one list:
|
|
122
124
|
~/.config/nomarmy/agents.yml (or NOMARMY_CONFIG_DIR).
|
|
@@ -1642,14 +1644,9 @@ function git(args) {
|
|
|
1642
1644
|
*/
|
|
1643
1645
|
async function cmdUpdate() {
|
|
1644
1646
|
const say = (s) => { if (!json) console.log(s); };
|
|
1645
|
-
// Installed from npm:
|
|
1646
|
-
//
|
|
1647
|
-
if (!fs.existsSync(path.join(nomarmyRoot, ".git")))
|
|
1648
|
-
const how = "npm install -g nomarmy@alpha && nomarmy connect";
|
|
1649
|
-
if (json) return out({ error: "installed from npm, not a git checkout", fix: how });
|
|
1650
|
-
console.log(`This nomArmy was installed from npm, so there's nothing to pull. Update with:\n ${how}`);
|
|
1651
|
-
return;
|
|
1652
|
-
}
|
|
1647
|
+
// Installed from npm: npm updates the package, then connect resyncs the
|
|
1648
|
+
// copy each coordinator runs.
|
|
1649
|
+
if (!fs.existsSync(path.join(nomarmyRoot, ".git"))) return updateFromNpm();
|
|
1653
1650
|
const status = git(["status", "--porcelain"]);
|
|
1654
1651
|
if (status) {
|
|
1655
1652
|
if (json) { out({ error: "working tree is not clean; refusing to pull over local changes", status }); process.exit(1); }
|
|
@@ -1704,6 +1701,45 @@ async function cmdUpdate() {
|
|
|
1704
1701
|
console.log(c.yellow("\nThe MCP server is a per-session child process: every open Claude Code / Codex / Cursor session needs a restart to pick this up, not just this one."));
|
|
1705
1702
|
}
|
|
1706
1703
|
|
|
1704
|
+
// The coordinators nomArmy is registered with. Cursor has no CLI to probe,
|
|
1705
|
+
// so it counts when its own config already lists nomArmy.
|
|
1706
|
+
function connectedTargets() {
|
|
1707
|
+
return [commandExists("claude") && "claude", commandExists("codex") && "codex", cursorAlreadyConnected() && "cursor"].filter(Boolean);
|
|
1708
|
+
}
|
|
1709
|
+
|
|
1710
|
+
async function updateFromNpm() {
|
|
1711
|
+
const current = readPackageVersion(nomarmyRoot);
|
|
1712
|
+
let latest = null;
|
|
1713
|
+
try { latest = execFileSync("npm", ["view", "nomarmy", "dist-tags.alpha"], { encoding: "utf8", timeout: 20000 }).trim(); } catch { /* offline */ }
|
|
1714
|
+
if (!latest) {
|
|
1715
|
+
const fix = "npm install -g nomarmy@alpha && nomarmy connect claude";
|
|
1716
|
+
if (json) { out({ error: "could not read nomarmy's latest release from npm", fix }); process.exit(1); }
|
|
1717
|
+
console.log(c.red("Couldn't reach npm to find nomArmy's latest release.") + ` Update by hand:\n ${fix}`);
|
|
1718
|
+
process.exit(1);
|
|
1719
|
+
}
|
|
1720
|
+
const { copyVersion } = readInstallVersions(defaultInstallDir());
|
|
1721
|
+
const upgrade = compareVersions(current, latest) < 0;
|
|
1722
|
+
const staleCopy = !copyVersion || compareVersions(copyVersion, upgrade ? latest : current) < 0;
|
|
1723
|
+
if (!upgrade && !staleCopy) {
|
|
1724
|
+
if (json) return out({ updated: false, version: current, reason: "already up to date" });
|
|
1725
|
+
console.log(c.green(`✓ nomArmy ${current} is the latest, and your coordinators run it.`));
|
|
1726
|
+
return;
|
|
1727
|
+
}
|
|
1728
|
+
if (upgrade) {
|
|
1729
|
+
if (!json) console.log(c.bold(`🍪 Updating nomArmy ${current} → ${latest}\n`));
|
|
1730
|
+
execFileSync("npm", ["install", "-g", `nomarmy@${latest}`, "--no-audit", "--no-fund"], { stdio: json ? "ignore" : "inherit" });
|
|
1731
|
+
}
|
|
1732
|
+
// A child process, so the reconnect runs the code just installed rather
|
|
1733
|
+
// than the old code this process loaded.
|
|
1734
|
+
const targets = connectedTargets();
|
|
1735
|
+
if (targets.length) {
|
|
1736
|
+
if (!json) console.log(`\nReconnecting ${targets.join(", ")}...`);
|
|
1737
|
+
execFileSync(process.execPath, [path.join(nomarmyRoot, "bin", "nomarmy.mjs"), "connect", ...targets, ...(json ? ["--json"] : [])], { stdio: json ? "ignore" : "inherit" });
|
|
1738
|
+
}
|
|
1739
|
+
if (json) return out({ updated: upgrade, from: current, version: upgrade ? latest : current, resynced: targets });
|
|
1740
|
+
console.log(c.yellow("\nRestart every open Claude Code, Codex and Cursor session: each keeps the code it started with until then."));
|
|
1741
|
+
}
|
|
1742
|
+
|
|
1707
1743
|
function commandExists(cmd) {
|
|
1708
1744
|
try { execFileSync(process.platform === "win32" ? "where" : "which", [cmd], { stdio: "ignore" }); return true; }
|
|
1709
1745
|
catch { return false; }
|
package/docker/Dockerfile.go
CHANGED
|
@@ -9,15 +9,17 @@ RUN apt-get update && apt-get install -y --no-install-recommends \
|
|
|
9
9
|
# (verified live: apt installed 1.19 here while go.dev's stable was 1.27) --
|
|
10
10
|
# same reasoning as using rustup instead of apt for Rust in Dockerfile.rust.
|
|
11
11
|
# Pinned, not "latest", for a reproducible build; bump GO_VERSION by hand
|
|
12
|
-
# periodically against https://go.dev/dl
|
|
12
|
+
# periodically against https://go.dev/dl/, with both SHA256 values from
|
|
13
|
+
# https://go.dev/dl/?mode=json.
|
|
13
14
|
ENV GO_VERSION=1.27.1
|
|
14
15
|
RUN ARCH="$(dpkg --print-architecture)" \
|
|
15
16
|
&& case "$ARCH" in \
|
|
16
|
-
amd64) GOARCH=amd64 ;; \
|
|
17
|
-
arm64) GOARCH=arm64 ;; \
|
|
17
|
+
amd64) GOARCH=amd64; SHA256=63d339f0da5ab53635a56f2490a7984dfe12dfcff22ad749f63edaf590168445 ;; \
|
|
18
|
+
arm64) GOARCH=arm64; SHA256=3450b45a3f9ee8568792736a5c5e70a1f2e9b36c35a8f74958c03e51d7d92bec ;; \
|
|
18
19
|
*) echo "unsupported architecture for Go install: $ARCH" >&2; exit 1 ;; \
|
|
19
20
|
esac \
|
|
20
21
|
&& curl -fsSL "https://go.dev/dl/go${GO_VERSION}.linux-${GOARCH}.tar.gz" -o /tmp/go.tgz \
|
|
22
|
+
&& echo "${SHA256} /tmp/go.tgz" | sha256sum -c - \
|
|
21
23
|
&& tar -C /usr/local -xzf /tmp/go.tgz \
|
|
22
24
|
&& rm /tmp/go.tgz
|
|
23
25
|
ENV PATH="/usr/local/go/bin:${PATH}"
|
package/docker/Dockerfile.rust
CHANGED
|
@@ -14,6 +14,21 @@ USER node
|
|
|
14
14
|
ENV RUSTUP_HOME=/home/node/.rustup
|
|
15
15
|
ENV CARGO_HOME=/home/node/.cargo
|
|
16
16
|
ENV PATH="${CARGO_HOME}/bin:${PATH}"
|
|
17
|
-
|
|
17
|
+
# rustup-init is pinned by version and checksum instead of piping
|
|
18
|
+
# sh.rustup.rs into sh; rustup then verifies the toolchain it downloads.
|
|
19
|
+
# Bump RUSTUP_VERSION and both hashes together, from
|
|
20
|
+
# https://static.rust-lang.org/rustup/archive/<version>/<target>/rustup-init.sha256.
|
|
21
|
+
ENV RUSTUP_VERSION=1.29.1
|
|
22
|
+
RUN ARCH="$(dpkg --print-architecture)" \
|
|
23
|
+
&& case "$ARCH" in \
|
|
24
|
+
amd64) TARGET=x86_64-unknown-linux-gnu; SHA256=dda7234360b7f578ca8b0ddcb80145646fa61a67c1720a5abc7051b35c9fcb71 ;; \
|
|
25
|
+
arm64) TARGET=aarch64-unknown-linux-gnu; SHA256=15f6e4ce9f583b929c996c91562bad6d4454f3281de858b02cdfdef615fac433 ;; \
|
|
26
|
+
*) echo "unsupported architecture for Rust install: $ARCH" >&2; exit 1 ;; \
|
|
27
|
+
esac \
|
|
28
|
+
&& curl --proto '=https' --tlsv1.2 -fsSL "https://static.rust-lang.org/rustup/archive/${RUSTUP_VERSION}/${TARGET}/rustup-init" -o /tmp/rustup-init \
|
|
29
|
+
&& echo "${SHA256} /tmp/rustup-init" | sha256sum -c - \
|
|
30
|
+
&& chmod +x /tmp/rustup-init \
|
|
31
|
+
&& /tmp/rustup-init -y --no-modify-path --profile minimal --default-toolchain stable \
|
|
32
|
+
&& rm /tmp/rustup-init
|
|
18
33
|
WORKDIR /workspace
|
|
19
34
|
CMD ["sleep", "infinity"]
|
package/install.sh
CHANGED
|
@@ -99,9 +99,23 @@ fi
|
|
|
99
99
|
# on a fresh macOS install nothing has initialized that VM yet at this point.
|
|
100
100
|
if nomarmy_is_cloud; then need aws || { echo 'ERROR: the AWS CLI is required for cloud profiles.'; exit 1; }; fi
|
|
101
101
|
"$ROOT/scripts/install-llama-cpp.sh" "$NOMARMY_PROFILE"
|
|
102
|
+
# A pinned npm release, which npm checks against the registry's integrity
|
|
103
|
+
# hash, rather than piping a remote installer script into bash.
|
|
104
|
+
OPENCLAW_VERSION="${NOMARMY_OPENCLAW_VERSION:-2026.9.6}"
|
|
102
105
|
if ! command -v openclaw >/dev/null 2>&1; then
|
|
103
|
-
|
|
104
|
-
|
|
106
|
+
if ! node -e 'const [a, b] = process.versions.node.split(".").map(Number); process.exit((a === 24 && b >= 16) || (a === 26 && b >= 1) || a > 26 ? 0 : 1)'; then
|
|
107
|
+
echo "ERROR: OpenClaw $OPENCLAW_VERSION needs Node 24.16+ or 26.1+; this is Node $(node -v)."
|
|
108
|
+
echo ' Upgrade Node and re-run this installer, or install OpenClaw yourself first:'
|
|
109
|
+
echo ' https://openclaw.ai (then re-run this installer)'
|
|
110
|
+
exit 1
|
|
111
|
+
fi
|
|
112
|
+
echo "==> Installing OpenClaw $OPENCLAW_VERSION from npm"
|
|
113
|
+
if [[ -w "$(npm prefix -g)" ]]; then
|
|
114
|
+
npm install -g "openclaw@$OPENCLAW_VERSION" --no-audit --no-fund
|
|
115
|
+
else
|
|
116
|
+
# No writable global prefix: install for this user, as OpenClaw's own installer does.
|
|
117
|
+
npm install -g --prefix "$HOME/.npm-global" "openclaw@$OPENCLAW_VERSION" --no-audit --no-fund
|
|
118
|
+
fi
|
|
105
119
|
export PATH="$HOME/.local/bin:$HOME/.npm-global/bin:$PATH"
|
|
106
120
|
fi
|
|
107
121
|
need openclaw
|
package/lib/admission.mjs
CHANGED
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
import fs from "node:fs";
|
|
2
2
|
import path from "node:path";
|
|
3
3
|
import { executionMode } from "./execution.mjs";
|
|
4
|
+
import { jobLabel, jobElapsedSeconds } from "./job-format.mjs";
|
|
4
5
|
import { loadConfig } from "./config.mjs";
|
|
5
6
|
import { clampInt } from "./budget-state.mjs";
|
|
6
7
|
import { checkBrief, assessAdmission, describeBudgets } from "./budget.mjs";
|
|
@@ -160,7 +161,7 @@ export function createJobRuntime(deps) {
|
|
|
160
161
|
const who = entry.mode === "verify" ? "verification runner" : entry.agent ? `${entry.agent}${entry.model ? `/${entry.model}` : ""}` : "local model";
|
|
161
162
|
const took = Math.round((Date.now() - Date.parse(entry.startedAt)) / 60000);
|
|
162
163
|
const ok = !error && (result?.ok || m.coordinatorStatus === "complete");
|
|
163
|
-
notify(`nomArmy: ${entry.role ?? entry.mode ?? "job"} ${ok ? "done" : outcome}`, `${entry.workerId ?? entry.jobId} on ${who}: ${outcome} after ${took}m. ${ok ? "Ready for the General's review." : "Needs a look."}`);
|
|
164
|
+
notify(`nomArmy: ${entry.role ?? entry.mode ?? "job"} ${ok ? "done" : outcome}`, `${entry.label ? `${entry.label} (${entry.workerId ?? entry.jobId})` : entry.workerId ?? entry.jobId} on ${who}: ${outcome} after ${took}m. ${ok ? "Ready for the General's review." : "Needs a look."}`);
|
|
164
165
|
}
|
|
165
166
|
function admissionHardware() {
|
|
166
167
|
return executionMode(deps.env).managesModelServer ? deps.budgetState.hardwareSnapshot : null;
|
|
@@ -349,7 +350,7 @@ export function createJobRuntime(deps) {
|
|
|
349
350
|
function launch(args) {
|
|
350
351
|
const workerId = args.worker_id || null;
|
|
351
352
|
const jobId = slug(workerId || (args.mode === "scout" ? "scout" : "worker"));
|
|
352
|
-
return trackInRun(args, track(jobId, { mode: args.mode, workerId: workerId || jobId, lane: jobLane(args), agent: args.agentName ?? null, runId: args.run_id ?? null, role: args.armyRole ?? null, model: args.model ?? null },
|
|
353
|
+
return trackInRun(args, track(jobId, { mode: args.mode, workerId: workerId || jobId, lane: jobLane(args), agent: args.agentName ?? null, runId: args.run_id ?? null, role: args.armyRole ?? null, model: args.model ?? null, label: jobLabel(args) },
|
|
353
354
|
withAgentSlot(args, jobId, () => executeJob({ ...jobArgs(args, workerId), jobId }))));
|
|
354
355
|
}
|
|
355
356
|
// Best-effort progress signal for a job still mid-run: a plain "phase: worker,
|
|
@@ -395,7 +396,7 @@ export function createJobRuntime(deps) {
|
|
|
395
396
|
|
|
396
397
|
async function summarize(entry, files, jobDir = null) {
|
|
397
398
|
const status = files.status, meta = files.meta ?? files.failure;
|
|
398
|
-
const elapsedSeconds =
|
|
399
|
+
const elapsedSeconds = jobElapsedSeconds({ status, meta, entry });
|
|
399
400
|
const out = { jobId: entry?.jobId ?? status?.jobId ?? meta?.jobId ?? null, workerId: entry?.workerId ?? status?.workerId ?? meta?.workerId ?? null,
|
|
400
401
|
mode: entry?.mode ?? status?.mode ?? meta?.mode ?? null, state: null, phase: status?.phase ?? "starting", elapsedSeconds,
|
|
401
402
|
timeoutSeconds: status?.timeoutSeconds ?? null, coordinatorStatus: meta?.coordinatorStatus ?? null, outcome: meta?.outcome ?? null,
|
package/lib/agents.mjs
CHANGED
|
@@ -30,6 +30,7 @@ import fs from "node:fs";
|
|
|
30
30
|
import path from "node:path";
|
|
31
31
|
import YAML from "yaml";
|
|
32
32
|
import { z } from "zod";
|
|
33
|
+
import { typeError, isMissingField, isUnknownDiscriminator } from "./zod-issues.mjs";
|
|
33
34
|
|
|
34
35
|
import { providerEntrySchema, formatDispatchIssues, openclawProviderId, PROVIDER_TYPES, ID_RE, HOST_TOOL_PROVIDERS, agentRunsToolsOnHost } from "./dispatch-schema.mjs";
|
|
35
36
|
|
|
@@ -45,10 +46,10 @@ export const RESERVED_AGENT_NAMES = Object.freeze(["__proto__", "constructor", "
|
|
|
45
46
|
export const BUILTIN_LOCAL_AGENT = Object.freeze({ kind: "local", slot: "coder" });
|
|
46
47
|
|
|
47
48
|
const requiredString = () =>
|
|
48
|
-
z.string({
|
|
49
|
+
z.string({ error: typeError("a string") })
|
|
49
50
|
.refine((value) => value.trim().length > 0, { message: "must not be empty" });
|
|
50
51
|
const thinkingSchema = z.union([z.boolean(), z.enum(["low", "medium", "high"])]).default(true);
|
|
51
|
-
const positiveInt = () => z.number({
|
|
52
|
+
const positiveInt = () => z.number({ error: typeError("a number") }).int("must be a whole number").positive("must be a positive number");
|
|
52
53
|
|
|
53
54
|
const localAgentSchema = z.object({
|
|
54
55
|
kind: z.literal("local"),
|
|
@@ -62,7 +63,7 @@ const localAgentSchema = z.object({
|
|
|
62
63
|
// providerEntrySchema in agentsFileSchema's superRefine, so they exist once.
|
|
63
64
|
const apiAgentSchema = z.object({
|
|
64
65
|
kind: z.literal("api"),
|
|
65
|
-
provider: z.enum(API_PROVIDER_TYPES, {
|
|
66
|
+
provider: z.enum(API_PROVIDER_TYPES, { error: `must be one of ${API_PROVIDER_TYPES.join(", ")}` }),
|
|
66
67
|
model: requiredString().optional(),
|
|
67
68
|
auth_env: requiredString(),
|
|
68
69
|
base_url: z.string().optional(),
|
|
@@ -95,7 +96,7 @@ export function hostToolsImplementProblem(name, agent) {
|
|
|
95
96
|
}
|
|
96
97
|
|
|
97
98
|
export const agentSchema = z.discriminatedUnion("kind", [localAgentSchema, apiAgentSchema, subscriptionAgentSchema], {
|
|
98
|
-
|
|
99
|
+
error: (issue) => (isUnknownDiscriminator(issue) ? `kind must be one of ${AGENT_KINDS.join(", ")}` : undefined),
|
|
99
100
|
});
|
|
100
101
|
|
|
101
102
|
/** An api agent as the pool entry the dispatch path understands. `model` may be absent (see resolveAgentModel). */
|
|
@@ -148,7 +149,7 @@ export function formatAgentIssues(error) {
|
|
|
148
149
|
const lines = error.issues.map((issue) => {
|
|
149
150
|
const where = issue.path.length ? issue.path.join(".") : "config";
|
|
150
151
|
if (issue.code === "unrecognized_keys") return `${where}: unexpected field(s) ${issue.keys.map((k) => `"${k}"`).join(", ")}`;
|
|
151
|
-
if (issue
|
|
152
|
+
if (isMissingField(issue)) return `${where}: is required`;
|
|
152
153
|
return `${where}: ${issue.message}`;
|
|
153
154
|
});
|
|
154
155
|
return [...new Set(lines)];
|
package/lib/army.mjs
CHANGED
|
@@ -29,6 +29,7 @@ import os from "node:os";
|
|
|
29
29
|
import path from "node:path";
|
|
30
30
|
import YAML from "yaml";
|
|
31
31
|
import { z } from "zod";
|
|
32
|
+
import { typeError } from "./zod-issues.mjs";
|
|
32
33
|
import { agentRunsToolsOnHost } from "./dispatch-schema.mjs";
|
|
33
34
|
import { usageStatus } from "./usage-limits.mjs";
|
|
34
35
|
|
|
@@ -116,7 +117,7 @@ const roleSchema = z.object({
|
|
|
116
117
|
disabled: z.boolean().optional(),
|
|
117
118
|
}).strict();
|
|
118
119
|
|
|
119
|
-
const positiveNumber = () => z.number({
|
|
120
|
+
const positiveNumber = () => z.number({ error: typeError("a number") }).positive("must be positive");
|
|
120
121
|
// Limits on one /feature run (lib/runs.mjs). Personal, like the General:
|
|
121
122
|
// global or local only, never a committed project file.
|
|
122
123
|
const runLimitsSchema = z.object({
|
package/lib/connect.mjs
CHANGED
|
@@ -28,6 +28,7 @@ import { execFileSync } from "node:child_process";
|
|
|
28
28
|
import { loadAgents } from "./agents.mjs";
|
|
29
29
|
import { globalConfigDir } from "./army.mjs";
|
|
30
30
|
import { buildNotifierApp } from "./notifier-app.mjs";
|
|
31
|
+
import { recordCopySource } from "./install-freshness.mjs";
|
|
31
32
|
|
|
32
33
|
export function defaultInstallDir() {
|
|
33
34
|
return process.env.NOMARMY_AGENT_INSTALL_DIR || path.join(process.env.HOME ?? process.env.USERPROFILE ?? ".", ".local", "share", "nomarmy-local-worker");
|
|
@@ -142,6 +143,7 @@ export function installMcpCopy({ nomarmyRoot, installDir, run = defaultRun }) {
|
|
|
142
143
|
fs.mkdirSync(path.join(installDir, "mcp"), { recursive: true });
|
|
143
144
|
fs.copyFileSync(path.join(nomarmyRoot, "package.json"), path.join(installDir, "package.json"));
|
|
144
145
|
fs.copyFileSync(path.join(nomarmyRoot, "mcp", "server.mjs"), path.join(installDir, "mcp", "server.mjs"));
|
|
146
|
+
recordCopySource(installDir, nomarmyRoot);
|
|
145
147
|
fs.rmSync(path.join(installDir, "lib"), { recursive: true, force: true });
|
|
146
148
|
fs.cpSync(path.join(nomarmyRoot, "lib"), path.join(installDir, "lib"), { recursive: true });
|
|
147
149
|
// lib/sandbox-images.mjs resolves docker/ relative to its own location
|
package/lib/diff-checks.mjs
CHANGED
|
@@ -1,3 +1,5 @@
|
|
|
1
|
+
import path from "node:path";
|
|
2
|
+
|
|
1
3
|
// ---------------------------------------------------------------------------
|
|
2
4
|
// Test-change classification (plan 16). One tunable constant, on purpose:
|
|
3
5
|
// every heuristic about what counts as a test file lives here and nowhere else.
|
|
@@ -416,3 +418,91 @@ export function mergeUntrackedIntoNameStatus(nameStatus, untrackedFiles) {
|
|
|
416
418
|
return [...(nameStatus ?? []), ...extra];
|
|
417
419
|
}
|
|
418
420
|
|
|
421
|
+
|
|
422
|
+
// ---------------------------------------------------------------------------
|
|
423
|
+
// Verification inputs: .nomarmy.yml is read from the operator's checkout so a
|
|
424
|
+
// worker can't weaken its own checks, but the files its commands run come
|
|
425
|
+
// from the worker's worktree. Found live: a worker whose verification ran
|
|
426
|
+
// `node check.js` rewrote check.js to print PASS and exit 0, and the job
|
|
427
|
+
// committed as done. The revert check can't see it: reverting check.js along
|
|
428
|
+
// with the code makes verification fail, which reads as coverage.
|
|
429
|
+
//
|
|
430
|
+
// Blocking (the diff changes what a command runs): a changed non-test file a
|
|
431
|
+
// command names (`node check.js`, `bash scripts/verify.sh`), a changed
|
|
432
|
+
// Makefile or justfile under `make`/`just`, or a changed package.json script
|
|
433
|
+
// that `npm test`/`pnpm run x`/`yarn x`/`bun run x` calls. Test files named by
|
|
434
|
+
// a command are left to the test-change review, since adding cases is normal.
|
|
435
|
+
// Flagged only: test-runner configuration, which is often a legitimate edit.
|
|
436
|
+
// ---------------------------------------------------------------------------
|
|
437
|
+
|
|
438
|
+
const RUNNER_CONFIG_RE = /(^|\/)(conftest\.py|pytest\.ini|tox\.ini|setup\.cfg|\.coveragerc|(jest|vitest|vite|playwright|karma|cypress)\.config\.[cm]?[jt]s|\.mocharc(\.[a-z]+)?|phpunit\.xml(\.dist)?|\.rspec)$/;
|
|
439
|
+
const TASK_FILES = { make: ["Makefile", "makefile", "GNUmakefile"], just: ["justfile", "Justfile", ".justfile"] };
|
|
440
|
+
|
|
441
|
+
function shellWords(segment) {
|
|
442
|
+
return (segment.match(/"[^"]*"|'[^']*'|\S+/g) ?? []).map((w) => w.replace(/^["']|["']$/g, ""));
|
|
443
|
+
}
|
|
444
|
+
|
|
445
|
+
// Which package.json script a package-manager command runs, if any.
|
|
446
|
+
function packageScript(words) {
|
|
447
|
+
const [tool, first, second] = words;
|
|
448
|
+
if (!["npm", "pnpm", "yarn", "bun"].includes(tool) || !first) return null;
|
|
449
|
+
if (["run", "run-script"].includes(first)) return second && !second.startsWith("-") ? second : null;
|
|
450
|
+
if (tool === "npm") return ["test", "t", "tst"].includes(first) ? "test" : ["start", "stop", "restart"].includes(first) ? first : null;
|
|
451
|
+
if (tool === "bun") return null; // `bun test` is bun's own runner, not a script
|
|
452
|
+
return first.startsWith("-") || ["install", "add", "remove", "exec", "dlx", "x"].includes(first) ? null : first;
|
|
453
|
+
}
|
|
454
|
+
|
|
455
|
+
function scriptsOf(text) {
|
|
456
|
+
try { return JSON.parse(text ?? "")?.scripts ?? {}; } catch { return {}; }
|
|
457
|
+
}
|
|
458
|
+
|
|
459
|
+
function pytestSection(text) {
|
|
460
|
+
const match = /^\[tool\.pytest[^\]]*\]\s*$([\s\S]*?)(?=^\[|(?![\s\S]))/m.exec(String(text ?? ""));
|
|
461
|
+
return match ? match[1].trim() : null;
|
|
462
|
+
}
|
|
463
|
+
|
|
464
|
+
/**
|
|
465
|
+
* @param {{ commands: string[], changedFiles: string[], readBase: (file: string) => string|null|Promise<string|null>,
|
|
466
|
+
* readHead: (file: string) => string|null|Promise<string|null>, isTestPathFn?: (file: string) => boolean }} input
|
|
467
|
+
* @returns {Promise<{ blocked: {file: string, command: string, why: string}[], flagged: {file: string, why: string}[] } | null>}
|
|
468
|
+
*/
|
|
469
|
+
export async function detectVerificationInputChanges({ commands = [], changedFiles = [], readBase = () => null, readHead = () => null, isTestPathFn = isTestPath } = {}) {
|
|
470
|
+
const changed = new Set(changedFiles.map((f) => path.posix.normalize(f)));
|
|
471
|
+
if (!changed.size) return null;
|
|
472
|
+
const blocked = [], flagged = [], seen = new Set();
|
|
473
|
+
const block = (file, command, why) => { if (!seen.has(file)) { seen.add(file); blocked.push({ file, command, why }); } };
|
|
474
|
+
for (const command of commands) {
|
|
475
|
+
let dir = "";
|
|
476
|
+
for (const segment of String(command ?? "").split(/\s*(?:&&|\|\||;|\|)\s*/)) {
|
|
477
|
+
const all = shellWords(segment);
|
|
478
|
+
let lead = 0;
|
|
479
|
+
while (lead < all.length && /^[A-Za-z_][A-Za-z0-9_]*=/.test(all[lead])) lead++; // FOO=1 cmd ...
|
|
480
|
+
const words = all.slice(lead);
|
|
481
|
+
if (!words.length) continue;
|
|
482
|
+
const at = (p) => path.posix.normalize(path.posix.join(dir, p));
|
|
483
|
+
if (words[0] === "cd" && words[1]) { dir = at(words[1]); continue; }
|
|
484
|
+
for (const word of words) {
|
|
485
|
+
if (word.startsWith("-") || word.includes("$")) continue;
|
|
486
|
+
const file = at(word);
|
|
487
|
+
if (changed.has(file) && !isTestPathFn(file)) block(file, command, `run by \`${command}\``);
|
|
488
|
+
}
|
|
489
|
+
for (const name of TASK_FILES[words[0]] ?? []) {
|
|
490
|
+
const file = at(name);
|
|
491
|
+
if (changed.has(file)) block(file, command, `read by \`${command}\``);
|
|
492
|
+
}
|
|
493
|
+
const script = packageScript(words);
|
|
494
|
+
const manifest = at("package.json");
|
|
495
|
+
if (script && changed.has(manifest)) {
|
|
496
|
+
const before = scriptsOf(await readBase(manifest)), after = scriptsOf(await readHead(manifest));
|
|
497
|
+
const touched = [script, `pre${script}`, `post${script}`].filter((s) => before[s] !== after[s]);
|
|
498
|
+
if (touched.length) block(manifest, command, `its script${touched.length > 1 ? "s" : ""} ${touched.map((s) => `"${s}"`).join(", ")} changed, and \`${command}\` runs it`);
|
|
499
|
+
}
|
|
500
|
+
}
|
|
501
|
+
}
|
|
502
|
+
for (const file of changed) {
|
|
503
|
+
if (seen.has(file)) continue;
|
|
504
|
+
if (RUNNER_CONFIG_RE.test(file)) flagged.push({ file, why: "test-runner configuration" });
|
|
505
|
+
else if (/(^|\/)pyproject\.toml$/.test(file) && pytestSection(await readBase(file)) !== pytestSection(await readHead(file))) flagged.push({ file, why: "its [tool.pytest] settings changed" });
|
|
506
|
+
}
|
|
507
|
+
return blocked.length || flagged.length ? { blocked, flagged } : null;
|
|
508
|
+
}
|
package/lib/dispatch-schema.mjs
CHANGED
|
@@ -15,10 +15,11 @@
|
|
|
15
15
|
// never in a file nomArmy writes or reads back.
|
|
16
16
|
|
|
17
17
|
import { z } from "zod";
|
|
18
|
+
import { typeError, isMissingField, isUnknownDiscriminator } from "./zod-issues.mjs";
|
|
18
19
|
|
|
19
20
|
const requiredString = () =>
|
|
20
21
|
z
|
|
21
|
-
.string({
|
|
22
|
+
.string({ error: typeError("a string") })
|
|
22
23
|
.refine((value) => value.trim().length > 0, { message: "must not be empty" });
|
|
23
24
|
|
|
24
25
|
// Exported (not just used internally) so bin/nomarmy.mjs's interactive
|
|
@@ -101,11 +102,11 @@ export function isNativeProviderType(type) {
|
|
|
101
102
|
export const CUSTOM_ENDPOINT_PROVIDER_TYPES = Object.freeze(["bedrock", "azure-openai", "openai-compatible"]);
|
|
102
103
|
|
|
103
104
|
const idSchema = z
|
|
104
|
-
.string({
|
|
105
|
+
.string({ error: typeError("a string") })
|
|
105
106
|
.regex(ID_RE, "must be 1-64 characters of letters, numbers, dot, underscore or hyphen");
|
|
106
107
|
|
|
107
108
|
const weightSchema = z
|
|
108
|
-
.number({
|
|
109
|
+
.number({ error: typeError("a number") })
|
|
109
110
|
.positive("must be a positive number")
|
|
110
111
|
.finite("must be a finite number");
|
|
111
112
|
|
|
@@ -113,13 +114,13 @@ const weightSchema = z
|
|
|
113
114
|
// one entry at once -- a stand-in for real rate-limit-aware admission (see
|
|
114
115
|
// README's dispatch-pool section for why that's out of scope for now).
|
|
115
116
|
const maxConcurrentSchema = z
|
|
116
|
-
.number({
|
|
117
|
+
.number({ error: typeError("a number") })
|
|
117
118
|
.int("must be a whole number")
|
|
118
119
|
.positive("must be a positive number")
|
|
119
120
|
.default(2);
|
|
120
121
|
|
|
121
122
|
const authEnvSchema = z
|
|
122
|
-
.string({
|
|
123
|
+
.string({ error: typeError("a string") })
|
|
123
124
|
.regex(
|
|
124
125
|
AUTH_ENV_NAME_RE,
|
|
125
126
|
"must be an environment variable NAME (uppercase letters, digits, underscores), never the credential itself",
|
|
@@ -127,7 +128,7 @@ const authEnvSchema = z
|
|
|
127
128
|
|
|
128
129
|
const urlSchema = () =>
|
|
129
130
|
z
|
|
130
|
-
.string({
|
|
131
|
+
.string({ error: typeError("a string") })
|
|
131
132
|
.url("must be a valid URL");
|
|
132
133
|
|
|
133
134
|
// Whether/how the entry's model uses a thinking/reasoning mode. `true`
|
|
@@ -152,7 +153,7 @@ const thinkingSchema = z.union([z.boolean(), z.enum(["low", "medium", "high"])])
|
|
|
152
153
|
// knows about yet (see lib/model-catalog.mjs), or an operator who wants to
|
|
153
154
|
// be more conservative than the model's rated maximum.
|
|
154
155
|
const contextWindowSchema = z
|
|
155
|
-
.number({
|
|
156
|
+
.number({ error: typeError("a number") })
|
|
156
157
|
.int("must be a whole number")
|
|
157
158
|
.positive("must be a positive number")
|
|
158
159
|
.optional();
|
|
@@ -189,7 +190,7 @@ const llamaCppEntrySchema = z
|
|
|
189
190
|
// `agents add` installs it before registering the key.
|
|
190
191
|
const genericOpenclawEntrySchema = hostedProviderSchema("openclaw", { requireBaseUrl: false }).extend({
|
|
191
192
|
openclaw_provider: z
|
|
192
|
-
.string({
|
|
193
|
+
.string({ error: typeError("a string") })
|
|
193
194
|
.regex(OPENCLAW_PROVIDER_ID_RE, "must be an OpenClaw provider id (lowercase letters, digits, dot, underscore, hyphen)")
|
|
194
195
|
.refine((id) => !PROVIDER_TYPES.includes(id) || NATIVE_PROVIDER_TYPES.includes(id), "names a nomArmy provider type with its own setup -- use that type instead"),
|
|
195
196
|
plugin: z.string().regex(/^\S+$/, "must be one plugin install spec, e.g. clawhub:@openclaw/deepseek-provider").optional(),
|
|
@@ -205,7 +206,7 @@ export const providerEntrySchema = z.discriminatedUnion("provider", [
|
|
|
205
206
|
export const poolsSchema = z.record(
|
|
206
207
|
z.string().min(1, "pool name must not be empty"),
|
|
207
208
|
z
|
|
208
|
-
.array(providerEntrySchema, {
|
|
209
|
+
.array(providerEntrySchema, { error: typeError("an array of provider entries") })
|
|
209
210
|
.min(1, "must list at least one provider entry"),
|
|
210
211
|
);
|
|
211
212
|
|
|
@@ -250,11 +251,11 @@ export function formatDispatchIssues(error) {
|
|
|
250
251
|
lines.push(`${where}: unexpected field(s) ${keys}`);
|
|
251
252
|
continue;
|
|
252
253
|
}
|
|
253
|
-
if (issue
|
|
254
|
+
if (isUnknownDiscriminator(issue)) {
|
|
254
255
|
lines.push(`${where}: must be one of ${PROVIDER_TYPES.join(", ")}`);
|
|
255
256
|
continue;
|
|
256
257
|
}
|
|
257
|
-
if (issue
|
|
258
|
+
if (isMissingField(issue)) {
|
|
258
259
|
lines.push(`${where}: is required`);
|
|
259
260
|
continue;
|
|
260
261
|
}
|
package/lib/execute.mjs
CHANGED
|
@@ -17,7 +17,7 @@ import { describeRecoveryChanges, reportRecoveryPrompt } from "./worker-prompt.m
|
|
|
17
17
|
import { parseWorkerReport } from "./report.mjs";
|
|
18
18
|
import { OUTCOMES, COORDINATOR_STATUS_BY_OUTCOME } from "./outcomes.mjs";
|
|
19
19
|
import { resolveOutcome, finalText, workerMetadata, applyRefactorContract, applyVerificationPolicy } from "./outcome.mjs";
|
|
20
|
-
import { isTestPath, isDocumentationPath, detectScopedTestSelectionRisk, detectUnwiredNewDefinitions, detectMislabeledTestNames, detectPossibleSecrets } from "./diff-checks.mjs";
|
|
20
|
+
import { isTestPath, isDocumentationPath, detectScopedTestSelectionRisk, detectUnwiredNewDefinitions, detectMislabeledTestNames, detectPossibleSecrets, detectVerificationInputChanges } from "./diff-checks.mjs";
|
|
21
21
|
|
|
22
22
|
// ---------------------------------------------------------------------------
|
|
23
23
|
// Job status for polling. `status.json` is written at every phase transition
|
|
@@ -250,7 +250,7 @@ export function createExecutor(deps) {
|
|
|
250
250
|
// which reads as evidence about work that never happened.
|
|
251
251
|
independentVerification = normalizeVerification({ status: "not_run", basis: "not-applicable", reason: "the worker changed nothing, so there was none of its work to verify" }, verification ?? null);
|
|
252
252
|
} else if (verificationFlow.verificationRunner || !reportValidation.valid) {
|
|
253
|
-
independentVerification = await runIndependentVerification({ profile: verification ?? null, cwd, jobId, baseSha: base.sha, branch, mode, record: preCommit });
|
|
253
|
+
independentVerification = await runIndependentVerification({ profile: verification ?? null, cwd, jobId, baseSha: base.sha, branch, mode, record: preCommit, logFile: path.join(jobDir, "verification.log") });
|
|
254
254
|
}
|
|
255
255
|
|
|
256
256
|
// verify_regression: on by default whenever there's a verification
|
|
@@ -299,12 +299,19 @@ export function createExecutor(deps) {
|
|
|
299
299
|
// Cheap, always-on, additive: never changes commitAllowed/commitBlockedReason
|
|
300
300
|
// on its own (unlike the regression-check override above), only flags for
|
|
301
301
|
// review -- see detectScopedTestSelectionRisk's own doc comment for why.
|
|
302
|
-
let selectionRisk = null;
|
|
302
|
+
let selectionRisk = null, verificationInputs = null;
|
|
303
303
|
if (mode === "implement" && verification) {
|
|
304
304
|
try {
|
|
305
305
|
const loaded = loadConfig(projectDir); // the operator's contract; see registerVerificationRunner's call
|
|
306
306
|
const profileCommands = loaded.found ? (loaded.config?.verification?.[verification]?.commands ?? []) : [];
|
|
307
307
|
selectionRisk = detectScopedTestSelectionRisk({ commands: profileCommands, testChanges: preCommit.testChanges });
|
|
308
|
+
// A diff that changes what those commands run (see
|
|
309
|
+
// detectVerificationInputChanges): applied below, after the others.
|
|
310
|
+
verificationInputs = await detectVerificationInputChanges({
|
|
311
|
+
commands: profileCommands, changedFiles: preCommit.changedFiles ?? [],
|
|
312
|
+
readBase: (file) => gitRaw(["show", `${base.sha}:${file}`], cwd).catch(() => null),
|
|
313
|
+
readHead: (file) => { try { return fs.readFileSync(path.join(cwd, file), "utf8"); } catch { return null; } },
|
|
314
|
+
});
|
|
308
315
|
} catch { /* a config load failure here is the verification runner's own problem to report, not this check's */ }
|
|
309
316
|
}
|
|
310
317
|
const afterSelectionRisk = selectionRisk
|
|
@@ -399,7 +406,21 @@ export function createExecutor(deps) {
|
|
|
399
406
|
commitBlockedReason: `possible secret detected: ${possibleSecrets.reason}`,
|
|
400
407
|
reasons: [...afterHostInstalls.reasons, `POSSIBLE SECRET DETECTED: ${possibleSecrets.reason}`] }
|
|
401
408
|
: afterHostInstalls;
|
|
402
|
-
|
|
409
|
+
// A worker must not be judged by a check it rewrote: a changed script,
|
|
410
|
+
// Makefile or package.json script that a verification command runs
|
|
411
|
+
// blocks the commit; changed test-runner config only asks for review.
|
|
412
|
+
const blockedInputs = verificationInputs?.blocked ?? [], flaggedInputs = verificationInputs?.flagged ?? [];
|
|
413
|
+
const inputLine = blockedInputs.map((b) => `${b.file} (${b.why})`).join("; ");
|
|
414
|
+
const afterInputs = blockedInputs.length
|
|
415
|
+
? { ...afterSecrets, outcome: afterSecrets.commitAllowed || afterSecrets.outcome === OUTCOMES.WORKER_DONE ? OUTCOMES.NEEDS_REVIEW : afterSecrets.outcome,
|
|
416
|
+
reviewRequired: true, commitAllowed: false,
|
|
417
|
+
commitBlockedReason: afterSecrets.commitAllowed ? `the diff changes what verification runs: ${inputLine}` : afterSecrets.commitBlockedReason,
|
|
418
|
+
reasons: [...afterSecrets.reasons, `VERIFICATION INPUT CHANGED: the diff changes what profile '${verification}' runs, so its result can't be trusted: ${inputLine}`] }
|
|
419
|
+
: afterSecrets;
|
|
420
|
+
const afterConfig = flaggedInputs.length
|
|
421
|
+
? { ...afterInputs, reviewRequired: true, reasons: [...afterInputs.reasons, `TEST CONFIG CHANGED: ${flaggedInputs.map((c) => `${c.file} (${c.why})`).join("; ")}`] }
|
|
422
|
+
: afterInputs;
|
|
423
|
+
const finalOutcome = applyRefactorContract(applyVerificationPolicy(afterConfig, independentVerification.status, repoPolicy()),
|
|
403
424
|
{ refactor, verificationStatus: independentVerification.status, testChanges: preCommit.testChanges });
|
|
404
425
|
|
|
405
426
|
progress("commit");
|
|
@@ -411,6 +432,7 @@ export function createExecutor(deps) {
|
|
|
411
432
|
let coordinatorStatus = COORDINATOR_STATUS_BY_OUTCOME[finalOutcome.outcome] ?? "incomplete";
|
|
412
433
|
const issues = [...finalOutcome.reasons, ...(preCommit.issues ?? [])];
|
|
413
434
|
if (workerError) issues.push(`worker error: ${String(workerError).split("\n")[0]}`);
|
|
435
|
+
if ((result ?? attempted)?.salvaged) issues.push(`runner cleanup failed after the run (${(result ?? attempted).salvagedFrom}); the worker's report was recovered from the run's transcript`);
|
|
414
436
|
if (workerFailed || workerTimedOut) { const restarted = vmRestartIssue(vmStartedBefore, deps.podmanVmStartedAt?.() ?? null); if (restarted) issues.unshift(restarted); }
|
|
415
437
|
if (repositoryChanged && !commit.created) {
|
|
416
438
|
if (coordinatorStatus === "complete") coordinatorStatus = "incomplete";
|
|
@@ -575,6 +597,7 @@ export function createExecutor(deps) {
|
|
|
575
597
|
const worker = workerMetadata(result ?? attempted);
|
|
576
598
|
const issues = [...outcome.reasons];
|
|
577
599
|
if (workerError) issues.push(`scout error: ${String(workerError).split("\n")[0]}`);
|
|
600
|
+
if ((result ?? attempted)?.salvaged) issues.push(`runner cleanup failed after the run (${(result ?? attempted).salvagedFrom}); the scout's report was recovered from the run's transcript`);
|
|
578
601
|
const failures = worker.toolSummary?.failures ?? 0; if (failures > 0) issues.push(`scout recorded ${failures} tool failure(s)`);
|
|
579
602
|
if (dirty) issues.push(`snapshot changed: ${record.repoStatusFiles.join(", ")}`);
|
|
580
603
|
if (reportRecoveryAttempted) {
|
|
@@ -706,6 +729,7 @@ export function createExecutor(deps) {
|
|
|
706
729
|
const worker = workerMetadata(result ?? attempted);
|
|
707
730
|
const issues = [...outcome.reasons];
|
|
708
731
|
if (workerError) issues.push(`decompose error: ${String(workerError).split("\n")[0]}`);
|
|
732
|
+
if ((result ?? attempted)?.salvaged) issues.push(`runner cleanup failed after the run (${(result ?? attempted).salvagedFrom}); the decomposer's report was recovered from the run's transcript`);
|
|
709
733
|
const failures = worker.toolSummary?.failures ?? 0; if (failures > 0) issues.push(`decomposer recorded ${failures} tool failure(s)`);
|
|
710
734
|
if (dirty) issues.push(`snapshot changed: ${record.repoStatusFiles.join(", ")}`);
|
|
711
735
|
if (overlaps.length) issues.push(`${overlaps.length} subtask pair(s) claim overlapping files; not safe to dispatch as independent jobs as proposed`);
|
package/lib/health.mjs
CHANGED
|
@@ -11,6 +11,8 @@
|
|
|
11
11
|
// army a role or the General pointing at something unusable
|
|
12
12
|
// config agents.yml unloadable or unsafe
|
|
13
13
|
// leftovers retained worktrees, job storage, stale "running" jobs
|
|
14
|
+
// nomarmy-update a newer nomArmy on npm; nomarmy-copy: coordinators run
|
|
15
|
+
// an older copy than the installed CLI (install-freshness.mjs)
|
|
14
16
|
//
|
|
15
17
|
// Each issue: { id, severity: "error"|"warn"|"info", title, detail, fix,
|
|
16
18
|
// short }. `id` is stable across runs, so a notification goes out once per
|
|
@@ -23,6 +25,7 @@ import { executionMode } from "./execution.mjs";
|
|
|
23
25
|
import { modelRejection } from "./openclaw-errors.mjs";
|
|
24
26
|
import { providerConfigured, readOpenclawConfig } from "./openclaw-config.mjs";
|
|
25
27
|
import { readUsageSnapshots, usageStatus } from "./usage-limits.mjs";
|
|
28
|
+
import { freshnessIssues, readInstallVersions } from "./install-freshness.mjs";
|
|
26
29
|
|
|
27
30
|
const DAY = 86400000;
|
|
28
31
|
|
|
@@ -213,7 +216,7 @@ export function leftoverIssues({ retainedWorktrees = 0, jobsBytes = 0, staleRunn
|
|
|
213
216
|
* Run every check. `env` supplies what each needs, with real defaults;
|
|
214
217
|
* tests pass their own.
|
|
215
218
|
*/
|
|
216
|
-
export async function runHealthChecks({ now = Date.now(), mode = "local", openclawCmd = process.env.NOMARMY_OPENCLAW_CMD || "openclaw", run = runBounded, armySummary = null, agentsError = null, jobsRoot = null, pidAlive = () => true, agents = null, openclawConfig = null, vendors = {}, modelsInUse = null, autoPruned = null, usageSnapshots = null } = {}) {
|
|
219
|
+
export async function runHealthChecks({ now = Date.now(), mode = "local", openclawCmd = process.env.NOMARMY_OPENCLAW_CMD || "openclaw", run = runBounded, armySummary = null, agentsError = null, jobsRoot = null, pidAlive = () => true, agents = null, openclawConfig = null, vendors = {}, modelsInUse = null, autoPruned = null, usageSnapshots = null, install = null } = {}) {
|
|
217
220
|
const issues = [];
|
|
218
221
|
if (autoPruned?.freedBytes) issues.push({ id: `auto-prune:${new Date(now).toISOString()}`, severity: "info",
|
|
219
222
|
title: `Freed ${(autoPruned.freedBytes / 1024 ** 3).toFixed(2)} GB: ${[autoPruned.pruned ? `runtime data of ${autoPruned.pruned} finished job${autoPruned.pruned === 1 ? "" : "s"} older than ${autoPruned.olderThanHours}h` : null, autoPruned.scratchCleared ? `OpenClaw scratch files of ${autoPruned.scratchCleared} more` : null].filter(Boolean).join(", ")}`,
|
|
@@ -230,12 +233,14 @@ export async function runHealthChecks({ now = Date.now(), mode = "local", opencl
|
|
|
230
233
|
short });
|
|
231
234
|
}
|
|
232
235
|
}
|
|
233
|
-
const [auth, version, latest, plugins] = await Promise.all([
|
|
236
|
+
const [auth, version, latest, plugins, nomarmyLatest] = await Promise.all([
|
|
234
237
|
run(openclawCmd, ["models", "auth", "list", "--json"]),
|
|
235
238
|
run(openclawCmd, ["--version"]),
|
|
236
239
|
run("npm", ["view", "openclaw", "version"], { timeoutMs: 15000 }),
|
|
237
240
|
run(openclawCmd, ["plugins", "inspect", "codex"]),
|
|
241
|
+
install ? run("npm", ["view", "nomarmy", "dist-tags.alpha"], { timeoutMs: 15000 }) : null,
|
|
238
242
|
]);
|
|
243
|
+
if (install) issues.push(...freshnessIssues({ ...install, latestVersion: nomarmyLatest?.ok ? nomarmyLatest.stdout : null }));
|
|
239
244
|
if (auth.ok) { try { issues.push(...loginExpiryIssues(JSON.parse(auth.stdout.slice(auth.stdout.indexOf("{"))), { now })); } catch { /* unparseable: skip */ } }
|
|
240
245
|
const pluginVersion = /Version:\s*(\S+)/.exec(plugins.stdout ?? "")?.[1];
|
|
241
246
|
if (version.ok) issues.push(...versionIssues({ installed: version.stdout, latest: latest.ok ? latest.stdout : null, plugins: pluginVersion ? [{ id: "codex", version: pluginVersion }] : [] }));
|
|
@@ -318,8 +323,10 @@ export async function checkAndRecordHealth({ projectDir, stateRoot, configDir, n
|
|
|
318
323
|
let autoPruned = null;
|
|
319
324
|
if (ageMs !== null) { try { autoPruned = { ...pruneJobRuntime({ stateRoot, olderThanMs: ageMs, now }), olderThanHours: ageMs / 3600000 }; } catch { /* best-effort */ } }
|
|
320
325
|
const mode = executionMode(env).mode;
|
|
326
|
+
const { defaultInstallDir } = await import("./connect.mjs");
|
|
327
|
+
const install = readInstallVersions(defaultInstallDir());
|
|
321
328
|
const result = await runHealthChecks({ now, mode, armySummary, agentsError, jobsRoot: path.join(stateRoot, "jobs"), pidAlive,
|
|
322
|
-
agents, openclawConfig: readOpenclawConfig(), vendors: SUBSCRIPTION_VENDORS, modelsInUse, autoPruned, usageSnapshots: readUsageSnapshots(stateRoot) });
|
|
329
|
+
agents, openclawConfig: readOpenclawConfig(), vendors: SUBSCRIPTION_VENDORS, modelsInUse, autoPruned, usageSnapshots: readUsageSnapshots(stateRoot), install });
|
|
323
330
|
const toNotify = recordHealth(path.join(stateRoot, "health.json"), result, { now });
|
|
324
331
|
return { result, toNotify };
|
|
325
332
|
}
|
|
@@ -0,0 +1,83 @@
|
|
|
1
|
+
// Whether the nomArmy a coordinator runs is current. Claude Code, Codex and
|
|
2
|
+
// Cursor run a copy of the server that `nomarmy connect` puts in the install
|
|
3
|
+
// dir, and each session keeps the code it started with. So an upgrade can
|
|
4
|
+
// stall at three points, each checked here:
|
|
5
|
+
//
|
|
6
|
+
// nomarmy-update npm's alpha release is newer than the installed CLI
|
|
7
|
+
// nomarmy-copy the CLI is newer than the copy coordinators run
|
|
8
|
+
// (`nomarmy connect` wasn't re-run)
|
|
9
|
+
// restart the copy on disk changed after this server started
|
|
10
|
+
// (the session wasn't restarted); per session, so it is
|
|
11
|
+
// reported by that session's server, not in health.json
|
|
12
|
+
|
|
13
|
+
import fs from "node:fs";
|
|
14
|
+
import path from "node:path";
|
|
15
|
+
|
|
16
|
+
/** Written into the install dir by `nomarmy connect`: where the copy came from. */
|
|
17
|
+
export const SOURCE_FILE = "source.json";
|
|
18
|
+
|
|
19
|
+
export function readPackageVersion(root) {
|
|
20
|
+
try { return JSON.parse(fs.readFileSync(path.join(root, "package.json"), "utf8")).version ?? null; }
|
|
21
|
+
catch { return null; }
|
|
22
|
+
}
|
|
23
|
+
|
|
24
|
+
export function recordCopySource(installDir, nomarmyRoot) {
|
|
25
|
+
fs.writeFileSync(path.join(installDir, SOURCE_FILE), JSON.stringify({ root: nomarmyRoot, version: readPackageVersion(nomarmyRoot) }, null, 2) + "\n");
|
|
26
|
+
}
|
|
27
|
+
|
|
28
|
+
/** The copy's version and, when connect recorded it, the version now at its source. */
|
|
29
|
+
export function readInstallVersions(installDir) {
|
|
30
|
+
let source = null;
|
|
31
|
+
try { source = JSON.parse(fs.readFileSync(path.join(installDir, SOURCE_FILE), "utf8")); } catch { /* connected before source.json existed */ }
|
|
32
|
+
return { copyVersion: readPackageVersion(installDir), sourceVersion: source?.root ? readPackageVersion(source.root) : null };
|
|
33
|
+
}
|
|
34
|
+
|
|
35
|
+
/** Semver order, prerelease included (0.1.0-alpha.7 < 0.1.0-alpha.10 < 0.1.0). */
|
|
36
|
+
export function compareVersions(a, b) {
|
|
37
|
+
const split = (v) => { const [main, pre] = String(v).trim().replace(/^v/, "").split("-", 2); return { main: main.split(".").map(Number), pre: pre ? pre.split(".") : null }; };
|
|
38
|
+
const x = split(a), y = split(b);
|
|
39
|
+
for (let i = 0; i < 3; i++) if ((x.main[i] || 0) !== (y.main[i] || 0)) return (x.main[i] || 0) < (y.main[i] || 0) ? -1 : 1;
|
|
40
|
+
if (!x.pre || !y.pre) return x.pre === y.pre ? 0 : x.pre ? -1 : 1;
|
|
41
|
+
for (let i = 0; i < Math.max(x.pre.length, y.pre.length); i++) {
|
|
42
|
+
const p = x.pre[i], q = y.pre[i];
|
|
43
|
+
if (p === undefined || q === undefined) return p === undefined ? -1 : 1;
|
|
44
|
+
if (p === q) continue;
|
|
45
|
+
const pn = /^\d+$/.test(p), qn = /^\d+$/.test(q);
|
|
46
|
+
if (pn && qn) return Number(p) < Number(q) ? -1 : 1;
|
|
47
|
+
if (pn !== qn) return pn ? -1 : 1;
|
|
48
|
+
return p < q ? -1 : 1;
|
|
49
|
+
}
|
|
50
|
+
return 0;
|
|
51
|
+
}
|
|
52
|
+
|
|
53
|
+
const valid = (v) => typeof v === "string" && /^v?\d+\.\d+\.\d+/.test(v.trim());
|
|
54
|
+
|
|
55
|
+
/** Health issues for a stale CLI or a stale installed copy. */
|
|
56
|
+
export function freshnessIssues({ copyVersion = null, sourceVersion = null, latestVersion = null }) {
|
|
57
|
+
const issues = [];
|
|
58
|
+
if (valid(latestVersion) && valid(sourceVersion) && compareVersions(sourceVersion, latestVersion) < 0) {
|
|
59
|
+
const latest = latestVersion.trim();
|
|
60
|
+
issues.push({ id: `nomarmy-update:${latest}`, severity: "info", title: `nomArmy ${latest} is out (you have ${sourceVersion})`,
|
|
61
|
+
detail: "Updating installs it, reconnects your coordinators and tells you which sessions to restart.",
|
|
62
|
+
fix: "nomarmy update", short: "nomarmy update" });
|
|
63
|
+
}
|
|
64
|
+
if (valid(sourceVersion) && valid(copyVersion) && compareVersions(copyVersion, sourceVersion) < 0) {
|
|
65
|
+
issues.push({ id: `nomarmy-copy:${sourceVersion}`, severity: "warn", title: `Your coordinators run nomArmy ${copyVersion}, but ${sourceVersion} is installed`,
|
|
66
|
+
detail: "Claude Code, Codex and Cursor run a copy of nomArmy that only `nomarmy connect` refreshes.",
|
|
67
|
+
fix: "nomarmy connect claude (and codex, cursor), then restart those sessions", short: "nomarmy reconnect" });
|
|
68
|
+
}
|
|
69
|
+
return issues;
|
|
70
|
+
}
|
|
71
|
+
|
|
72
|
+
/**
|
|
73
|
+
* For the running server: a notice when its copy on disk changed after it
|
|
74
|
+
* started, so this session still runs the old code. Null when current.
|
|
75
|
+
*/
|
|
76
|
+
export function restartNotice({ serverFile, startedAtMs, runningVersion, stat = fs.statSync, readVersion = readPackageVersion }) {
|
|
77
|
+
let changedMs;
|
|
78
|
+
try { changedMs = stat(serverFile).mtimeMs; } catch { return null; }
|
|
79
|
+
if (!(changedMs > startedAtMs)) return null;
|
|
80
|
+
const onDisk = readVersion(path.join(path.dirname(serverFile), ".."));
|
|
81
|
+
const versions = onDisk && onDisk !== runningVersion ? ` (this session runs ${runningVersion}; ${onDisk} is installed)` : "";
|
|
82
|
+
return `nomArmy was updated after this session started${versions}. Restart this session to use the new version; until then it runs the old code.`;
|
|
83
|
+
}
|
package/lib/job-format.mjs
CHANGED
|
@@ -124,3 +124,27 @@ export function formatUnion(union) {
|
|
|
124
124
|
const artifacts = union.worktree ? `\n\nUnion artifacts: ${path.dirname(union.worktree)}\nWorktree retained for review: ${union.worktree}\nBranch retained for review: ${union.branch}` : "";
|
|
125
125
|
return `${banner}--- UNION RECORD ---\n${JSON.stringify(union, null, 2)}${artifacts}`;
|
|
126
126
|
}
|
|
127
|
+
|
|
128
|
+
/**
|
|
129
|
+
* What a person calls a job: its commit subject, else the task's first
|
|
130
|
+
* sentence, capped. Two jobs on one agent and model differ here even with no
|
|
131
|
+
* role, so a notification can tell them apart.
|
|
132
|
+
*/
|
|
133
|
+
export function jobLabel(args) {
|
|
134
|
+
const text = String(args?.commit_subject || args?.task || "").split("\n")[0].split(/(?<=\.)\s/)[0].trim();
|
|
135
|
+
return text.length > 60 ? `${text.slice(0, 57)}...` : text || null;
|
|
136
|
+
}
|
|
137
|
+
|
|
138
|
+
/**
|
|
139
|
+
* Seconds a job has run: to its finish once it has one, not to whenever it's
|
|
140
|
+
* asked about. The record's total_elapsed covers the whole job; an implement
|
|
141
|
+
* job's finishedAt marks only the worker's end, before verification and commit.
|
|
142
|
+
*/
|
|
143
|
+
export function jobElapsedSeconds({ status = null, meta = null, entry = null, now = Date.now() } = {}) {
|
|
144
|
+
const total = meta?.metrics?.total_elapsed;
|
|
145
|
+
if (Number.isFinite(total)) return Math.round(total / 1000);
|
|
146
|
+
const startedMs = Date.parse(status?.startedAt ?? entry?.startedAt ?? meta?.startedAt ?? "");
|
|
147
|
+
if (!Number.isFinite(startedMs)) return null;
|
|
148
|
+
const finishedMs = Date.parse((status?.state === "finished" ? status.updatedAt : null) ?? meta?.finishedAt ?? "");
|
|
149
|
+
return Math.round(((Number.isFinite(finishedMs) ? finishedMs : now) - startedMs) / 1000);
|
|
150
|
+
}
|
package/lib/outcome.mjs
CHANGED
|
@@ -7,6 +7,16 @@ import { parseWorkerReport } from "./report.mjs";
|
|
|
7
7
|
// failed. Recovery exists so that a mangled REPORT cannot destroy correct WORK.
|
|
8
8
|
// It does not exist to launder a failure into a success.
|
|
9
9
|
// ---------------------------------------------------------------------------
|
|
10
|
+
// A failed verification's own detail (the command, its exit code, the end of
|
|
11
|
+
// its output), capped so an issue line stays readable; the full output is in
|
|
12
|
+
// the job's verification.log.
|
|
13
|
+
const FAILURE_DETAIL_CHARS = 600;
|
|
14
|
+
function failureDetail(independentVerification) {
|
|
15
|
+
const detail = String(independentVerification?.detail ?? independentVerification?.reason ?? "").trim();
|
|
16
|
+
if (!detail) return "";
|
|
17
|
+
return `: ${detail.length > FAILURE_DETAIL_CHARS ? `${detail.slice(0, FAILURE_DETAIL_CHARS)}...` : detail}`;
|
|
18
|
+
}
|
|
19
|
+
|
|
10
20
|
export function resolveOutcome({ report, repositoryChanged = false, independentVerification = null, regressionCheck = null, workerFailed = false, workerTimedOut = false, mode = "implement" }) {
|
|
11
21
|
const verification = independentVerification?.status ?? "not_run";
|
|
12
22
|
const parsed = report ?? parseWorkerReport("");
|
|
@@ -34,7 +44,7 @@ export function resolveOutcome({ report, repositoryChanged = false, independentV
|
|
|
34
44
|
if (verification === "fail") {
|
|
35
45
|
return { ...base, outcome: OUTCOMES.NEEDS_REVIEW, reviewRequired: true,
|
|
36
46
|
commitBlockedReason: "independent verification failed despite a clean done/pass report",
|
|
37
|
-
reasons: [
|
|
47
|
+
reasons: [`worker claimed done/pass but independent verification failed${failureDetail(independentVerification)}`] };
|
|
38
48
|
}
|
|
39
49
|
// verify_regression: reverting just the production files and re-running
|
|
40
50
|
// the SAME verification profile still passed (or came back genuinely
|
|
@@ -92,7 +102,7 @@ export function resolveOutcome({ report, repositoryChanged = false, independentV
|
|
|
92
102
|
if (verification === "fail") {
|
|
93
103
|
return { ...recovery, outcome: OUTCOMES.WORKER_REPORT_INVALID,
|
|
94
104
|
commitBlockedReason: "independent verification failed; recovery cannot promote a failure",
|
|
95
|
-
reasons: [...recovery.reasons,
|
|
105
|
+
reasons: [...recovery.reasons, `independent verification FAILED${failureDetail(independentVerification)}`] };
|
|
96
106
|
}
|
|
97
107
|
if (verification === "pass") {
|
|
98
108
|
// A leniently recovered `done` plus a passing independent check is the
|
package/lib/schema.mjs
CHANGED
|
@@ -12,6 +12,7 @@
|
|
|
12
12
|
// approval before the job runs. See `collectElevated`.
|
|
13
13
|
|
|
14
14
|
import { z } from "zod";
|
|
15
|
+
import { typeError, isMissingField, isUnknownDiscriminator } from "./zod-issues.mjs";
|
|
15
16
|
import { armySchema } from "./army.mjs";
|
|
16
17
|
|
|
17
18
|
export const SERVICE_SOURCES = Object.freeze([
|
|
@@ -57,12 +58,12 @@ export const SERVICE_FIELDS = Object.freeze({
|
|
|
57
58
|
// the dotted path, so repeating it reads as stutter.
|
|
58
59
|
const requiredString = () =>
|
|
59
60
|
z
|
|
60
|
-
.string({
|
|
61
|
+
.string({ error: typeError("a string") })
|
|
61
62
|
.refine((value) => value.trim().length > 0, { message: "must not be empty" });
|
|
62
63
|
|
|
63
64
|
const enumOf = (values) =>
|
|
64
65
|
z.enum(values, {
|
|
65
|
-
|
|
66
|
+
error: `must be one of ${values.join(", ")}`,
|
|
66
67
|
});
|
|
67
68
|
|
|
68
69
|
// A plain hostname: no scheme, no path, no port, no whitespace.
|
|
@@ -90,7 +91,7 @@ export function hostnameProblem(value) {
|
|
|
90
91
|
}
|
|
91
92
|
|
|
92
93
|
const hostnameSchema = z
|
|
93
|
-
.string({
|
|
94
|
+
.string({ error: typeError("a string hostname") })
|
|
94
95
|
.superRefine((value, ctx) => {
|
|
95
96
|
const problem = hostnameProblem(value);
|
|
96
97
|
if (problem) ctx.addIssue({ code: z.ZodIssueCode.custom, message: problem });
|
|
@@ -162,7 +163,7 @@ export const serviceSchema = z.discriminatedUnion("source", [
|
|
|
162
163
|
export const pythonEnvironmentSchema = z
|
|
163
164
|
.object({
|
|
164
165
|
requirements: z
|
|
165
|
-
.array(requiredString(), {
|
|
166
|
+
.array(requiredString(), { error: typeError("an array of strings") })
|
|
166
167
|
.min(1, "must list at least one requirements file"),
|
|
167
168
|
})
|
|
168
169
|
.strict();
|
|
@@ -211,10 +212,7 @@ export const verificationProfileSchema = z
|
|
|
211
212
|
.object({
|
|
212
213
|
environment: enumOf(ENVIRONMENT_LEVELS).default("none"),
|
|
213
214
|
commands: z
|
|
214
|
-
.array(requiredString(), {
|
|
215
|
-
required_error: "is required",
|
|
216
|
-
invalid_type_error: "must be an array of strings",
|
|
217
|
-
})
|
|
215
|
+
.array(requiredString(), { error: typeError("an array of strings") })
|
|
218
216
|
.min(1, "must list at least one command"),
|
|
219
217
|
})
|
|
220
218
|
.strict();
|
|
@@ -235,7 +233,7 @@ export const environmentRetentionSchema = z
|
|
|
235
233
|
debug: enumOf(RETENTION_ACTIONS).default(DEFAULT_RETENTION.debug),
|
|
236
234
|
})
|
|
237
235
|
.strict()
|
|
238
|
-
.
|
|
236
|
+
.prefault({});
|
|
239
237
|
|
|
240
238
|
// ---------------------------------------------------------------------------
|
|
241
239
|
// root
|
|
@@ -295,11 +293,11 @@ export function formatIssues(error) {
|
|
|
295
293
|
lines.push(`${where}: unexpected field(s) ${keys}`);
|
|
296
294
|
continue;
|
|
297
295
|
}
|
|
298
|
-
if (issue
|
|
296
|
+
if (isUnknownDiscriminator(issue)) {
|
|
299
297
|
lines.push(`${where}: must be one of ${SERVICE_SOURCES.join(", ")}`);
|
|
300
298
|
continue;
|
|
301
299
|
}
|
|
302
|
-
if (issue
|
|
300
|
+
if (isMissingField(issue)) {
|
|
303
301
|
lines.push(`${where}: is required`);
|
|
304
302
|
continue;
|
|
305
303
|
}
|
|
@@ -0,0 +1,15 @@
|
|
|
1
|
+
// Zod 4 helpers shared by nomArmy's config schemas and their issue formatters.
|
|
2
|
+
|
|
3
|
+
/**
|
|
4
|
+
* The `error` option for a primitive schema: "is required" when the field is
|
|
5
|
+
* absent, "must be <expected>" when it has the wrong type. Zod 4 replaced
|
|
6
|
+
* zod 3's required_error and invalid_type_error with this single function.
|
|
7
|
+
* @param {string} expected e.g. "a string"
|
|
8
|
+
*/
|
|
9
|
+
export const typeError = (expected) => (issue) => (issue.input === undefined ? "is required" : `must be ${expected}`);
|
|
10
|
+
|
|
11
|
+
/** A field that is absent entirely, with no custom message of its own. */
|
|
12
|
+
export const isMissingField = (issue) => issue.code === "invalid_type" && / received undefined$/.test(issue.message);
|
|
13
|
+
|
|
14
|
+
/** A discriminated union whose discriminator matched none of its options. */
|
|
15
|
+
export const isUnknownDiscriminator = (issue) => issue.code === "invalid_union" && issue.note === "No matching discriminator";
|
package/mcp/server.mjs
CHANGED
|
@@ -35,8 +35,9 @@ import { OUTCOMES, COORDINATOR_STATUS_BY_OUTCOME } from "../lib/outcomes.mjs";
|
|
|
35
35
|
import { readUsageSnapshots, usageStatus } from "../lib/usage-limits.mjs";
|
|
36
36
|
import { modelRefusals } from "../lib/health.mjs";
|
|
37
37
|
import { podmanProblem, podmanVmStartedAt } from "../lib/podman-health.mjs";
|
|
38
|
+
import { restartNotice } from "../lib/install-freshness.mjs";
|
|
38
39
|
import { createBuildMetrics, resolveOutcome, finalText, workerMetadata, usageMetrics, policyAdmissionProblems, applyRefactorContract, applyVerificationPolicy, resolveVerifyRegression } from "../lib/outcome.mjs";
|
|
39
|
-
import { compactJobRecord, formatResult, formatUnion, testChangeBanner, regressionCheckBanner, decomposeOverlapBanner } from "../lib/job-format.mjs";
|
|
40
|
+
import { jobLabel, compactJobRecord, formatResult, formatUnion, testChangeBanner, regressionCheckBanner, decomposeOverlapBanner } from "../lib/job-format.mjs";
|
|
40
41
|
|
|
41
42
|
export { run, mapLimit };
|
|
42
43
|
export { readsMeasurable, measureReads };
|
|
@@ -57,6 +58,7 @@ export { TEST_PATH_PATTERNS, isTestPath, testPatternFor, classifyTestChanges, de
|
|
|
57
58
|
// package.json to the same relative location next to the installed
|
|
58
59
|
// mcp/server.mjs, so this resolves identically in a dev checkout or an
|
|
59
60
|
// installed copy.
|
|
61
|
+
const SERVER_STARTED_MS = Date.now();
|
|
60
62
|
const VERSION = JSON.parse(fs.readFileSync(path.join(path.dirname(fileURLToPath(import.meta.url)), "..", "package.json"), "utf8")).version;
|
|
61
63
|
// Sent to every coordinator on connect, so no project needs a copied CLAUDE.md.
|
|
62
64
|
const server = new McpServer({ name: "nomarmy-local-worker", version: VERSION }, { instructions: COORDINATOR_INSTRUCTIONS });
|
|
@@ -406,9 +408,20 @@ server.tool("local_worker_status", `Status of one job started by this server: ph
|
|
|
406
408
|
if (full && files.meta) return toolText(JSON.stringify(files.meta, null, 2), summary.coordinatorStatus !== "complete");
|
|
407
409
|
return toolText(JSON.stringify({ ...summary, jobDir, hint: entry?.result || files.meta ? "call again with full=true for the complete report" : null }, null, 2), summary.state === "orphaned" || summary.state === "failed");
|
|
408
410
|
});
|
|
411
|
+
// Set when this session's copy of nomArmy changed on disk after it started
|
|
412
|
+
// (nomarmy connect or update ran): shown first in army and capacity, and
|
|
413
|
+
// notified once, since only a restart of this session picks it up.
|
|
414
|
+
let restartNotified = false;
|
|
415
|
+
function currentRestartNotice() {
|
|
416
|
+
const notice = restartNotice({ serverFile: fileURLToPath(import.meta.url), startedAtMs: SERVER_STARTED_MS, runningVersion: VERSION });
|
|
417
|
+
if (notice && !restartNotified) { restartNotified = true; try { notify("nomArmy: restart this session", notice); } catch { /* best-effort */ } }
|
|
418
|
+
return notice;
|
|
419
|
+
}
|
|
420
|
+
const withRestartNotice = (value) => { const notice = currentRestartNotice(); return notice ? { restartNeeded: notice, ...value } : value; };
|
|
421
|
+
|
|
409
422
|
server.tool("local_worker_capacity", "What this host can take right now: context per nom and the brief/report budgets derived from it, memory pressure and whether another job would be admitted, and the jobs currently running. Read-only.", {}, async () => {
|
|
410
423
|
await budgetState.refresh();
|
|
411
|
-
return toolText(JSON.stringify(capacitySnapshot(), null, 2));
|
|
424
|
+
return toolText(JSON.stringify(withRestartNotice(capacitySnapshot()), null, 2));
|
|
412
425
|
});
|
|
413
426
|
// The only way to know what `verification`/`union_verification`/
|
|
414
427
|
// `verify_regression` profile names are actually valid for this repo used to
|
|
@@ -515,7 +528,7 @@ server.tool("army", "Who you, the General, are and who you call for what in this
|
|
|
515
528
|
role.modelNote = `${role.model} isn't in OpenClaw's catalog for ${role.agent}; \`army assign\` checked it with a real test call when it was set, and the catalog can lag new models. Use it as assigned; if a job reports "Unknown model", reassign.`;
|
|
516
529
|
}
|
|
517
530
|
}
|
|
518
|
-
return toolText(JSON.stringify(summary, null, 2));
|
|
531
|
+
return toolText(JSON.stringify(withRestartNotice(summary), null, 2));
|
|
519
532
|
} catch (error) {
|
|
520
533
|
return toolText(error.message, true);
|
|
521
534
|
}
|
|
@@ -569,7 +582,7 @@ server.tool("local_workers", "Run independent jobs (implement or scout) with bou
|
|
|
569
582
|
// so it was invisible to both ceilings while it ran.
|
|
570
583
|
// A batch job waits for its agent's slot (up to its own timeout) rather
|
|
571
584
|
// than failing because an earlier job in the same batch holds it.
|
|
572
|
-
return trackInRun(j, track(jobId, { mode: j.mode, workerId, lane: jobLane(j), agent: j.agentName ?? null, runId: j.run_id ?? null, role: j.armyRole ?? null, model: j.model ?? null },
|
|
585
|
+
return trackInRun(j, track(jobId, { mode: j.mode, workerId, lane: jobLane(j), agent: j.agentName ?? null, runId: j.run_id ?? null, role: j.armyRole ?? null, model: j.model ?? null, label: jobLabel(j) },
|
|
573
586
|
withAgentSlot(j, jobId, () => executeJob({ ...jobArgs(effectiveJob, workerId), jobId }), { waitMs: (j.timeout_seconds ?? 600) * 1000 }))).promise;
|
|
574
587
|
}, { staggerMs: WORKER_START_STAGGER_MS });
|
|
575
588
|
indices.forEach((i, laneI) => { results[i] = laneResults[laneI]; });
|
package/package.json
CHANGED
|
@@ -1,9 +1,9 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "nomarmy",
|
|
3
|
-
"description": "
|
|
3
|
+
"description": "Every byte verified: a harness for AI coding workers whose claims are never trusted. Your coding assistant stays in charge while workers implement and test in sandboxes, and nomArmy checks every change before it is committed.",
|
|
4
4
|
"author": "Rayson Technologies",
|
|
5
5
|
"license": "Apache-2.0",
|
|
6
|
-
"version": "0.1.0-alpha.
|
|
6
|
+
"version": "0.1.0-alpha.8",
|
|
7
7
|
"private": false,
|
|
8
8
|
"type": "module",
|
|
9
9
|
"engines": {
|
|
@@ -21,11 +21,11 @@
|
|
|
21
21
|
"tag": "alpha"
|
|
22
22
|
},
|
|
23
23
|
"dependencies": {
|
|
24
|
-
"@modelcontextprotocol/sdk": "^1.
|
|
24
|
+
"@modelcontextprotocol/sdk": "^1.30.1",
|
|
25
25
|
"@secretlint/node": "^13.0.5",
|
|
26
26
|
"@secretlint/secretlint-rule-preset-recommend": "^13.0.5",
|
|
27
27
|
"yaml": "^2.5.0",
|
|
28
|
-
"zod": "^
|
|
28
|
+
"zod": "^4.6.5"
|
|
29
29
|
},
|
|
30
30
|
"scripts": {
|
|
31
31
|
"test": "node --test tests/*.test.mjs",
|