nomarmy 0.1.0-alpha.3 → 0.1.0-alpha.4

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -3,7 +3,7 @@
3
3
  <p align="center">
4
4
  <a href="https://github.com/rayson-tech/nomarmy/actions/workflows/ci.yml"><img src="https://github.com/rayson-tech/nomarmy/actions/workflows/ci.yml/badge.svg" alt="CI"></a>
5
5
  <a href="https://www.npmjs.com/package/nomarmy"><img src="https://img.shields.io/npm/v/nomarmy/alpha?label=npm%40alpha" alt="npm"></a>
6
- <a href="LICENSE"><img src="https://img.shields.io/badge/License-Apache%202.0-blue.svg" alt="License: Apache 2.0"></a>
6
+ <a href="https://github.com/rayson-tech/nomarmy/blob/main/LICENSE"><img src="https://img.shields.io/badge/License-Apache%202.0-blue.svg" alt="License: Apache 2.0"></a>
7
7
  </p>
8
8
 
9
9
  <p align="center"><em>Tiny coders, big appetites for bounded tickets.</em> 🍪</p>
@@ -13,18 +13,28 @@
13
13
  ## TL;DR
14
14
 
15
15
  1. **Have** Git, Node 20+ and [Podman](https://podman.io) (on macOS: `brew install podman && podman machine init && podman machine start`).
16
- 2. **Install** (OpenClaw and the sandbox, and registers nomArmy with Claude Code):
16
+ 2. **Install and set up:**
17
17
  ```bash
18
- git clone https://github.com/rayson-tech/nomarmy.git && cd nomarmy
19
- npm install && npm link
20
- nomarmy setup --hosted
21
- ./install.sh --profile hosted
18
+ npm install -g nomarmy@alpha
19
+ cd your-project
20
+ nomarmy setup
22
21
  ```
23
- 3. **Add your workers:** `nomarmy agents add` (an API key, or your ChatGPT or Muse Code subscription), then `nomarmy army init --agent <name>` to put every role on it. `nomarmy doctor` checks the lot.
24
- 4. **Set up your repo:** in the project, run `nomarmy init`. It proposes a `.nomarmy.yml` with your test command.
25
- 5. **Use it:** restart Claude Code in that project and ask it to use nomArmy for one small bug that has a test. When that works, try `/feature <what you want built>`.
22
+ `nomarmy setup` is the playbook. It shows a checklist and runs the next step each time you say yes:
23
+ ```text
24
+ ✓ Where models run: hosted
25
+ ✓ Installed: OpenClaw 2026.9.6
26
+ → Agents: add a hosted agent
27
+ Roles: add an agent first
28
+ This repo: configure this project
29
+ Check: verify the installation
30
+ Run `nomarmy agents add` now? [Y/n]
31
+ ```
32
+ In order: pick where models run (API keys and subscriptions for most people), install OpenClaw and the sandbox, add your agents (an API key, or your ChatGPT or Muse Code subscription), put the roles on them, write this repo's `.nomarmy.yml`, then check it all. Stop anytime; `nomarmy setup` picks up where you left off.
33
+
34
+ **Want every step spelled out?** [Example setup: Claude Code, Codex and an API key](https://github.com/rayson-tech/nomarmy/blob/main/docs/setup/example.md) walks through a complete setup, command by command.
35
+ 3. **Use it:** restart Claude Code in the project and ask it to use nomArmy for one small bug that has a test. When that works, try `/feature <what you want built>`.
26
36
 
27
- **Have a GPU or a Mac with plenty of memory?** Workers can also run free on a local model: `./install.sh --profile macbook-pro` (or `nvidia-linux`, `cpu-linux`, `dgx-spark`) builds llama.cpp and starts it; see [Install](#install). A team GPU server works too: [a shared model server](#a-shared-model-server). Codex or Cursor as the coordinator: `nomarmy connect codex cursor`.
37
+ **Have a GPU or a Mac with plenty of memory?** Choose "a local model" in `nomarmy setup` and workers run on llama.cpp on your own machine: no per-token bill and your code stays home, but you pay in hardware, power and speed, and a model too big for your memory crawls. `nomarmy sizing` tells you what fits; see [Install](https://github.com/rayson-tech/nomarmy/blob/main/docs/install.md#install). A team GPU server works too: [a shared model server](https://github.com/rayson-tech/nomarmy/blob/main/docs/install.md#a-shared-model-server). Codex or Cursor as the coordinator: `nomarmy connect codex cursor`.
28
38
 
29
39
  Stuck? `nomarmy doctor` checks the machine, and `nomarmy health` checks everything nomArmy runs on.
30
40
 
@@ -32,14 +42,14 @@ Stuck? `nomarmy doctor` checks the machine, and `nomarmy health` checks everythi
32
42
 
33
43
  Your coding assistant (Claude Code, Codex or Cursor) stays in charge as the **General**: it decides what gets built and whether the result is acceptable. The work goes to **noms**, workers that implement, test and repair in their own git worktree and sandbox, on an API key, your own ChatGPT or Muse Code subscription, or a local model. nomArmy owns everything in between: worktrees, git, sandboxes, verification, and the evidence that decides whether work is accepted.
34
44
 
35
- **What you get is work you don't have to take on faith**, not cheaper work. Delegating costs the General tokens too: briefing and reviewing. On small, already-diagnosed tickets we measured 4 to 8 times more of the General's tokens than fixing the bug directly, and break-even at roughly 150 lines of context a fix needs to read ([the measurements](docs/experiments/2026-09-20-model-bakeoff-and-economics.md)). It pays off on bigger tickets, on parallel work, and anywhere you'd otherwise have to trust an agent's say-so.
45
+ **What you get is work you don't have to take on faith**, not cheaper work. Delegating costs the General tokens too: briefing and reviewing. On small, already-diagnosed tickets we measured 4 to 8 times more of the General's tokens than fixing the bug directly, and break-even at roughly 150 lines of context a fix needs to read ([the measurements](https://github.com/rayson-tech/nomarmy/blob/main/docs/experiments/2026-09-20-model-bakeoff-and-economics.md)). It pays off on bigger tickets, on parallel work, and anywhere you'd otherwise have to trust an agent's say-so.
36
46
 
37
47
  Developed and maintained by Rayson Technologies. This is an alpha (`0.1.0-alpha`).
38
48
 
39
49
  ## How it works
40
50
 
41
51
  1. The General briefs a job: a task, acceptance criteria, the tests that prove it.
42
- 2. nomArmy creates a worktree from your branch and runs the worker in a Podman sandbox with no network and no host credentials. (One exception, the Claude subscription: see [Security posture](#security-posture).)
52
+ 2. nomArmy creates a worktree from your branch and runs the worker in a Podman sandbox with no network and no host credentials. (One exception, the Claude subscription: see [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md#security-posture).)
43
53
  3. The worker edits, runs tests, and ends with a four-line report: `STATUS`, `TESTS`, `NOT_DONE`, `NOTE`.
44
54
  4. nomArmy treats that report as a claim. It reads the real diff from git, runs your verification profile itself in a fresh sandbox, reverts the production change to check the tests actually fail without it, and scans for secrets.
45
55
  5. Only then does it commit, on the worker's own branch. It never merges into yours: reviewing and integrating stay with the General, and with you.
@@ -48,435 +58,23 @@ A malformed report isn't automatically a failure: if the repository changed, nom
48
58
 
49
59
  Around that core: **agents** say where a job can run, the **army** says which role runs on which agent, and **`/feature`** runs a whole feature end to end, from plan through build, review and acceptance, handing you a branch to merge.
50
60
 
51
- ## Install
61
+ ## Docs
52
62
 
53
- | Setup | Guide |
63
+ | | |
54
64
  |---|---|
55
- | API keys and subscriptions, no local model (most people) | [Hosted workers only](#hosted-workers-only) |
56
- | A local model on macOS (Apple Silicon) | [macOS](#macos-apple-silicon) |
57
- | A local model on Linux, with or without an NVIDIA GPU | [Linux](#linux) |
58
- | Windows | [Windows](#windows) |
59
- | NVIDIA DGX Spark | [DGX Spark](#dgx-spark) |
60
- | A shared GPU server (or a tunnel to one) | [A shared model server](#a-shared-model-server) |
61
- | No GPU, with Bedrock | [Cloud (Bedrock)](#cloud-bedrock) |
62
-
63
- Every platform needs Git and Podman. `nomarmy doctor` checks the host and prints a fix for anything missing.
64
-
65
- `install.sh` builds llama.cpp when you run a local model, installs and configures [OpenClaw](https://github.com/openclaw/openclaw) (the host-side broker every model call goes through), builds the sandbox image, and registers the MCP server if Claude Code is installed. `nomarmy connect` (run by `install.sh`, or by hand for Codex and Cursor) also installs the `/feature` command, Claude Code's status line and, on macOS, nomArmy's notifier. The coordinator gets nomArmy's instructions from the MCP server itself, so there's nothing to copy into your projects.
66
-
67
- **From npm:** `npm install -g nomarmy@alpha` gives you the `nomarmy` command; `nomarmy setup` then picks a profile and model and prints the `install.sh` command to run (`nomarmy setup --hosted` or `--llama-url <server>` without a local model). Installing from a clone, as in the TL;DR, is the most tested path.
68
-
69
- ### Hosted workers only
70
-
71
- No GPU and no local model: every job runs on an API key or a subscription (ChatGPT, Muse Code) you add as an agent. Git worktrees, the sandbox and verification still run on your machine, so you still need Git, Node and Podman.
72
-
73
- ```bash
74
- git clone https://github.com/rayson-tech/nomarmy.git && cd nomarmy
75
- npm install && npm link # or: npm install -g nomarmy@alpha
76
- nomarmy setup --hosted # records that this install has no local model
77
- ./install.sh --profile hosted # OpenClaw, the sandbox, and the Claude Code registration; no llama.cpp
78
- nomarmy agents add # an API key or a subscription login
79
- nomarmy army init --agent <name> # every role on that agent (add --model <model> to pick one)
80
- nomarmy doctor
81
- ```
82
-
83
- A hosted install refuses a job that names no role or agent, rather than falling back to a local model that isn't there. `nomarmy health` warns about any role still on `local`. `e2e.sh` tests the local model, so it has nothing to do here; `nomarmy army assign` tests each role's route instead.
84
-
85
- ### macOS (Apple Silicon)
86
-
87
- ```bash
88
- git clone https://github.com/rayson-tech/nomarmy.git
89
- cd nomarmy
90
- chmod +x install.sh e2e.sh scripts/*.sh
91
- ./install.sh --profile macbook-pro
92
- ./e2e.sh --profile macbook-pro
93
- ```
94
-
95
- `install.sh` builds llama.cpp with Metal, installs and configures OpenClaw, starts the model, builds the sandbox, and registers the MCP server if Claude Code is installed. `e2e.sh` should end with:
96
-
97
- ```text
98
- PASS inference health
99
- PASS model discovery
100
- PASS worker report contract
101
- PASS autonomous edit + verification
102
- === E2E PASS ===
103
- ```
104
-
105
- If `e2e.sh` says `No API key found for provider "llama-cpp"`, run `./scripts/configure-openclaw.sh macbook-pro` and try again.
106
-
107
- ### Linux
108
-
109
- `install.sh` installs the toolchain, CMake, Git and Node when they're missing. Install Podman yourself first (`apt`, `dnf`, `zypper`, `pacman` or `apk install podman`).
110
-
111
- ```bash
112
- # no NVIDIA GPU
113
- ./install.sh --profile cpu-linux --no-claude && ./e2e.sh --profile cpu-linux
114
-
115
- # an NVIDIA GPU (not a DGX Spark: see below)
116
- ./install.sh --profile nvidia-linux --no-claude && ./e2e.sh --profile nvidia-linux
117
- ```
118
-
119
- CPU-only Linux works but is slow for interactive use: see [Sizing](#sizing).
120
-
121
- ### Windows
122
-
123
- `install.sh` doesn't run natively. In order of preference:
124
-
125
- 1. **WSL2** (the supported path): install a Linux distro under WSL2, install Podman inside it, and follow the [Linux](#linux) guide entirely inside the distro. Watch WSL2's default cap of about half your RAM (`.wslconfig`), and set `git config --global core.longpaths true` (`nomarmy doctor` checks this).
126
- 2. **Native llama.cpp on Windows**, built from source: more RAM, more setup. Prebuilt binaries aren't a safe shortcut; some CPUs crash every backend at startup.
127
- 3. **No local inference**: [hosted workers only](#hosted-workers-only), or [Bedrock](#cloud-bedrock).
128
-
129
- `nomarmy sizing` reports what your hardware can support.
130
-
131
- ### DGX Spark
132
-
133
- ```bash
134
- git clone https://github.com/rayson-tech/nomarmy.git
135
- cd nomarmy
136
- chmod +x install.sh e2e.sh scripts/*.sh
137
- ./install.sh --profile dgx-spark --no-claude
138
- ./e2e.sh --profile dgx-spark
139
- ```
140
-
141
- Moving from a Mac install? Don't copy a Mac binary or model cache over: clone fresh and let `install.sh` build llama.cpp for CUDA on that machine.
142
-
143
- ### A shared model server
144
-
145
- A team GPU box (a DGX, a workstation) runs one llama-server; everyone else points nomArmy at it. That works through an SSH tunnel too (`ssh -L 8080:localhost:8080 gpu-box`, then `http://127.0.0.1:8080`).
146
-
147
- ```bash
148
- nomarmy setup --llama-url http://gpu-box:8080 # checks /health, records the address
149
- ./install.sh --profile remote # no llama.cpp build; OpenClaw points at that server
150
- ./e2e.sh --profile remote
151
- ```
152
-
153
- nomArmy doesn't start, stop or size that server: whoever runs it sets its model, context and slots, and `install.sh` reads the model name and context from the server. Your machine still runs the sandbox and verification for your jobs. Several people sharing one server share its slots, so keep `NOMARMY_MAX_WORKERS` low (the profile sets 1). Run the server itself with any local profile's `install.sh` on the GPU machine, with `NOMARMY_LLAMA_HOST=0.0.0.0` so others can reach it, on a network you trust: llama-server has no authentication.
154
-
155
- ### Cloud (Bedrock)
156
-
157
- No GPU, no local build:
158
-
159
- ```bash
160
- export AWS_BEARER_TOKEN_BEDROCK=... # or configure an AWS profile
161
- ./install.sh --profile bedrock
162
- ./e2e.sh --profile bedrock
163
- ```
164
-
165
- Enable the models you intend to use in the Bedrock console first. `bedrock-cheap` runs the coordinator on the same open-weight model as the workers: real savings, a materially weaker guarantee (see `policies/reviewer.md`).
166
-
167
- ## Agents: where a job can run
168
-
169
- `~/.config/nomarmy/agents.yml` (or `$NOMARMY_CONFIG_DIR`) lists every account a job can run on. An api or subscription agent is the **account**, not a model: the model is chosen per role or per job, and a `model` on the agent is only an optional default.
170
-
171
- | Kind | What it is | Set up with |
172
- |---|---|---|
173
- | `local` | Your local llama-server model. `local` is built in; define another with `slot: gpt` for a second loaded model (see `config/agents.yml.example`) | nothing |
174
- | `api` | A metered API key: `xai`, `openai`, `anthropic`, `deepinfra`, `bedrock`, `azure-openai`, `openai-compatible` (with `base_url`), or `openclaw` for any other OpenClaw provider by id | `nomarmy agents add api` |
175
- | `subscription` | **One person's own** Claude, ChatGPT or Muse Code plan. Never pooled; every job on it names its owner. A Claude subscription's tools run on your machine, so it's for scouts and reviews unless you allow more (see [Security posture](#security-posture)) | `nomarmy agents add subscription claude\|codex\|meta` |
176
-
177
- ```bash
178
- nomarmy agents add # asks which kind, then walks through it
179
- nomarmy agents list # each agent, and which roles use it
180
- nomarmy agents update codex --max-concurrent 3
181
- ```
182
-
183
- `agents add` does the whole setup and asks only when something needs you:
184
-
185
- - **An api agent**: it registers the key with OpenClaw over stdin, typed once or read from an environment variable you name, and makes a test call. `agents.yml` only ever holds the variable's *name*.
186
- - **A subscription**: it installs the vendor's CLI if it's missing, runs that CLI's login, installs the OpenClaw plugin the vendor needs, lists the models your plan can use, defaults the owner to the account you logged in as, and makes a real test call before saving. For Meta it also copies the key Muse Code's login stored in your macOS keychain into OpenClaw (only that auto-connected key is flat-rate, per Meta's docs), and adds Meta's provider entry to OpenClaw's config.
187
-
188
- nomArmy starts each login; the vendor's CLI and OpenClaw do the authenticating.
189
-
190
- ```yaml
191
- # ~/.config/nomarmy/agents.yml (see config/agents.yml.example)
192
- agents:
193
- claude: { kind: subscription, provider: claude-cli, owner: you@example.com }
194
- codex: { kind: subscription, provider: openai, owner: you@example.com, max_concurrent: 3 }
195
- grok: { kind: api, provider: xai, auth_env: NOMARMY_XAI_API_KEY, thinking: high }
196
- ```
197
-
198
- **Which model runs**, first to last: the job's own `model` (the General's choice), then the role's model unless it's `auto`, then the agent's default. No model at all is refused, never guessed, and nothing picks between agents at random. A job with no role and no agent runs on `local`.
199
-
200
- **Settings**:
201
- - `max_concurrent`: how many jobs run on this agent at once, counted across every session on the machine. Defaults are 1 for a subscription and 2 for an api key; raising it spends your plan's usage limits faster. Api and subscription jobs together are also capped by `NOMARMY_MAX_POOL_WORKERS` (default 4), separately from local workers.
202
- - `thinking`: `true` follows the job's level, `false` is off, or a fixed `low`, `medium` or `high`.
203
- - `context_window`: overrides OpenClaw's catalog for a model newer than it knows.
204
-
205
- Changes apply to the next job with no restart. The exception is a **new** api agent, which needs one `nomarmy connect claude` so the MCP server sees its key variable (`agents add` offers to do it).
206
-
207
- **Why subscriptions are individual, never pooled.** Anthropic's terms separate "individual experimentation and automation" (sanctioned, including third-party apps through the Agent SDK, per seat and non-transferable) from "teams running shared production automation" (use the metered API). So a subscription agent belongs to one named person, a job on it must say `on_behalf_of: "<owner>"` or it's refused, and nothing ever load-balances across subscriptions. xAI subscriptions are out of scope until its terms are clear; an xAI API key is fine.
208
-
209
- **One credential per provider id.** OpenAI, Meta and xAI keep a subscription and an API key under the same OpenClaw provider id, so `agents.yml` refuses an api agent and a subscription agent on the same one. Claude never collides: its subscription is `claude-cli`, its API key `anthropic`. A ChatGPT plan runs as `openai/<model>` through the Codex login.
210
-
211
- **Your plan decides which models run.** A model can be listed and still refused: on a ChatGPT plan, the Codex route runs gpt-6-astra and the gpt-5.6 models but refuses gpt-6-sol and gpt-6-luna. `army assign` and `agents update --probe` test the exact route a job takes, so they catch this before a job does.
212
-
213
- **Usage limits.** nomArmy reads how much of a subscription's limit is used where the vendor reports it: Codex in every job's session log, Claude in what Claude Code passes to the status line (Pro and Max plans). It shows in `army`, `local_worker_capacity` and the status line (`⚠ openai 85% wk`), and `nomarmy health` warns from 80%. At the limit, a job on that agent is held rather than left to fail: the General asks you, and resubmits with `confirm_over_limit: true` if you say go. Muse, Grok and API keys don't report their limits; a job that hits one in a `/feature` run pauses that agent for the run.
214
-
215
- **Vendor terms and platform risk.** Every model call goes through [OpenClaw](https://github.com/openclaw/openclaw), and subscriptions are reached through each vendor's own CLI or login. We've read the terms that apply (see above), but using a personal subscription through a harness is exactly the kind of use vendors tighten, and a change in a vendor's terms or in OpenClaw can stop a subscription agent from working. Local models and API keys don't carry that risk. Plan on subscriptions as a convenience, not the only way your roles can run.
216
-
217
- **Picking an agent.** Build work goes to a sandboxed agent: `local`, an api key, Codex or Muse. `local` for a bounded change against a written spec with a test; your code never leaves your machine. An api or subscription agent when the work needs more than the local model, knowing it sends code to that vendor. That's a decision about where your source travels, separate from the trust boundary, which is the same for every agent. The General itself when the answer isn't known yet.
218
-
219
- ## The army: who does what
220
-
221
- The **General** is your coordinator session. Its charter is fixed by nomArmy: it plans and decomposes, makes architecture and security decisions, briefs and dispatches each role, reviews every result against nomArmy's verified record, owns git and integration, and gives final acceptance. It runs outside every sandbox and is never dispatched to. What you define is which agent it is:
222
-
223
- ```bash
224
- nomarmy army general claude # the agent for your own Claude seat
225
- ```
226
-
227
- That lets nomArmy flag a role that runs on the General's own agent (a review that isn't independent) or on the same subscription (the same usage limit).
228
-
229
- Every other role is yours: a name, when the General calls it, a phase, and the agent it runs on.
230
-
231
- ```bash
232
- nomarmy army init # the default roster
233
- nomarmy army assign sr-dev codex gpt-6-astra # an agent and a model
234
- nomarmy army assign pm codex auto # the General picks the model per job
235
- nomarmy army assign ui-ux codex gpt-6-astra --project # this repo, committed
236
- nomarmy army assign security-analyst grok grok-4.7 --local # just you, just this repo
237
- nomarmy army show # the General, the roster, and who set what
238
- ```
239
-
240
- `army assign` makes a real test call on the job's route before it saves, and refuses a model that doesn't run.
241
-
242
- The default roster follows a normal delivery cycle. The **Sr Dev** does the first cut and keeps the harder implementation, handing simple, well-specified work to **Jr Devs**; **UI/UX** gets UI work. Once the build is done, the specialists who apply review it (**data architect** for star schema and medallion design, **security analyst**), the **PM** reviews against the plan, and the **PO** and **stakeholders** test end to end. Not every role runs every time. Every role starts on `local`.
243
-
244
- A job dispatches with `army_role: "security-analyst"`: nomArmy runs it on that role's agent and opens the brief with the role's description. The General reads the whole roster through the read-only `army` MCP tool.
245
-
246
- **Config layers**, merged like Claude Code's settings (later wins, field by field):
247
-
248
- | Layer | File | Committed | Holds |
249
- |---|---|---|---|
250
- | global | `~/.config/nomarmy/config.yml` | no | your default army, your General, your run limits |
251
- | project | `<repo>/.nomarmy.yml` (`army:`, beside `verification:`) | yes | the team's roles for this repo (never the General) |
252
- | local | `<repo>/.nomarmy.local.yml` | no; a tracked copy is refused | your overrides for this repo |
253
-
254
- An army section can only **name agents**. It has no field for a credential, endpoint, owner or provider, so a hostile `.nomarmy.yml` in a cloned repo can at worst route a job to one of your own agents. It's also why a project file uses generic agent names (`codex`, not `jason-codex`): each teammate defines an agent by that name on their own login.
255
-
256
- **Shared machines** (a team DGX Spark): give each person their own OS account. Subscription logins live in that account's home directory and keychain, never in a nomArmy file, so each person's coordinator only reaches their own subscriptions, and nomArmy refuses an `agents.yml` another account owns or can write. One OS account shared by several people is the pooling this design exists to prevent, and `on_behalf_of` can't detect it.
257
-
258
- ## `/feature`: a feature, end to end
259
-
260
- ```
261
- /feature add join partners to the schema context
262
- ```
263
-
264
- The General runs the army's whole workflow on its own and comes back when it's done: a plan, the build (Sr Dev, Jr Dev, UI/UX), the reviews that apply (data architect, security analyst, then the PM against the plan) and acceptance (PO, stakeholder), sending fixes back to the builders along the way. It ends with **a branch for you to review and merge**. nomArmy never merges or pushes, and never deploys or touches cloud credentials; those are hard stops. Any other decision it would normally ask you about, it makes conservatively, records, and keeps going, and every such decision is in the final report.
65
+ | [Install](https://github.com/rayson-tech/nomarmy/blob/main/docs/install.md) | Every setup: hosted (API keys and subscriptions), a local model on macOS, Linux, Windows or DGX Spark, a shared model server, Bedrock |
66
+ | [Example setup](https://github.com/rayson-tech/nomarmy/blob/main/docs/setup/example.md) | Claude Code, Codex and an API key, command by command |
67
+ | [Agents and the army](https://github.com/rayson-tech/nomarmy/blob/main/docs/agents-and-army.md) | Where a job can run, who does what, usage limits, picking an agent |
68
+ | [`/feature` runs](https://github.com/rayson-tech/nomarmy/blob/main/docs/feature-runs.md) | A feature end to end, and watching what nomArmy is doing |
69
+ | [Your repository](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md) | `.nomarmy.yml`, verification, languages and dependencies, what nomArmy checks |
70
+ | [Configuration](https://github.com/rayson-tech/nomarmy/blob/main/docs/configuration.md) | Settings, swapping the local model, sizing, admission |
71
+ | [Reference](https://github.com/rayson-tech/nomarmy/blob/main/docs/reference.md) | Every CLI command and MCP tool |
72
+ | [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md) | What the sandbox holds back, and the one exception |
73
+ | [Troubleshooting](https://github.com/rayson-tech/nomarmy/blob/main/docs/troubleshooting.md) | Symptoms and fixes |
265
74
 
266
- `nomarmy connect` installs it: `/feature` in Claude Code, a `nomarmy-feature` skill in Codex, and `/feature` in Cursor (Cursor's path follows its documentation and hasn't been tested against a real install). A command of your own with the same name is never overwritten.
75
+ ## Security
267
76
 
268
- **Limits.** Each feature is a *run* (`run_start`), and its jobs join it automatically. Admission enforces the run's limits and warns at 80%:
269
-
270
- ```yaml
271
- # ~/.config/nomarmy/config.yml (or .nomarmy.local.yml; never the committed .nomarmy.yml)
272
- army:
273
- run_limits:
274
- max_jobs: 40 # the defaults
275
- max_api_usd: 10 # api agents only; a subscription isn't billed per call
276
- max_hours: 6
277
- warn_at: 0.8
278
- ```
279
-
280
- The General can lower these for one run, never raise them. When a vendor answers with a usage-limit error, that agent is paused for the rest of the run and the General stops and tells you; it never moves the role to another vendor to get around it. The one limit no tool can see is your coordinator's own seat. If that runs out mid-feature, the run log (kept current after every phase) lets `/feature resume <run-id>` in a fresh session carry on.
281
-
282
- ## Watching what nomArmy is doing
283
-
284
- - **Claude Code's status line** shows what's running in this repo, a count for other repos, the open run, and the most serious health warning: `Opus 5.5 · rayson-senti │ 🍪 2: sr-dev codex 9m 10f · scout grok 1m │ run 3/14 $0.41`. `nomarmy connect claude` installs it unless you have your own; then `nomarmy statusline` prints nomArmy's part for you to add.
285
- - **Desktop notifications** when a job finishes, a run crosses a limit or pauses an agent, or a health check finds a new problem. They come from nomArmy itself, so they work with any coordinator; on macOS they carry nomArmy's icon. `NOMARMY_NOTIFY=0` turns them off.
286
- - **`nomarmy jobs --watch`** is a live table; **`nomarmy jobs --events`** prints one line per job start, phase change and finish, which the General watches in the background instead of polling.
287
- - **`run_status`** lists a run's running and finished jobs, with each one's phase, last tool call and files changed so far.
288
- - **Health checks** run a minute after each server starts and every 6 hours after that: logins about to expire, OpenClaw and plugin versions, roles that can't run, providers missing from OpenClaw's config, models refused on a real job, and storage. `nomarmy health` runs them now.
289
-
290
- **Storage** looks after itself. OpenClaw's scratch files (about 1.2 GB for a Codex job) go when each call ends, and each health check removes finished jobs' remaining runtime data after a day (`NOMARMY_AUTO_PRUNE_HOURS`, `0` to turn it off), keeping every job's record and report. `nomarmy jobs --prune --older-than 0` does it for every finished job now.
291
-
292
- ## Your repository
293
-
294
- ### `.nomarmy.yml`
295
-
296
- A repo's execution contract: its verification profiles and, optionally, its army and dependency settings. `nomarmy init` proposes one from what it finds (compose files, CI steps, requirements files, test commands) and writes it only after you confirm. nomArmy reads it from your checkout, never from a job's worktree, so a worker can't weaken its own checks.
297
-
298
- ```yaml
299
- verification:
300
- quick:
301
- environment: none
302
- commands:
303
- - npm test
304
- ```
305
-
306
- `nomarmy validate` checks the file against the schema; `nomarmy scan --check` diffs it against what the repo actually contains.
307
-
308
- **Policy: what no job can skip.**
309
-
310
- ```yaml
311
- policy:
312
- require_verification: true # every implement job needs a verification profile; only passing work commits
313
- require_regression_check: true # verify_regression can't be switched off per job
314
- ```
315
-
316
- `nomarmy init` proposes both for new repos. Without them, a job with no verification profile still commits (flagged for review, not blocked), and the General decides per job whether to run the revert check. With them, those are the repo's rules, not the General's judgment calls, and since nomArmy reads this file only from your checkout, neither the General nor a worker can relax it.
317
-
318
- **Refactors.** Reverting a behavior-preserving change restores code that works, so the revert check can't prove anything about it. A job can declare `refactor: true` instead: nomArmy then commits it only if verification passes **and no test file was added, changed or deleted**. The existing tests passing unchanged is the evidence. A change that alters behavior has to alter tests to show it, so it can't pass as a refactor.
319
-
320
- **Add a check for what unit tests can't see.** A module left out of a deploy bundle passes every unit test and crashes at deploy. When `nomarmy init` sees a bundle or packaging step (Lambda asset scripts, SAM, Serverless, CDK), it suggests a profile that runs it and then imports each entry point from the built bundle.
321
-
322
- ### Languages and dependencies
323
-
324
- The sandbox has no network, so dependencies are installed when its image is built, on your machine, and the image is cached by a hash of the dependency files.
325
-
326
- | Repo | Detected by | Sandbox |
327
- |---|---|---|
328
- | Node | every package with its own `package-lock.json` or `npm-shrinkwrap.json`: the root, and any others (a `ui/`, a `lambda/api/`) | `npm ci` for each at build time, under `/deps` at the same path. The root's packages are at `/node_modules`; each other package gets a `node_modules` link into the image, which nomArmy never commits |
329
- | Python | `requirements.txt`, or `environment.python.requirements` | `pip install` at build time |
330
- | Both | both of the above | one image with both |
331
- | Go, Rust | `go.mod`, `Cargo.toml` | the toolchain, built once and cached |
332
- | anything else | | the base image: Node, Python 3, git, ripgrep |
333
-
334
- The worker's own tool calls and nomArmy's verification use the same image. `NOMARMY_AGENT_IMAGE` overrides detection, and `environment.node.install: false` turns the Node install off.
335
-
336
- ```yaml
337
- environment:
338
- python:
339
- requirements:
340
- - requirements-dev.txt
341
- - lambda/requirements.txt
342
- ```
343
-
344
- A package whose install fails (a private registry, say) is marked and skipped; the rest still install. npm workspaces, yarn, pnpm and bun aren't installed yet. For those repos, verification borrows your own checkout's `node_modules` read-only, which works for plain JavaScript packages but not for ones with native binaries built for your host.
345
-
346
- ### Scoping verification to the diff
347
-
348
- A command scoped by a hand-maintained filter (`pytest -k`) can silently skip the file a worker changed. Every verification command gets two variables to scope by instead:
349
-
350
- | Variable | Contents |
351
- |---|---|
352
- | `NOMARMY_CHANGED_TEST_FILES` | New and modified test files, space-separated |
353
- | `NOMARMY_CHANGED_PRODUCTION_FILES` | Non-test files touched, space-separated |
354
-
355
- ```yaml
356
- verification:
357
- python:
358
- commands:
359
- - 'if [ -n "$NOMARMY_CHANGED_TEST_FILES" ]; then python3 -m pytest $NOMARMY_CHANGED_TEST_FILES -q; fi'
360
- - 'if [ -n "$NOMARMY_CHANGED_PRODUCTION_FILES" ]; then python3 -m pytest lambda/tests/ -q; fi'
361
- ```
362
-
363
- Run the narrow pass on the touched files for a fast, sharp signal, *and* the broad pass whenever production code changed: a shared module can have far more dependents than the files a diff happens to touch. Running the changed tests on their own also catches a test that only passes inside the full suite.
364
-
365
- ### What else nomArmy checks
366
-
367
- - **Tests that prove nothing.** With `verify_regression` (on whenever a job has a verification profile), nomArmy reverts the production change and re-runs the tests: a test that still passes is flagged.
368
- - **Tests made to pass.** New skip markers, stubbed imports, fake modules named like a dependency, and stray backup files are flagged for review.
369
- - **Code wired to nothing.** A new function or class that nothing outside its own test calls is flagged (heuristic and review-only).
370
- - **Secrets.** Every diff and report is scanned for known secret shapes (secretlint's recommended preset) before a commit is allowed; a match blocks it.
371
-
372
- ## Configuration
373
-
374
- Local inference is configured in two files, and a shell-exported variable overrides both. Every variable is documented where it's set.
375
-
376
- - **`config/common.env`**: defaults shared by every profile.
377
- - **`config/profiles/<name>.env`**: per-machine overrides (context size, GPU layers, threads, workers).
378
-
379
- | Variable | Lives in | Default | Controls |
380
- |---|---|---|---|
381
- | `NOMARMY_WORKER_MODEL` | common.env | `gpt-oss-20b` | Which local model runs jobs |
382
- | `NOMARMY_MODEL_REPO` / `_QUANT` | common.env | `ggml-org/gpt-oss-20b-GGUF` / `MXFP4` | Which GGUF to download |
383
- | `NOMARMY_LLAMA_CONTEXT` | profile | `65536` | **Total** context across all slots |
384
- | `NOMARMY_LLAMA_PARALLEL` | profile | `1` | Inference slots (the context is divided across them) |
385
- | `NOMARMY_MAX_WORKERS` | profile | `1` | How many local jobs run at once |
386
- | `NOMARMY_MAX_POOL_WORKERS` | environment | `4` | How many api and subscription jobs run at once |
387
- | `NOMARMY_EXECUTION` | common.env | `local` | `local` or `bedrock` |
388
- | `NOMARMY_ORCHESTRATOR_TRUST` | common.env | `frontier` | `frontier` or `degraded`: see `policies/reviewer.md` |
389
-
390
- **The coupling that trips people up**: `NOMARMY_LLAMA_CONTEXT` is divided by `NOMARMY_LLAMA_PARALLEL`, not given to each slot whole. `65536` across 2 slots is `32768` per nom.
391
-
392
- A change needs an inference restart: `nomarmy stop && nomarmy start` (logs go to `~/.local/share/nomarmy-local-agents/logs/`). `nomarmy config paths` shows where every config file lives.
393
-
394
- ### Swapping the local model
395
-
396
- ```bash
397
- nomarmy model
398
- ```
399
-
400
- It offers three measured choices (gpt-oss-20b, Qwen3-Coder-Next and Qwen3.6-27B: see the [model bake-off](docs/experiments/2026-09-20-model-bakeoff-and-economics.md)) or a Hugging Face search. It also offers to update the MCP registration and restart inference, since the config, the registration and the running model all need to agree.
401
-
402
- ### Sizing
403
-
404
- Three coupled settings: `NOMARMY_LLAMA_CONTEXT` (total context), `NOMARMY_LLAMA_PARALLEL` (slots) and `NOMARMY_MAX_WORKERS` (concurrent local jobs). `nomarmy sizing` reads your hardware and the model's GGUF metadata, checks live memory pressure, and recommends a combination.
405
-
406
- - **Aim for 64K context per nom.** An autonomous explore, implement, test and repair loop needs more room than a one-shot edit.
407
- - **More noms isn't automatically faster.** On one Apple Silicon machine, 4 parallel local workers produced no more accepted work than 1. Measure before raising it.
408
- - **Speed matters more than fit.** CPU-only inference measured about 3.8 tokens per second on a 20-core i7 (an 11.5-minute job for two turns). Use a GPU or a hosted agent for interactive work; CPU-only is a correctness testbed.
409
-
410
- Advanced llama-server tuning (`NOMARMY_LLAMA_CACHE_TYPE_K/V`, `_FLASH_ATTN`, `_REASONING_BUDGET`, `_REASONING_PRESERVE`, `_EXTRA_ARGS`) is documented in `config/profiles/*.env`; all unset by default.
411
-
412
- ### Admission and budgets
413
-
414
- A job is admitted only when there's context and memory for it, and briefs and reports are sized to the agent. The local model keeps caps calibrated on a 20B model, where longer briefs made it thrash: a 3,000-character brief, 6,000 characters of evidence, a 512-token implement report. An api or subscription agent gets ceilings that scale with its model's context, up to a 16,000-character brief and 24,000 characters of evidence. A job's `report` (`brief`, `standard`, `full`) sets how much comes back, up to about 2k tokens for an implement job and 4k for a scout. The report lands in the General's own context, so that's its call per job.
415
-
416
- ## Command reference
417
-
418
- ### The `nomarmy` CLI
419
-
420
- Every command proposes before it writes: a `[y/N]` prompt, or an explicit flag under `--json`. All take `--json` and `--repo <dir>`.
421
-
422
- | Command | What it does |
423
- |---|---|
424
- | `nomarmy doctor` | Checks this machine is ready, with a fix for anything missing. Start here. |
425
- | `nomarmy setup` | Detects the machine, recommends a profile, picks a model, writes config. Prints (never runs) `install.sh`. |
426
- | `nomarmy connect [claude] [codex] [cursor]` | Registers nomArmy with each coordinator and installs `/feature`, the status line and the notifier. No target: pick interactively. |
427
- | `nomarmy init` | Proposes a `.nomarmy.yml` from what the repo contains. |
428
- | `nomarmy agents list\|add\|update\|remove` | Where jobs can run. See [Agents](#agents-where-a-job-can-run). |
429
- | `nomarmy army show\|init\|assign\|general` | The General and the roster. See [The army](#the-army-who-does-what). |
430
- | `nomarmy jobs [--watch\|--events\|--prune]` | What's running across every session, and what just finished. |
431
- | `nomarmy health` | Runs the health checks now. |
432
- | `nomarmy statusline` | nomArmy's part of Claude Code's status line. |
433
- | `nomarmy config paths` | Where each config file lives. |
434
- | `nomarmy model` | Changes the local model. |
435
- | `nomarmy sizing` | Recommends context, slots and workers. `--check` evaluates the loaded profile; `--noms N` sizes for a count. |
436
- | `nomarmy start` / `stop` | Starts or stops local inference. |
437
- | `nomarmy scan` | Reports the repo's execution environment. `--check` diffs it against `.nomarmy.yml`. |
438
- | `nomarmy validate` | Validates `.nomarmy.yml`. |
439
- | `nomarmy update` | From a clone: pulls (fast-forward only) and resyncs what each coordinator runs. From npm: tells you the npm command. |
440
- | `nomarmy uninstall` | Removes the MCP registration and install. `--clear-agents`, `--clear-models` or `--all` go further. |
441
-
442
- ### MCP tools
443
-
444
- What the General uses. Every job takes the same shape: a `task`, optional `acceptance`, a `mode` and a timeout.
445
-
446
- | Tool | What it does |
447
- |---|---|
448
- | `local_worker` | Runs one job and waits for it. |
449
- | `local_worker_start` / `local_worker_status` | Starts a job in the background / waits for its result. |
450
- | `local_workers` | Runs a batch of independent jobs in parallel. `auto_union: true` merges them into one integration branch for review. |
451
- | `repo_evidence` | Deterministic answers (definitions, references, outlines, grep, files) with `[path:line]` on every hit, no model. |
452
- | `army` | The General's charter and agent, then this repo's roles and who runs each. |
453
- | `run_start` / `run_status` / `run_finish` | A `/feature` run: its limits, usage per agent, warnings, paused agents and log. |
454
- | `local_worker_capacity` | Context, budgets, memory pressure and what's running. |
455
- | `local_worker_jobs` | Recent job records; `full: true` for the complete manifest. |
456
- | `local_worker_config` | This repo's verification profiles. |
457
- | `local_worker_cleanup` | Removes one worktree and branch. Recognizes a cherry-picked branch as integrated by content. |
458
- | `local_worker_sweep` | Removes worktrees that are provably empty. `dry_run` previews. |
459
-
460
- A job's `mode` is `implement` (edits, then nomArmy verifies and commits), `scout` (read-only research, every claim cited as `[path:start-end]` and checked against the base commit) or `decompose` (read-only, proposes independent subtasks for the General to dispatch).
461
-
462
- ## Security posture
463
-
464
- The worker gets a writable worktree inside Podman and nothing else: no Podman socket, no host credentials, no network. Repository content is untrusted input, and `.nomarmy.yml` is data to validate, never authority.
465
-
466
- **The exception: a Claude subscription runs its tools on your machine.** OpenClaw reaches a Claude plan by running the real `claude` command on the host, and Claude Code's own tools (Bash, Edit, Write) run there, with your files and the network, not in the sandbox. We checked each route by having a worker report where its shell ran:
467
-
468
- | Agent | Its tools run |
469
- |---|---|
470
- | `local`, api keys (xAI, OpenAI, Anthropic, ...), ChatGPT via Codex, Muse Code | in the sandbox: Linux, `/workspace`, no network |
471
- | Claude subscription (`claude-cli`) | **on this machine**: your real paths, with network |
472
-
473
- So nomArmy refuses **implement** jobs on a Claude subscription unless that agent says `allow_host_tools: true` in `agents.yml`; scouts and reviews still run, labeled. `nomarmy health` and the `army` tool flag any build role on it, and `agents list` says so. If a job's worktree comes back with a real `node_modules` where nomArmy's dependency link was (packages installed where the sandbox couldn't have), nomArmy flags the job for review and verifies against the sandbox's own dependencies. Sandboxing the Claude route properly needs OpenClaw to run it with only OpenClaw's own (sandboxed) tools, which it supports internally but doesn't expose yet. An Anthropic **api key** runs through OpenClaw's own loop and is sandboxed like the rest.
474
-
475
- **Never hand a worker** AWS or production credentials, deployment access, SSH keys, Kubernetes contexts or Terraform state.
476
-
477
- Every model call, local, api or subscription, is made by OpenClaw on the host, never from inside the sandbox. A subscription is reached through the vendor's own logged-in session; nomArmy never reads or stores the token. What changes with a hosted agent or a Bedrock profile is where your code goes (to that vendor), not what the sandbox can reach.
478
-
479
- `on_behalf_of` is a self-reported attestation, not a verified identity: nomArmy has no caller-identity boundary. The secret scan catches known secret shapes, not steered content with no recognizable shape. Both are covered in [SECURITY.md](SECURITY.md), which is also where to report a vulnerability.
77
+ A worker gets a writable git worktree inside a Podman sandbox and nothing else: no network, no host credentials, no Podman socket. Every model call is made by OpenClaw on your machine, never from inside the sandbox. **The exception is a Claude subscription**, whose tools run on your machine, so nomArmy refuses build jobs on it unless you allow it. Never hand a worker production credentials, deployment access or SSH keys. Details: [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md); to report a vulnerability, [SECURITY.md](https://github.com/rayson-tech/nomarmy/blob/main/SECURITY.md).
480
78
 
481
79
  ## Status
482
80
 
@@ -494,38 +92,26 @@ Every model call, local, api or subscription, is made by OpenClaw on the host, n
494
92
  | Disposable per-job services (Postgres, mocks) | Not built |
495
93
  | Browser/E2E testing inside a nom | Not built |
496
94
 
497
- What we've learned from real runs, including where delegating pays and where it doesn't, is in [docs/findings.md](docs/findings.md).
95
+ What we've learned from real runs, including where delegating pays and where it doesn't, is in [docs/findings.md](https://github.com/rayson-tech/nomarmy/blob/main/docs/findings.md).
498
96
 
499
97
  ### Known limitations
500
98
 
501
- - **A Claude subscription isn't sandboxed.** Its tools run on your machine, so implement jobs on it are refused unless you set `allow_host_tools: true`. See [Security posture](#security-posture).
99
+ - **A Claude subscription isn't sandboxed.** Its tools run on your machine, so implement jobs on it are refused unless you set `allow_host_tools: true`. See [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md#security-posture).
502
100
  - **A refused model costs one job.** When a vendor refuses a model at run time that OpenClaw lists (gpt-6-sol on a ChatGPT plan), the first job on it fails with `model_not_found`. After that nomArmy refuses to dispatch it until a job or test call on it works. `army assign` tests the job's route and catches this before any job.
503
101
  - **Claude subscription token counts** come from the Claude CLI's own session log, since OpenClaw sees only the final reply. Totals include cache reads and writes, which make up most of an agent's prompt; each part is also kept separately.
504
102
  - **Test-workaround detection is a flag, not a verdict**: a legitimate new skip still gets flagged.
505
- - **Deploy-time failures need your own check.** See [Add a check for what unit tests can't see](#nomarmyyml).
103
+ - **Deploy-time failures need your own check.** See [Add a check for what unit tests can't see](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md#nomarmyyml).
506
104
  - **Node dependencies install only from npm lockfiles**, one per package (no workspaces, yarn, pnpm or bun yet), and only from the public registry: the image build has no credentials for a private one.
507
105
  - **Verification needing services** (a database, a mock server) reports `not_run` instead of running without them. The compose parser doesn't resolve YAML anchors.
508
- - **Same-host sandboxes**: the MCP server, OpenClaw and every job's sandbox run on the machine with the coordinator. Only the model can be elsewhere (an agent, or [a shared model server](#a-shared-model-server)).
509
-
510
- ## Troubleshooting
511
-
512
- | Symptom | Fix |
513
- |---|---|
514
- | Anything unclear about the machine | `nomarmy doctor`, then `nomarmy health` |
515
- | `No API key found for provider "llama-cpp"` during `e2e.sh` | `./scripts/configure-openclaw.sh <profile>`, then rerun |
516
- | A job fails with `model_not_found` | Your plan or OpenClaw can't run that model. `nomarmy agents list` shows which roles use it; `nomarmy army assign <role> <agent> <model>` moves the role and tests the new model |
517
- | Verification fails on a missing package | Node: commit each package's `package-lock.json` (yarn, pnpm and workspaces aren't installed yet). Python: list your requirements files under `environment.python.requirements` |
518
- | A config change didn't take effect | `nomarmy stop && nomarmy start` for inference; restart your coordinator after `nomarmy connect` or an update |
519
- | `git worktree add` fails with `Filename too long` (Windows/WSL2) | `git config --global core.longpaths true` |
520
- | Not sure a setting fits your hardware | `nomarmy sizing`, or `nomarmy sizing --check` |
106
+ - **Same-host sandboxes**: the MCP server, OpenClaw and every job's sandbox run on the machine with the coordinator. Only the model can be elsewhere (an agent, or [a shared model server](https://github.com/rayson-tech/nomarmy/blob/main/docs/install.md#a-shared-model-server)).
521
107
 
522
108
  ## More
523
109
 
524
- - [docs/findings.md](docs/findings.md): what real runs taught us.
525
- - [docs/experiments](docs/experiments): the measured write-ups behind them.
526
- - [policies](policies): how scouts and reviewers are held to account.
527
- - [CONTRIBUTING.md](CONTRIBUTING.md), [SECURITY.md](SECURITY.md), [CODE_OF_CONDUCT.md](CODE_OF_CONDUCT.md).
110
+ - [docs/findings.md](https://github.com/rayson-tech/nomarmy/blob/main/docs/findings.md): what real runs taught us.
111
+ - [docs/experiments](https://github.com/rayson-tech/nomarmy/tree/main/docs/experiments): the measured write-ups behind them.
112
+ - [policies](https://github.com/rayson-tech/nomarmy/tree/main/policies): how scouts and reviewers are held to account.
113
+ - [CONTRIBUTING.md](https://github.com/rayson-tech/nomarmy/blob/main/CONTRIBUTING.md), [SECURITY.md](https://github.com/rayson-tech/nomarmy/blob/main/SECURITY.md), [CODE_OF_CONDUCT.md](https://github.com/rayson-tech/nomarmy/blob/main/CODE_OF_CONDUCT.md).
528
114
 
529
115
  Configuration variables are `NOMARMY_*`; a legacy `RAYSON_*` variable is translated once at load, with a deprecation warning.
530
116
 
531
- Licensed under [Apache 2.0](LICENSE).
117
+ Licensed under [Apache 2.0](https://github.com/rayson-tech/nomarmy/blob/main/LICENSE).
package/bin/nomarmy.mjs CHANGED
@@ -2,10 +2,10 @@
2
2
  // nomArmy CLI. Every command proposes before it writes anything -- init,
3
3
  // setup, model and update all show exactly what would change and write only
4
4
  // after explicit confirmation ([y/N]) or an explicit non-interactive flag
5
- // (--write, --json with the required choices given up front). Nothing here
6
- // provisions SYSTEM-level infrastructure on its own: install.sh (builds
7
- // llama.cpp, installs OpenClaw, configures the sandbox) stays a separate,
8
- // manual step in every case, printed but never run.
5
+ // (--write, --json with the required choices given up front). System-level
6
+ // setup (install.sh: OpenClaw, the sandbox, llama.cpp for a local model)
7
+ // runs only when asked: `nomarmy install`, or `nomarmy setup` after it has
8
+ // shown the exact command and the operator said yes.
9
9
  import fs from "node:fs";
10
10
  import os from "node:os";
11
11
  import path from "node:path";
@@ -20,9 +20,11 @@ import { readGGUFMetadata, resolveModelPath, totalSplitBytes } from "../lib/gguf
20
20
  import { recommend, customRecommendation, evaluateConfig, bytesPerKvElementForCacheTypes, MIN_CONTEXT_PER_NOM } from "../lib/sizing.mjs";
21
21
  import { connectClaude, connectCodex, connectCursor, cursorAlreadyConnected, deriveWorkerModelEnv } from "../lib/connect.mjs";
22
22
  import { ID_RE, AUTH_ENV_NAME_RE, OPENCLAW_PROVIDER_ID_RE, openclawProviderId, isNativeProviderType } from "../lib/dispatch-schema.mjs";
23
- import { loadAgents, readAgentsFile, writeAgentsFile, agentsConfigPath, apiAgentAsPoolEntry, describeAgent as describeAgentLabel, agentRunsToolsOnHost, AGENT_KINDS, API_PROVIDER_TYPES, RESERVED_AGENT_NAMES, BUILTIN_LOCAL_AGENT } from "../lib/agents.mjs";
24
- import { loadArmy, describeArmy, readArmyFile, updateArmyInFile, assignRoleInFile, parseTargetSpec, armyLayerPath, globalConfigDir, DEFAULT_ARMY, ARMY_PHASES, LOCAL_CONFIG_FILENAME } from "../lib/army.mjs";
23
+ import { loadAgents, readAgentsFile, writeAgentsFile, agentsConfigPath, apiAgentAsPoolEntry, describeAgent as describeAgentLabel, agentRunsToolsOnHost, agentProviderId, AGENT_KINDS, API_PROVIDER_TYPES, RESERVED_AGENT_NAMES, BUILTIN_LOCAL_AGENT } from "../lib/agents.mjs";
24
+ import { loadArmy, mergeArmy, describeArmy, readArmyFile, updateArmyInFile, assignRoleInFile, parseTargetSpec, armyLayerPath, globalConfigDir, DEFAULT_ARMY, ARMY_PHASES, LOCAL_CONFIG_FILENAME } from "../lib/army.mjs";
25
25
  import { parseLlamaUrl } from "../lib/execution.mjs";
26
+ import { setupSteps, formatSetupSteps, runSetupPlaybook } from "../lib/setup-steps.mjs";
27
+ import { readUsageSnapshots } from "../lib/usage-limits.mjs";
26
28
  import { ensureProviderConfig } from "../lib/openclaw-config.mjs";
27
29
  import { recordProbeSuccess } from "../lib/health.mjs";
28
30
  import { pruneJobRuntime } from "../lib/prune.mjs";
@@ -89,13 +91,16 @@ Usage: nomarmy <command> [options]
89
91
  and write it after confirmation.
90
92
  --force overwrite an existing .nomarmy.yml
91
93
  --write with --json, write without prompting (needs a valid proposal)
92
- setup Detect this machine, recommend a profile (offering "more
93
- noms" vs "nominal" when they differ), choose a model, and
94
- write config/profiles/<name>.env (+ config/common.env).
95
- Prints the install.sh command; never runs it.
96
- --tier <more|nominal> with --json, skip the prompt
94
+ setup Show setup progress and run the next unfinished step.
95
+ --status [--json] print the checklist only
96
+ --choose choose hosted/local/remote/Bedrock
97
97
  --hosted skip hardware/model questions
98
- --llama-url <url> use a llama-server on another machine
98
+ --llama-url <url> use a shared llama-server
99
+ --json --profile-name <name> [--model <model>] [--tier <more|nominal>]
100
+ write a profile non-interactively
101
+ install Run the bundled installer for the chosen setup profile.
102
+ --profile <name> override NOMARMY_SETUP_PROFILE
103
+ --no-claude skip Claude Code registration
99
104
  model Change the configured model later, without the rest of
100
105
  setup's questions. Offers to also resync the MCP
101
106
  registration's worker-routing env vars, and to restart
@@ -509,18 +514,108 @@ async function chooseModel(rl) {
509
514
  return { kind: "search", thinking: thinkingAnswer === "y" || thinkingAnswer === "yes" };
510
515
  }
511
516
 
512
- /**
513
- * `nomarmy setup`: detect hardware, recommend a profile the same way
514
- * `nomarmy sizing` already does, let the user pick a model, then write the
515
- * result to config/profiles/<name>.env and (for the two curated model
516
- * choices) config/common.env. Stops there -- prints the exact `install.sh`
517
- * command rather than running it. install.sh builds llama.cpp, curl-pipes an
518
- * installer and touches sandbox/provider config; that is not a proportionate
519
- * thing for an opt-in flag on a CLI whose whole brand is "reports or
520
- * proposes" to cross, unlike the cheap, reversible, single-file writes this
521
- * command itself does.
522
- */
517
+ function setupProjectDir() {
518
+ let project = process.cwd();
519
+ while (!fs.existsSync(path.join(project, ".git"))) {
520
+ const parent = path.dirname(project);
521
+ if (parent === project) { project = null; break; }
522
+ project = parent;
523
+ }
524
+ return project;
525
+ }
526
+
527
+ /** The profile install.sh picks when given none (scripts/lib.sh load_profile). */
528
+ function defaultLocalProfile() {
529
+ if (process.platform === "darwin") return "macbook-pro";
530
+ return spawnSync("nvidia-smi", ["-L"], { stdio: "ignore", timeout: 5000 }).status === 0 ? "nvidia-linux" : "cpu-linux";
531
+ }
532
+
533
+ function setupChecklist() {
534
+ const common = path.join(nomarmyRoot, "config", "common.env");
535
+ const chosen = readEnvValue(common, "NOMARMY_SETUP_PROFILE");
536
+ const probeCommand = (binary, args) => {
537
+ const result = spawnSync(binary, args, { encoding: "utf8", timeout: 10000 });
538
+ return result.status === 0 ? result.stdout.trim() : "";
539
+ };
540
+ const project = setupProjectDir();
541
+ const profileFile = chosen ? path.join(nomarmyRoot, "config", "profiles", `${chosen}.env`) : null;
542
+ const root = (process.env.NOMARMY_INSTALL_ROOT || (profileFile && readEnvValue(profileFile, "NOMARMY_INSTALL_ROOT")) || readEnvValue(common, "NOMARMY_INSTALL_ROOT") || "$HOME/.local/share/nomarmy-local-agents").replace(/\$HOME|\$\{HOME\}/g, os.homedir());
543
+ let marker = null;
544
+ try { marker = JSON.parse(fs.readFileSync(path.join(root, "install.json"), "utf8")); } catch (error) { if (error.code !== "ENOENT") marker = {}; }
545
+ const version = probeCommand("openclaw", ["--version"]);
546
+ // `mcp list` would connect to every server; `mcp get` only looks this one up.
547
+ const registered = !marker && Boolean(version) && ["claude", "codex"].some((name) => spawnSync(name, ["mcp", "get", "nomarmy-local-worker"], { stdio: "ignore", timeout: 10000 }).status === 0);
548
+ // An install from before setup recorded its profile: the marker's, else the
549
+ // execution mode's, else (a working local install) install.sh's default.
550
+ const execution = readEnvValue(common, "NOMARMY_EXECUTION");
551
+ const profile = chosen ?? marker?.profile
552
+ ?? (["hosted", "remote", "bedrock"].includes(execution) ? execution : null)
553
+ ?? (registered ? defaultLocalProfile() : null);
554
+ return setupSteps({
555
+ mode: () => ({ profile, host: readEnvValue(common, "NOMARMY_LLAMA_HOST"), port: readEnvValue(common, "NOMARMY_LLAMA_PORT") }),
556
+ install: () => ({ marker, version, registered }),
557
+ agents: () => Object.keys(loadAgents(globalConfigDir()).agents),
558
+ army: () => {
559
+ const global = readArmyFile(armyLayerPath("global"), { armyOnly: true });
560
+ const local = project ? readArmyFile(armyLayerPath("project", { projectDir: project })) : null;
561
+ const merged = mergeArmy([{ layer: "global", army: global }, { layer: "project", army: local }]).army;
562
+ // Repair the layer that would otherwise keep overriding a global init.
563
+ const repairLayer = profile === "hosted" && Object.values(local?.roles ?? {}).some((role) => role.agent === "local") ? "project" : null;
564
+ return { ...merged, repairLayer };
565
+ },
566
+ repo: () => ({ inside: Boolean(project), configured: Boolean(project && fs.existsSync(path.join(project, ".nomarmy.yml"))) }),
567
+ });
568
+ }
569
+
570
+ function runSetupChild(args) {
571
+ // Work from the repository root even when setup was started in a subdirectory;
572
+ // the child's cwd remains unchanged, as it does for every other step.
573
+ const project = setupProjectDir();
574
+ if (project && ["init", "army"].includes(args[0])) args = [...args, "--repo", project];
575
+ const result = spawnSync(process.execPath, [path.join(nomarmyRoot, "bin", "nomarmy.mjs"), ...args], { stdio: "inherit", cwd: process.cwd() });
576
+ return result.status ?? 1;
577
+ }
578
+
579
+ function cmdInstall() {
580
+ const profile = value("profile", readEnvValue(path.join(nomarmyRoot, "config", "common.env"), "NOMARMY_SETUP_PROFILE"));
581
+ if (!profile) throw new Error("Choose a profile first: nomarmy setup --choose (or install --profile <name>).");
582
+ const result = spawnSync("bash", [path.join(nomarmyRoot, "install.sh"), "--profile", profile, ...(flag("no-claude") ? ["--no-claude"] : [])], { stdio: "inherit", cwd: nomarmyRoot });
583
+ process.exitCode = result.status ?? 1;
584
+ }
585
+
523
586
  async function cmdSetup() {
587
+ if (flag("status") || (!flag("choose") && !flag("hosted") && !flag("llama-url") && !json)) {
588
+ if (flag("status") || !process.stdin.isTTY) {
589
+ const steps = setupChecklist();
590
+ return json ? out(steps) : console.log(formatSetupSteps(steps));
591
+ }
592
+ process.exitCode = await runSetupPlaybook({
593
+ evaluate: setupChecklist, print: console.log, run: runSetupChild,
594
+ ask: async (prompt) => {
595
+ const rl = createInterface({ input, output });
596
+ try { return await rl.question(prompt); } finally { rl.close(); }
597
+ },
598
+ });
599
+ return;
600
+ }
601
+ if (flag("choose")) {
602
+ if (!process.stdin.isTTY) throw new Error("setup --choose needs an interactive terminal.");
603
+ const rl = createInterface({ input, output });
604
+ try {
605
+ console.log("1. hosted (API keys and subscriptions, most people)\n2. a local model on this machine\n3. a shared model server\n4. Bedrock");
606
+ const choice = await askUntilValid(rl, "Choice [1]: ", { pattern: /^[1-4]$/, invalidMessage: "Choose 1, 2, 3 or 4.", allowEmpty: true, fallback: "1" });
607
+ if (choice === "1") argv.push("--hosted");
608
+ if (choice === "3") argv.push("--llama-url", (await rl.question("Server URL: ")).trim());
609
+ if (choice === "4") {
610
+ const common = path.join(nomarmyRoot, "config", "common.env");
611
+ fs.mkdirSync(path.dirname(common), { recursive: true });
612
+ writeEnvLine(common, "NOMARMY_EXECUTION", "bedrock");
613
+ writeEnvLine(common, "NOMARMY_SETUP_PROFILE", "bedrock");
614
+ console.log("Next: nomarmy install");
615
+ return;
616
+ }
617
+ } finally { rl.close(); }
618
+ }
524
619
  const hosted = flag("hosted");
525
620
  const hasLlamaUrl = flag("llama-url");
526
621
  if (hosted && hasLlamaUrl) throw new Error("--hosted and --llama-url cannot be used together.");
@@ -530,12 +625,13 @@ async function cmdSetup() {
530
625
 
531
626
  if (hosted) {
532
627
  const next = [
533
- `${path.join(nomarmyRoot, "install.sh")} --profile hosted`,
628
+ "nomarmy install",
534
629
  "nomarmy agents add",
535
630
  "nomarmy army init --agent <name>",
536
631
  ];
537
632
  fs.mkdirSync(path.dirname(commonPath), { recursive: true });
538
633
  writeEnvLine(commonPath, "NOMARMY_EXECUTION", "hosted");
634
+ writeEnvLine(commonPath, "NOMARMY_SETUP_PROFILE", "hosted");
539
635
  if (json) return out({ written: commonPath, execution: "hosted", next });
540
636
  console.log(c.green(`✓ Wrote NOMARMY_EXECUTION=hosted to ${path.relative(nomarmyRoot, commonPath)}.`));
541
637
  console.log(c.dim("\nNext:"));
@@ -557,9 +653,10 @@ async function cmdSetup() {
557
653
  }
558
654
  fs.mkdirSync(path.dirname(commonPath), { recursive: true });
559
655
  writeEnvLine(commonPath, "NOMARMY_EXECUTION", "remote");
656
+ writeEnvLine(commonPath, "NOMARMY_SETUP_PROFILE", "remote");
560
657
  writeEnvLine(commonPath, "NOMARMY_LLAMA_HOST", llamaHost);
561
658
  writeEnvLine(commonPath, "NOMARMY_LLAMA_PORT", llamaPort);
562
- const next = `${path.join(nomarmyRoot, "install.sh")} --profile remote`;
659
+ const next = "nomarmy install";
563
660
  if (json) return out({ written: commonPath, execution: "remote", llamaHost, llamaPort, reachable, next });
564
661
  console.log(reachable
565
662
  ? c.green(`✓ llama-server is reachable at ${healthUrl}.`)
@@ -570,7 +667,7 @@ async function cmdSetup() {
570
667
  return;
571
668
  }
572
669
 
573
- const execution = value("execution", process.env.NOMARMY_EXECUTION || "local");
670
+ const execution = flag("choose") ? "local" : value("execution", process.env.NOMARMY_EXECUTION || "local");
574
671
  const isCloud = execution !== "local";
575
672
  const hardware = isCloud ? null : await detectHardware();
576
673
  const modelPath = isCloud ? null : findModel();
@@ -666,6 +763,9 @@ async function cmdSetup() {
666
763
 
667
764
  fs.mkdirSync(path.dirname(profilePath), { recursive: true });
668
765
  for (const [k, v] of Object.entries(profileWrites)) writeEnvLine(profilePath, k, v);
766
+ writeEnvLine(commonPath, "NOMARMY_EXECUTION", execution);
767
+ if (execution === "local") writeEnvLine(commonPath, "NOMARMY_LLAMA_HOST", "127.0.0.1");
768
+ writeEnvLine(commonPath, "NOMARMY_SETUP_PROFILE", execution === "bedrock" ? "bedrock" : profileName);
669
769
  if (model?.kind === "known" && model.repo) {
670
770
  writeEnvLine(commonPath, "NOMARMY_MODEL_REPO", model.repo);
671
771
  writeEnvLine(commonPath, "NOMARMY_MODEL_QUANT", model.quant);
@@ -686,7 +786,7 @@ async function cmdSetup() {
686
786
  }
687
787
  }
688
788
 
689
- const installCmd = `./install.sh --profile ${profileName}${isCloud ? "" : ""}`;
789
+ const installCmd = "nomarmy install";
690
790
  if (json) return out({ written: { profile: profilePath, common: model?.kind === "known" ? commonPath : null }, env: profileWrites, sizingTier, installCommand: installCmd });
691
791
  console.log(c.green(`\n✓ Wrote ${path.relative(nomarmyRoot, profilePath)}${model?.kind === "known" ? ` and ${path.relative(nomarmyRoot, commonPath)}` : ""}.`));
692
792
  console.log(c.dim("\nThis proposes; it does not install. Run:\n"));
@@ -1205,7 +1305,7 @@ async function cmdAgentsAdd() {
1205
1305
  let kind = argv[2];
1206
1306
  if (!AGENT_KINDS.includes(kind)) {
1207
1307
  console.log("\n" + c.bold("What kind of agent?"));
1208
- console.log(` ${c.cyan("1.")} local ${c.dim("the local model on this machine (free, private)")}`);
1308
+ console.log(` ${c.cyan("1.")} local ${c.dim("the local model on this machine (no per-token bill, private; slower)")}`);
1209
1309
  console.log(` ${c.cyan("2.")} api ${c.dim("a metered API key (xAI, OpenAI, Anthropic, DeepSeek, ...)")}`);
1210
1310
  console.log(` ${c.cyan("3.")} subscription ${c.dim("your own Claude, ChatGPT or Muse Code plan (never shared)")}`);
1211
1311
  kind = AGENT_KINDS[Number((await rl.question(c.bold("Choice: "))).trim()) - 1];
@@ -1428,7 +1528,9 @@ async function cmdAgentsUpdate() {
1428
1528
 
1429
1529
  const changes = {};
1430
1530
  let probe = false;
1431
- if (json) {
1531
+ // Flags alone are enough; without any, it asks field by field.
1532
+ const flagged = ["model", "no-model", "slot", "auth-env", "base-url", "max-concurrent", "context-window", "thinking", "no-thinking", "probe"].some((f) => flag(f));
1533
+ if (json || flagged) {
1432
1534
  const num = (flagName) => (value(flagName) !== null ? Number(value(flagName)) : undefined);
1433
1535
  if (value("model") !== null) changes.model = value("model");
1434
1536
  // Back to no default model: every role (or job) then names its own.
@@ -1446,7 +1548,7 @@ async function cmdAgentsUpdate() {
1446
1548
  probe = flag("probe") && current.kind === "subscription";
1447
1549
  if (!Object.keys(changes).length) throw new Error("Nothing to update -- pass at least one field flag (see `nomarmy agents update` usage).");
1448
1550
  } else {
1449
- if (!process.stdin.isTTY) throw new Error("nomarmy agents update needs an interactive terminal, or --json with explicit flags.");
1551
+ if (!process.stdin.isTTY) throw new Error("nomarmy agents update needs an interactive terminal, or the flags to change (e.g. --max-concurrent 3).");
1450
1552
  const rl = createInterface({ input, output });
1451
1553
  try {
1452
1554
  console.log(c.bold("🍪 nomArmy agents update") + c.dim(` (${name}: ${describeAgentLabel(current)})`));
@@ -1955,7 +2057,8 @@ function armyLayerFlag(fallback = "global") {
1955
2057
  function loadArmyForCli() {
1956
2058
  const agents = loadAgentsOrExit().agents;
1957
2059
  const loaded = loadArmy({ projectDir: repoDir });
1958
- return { loaded, agents, summary: describeArmy(loaded, { agents, describeAgent: describeAgentLabel }) };
2060
+ const usageSnapshots = readUsageSnapshots(process.env.NOMARMY_AGENT_STATE || path.join(os.homedir(), ".local", "share", "nomarmy-local-agents"));
2061
+ return { loaded, agents, summary: describeArmy(loaded, { agents, describeAgent: describeAgentLabel, usageSnapshots, agentProviderId }) };
1959
2062
  }
1960
2063
 
1961
2064
  // Claude Code adds settings.local.json to .gitignore for the same reason:
@@ -1975,6 +2078,12 @@ function agentCell(name, runsOn, role = null) {
1975
2078
  return `${name}${model ? ` ${c.bold(model)}` : ""}${runsOn ? c.dim(` ${runsOn}`) : ""}`;
1976
2079
  }
1977
2080
 
2081
+ /** An agent's usage-limit reading (lib/usage-limits.mjs), colored by level. */
2082
+ function usageLine(usage, indent) {
2083
+ const text = `${indent}usage: ${usage.text}`;
2084
+ return usage.level === "over" ? c.red(`${text} (at the limit)`) : usage.level === "high" ? c.yellow(text) : c.dim(text);
2085
+ }
2086
+
1978
2087
  async function cmdArmyShow() {
1979
2088
  const { summary } = loadArmyForCli();
1980
2089
  if (json) return out(summary);
@@ -1984,6 +2093,7 @@ async function cmdArmyShow() {
1984
2093
  console.log(c.dim(` ${g.who}`));
1985
2094
  for (const line of g.responsibilities) console.log(c.dim(` - ${line}`));
1986
2095
  if (g.problem) console.log(c.yellow(` ⚠ ${g.problem}`));
2096
+ if (g.usage) console.log(usageLine(g.usage, " "));
1987
2097
  if (summary.workflow) console.log(`\n${c.bold("Workflow")}\n${summary.workflow.split("\n").map((l) => ` ${l}`).join("\n")}`);
1988
2098
  const names = Object.keys(summary.roles);
1989
2099
  if (!names.length) {
@@ -2003,6 +2113,7 @@ async function cmdArmyShow() {
2003
2113
  if (role.description) console.log(c.dim(` ${role.description}`));
2004
2114
  if (role.problem && role.agent) console.log(c.red(` ✗ ${role.problem}`));
2005
2115
  if (role.overlapsGeneral) console.log(c.yellow(` ⚠ ${role.overlapsGeneral}`));
2116
+ if (role.usage) console.log(usageLine(role.usage, " "));
2006
2117
  }
2007
2118
  }
2008
2119
  }
@@ -2320,7 +2431,7 @@ async function cmdStatusline() {
2320
2431
  process.stdout.write(`${statusLineText({ session })}\n`);
2321
2432
  }
2322
2433
 
2323
- const commands = { scan: cmdScan, validate: cmdValidate, sizing: cmdSizing, init: cmdInit, setup: cmdSetup, model: cmdModel, agents: cmdAgents, army: cmdArmy, jobs: cmdJobs, statusline: cmdStatusline, health: cmdHealth, config: cmdConfig, update: cmdUpdate, connect: cmdConnect, start: cmdStart, stop: cmdStop, uninstall: cmdUninstall, help: () => usage(0) };
2434
+ const commands = { scan: cmdScan, validate: cmdValidate, sizing: cmdSizing, init: cmdInit, setup: cmdSetup, install: cmdInstall, model: cmdModel, agents: cmdAgents, army: cmdArmy, jobs: cmdJobs, statusline: cmdStatusline, health: cmdHealth, config: cmdConfig, update: cmdUpdate, connect: cmdConnect, start: cmdStart, stop: cmdStop, uninstall: cmdUninstall, help: () => usage(0) };
2324
2435
  // doctor command
2325
2436
  async function cmdDoctor() {
2326
2437
  // Import lazily to avoid circular dependencies
package/install.sh CHANGED
@@ -72,7 +72,7 @@ MSG
72
72
  exit 1
73
73
  ;;
74
74
  *)
75
- echo "ERROR: $OS_NAME is not a supported host. See the README for supported platforms." >&2
75
+ echo "ERROR: $OS_NAME is not a supported host. See docs/install.md for supported platforms." >&2
76
76
  exit 1
77
77
  ;;
78
78
  esac
@@ -127,3 +127,12 @@ if [[ "$(nomarmy_execution_mode)" == hosted ]]; then
127
127
  else
128
128
  echo "==> Install complete. Run: ./e2e.sh --profile $NOMARMY_PROFILE"
129
129
  fi
130
+ # Record only a completed, verified installation.
131
+ node --input-type=module - "$ROOT/package.json" "$NOMARMY_INSTALL_ROOT" "$NOMARMY_PROFILE" "$(nomarmy_execution_mode)" <<'JS'
132
+ import fs from 'node:fs';
133
+ import path from 'node:path';
134
+ const [pkg, root, profile, execution] = process.argv.slice(2);
135
+ const { version } = JSON.parse(fs.readFileSync(pkg, 'utf8'));
136
+ fs.mkdirSync(root, { recursive: true });
137
+ fs.writeFileSync(path.join(root, 'install.json'), JSON.stringify({ profile, execution, installedAt: new Date().toISOString(), version }, null, 2) + '\n');
138
+ JS
@@ -31,18 +31,22 @@ const TEXT_EXT_HINT = /\.(mjs|cjs|js|jsx|ts|tsx|py|go|rs|java|kt|kts|cs|rb|php|s
31
31
  // ---------------------------------------------------------------------------
32
32
  // File walk
33
33
  // ---------------------------------------------------------------------------
34
- function globToRegex(glob) {
35
- // Minimal glob: ** any path, * within a segment, ? one char, {a,b} alternation.
34
+ // Minimal glob: ** any path, * within a segment, ? one char, {a,b}
35
+ // alternation. Everything else is literal, so the regex built here can only
36
+ // ever contain the fragments below, whatever the glob says.
37
+ function globBody(glob) {
36
38
  let re = "";
37
39
  for (let i = 0; i < glob.length; i++) {
38
40
  const ch = glob[i];
39
41
  if (ch === "*") { if (glob[i + 1] === "*") { re += "(?:.*\\/)?"; i++; if (glob[i + 1] === "/") { i++; } } else re += "[^/]*"; }
40
42
  else if (ch === "?") re += "[^/]";
41
- else if (ch === "{") { const j = glob.indexOf("}", i); if (j > i) { re += "(" + glob.slice(i + 1, j).split(",").map(s => s.replace(/[.+^$()|[\]\\]/g, "\\$&")).join("|") + ")"; i = j; } else re += "\\{"; }
42
- else if (/[.+^$()|[\]\\]/.test(ch)) re += "\\" + ch;
43
- else re += ch;
43
+ else if (ch === "{") { const j = glob.indexOf("}", i); if (j > i) { re += "(?:" + glob.slice(i + 1, j).split(",").map(globBody).join("|") + ")"; i = j; } else re += "\\{"; }
44
+ else re += ch.replace(/[.*+?^${}()|[\]\\/]/g, "\\$&");
44
45
  }
45
- return new RegExp("^(?:.*/)?" + re + "$");
46
+ return re;
47
+ }
48
+ function globToRegex(glob) {
49
+ return new RegExp("^(?:.*/)?" + globBody(glob) + "$");
46
50
  }
47
51
  export function matchesGlob(rel, glob) {
48
52
  if (!glob) return true;
@@ -98,6 +102,8 @@ export function grepRepo(root, { pattern, regex = true, ignoreCase = false, whol
98
102
  let src = regex ? pattern : pattern.replace(/[.*+?^${}()|[\]\\]/g, "\\$&");
99
103
  if (wholeWord) src = `\\b(?:${src})\\b`;
100
104
  let re;
105
+ // A regex from the caller is the point of grep (repo_evidence); the
106
+ // caller is the operator's own coordinator session.
101
107
  try { re = new RegExp(src, ignoreCase ? "i" : ""); } catch (e) { throw new Error(`invalid pattern: ${e.message}`); }
102
108
  const { files, truncated: filesTruncated } = listFiles(root, { glob, ignores });
103
109
  const hits = [];
@@ -0,0 +1,55 @@
1
+ // `nomarmy setup`: the setup playbook as a checklist. Each step says whether
2
+ // it's done and which `nomarmy` command does it; the runner shows the list,
3
+ // runs the first unfinished step once the operator says yes, and repeats.
4
+ // All observations are injected functions; importing this module reads no state.
5
+ export function setupSteps(probes) {
6
+ const mode = probes.mode();
7
+ const profile = mode.profile;
8
+ const hosted = profile === "hosted";
9
+ const installation = probes.install(profile);
10
+ const agents = probes.agents().filter((name) => name !== "local");
11
+ const army = probes.army();
12
+ const roles = Object.values(army?.roles ?? {});
13
+ const rolesDone = roles.length > 0 && (!hosted || roles.every((role) => role.agent !== "local"));
14
+ const repo = probes.repo();
15
+ const step = (id, title, status, detail, command) => ({ id, title, status, detail, command });
16
+ const modeDetail = !profile ? "choose where models run" : profile === "remote" ? `remote ${mode.host}:${mode.port}` : ["hosted", "bedrock"].includes(profile) ? profile : `local ${profile}`;
17
+ const installed = installation.marker ? installation.marker.profile === profile : Boolean(installation.version && installation.registered);
18
+ return [
19
+ step("mode", "Where models run", profile ? "done" : "todo", modeDetail, ["setup", "--choose"]),
20
+ step("install", "Installed", installed ? "done" : "todo", installation.version || "OpenClaw not found", ["install"]),
21
+ step("agents", "Agents", agents.length ? "done" : hosted ? "todo" : "skipped", agents.length ? agents.join(", ") : hosted ? "add a hosted agent" : "optional: the local model works without one", ["agents", "add"]),
22
+ step("roles", "Roles", rolesDone ? "done" : "todo", rolesDone ? `${roles.length} role${roles.length === 1 ? "" : "s"}` : !agents.length ? (hosted ? "add an agent first" : "every role on the local model") : hosted && roles.length ? "roles must use a hosted agent" : "define an army", agents.length ? ["army", "init", "--agent", agents[0], "--force", ...(army?.repairLayer ? [`--${army.repairLayer}`] : [])] : ["army", "init", "--agent", "local", "--force"]),
23
+ step("repo", "This repo", !repo.inside ? "skipped" : repo.configured ? "done" : "todo", !repo.inside ? "run nomarmy setup inside a project to set it up" : repo.configured ? ".nomarmy.yml exists" : "configure this project", ["init"]),
24
+ step("check", "Check", "todo", "verify the installation", ["doctor"]),
25
+ ];
26
+ }
27
+
28
+ export function formatSetupSteps(steps) {
29
+ const first = steps.findIndex((step) => step.status === "todo");
30
+ return steps.map((step, i) => `${step.status === "done" ? "✓" : step.status === "skipped" ? "–" : i === first ? "→" : " "} ${step.title}: ${step.detail}`).join("\n");
31
+ }
32
+
33
+ export async function runSetupPlaybook({ evaluate, ask, run, print }) {
34
+ for (;;) {
35
+ const steps = await evaluate();
36
+ print(formatSetupSteps(steps));
37
+ const next = steps.find((step) => step.status === "todo");
38
+ if (!next) return 0;
39
+ const answer = (await ask(`Run \`nomarmy ${next.command.join(" ")}\` now? [Y/n] `)).trim().toLowerCase();
40
+ if (answer && answer !== "y" && answer !== "yes") {
41
+ print("Resume with: nomarmy setup");
42
+ return 0;
43
+ }
44
+ const code = await run(next.command);
45
+ if (code !== 0) {
46
+ print(`Setup step failed: ${next.title} (exit ${code})`);
47
+ return code;
48
+ }
49
+ if (next.id === "check") {
50
+ print("Setup complete");
51
+ print("Restart Claude Code in the project and ask it to use nomArmy.");
52
+ return 0;
53
+ }
54
+ }
55
+ }
package/package.json CHANGED
@@ -3,7 +3,7 @@
3
3
  "description": "A harness for AI coding workers whose claims are never trusted: your coding assistant stays in charge while workers implement and test in sandboxes, on local models, API keys or your own subscriptions.",
4
4
  "author": "Rayson Technologies",
5
5
  "license": "Apache-2.0",
6
- "version": "0.1.0-alpha.3",
6
+ "version": "0.1.0-alpha.4",
7
7
  "private": false,
8
8
  "type": "module",
9
9
  "engines": {
@@ -15,8 +15,8 @@ if [[ -f "$PID" ]] && kill -0 "$(cat "$PID")" 2>/dev/null; then echo "llama-serv
15
15
  BIN="$NOMARMY_INSTALL_ROOT/llama-server"; [[ -x "$BIN" ]] || { echo "ERROR: $BIN missing"; exit 1; }
16
16
  ARGS=(-hf "$NOMARMY_MODEL_REPO:$NOMARMY_MODEL_QUANT" --alias "$NOMARMY_MODEL_ALIAS" --host "$NOMARMY_LLAMA_HOST" --port "$NOMARMY_LLAMA_PORT" -c "$NOMARMY_LLAMA_CONTEXT" -np "$NOMARMY_LLAMA_PARALLEL" -ngl "$NOMARMY_LLAMA_GPU_LAYERS" -t "$NOMARMY_LLAMA_THREADS" --metrics)
17
17
  # Everything below is optional advanced tuning, unset by default: nothing
18
- # here changes behavior for anyone who hasn't set these. See README's
19
- # "Advanced llama-server tuning" section for what each one is for.
18
+ # here changes behavior for anyone who hasn't set these. See "Advanced
19
+ # llama-server tuning" in docs/configuration.md for what each one is for.
20
20
  [[ -n "${NOMARMY_LLAMA_CACHE_TYPE_K:-}" ]] && ARGS+=(--cache-type-k "$NOMARMY_LLAMA_CACHE_TYPE_K")
21
21
  [[ -n "${NOMARMY_LLAMA_CACHE_TYPE_V:-}" ]] && ARGS+=(--cache-type-v "$NOMARMY_LLAMA_CACHE_TYPE_V")
22
22
  [[ -n "${NOMARMY_LLAMA_FLASH_ATTN:-}" ]] && ARGS+=(--flash-attn "$NOMARMY_LLAMA_FLASH_ATTN")