nomarmy 0.1.0-alpha.1 → 0.1.0-alpha.3
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +50 -15
- package/bin/nomarmy.mjs +98 -7
- package/config/common.env +4 -2
- package/config/profiles/hosted.env +6 -0
- package/config/profiles/remote.env +9 -0
- package/e2e.sh +6 -0
- package/install.sh +9 -5
- package/lib/admission.mjs +30 -7
- package/lib/army.mjs +13 -3
- package/lib/budget-state.mjs +5 -2
- package/lib/budget.mjs +4 -1
- package/lib/connect.mjs +4 -0
- package/lib/coordinator-instructions.mjs +1 -0
- package/lib/doctor.mjs +49 -5
- package/lib/execution.mjs +57 -0
- package/lib/health.mjs +24 -6
- package/lib/statusline.mjs +32 -6
- package/lib/usage-limits.mjs +126 -0
- package/mcp/server.mjs +11 -4
- package/package.json +1 -1
- package/scripts/configure-openclaw.sh +40 -4
- package/scripts/install-llama-cpp.sh +8 -1
- package/scripts/lib.sh +30 -3
- package/scripts/start-inference.sh +8 -1
- package/scripts/stop-inference.sh +8 -1
- package/scripts/verify-install.sh +16 -6
package/README.md
CHANGED
|
@@ -13,21 +13,24 @@
|
|
|
13
13
|
## TL;DR
|
|
14
14
|
|
|
15
15
|
1. **Have** Git, Node 20+ and [Podman](https://podman.io) (on macOS: `brew install podman && podman machine init && podman machine start`).
|
|
16
|
-
2. **Install** (
|
|
16
|
+
2. **Install** (OpenClaw and the sandbox, and registers nomArmy with Claude Code):
|
|
17
17
|
```bash
|
|
18
18
|
git clone https://github.com/rayson-tech/nomarmy.git && cd nomarmy
|
|
19
|
-
|
|
20
|
-
|
|
19
|
+
npm install && npm link
|
|
20
|
+
nomarmy setup --hosted
|
|
21
|
+
./install.sh --profile hosted
|
|
21
22
|
```
|
|
22
|
-
3. **
|
|
23
|
-
4. **
|
|
24
|
-
5. **
|
|
23
|
+
3. **Add your workers:** `nomarmy agents add` (an API key, or your ChatGPT or Muse Code subscription), then `nomarmy army init --agent <name>` to put every role on it. `nomarmy doctor` checks the lot.
|
|
24
|
+
4. **Set up your repo:** in the project, run `nomarmy init`. It proposes a `.nomarmy.yml` with your test command.
|
|
25
|
+
5. **Use it:** restart Claude Code in that project and ask it to use nomArmy for one small bug that has a test. When that works, try `/feature <what you want built>`.
|
|
26
|
+
|
|
27
|
+
**Have a GPU or a Mac with plenty of memory?** Workers can also run free on a local model: `./install.sh --profile macbook-pro` (or `nvidia-linux`, `cpu-linux`, `dgx-spark`) builds llama.cpp and starts it; see [Install](#install). A team GPU server works too: [a shared model server](#a-shared-model-server). Codex or Cursor as the coordinator: `nomarmy connect codex cursor`.
|
|
25
28
|
|
|
26
29
|
Stuck? `nomarmy doctor` checks the machine, and `nomarmy health` checks everything nomArmy runs on.
|
|
27
30
|
|
|
28
31
|
## What it is
|
|
29
32
|
|
|
30
|
-
Your coding assistant (Claude Code, Codex or Cursor) stays in charge as the **General**: it decides what gets built and whether the result is acceptable. The work goes to **noms**, workers that implement, test and repair in their own git worktree and sandbox, on
|
|
33
|
+
Your coding assistant (Claude Code, Codex or Cursor) stays in charge as the **General**: it decides what gets built and whether the result is acceptable. The work goes to **noms**, workers that implement, test and repair in their own git worktree and sandbox, on an API key, your own ChatGPT or Muse Code subscription, or a local model. nomArmy owns everything in between: worktrees, git, sandboxes, verification, and the evidence that decides whether work is accepted.
|
|
31
34
|
|
|
32
35
|
**What you get is work you don't have to take on faith**, not cheaper work. Delegating costs the General tokens too: briefing and reviewing. On small, already-diagnosed tickets we measured 4 to 8 times more of the General's tokens than fixing the bug directly, and break-even at roughly 150 lines of context a fix needs to read ([the measurements](docs/experiments/2026-09-20-model-bakeoff-and-economics.md)). It pays off on bigger tickets, on parallel work, and anywhere you'd otherwise have to trust an agent's say-so.
|
|
33
36
|
|
|
@@ -47,19 +50,37 @@ Around that core: **agents** say where a job can run, the **army** says which ro
|
|
|
47
50
|
|
|
48
51
|
## Install
|
|
49
52
|
|
|
50
|
-
|
|
|
53
|
+
| Setup | Guide |
|
|
51
54
|
|---|---|
|
|
52
|
-
|
|
|
53
|
-
|
|
|
55
|
+
| API keys and subscriptions, no local model (most people) | [Hosted workers only](#hosted-workers-only) |
|
|
56
|
+
| A local model on macOS (Apple Silicon) | [macOS](#macos-apple-silicon) |
|
|
57
|
+
| A local model on Linux, with or without an NVIDIA GPU | [Linux](#linux) |
|
|
54
58
|
| Windows | [Windows](#windows) |
|
|
55
59
|
| NVIDIA DGX Spark | [DGX Spark](#dgx-spark) |
|
|
56
|
-
|
|
|
60
|
+
| A shared GPU server (or a tunnel to one) | [A shared model server](#a-shared-model-server) |
|
|
61
|
+
| No GPU, with Bedrock | [Cloud (Bedrock)](#cloud-bedrock) |
|
|
57
62
|
|
|
58
63
|
Every platform needs Git and Podman. `nomarmy doctor` checks the host and prints a fix for anything missing.
|
|
59
64
|
|
|
60
|
-
`install.sh` builds llama.cpp
|
|
65
|
+
`install.sh` builds llama.cpp when you run a local model, installs and configures [OpenClaw](https://github.com/openclaw/openclaw) (the host-side broker every model call goes through), builds the sandbox image, and registers the MCP server if Claude Code is installed. `nomarmy connect` (run by `install.sh`, or by hand for Codex and Cursor) also installs the `/feature` command, Claude Code's status line and, on macOS, nomArmy's notifier. The coordinator gets nomArmy's instructions from the MCP server itself, so there's nothing to copy into your projects.
|
|
66
|
+
|
|
67
|
+
**From npm:** `npm install -g nomarmy@alpha` gives you the `nomarmy` command; `nomarmy setup` then picks a profile and model and prints the `install.sh` command to run (`nomarmy setup --hosted` or `--llama-url <server>` without a local model). Installing from a clone, as in the TL;DR, is the most tested path.
|
|
68
|
+
|
|
69
|
+
### Hosted workers only
|
|
70
|
+
|
|
71
|
+
No GPU and no local model: every job runs on an API key or a subscription (ChatGPT, Muse Code) you add as an agent. Git worktrees, the sandbox and verification still run on your machine, so you still need Git, Node and Podman.
|
|
72
|
+
|
|
73
|
+
```bash
|
|
74
|
+
git clone https://github.com/rayson-tech/nomarmy.git && cd nomarmy
|
|
75
|
+
npm install && npm link # or: npm install -g nomarmy@alpha
|
|
76
|
+
nomarmy setup --hosted # records that this install has no local model
|
|
77
|
+
./install.sh --profile hosted # OpenClaw, the sandbox, and the Claude Code registration; no llama.cpp
|
|
78
|
+
nomarmy agents add # an API key or a subscription login
|
|
79
|
+
nomarmy army init --agent <name> # every role on that agent (add --model <model> to pick one)
|
|
80
|
+
nomarmy doctor
|
|
81
|
+
```
|
|
61
82
|
|
|
62
|
-
|
|
83
|
+
A hosted install refuses a job that names no role or agent, rather than falling back to a local model that isn't there. `nomarmy health` warns about any role still on `local`. `e2e.sh` tests the local model, so it has nothing to do here; `nomarmy army assign` tests each role's route instead.
|
|
63
84
|
|
|
64
85
|
### macOS (Apple Silicon)
|
|
65
86
|
|
|
@@ -103,7 +124,7 @@ CPU-only Linux works but is slow for interactive use: see [Sizing](#sizing).
|
|
|
103
124
|
|
|
104
125
|
1. **WSL2** (the supported path): install a Linux distro under WSL2, install Podman inside it, and follow the [Linux](#linux) guide entirely inside the distro. Watch WSL2's default cap of about half your RAM (`.wslconfig`), and set `git config --global core.longpaths true` (`nomarmy doctor` checks this).
|
|
105
126
|
2. **Native llama.cpp on Windows**, built from source: more RAM, more setup. Prebuilt binaries aren't a safe shortcut; some CPUs crash every backend at startup.
|
|
106
|
-
3. **No local inference**:
|
|
127
|
+
3. **No local inference**: [hosted workers only](#hosted-workers-only), or [Bedrock](#cloud-bedrock).
|
|
107
128
|
|
|
108
129
|
`nomarmy sizing` reports what your hardware can support.
|
|
109
130
|
|
|
@@ -119,6 +140,18 @@ chmod +x install.sh e2e.sh scripts/*.sh
|
|
|
119
140
|
|
|
120
141
|
Moving from a Mac install? Don't copy a Mac binary or model cache over: clone fresh and let `install.sh` build llama.cpp for CUDA on that machine.
|
|
121
142
|
|
|
143
|
+
### A shared model server
|
|
144
|
+
|
|
145
|
+
A team GPU box (a DGX, a workstation) runs one llama-server; everyone else points nomArmy at it. That works through an SSH tunnel too (`ssh -L 8080:localhost:8080 gpu-box`, then `http://127.0.0.1:8080`).
|
|
146
|
+
|
|
147
|
+
```bash
|
|
148
|
+
nomarmy setup --llama-url http://gpu-box:8080 # checks /health, records the address
|
|
149
|
+
./install.sh --profile remote # no llama.cpp build; OpenClaw points at that server
|
|
150
|
+
./e2e.sh --profile remote
|
|
151
|
+
```
|
|
152
|
+
|
|
153
|
+
nomArmy doesn't start, stop or size that server: whoever runs it sets its model, context and slots, and `install.sh` reads the model name and context from the server. Your machine still runs the sandbox and verification for your jobs. Several people sharing one server share its slots, so keep `NOMARMY_MAX_WORKERS` low (the profile sets 1). Run the server itself with any local profile's `install.sh` on the GPU machine, with `NOMARMY_LLAMA_HOST=0.0.0.0` so others can reach it, on a network you trust: llama-server has no authentication.
|
|
154
|
+
|
|
122
155
|
### Cloud (Bedrock)
|
|
123
156
|
|
|
124
157
|
No GPU, no local build:
|
|
@@ -177,6 +210,8 @@ Changes apply to the next job with no restart. The exception is a **new** api ag
|
|
|
177
210
|
|
|
178
211
|
**Your plan decides which models run.** A model can be listed and still refused: on a ChatGPT plan, the Codex route runs gpt-6-astra and the gpt-5.6 models but refuses gpt-6-sol and gpt-6-luna. `army assign` and `agents update --probe` test the exact route a job takes, so they catch this before a job does.
|
|
179
212
|
|
|
213
|
+
**Usage limits.** nomArmy reads how much of a subscription's limit is used where the vendor reports it: Codex in every job's session log, Claude in what Claude Code passes to the status line (Pro and Max plans). It shows in `army`, `local_worker_capacity` and the status line (`⚠ openai 85% wk`), and `nomarmy health` warns from 80%. At the limit, a job on that agent is held rather than left to fail: the General asks you, and resubmits with `confirm_over_limit: true` if you say go. Muse, Grok and API keys don't report their limits; a job that hits one in a `/feature` run pauses that agent for the run.
|
|
214
|
+
|
|
180
215
|
**Vendor terms and platform risk.** Every model call goes through [OpenClaw](https://github.com/openclaw/openclaw), and subscriptions are reached through each vendor's own CLI or login. We've read the terms that apply (see above), but using a personal subscription through a harness is exactly the kind of use vendors tighten, and a change in a vendor's terms or in OpenClaw can stop a subscription agent from working. Local models and API keys don't carry that risk. Plan on subscriptions as a convenience, not the only way your roles can run.
|
|
181
216
|
|
|
182
217
|
**Picking an agent.** Build work goes to a sandboxed agent: `local`, an api key, Codex or Muse. `local` for a bounded change against a written spec with a test; your code never leaves your machine. An api or subscription agent when the work needs more than the local model, knowing it sends code to that vendor. That's a decision about where your source travels, separate from the trust boundary, which is the same for every agent. The General itself when the answer isn't known yet.
|
|
@@ -470,7 +505,7 @@ What we've learned from real runs, including where delegating pays and where it
|
|
|
470
505
|
- **Deploy-time failures need your own check.** See [Add a check for what unit tests can't see](#nomarmyyml).
|
|
471
506
|
- **Node dependencies install only from npm lockfiles**, one per package (no workspaces, yarn, pnpm or bun yet), and only from the public registry: the image build has no credentials for a private one.
|
|
472
507
|
- **Verification needing services** (a database, a mock server) reports `not_run` instead of running without them. The compose parser doesn't resolve YAML anchors.
|
|
473
|
-
- **Same-host
|
|
508
|
+
- **Same-host sandboxes**: the MCP server, OpenClaw and every job's sandbox run on the machine with the coordinator. Only the model can be elsewhere (an agent, or [a shared model server](#a-shared-model-server)).
|
|
474
509
|
|
|
475
510
|
## Troubleshooting
|
|
476
511
|
|
package/bin/nomarmy.mjs
CHANGED
|
@@ -18,10 +18,11 @@ import { buildConfigProposal } from "../lib/propose.mjs";
|
|
|
18
18
|
import { detectHardware } from "../lib/hardware.mjs";
|
|
19
19
|
import { readGGUFMetadata, resolveModelPath, totalSplitBytes } from "../lib/gguf.mjs";
|
|
20
20
|
import { recommend, customRecommendation, evaluateConfig, bytesPerKvElementForCacheTypes, MIN_CONTEXT_PER_NOM } from "../lib/sizing.mjs";
|
|
21
|
-
import { connectClaude, connectCodex, connectCursor, cursorAlreadyConnected } from "../lib/connect.mjs";
|
|
21
|
+
import { connectClaude, connectCodex, connectCursor, cursorAlreadyConnected, deriveWorkerModelEnv } from "../lib/connect.mjs";
|
|
22
22
|
import { ID_RE, AUTH_ENV_NAME_RE, OPENCLAW_PROVIDER_ID_RE, openclawProviderId, isNativeProviderType } from "../lib/dispatch-schema.mjs";
|
|
23
|
-
import { loadAgents, readAgentsFile, writeAgentsFile, agentsConfigPath, apiAgentAsPoolEntry, describeAgent as describeAgentLabel, AGENT_KINDS, API_PROVIDER_TYPES, RESERVED_AGENT_NAMES, BUILTIN_LOCAL_AGENT } from "../lib/agents.mjs";
|
|
23
|
+
import { loadAgents, readAgentsFile, writeAgentsFile, agentsConfigPath, apiAgentAsPoolEntry, describeAgent as describeAgentLabel, agentRunsToolsOnHost, AGENT_KINDS, API_PROVIDER_TYPES, RESERVED_AGENT_NAMES, BUILTIN_LOCAL_AGENT } from "../lib/agents.mjs";
|
|
24
24
|
import { loadArmy, describeArmy, readArmyFile, updateArmyInFile, assignRoleInFile, parseTargetSpec, armyLayerPath, globalConfigDir, DEFAULT_ARMY, ARMY_PHASES, LOCAL_CONFIG_FILENAME } from "../lib/army.mjs";
|
|
25
|
+
import { parseLlamaUrl } from "../lib/execution.mjs";
|
|
25
26
|
import { ensureProviderConfig } from "../lib/openclaw-config.mjs";
|
|
26
27
|
import { recordProbeSuccess } from "../lib/health.mjs";
|
|
27
28
|
import { pruneJobRuntime } from "../lib/prune.mjs";
|
|
@@ -93,6 +94,8 @@ Usage: nomarmy <command> [options]
|
|
|
93
94
|
write config/profiles/<name>.env (+ config/common.env).
|
|
94
95
|
Prints the install.sh command; never runs it.
|
|
95
96
|
--tier <more|nominal> with --json, skip the prompt
|
|
97
|
+
--hosted skip hardware/model questions
|
|
98
|
+
--llama-url <url> use a llama-server on another machine
|
|
96
99
|
model Change the configured model later, without the rest of
|
|
97
100
|
setup's questions. Offers to also resync the MCP
|
|
98
101
|
registration's worker-routing env vars, and to restart
|
|
@@ -168,7 +171,9 @@ Usage: nomarmy <command> [options]
|
|
|
168
171
|
UI/UX, data architect, security analyst, PM,
|
|
169
172
|
PO, stakeholder, all on \`local\`) to --global
|
|
170
173
|
(default), --project or --local; --force
|
|
171
|
-
replaces an existing one
|
|
174
|
+
replaces an existing one; --agent <name>
|
|
175
|
+
starts every role there, optionally with
|
|
176
|
+
--model <model|auto>
|
|
172
177
|
assign <role> <agent|none> [model|auto]
|
|
173
178
|
give a role an agent, and optionally the model
|
|
174
179
|
to run on it ("auto" lets the General pick per
|
|
@@ -516,6 +521,55 @@ async function chooseModel(rl) {
|
|
|
516
521
|
* command itself does.
|
|
517
522
|
*/
|
|
518
523
|
async function cmdSetup() {
|
|
524
|
+
const hosted = flag("hosted");
|
|
525
|
+
const hasLlamaUrl = flag("llama-url");
|
|
526
|
+
if (hosted && hasLlamaUrl) throw new Error("--hosted and --llama-url cannot be used together.");
|
|
527
|
+
|
|
528
|
+
if (hosted || hasLlamaUrl) {
|
|
529
|
+
const commonPath = path.join(nomarmyRoot, "config", "common.env");
|
|
530
|
+
|
|
531
|
+
if (hosted) {
|
|
532
|
+
const next = [
|
|
533
|
+
`${path.join(nomarmyRoot, "install.sh")} --profile hosted`,
|
|
534
|
+
"nomarmy agents add",
|
|
535
|
+
"nomarmy army init --agent <name>",
|
|
536
|
+
];
|
|
537
|
+
fs.mkdirSync(path.dirname(commonPath), { recursive: true });
|
|
538
|
+
writeEnvLine(commonPath, "NOMARMY_EXECUTION", "hosted");
|
|
539
|
+
if (json) return out({ written: commonPath, execution: "hosted", next });
|
|
540
|
+
console.log(c.green(`✓ Wrote NOMARMY_EXECUTION=hosted to ${path.relative(nomarmyRoot, commonPath)}.`));
|
|
541
|
+
console.log(c.dim("\nNext:"));
|
|
542
|
+
for (const step of next) console.log(` ${c.bold(step)}`);
|
|
543
|
+
return;
|
|
544
|
+
}
|
|
545
|
+
|
|
546
|
+
const llamaInput = value("llama-url");
|
|
547
|
+
if (!llamaInput) throw new Error("--llama-url requires an http:// URL, for example http://server:8080.");
|
|
548
|
+
const { host: llamaHost, port: llamaPort } = parseLlamaUrl(llamaInput);
|
|
549
|
+
const urlHost = llamaHost.includes(":") ? `[${llamaHost}]` : llamaHost;
|
|
550
|
+
const healthUrl = `http://${urlHost}:${llamaPort}/health`;
|
|
551
|
+
let reachable = false;
|
|
552
|
+
try {
|
|
553
|
+
await fetch(healthUrl, { signal: AbortSignal.timeout(5000) });
|
|
554
|
+
reachable = true;
|
|
555
|
+
} catch {
|
|
556
|
+
// A server may simply be offline during setup; retain its validated address.
|
|
557
|
+
}
|
|
558
|
+
fs.mkdirSync(path.dirname(commonPath), { recursive: true });
|
|
559
|
+
writeEnvLine(commonPath, "NOMARMY_EXECUTION", "remote");
|
|
560
|
+
writeEnvLine(commonPath, "NOMARMY_LLAMA_HOST", llamaHost);
|
|
561
|
+
writeEnvLine(commonPath, "NOMARMY_LLAMA_PORT", llamaPort);
|
|
562
|
+
const next = `${path.join(nomarmyRoot, "install.sh")} --profile remote`;
|
|
563
|
+
if (json) return out({ written: commonPath, execution: "remote", llamaHost, llamaPort, reachable, next });
|
|
564
|
+
console.log(reachable
|
|
565
|
+
? c.green(`✓ llama-server is reachable at ${healthUrl}.`)
|
|
566
|
+
: c.yellow(`⚠ llama-server is not reachable at ${healthUrl} right now; configuration was still written.`));
|
|
567
|
+
console.log(c.green(`✓ Wrote the remote llama-server settings to ${path.relative(nomarmyRoot, commonPath)}.`));
|
|
568
|
+
console.log(c.dim("\nNext:"));
|
|
569
|
+
console.log(` ${c.bold(next)}`);
|
|
570
|
+
return;
|
|
571
|
+
}
|
|
572
|
+
|
|
519
573
|
const execution = value("execution", process.env.NOMARMY_EXECUTION || "local");
|
|
520
574
|
const isCloud = execution !== "local";
|
|
521
575
|
const hardware = isCloud ? null : await detectHardware();
|
|
@@ -1962,16 +2016,45 @@ async function cmdArmyShow() {
|
|
|
1962
2016
|
async function cmdArmyInit() {
|
|
1963
2017
|
const layer = armyLayerFlag("global");
|
|
1964
2018
|
const filePath = armyLayerPath(layer, { projectDir: repoDir });
|
|
2019
|
+
const agentName = value("agent");
|
|
2020
|
+
let selectedAgent = null;
|
|
2021
|
+
let roleModel = null;
|
|
2022
|
+
if (flag("agent") && !agentName) throw new Error("--agent requires a name from agents.yml.");
|
|
2023
|
+
if (agentName) {
|
|
2024
|
+
const agents = loadAgentsOrExit().agents;
|
|
2025
|
+
if (!Object.prototype.hasOwnProperty.call(agents, agentName)) {
|
|
2026
|
+
throw new Error(`Unknown agent "${agentName}". Pick one of: ${Object.keys(agents).join(", ")}`);
|
|
2027
|
+
}
|
|
2028
|
+
selectedAgent = agents[agentName];
|
|
2029
|
+
if (flag("model") && !value("model")) throw new Error("--model requires a model name or auto.");
|
|
2030
|
+
roleModel = value("model") ?? (selectedAgent.model ? null : "auto");
|
|
2031
|
+
}
|
|
1965
2032
|
const existing = readArmyFile(filePath, { armyOnly: layer !== "project" });
|
|
1966
2033
|
if (existing?.roles && Object.keys(existing.roles).length && !flag("force")) {
|
|
1967
2034
|
throw new Error(`${filePath} already defines an army (${Object.keys(existing.roles).join(", ")}). Re-run with --force to replace it.`);
|
|
1968
2035
|
}
|
|
1969
2036
|
// Keep a General already defined in this layer; the roster is what init resets.
|
|
1970
|
-
|
|
2037
|
+
const roster = structuredClone(DEFAULT_ARMY);
|
|
2038
|
+
if (agentName) {
|
|
2039
|
+
for (const role of Object.values(roster.roles)) {
|
|
2040
|
+
role.agent = agentName;
|
|
2041
|
+
if (roleModel) role.model = roleModel;
|
|
2042
|
+
}
|
|
2043
|
+
}
|
|
2044
|
+
updateArmyInFile(filePath, (army) => ({ ...roster, ...(army.general ? { general: army.general } : {}) }));
|
|
1971
2045
|
if (layer === "local") ensureLocalLayerIgnored();
|
|
1972
2046
|
if (json) return out({ written: filePath, layer, roles: Object.keys(DEFAULT_ARMY.roles) });
|
|
1973
2047
|
console.log(c.green(`✓ Wrote the default army to ${filePath} (${layer}).`));
|
|
1974
|
-
|
|
2048
|
+
if (!agentName) {
|
|
2049
|
+
console.log(c.dim("Every role starts on the local model. Next: `nomarmy army general <agent>` (the agent your coordinator session runs on), then `nomarmy army assign <role> <agent>` for any role you want elsewhere."));
|
|
2050
|
+
return;
|
|
2051
|
+
}
|
|
2052
|
+
console.log(c.dim(roleModel === "auto"
|
|
2053
|
+
? `Every role starts on ${agentName} with model auto; the coordinator picks a model per job.`
|
|
2054
|
+
: `Every role starts on ${agentName}${roleModel ? ` with model ${roleModel}` : ` using its default model ${selectedAgent.model}`}.`));
|
|
2055
|
+
if (agentRunsToolsOnHost(selectedAgent)) {
|
|
2056
|
+
console.log(c.yellow(`⚠ ${agentName} runs its tools on the host. Build roles on it will be refused unless allow_host_tools is set in agents.yml.`));
|
|
2057
|
+
}
|
|
1975
2058
|
}
|
|
1976
2059
|
|
|
1977
2060
|
async function cmdArmyAssign() {
|
|
@@ -2206,10 +2289,18 @@ async function cmdJobs() {
|
|
|
2206
2289
|
|
|
2207
2290
|
// `nomarmy health`: run the checks now (the MCP server also runs them every
|
|
2208
2291
|
// 6 hours) and record them, which also refreshes the status line's warning.
|
|
2292
|
+
// This install's settings from config/common.env (the execution mode and the
|
|
2293
|
+
// model server's address, as `nomarmy connect` gives the MCP server), under
|
|
2294
|
+
// anything set (non-empty) in the environment.
|
|
2295
|
+
function installEnv() {
|
|
2296
|
+
const set = Object.fromEntries(Object.entries(process.env).filter(([, v]) => v !== ""));
|
|
2297
|
+
return { ...deriveWorkerModelEnv(nomarmyRoot), ...set };
|
|
2298
|
+
}
|
|
2299
|
+
|
|
2209
2300
|
async function cmdHealth() {
|
|
2210
2301
|
const { checkAndRecordHealth } = await import("../lib/health.mjs");
|
|
2211
2302
|
const stateRoot = process.env.NOMARMY_AGENT_STATE || path.join(os.homedir(), ".local", "share", "nomarmy-local-agents");
|
|
2212
|
-
const { result } = await checkAndRecordHealth({ projectDir: repoDir, stateRoot, configDir: globalConfigDir() });
|
|
2303
|
+
const { result } = await checkAndRecordHealth({ projectDir: repoDir, stateRoot, configDir: globalConfigDir(), env: installEnv() });
|
|
2213
2304
|
if (json) return out(result);
|
|
2214
2305
|
console.log(c.bold("🍪 nomArmy health") + c.dim(` ${new Date(result.checkedAt).toLocaleString()}`));
|
|
2215
2306
|
if (!result.issues.length) { console.log(c.green("\n✓ Nothing to fix.")); return; }
|
|
@@ -2234,7 +2325,7 @@ const commands = { scan: cmdScan, validate: cmdValidate, sizing: cmdSizing, init
|
|
|
2234
2325
|
async function cmdDoctor() {
|
|
2235
2326
|
// Import lazily to avoid circular dependencies
|
|
2236
2327
|
const { runDoctor } = await import("../lib/doctor.mjs");
|
|
2237
|
-
await runDoctor({ json, exit: true });
|
|
2328
|
+
await runDoctor({ json, exit: true, env: installEnv() });
|
|
2238
2329
|
}
|
|
2239
2330
|
commands.doctor = cmdDoctor;
|
|
2240
2331
|
if (!command || flag("help") || !commands[command]) usage(command && !commands[command] ? 2 : 0);
|
package/config/common.env
CHANGED
|
@@ -16,8 +16,10 @@ NOMARMY_MAX_WORKERS=1
|
|
|
16
16
|
NOMARMY_AGENT_IMAGE=openclaw-nomarmy-coder:bookworm
|
|
17
17
|
NOMARMY_INSTALL_ROOT=$HOME/.local/share/nomarmy-local-agents
|
|
18
18
|
|
|
19
|
-
#
|
|
20
|
-
#
|
|
19
|
+
# Where models run: 'local' runs llama-server on this machine, 'remote' uses
|
|
20
|
+
# one nomArmy doesn't run (nomarmy setup --llama-url), 'hosted' has no local
|
|
21
|
+
# model at all (nomarmy setup --hosted), 'bedrock' calls Bedrock. Profiles
|
|
22
|
+
# override this.
|
|
21
23
|
NOMARMY_EXECUTION=local
|
|
22
24
|
|
|
23
25
|
# Worker model routing. The MCP server composes "<provider>/<model>" for
|
|
@@ -0,0 +1,6 @@
|
|
|
1
|
+
# Hosted workers only: no local model on this machine. Every job runs on an
|
|
2
|
+
# api or subscription agent (nomarmy agents add); the sandbox, git worktrees
|
|
3
|
+
# and verification still run here. `nomarmy setup --hosted` writes this
|
|
4
|
+
# mode into config/common.env as well, so the MCP server knows it.
|
|
5
|
+
NOMARMY_PROFILE=hosted
|
|
6
|
+
NOMARMY_EXECUTION=hosted
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# A llama-server nomArmy doesn't run, e.g. a team's GPU server or an SSH
|
|
2
|
+
# tunnel to one. Set its address with `nomarmy setup --llama-url
|
|
3
|
+
# http://host:8080`, which writes NOMARMY_LLAMA_HOST / NOMARMY_LLAMA_PORT
|
|
4
|
+
# into config/common.env. nomArmy
|
|
5
|
+
# doesn't start, stop or size that server; the sandbox, git worktrees and
|
|
6
|
+
# verification still run on this machine.
|
|
7
|
+
NOMARMY_PROFILE=remote
|
|
8
|
+
NOMARMY_EXECUTION=remote
|
|
9
|
+
NOMARMY_MAX_WORKERS=1
|
package/e2e.sh
CHANGED
|
@@ -24,6 +24,12 @@ load_profile "$PROFILE"
|
|
|
24
24
|
|
|
25
25
|
echo "=== nomArmy E2E: $NOMARMY_PROFILE ==="
|
|
26
26
|
|
|
27
|
+
if [[ "$(nomarmy_execution_mode)" == hosted ]]; then
|
|
28
|
+
echo "e2e.sh runs a job on the local model, and profile '$NOMARMY_PROFILE' has none."
|
|
29
|
+
echo "Check each agent with: nomarmy doctor"
|
|
30
|
+
exit 0
|
|
31
|
+
fi
|
|
32
|
+
|
|
27
33
|
"$ROOT/scripts/start-inference.sh" "$NOMARMY_PROFILE"
|
|
28
34
|
|
|
29
35
|
if nomarmy_is_cloud; then
|
package/install.sh
CHANGED
|
@@ -15,7 +15,7 @@ case "$OS_NAME" in
|
|
|
15
15
|
command -v curl >/dev/null 2>&1 || brew install curl
|
|
16
16
|
command -v git >/dev/null 2>&1 || brew install git
|
|
17
17
|
command -v node >/dev/null 2>&1 || brew install node
|
|
18
|
-
if
|
|
18
|
+
if nomarmy_manages_model_server; then
|
|
19
19
|
command -v xcode-select >/dev/null 2>&1 && xcode-select -p >/dev/null 2>&1 || { echo 'ERROR: Xcode Command Line Tools are required. Run: xcode-select --install'; exit 1; }
|
|
20
20
|
command -v cmake >/dev/null 2>&1 || brew install cmake
|
|
21
21
|
fi
|
|
@@ -24,8 +24,8 @@ case "$OS_NAME" in
|
|
|
24
24
|
command -v podman >/dev/null 2>&1 || brew install podman
|
|
25
25
|
;;
|
|
26
26
|
Linux)
|
|
27
|
-
if
|
|
28
|
-
# No
|
|
27
|
+
if ! nomarmy_manages_model_server; then
|
|
28
|
+
# No model is built or served here, so the C++ toolchain is not needed.
|
|
29
29
|
need curl; need git; need node; need npm
|
|
30
30
|
else
|
|
31
31
|
if ! command -v cmake >/dev/null 2>&1 || ! command -v c++ >/dev/null 2>&1 || ! command -v curl >/dev/null 2>&1 || ! command -v git >/dev/null 2>&1 || ! command -v node >/dev/null 2>&1 || ! command -v npm >/dev/null 2>&1; then install_build_dependencies; fi
|
|
@@ -105,7 +105,7 @@ if ! command -v openclaw >/dev/null 2>&1; then
|
|
|
105
105
|
export PATH="$HOME/.local/bin:$HOME/.npm-global/bin:$PATH"
|
|
106
106
|
fi
|
|
107
107
|
need openclaw
|
|
108
|
-
if
|
|
108
|
+
if nomarmy_has_local_model; then openclaw plugins install @openclaw/llama-cpp-provider || true; fi
|
|
109
109
|
"$ROOT/scripts/start-inference.sh" "$NOMARMY_PROFILE"
|
|
110
110
|
# Sandbox before provider config: configure-openclaw.sh refuses to store a real
|
|
111
111
|
# Bedrock credential unless the coder sandbox is already network-isolated.
|
|
@@ -122,4 +122,8 @@ if nomarmy_is_cloud && [[ "${NOMARMY_ORCHESTRATOR_RUNTIME:-}" == "claude-code" ]
|
|
|
122
122
|
echo " ./scripts/configure-orchestrator.sh $NOMARMY_PROFILE # print the settings"
|
|
123
123
|
echo " ./scripts/configure-orchestrator.sh $NOMARMY_PROFILE --apply # write them to Claude Code"
|
|
124
124
|
fi
|
|
125
|
-
|
|
125
|
+
if [[ "$(nomarmy_execution_mode)" == hosted ]]; then
|
|
126
|
+
echo "==> Install complete. Add an agent (nomarmy agents add), give roles to it (nomarmy army init --agent <name>), then run: nomarmy doctor"
|
|
127
|
+
else
|
|
128
|
+
echo "==> Install complete. Run: ./e2e.sh --profile $NOMARMY_PROFILE"
|
|
129
|
+
fi
|
package/lib/admission.mjs
CHANGED
|
@@ -1,5 +1,6 @@
|
|
|
1
1
|
import fs from "node:fs";
|
|
2
2
|
import path from "node:path";
|
|
3
|
+
import { executionMode } from "./execution.mjs";
|
|
3
4
|
import { clampInt } from "./budget-state.mjs";
|
|
4
5
|
import { checkBrief, assessAdmission, describeBudgets } from "./budget.mjs";
|
|
5
6
|
import { parseStatusPorcelainZ, isRuntimeJunk } from "./git-record.mjs";
|
|
@@ -7,6 +8,7 @@ import { readJson } from "./openclaw-run.mjs";
|
|
|
7
8
|
import { readOpenClawTranscriptTail } from "./transcript.mjs";
|
|
8
9
|
import { readClaudeSessionTranscript } from "./claude-transcript.mjs";
|
|
9
10
|
import { notify } from "./notify.mjs";
|
|
11
|
+
import { readCodexRateLimits, recordUsageSnapshot, readUsageSnapshots, usageStatus } from "./usage-limits.mjs";
|
|
10
12
|
import { recentModelRefusal } from "./health.mjs";
|
|
11
13
|
import { writeLease, removeLease, liveLeases, liveSlots, acquireSlot } from "./slots.mjs";
|
|
12
14
|
import { loadRun, runTotals, runAdmissionProblems, recordRunJob, detectUsageLimit } from "./runs.mjs";
|
|
@@ -139,6 +141,11 @@ export function createJobRuntime(deps) {
|
|
|
139
141
|
* watching hears about it from any coordinator without polling.
|
|
140
142
|
*/
|
|
141
143
|
function notifyJobFinished(entry, result, error) {
|
|
144
|
+
try {
|
|
145
|
+
const provider = entry.agent ? agentProviderId(agentsConfig().agents[entry.agent]) : null;
|
|
146
|
+
const snapshot = provider ? readCodexRateLimits(path.join(jobsRoot, entry.jobId)) : null;
|
|
147
|
+
if (snapshot) recordUsageSnapshot(stateRoot, provider, snapshot);
|
|
148
|
+
} catch { /* Usage telemetry must never affect the job result. */ }
|
|
142
149
|
if (!entry.lane) return; // only tracked jobs, never internal helpers
|
|
143
150
|
const m = result?.manifest ?? {};
|
|
144
151
|
const outcome = error ? "failed" : String(m.outcome ?? (result?.ok ? "done" : "finished")).toLowerCase().replace(/_/g, " ");
|
|
@@ -147,8 +154,11 @@ export function createJobRuntime(deps) {
|
|
|
147
154
|
const ok = !error && (result?.ok || m.coordinatorStatus === "complete");
|
|
148
155
|
notify(`nomArmy: ${entry.role ?? entry.mode ?? "job"} ${ok ? "done" : outcome}`, `${entry.workerId ?? entry.jobId} on ${who}: ${outcome} after ${took}m. ${ok ? "Ready for the General's review." : "Needs a look."}`);
|
|
149
156
|
}
|
|
157
|
+
function admissionHardware() {
|
|
158
|
+
return executionMode(deps.env).managesModelServer ? deps.budgetState.hardwareSnapshot : null;
|
|
159
|
+
}
|
|
150
160
|
function capacitySnapshot() {
|
|
151
|
-
const admission = assessAdmission({ hardware:
|
|
161
|
+
const admission = assessAdmission({ hardware: admissionHardware(), runningJobs: runningCount("local"), slots: deps.budgetState.contextInfo.slots, maxWorkers: currentMaxWorkers() });
|
|
152
162
|
return {
|
|
153
163
|
// The local model's budget. An api or subscription job's scales with
|
|
154
164
|
// its own model; local_worker_start reports that job's.
|
|
@@ -157,6 +167,7 @@ export function createJobRuntime(deps) {
|
|
|
157
167
|
admission,
|
|
158
168
|
memory: deps.budgetState.hardwareSnapshot?.memory ?? null,
|
|
159
169
|
running: [...activeJobs.values()].filter(j => !j.settled).map(j => ({ jobId: j.jobId, workerId: j.workerId, mode: j.mode, lane: j.lane, startedAt: j.startedAt, phase: readJson(path.join(jobsRoot, j.jobId, "status.json"))?.phase ?? "starting" })),
|
|
170
|
+
usageLimits: Object.fromEntries(Object.entries(readUsageSnapshots(stateRoot)).map(([provider, snapshot]) => [provider, usageStatus(snapshot)])),
|
|
160
171
|
maxWorkers: currentMaxWorkers(),
|
|
161
172
|
remote: { running: runningCount("remote"), maxWorkers: currentMaxPoolWorkers(), note: "api and subscription agents; each agent's own max_concurrent also applies" }
|
|
162
173
|
};
|
|
@@ -165,6 +176,16 @@ export function createJobRuntime(deps) {
|
|
|
165
176
|
await deps.budgetState.refresh();
|
|
166
177
|
if (jobs.some((j) => jobLane(j) === "remote")) await modelCatalogReady();
|
|
167
178
|
const problems = [];
|
|
179
|
+
const snapshots = readUsageSnapshots(stateRoot);
|
|
180
|
+
jobs.forEach((j, i) => {
|
|
181
|
+
if (!j.agentName || j.confirm_over_limit === true) return;
|
|
182
|
+
let provider;
|
|
183
|
+
try { provider = agentProviderId(agentsConfig().agents[j.agentName]); } catch { return; }
|
|
184
|
+
const snapshot = snapshots[provider];
|
|
185
|
+
if (!snapshot) return;
|
|
186
|
+
const status = usageStatus(snapshot);
|
|
187
|
+
if (status.level === "over") problems.push(`${jobs.length > 1 ? `job ${i + 1}: ` : ""}agent "${j.agentName}" is held at its usage limit: ${status.text} (reading ${status.ageMinutes} minutes old). Ask the operator before resubmitting with confirm_over_limit: true, or send the job to another agent.`);
|
|
188
|
+
});
|
|
168
189
|
// A pool-routed job is checked against that pool's OWN (model-dependent)
|
|
169
190
|
// budget, not the local-derived global one -- see budgetsForPool. Which
|
|
170
191
|
// specific entry pickProvider will land on isn't known yet at admission
|
|
@@ -241,14 +262,16 @@ export function createJobRuntime(deps) {
|
|
|
241
262
|
} catch (error) { problems.push(jobs.length > 1 ? `job ${i + 1}: ${error.message}` : error.message); }
|
|
242
263
|
});
|
|
243
264
|
// Slot capacity only concerns local jobs: a remote job's inference runs
|
|
244
|
-
// at its vendor and never competes for llama-server's slots.
|
|
245
|
-
//
|
|
246
|
-
//
|
|
247
|
-
//
|
|
265
|
+
// at its vendor and never competes for llama-server's slots. When this
|
|
266
|
+
// install runs llama-server itself, free memory applies to every job
|
|
267
|
+
// (the model and each job's sandbox share it), so a remote-only batch is
|
|
268
|
+
// checked for memory alone. A remote or hosted install has no model
|
|
269
|
+
// here to size memory against. Remote jobs have their own, additive
|
|
270
|
+
// ceiling (currentMaxPoolWorkers).
|
|
248
271
|
const anyLocal = jobs.some((j) => jobLane(j) === "local");
|
|
249
272
|
const admission = anyLocal
|
|
250
|
-
? assessAdmission({ hardware:
|
|
251
|
-
: assessAdmission({ hardware:
|
|
273
|
+
? assessAdmission({ hardware: admissionHardware(), runningJobs: runningCount("local"), slots: deps.budgetState.contextInfo.slots, maxWorkers: currentMaxWorkers() })
|
|
274
|
+
: assessAdmission({ hardware: admissionHardware(), runningJobs: 0, slots: null, maxWorkers: Infinity });
|
|
252
275
|
if (!admission.admit) problems.push(...admission.reasons.map(r => `not admitted (${admission.level}): ${r}`));
|
|
253
276
|
if (jobs.some((j) => jobLane(j) === "remote")) {
|
|
254
277
|
const remoteCeiling = currentMaxPoolWorkers(), runningRemote = runningCount("remote");
|
package/lib/army.mjs
CHANGED
|
@@ -30,6 +30,7 @@ import path from "node:path";
|
|
|
30
30
|
import YAML from "yaml";
|
|
31
31
|
import { z } from "zod";
|
|
32
32
|
import { agentRunsToolsOnHost } from "./dispatch-schema.mjs";
|
|
33
|
+
import { usageStatus } from "./usage-limits.mjs";
|
|
33
34
|
|
|
34
35
|
export const GLOBAL_CONFIG_FILENAME = "config.yml";
|
|
35
36
|
export const PROJECT_CONFIG_FILENAMES = Object.freeze([".nomarmy.yml", ".nomarmy.yaml"]);
|
|
@@ -361,11 +362,20 @@ export function generalOverlap(army, agents = {}) {
|
|
|
361
362
|
* description, phase, suggested mode, agent, any problem or overlap with
|
|
362
363
|
* the General, and which layer set each field.
|
|
363
364
|
*/
|
|
364
|
-
export function describeArmy(loaded, { agents = {}, describeAgent = null } = {}) {
|
|
365
|
+
export function describeArmy(loaded, { agents = {}, describeAgent = null, usageSnapshots = null, agentProviderId = null, now = Date.now() } = {}) {
|
|
365
366
|
const army = loaded.army;
|
|
366
367
|
const problems = armyTargetProblems(army, agents);
|
|
367
368
|
const overlap = generalOverlap(army, agents);
|
|
368
369
|
const runsOn = (name) => (name && agents[name] && describeAgent ? describeAgent(agents[name]) : null);
|
|
370
|
+
const usage = (name) => {
|
|
371
|
+
if (!name || !agents[name] || !usageSnapshots || !agentProviderId) return null;
|
|
372
|
+
let provider;
|
|
373
|
+
try { provider = agentProviderId(agents[name]); } catch { return null; }
|
|
374
|
+
const snapshot = usageSnapshots[provider];
|
|
375
|
+
if (!snapshot) return null;
|
|
376
|
+
const { level, text } = usageStatus(snapshot, now);
|
|
377
|
+
return { level, text };
|
|
378
|
+
};
|
|
369
379
|
// A subscription whose owner isn't the General's own is worth a word,
|
|
370
380
|
// not a refusal: it may be the same person's other account (a real Senti
|
|
371
381
|
// General skipped a role's agent for exactly this, unsure whose it was).
|
|
@@ -378,7 +388,7 @@ export function describeArmy(loaded, { agents = {}, describeAgent = null } = {})
|
|
|
378
388
|
const roles = Object.fromEntries(Object.entries(army.roles).map(([name, role]) => [name, {
|
|
379
389
|
description: role.description ?? null, phase: role.phase ?? null, mode: role.mode ?? null,
|
|
380
390
|
agent: role.agent ?? null, model: role.model ?? agents[role.agent]?.model ?? null, modelIsAuto: role.model === "auto",
|
|
381
|
-
agentRunsOn: runsOn(role.agent),
|
|
391
|
+
agentRunsOn: runsOn(role.agent), usage: usage(role.agent),
|
|
382
392
|
// Its agent's own tools run on this machine (dispatch-schema.mjs): an
|
|
383
393
|
// implement role there is refused per job unless the agent allows it.
|
|
384
394
|
hostTools: agentRunsToolsOnHost(agents[role.agent]) ? { allowed: agents[role.agent].allow_host_tools === true, implementRole: (role.mode ?? "implement") === "implement" } : null,
|
|
@@ -388,7 +398,7 @@ export function describeArmy(loaded, { agents = {}, describeAgent = null } = {})
|
|
|
388
398
|
? "not defined -- run `nomarmy army general <agent>` so nomArmy can flag roles that share the General's model or usage"
|
|
389
399
|
: !Object.prototype.hasOwnProperty.call(agents, army.general) ? `agent "${army.general}" is not defined in your agents.yml` : null;
|
|
390
400
|
return {
|
|
391
|
-
general: { ...GENERAL, agent: army.general, agentRunsOn: runsOn(army.general), problem: generalProblem, setBy: loaded.sources.general },
|
|
401
|
+
general: { ...GENERAL, agent: army.general, agentRunsOn: runsOn(army.general), usage: usage(army.general), problem: generalProblem, setBy: loaded.sources.general },
|
|
392
402
|
workflow: army.workflow,
|
|
393
403
|
runLimits: army.runLimits ?? {},
|
|
394
404
|
roles,
|
package/lib/budget-state.mjs
CHANGED
|
@@ -1,4 +1,5 @@
|
|
|
1
1
|
import { deriveBudgets, resolveContextPerNom } from "./budget.mjs";
|
|
2
|
+
import { executionMode } from "./execution.mjs";
|
|
2
3
|
|
|
3
4
|
export function clampInt(value, min, max, fallback) {
|
|
4
5
|
const n = Number.parseInt(value ?? "", 10);
|
|
@@ -12,8 +13,10 @@ export function createBudgetState({ env = process.env } = {}) {
|
|
|
12
13
|
function currentBudgets() { return budgets; }
|
|
13
14
|
async function refresh() {
|
|
14
15
|
try {
|
|
15
|
-
|
|
16
|
-
|
|
16
|
+
if (executionMode(env).hasLocalModel) {
|
|
17
|
+
contextInfo = await resolveContextPerNom({ env });
|
|
18
|
+
budgets = deriveBudgets({ contextPerNom: contextInfo.contextPerNom, source: contextInfo.source, env });
|
|
19
|
+
}
|
|
17
20
|
} catch { /* keep the previous budgets; a failed probe is not a reason to refuse work */ }
|
|
18
21
|
try {
|
|
19
22
|
const { detectHardware } = await import("./hardware.mjs");
|
package/lib/budget.mjs
CHANGED
|
@@ -16,6 +16,7 @@
|
|
|
16
16
|
// Everything here is pure except `resolveContextPerNom`, which may ask a
|
|
17
17
|
// running llama-server what its slots actually are; the probe is injectable.
|
|
18
18
|
import { DEFAULT_TARGET_CONTEXT_PER_NOM, MIN_CONTEXT_PER_NOM, RESERVES, GIB, formatBytes, isCloudExecution } from "./sizing.mjs";
|
|
19
|
+
import { executionMode } from "./execution.mjs";
|
|
19
20
|
|
|
20
21
|
/**
|
|
21
22
|
* Hard ceilings the MCP tool schema enforces regardless of hardware. The
|
|
@@ -309,7 +310,9 @@ async function defaultProbe(env) {
|
|
|
309
310
|
export async function resolveContextPerNom({ env = process.env, probe = defaultProbe } = {}) {
|
|
310
311
|
const direct = envInt(env, "NOMARMY_CONTEXT_PER_NOM");
|
|
311
312
|
if (direct) return { contextPerNom: direct, slots: envInt(env, "NOMARMY_LLAMA_PARALLEL"), source: "profile (NOMARMY_CONTEXT_PER_NOM)" };
|
|
312
|
-
|
|
313
|
+
// A remote server's context is its own, not this machine's settings.
|
|
314
|
+
const remote = executionMode(env).mode === "remote";
|
|
315
|
+
const total = remote ? null : envInt(env, "NOMARMY_LLAMA_CONTEXT"), parallel = remote ? null : envInt(env, "NOMARMY_LLAMA_PARALLEL");
|
|
313
316
|
if (total && parallel) return { contextPerNom: Math.floor(total / parallel), slots: parallel, source: "profile (NOMARMY_LLAMA_CONTEXT / NOMARMY_LLAMA_PARALLEL)" };
|
|
314
317
|
if (!isCloudExecution(env.NOMARMY_EXECUTION || "local")) {
|
|
315
318
|
const probed = await probe(env);
|
package/lib/connect.mjs
CHANGED
|
@@ -240,6 +240,10 @@ export function deriveWorkerModelEnv(nomarmyRoot) {
|
|
|
240
240
|
const env = {};
|
|
241
241
|
if (model) env.NOMARMY_WORKER_MODEL = model;
|
|
242
242
|
if (thinking !== null) env.NOMARMY_WORKER_MODEL_THINKING = thinking;
|
|
243
|
+
for (const key of ["NOMARMY_EXECUTION", "NOMARMY_LLAMA_HOST", "NOMARMY_LLAMA_PORT"]) {
|
|
244
|
+
const value = readEnvValue(commonPath, key);
|
|
245
|
+
if (value !== null) env[key] = value;
|
|
246
|
+
}
|
|
243
247
|
return env;
|
|
244
248
|
}
|
|
245
249
|
|
|
@@ -11,6 +11,7 @@ export const COORDINATOR_INSTRUCTIONS = `nomArmy runs bounded engineering jobs (
|
|
|
11
11
|
Before dispatching:
|
|
12
12
|
- Answer where-is / who-calls / grep questions with repo_evidence (deterministic, [path:line] on every hit). Use mode: scout only for read-only research that would otherwise pull many files into your own context.
|
|
13
13
|
- Call army to see this repo's roles and which agent each runs on; dispatch by army_role when a role fits. A job on a subscription agent needs on_behalf_of set to that agent's owner.
|
|
14
|
+
- Agents' usage limits show in army and local_worker_capacity; a job on an agent at its limit is held. Ask the operator before resubmitting with confirm_over_limit: true, or move the job to another agent. Never set it on your own.
|
|
14
15
|
- A Claude subscription agent (claude-cli) runs its tools on this machine, outside the sandbox: use it for scout and review work. nomArmy refuses implement jobs on it unless the operator set allow_host_tools; send build work to a sandboxed agent.
|
|
15
16
|
- Brief outcomes, not edits: a task, explicit acceptance criteria, and the tests that prove it. Put facts you've already resolved in evidence.
|
|
16
17
|
- Prefer local_worker_start + local_worker_status for anything longer than a few minutes. For a whole feature, use /feature (run_start keeps a run's jobs, spend and hours bounded).
|