nomarmy 0.1.0-alpha.1 → 0.1.0-alpha.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +35 -4
- package/bin/nomarmy.mjs +98 -7
- package/config/common.env +4 -2
- package/config/profiles/hosted.env +6 -0
- package/config/profiles/remote.env +9 -0
- package/e2e.sh +6 -0
- package/install.sh +9 -5
- package/lib/admission.mjs +13 -7
- package/lib/budget-state.mjs +5 -2
- package/lib/budget.mjs +4 -1
- package/lib/connect.mjs +4 -0
- package/lib/doctor.mjs +49 -5
- package/lib/execution.mjs +57 -0
- package/lib/health.mjs +11 -5
- package/mcp/server.mjs +4 -2
- package/package.json +1 -1
- package/scripts/configure-openclaw.sh +40 -4
- package/scripts/install-llama-cpp.sh +8 -1
- package/scripts/lib.sh +30 -3
- package/scripts/start-inference.sh +8 -1
- package/scripts/stop-inference.sh +8 -1
- package/scripts/verify-install.sh +16 -6
package/README.md
CHANGED
|
@@ -19,6 +19,7 @@
|
|
|
19
19
|
./install.sh --profile macbook-pro # or nvidia-linux, cpu-linux, dgx-spark, bedrock
|
|
20
20
|
./e2e.sh --profile macbook-pro # should end with: === E2E PASS ===
|
|
21
21
|
```
|
|
22
|
+
No local model? See [hosted workers only](#hosted-workers-only) or [a shared model server](#a-shared-model-server).
|
|
22
23
|
3. **Set up your repo:** in the project, run `nomarmy init`. It proposes a `.nomarmy.yml` with your test command.
|
|
23
24
|
4. **Use it:** restart Claude Code in that project and ask it to use nomArmy for one small bug that has a test. When that works, try `/feature <what you want built>`.
|
|
24
25
|
5. **Optional:** add hosted workers with `nomarmy agents add`, give roles to them with `nomarmy army init`, or connect Codex or Cursor with `nomarmy connect codex cursor`.
|
|
@@ -53,13 +54,15 @@ Around that core: **agents** say where a job can run, the **army** says which ro
|
|
|
53
54
|
| Linux, with or without an NVIDIA GPU | [Linux](#linux) |
|
|
54
55
|
| Windows | [Windows](#windows) |
|
|
55
56
|
| NVIDIA DGX Spark | [DGX Spark](#dgx-spark) |
|
|
56
|
-
| No
|
|
57
|
+
| No local model: API keys and subscriptions only | [Hosted workers only](#hosted-workers-only) |
|
|
58
|
+
| A shared GPU server (or a tunnel to one) | [A shared model server](#a-shared-model-server) |
|
|
59
|
+
| No GPU, with Bedrock | [Cloud (Bedrock)](#cloud-bedrock) |
|
|
57
60
|
|
|
58
61
|
Every platform needs Git and Podman. `nomarmy doctor` checks the host and prints a fix for anything missing.
|
|
59
62
|
|
|
60
63
|
`install.sh` builds llama.cpp for local inference, installs and configures [OpenClaw](https://github.com/openclaw/openclaw) (the host-side broker every model call goes through), builds the sandbox image, and registers the MCP server if Claude Code is installed. `nomarmy connect` (run by `install.sh`, or by hand for Codex and Cursor) also installs the `/feature` command, Claude Code's status line and, on macOS, nomArmy's notifier. The coordinator gets nomArmy's instructions from the MCP server itself, so there's nothing to copy into your projects.
|
|
61
64
|
|
|
62
|
-
**From npm:** `npm install -g nomarmy@alpha` gives you the `nomarmy` command; `nomarmy setup` then picks a profile and model and prints the `install.sh` command to run. Installing from a clone, as in the TL;DR, is the most tested path.
|
|
65
|
+
**From npm:** `npm install -g nomarmy@alpha` gives you the `nomarmy` command; `nomarmy setup` then picks a profile and model and prints the `install.sh` command to run (`nomarmy setup --hosted` or `--llama-url <server>` without a local model). Installing from a clone, as in the TL;DR, is the most tested path.
|
|
63
66
|
|
|
64
67
|
### macOS (Apple Silicon)
|
|
65
68
|
|
|
@@ -103,7 +106,7 @@ CPU-only Linux works but is slow for interactive use: see [Sizing](#sizing).
|
|
|
103
106
|
|
|
104
107
|
1. **WSL2** (the supported path): install a Linux distro under WSL2, install Podman inside it, and follow the [Linux](#linux) guide entirely inside the distro. Watch WSL2's default cap of about half your RAM (`.wslconfig`), and set `git config --global core.longpaths true` (`nomarmy doctor` checks this).
|
|
105
108
|
2. **Native llama.cpp on Windows**, built from source: more RAM, more setup. Prebuilt binaries aren't a safe shortcut; some CPUs crash every backend at startup.
|
|
106
|
-
3. **No local inference**:
|
|
109
|
+
3. **No local inference**: [hosted workers only](#hosted-workers-only), or [Bedrock](#cloud-bedrock).
|
|
107
110
|
|
|
108
111
|
`nomarmy sizing` reports what your hardware can support.
|
|
109
112
|
|
|
@@ -119,6 +122,34 @@ chmod +x install.sh e2e.sh scripts/*.sh
|
|
|
119
122
|
|
|
120
123
|
Moving from a Mac install? Don't copy a Mac binary or model cache over: clone fresh and let `install.sh` build llama.cpp for CUDA on that machine.
|
|
121
124
|
|
|
125
|
+
### Hosted workers only
|
|
126
|
+
|
|
127
|
+
No GPU and no local model: every job runs on an API key or a subscription (ChatGPT, Muse Code) you add as an agent. Git worktrees, the sandbox and verification still run on your machine, so you still need Git, Node and Podman.
|
|
128
|
+
|
|
129
|
+
```bash
|
|
130
|
+
git clone https://github.com/rayson-tech/nomarmy.git && cd nomarmy
|
|
131
|
+
npm install && npm link # or: npm install -g nomarmy@alpha
|
|
132
|
+
nomarmy setup --hosted # records that this install has no local model
|
|
133
|
+
./install.sh --profile hosted # OpenClaw, the sandbox, and the Claude Code registration; no llama.cpp
|
|
134
|
+
nomarmy agents add # an API key or a subscription login
|
|
135
|
+
nomarmy army init --agent <name> # every role on that agent (add --model <model> to pick one)
|
|
136
|
+
nomarmy doctor
|
|
137
|
+
```
|
|
138
|
+
|
|
139
|
+
A hosted install refuses a job that names no role or agent, rather than falling back to a local model that isn't there. `nomarmy health` warns about any role still on `local`. `e2e.sh` tests the local model, so it has nothing to do here; `nomarmy army assign` tests each role's route instead.
|
|
140
|
+
|
|
141
|
+
### A shared model server
|
|
142
|
+
|
|
143
|
+
A team GPU box (a DGX, a workstation) runs one llama-server; everyone else points nomArmy at it. That works through an SSH tunnel too (`ssh -L 8080:localhost:8080 gpu-box`, then `http://127.0.0.1:8080`).
|
|
144
|
+
|
|
145
|
+
```bash
|
|
146
|
+
nomarmy setup --llama-url http://gpu-box:8080 # checks /health, records the address
|
|
147
|
+
./install.sh --profile remote # no llama.cpp build; OpenClaw points at that server
|
|
148
|
+
./e2e.sh --profile remote
|
|
149
|
+
```
|
|
150
|
+
|
|
151
|
+
nomArmy doesn't start, stop or size that server: whoever runs it sets its model, context and slots, and `install.sh` reads the model name and context from the server. Your machine still runs the sandbox and verification for your jobs. Several people sharing one server share its slots, so keep `NOMARMY_MAX_WORKERS` low (the profile sets 1). Run the server itself with any local profile's `install.sh` on the GPU machine, with `NOMARMY_LLAMA_HOST=0.0.0.0` so others can reach it, on a network you trust: llama-server has no authentication.
|
|
152
|
+
|
|
122
153
|
### Cloud (Bedrock)
|
|
123
154
|
|
|
124
155
|
No GPU, no local build:
|
|
@@ -470,7 +501,7 @@ What we've learned from real runs, including where delegating pays and where it
|
|
|
470
501
|
- **Deploy-time failures need your own check.** See [Add a check for what unit tests can't see](#nomarmyyml).
|
|
471
502
|
- **Node dependencies install only from npm lockfiles**, one per package (no workspaces, yarn, pnpm or bun yet), and only from the public registry: the image build has no credentials for a private one.
|
|
472
503
|
- **Verification needing services** (a database, a mock server) reports `not_run` instead of running without them. The compose parser doesn't resolve YAML anchors.
|
|
473
|
-
- **Same-host
|
|
504
|
+
- **Same-host sandboxes**: the MCP server, OpenClaw and every job's sandbox run on the machine with the coordinator. Only the model can be elsewhere (an agent, or [a shared model server](#a-shared-model-server)).
|
|
474
505
|
|
|
475
506
|
## Troubleshooting
|
|
476
507
|
|
package/bin/nomarmy.mjs
CHANGED
|
@@ -18,10 +18,11 @@ import { buildConfigProposal } from "../lib/propose.mjs";
|
|
|
18
18
|
import { detectHardware } from "../lib/hardware.mjs";
|
|
19
19
|
import { readGGUFMetadata, resolveModelPath, totalSplitBytes } from "../lib/gguf.mjs";
|
|
20
20
|
import { recommend, customRecommendation, evaluateConfig, bytesPerKvElementForCacheTypes, MIN_CONTEXT_PER_NOM } from "../lib/sizing.mjs";
|
|
21
|
-
import { connectClaude, connectCodex, connectCursor, cursorAlreadyConnected } from "../lib/connect.mjs";
|
|
21
|
+
import { connectClaude, connectCodex, connectCursor, cursorAlreadyConnected, deriveWorkerModelEnv } from "../lib/connect.mjs";
|
|
22
22
|
import { ID_RE, AUTH_ENV_NAME_RE, OPENCLAW_PROVIDER_ID_RE, openclawProviderId, isNativeProviderType } from "../lib/dispatch-schema.mjs";
|
|
23
|
-
import { loadAgents, readAgentsFile, writeAgentsFile, agentsConfigPath, apiAgentAsPoolEntry, describeAgent as describeAgentLabel, AGENT_KINDS, API_PROVIDER_TYPES, RESERVED_AGENT_NAMES, BUILTIN_LOCAL_AGENT } from "../lib/agents.mjs";
|
|
23
|
+
import { loadAgents, readAgentsFile, writeAgentsFile, agentsConfigPath, apiAgentAsPoolEntry, describeAgent as describeAgentLabel, agentRunsToolsOnHost, AGENT_KINDS, API_PROVIDER_TYPES, RESERVED_AGENT_NAMES, BUILTIN_LOCAL_AGENT } from "../lib/agents.mjs";
|
|
24
24
|
import { loadArmy, describeArmy, readArmyFile, updateArmyInFile, assignRoleInFile, parseTargetSpec, armyLayerPath, globalConfigDir, DEFAULT_ARMY, ARMY_PHASES, LOCAL_CONFIG_FILENAME } from "../lib/army.mjs";
|
|
25
|
+
import { parseLlamaUrl } from "../lib/execution.mjs";
|
|
25
26
|
import { ensureProviderConfig } from "../lib/openclaw-config.mjs";
|
|
26
27
|
import { recordProbeSuccess } from "../lib/health.mjs";
|
|
27
28
|
import { pruneJobRuntime } from "../lib/prune.mjs";
|
|
@@ -93,6 +94,8 @@ Usage: nomarmy <command> [options]
|
|
|
93
94
|
write config/profiles/<name>.env (+ config/common.env).
|
|
94
95
|
Prints the install.sh command; never runs it.
|
|
95
96
|
--tier <more|nominal> with --json, skip the prompt
|
|
97
|
+
--hosted skip hardware/model questions
|
|
98
|
+
--llama-url <url> use a llama-server on another machine
|
|
96
99
|
model Change the configured model later, without the rest of
|
|
97
100
|
setup's questions. Offers to also resync the MCP
|
|
98
101
|
registration's worker-routing env vars, and to restart
|
|
@@ -168,7 +171,9 @@ Usage: nomarmy <command> [options]
|
|
|
168
171
|
UI/UX, data architect, security analyst, PM,
|
|
169
172
|
PO, stakeholder, all on \`local\`) to --global
|
|
170
173
|
(default), --project or --local; --force
|
|
171
|
-
replaces an existing one
|
|
174
|
+
replaces an existing one; --agent <name>
|
|
175
|
+
starts every role there, optionally with
|
|
176
|
+
--model <model|auto>
|
|
172
177
|
assign <role> <agent|none> [model|auto]
|
|
173
178
|
give a role an agent, and optionally the model
|
|
174
179
|
to run on it ("auto" lets the General pick per
|
|
@@ -516,6 +521,55 @@ async function chooseModel(rl) {
|
|
|
516
521
|
* command itself does.
|
|
517
522
|
*/
|
|
518
523
|
async function cmdSetup() {
|
|
524
|
+
const hosted = flag("hosted");
|
|
525
|
+
const hasLlamaUrl = flag("llama-url");
|
|
526
|
+
if (hosted && hasLlamaUrl) throw new Error("--hosted and --llama-url cannot be used together.");
|
|
527
|
+
|
|
528
|
+
if (hosted || hasLlamaUrl) {
|
|
529
|
+
const commonPath = path.join(nomarmyRoot, "config", "common.env");
|
|
530
|
+
|
|
531
|
+
if (hosted) {
|
|
532
|
+
const next = [
|
|
533
|
+
`${path.join(nomarmyRoot, "install.sh")} --profile hosted`,
|
|
534
|
+
"nomarmy agents add",
|
|
535
|
+
"nomarmy army init --agent <name>",
|
|
536
|
+
];
|
|
537
|
+
fs.mkdirSync(path.dirname(commonPath), { recursive: true });
|
|
538
|
+
writeEnvLine(commonPath, "NOMARMY_EXECUTION", "hosted");
|
|
539
|
+
if (json) return out({ written: commonPath, execution: "hosted", next });
|
|
540
|
+
console.log(c.green(`✓ Wrote NOMARMY_EXECUTION=hosted to ${path.relative(nomarmyRoot, commonPath)}.`));
|
|
541
|
+
console.log(c.dim("\nNext:"));
|
|
542
|
+
for (const step of next) console.log(` ${c.bold(step)}`);
|
|
543
|
+
return;
|
|
544
|
+
}
|
|
545
|
+
|
|
546
|
+
const llamaInput = value("llama-url");
|
|
547
|
+
if (!llamaInput) throw new Error("--llama-url requires an http:// URL, for example http://server:8080.");
|
|
548
|
+
const { host: llamaHost, port: llamaPort } = parseLlamaUrl(llamaInput);
|
|
549
|
+
const urlHost = llamaHost.includes(":") ? `[${llamaHost}]` : llamaHost;
|
|
550
|
+
const healthUrl = `http://${urlHost}:${llamaPort}/health`;
|
|
551
|
+
let reachable = false;
|
|
552
|
+
try {
|
|
553
|
+
await fetch(healthUrl, { signal: AbortSignal.timeout(5000) });
|
|
554
|
+
reachable = true;
|
|
555
|
+
} catch {
|
|
556
|
+
// A server may simply be offline during setup; retain its validated address.
|
|
557
|
+
}
|
|
558
|
+
fs.mkdirSync(path.dirname(commonPath), { recursive: true });
|
|
559
|
+
writeEnvLine(commonPath, "NOMARMY_EXECUTION", "remote");
|
|
560
|
+
writeEnvLine(commonPath, "NOMARMY_LLAMA_HOST", llamaHost);
|
|
561
|
+
writeEnvLine(commonPath, "NOMARMY_LLAMA_PORT", llamaPort);
|
|
562
|
+
const next = `${path.join(nomarmyRoot, "install.sh")} --profile remote`;
|
|
563
|
+
if (json) return out({ written: commonPath, execution: "remote", llamaHost, llamaPort, reachable, next });
|
|
564
|
+
console.log(reachable
|
|
565
|
+
? c.green(`✓ llama-server is reachable at ${healthUrl}.`)
|
|
566
|
+
: c.yellow(`⚠ llama-server is not reachable at ${healthUrl} right now; configuration was still written.`));
|
|
567
|
+
console.log(c.green(`✓ Wrote the remote llama-server settings to ${path.relative(nomarmyRoot, commonPath)}.`));
|
|
568
|
+
console.log(c.dim("\nNext:"));
|
|
569
|
+
console.log(` ${c.bold(next)}`);
|
|
570
|
+
return;
|
|
571
|
+
}
|
|
572
|
+
|
|
519
573
|
const execution = value("execution", process.env.NOMARMY_EXECUTION || "local");
|
|
520
574
|
const isCloud = execution !== "local";
|
|
521
575
|
const hardware = isCloud ? null : await detectHardware();
|
|
@@ -1962,16 +2016,45 @@ async function cmdArmyShow() {
|
|
|
1962
2016
|
async function cmdArmyInit() {
|
|
1963
2017
|
const layer = armyLayerFlag("global");
|
|
1964
2018
|
const filePath = armyLayerPath(layer, { projectDir: repoDir });
|
|
2019
|
+
const agentName = value("agent");
|
|
2020
|
+
let selectedAgent = null;
|
|
2021
|
+
let roleModel = null;
|
|
2022
|
+
if (flag("agent") && !agentName) throw new Error("--agent requires a name from agents.yml.");
|
|
2023
|
+
if (agentName) {
|
|
2024
|
+
const agents = loadAgentsOrExit().agents;
|
|
2025
|
+
if (!Object.prototype.hasOwnProperty.call(agents, agentName)) {
|
|
2026
|
+
throw new Error(`Unknown agent "${agentName}". Pick one of: ${Object.keys(agents).join(", ")}`);
|
|
2027
|
+
}
|
|
2028
|
+
selectedAgent = agents[agentName];
|
|
2029
|
+
if (flag("model") && !value("model")) throw new Error("--model requires a model name or auto.");
|
|
2030
|
+
roleModel = value("model") ?? (selectedAgent.model ? null : "auto");
|
|
2031
|
+
}
|
|
1965
2032
|
const existing = readArmyFile(filePath, { armyOnly: layer !== "project" });
|
|
1966
2033
|
if (existing?.roles && Object.keys(existing.roles).length && !flag("force")) {
|
|
1967
2034
|
throw new Error(`${filePath} already defines an army (${Object.keys(existing.roles).join(", ")}). Re-run with --force to replace it.`);
|
|
1968
2035
|
}
|
|
1969
2036
|
// Keep a General already defined in this layer; the roster is what init resets.
|
|
1970
|
-
|
|
2037
|
+
const roster = structuredClone(DEFAULT_ARMY);
|
|
2038
|
+
if (agentName) {
|
|
2039
|
+
for (const role of Object.values(roster.roles)) {
|
|
2040
|
+
role.agent = agentName;
|
|
2041
|
+
if (roleModel) role.model = roleModel;
|
|
2042
|
+
}
|
|
2043
|
+
}
|
|
2044
|
+
updateArmyInFile(filePath, (army) => ({ ...roster, ...(army.general ? { general: army.general } : {}) }));
|
|
1971
2045
|
if (layer === "local") ensureLocalLayerIgnored();
|
|
1972
2046
|
if (json) return out({ written: filePath, layer, roles: Object.keys(DEFAULT_ARMY.roles) });
|
|
1973
2047
|
console.log(c.green(`✓ Wrote the default army to ${filePath} (${layer}).`));
|
|
1974
|
-
|
|
2048
|
+
if (!agentName) {
|
|
2049
|
+
console.log(c.dim("Every role starts on the local model. Next: `nomarmy army general <agent>` (the agent your coordinator session runs on), then `nomarmy army assign <role> <agent>` for any role you want elsewhere."));
|
|
2050
|
+
return;
|
|
2051
|
+
}
|
|
2052
|
+
console.log(c.dim(roleModel === "auto"
|
|
2053
|
+
? `Every role starts on ${agentName} with model auto; the coordinator picks a model per job.`
|
|
2054
|
+
: `Every role starts on ${agentName}${roleModel ? ` with model ${roleModel}` : ` using its default model ${selectedAgent.model}`}.`));
|
|
2055
|
+
if (agentRunsToolsOnHost(selectedAgent)) {
|
|
2056
|
+
console.log(c.yellow(`⚠ ${agentName} runs its tools on the host. Build roles on it will be refused unless allow_host_tools is set in agents.yml.`));
|
|
2057
|
+
}
|
|
1975
2058
|
}
|
|
1976
2059
|
|
|
1977
2060
|
async function cmdArmyAssign() {
|
|
@@ -2206,10 +2289,18 @@ async function cmdJobs() {
|
|
|
2206
2289
|
|
|
2207
2290
|
// `nomarmy health`: run the checks now (the MCP server also runs them every
|
|
2208
2291
|
// 6 hours) and record them, which also refreshes the status line's warning.
|
|
2292
|
+
// This install's settings from config/common.env (the execution mode and the
|
|
2293
|
+
// model server's address, as `nomarmy connect` gives the MCP server), under
|
|
2294
|
+
// anything set (non-empty) in the environment.
|
|
2295
|
+
function installEnv() {
|
|
2296
|
+
const set = Object.fromEntries(Object.entries(process.env).filter(([, v]) => v !== ""));
|
|
2297
|
+
return { ...deriveWorkerModelEnv(nomarmyRoot), ...set };
|
|
2298
|
+
}
|
|
2299
|
+
|
|
2209
2300
|
async function cmdHealth() {
|
|
2210
2301
|
const { checkAndRecordHealth } = await import("../lib/health.mjs");
|
|
2211
2302
|
const stateRoot = process.env.NOMARMY_AGENT_STATE || path.join(os.homedir(), ".local", "share", "nomarmy-local-agents");
|
|
2212
|
-
const { result } = await checkAndRecordHealth({ projectDir: repoDir, stateRoot, configDir: globalConfigDir() });
|
|
2303
|
+
const { result } = await checkAndRecordHealth({ projectDir: repoDir, stateRoot, configDir: globalConfigDir(), env: installEnv() });
|
|
2213
2304
|
if (json) return out(result);
|
|
2214
2305
|
console.log(c.bold("🍪 nomArmy health") + c.dim(` ${new Date(result.checkedAt).toLocaleString()}`));
|
|
2215
2306
|
if (!result.issues.length) { console.log(c.green("\n✓ Nothing to fix.")); return; }
|
|
@@ -2234,7 +2325,7 @@ const commands = { scan: cmdScan, validate: cmdValidate, sizing: cmdSizing, init
|
|
|
2234
2325
|
async function cmdDoctor() {
|
|
2235
2326
|
// Import lazily to avoid circular dependencies
|
|
2236
2327
|
const { runDoctor } = await import("../lib/doctor.mjs");
|
|
2237
|
-
await runDoctor({ json, exit: true });
|
|
2328
|
+
await runDoctor({ json, exit: true, env: installEnv() });
|
|
2238
2329
|
}
|
|
2239
2330
|
commands.doctor = cmdDoctor;
|
|
2240
2331
|
if (!command || flag("help") || !commands[command]) usage(command && !commands[command] ? 2 : 0);
|
package/config/common.env
CHANGED
|
@@ -16,8 +16,10 @@ NOMARMY_MAX_WORKERS=1
|
|
|
16
16
|
NOMARMY_AGENT_IMAGE=openclaw-nomarmy-coder:bookworm
|
|
17
17
|
NOMARMY_INSTALL_ROOT=$HOME/.local/share/nomarmy-local-agents
|
|
18
18
|
|
|
19
|
-
#
|
|
20
|
-
#
|
|
19
|
+
# Where models run: 'local' runs llama-server on this machine, 'remote' uses
|
|
20
|
+
# one nomArmy doesn't run (nomarmy setup --llama-url), 'hosted' has no local
|
|
21
|
+
# model at all (nomarmy setup --hosted), 'bedrock' calls Bedrock. Profiles
|
|
22
|
+
# override this.
|
|
21
23
|
NOMARMY_EXECUTION=local
|
|
22
24
|
|
|
23
25
|
# Worker model routing. The MCP server composes "<provider>/<model>" for
|
|
@@ -0,0 +1,6 @@
|
|
|
1
|
+
# Hosted workers only: no local model on this machine. Every job runs on an
|
|
2
|
+
# api or subscription agent (nomarmy agents add); the sandbox, git worktrees
|
|
3
|
+
# and verification still run here. `nomarmy setup --hosted` writes this
|
|
4
|
+
# mode into config/common.env as well, so the MCP server knows it.
|
|
5
|
+
NOMARMY_PROFILE=hosted
|
|
6
|
+
NOMARMY_EXECUTION=hosted
|
|
@@ -0,0 +1,9 @@
|
|
|
1
|
+
# A llama-server nomArmy doesn't run, e.g. a team's GPU server or an SSH
|
|
2
|
+
# tunnel to one. Set its address with `nomarmy setup --llama-url
|
|
3
|
+
# http://host:8080`, which writes NOMARMY_LLAMA_HOST / NOMARMY_LLAMA_PORT
|
|
4
|
+
# into config/common.env. nomArmy
|
|
5
|
+
# doesn't start, stop or size that server; the sandbox, git worktrees and
|
|
6
|
+
# verification still run on this machine.
|
|
7
|
+
NOMARMY_PROFILE=remote
|
|
8
|
+
NOMARMY_EXECUTION=remote
|
|
9
|
+
NOMARMY_MAX_WORKERS=1
|
package/e2e.sh
CHANGED
|
@@ -24,6 +24,12 @@ load_profile "$PROFILE"
|
|
|
24
24
|
|
|
25
25
|
echo "=== nomArmy E2E: $NOMARMY_PROFILE ==="
|
|
26
26
|
|
|
27
|
+
if [[ "$(nomarmy_execution_mode)" == hosted ]]; then
|
|
28
|
+
echo "e2e.sh runs a job on the local model, and profile '$NOMARMY_PROFILE' has none."
|
|
29
|
+
echo "Check each agent with: nomarmy doctor"
|
|
30
|
+
exit 0
|
|
31
|
+
fi
|
|
32
|
+
|
|
27
33
|
"$ROOT/scripts/start-inference.sh" "$NOMARMY_PROFILE"
|
|
28
34
|
|
|
29
35
|
if nomarmy_is_cloud; then
|
package/install.sh
CHANGED
|
@@ -15,7 +15,7 @@ case "$OS_NAME" in
|
|
|
15
15
|
command -v curl >/dev/null 2>&1 || brew install curl
|
|
16
16
|
command -v git >/dev/null 2>&1 || brew install git
|
|
17
17
|
command -v node >/dev/null 2>&1 || brew install node
|
|
18
|
-
if
|
|
18
|
+
if nomarmy_manages_model_server; then
|
|
19
19
|
command -v xcode-select >/dev/null 2>&1 && xcode-select -p >/dev/null 2>&1 || { echo 'ERROR: Xcode Command Line Tools are required. Run: xcode-select --install'; exit 1; }
|
|
20
20
|
command -v cmake >/dev/null 2>&1 || brew install cmake
|
|
21
21
|
fi
|
|
@@ -24,8 +24,8 @@ case "$OS_NAME" in
|
|
|
24
24
|
command -v podman >/dev/null 2>&1 || brew install podman
|
|
25
25
|
;;
|
|
26
26
|
Linux)
|
|
27
|
-
if
|
|
28
|
-
# No
|
|
27
|
+
if ! nomarmy_manages_model_server; then
|
|
28
|
+
# No model is built or served here, so the C++ toolchain is not needed.
|
|
29
29
|
need curl; need git; need node; need npm
|
|
30
30
|
else
|
|
31
31
|
if ! command -v cmake >/dev/null 2>&1 || ! command -v c++ >/dev/null 2>&1 || ! command -v curl >/dev/null 2>&1 || ! command -v git >/dev/null 2>&1 || ! command -v node >/dev/null 2>&1 || ! command -v npm >/dev/null 2>&1; then install_build_dependencies; fi
|
|
@@ -105,7 +105,7 @@ if ! command -v openclaw >/dev/null 2>&1; then
|
|
|
105
105
|
export PATH="$HOME/.local/bin:$HOME/.npm-global/bin:$PATH"
|
|
106
106
|
fi
|
|
107
107
|
need openclaw
|
|
108
|
-
if
|
|
108
|
+
if nomarmy_has_local_model; then openclaw plugins install @openclaw/llama-cpp-provider || true; fi
|
|
109
109
|
"$ROOT/scripts/start-inference.sh" "$NOMARMY_PROFILE"
|
|
110
110
|
# Sandbox before provider config: configure-openclaw.sh refuses to store a real
|
|
111
111
|
# Bedrock credential unless the coder sandbox is already network-isolated.
|
|
@@ -122,4 +122,8 @@ if nomarmy_is_cloud && [[ "${NOMARMY_ORCHESTRATOR_RUNTIME:-}" == "claude-code" ]
|
|
|
122
122
|
echo " ./scripts/configure-orchestrator.sh $NOMARMY_PROFILE # print the settings"
|
|
123
123
|
echo " ./scripts/configure-orchestrator.sh $NOMARMY_PROFILE --apply # write them to Claude Code"
|
|
124
124
|
fi
|
|
125
|
-
|
|
125
|
+
if [[ "$(nomarmy_execution_mode)" == hosted ]]; then
|
|
126
|
+
echo "==> Install complete. Add an agent (nomarmy agents add), give roles to it (nomarmy army init --agent <name>), then run: nomarmy doctor"
|
|
127
|
+
else
|
|
128
|
+
echo "==> Install complete. Run: ./e2e.sh --profile $NOMARMY_PROFILE"
|
|
129
|
+
fi
|
package/lib/admission.mjs
CHANGED
|
@@ -1,5 +1,6 @@
|
|
|
1
1
|
import fs from "node:fs";
|
|
2
2
|
import path from "node:path";
|
|
3
|
+
import { executionMode } from "./execution.mjs";
|
|
3
4
|
import { clampInt } from "./budget-state.mjs";
|
|
4
5
|
import { checkBrief, assessAdmission, describeBudgets } from "./budget.mjs";
|
|
5
6
|
import { parseStatusPorcelainZ, isRuntimeJunk } from "./git-record.mjs";
|
|
@@ -147,8 +148,11 @@ export function createJobRuntime(deps) {
|
|
|
147
148
|
const ok = !error && (result?.ok || m.coordinatorStatus === "complete");
|
|
148
149
|
notify(`nomArmy: ${entry.role ?? entry.mode ?? "job"} ${ok ? "done" : outcome}`, `${entry.workerId ?? entry.jobId} on ${who}: ${outcome} after ${took}m. ${ok ? "Ready for the General's review." : "Needs a look."}`);
|
|
149
150
|
}
|
|
151
|
+
function admissionHardware() {
|
|
152
|
+
return executionMode(deps.env).managesModelServer ? deps.budgetState.hardwareSnapshot : null;
|
|
153
|
+
}
|
|
150
154
|
function capacitySnapshot() {
|
|
151
|
-
const admission = assessAdmission({ hardware:
|
|
155
|
+
const admission = assessAdmission({ hardware: admissionHardware(), runningJobs: runningCount("local"), slots: deps.budgetState.contextInfo.slots, maxWorkers: currentMaxWorkers() });
|
|
152
156
|
return {
|
|
153
157
|
// The local model's budget. An api or subscription job's scales with
|
|
154
158
|
// its own model; local_worker_start reports that job's.
|
|
@@ -241,14 +245,16 @@ export function createJobRuntime(deps) {
|
|
|
241
245
|
} catch (error) { problems.push(jobs.length > 1 ? `job ${i + 1}: ${error.message}` : error.message); }
|
|
242
246
|
});
|
|
243
247
|
// Slot capacity only concerns local jobs: a remote job's inference runs
|
|
244
|
-
// at its vendor and never competes for llama-server's slots.
|
|
245
|
-
//
|
|
246
|
-
//
|
|
247
|
-
//
|
|
248
|
+
// at its vendor and never competes for llama-server's slots. When this
|
|
249
|
+
// install runs llama-server itself, free memory applies to every job
|
|
250
|
+
// (the model and each job's sandbox share it), so a remote-only batch is
|
|
251
|
+
// checked for memory alone. A remote or hosted install has no model
|
|
252
|
+
// here to size memory against. Remote jobs have their own, additive
|
|
253
|
+
// ceiling (currentMaxPoolWorkers).
|
|
248
254
|
const anyLocal = jobs.some((j) => jobLane(j) === "local");
|
|
249
255
|
const admission = anyLocal
|
|
250
|
-
? assessAdmission({ hardware:
|
|
251
|
-
: assessAdmission({ hardware:
|
|
256
|
+
? assessAdmission({ hardware: admissionHardware(), runningJobs: runningCount("local"), slots: deps.budgetState.contextInfo.slots, maxWorkers: currentMaxWorkers() })
|
|
257
|
+
: assessAdmission({ hardware: admissionHardware(), runningJobs: 0, slots: null, maxWorkers: Infinity });
|
|
252
258
|
if (!admission.admit) problems.push(...admission.reasons.map(r => `not admitted (${admission.level}): ${r}`));
|
|
253
259
|
if (jobs.some((j) => jobLane(j) === "remote")) {
|
|
254
260
|
const remoteCeiling = currentMaxPoolWorkers(), runningRemote = runningCount("remote");
|
package/lib/budget-state.mjs
CHANGED
|
@@ -1,4 +1,5 @@
|
|
|
1
1
|
import { deriveBudgets, resolveContextPerNom } from "./budget.mjs";
|
|
2
|
+
import { executionMode } from "./execution.mjs";
|
|
2
3
|
|
|
3
4
|
export function clampInt(value, min, max, fallback) {
|
|
4
5
|
const n = Number.parseInt(value ?? "", 10);
|
|
@@ -12,8 +13,10 @@ export function createBudgetState({ env = process.env } = {}) {
|
|
|
12
13
|
function currentBudgets() { return budgets; }
|
|
13
14
|
async function refresh() {
|
|
14
15
|
try {
|
|
15
|
-
|
|
16
|
-
|
|
16
|
+
if (executionMode(env).hasLocalModel) {
|
|
17
|
+
contextInfo = await resolveContextPerNom({ env });
|
|
18
|
+
budgets = deriveBudgets({ contextPerNom: contextInfo.contextPerNom, source: contextInfo.source, env });
|
|
19
|
+
}
|
|
17
20
|
} catch { /* keep the previous budgets; a failed probe is not a reason to refuse work */ }
|
|
18
21
|
try {
|
|
19
22
|
const { detectHardware } = await import("./hardware.mjs");
|
package/lib/budget.mjs
CHANGED
|
@@ -16,6 +16,7 @@
|
|
|
16
16
|
// Everything here is pure except `resolveContextPerNom`, which may ask a
|
|
17
17
|
// running llama-server what its slots actually are; the probe is injectable.
|
|
18
18
|
import { DEFAULT_TARGET_CONTEXT_PER_NOM, MIN_CONTEXT_PER_NOM, RESERVES, GIB, formatBytes, isCloudExecution } from "./sizing.mjs";
|
|
19
|
+
import { executionMode } from "./execution.mjs";
|
|
19
20
|
|
|
20
21
|
/**
|
|
21
22
|
* Hard ceilings the MCP tool schema enforces regardless of hardware. The
|
|
@@ -309,7 +310,9 @@ async function defaultProbe(env) {
|
|
|
309
310
|
export async function resolveContextPerNom({ env = process.env, probe = defaultProbe } = {}) {
|
|
310
311
|
const direct = envInt(env, "NOMARMY_CONTEXT_PER_NOM");
|
|
311
312
|
if (direct) return { contextPerNom: direct, slots: envInt(env, "NOMARMY_LLAMA_PARALLEL"), source: "profile (NOMARMY_CONTEXT_PER_NOM)" };
|
|
312
|
-
|
|
313
|
+
// A remote server's context is its own, not this machine's settings.
|
|
314
|
+
const remote = executionMode(env).mode === "remote";
|
|
315
|
+
const total = remote ? null : envInt(env, "NOMARMY_LLAMA_CONTEXT"), parallel = remote ? null : envInt(env, "NOMARMY_LLAMA_PARALLEL");
|
|
313
316
|
if (total && parallel) return { contextPerNom: Math.floor(total / parallel), slots: parallel, source: "profile (NOMARMY_LLAMA_CONTEXT / NOMARMY_LLAMA_PARALLEL)" };
|
|
314
317
|
if (!isCloudExecution(env.NOMARMY_EXECUTION || "local")) {
|
|
315
318
|
const probed = await probe(env);
|
package/lib/connect.mjs
CHANGED
|
@@ -240,6 +240,10 @@ export function deriveWorkerModelEnv(nomarmyRoot) {
|
|
|
240
240
|
const env = {};
|
|
241
241
|
if (model) env.NOMARMY_WORKER_MODEL = model;
|
|
242
242
|
if (thinking !== null) env.NOMARMY_WORKER_MODEL_THINKING = thinking;
|
|
243
|
+
for (const key of ["NOMARMY_EXECUTION", "NOMARMY_LLAMA_HOST", "NOMARMY_LLAMA_PORT"]) {
|
|
244
|
+
const value = readEnvValue(commonPath, key);
|
|
245
|
+
if (value !== null) env[key] = value;
|
|
246
|
+
}
|
|
243
247
|
return env;
|
|
244
248
|
}
|
|
245
249
|
|
package/lib/doctor.mjs
CHANGED
|
@@ -9,6 +9,7 @@ import fs from "node:fs";
|
|
|
9
9
|
import path from "node:path";
|
|
10
10
|
import os from "node:os";
|
|
11
11
|
import { spawn } from "node:child_process";
|
|
12
|
+
import { executionMode } from "./execution.mjs";
|
|
12
13
|
|
|
13
14
|
const MIN_NODE_MAJOR = 18;
|
|
14
15
|
const DEFAULT_LLAMA_HOST = "127.0.0.1";
|
|
@@ -189,11 +190,28 @@ export function checkPodmanDaemon(facts) {
|
|
|
189
190
|
export function checkEndpoint(facts) {
|
|
190
191
|
const { execution, endpoint } = facts;
|
|
191
192
|
if (execution === "bedrock") return checkBedrockEndpoint(endpoint);
|
|
193
|
+
if (execution === "hosted") {
|
|
194
|
+
return {
|
|
195
|
+
ok: true,
|
|
196
|
+
message: "No local model by design; jobs run on API and subscription agents (run 'nomarmy agents list' to see them).",
|
|
197
|
+
};
|
|
198
|
+
}
|
|
199
|
+
if (execution === "remote") return checkRemoteEndpoint(endpoint);
|
|
192
200
|
if (execution === "local") return checkLocalEndpoint(endpoint);
|
|
193
201
|
return {
|
|
194
202
|
ok: false,
|
|
195
203
|
message: `NOMARMY_EXECUTION '${execution}' is not recognized.`,
|
|
196
|
-
fix: "Set NOMARMY_EXECUTION to 'local' or 'bedrock' (see config/common.env).",
|
|
204
|
+
fix: "Set NOMARMY_EXECUTION to 'local', 'remote', 'hosted', or 'bedrock' (see config/common.env).",
|
|
205
|
+
};
|
|
206
|
+
}
|
|
207
|
+
|
|
208
|
+
function checkRemoteEndpoint(endpoint) {
|
|
209
|
+
const { url, healthy, error } = endpoint;
|
|
210
|
+
if (healthy) return { ok: true, message: `Remote model server is healthy at ${url}.` };
|
|
211
|
+
return {
|
|
212
|
+
ok: false,
|
|
213
|
+
message: `Remote model server health check failed at ${url}${error ? `: ${error}` : "."}`,
|
|
214
|
+
fix: "Check that the remote server is running and reachable from this machine, then re-run 'nomarmy doctor'.",
|
|
197
215
|
};
|
|
198
216
|
}
|
|
199
217
|
|
|
@@ -324,6 +342,21 @@ async function probeLocalEndpoint(host, port) {
|
|
|
324
342
|
}
|
|
325
343
|
}
|
|
326
344
|
|
|
345
|
+
/** Probe a remotely managed model server's health endpoint. */
|
|
346
|
+
export async function probeRemoteEndpoint(llamaUrl, fetchImpl = fetch) {
|
|
347
|
+
const url = `${llamaUrl}/health`;
|
|
348
|
+
const controller = new AbortController();
|
|
349
|
+
const timer = setTimeout(() => controller.abort(), 2000);
|
|
350
|
+
try {
|
|
351
|
+
const res = await fetchImpl(url, { signal: controller.signal });
|
|
352
|
+
return { mode: "remote", url, healthy: res.ok, error: res.ok ? null : `HTTP ${res.status}` };
|
|
353
|
+
} catch (err) {
|
|
354
|
+
return { mode: "remote", url, healthy: false, error: err.name === "AbortError" ? "timed out" : err.message };
|
|
355
|
+
} finally {
|
|
356
|
+
clearTimeout(timer);
|
|
357
|
+
}
|
|
358
|
+
}
|
|
359
|
+
|
|
327
360
|
/**
|
|
328
361
|
* Gather the (non-network-secret) facts a bedrock profile needs: is a region
|
|
329
362
|
* configured and syntactically valid, and are credentials discoverable at
|
|
@@ -344,6 +377,20 @@ function collectBedrockFacts(env) {
|
|
|
344
377
|
return { mode: "bedrock", region, baseUrl, regionValid, credentialsPresent };
|
|
345
378
|
}
|
|
346
379
|
|
|
380
|
+
/** Resolve execution mode and collect only its endpoint facts. */
|
|
381
|
+
export async function collectEndpointFacts(env = process.env, fetchImpl = fetch) {
|
|
382
|
+
const configured = String(env.NOMARMY_EXECUTION || "local").trim().toLowerCase();
|
|
383
|
+
const resolved = executionMode(env);
|
|
384
|
+
const execution = ["local", "remote", "hosted", "bedrock"].includes(configured) ? resolved.mode : configured;
|
|
385
|
+
if (execution === "bedrock") return { execution, endpoint: collectBedrockFacts(env) };
|
|
386
|
+
if (execution === "hosted") return { execution, endpoint: { mode: "hosted" } };
|
|
387
|
+
if (execution === "remote") return { execution, endpoint: await probeRemoteEndpoint(resolved.llamaUrl, fetchImpl) };
|
|
388
|
+
return {
|
|
389
|
+
execution,
|
|
390
|
+
endpoint: await probeLocalEndpoint(env.NOMARMY_LLAMA_HOST || DEFAULT_LLAMA_HOST, env.NOMARMY_LLAMA_PORT || DEFAULT_LLAMA_PORT),
|
|
391
|
+
};
|
|
392
|
+
}
|
|
393
|
+
|
|
347
394
|
/**
|
|
348
395
|
* Collect every fact the checks need. Side effects only - no decisions here.
|
|
349
396
|
* @param {NodeJS.ProcessEnv} [env]
|
|
@@ -382,10 +429,7 @@ export async function collectFacts(env = process.env) {
|
|
|
382
429
|
const podmanDaemon = podmanPath
|
|
383
430
|
? await probePodmanDaemon(podmanPath)
|
|
384
431
|
: { reachable: false, error: "podman not found" };
|
|
385
|
-
const execution = (env
|
|
386
|
-
const endpoint = execution === "bedrock"
|
|
387
|
-
? collectBedrockFacts(env)
|
|
388
|
-
: await probeLocalEndpoint(env.NOMARMY_LLAMA_HOST || DEFAULT_LLAMA_HOST, env.NOMARMY_LLAMA_PORT || DEFAULT_LLAMA_PORT);
|
|
432
|
+
const { execution, endpoint } = await collectEndpointFacts(env);
|
|
389
433
|
return {
|
|
390
434
|
nodeVersion: process.version,
|
|
391
435
|
platform,
|
|
@@ -0,0 +1,57 @@
|
|
|
1
|
+
// Where this install's models run, in one place, so the MCP server, doctor,
|
|
2
|
+
// health and the CLI agree.
|
|
3
|
+
//
|
|
4
|
+
// local a llama-server on this machine (the default)
|
|
5
|
+
// remote a llama-server nomArmy doesn't run, e.g. a team's GPU server or
|
|
6
|
+
// an SSH tunnel to one: NOMARMY_EXECUTION=remote (what `nomarmy
|
|
7
|
+
// setup --llama-url` writes), or local with a non-loopback
|
|
8
|
+
// NOMARMY_LLAMA_HOST
|
|
9
|
+
// hosted no local model at all: every job runs on an api or subscription
|
|
10
|
+
// agent (NOMARMY_EXECUTION=hosted)
|
|
11
|
+
// bedrock the Bedrock cloud profile (NOMARMY_EXECUTION=bedrock)
|
|
12
|
+
//
|
|
13
|
+
// Settings come from config/common.env (and a profile), which `nomarmy
|
|
14
|
+
// connect` forwards to the MCP server's environment.
|
|
15
|
+
|
|
16
|
+
const LOOPBACK = new Set(["127.0.0.1", "localhost", "::1", "[::1]", "0.0.0.0"]);
|
|
17
|
+
|
|
18
|
+
/** Whether a host name or address points at this machine. */
|
|
19
|
+
export function isLoopbackHost(host) {
|
|
20
|
+
const h = String(host ?? "").trim().toLowerCase();
|
|
21
|
+
return h === "" || LOOPBACK.has(h) || /^127\./.test(h);
|
|
22
|
+
}
|
|
23
|
+
|
|
24
|
+
/**
|
|
25
|
+
* This install's execution mode.
|
|
26
|
+
* @returns {{ mode: "local"|"remote"|"hosted"|"bedrock", hasLocalModel: boolean, managesModelServer: boolean, llamaHost: string, llamaPort: string, llamaUrl: string|null }}
|
|
27
|
+
* hasLocalModel: jobs can run on the `local` agent (local or remote).
|
|
28
|
+
* managesModelServer: nomArmy starts/stops llama-server and sizes it
|
|
29
|
+
* against this machine's memory (local only).
|
|
30
|
+
*/
|
|
31
|
+
export function executionMode(env = process.env) {
|
|
32
|
+
const execution = String(env.NOMARMY_EXECUTION || "local").trim().toLowerCase();
|
|
33
|
+
const llamaHost = String(env.NOMARMY_LLAMA_HOST || "127.0.0.1").trim();
|
|
34
|
+
const llamaPort = String(env.NOMARMY_LLAMA_PORT || "8080").trim();
|
|
35
|
+
const llamaUrl = `http://${llamaHost.includes(":") && !llamaHost.startsWith("[") ? `[${llamaHost}]` : llamaHost}:${llamaPort}`;
|
|
36
|
+
if (execution === "hosted") return { mode: "hosted", hasLocalModel: false, managesModelServer: false, llamaHost, llamaPort, llamaUrl: null };
|
|
37
|
+
if (execution === "bedrock") return { mode: "bedrock", hasLocalModel: false, managesModelServer: false, llamaHost, llamaPort, llamaUrl: null };
|
|
38
|
+
const remote = execution === "remote" || !isLoopbackHost(llamaHost);
|
|
39
|
+
return { mode: remote ? "remote" : "local", hasLocalModel: true, managesModelServer: !remote, llamaHost, llamaPort, llamaUrl };
|
|
40
|
+
}
|
|
41
|
+
|
|
42
|
+
/**
|
|
43
|
+
* Parse a model-server URL given to `nomarmy setup --llama-url` into the
|
|
44
|
+
* NOMARMY_LLAMA_HOST / NOMARMY_LLAMA_PORT pair. http only (llama-server
|
|
45
|
+
* speaks plain HTTP; put TLS in front of it and use its host if needed),
|
|
46
|
+
* port defaulting to 8080. Throws a clear error on anything else.
|
|
47
|
+
*/
|
|
48
|
+
export function parseLlamaUrl(input) {
|
|
49
|
+
let url;
|
|
50
|
+
try { url = new URL(String(input ?? "").trim()); } catch { throw new Error(`"${input}" isn't a URL; use the form http://host:8080`); }
|
|
51
|
+
if (url.protocol !== "http:") throw new Error(`use an http:// URL for llama-server (got ${url.protocol}//)`);
|
|
52
|
+
if (url.username || url.password) throw new Error("the URL can't carry a username or password");
|
|
53
|
+
if (url.pathname && url.pathname !== "/") throw new Error(`give the server's address only, without a path (got ${url.pathname})`);
|
|
54
|
+
const host = url.hostname.replace(/^\[|\]$/g, "");
|
|
55
|
+
if (!host) throw new Error("the URL has no host");
|
|
56
|
+
return { host, port: url.port || "8080" };
|
|
57
|
+
}
|
package/lib/health.mjs
CHANGED
|
@@ -19,6 +19,7 @@
|
|
|
19
19
|
import { execFile } from "node:child_process";
|
|
20
20
|
import fs from "node:fs";
|
|
21
21
|
import path from "node:path";
|
|
22
|
+
import { executionMode } from "./execution.mjs";
|
|
22
23
|
import { modelRejection } from "./openclaw-errors.mjs";
|
|
23
24
|
import { providerConfigured, readOpenclawConfig } from "./openclaw-config.mjs";
|
|
24
25
|
|
|
@@ -70,11 +71,15 @@ export function versionIssues({ installed, latest, plugins = [] }) {
|
|
|
70
71
|
}
|
|
71
72
|
|
|
72
73
|
/** Roles, and the General, pointing at something unusable (from describeArmy's summary). */
|
|
73
|
-
export function armyIssues(summary) {
|
|
74
|
+
export function armyIssues(summary, mode = "local") {
|
|
74
75
|
const issues = [];
|
|
75
76
|
if (summary?.general?.problem) issues.push({ id: "army:general", severity: "warn", title: "The General's agent isn't set up", detail: summary.general.problem, fix: "nomarmy army general <agent>", short: null });
|
|
76
77
|
for (const [role, r] of Object.entries(summary?.roles ?? {})) {
|
|
77
78
|
if (r.problem) issues.push({ id: `army:role:${role}:${r.problem}`, severity: "warn", title: `Role ${role} can't be dispatched`, detail: r.problem, fix: `nomarmy army assign ${role} <agent> [model|auto]`, short: `${role} unusable` });
|
|
79
|
+
else if (mode === "hosted" && r.agent === "local") issues.push({ id: `army:hosted-local:${role}`, severity: "warn",
|
|
80
|
+
title: `Role ${role} uses the local agent in hosted mode`,
|
|
81
|
+
detail: "Hosted installs have no local model, so jobs for this role will be refused.",
|
|
82
|
+
fix: `nomarmy army assign ${role} <agent> [model] or nomarmy army init --agent <name>`, short: `${role} uses local` });
|
|
78
83
|
else if (r.hostTools?.implementRole && !r.hostTools.allowed) issues.push({ id: `army:host-tools:${role}:${r.agent}`, severity: "warn",
|
|
79
84
|
title: `Role ${role} builds on ${r.agent}, whose tools run on this machine`,
|
|
80
85
|
detail: `${r.agent}'s worker runs its own shell on your machine, with your files and the network, outside nomArmy's sandbox, so nomArmy refuses implement jobs on it.`,
|
|
@@ -178,7 +183,7 @@ export function leftoverIssues({ retainedWorktrees = 0, jobsBytes = 0, staleRunn
|
|
|
178
183
|
* Run every check. `env` supplies what each needs, with real defaults;
|
|
179
184
|
* tests pass their own.
|
|
180
185
|
*/
|
|
181
|
-
export async function runHealthChecks({ now = Date.now(), openclawCmd = process.env.NOMARMY_OPENCLAW_CMD || "openclaw", run = runBounded, armySummary = null, agentsError = null, jobsRoot = null, pidAlive = () => true, agents = null, openclawConfig = null, vendors = {}, modelsInUse = null, autoPruned = null } = {}) {
|
|
186
|
+
export async function runHealthChecks({ now = Date.now(), mode = "local", openclawCmd = process.env.NOMARMY_OPENCLAW_CMD || "openclaw", run = runBounded, armySummary = null, agentsError = null, jobsRoot = null, pidAlive = () => true, agents = null, openclawConfig = null, vendors = {}, modelsInUse = null, autoPruned = null } = {}) {
|
|
182
187
|
const issues = [];
|
|
183
188
|
if (autoPruned?.freedBytes) issues.push({ id: `auto-prune:${new Date(now).toISOString()}`, severity: "info",
|
|
184
189
|
title: `Freed ${(autoPruned.freedBytes / 1024 ** 3).toFixed(2)} GB: ${[autoPruned.pruned ? `runtime data of ${autoPruned.pruned} finished job${autoPruned.pruned === 1 ? "" : "s"} older than ${autoPruned.olderThanHours}h` : null, autoPruned.scratchCleared ? `OpenClaw scratch files of ${autoPruned.scratchCleared} more` : null].filter(Boolean).join(", ")}`,
|
|
@@ -194,7 +199,7 @@ export async function runHealthChecks({ now = Date.now(), openclawCmd = process.
|
|
|
194
199
|
const pluginVersion = /Version:\s*(\S+)/.exec(plugins.stdout ?? "")?.[1];
|
|
195
200
|
if (version.ok) issues.push(...versionIssues({ installed: version.stdout, latest: latest.ok ? latest.stdout : null, plugins: pluginVersion ? [{ id: "codex", version: pluginVersion }] : [] }));
|
|
196
201
|
if (agentsError) issues.push({ id: `config:agents:${agentsError}`, severity: "error", title: "agents.yml can't be loaded", detail: agentsError, fix: "nomarmy agents list (shows the problem)", short: "agents.yml broken" });
|
|
197
|
-
if (armySummary) issues.push(...armyIssues(armySummary));
|
|
202
|
+
if (armySummary) issues.push(...armyIssues(armySummary, mode));
|
|
198
203
|
if (jobsRoot) {
|
|
199
204
|
let retainedWorktrees = 0, jobsBytes = 0, staleRunning = 0;
|
|
200
205
|
const records = [];
|
|
@@ -245,7 +250,7 @@ export function recordHealth(file, result, { now = Date.now() } = {}) {
|
|
|
245
250
|
* periodic run and `nomarmy health`, then run and recorded to the shared
|
|
246
251
|
* health.json. Returns { result, toNotify }.
|
|
247
252
|
*/
|
|
248
|
-
export async function checkAndRecordHealth({ projectDir, stateRoot, configDir, now = Date.now() }) {
|
|
253
|
+
export async function checkAndRecordHealth({ projectDir, stateRoot, configDir, now = Date.now(), env = process.env }) {
|
|
249
254
|
const { loadAgents, describeAgent, agentProviderId } = await import("./agents.mjs");
|
|
250
255
|
const { loadArmy, describeArmy } = await import("./army.mjs");
|
|
251
256
|
const { pidAlive } = await import("./slots.mjs");
|
|
@@ -271,7 +276,8 @@ export async function checkAndRecordHealth({ projectDir, stateRoot, configDir, n
|
|
|
271
276
|
const ageMs = autoPruneAgeMs();
|
|
272
277
|
let autoPruned = null;
|
|
273
278
|
if (ageMs !== null) { try { autoPruned = { ...pruneJobRuntime({ stateRoot, olderThanMs: ageMs, now }), olderThanHours: ageMs / 3600000 }; } catch { /* best-effort */ } }
|
|
274
|
-
const
|
|
279
|
+
const mode = executionMode(env).mode;
|
|
280
|
+
const result = await runHealthChecks({ now, mode, armySummary, agentsError, jobsRoot: path.join(stateRoot, "jobs"), pidAlive,
|
|
275
281
|
agents, openclawConfig: readOpenclawConfig(), vendors: SUBSCRIPTION_VENDORS, modelsInUse, autoPruned });
|
|
276
282
|
const toNotify = recordHealth(path.join(stateRoot, "health.json"), result, { now });
|
|
277
283
|
return { result, toNotify };
|
package/mcp/server.mjs
CHANGED
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
import { createJobRuntime, jobLane, currentMaxPoolWorkers, splitJobsByLane, toolText, refusalText } from "../lib/admission.mjs";
|
|
2
2
|
export { jobLane, currentMaxPoolWorkers, splitJobsByLane, refusalText };
|
|
3
3
|
import { createExecutor, sleep } from "../lib/execute.mjs";
|
|
4
|
+
import { executionMode } from "../lib/execution.mjs";
|
|
4
5
|
export { JOB_PHASES } from "../lib/execute.mjs";
|
|
5
6
|
import { createVerificationFlow } from "../lib/verification-flow.mjs";
|
|
6
7
|
export { gitShowBuffer, gitModeAtBase, planProductionRevert, revertToBase, restoreWorkerVersion, blobHash, currentBlobHash } from "../lib/verification-flow.mjs";
|
|
@@ -207,7 +208,7 @@ export function makeHeartbeatTick(jobDir) { return heartbeatTick(jobDir, livePro
|
|
|
207
208
|
// Senti run none were tagged, so a 4-hour run went 8.46 hours unchecked.
|
|
208
209
|
let activeRunId = null;
|
|
209
210
|
|
|
210
|
-
export function expandJobs(jobs, { getArmy = currentArmy, getAgents = () => agentsConfig().agents, getActiveRun = () => activeRunId } = {}) {
|
|
211
|
+
export function expandJobs(jobs, { getArmy = currentArmy, getAgents = () => agentsConfig().agents, getActiveRun = () => activeRunId, env = process.env } = {}) {
|
|
211
212
|
const problems = [];
|
|
212
213
|
let army = null, agents = null;
|
|
213
214
|
const runId = getActiveRun();
|
|
@@ -217,6 +218,7 @@ export function expandJobs(jobs, { getArmy = currentArmy, getAgents = () => agen
|
|
|
217
218
|
if (j.army_role) { army ??= getArmy().army; j = expandArmyRole(j, army); }
|
|
218
219
|
const { agent, roleModel = null, ...rest } = j;
|
|
219
220
|
if (!agent) {
|
|
221
|
+
if (executionMode(env).mode === "hosted") throw new Error("this install has no local model (NOMARMY_EXECUTION=hosted): give the job an army_role or an agent (the army tool lists them)");
|
|
220
222
|
if (rest.model) throw new Error(`model "${rest.model}" needs an agent to run on: add agent (or army_role), or drop model to use the local model`);
|
|
221
223
|
return { ...rest, profile: rest.profile ?? "coder" };
|
|
222
224
|
}
|
|
@@ -254,7 +256,7 @@ export { executeJob };
|
|
|
254
256
|
|
|
255
257
|
const { WORKER_START_STAGGER_MS, activeJobs, runningCount, agentMaxConcurrent, withAgentSlot, track, notifyJobFinished, capacitySnapshot, admit, refusal, runBrief, recordJobInRun, trackInRun, launch, liveProgress, summarize } = createJobRuntime({
|
|
256
258
|
projectDir, stateRoot, jobsRoot, runsRoot, leasesRoot, slotsRoot, run, currentMaxWorkers, slug, agentsConfig, modelCatalogReady, budgetsForJob, resolveSubscriptionSelection, executeJob, subscriptionJobFieldProblems, repoPolicy, jobArgs,
|
|
257
|
-
budgetState, getActiveRunId: () => activeRunId,
|
|
259
|
+
env: process.env, budgetState, getActiveRunId: () => activeRunId,
|
|
258
260
|
});
|
|
259
261
|
export { runningCount, track };
|
|
260
262
|
|
package/package.json
CHANGED
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
"description": "A harness for AI coding workers whose claims are never trusted: your coding assistant stays in charge while workers implement and test in sandboxes, on local models, API keys or your own subscriptions.",
|
|
4
4
|
"author": "Rayson Technologies",
|
|
5
5
|
"license": "Apache-2.0",
|
|
6
|
-
"version": "0.1.0-alpha.
|
|
6
|
+
"version": "0.1.0-alpha.2",
|
|
7
7
|
"private": false,
|
|
8
8
|
"type": "module",
|
|
9
9
|
"engines": {
|
|
@@ -2,6 +2,14 @@
|
|
|
2
2
|
set -euo pipefail
|
|
3
3
|
ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"; source "$ROOT/scripts/lib.sh"; load_profile "${1:-}"
|
|
4
4
|
|
|
5
|
+
if [[ "$(nomarmy_execution_mode)" == hosted ]]; then
|
|
6
|
+
# No model provider to register: each api or subscription agent is set up
|
|
7
|
+
# with `nomarmy agents add`, which configures its own OpenClaw provider.
|
|
8
|
+
echo "Profile '$NOMARMY_PROFILE' runs every job on an api or subscription agent; no local model provider to configure."
|
|
9
|
+
echo "Add one with: nomarmy agents add"
|
|
10
|
+
exit 0
|
|
11
|
+
fi
|
|
12
|
+
|
|
5
13
|
PROVIDER="$NOMARMY_WORKER_PROVIDER"
|
|
6
14
|
PROFILE_ID="$PROVIDER:nomarmy-$NOMARMY_PROFILE"
|
|
7
15
|
|
|
@@ -39,7 +47,34 @@ else
|
|
|
39
47
|
MODEL_ID="$NOMARMY_MODEL_ALIAS"
|
|
40
48
|
AUTH_CHOICE="llama-cpp-existing-server"
|
|
41
49
|
PROFILE_ID="llama-cpp:nomarmy-local"
|
|
42
|
-
|
|
50
|
+
if [[ "$(nomarmy_execution_mode)" == remote ]]; then
|
|
51
|
+
# Someone else runs this server, so its model name and context come from
|
|
52
|
+
# the server itself rather than this machine's settings.
|
|
53
|
+
SERVER="http://$NOMARMY_LLAMA_HOST:$NOMARMY_LLAMA_PORT"
|
|
54
|
+
curl -fsS --max-time 5 "$SERVER/health" >/dev/null || { echo "ERROR: no llama-server answering at $SERVER/health. Set its address with: nomarmy setup --llama-url http://<host>:8080" >&2; exit 1; }
|
|
55
|
+
REMOTE_MODEL="$(curl -fsS --max-time 5 "$SERVER/v1/models" | node -e 'let s="";process.stdin.on("data",d=>s+=d).on("end",()=>{try{process.stdout.write(JSON.parse(s).data?.[0]?.id??"")}catch{}})' || true)"
|
|
56
|
+
[[ -n "$REMOTE_MODEL" ]] && MODEL_ID="$REMOTE_MODEL"
|
|
57
|
+
if [[ -n "$REMOTE_MODEL" && "$REMOTE_MODEL" != "${NOMARMY_WORKER_MODEL:-}" ]]; then
|
|
58
|
+
# Jobs ask for NOMARMY_WORKER_MODEL, so record the name this server
|
|
59
|
+
# actually serves (`nomarmy connect`, run next by install.sh, reads it).
|
|
60
|
+
COMMON="$ROOT/config/common.env"
|
|
61
|
+
for key in NOMARMY_MODEL_ALIAS NOMARMY_WORKER_MODEL; do
|
|
62
|
+
if grep -q "^$key=" "$COMMON"; then
|
|
63
|
+
KEY="$key" VALUE="$REMOTE_MODEL" node -e 'const fs=require("fs"),f=process.argv[1];fs.writeFileSync(f,fs.readFileSync(f,"utf8").replace(new RegExp(`^${process.env.KEY}=.*$`,"m"),`${process.env.KEY}=${process.env.VALUE}`))' "$COMMON"
|
|
64
|
+
else
|
|
65
|
+
printf '%s=%s\n' "$key" "$REMOTE_MODEL" >>"$COMMON"
|
|
66
|
+
fi
|
|
67
|
+
done
|
|
68
|
+
export NOMARMY_MODEL_ALIAS="$REMOTE_MODEL" NOMARMY_WORKER_MODEL="$REMOTE_MODEL"
|
|
69
|
+
echo "==> The server serves '$REMOTE_MODEL'; recorded it in config/common.env"
|
|
70
|
+
fi
|
|
71
|
+
# llama-server reports the context of one slot, which is one nom's share.
|
|
72
|
+
REMOTE_CTX="$(curl -fsS --max-time 5 "$SERVER/props" | node -e 'let s="";process.stdin.on("data",d=>s+=d).on("end",()=>{try{const n=JSON.parse(s).default_generation_settings?.n_ctx;if(Number.isInteger(n)&&n>0)process.stdout.write(String(n))}catch{}})' || true)"
|
|
73
|
+
[[ -n "$REMOTE_CTX" ]] && export NOMARMY_CONTEXT_PER_NOM="$REMOTE_CTX"
|
|
74
|
+
echo "==> Configuring OpenClaw against the llama-server at $SERVER, model $MODEL_ID${REMOTE_CTX:+, context $REMOTE_CTX per nom}"
|
|
75
|
+
else
|
|
76
|
+
echo "==> Configuring OpenClaw against local llama-server, model $MODEL_ID"
|
|
77
|
+
fi
|
|
43
78
|
fi
|
|
44
79
|
|
|
45
80
|
openclaw onboard --non-interactive --accept-risk \
|
|
@@ -61,9 +96,10 @@ if ! nomarmy_is_cloud; then
|
|
|
61
96
|
# against a freshly-resized 65536-token nom still overflowed at ~20K tokens
|
|
62
97
|
# of prompt, on literally the first turn, because OpenClaw was still
|
|
63
98
|
# enforcing its onboarding-time guess. NOMARMY_CONTEXT_PER_NOM (exported by
|
|
64
|
-
# nomarmy_validate_local, above, in load_profile
|
|
65
|
-
# computed truth for this exact
|
|
66
|
-
#
|
|
99
|
+
# nomarmy_validate_local, above, in load_profile, or read from a remote
|
|
100
|
+
# server's /props) is nomArmy's own already-computed truth for this exact
|
|
101
|
+
# number; write it back so OpenClaw's model registration cannot drift from
|
|
102
|
+
# the server it is actually talking to.
|
|
67
103
|
# models[0] assumes exactly the one custom local model this script just
|
|
68
104
|
# onboarded, which is what onboard --custom-model-id always produces here.
|
|
69
105
|
CONTEXT_WINDOW="${NOMARMY_CONTEXT_PER_NOM:-24576}"
|
|
@@ -1,7 +1,14 @@
|
|
|
1
1
|
#!/usr/bin/env bash
|
|
2
2
|
set -euo pipefail
|
|
3
3
|
ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"; source "$ROOT/scripts/lib.sh"; load_profile "${1:-}"
|
|
4
|
-
if
|
|
4
|
+
if ! nomarmy_manages_model_server; then
|
|
5
|
+
case "$(nomarmy_execution_mode)" in
|
|
6
|
+
remote) echo "Profile '$NOMARMY_PROFILE' uses the model server at $NOMARMY_LLAMA_HOST:$NOMARMY_LLAMA_PORT; skipping llama.cpp build." ;;
|
|
7
|
+
hosted) echo "Profile '$NOMARMY_PROFILE' runs every job on an api or subscription agent; skipping llama.cpp build." ;;
|
|
8
|
+
*) echo "Profile '$NOMARMY_PROFILE' uses hosted inference at ${NOMARMY_BEDROCK_BASE_URL:-Bedrock}; skipping llama.cpp build." ;;
|
|
9
|
+
esac
|
|
10
|
+
exit 0
|
|
11
|
+
fi
|
|
5
12
|
SRC="$NOMARMY_INSTALL_ROOT/llama.cpp"; mkdir -p "$NOMARMY_INSTALL_ROOT"
|
|
6
13
|
if [[ ! -d "$SRC/.git" ]]; then git clone --depth 1 https://github.com/ggml-org/llama.cpp.git "$SRC"; else git -C "$SRC" pull --ff-only; fi
|
|
7
14
|
if [[ "$(uname -s)" == Darwin ]]; then
|
package/scripts/lib.sh
CHANGED
|
@@ -24,9 +24,33 @@ nomarmy_available_profiles(){
|
|
|
24
24
|
find "$root/config/profiles" -name '*.env' -exec basename {} .env \; | sort | tr '\n' ' '
|
|
25
25
|
}
|
|
26
26
|
|
|
27
|
-
#
|
|
28
|
-
#
|
|
29
|
-
|
|
27
|
+
# Where this install's models run (lib/execution.mjs is the same logic for the
|
|
28
|
+
# MCP server). Scripts branch on these instead of on the profile name.
|
|
29
|
+
# local a llama-server on this machine, started and sized by nomArmy
|
|
30
|
+
# remote a llama-server nomArmy doesn't run, e.g. a team's GPU server or an
|
|
31
|
+
# SSH tunnel to one (NOMARMY_EXECUTION=remote, or local with a
|
|
32
|
+
# non-loopback NOMARMY_LLAMA_HOST)
|
|
33
|
+
# hosted no local model: every job runs on an api or subscription agent
|
|
34
|
+
# bedrock the Bedrock cloud profiles
|
|
35
|
+
nomarmy_is_loopback_host(){
|
|
36
|
+
local h
|
|
37
|
+
h="$(printf '%s' "${1:-}" | tr '[:upper:]' '[:lower:]')"
|
|
38
|
+
[[ -z "$h" || "$h" == localhost || "$h" == ::1 || "$h" == "[::1]" || "$h" == 0.0.0.0 || "$h" == 127.* ]]
|
|
39
|
+
}
|
|
40
|
+
nomarmy_execution_mode(){
|
|
41
|
+
case "${NOMARMY_EXECUTION:-local}" in
|
|
42
|
+
hosted) echo hosted ;;
|
|
43
|
+
bedrock) echo bedrock ;;
|
|
44
|
+
remote) echo remote ;;
|
|
45
|
+
*) if nomarmy_is_loopback_host "${NOMARMY_LLAMA_HOST:-127.0.0.1}"; then echo local; else echo remote; fi ;;
|
|
46
|
+
esac
|
|
47
|
+
}
|
|
48
|
+
# True on the Bedrock profiles: a Bedrock credential, the AWS CLI and a region.
|
|
49
|
+
nomarmy_is_cloud(){ [[ "$(nomarmy_execution_mode)" == bedrock ]]; }
|
|
50
|
+
# True when jobs can run on the `local` agent (a llama-server here or remote).
|
|
51
|
+
nomarmy_has_local_model(){ local m; m="$(nomarmy_execution_mode)"; [[ "$m" == local || "$m" == remote ]]; }
|
|
52
|
+
# True when nomArmy builds, starts, stops and sizes llama-server on this machine.
|
|
53
|
+
nomarmy_manages_model_server(){ [[ "$(nomarmy_execution_mode)" == local ]]; }
|
|
30
54
|
|
|
31
55
|
nomarmy_validate_cloud(){
|
|
32
56
|
local missing=()
|
|
@@ -124,6 +148,9 @@ load_profile(){
|
|
|
124
148
|
# Cloud profiles host no local model, so they carry no hardware or OS
|
|
125
149
|
# requirement and are the only profiles usable on a machine without a GPU.
|
|
126
150
|
nomarmy_validate_cloud
|
|
151
|
+
elif ! nomarmy_manages_model_server; then
|
|
152
|
+
# hosted and remote run no model here either, so any machine will do.
|
|
153
|
+
:
|
|
127
154
|
else
|
|
128
155
|
if [[ "$os" == Darwin && "$profile" != macbook-pro ]]; then
|
|
129
156
|
echo "ERROR: profile '$profile' requires Linux. Use macbook-pro on macOS." >&2
|
|
@@ -1,7 +1,14 @@
|
|
|
1
1
|
#!/usr/bin/env bash
|
|
2
2
|
set -euo pipefail
|
|
3
3
|
ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"; source "$ROOT/scripts/lib.sh"; load_profile "${1:-}"
|
|
4
|
-
if
|
|
4
|
+
if ! nomarmy_manages_model_server; then
|
|
5
|
+
case "$(nomarmy_execution_mode)" in
|
|
6
|
+
remote) echo "Profile '$NOMARMY_PROFILE' uses the model server at $NOMARMY_LLAMA_HOST:$NOMARMY_LLAMA_PORT; no llama-server to start." ;;
|
|
7
|
+
hosted) echo "Profile '$NOMARMY_PROFILE' runs every job on an api or subscription agent; no llama-server to start." ;;
|
|
8
|
+
*) echo "Profile '$NOMARMY_PROFILE' uses hosted inference at ${NOMARMY_BEDROCK_BASE_URL:-Bedrock}; no llama-server to start." ;;
|
|
9
|
+
esac
|
|
10
|
+
exit 0
|
|
11
|
+
fi
|
|
5
12
|
mkdir -p "$NOMARMY_INSTALL_ROOT/run" "$NOMARMY_INSTALL_ROOT/logs"
|
|
6
13
|
PID="$NOMARMY_INSTALL_ROOT/run/llama.pid"; LOG="$NOMARMY_INSTALL_ROOT/logs/llama-server.log"
|
|
7
14
|
if [[ -f "$PID" ]] && kill -0 "$(cat "$PID")" 2>/dev/null; then echo "llama-server already running PID $(cat "$PID")"; exit 0; fi
|
|
@@ -1,5 +1,12 @@
|
|
|
1
1
|
#!/usr/bin/env bash
|
|
2
2
|
set -euo pipefail
|
|
3
3
|
ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)"; source "$ROOT/scripts/lib.sh"; load_profile "${1:-}"
|
|
4
|
-
if
|
|
4
|
+
if ! nomarmy_manages_model_server; then
|
|
5
|
+
case "$(nomarmy_execution_mode)" in
|
|
6
|
+
remote) echo "Profile '$NOMARMY_PROFILE' uses the model server at $NOMARMY_LLAMA_HOST:$NOMARMY_LLAMA_PORT; nothing local to stop." ;;
|
|
7
|
+
hosted) echo "Profile '$NOMARMY_PROFILE' runs every job on an api or subscription agent; nothing local to stop." ;;
|
|
8
|
+
*) echo "Profile '$NOMARMY_PROFILE' uses hosted inference at ${NOMARMY_BEDROCK_BASE_URL:-Bedrock}; nothing local to stop." ;;
|
|
9
|
+
esac
|
|
10
|
+
exit 0
|
|
11
|
+
fi
|
|
5
12
|
PID="$NOMARMY_INSTALL_ROOT/run/llama.pid"; if [[ -f "$PID" ]] && kill -0 "$(cat "$PID")" 2>/dev/null; then kill "$(cat "$PID")"; fi; rm -f "$PID"; echo stopped
|
|
@@ -24,6 +24,11 @@ if nomarmy_is_cloud; then
|
|
|
24
24
|
echo "FAIL Bedrock model not listed in $NOMARMY_BEDROCK_REGION: $model"; fail=1
|
|
25
25
|
fi
|
|
26
26
|
done
|
|
27
|
+
elif [[ "$(nomarmy_execution_mode)" == hosted ]]; then
|
|
28
|
+
echo "INFO hosted profile '$NOMARMY_PROFILE': no local model; jobs run on api or subscription agents"
|
|
29
|
+
elif [[ "$(nomarmy_execution_mode)" == remote ]]; then
|
|
30
|
+
echo "INFO profile '$NOMARMY_PROFILE' uses the model server at $NOMARMY_LLAMA_HOST:$NOMARMY_LLAMA_PORT"
|
|
31
|
+
check curl -fsS --max-time 5 "http://$NOMARMY_LLAMA_HOST:$NOMARMY_LLAMA_PORT/health"
|
|
27
32
|
else
|
|
28
33
|
check "$NOMARMY_INSTALL_ROOT/llama-server" --version
|
|
29
34
|
if [[ "$(uname -s)" == Darwin ]]; then
|
|
@@ -36,9 +41,13 @@ else
|
|
|
36
41
|
check curl -fsS "http://$NOMARMY_LLAMA_HOST:$NOMARMY_LLAMA_PORT/health"
|
|
37
42
|
fi
|
|
38
43
|
|
|
39
|
-
|
|
40
|
-
|
|
41
|
-
|
|
44
|
+
if [[ "$(nomarmy_execution_mode)" == hosted ]]; then
|
|
45
|
+
echo "INFO no worker model to check; nomarmy doctor checks each agent"
|
|
46
|
+
else
|
|
47
|
+
openclaw models list --provider "$NOMARMY_WORKER_PROVIDER" | grep -q "$NOMARMY_WORKER_MODEL" \
|
|
48
|
+
&& echo "PASS OpenClaw worker model ($NOMARMY_WORKER_PROVIDER/$NOMARMY_WORKER_MODEL)" \
|
|
49
|
+
|| { echo "FAIL OpenClaw worker model ($NOMARMY_WORKER_PROVIDER/$NOMARMY_WORKER_MODEL)"; fail=1; }
|
|
50
|
+
fi
|
|
42
51
|
# The sandbox config keeps its "docker" sub-key namespace regardless of
|
|
43
52
|
# backend, so grepping the whole block for "docker" would falsely pass even
|
|
44
53
|
# when the backend is podman. Check the actual backend value instead.
|
|
@@ -46,12 +55,13 @@ SANDBOX_BACKEND="$(openclaw config get agents.defaults.sandbox.backend 2>/dev/nu
|
|
|
46
55
|
[[ "$SANDBOX_BACKEND" == "podman" ]] && echo 'PASS Podman sandbox configured' || { echo "FAIL Podman sandbox (backend is '$SANDBOX_BACKEND')"; fail=1; }
|
|
47
56
|
|
|
48
57
|
# The no-network sandbox is what keeps repository content away from any
|
|
49
|
-
# credential the host process holds. It is a hard requirement on cloud
|
|
58
|
+
# credential the host process holds. It is a hard requirement on cloud and
|
|
59
|
+
# hosted profiles, where every job runs against a real credential.
|
|
50
60
|
SANDBOX_NET="$(openclaw config get agents.defaults.sandbox.docker.network 2>/dev/null | tr -d '[:space:]"' || true)"
|
|
51
61
|
if [[ "$SANDBOX_NET" == "none" ]]; then
|
|
52
62
|
echo 'PASS Sandbox network isolated'
|
|
53
|
-
elif
|
|
54
|
-
echo "FAIL Sandbox network is '$SANDBOX_NET', must be 'none' on a
|
|
63
|
+
elif ! nomarmy_has_local_model; then
|
|
64
|
+
echo "FAIL Sandbox network is '$SANDBOX_NET', must be 'none' on a ${NOMARMY_EXECUTION} profile"; fail=1
|
|
55
65
|
else
|
|
56
66
|
echo "WARN Sandbox network is '$SANDBOX_NET', expected 'none'"
|
|
57
67
|
fi
|