nomarmy 0.1.0-alpha.12 → 0.1.0-alpha.13
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/README.md +6 -0
- package/bin/nomarmy.mjs +41 -5
- package/lib/coordinator-instructions.mjs +2 -2
- package/lib/execute.mjs +17 -5
- package/lib/git-record.mjs +25 -2
- package/lib/openclaw-run.mjs +31 -1
- package/lib/process.mjs +4 -1
- package/lib/scout.mjs +15 -2
- package/lib/transcript.mjs +3 -0
- package/mcp/server.mjs +9 -0
- package/package.json +1 -1
- package/playbooks/feature.md +1 -1
package/README.md
CHANGED
|
@@ -52,6 +52,8 @@ Failing verification stays failed, unconditionally. A malformed report isn't aut
|
|
|
52
52
|
|
|
53
53
|
**Checking without building** costs nothing: `mode: verify` runs a verification profile against any branch, with no worker and no model tokens.
|
|
54
54
|
|
|
55
|
+
**Want deeper checks?** Three optional [validators](https://github.com/rayson-tech/nomarmy/blob/main/docs/validators.md) go further, each only adding review flags: mutation testing (do the tests pin down the changed lines?), Jev (do a scout's citations support its findings, does a report match its diff?) and a model judge (acceptance criteria, weakened tests).
|
|
56
|
+
|
|
55
57
|
## Where the work runs
|
|
56
58
|
|
|
57
59
|
- **Agents** say where a job can run: an API key, your own ChatGPT or Muse Code subscription, or a local model on llama.cpp.
|
|
@@ -74,10 +76,12 @@ Developed and maintained by Rayson Technologies. This is an alpha (`0.1.0-alpha`
|
|
|
74
76
|
| [Agents and the army](https://github.com/rayson-tech/nomarmy/blob/main/docs/agents-and-army.md) | Where a job can run, who does what, usage limits, picking an agent |
|
|
75
77
|
| [`/feature` runs](https://github.com/rayson-tech/nomarmy/blob/main/docs/feature-runs.md) | A feature end to end, and watching what nomArmy is doing |
|
|
76
78
|
| [Your repository](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md) | `.nomarmy.yml`, verification, dependencies, private registries, what nomArmy checks |
|
|
79
|
+
| [Validators](https://github.com/rayson-tech/nomarmy/blob/main/docs/validators.md) | Optional deeper checks: mutation testing, Jev, a model judge |
|
|
77
80
|
| [Harnesses](https://github.com/rayson-tech/nomarmy/blob/main/docs/harnesses.md) | Ecosystem registry, detection, network levels, and requirements |
|
|
78
81
|
| [Configuration](https://github.com/rayson-tech/nomarmy/blob/main/docs/configuration.md) | Settings, swapping the local model, sizing, admission |
|
|
79
82
|
| [Reference](https://github.com/rayson-tech/nomarmy/blob/main/docs/reference.md) | Every CLI command and MCP tool |
|
|
80
83
|
| [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md) | What the sandbox holds back, and the one exception |
|
|
84
|
+
| [FAQ](https://github.com/rayson-tech/nomarmy/blob/main/docs/faq.md) | Which model for which role, switching models, usage limits |
|
|
81
85
|
| [Troubleshooting](https://github.com/rayson-tech/nomarmy/blob/main/docs/troubleshooting.md) | Symptoms and fixes |
|
|
82
86
|
|
|
83
87
|
## Security
|
|
@@ -97,6 +101,8 @@ A nom gets a writable git worktree inside a Podman sandbox and nothing else: no
|
|
|
97
101
|
| The army and `/feature` | Driven by a real Claude Code General across three runs, about 18 implement jobs |
|
|
98
102
|
| Harnesses: Go, Rust, Python, Node and mixed repos; Playwright; fake services | Live-verified offline |
|
|
99
103
|
| Private registries and a verification-only network allowlist | Live-verified; each passed an independent security review |
|
|
104
|
+
| Validators: mutation testing, Jev, a model judge | Unit and live tested; Jev and the judge evaluated on real job records |
|
|
105
|
+
| `nomarmy stats` | Checked against a hand-built report on real job records |
|
|
100
106
|
|
|
101
107
|
What we've learned from real runs, including where delegating pays and where it doesn't, is in [docs/findings.md](https://github.com/rayson-tech/nomarmy/blob/main/docs/findings.md).
|
|
102
108
|
|
package/bin/nomarmy.mjs
CHANGED
|
@@ -21,6 +21,7 @@ import { recommend, customRecommendation, evaluateConfig, bytesPerKvElementForCa
|
|
|
21
21
|
import { connectClaude, connectCodex, connectCursor, cursorAlreadyConnected, deriveWorkerModelEnv, defaultInstallDir, installMcpCopy, SCOPES, claudeUserScoped, portableServerLaunch } from "../lib/connect.mjs";
|
|
22
22
|
import { compareVersions, readPackageVersion, readInstallVersions, copyIsStale } from "../lib/install-freshness.mjs";
|
|
23
23
|
import { loadJobRecords, computeStats, formatStats, parseSince, resolveRepo } from "../lib/stats.mjs";
|
|
24
|
+
import { requestJobStop } from "../lib/openclaw-run.mjs";
|
|
24
25
|
import { loadValidators, saveJevKey, removeJev, jevSettings, askJev, validatorsPath, JEV_CHECKS, saveJudge, removeJudge, judgeSettings } from "../lib/validators.mjs";
|
|
25
26
|
import { probeModel } from "../lib/model-probe.mjs";
|
|
26
27
|
import { ID_RE, AUTH_ENV_NAME_RE, OPENCLAW_PROVIDER_ID_RE, openclawProviderId, isNativeProviderType } from "../lib/dispatch-schema.mjs";
|
|
@@ -198,17 +199,25 @@ Usage: nomarmy <command> [options]
|
|
|
198
199
|
which agent the General is, in --global
|
|
199
200
|
(default) or --local
|
|
200
201
|
config paths where agents.yml and the three army layers live
|
|
201
|
-
jobs [--watch|--events|--prune|--wait <jobId>] [--interval N] [--older-than DAYS]
|
|
202
|
+
jobs [--watch|--events [--until-done]|--prune|--wait <jobId>|--stop <jobId> [--reason <text>]] [--interval N] [--older-than DAYS]
|
|
202
203
|
what's running across every session (agent, model, phase,
|
|
203
204
|
last tool call, files changed, heartbeat) and what just
|
|
204
205
|
finished; --watch redraws every N seconds (default 3);
|
|
205
206
|
--events prints one line per start, phase change and
|
|
206
|
-
finish (for
|
|
207
|
-
|
|
207
|
+
finish (--json for JSON lines). It's a stream: read it
|
|
208
|
+
with a monitor that wakes on each line. A background
|
|
209
|
+
command is only reported when it exits, so there use
|
|
210
|
+
--events --until-done, which exits once every job it saw
|
|
211
|
+
running has finished (or --wait for one job). The plain
|
|
212
|
+
stream ends on its own after 30 minutes with nothing
|
|
213
|
+
running (--idle-minutes N); --prune removes the bulky runtime data
|
|
208
214
|
from finished jobs older than DAYS (default 2), keeping
|
|
209
215
|
their records, reports and any retained worktree;
|
|
210
216
|
--wait <jobId> [--timeout <seconds>] blocks for one job
|
|
211
|
-
to finish (default timeout 1800; --json is supported)
|
|
217
|
+
to finish (default timeout 1800; --json is supported);
|
|
218
|
+
--stop <jobId> stops a running job's worker (no report
|
|
219
|
+
recovery, no verification), keeping its worktree for
|
|
220
|
+
continue_from
|
|
212
221
|
health check what's likely to break a run before it does:
|
|
213
222
|
expiring logins, an outdated OpenClaw or plugin, roles
|
|
214
223
|
that can't be dispatched, an unloadable agents.yml,
|
|
@@ -2518,6 +2527,12 @@ function renderJobs({ running, recent }) {
|
|
|
2518
2527
|
*/
|
|
2519
2528
|
async function streamJobEvents() {
|
|
2520
2529
|
const interval = Math.max(1, Number(value("interval", "3")) || 3) * 1000;
|
|
2530
|
+
// A stream nobody reads must still end: a General that ran this as a
|
|
2531
|
+
// background command (reported only on exit) was never told jobs had
|
|
2532
|
+
// finished, and eight of these streams were left running for days.
|
|
2533
|
+
const untilDone = flag("until-done");
|
|
2534
|
+
const idleLimitMs = Math.max(1, Number(value("idle-minutes", "30")) || 30) * 60000;
|
|
2535
|
+
let idleSinceMs = Date.now(), sawRunning = false;
|
|
2521
2536
|
const seen = new Map();
|
|
2522
2537
|
const emit = (event, job, detail = "") => {
|
|
2523
2538
|
if (json) console.log(JSON.stringify({ at: new Date().toISOString(), event, jobId: job.jobId, agent: job.agent, model: job.model, phase: job.phase, detail }));
|
|
@@ -2540,10 +2555,23 @@ async function streamJobEvents() {
|
|
|
2540
2555
|
seen.clear();
|
|
2541
2556
|
for (const [id, j] of now) seen.set(id, j);
|
|
2542
2557
|
first = false;
|
|
2558
|
+
if (running.length) { sawRunning = true; idleSinceMs = Date.now(); }
|
|
2559
|
+
else if (untilDone && sawRunning) {
|
|
2560
|
+
if (json) console.log(JSON.stringify({ at: new Date().toISOString(), event: "done", detail: "every job seen running has finished" }));
|
|
2561
|
+
else console.log(`${new Date().toLocaleTimeString()} done every job seen running has finished`);
|
|
2562
|
+
return;
|
|
2563
|
+
} else if (Date.now() - idleSinceMs >= (untilDone ? Math.min(idleLimitMs, 120000) : idleLimitMs)) {
|
|
2564
|
+
const why = untilDone ? "no job was running to wait for" : `nothing has run for ${Math.round(idleLimitMs / 60000)} minutes`;
|
|
2565
|
+
if (json) console.log(JSON.stringify({ at: new Date().toISOString(), event: "idle", detail: why }));
|
|
2566
|
+
else console.log(`${new Date().toLocaleTimeString()} idle ${why}; exiting`);
|
|
2567
|
+
return;
|
|
2568
|
+
}
|
|
2543
2569
|
await new Promise((r) => setTimeout(r, interval));
|
|
2544
2570
|
}
|
|
2545
2571
|
}
|
|
2546
2572
|
|
|
2573
|
+
const commitSha = (commit) => (typeof commit === "string" ? commit : typeof commit?.sha === "string" ? commit.sha : null);
|
|
2574
|
+
|
|
2547
2575
|
/** Wait for one job in the shared, cross-session state directory. */
|
|
2548
2576
|
async function waitForJobCli() {
|
|
2549
2577
|
const requested = value("wait");
|
|
@@ -2580,7 +2608,8 @@ async function waitForJobCli() {
|
|
|
2580
2608
|
outcome: meta.outcome ?? status.outcome ?? null,
|
|
2581
2609
|
coordinatorStatus: meta.coordinatorStatus ?? status.coordinatorStatus ?? null,
|
|
2582
2610
|
branch: meta.branch ?? status.branch ?? null,
|
|
2583
|
-
commit:
|
|
2611
|
+
// A job that made no commit has commit: { created: false, sha: null }; only a sha is a commit.
|
|
2612
|
+
commit: commitSha(meta.commit) ?? commitSha(status.commit),
|
|
2584
2613
|
issues,
|
|
2585
2614
|
};
|
|
2586
2615
|
if (json) out(result);
|
|
@@ -2621,6 +2650,13 @@ function pruneJobRuntimeCli() {
|
|
|
2621
2650
|
|
|
2622
2651
|
async function cmdJobs() {
|
|
2623
2652
|
if (flag("wait")) return waitForJobCli();
|
|
2653
|
+
if (flag("stop")) {
|
|
2654
|
+
const r = requestJobStop({ jobsRoot: jobsRootDir(), jobId: value("stop"), reason: value("reason") });
|
|
2655
|
+
if (json) return out(r);
|
|
2656
|
+
console.log(r.ok ? c.green(`✓ ${r.message}`) : c.red(`✗ ${r.message}`));
|
|
2657
|
+
if (!r.ok) process.exitCode = 1;
|
|
2658
|
+
return;
|
|
2659
|
+
}
|
|
2624
2660
|
if (flag("events")) return streamJobEvents();
|
|
2625
2661
|
if (flag("prune")) return pruneJobRuntimeCli();
|
|
2626
2662
|
if (json) return out(collectJobs());
|
|
@@ -15,12 +15,12 @@ Before dispatching:
|
|
|
15
15
|
- A Claude subscription agent (claude-cli) runs its tools on this machine, outside the sandbox: use it for scout and review work. nomArmy refuses implement jobs on it unless the operator set allow_host_tools; send build work to a sandboxed agent.
|
|
16
16
|
- To run tests without changing anything, use mode: verify; it costs no model usage.
|
|
17
17
|
- Brief outcomes, not edits: a task, explicit acceptance criteria, and the tests that prove it. Put facts you've already resolved in evidence.
|
|
18
|
-
- Prefer local_worker_start for anything longer than a few minutes. Right after, if you can run a background command, run \`nomarmy jobs --wait <job_id>\` in the background so you're told the moment it finishes and can tell the operator; otherwise poll local_worker_status with wait_seconds. For a whole feature, use /feature (run_start keeps a run's jobs, spend and hours bounded).
|
|
18
|
+
- Prefer local_worker_start for anything longer than a few minutes. Right after, if you can run a background command, run \`nomarmy jobs --wait <job_id>\` in the background so you're told the moment it finishes and can tell the operator (for several jobs, \`nomarmy jobs --events --until-done\`, which exits when they've all finished); otherwise poll local_worker_status with wait_seconds. Never run the plain \`nomarmy jobs --events\` stream as a background command: it only reports when it exits, so you'd never hear; it's for a monitor that reads each line. For a whole feature, use /feature (run_start keeps a run's jobs, spend and hours bounded).
|
|
19
19
|
|
|
20
20
|
Trust boundary:
|
|
21
21
|
- A worker's four-line report is a claim; nomArmy's verified git record and independent verification are the evidence. A job isn't complete if its report is missing or malformed, its STATUS is partial or blocked, STATUS done lacks VERIFICATION pass, or its changes aren't committed by nomArmy.
|
|
22
22
|
- Read the diff of anything material before integrating it. nomArmy commits on the worker's branch and never merges into yours: integration, conflicts and pushes are yours.
|
|
23
|
-
- To finish a job that came back partial, blocked
|
|
23
|
+
- A job on the wrong track can be stopped with local_worker_stop; it keeps its worktree. To finish a job that came back partial, blocked, failing verification or stopped, dispatch the correction with continue_from: <that job id> (a cheaper model is fine). The new job starts with its unfinished work in place and verifies the whole. Don't fix and commit a worker's files yourself: that lands them unverified. If you must, commit them on a branch, run mode: verify on it before building on it, and say so in your report.
|
|
24
24
|
- Failed and incomplete worktrees are kept for review; clean up with local_worker_cleanup or local_worker_sweep once you've decided.
|
|
25
25
|
|
|
26
26
|
Never delegate deployments, production access, cloud or SSH credentials, secrets, Terraform state or kubectl contexts to a worker.`;
|
package/lib/execute.mjs
CHANGED
|
@@ -1,4 +1,4 @@
|
|
|
1
|
-
import { writeStatus, shouldRetryTransientAbort, shouldAttemptScoutRecovery } from "./openclaw-run.mjs";
|
|
1
|
+
import { writeStatus, shouldRetryTransientAbort, shouldAttemptScoutRecovery, readStopRequest } from "./openclaw-run.mjs";
|
|
2
2
|
import fs from "node:fs";
|
|
3
3
|
import path from "node:path";
|
|
4
4
|
import { fileURLToPath } from "node:url";
|
|
@@ -12,7 +12,7 @@ import { continuationProblem, continuationBase, snapshotRetainedWork, applyRetai
|
|
|
12
12
|
import { checkScoutCitations, checkReportClaims } from "./jev-checks.mjs";
|
|
13
13
|
import { runJudge } from "./judge.mjs";
|
|
14
14
|
import { pickMutants, runMutants, describeSurvivors } from "./mutation.mjs";
|
|
15
|
-
import { estimateDisplacement } from "./transcript.mjs";
|
|
15
|
+
import { estimateDisplacement, readOpenClawTranscript } from "./transcript.mjs";
|
|
16
16
|
import { outlineFile, findReferences } from "./repo-query.mjs";
|
|
17
17
|
import { loadConfig } from "./config.mjs";
|
|
18
18
|
import { linkNodePackages, nodeModulesState, repairHostInstalls } from "./sandbox-images.mjs";
|
|
@@ -275,6 +275,9 @@ export function createExecutor(deps) {
|
|
|
275
275
|
// failed job that changed nothing was recorded "pass" (a Senti run),
|
|
276
276
|
// which reads as evidence about work that never happened.
|
|
277
277
|
independentVerification = normalizeVerification({ status: "not_run", basis: "not-applicable", reason: "the worker changed nothing, so there was none of its work to verify" }, verification ?? null);
|
|
278
|
+
} else if (workerStopReason === "stopped") {
|
|
279
|
+
// Stopped on request: end now, without spending time on tests.
|
|
280
|
+
independentVerification = normalizeVerification({ status: "not_run", basis: "not-applicable", reason: "the job was stopped on request" }, verification ?? null);
|
|
278
281
|
} else if (verificationFlow.verificationRunner || !reportValidation.valid) {
|
|
279
282
|
independentVerification = await runIndependentVerification({ profile: verification ?? null, cwd, jobId, baseSha: base.sha, branch, mode, record: preCommit, logFile: path.join(jobDir, "verification.log") });
|
|
280
283
|
}
|
|
@@ -533,7 +536,10 @@ export function createExecutor(deps) {
|
|
|
533
536
|
|
|
534
537
|
let coordinatorStatus = COORDINATOR_STATUS_BY_OUTCOME[finalOutcome.outcome] ?? "incomplete";
|
|
535
538
|
const issues = [...finalOutcome.reasons, ...(preCommit.issues ?? [])];
|
|
536
|
-
if (
|
|
539
|
+
if (workerStopReason === "stopped") {
|
|
540
|
+
const request = readStopRequest(jobDir);
|
|
541
|
+
issues.push(`stopped on request${request?.reason ? `: ${request.reason}` : ""}; the worktree is kept, so continue_from can pick the work up (on another model too)`);
|
|
542
|
+
} else if (workerError) issues.push(`worker error: ${String(workerError).split("\n")[0]}`);
|
|
537
543
|
if ((result ?? attempted)?.salvaged) issues.push(`runner cleanup failed after the run (${(result ?? attempted).salvagedFrom}); the worker's report was recovered from the run's transcript`);
|
|
538
544
|
if (jevClaims?.error) issues.push(`Jev report check skipped (${jevClaims.error}); this job's result doesn't depend on it`);
|
|
539
545
|
if (judged?.error) issues.push(`Judge ${judged.skipped ? "skipped" : "didn't answer"} (${judged.error}); this job's result doesn't depend on it`);
|
|
@@ -547,7 +553,7 @@ export function createExecutor(deps) {
|
|
|
547
553
|
// the diffstat right in the issue a caller actually reads -- not just
|
|
548
554
|
// buried in the full manifest's git record -- is what makes "go look at
|
|
549
555
|
// the worktree" worth doing instead of discarding the job.
|
|
550
|
-
issues.push(`repository changes remain uncommitted (${record.filesChanged} file(s), +${record.additions}/-${record.deletions}): ${commit.reason}`);
|
|
556
|
+
if (record.filesChanged > 0) issues.push(`repository changes remain uncommitted (${record.filesChanged} file(s), +${record.additions}/-${record.deletions}): ${commit.reason}`);
|
|
551
557
|
}
|
|
552
558
|
const failures = worker.toolSummary?.failures ?? 0; if (failures > 0) issues.push(`worker recorded ${failures} tool failure(s)`);
|
|
553
559
|
if (record.ignoredRuntimeJunk.length) issues.push(`runtime junk ignored: ${record.ignoredRuntimeJunk.join(", ")}`);
|
|
@@ -669,12 +675,18 @@ export function createExecutor(deps) {
|
|
|
669
675
|
const remainingSeconds = timeoutSeconds - Math.round(workerElapsedMs / 1000);
|
|
670
676
|
if (shouldAttemptScoutRecovery({ workerFailed, workerTimedOut, report, remainingSeconds })) {
|
|
671
677
|
reportRecoveryAttempted = true;
|
|
678
|
+
// What the first run left, in case the follow-up can't see its session.
|
|
679
|
+
let filesRead = [], earlierReply = reportText;
|
|
680
|
+
try {
|
|
681
|
+
const first = await readOpenClawTranscript(path.join(runtimeDir, "state"));
|
|
682
|
+
if (first.available) { filesRead = first.filesRead ?? []; earlierReply = earlierReply || first.lastAssistantText || ""; }
|
|
683
|
+
} catch { /* the reply alone still helps */ }
|
|
672
684
|
try {
|
|
673
685
|
const recoveryResult = await runOpenClaw({
|
|
674
686
|
task, acceptance, verification: null, mode, cwd: worktree, baseRef: base.ref, baseSha: base.sha,
|
|
675
687
|
timeoutSeconds: remainingSeconds, runtimeDir, profile, reasoning, pool, subscriptionWorker, onBehalfOf, model, reportSize, jobDir, workerId: workerId || jobId,
|
|
676
688
|
evidenceTool: evidencePlaced ? evidenceTool : null,
|
|
677
|
-
overridePrompt: scoutReportRecoveryPrompt({ report: used.report.scout, question: task, acceptance }), logSuffix: "-recovery",
|
|
689
|
+
overridePrompt: scoutReportRecoveryPrompt({ report: used.report.scout, question: task, acceptance, earlierReply, filesRead }), logSuffix: "-recovery",
|
|
678
690
|
});
|
|
679
691
|
const recoveryReport = parseScoutReport(finalText(recoveryResult), (recoveryResult?.budgetsUsed ?? used).scout);
|
|
680
692
|
if (!isScoutReportUnusable(recoveryReport)) {
|
package/lib/git-record.mjs
CHANGED
|
@@ -139,9 +139,26 @@ export function createGitRecord({ run, git, gitRaw }) {
|
|
|
139
139
|
// idleMinElapsedMs of the work phase has passed (an early snapshot mid-first-
|
|
140
140
|
// edit looks identical to no edit at all). A worktree read failing mid-write
|
|
141
141
|
// is expected, not an error; it just means "nothing to report this tick."
|
|
142
|
-
|
|
143
|
-
|
|
142
|
+
//
|
|
143
|
+
// An unchanged worktree isn't enough on its own: a frontier worker running
|
|
144
|
+
// a suite of thousands of tests, or reading before its next edit, changes
|
|
145
|
+
// no file for minutes. Three Senti jobs were cut off at 7 to 13 minutes
|
|
146
|
+
// while still working. With `activity` (the worker's transcript: event
|
|
147
|
+
// count and whether a tool call is still running), the breaker waits while
|
|
148
|
+
// the worker is active, and stops only when it has also gone quiet for
|
|
149
|
+
// idleMs, or when nothing has changed for ACTIVE_CAP times idleMs however
|
|
150
|
+
// busy it looks (a worker looping without progress).
|
|
151
|
+
function makeIdleDiffTick(cwd, { idleMs, minElapsedMs, activity = null }) {
|
|
152
|
+
const ACTIVE_CAP = 3;
|
|
153
|
+
let lastHash = null, lastChangeAtMs = 0, sawChange = false, lastEvents = null, lastActivityAtMs = 0, inFlight = false;
|
|
144
154
|
return async elapsedMs => {
|
|
155
|
+
if (activity) {
|
|
156
|
+
try {
|
|
157
|
+
const a = await activity();
|
|
158
|
+
if (a && a.events !== lastEvents) { lastEvents = a.events; lastActivityAtMs = elapsedMs; }
|
|
159
|
+
inFlight = Boolean(a?.toolInFlight);
|
|
160
|
+
} catch { /* a failed read just means no activity signal this tick */ }
|
|
161
|
+
}
|
|
145
162
|
let statusOut;
|
|
146
163
|
try { statusOut = await gitRaw(["status", "--porcelain=v1", "-z", "--untracked-files=all"], cwd); }
|
|
147
164
|
catch { return { stop: false }; }
|
|
@@ -182,6 +199,12 @@ export function createGitRecord({ run, git, gitRaw }) {
|
|
|
182
199
|
if (!sawChange || elapsedMs < minElapsedMs) return { stop: false };
|
|
183
200
|
const idleForMs = elapsedMs - lastChangeAtMs;
|
|
184
201
|
if (idleForMs < idleMs) return { stop: false };
|
|
202
|
+
if (activity) {
|
|
203
|
+
const quietForMs = elapsedMs - lastActivityAtMs;
|
|
204
|
+
const active = inFlight || quietForMs < idleMs;
|
|
205
|
+
if (active && idleForMs < idleMs * ACTIVE_CAP) return { stop: false };
|
|
206
|
+
if (active) return { stop: true, reason: "idle_diff", detail: `worktree unchanged for ${Math.round(idleForMs / 1000)}s while the worker kept working (${ACTIVE_CAP}x the idle limit)` };
|
|
207
|
+
}
|
|
185
208
|
return { stop: true, reason: "idle_diff", detail: `worktree unchanged for ${Math.round(idleForMs / 1000)}s` };
|
|
186
209
|
};
|
|
187
210
|
}
|
package/lib/openclaw-run.mjs
CHANGED
|
@@ -61,6 +61,35 @@ export function makeAbandonedBackgroundProcessTick(stateDir, { idleMs, minElapse
|
|
|
61
61
|
* tool call and files changed so far (liveProgress), and when. Never asks
|
|
62
62
|
* to stop; a failed read just skips that beat.
|
|
63
63
|
*/
|
|
64
|
+
// local_worker_stop (or `nomarmy jobs --stop`) writes this into the job's
|
|
65
|
+
// folder, so any session can stop any job; the next tick ends the worker.
|
|
66
|
+
export const STOP_REQUEST_FILE = "stop-request.json";
|
|
67
|
+
export function makeStopRequestTick(jobDir) {
|
|
68
|
+
return async () => (fs.existsSync(path.join(jobDir, STOP_REQUEST_FILE)) ? { stop: true, reason: "stopped" } : { stop: false });
|
|
69
|
+
}
|
|
70
|
+
/**
|
|
71
|
+
* Ask a running job to stop. Refuses what isn't running; never kills
|
|
72
|
+
* anything itself (the job's own tick does, within one tick).
|
|
73
|
+
* @returns {{ ok: boolean, message: string }}
|
|
74
|
+
*/
|
|
75
|
+
export function requestJobStop({ jobsRoot, jobId, reason = null, now = new Date() }) {
|
|
76
|
+
if (!/^[A-Za-z0-9._-]{1,120}$/.test(String(jobId ?? ""))) return { ok: false, message: "not a job id" };
|
|
77
|
+
const jobDir = path.join(jobsRoot, jobId);
|
|
78
|
+
let status = null;
|
|
79
|
+
try { status = JSON.parse(fs.readFileSync(path.join(jobDir, "status.json"), "utf8")); } catch { return { ok: false, message: `no job ${jobId} (see local_worker_jobs)` }; }
|
|
80
|
+
if (status.state !== "running") return { ok: false, message: `job ${jobId} isn't running (it's ${status.state ?? "unknown"})` };
|
|
81
|
+
if (status.phase && status.phase !== "worker" && status.phase !== "starting" && status.phase !== "worktree") {
|
|
82
|
+
return { ok: false, message: `job ${jobId} is past its worker (phase ${status.phase}); it's finishing on its own and spends no more model usage` };
|
|
83
|
+
}
|
|
84
|
+
if (fs.existsSync(path.join(jobDir, STOP_REQUEST_FILE))) return { ok: true, message: `a stop was already requested for ${jobId}` };
|
|
85
|
+
fs.writeFileSync(path.join(jobDir, STOP_REQUEST_FILE), JSON.stringify({ at: now.toISOString(), reason: reason ? String(reason).slice(0, 300) : null }));
|
|
86
|
+
return { ok: true, message: `stop requested for ${jobId}: its worker ends within about 15 seconds, verification is skipped, and the worktree is kept uncommitted, so continue_from can pick the work up (on another model too)` };
|
|
87
|
+
}
|
|
88
|
+
|
|
89
|
+
export function readStopRequest(jobDir) {
|
|
90
|
+
try { return JSON.parse(fs.readFileSync(path.join(jobDir, STOP_REQUEST_FILE), "utf8")); } catch { return null; }
|
|
91
|
+
}
|
|
92
|
+
|
|
64
93
|
export function makeHeartbeatTick(jobDir, liveProgress) {
|
|
65
94
|
// Never two beats at once: a slow beat used to overlap the next.
|
|
66
95
|
let busy = false;
|
|
@@ -475,8 +504,9 @@ export function createOpenClawRunner(deps) {
|
|
|
475
504
|
// The heartbeat runs for every job, so a running job is always watchable
|
|
476
505
|
// (status.json used to keep its launch-time updatedAt until the end).
|
|
477
506
|
const onTick = combineTicks([
|
|
507
|
+
makeStopRequestTick(jobDir),
|
|
478
508
|
makeHeartbeatTick(jobDir),
|
|
479
|
-
...(idleDiff ? [makeIdleDiffTick(cwd, idleDiff), makeAbandonedBackgroundProcessTick(stateDir, idleDiff)] : []),
|
|
509
|
+
...(idleDiff ? [makeIdleDiffTick(cwd, { ...idleDiff, activity: async () => { const t = await readOpenClawTranscriptTail(stateDir, { limit: 5 }); return t.available ? { events: t.events, toolInFlight: t.toolInFlight } : null; } }), makeAbandonedBackgroundProcessTick(stateDir, idleDiff)] : []),
|
|
480
510
|
]);
|
|
481
511
|
const liveLogs = { stdout: path.join(jobDir, `openclaw${logSuffix}.stdout.log`), stderr: path.join(jobDir, `openclaw${logSuffix}.stderr.log`) };
|
|
482
512
|
// Where this call's own transcript events will start (see salvageFinishedRun).
|
package/lib/process.mjs
CHANGED
|
@@ -73,7 +73,10 @@ export function createProcess(ctx) {
|
|
|
73
73
|
if (ticker) clearInterval(ticker);
|
|
74
74
|
child.kill("SIGTERM");
|
|
75
75
|
const error = new Error(message);
|
|
76
|
-
|
|
76
|
+
// A stop someone asked for (local_worker_stop) isn't a timeout: a
|
|
77
|
+
// timeout invites a report-recovery call, which would spend exactly
|
|
78
|
+
// the usage the stop was meant to save.
|
|
79
|
+
error.timedOut = stopReason !== "stopped";
|
|
77
80
|
if (stopReason) error.stopReason = stopReason;
|
|
78
81
|
reject(error);
|
|
79
82
|
};
|
package/lib/scout.mjs
CHANGED
|
@@ -438,9 +438,22 @@ export function isScoutReportUnusable(report) {
|
|
|
438
438
|
// The recovery call carries the question itself: a reply cut off mid-run can
|
|
439
439
|
// leave the resumed session without it, and a scout told only to "restate the
|
|
440
440
|
// question" then came back with an empty report (a live PM scout, twice).
|
|
441
|
-
|
|
441
|
+
// A recovery call can't count on seeing the first run's work: when OpenClaw's
|
|
442
|
+
// cleanup crashes ("state ownership retained"), the follow-up starts a fresh
|
|
443
|
+
// session, and a scout told "report only what you already found" came back
|
|
444
|
+
// with nothing after nine minutes of real research. So the recovery carries
|
|
445
|
+
// what the first run left: its cut-off reply and the files it read, which it
|
|
446
|
+
// may re-read to confirm line numbers, and nothing else.
|
|
447
|
+
const EARLIER_REPLY_CHARS = 12000;
|
|
448
|
+
export function scoutReportRecoveryPrompt({ report = { targetTokens: 600, hardCapTokens: 1024 }, question = null, acceptance = [], earlierReply = "", filesRead = [] } = {}) {
|
|
442
449
|
const asked = question ? `\n\nThe question you were answering:\n${String(question).trim()}${acceptance?.length ? `\n\nA complete answer covers:\n${acceptance.map((a) => `- ${a}`).join("\n")}` : ""}` : "";
|
|
443
|
-
|
|
450
|
+
const reply = String(earlierReply ?? "").trim();
|
|
451
|
+
const earlier = reply ? `\n\nYour earlier reply, as far as it got${reply.length > EARLIER_REPLY_CHARS ? " (its last part)" : ""}:\n${reply.length > EARLIER_REPLY_CHARS ? reply.slice(-EARLIER_REPLY_CHARS) : reply}` : "";
|
|
452
|
+
const files = filesRead?.length ? `\n\nFiles you read during that work: ${filesRead.slice(0, 60).join(", ")}` : "";
|
|
453
|
+
const tools = filesRead?.length
|
|
454
|
+
? "Do not explore further. You may re-read the files listed above to confirm exact line numbers for what you found; call no other tool."
|
|
455
|
+
: "Do not repeat, redo, retry, or explore further. Do not call any tool.";
|
|
456
|
+
return `Your previous reply ended without the required SCOUT REPORT, or was cut off before reaching END.${asked}${earlier}${files}\n\n${tools} Based on what you already found, reply with ONLY the report below, nothing before it, nothing after it:\n\nSCOUT REPORT\nQUESTION: <the question restated in one line>\nCONFIDENCE: high | medium | low\nFINDING: <one sentence> [src/example.js:10-24]\nNOT_FOUND: none | <what you looked for and could not find>\nEND\n\nIf you did not actually find anything worth a FINDING, say so under NOT_FOUND rather than inventing one. Target ${report.targetTokens} tokens; ${report.hardCapTokens} is the hard cap.`;
|
|
444
457
|
}
|
|
445
458
|
|
|
446
459
|
// ---------------------------------------------------------------------------
|
package/lib/transcript.mjs
CHANGED
|
@@ -99,6 +99,9 @@ export function summarizeTranscriptEvents(events) {
|
|
|
99
99
|
}
|
|
100
100
|
}
|
|
101
101
|
out.filesRead = [...new Set(out.filesRead)];
|
|
102
|
+
// A tool call with no result yet: the worker is waiting on it (a long test
|
|
103
|
+
// run writes nothing to the transcript until it returns).
|
|
104
|
+
out.toolInFlight = pending.length > 0;
|
|
102
105
|
return out;
|
|
103
106
|
}
|
|
104
107
|
|
package/mcp/server.mjs
CHANGED
|
@@ -36,6 +36,7 @@ import { readUsageSnapshots, usageStatus } from "../lib/usage-limits.mjs";
|
|
|
36
36
|
import { modelRefusals } from "../lib/health.mjs";
|
|
37
37
|
import { podmanProblem, podmanVmStartedAt } from "../lib/podman-health.mjs";
|
|
38
38
|
import { restartNotice } from "../lib/install-freshness.mjs";
|
|
39
|
+
import { requestJobStop } from "../lib/openclaw-run.mjs";
|
|
39
40
|
import { loadJobRecords, computeStats, formatStats, parseSince, resolveRepo } from "../lib/stats.mjs";
|
|
40
41
|
import { probeModel } from "../lib/model-probe.mjs";
|
|
41
42
|
import { jevSettings, judgeSettings } from "../lib/validators.mjs";
|
|
@@ -443,6 +444,14 @@ server.tool("stats", "What nomArmy's own job records show for this repository (o
|
|
|
443
444
|
} catch (error) { return toolText(error.message, true); }
|
|
444
445
|
});
|
|
445
446
|
|
|
447
|
+
server.tool("local_worker_stop", "Stop a running job's worker, for example one burning a frontier model's usage on the wrong track. Its worker ends within about 15 seconds, without the report-recovery call a timeout gets and without running verification; its worktree is kept uncommitted, so a new job with continue_from: <job_id> (and a cheaper model if you like) can finish the work. Works for a job started by any session. Refuses a job that isn't running or is already past its worker.", {
|
|
448
|
+
job_id: z.string().regex(/^[A-Za-z0-9._-]{1,120}$/),
|
|
449
|
+
reason: z.string().max(300).optional().describe("Why it's being stopped; recorded in the job's issues."),
|
|
450
|
+
}, async ({ job_id, reason }) => {
|
|
451
|
+
const r = requestJobStop({ jobsRoot, jobId: job_id, reason: reason ?? null });
|
|
452
|
+
return toolText(r.message, !r.ok);
|
|
453
|
+
});
|
|
454
|
+
|
|
446
455
|
server.tool("local_worker_capacity", "What this host can take right now: context per nom and the brief/report budgets derived from it, memory pressure and whether another job would be admitted, and the jobs currently running. Read-only.", {}, async () => {
|
|
447
456
|
await budgetState.refresh();
|
|
448
457
|
return toolText(JSON.stringify(withRestartNotice(capacitySnapshot()), null, 2));
|
package/package.json
CHANGED
|
@@ -3,7 +3,7 @@
|
|
|
3
3
|
"description": "Every byte verified: a harness for AI coding workers whose claims are never trusted. Your coding assistant stays in charge while workers implement and test in sandboxes, and nomArmy checks every change before it is committed.",
|
|
4
4
|
"author": "Rayson Technologies",
|
|
5
5
|
"license": "Apache-2.0",
|
|
6
|
-
"version": "0.1.0-alpha.
|
|
6
|
+
"version": "0.1.0-alpha.13",
|
|
7
7
|
"private": false,
|
|
8
8
|
"type": "module",
|
|
9
9
|
"engines": {
|
package/playbooks/feature.md
CHANGED
|
@@ -26,7 +26,7 @@ Stop and ask only for something irreversible or outside this repository: merging
|
|
|
26
26
|
|
|
27
27
|
## Watching jobs
|
|
28
28
|
|
|
29
|
-
Prefer `local_worker_start`. Right after starting a job, if your coordinator can run a background command, run `nomarmy jobs --wait <job_id>` in the background so you're told the moment it finishes and can tell the operator. Otherwise poll `local_worker_status` with the longest `wait_seconds` it allows. `nomarmy jobs --events`
|
|
29
|
+
Prefer `local_worker_start`. Right after starting a job, if your coordinator can run a background command, run `nomarmy jobs --wait <job_id>` in the background so you're told the moment it finishes and can tell the operator. Otherwise poll `local_worker_status` with the longest `wait_seconds` it allows. For several jobs at once, `nomarmy jobs --events --until-done` exits when every one has finished. The plain `nomarmy jobs --events` stream never exits on its own while jobs run, so don't run it as a background command (you'd only hear when it exits); use it only with a monitor that wakes on each line. `run_status` lists the run's running jobs as well as finished ones. The operator gets a desktop notification whenever a job finishes and whenever the run crosses a limit, so you don't need to relay each one.
|
|
30
30
|
|
|
31
31
|
## Limits
|
|
32
32
|
|