nomarmy 0.1.0-alpha.12 → 0.1.0-alpha.14

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -52,6 +52,8 @@ Failing verification stays failed, unconditionally. A malformed report isn't aut
52
52
 
53
53
  **Checking without building** costs nothing: `mode: verify` runs a verification profile against any branch, with no worker and no model tokens.
54
54
 
55
+ **Want deeper checks?** Three optional [validators](https://github.com/rayson-tech/nomarmy/blob/main/docs/validators.md) go further, each only adding review flags: mutation testing (do the tests pin down the changed lines?), Jev (do a scout's citations support its findings, does a report match its diff?) and a model judge (acceptance criteria, weakened tests).
56
+
55
57
  ## Where the work runs
56
58
 
57
59
  - **Agents** say where a job can run: an API key, your own ChatGPT or Muse Code subscription, or a local model on llama.cpp.
@@ -74,10 +76,12 @@ Developed and maintained by Rayson Technologies. This is an alpha (`0.1.0-alpha`
74
76
  | [Agents and the army](https://github.com/rayson-tech/nomarmy/blob/main/docs/agents-and-army.md) | Where a job can run, who does what, usage limits, picking an agent |
75
77
  | [`/feature` runs](https://github.com/rayson-tech/nomarmy/blob/main/docs/feature-runs.md) | A feature end to end, and watching what nomArmy is doing |
76
78
  | [Your repository](https://github.com/rayson-tech/nomarmy/blob/main/docs/your-repo.md) | `.nomarmy.yml`, verification, dependencies, private registries, what nomArmy checks |
79
+ | [Validators](https://github.com/rayson-tech/nomarmy/blob/main/docs/validators.md) | Optional deeper checks: mutation testing, Jev, a model judge |
77
80
  | [Harnesses](https://github.com/rayson-tech/nomarmy/blob/main/docs/harnesses.md) | Ecosystem registry, detection, network levels, and requirements |
78
81
  | [Configuration](https://github.com/rayson-tech/nomarmy/blob/main/docs/configuration.md) | Settings, swapping the local model, sizing, admission |
79
82
  | [Reference](https://github.com/rayson-tech/nomarmy/blob/main/docs/reference.md) | Every CLI command and MCP tool |
80
83
  | [Security posture](https://github.com/rayson-tech/nomarmy/blob/main/docs/security.md) | What the sandbox holds back, and the one exception |
84
+ | [FAQ](https://github.com/rayson-tech/nomarmy/blob/main/docs/faq.md) | Which model for which role, switching models, usage limits |
81
85
  | [Troubleshooting](https://github.com/rayson-tech/nomarmy/blob/main/docs/troubleshooting.md) | Symptoms and fixes |
82
86
 
83
87
  ## Security
@@ -97,6 +101,8 @@ A nom gets a writable git worktree inside a Podman sandbox and nothing else: no
97
101
  | The army and `/feature` | Driven by a real Claude Code General across three runs, about 18 implement jobs |
98
102
  | Harnesses: Go, Rust, Python, Node and mixed repos; Playwright; fake services | Live-verified offline |
99
103
  | Private registries and a verification-only network allowlist | Live-verified; each passed an independent security review |
104
+ | Validators: mutation testing, Jev, a model judge | Unit and live tested; Jev and the judge evaluated on real job records |
105
+ | `nomarmy stats` | Checked against a hand-built report on real job records |
100
106
 
101
107
  What we've learned from real runs, including where delegating pays and where it doesn't, is in [docs/findings.md](https://github.com/rayson-tech/nomarmy/blob/main/docs/findings.md).
102
108
 
package/bin/nomarmy.mjs CHANGED
@@ -20,7 +20,8 @@ import { readGGUFMetadata, resolveModelPath, totalSplitBytes } from "../lib/gguf
20
20
  import { recommend, customRecommendation, evaluateConfig, bytesPerKvElementForCacheTypes, MIN_CONTEXT_PER_NOM } from "../lib/sizing.mjs";
21
21
  import { connectClaude, connectCodex, connectCursor, cursorAlreadyConnected, deriveWorkerModelEnv, defaultInstallDir, installMcpCopy, SCOPES, claudeUserScoped, portableServerLaunch } from "../lib/connect.mjs";
22
22
  import { compareVersions, readPackageVersion, readInstallVersions, copyIsStale } from "../lib/install-freshness.mjs";
23
- import { loadJobRecords, computeStats, formatStats, parseSince, resolveRepo } from "../lib/stats.mjs";
23
+ import { loadJobRecords, computeStats, formatStats, parseSince, resolveRepo, agentLookup } from "../lib/stats.mjs";
24
+ import { requestJobStop } from "../lib/openclaw-run.mjs";
24
25
  import { loadValidators, saveJevKey, removeJev, jevSettings, askJev, validatorsPath, JEV_CHECKS, saveJudge, removeJudge, judgeSettings } from "../lib/validators.mjs";
25
26
  import { probeModel } from "../lib/model-probe.mjs";
26
27
  import { ID_RE, AUTH_ENV_NAME_RE, OPENCLAW_PROVIDER_ID_RE, openclawProviderId, isNativeProviderType } from "../lib/dispatch-schema.mjs";
@@ -198,17 +199,25 @@ Usage: nomarmy <command> [options]
198
199
  which agent the General is, in --global
199
200
  (default) or --local
200
201
  config paths where agents.yml and the three army layers live
201
- jobs [--watch|--events|--prune|--wait <jobId>] [--interval N] [--older-than DAYS]
202
+ jobs [--watch|--events [--until-done]|--prune|--wait <jobId>|--stop <jobId> [--reason <text>]] [--interval N] [--older-than DAYS]
202
203
  what's running across every session (agent, model, phase,
203
204
  last tool call, files changed, heartbeat) and what just
204
205
  finished; --watch redraws every N seconds (default 3);
205
206
  --events prints one line per start, phase change and
206
- finish (for Claude Code's background monitor; --json for
207
- JSON lines); --prune removes the bulky runtime data
207
+ finish (--json for JSON lines). It's a stream: read it
208
+ with a monitor that wakes on each line. A background
209
+ command is only reported when it exits, so there use
210
+ --events --until-done, which exits once every job it saw
211
+ running has finished (or --wait for one job). The plain
212
+ stream ends on its own after 30 minutes with nothing
213
+ running (--idle-minutes N); --prune removes the bulky runtime data
208
214
  from finished jobs older than DAYS (default 2), keeping
209
215
  their records, reports and any retained worktree;
210
216
  --wait <jobId> [--timeout <seconds>] blocks for one job
211
- to finish (default timeout 1800; --json is supported)
217
+ to finish (default timeout 1800; --json is supported);
218
+ --stop <jobId> stops a running job's worker (no report
219
+ recovery, no verification), keeping its worktree for
220
+ continue_from
212
221
  health check what's likely to break a run before it does:
213
222
  expiring logins, an outdated OpenClaw or plugin, roles
214
223
  that can't be dispatched, an unloadable agents.yml,
@@ -2518,6 +2527,12 @@ function renderJobs({ running, recent }) {
2518
2527
  */
2519
2528
  async function streamJobEvents() {
2520
2529
  const interval = Math.max(1, Number(value("interval", "3")) || 3) * 1000;
2530
+ // A stream nobody reads must still end: a General that ran this as a
2531
+ // background command (reported only on exit) was never told jobs had
2532
+ // finished, and eight of these streams were left running for days.
2533
+ const untilDone = flag("until-done");
2534
+ const idleLimitMs = Math.max(1, Number(value("idle-minutes", "30")) || 30) * 60000;
2535
+ let idleSinceMs = Date.now(), sawRunning = false;
2521
2536
  const seen = new Map();
2522
2537
  const emit = (event, job, detail = "") => {
2523
2538
  if (json) console.log(JSON.stringify({ at: new Date().toISOString(), event, jobId: job.jobId, agent: job.agent, model: job.model, phase: job.phase, detail }));
@@ -2540,10 +2555,23 @@ async function streamJobEvents() {
2540
2555
  seen.clear();
2541
2556
  for (const [id, j] of now) seen.set(id, j);
2542
2557
  first = false;
2558
+ if (running.length) { sawRunning = true; idleSinceMs = Date.now(); }
2559
+ else if (untilDone && sawRunning) {
2560
+ if (json) console.log(JSON.stringify({ at: new Date().toISOString(), event: "done", detail: "every job seen running has finished" }));
2561
+ else console.log(`${new Date().toLocaleTimeString()} done every job seen running has finished`);
2562
+ return;
2563
+ } else if (Date.now() - idleSinceMs >= (untilDone ? Math.min(idleLimitMs, 120000) : idleLimitMs)) {
2564
+ const why = untilDone ? "no job was running to wait for" : `nothing has run for ${Math.round(idleLimitMs / 60000)} minutes`;
2565
+ if (json) console.log(JSON.stringify({ at: new Date().toISOString(), event: "idle", detail: why }));
2566
+ else console.log(`${new Date().toLocaleTimeString()} idle ${why}; exiting`);
2567
+ return;
2568
+ }
2543
2569
  await new Promise((r) => setTimeout(r, interval));
2544
2570
  }
2545
2571
  }
2546
2572
 
2573
+ const commitSha = (commit) => (typeof commit === "string" ? commit : typeof commit?.sha === "string" ? commit.sha : null);
2574
+
2547
2575
  /** Wait for one job in the shared, cross-session state directory. */
2548
2576
  async function waitForJobCli() {
2549
2577
  const requested = value("wait");
@@ -2580,7 +2608,8 @@ async function waitForJobCli() {
2580
2608
  outcome: meta.outcome ?? status.outcome ?? null,
2581
2609
  coordinatorStatus: meta.coordinatorStatus ?? status.coordinatorStatus ?? null,
2582
2610
  branch: meta.branch ?? status.branch ?? null,
2583
- commit: meta.commit?.sha ?? meta.commit ?? status.commit?.sha ?? status.commit ?? null,
2611
+ // A job that made no commit has commit: { created: false, sha: null }; only a sha is a commit.
2612
+ commit: commitSha(meta.commit) ?? commitSha(status.commit),
2584
2613
  issues,
2585
2614
  };
2586
2615
  if (json) out(result);
@@ -2621,6 +2650,13 @@ function pruneJobRuntimeCli() {
2621
2650
 
2622
2651
  async function cmdJobs() {
2623
2652
  if (flag("wait")) return waitForJobCli();
2653
+ if (flag("stop")) {
2654
+ const r = requestJobStop({ jobsRoot: jobsRootDir(), jobId: value("stop"), reason: value("reason") });
2655
+ if (json) return out(r);
2656
+ console.log(r.ok ? c.green(`✓ ${r.message}`) : c.red(`✗ ${r.message}`));
2657
+ if (!r.ok) process.exitCode = 1;
2658
+ return;
2659
+ }
2624
2660
  if (flag("events")) return streamJobEvents();
2625
2661
  if (flag("prune")) return pruneJobRuntimeCli();
2626
2662
  if (json) return out(collectJobs());
@@ -2676,7 +2712,9 @@ function cmdStats() {
2676
2712
  try { repo = execFileSync("git", ["rev-parse", "--show-toplevel"], { cwd: repoDir, encoding: "utf8", stdio: ["ignore", "pipe", "ignore"] }).trim(); }
2677
2713
  catch { throw new Error(`${repoDir} isn't inside a git repository; run nomarmy stats from one, or pass --repo <name> or --all-repos`); }
2678
2714
  }
2679
- const stats = computeStats(records, { repo, sinceMs: parseSince(value("since")), untilMs: parseSince(value("until")), role: value("role"), model: value("model") });
2715
+ let agentFor = () => null;
2716
+ try { agentFor = agentLookup(loadAgents(globalConfigDir()).agents, agentProviderId); } catch { /* no agents.yml: commands name <agent> */ }
2717
+ const stats = computeStats(records, { repo, sinceMs: parseSince(value("since")), untilMs: parseSince(value("until")), role: value("role"), model: value("model"), agentFor });
2680
2718
  if (json) return out(stats);
2681
2719
  console.log(formatStats(stats));
2682
2720
  }
package/lib/admission.mjs CHANGED
@@ -236,6 +236,15 @@ export function createJobRuntime(deps) {
236
236
  // verify_regression re-runs `verification`; with no profile set there is
237
237
  // nothing to re-run. Refuse before starting anything, matching every other
238
238
  // admission check here, rather than silently no-op at runtime.
239
+ // stakes applies to what gets committed; reviews names the job a scout reviews.
240
+ jobs.forEach((j, i) => {
241
+ const at = jobs.length > 1 ? `job ${i + 1}: ` : "";
242
+ if (j.stakes === "high" && (j.mode ?? "implement") !== "implement") problems.push(`${at}stakes: high applies to implement jobs; for a review, send a scout with reviews: <job id>`);
243
+ if (j.reviews) {
244
+ if (j.mode !== "scout") problems.push(`${at}reviews: <job id> marks a scout as a review of that job; this job is ${j.mode ?? "implement"}`);
245
+ else if (!fs.existsSync(path.join(jobsRoot, j.reviews, "metadata.json"))) problems.push(`${at}reviews: no finished job ${j.reviews} to review (see local_worker_jobs)`);
246
+ }
247
+ });
239
248
  // continue_from: the retained job must exist, be unfinished and uncommitted,
240
249
  // and belong to this repo; see lib/continue-from.mjs.
241
250
  jobs.forEach((j, i) => {
@@ -14,13 +14,14 @@ Before dispatching:
14
14
  - Agents' usage limits show in army and local_worker_capacity; a job on an agent at its limit is held. Ask the operator before resubmitting with confirm_over_limit: true, or move the job to another agent. Never set it on your own.
15
15
  - A Claude subscription agent (claude-cli) runs its tools on this machine, outside the sandbox: use it for scout and review work. nomArmy refuses implement jobs on it unless the operator set allow_host_tools; send build work to a sandboxed agent.
16
16
  - To run tests without changing anything, use mode: verify; it costs no model usage.
17
+ - Mark an implement job stakes: high when a mistake would be costly (security, access control, personal or tenant data, data loss, money, irreversible changes), however small it is. It then needs a verification profile, keeps the revert check, and needs an independent review before you accept it: a scout on another vendor with reviews: <job id>.
17
18
  - Brief outcomes, not edits: a task, explicit acceptance criteria, and the tests that prove it. Put facts you've already resolved in evidence.
18
- - Prefer local_worker_start for anything longer than a few minutes. Right after, if you can run a background command, run \`nomarmy jobs --wait <job_id>\` in the background so you're told the moment it finishes and can tell the operator; otherwise poll local_worker_status with wait_seconds. For a whole feature, use /feature (run_start keeps a run's jobs, spend and hours bounded).
19
+ - Prefer local_worker_start for anything longer than a few minutes. Right after, if you can run a background command, run \`nomarmy jobs --wait <job_id>\` in the background so you're told the moment it finishes and can tell the operator (for several jobs, \`nomarmy jobs --events --until-done\`, which exits when they've all finished); otherwise poll local_worker_status with wait_seconds. Never run the plain \`nomarmy jobs --events\` stream as a background command: it only reports when it exits, so you'd never hear; it's for a monitor that reads each line. For a whole feature, use /feature (run_start keeps a run's jobs, spend and hours bounded).
19
20
 
20
21
  Trust boundary:
21
22
  - A worker's four-line report is a claim; nomArmy's verified git record and independent verification are the evidence. A job isn't complete if its report is missing or malformed, its STATUS is partial or blocked, STATUS done lacks VERIFICATION pass, or its changes aren't committed by nomArmy.
22
23
  - Read the diff of anything material before integrating it. nomArmy commits on the worker's branch and never merges into yours: integration, conflicts and pushes are yours.
23
- - To finish a job that came back partial, blocked or failing verification, dispatch the correction with continue_from: <that job id>. The new job starts with its unfinished work in place and verifies the whole. Don't fix and commit a worker's files yourself: that lands them unverified. If you must, commit them on a branch, run mode: verify on it before building on it, and say so in your report.
24
+ - A job on the wrong track can be stopped with local_worker_stop; it keeps its worktree. To finish a job that came back partial, blocked, failing verification or stopped, dispatch the correction with continue_from: <that job id> (a cheaper model is fine). The new job starts with its unfinished work in place and verifies the whole. Don't fix and commit a worker's files yourself: that lands them unverified. If you must, commit them on a branch, run mode: verify on it before building on it, and say so in your report.
24
25
  - Failed and incomplete worktrees are kept for review; clean up with local_worker_cleanup or local_worker_sweep once you've decided.
25
26
 
26
27
  Never delegate deployments, production access, cloud or SSH credentials, secrets, Terraform state or kubectl contexts to a worker.`;
package/lib/execute.mjs CHANGED
@@ -1,4 +1,4 @@
1
- import { writeStatus, shouldRetryTransientAbort, shouldAttemptScoutRecovery } from "./openclaw-run.mjs";
1
+ import { writeStatus, shouldRetryTransientAbort, shouldAttemptScoutRecovery, readStopRequest } from "./openclaw-run.mjs";
2
2
  import fs from "node:fs";
3
3
  import path from "node:path";
4
4
  import { fileURLToPath } from "node:url";
@@ -12,7 +12,7 @@ import { continuationProblem, continuationBase, snapshotRetainedWork, applyRetai
12
12
  import { checkScoutCitations, checkReportClaims } from "./jev-checks.mjs";
13
13
  import { runJudge } from "./judge.mjs";
14
14
  import { pickMutants, runMutants, describeSurvivors } from "./mutation.mjs";
15
- import { estimateDisplacement } from "./transcript.mjs";
15
+ import { estimateDisplacement, readOpenClawTranscript } from "./transcript.mjs";
16
16
  import { outlineFile, findReferences } from "./repo-query.mjs";
17
17
  import { loadConfig } from "./config.mjs";
18
18
  import { linkNodePackages, nodeModulesState, repairHostInstalls } from "./sandbox-images.mjs";
@@ -20,7 +20,7 @@ import { detectTestSabotage, addedLinesOf, loadDependencyNames } from "./sabotag
20
20
  import { describeRecoveryChanges, reportRecoveryPrompt } from "./worker-prompt.mjs";
21
21
  import { parseWorkerReport } from "./report.mjs";
22
22
  import { OUTCOMES, COORDINATOR_STATUS_BY_OUTCOME } from "./outcomes.mjs";
23
- import { resolveOutcome, finalText, workerMetadata, applyRefactorContract, applyVerificationPolicy } from "./outcome.mjs";
23
+ import { resolveOutcome, finalText, workerMetadata, applyRefactorContract, applyVerificationPolicy, HIGH_STAKES_NOTE } from "./outcome.mjs";
24
24
  import { parseAddedLineNumbers, isTestPath, isDocumentationPath, detectScopedTestSelectionRisk, detectUnwiredNewDefinitions, detectMislabeledTestNames, detectPossibleSecrets, detectVerificationInputChanges } from "./diff-checks.mjs";
25
25
 
26
26
  // ---------------------------------------------------------------------------
@@ -42,7 +42,7 @@ export function createExecutor(deps) {
42
42
  // Optional Jev checks (lib/validators.mjs): the server wires the settings; without them (tests), none run.
43
43
  const jevSettingsFor = (check) => { try { const s = deps.jevSettings?.(); return s?.checks?.includes(check) ? s : null; } catch { return null; } };
44
44
 
45
- async function executeJob({ task, acceptance, verification, mode = "implement", baseRef, timeoutSeconds = 600, profile = "coder", reasoning = "high", pool = null, subscriptionWorker = null, onBehalfOf = null, model = null, reportSize = null, workerId, evidence = null, verifyRegression = false, commitSubject = null, refactor = false, continueFrom = null, jobId: presetJobId = null }) {
45
+ async function executeJob({ task, acceptance, verification, mode = "implement", baseRef, timeoutSeconds = 600, profile = "coder", reasoning = "high", pool = null, subscriptionWorker = null, onBehalfOf = null, model = null, reportSize = null, workerId, evidence = null, verifyRegression = false, commitSubject = null, refactor = false, continueFrom = null, stakes = null, reviews = null, jobId: presetJobId = null }) {
46
46
  await assertRepo();
47
47
  // A continuation starts from the retained job's own base commit.
48
48
  let continuation = null;
@@ -66,9 +66,9 @@ export function createExecutor(deps) {
66
66
  progress("starting", { startedAt: new Date().toISOString(), agent: mode === "verify" ? null : pool ?? subscriptionWorker ?? "local", model: model ?? null });
67
67
  const common = { task, acceptance, base, jobId, jobDir, runtimeDir, timeoutSeconds, profile, reasoning, pool, subscriptionWorker, onBehalfOf, model, reportSize, workerId, progress, jobStartedMs };
68
68
  if (mode === "verify") return executeVerify({ ...common, verification });
69
- if (mode === "scout") return executeScout(common);
69
+ if (mode === "scout") return executeScout({ ...common, reviews });
70
70
  if (mode === "decompose") return executeDecompose(common);
71
- return executeImplement({ ...common, verification, evidence, verifyRegression, commitSubject, refactor, continuation });
71
+ return executeImplement({ ...common, verification, evidence, verifyRegression, commitSubject, refactor, continuation, stakes });
72
72
  }
73
73
 
74
74
  async function executeVerify({ task, verification: profile, base, jobId, jobDir, workerId, progress, jobStartedMs }) {
@@ -103,7 +103,7 @@ export function createExecutor(deps) {
103
103
  return { ok: outcome === OUTCOMES.VERIFIED, manifest, jobDir, report: "" };
104
104
  }
105
105
 
106
- async function executeImplement({ task, acceptance, verification, base, jobId, jobDir, runtimeDir, timeoutSeconds, profile, reasoning, pool = null, subscriptionWorker = null, onBehalfOf = null, model = null, reportSize = null, workerId, evidence, verifyRegression = false, commitSubject = null, refactor = false, continuation = null, progress, jobStartedMs }) {
106
+ async function executeImplement({ task, acceptance, verification, base, jobId, jobDir, runtimeDir, timeoutSeconds, profile, reasoning, pool = null, subscriptionWorker = null, onBehalfOf = null, model = null, reportSize = null, workerId, evidence, verifyRegression = false, commitSubject = null, refactor = false, continuation = null, stakes = null, progress, jobStartedMs }) {
107
107
  const mode = "implement";
108
108
  let branch = `agent/${jobId}`, worktree = path.join(jobDir, "worktree");
109
109
  try {
@@ -275,6 +275,9 @@ export function createExecutor(deps) {
275
275
  // failed job that changed nothing was recorded "pass" (a Senti run),
276
276
  // which reads as evidence about work that never happened.
277
277
  independentVerification = normalizeVerification({ status: "not_run", basis: "not-applicable", reason: "the worker changed nothing, so there was none of its work to verify" }, verification ?? null);
278
+ } else if (workerStopReason === "stopped") {
279
+ // Stopped on request: end now, without spending time on tests.
280
+ independentVerification = normalizeVerification({ status: "not_run", basis: "not-applicable", reason: "the job was stopped on request" }, verification ?? null);
278
281
  } else if (verificationFlow.verificationRunner || !reportValidation.valid) {
279
282
  independentVerification = await runIndependentVerification({ profile: verification ?? null, cwd, jobId, baseSha: base.sha, branch, mode, record: preCommit, logFile: path.join(jobDir, "verification.log") });
280
283
  }
@@ -522,7 +525,8 @@ export function createExecutor(deps) {
522
525
  : mutation?.status === "survivors"
523
526
  ? { ...afterConfig, reviewRequired: true, reasons: [...afterConfig.reasons, describeSurvivors(mutation, verification)] }
524
527
  : afterConfig;
525
- const finalOutcome = applyRefactorContract(applyVerificationPolicy(afterMutation, independentVerification.status, repoPolicy()),
528
+ const afterStakes = stakes === "high" ? { ...afterMutation, reviewRequired: true, reasons: [...afterMutation.reasons, HIGH_STAKES_NOTE] } : afterMutation;
529
+ const finalOutcome = applyRefactorContract(applyVerificationPolicy(afterStakes, independentVerification.status, repoPolicy()),
526
530
  { refactor, verificationStatus: independentVerification.status, testChanges: preCommit.testChanges });
527
531
 
528
532
  progress("commit");
@@ -533,7 +537,10 @@ export function createExecutor(deps) {
533
537
 
534
538
  let coordinatorStatus = COORDINATOR_STATUS_BY_OUTCOME[finalOutcome.outcome] ?? "incomplete";
535
539
  const issues = [...finalOutcome.reasons, ...(preCommit.issues ?? [])];
536
- if (workerError) issues.push(`worker error: ${String(workerError).split("\n")[0]}`);
540
+ if (workerStopReason === "stopped") {
541
+ const request = readStopRequest(jobDir);
542
+ issues.push(`stopped on request${request?.reason ? `: ${request.reason}` : ""}; the worktree is kept, so continue_from can pick the work up (on another model too)`);
543
+ } else if (workerError) issues.push(`worker error: ${String(workerError).split("\n")[0]}`);
537
544
  if ((result ?? attempted)?.salvaged) issues.push(`runner cleanup failed after the run (${(result ?? attempted).salvagedFrom}); the worker's report was recovered from the run's transcript`);
538
545
  if (jevClaims?.error) issues.push(`Jev report check skipped (${jevClaims.error}); this job's result doesn't depend on it`);
539
546
  if (judged?.error) issues.push(`Judge ${judged.skipped ? "skipped" : "didn't answer"} (${judged.error}); this job's result doesn't depend on it`);
@@ -547,7 +554,7 @@ export function createExecutor(deps) {
547
554
  // the diffstat right in the issue a caller actually reads -- not just
548
555
  // buried in the full manifest's git record -- is what makes "go look at
549
556
  // the worktree" worth doing instead of discarding the job.
550
- issues.push(`repository changes remain uncommitted (${record.filesChanged} file(s), +${record.additions}/-${record.deletions}): ${commit.reason}`);
557
+ if (record.filesChanged > 0) issues.push(`repository changes remain uncommitted (${record.filesChanged} file(s), +${record.additions}/-${record.deletions}): ${commit.reason}`);
551
558
  }
552
559
  const failures = worker.toolSummary?.failures ?? 0; if (failures > 0) issues.push(`worker recorded ${failures} tool failure(s)`);
553
560
  if (record.ignoredRuntimeJunk.length) issues.push(`runtime junk ignored: ${record.ignoredRuntimeJunk.join(", ")}`);
@@ -565,9 +572,9 @@ export function createExecutor(deps) {
565
572
 
566
573
  const metrics = buildMetrics({ result: result ?? attempted, record, reportValidation, outcome: finalOutcome, workerElapsedMs, totalElapsedMs: Date.now() - jobStartedMs, regressionCheckElapsedMs, transientAbortRetried });
567
574
  const manifest = { version: VERSION, jobId, workerId: workerId || jobId, mode, projectDir, worktree, branch, startedAt, finishedAt,
568
- objective: task, acceptance: acceptance ?? [], verificationProfile: verification ?? null, ...(continuedFrom ? { continuedFrom } : {}),
575
+ objective: task, acceptance: acceptance ?? [], verificationProfile: verification ?? null, ...(continuedFrom ? { continuedFrom } : {}), ...(stakes ? { stakes } : {}),
569
576
  ...(mutation ? { mutation: { ...mutation, elapsedSeconds: Math.round((mutationElapsedMs ?? 0) / 1000) } } : {}),
570
- ...(jevClaims || judged ? { validators: { ...(jevClaims ? { jev: { check: "report-claims", verdict: jevClaims.verdict, flagged: Boolean(jevClaims.flag), error: jevClaims.error, truncated: Boolean(jevClaims.truncated), inputTokens: jevClaims.usage } } : {}), ...(judged ? { judge: { agent: judge?.agent ?? null, model: judge?.model ?? null, answer: judged.answer, flags: judged.flags, error: judged.error } } : {}) } } : {}),
577
+ ...(jevClaims || judged ? { validators: { ...(jevClaims ? { jev: { check: "report-claims", verdict: jevClaims.verdict, flagged: Boolean(jevClaims.flag), error: jevClaims.error, truncated: Boolean(jevClaims.truncated), inputTokens: jevClaims.usage } } : {}), ...(judged ? { judge: { agent: judge?.agent ?? null, provider: judge?.provider ?? null, model: judge?.model ?? null, answer: judged.answer, flags: judged.flags, error: judged.error } } : {}) } } : {}),
571
578
  outcome: finalOutcome.outcome, recovered: finalOutcome.recovered, recoveryAttempted: finalOutcome.recoveryAttempted,
572
579
  reportRecoveryAttempted, reportRecovered,
573
580
  reviewRequired: finalOutcome.reviewRequired || record.testChanges.reviewRequired,
@@ -610,7 +617,7 @@ export function createExecutor(deps) {
610
617
  // the worktree, so a scout that wrote to its snapshot cannot forge evidence.
611
618
  // A clean scout worktree holds no work and is removed; a dirty one is retained
612
619
  // because a scout that wrote is a scout that misbehaved, and that is worth a look.
613
- async function executeScout({ task, acceptance, base, jobId, jobDir, runtimeDir, timeoutSeconds, profile, reasoning, pool = null, subscriptionWorker = null, onBehalfOf = null, model = null, reportSize = null, workerId, progress, jobStartedMs }) {
620
+ async function executeScout({ task, acceptance, base, jobId, jobDir, runtimeDir, timeoutSeconds, profile, reasoning, pool = null, subscriptionWorker = null, onBehalfOf = null, model = null, reportSize = null, workerId, progress, jobStartedMs, reviews = null }) {
614
621
  const mode = "scout", worktree = path.join(jobDir, "worktree");
615
622
  let worktreeRetained = false;
616
623
  try {
@@ -669,12 +676,18 @@ export function createExecutor(deps) {
669
676
  const remainingSeconds = timeoutSeconds - Math.round(workerElapsedMs / 1000);
670
677
  if (shouldAttemptScoutRecovery({ workerFailed, workerTimedOut, report, remainingSeconds })) {
671
678
  reportRecoveryAttempted = true;
679
+ // What the first run left, in case the follow-up can't see its session.
680
+ let filesRead = [], earlierReply = reportText;
681
+ try {
682
+ const first = await readOpenClawTranscript(path.join(runtimeDir, "state"));
683
+ if (first.available) { filesRead = first.filesRead ?? []; earlierReply = earlierReply || first.lastAssistantText || ""; }
684
+ } catch { /* the reply alone still helps */ }
672
685
  try {
673
686
  const recoveryResult = await runOpenClaw({
674
687
  task, acceptance, verification: null, mode, cwd: worktree, baseRef: base.ref, baseSha: base.sha,
675
688
  timeoutSeconds: remainingSeconds, runtimeDir, profile, reasoning, pool, subscriptionWorker, onBehalfOf, model, reportSize, jobDir, workerId: workerId || jobId,
676
689
  evidenceTool: evidencePlaced ? evidenceTool : null,
677
- overridePrompt: scoutReportRecoveryPrompt({ report: used.report.scout, question: task, acceptance }), logSuffix: "-recovery",
690
+ overridePrompt: scoutReportRecoveryPrompt({ report: used.report.scout, question: task, acceptance, earlierReply, filesRead }), logSuffix: "-recovery",
678
691
  });
679
692
  const recoveryReport = parseScoutReport(finalText(recoveryResult), (recoveryResult?.budgetsUsed ?? used).scout);
680
693
  if (!isScoutReportUnusable(recoveryReport)) {
@@ -753,7 +766,7 @@ export function createExecutor(deps) {
753
766
  displacement_verdict: displacement.verdict
754
767
  };
755
768
  const manifest = { version: VERSION, jobId, workerId: workerId || jobId, mode, projectDir, worktree: worktreeRetained ? worktree : null, branch: null, baseSha: base.sha, startedAt, finishedAt,
756
- objective: task, mustCover: acceptance ?? [],
769
+ objective: task, mustCover: acceptance ?? [], ...(reviews ? { reviews } : {}),
757
770
  ...(jevCitations ? { validators: { jev: { check: "scout-citations", checked: jevCitations.checked, flags: jevCitations.flags, errors: jevCitations.errors, inputTokens: jevCitations.usage } } } : {}),
758
771
  outcome: outcome.outcome, coordinatorStatus: outcome.coordinatorStatus, reviewRequired: outcome.reviewRequired, issues,
759
772
  scout: { question: report.question, confidence: report.confidence, notFound: report.notFound,
@@ -139,9 +139,26 @@ export function createGitRecord({ run, git, gitRaw }) {
139
139
  // idleMinElapsedMs of the work phase has passed (an early snapshot mid-first-
140
140
  // edit looks identical to no edit at all). A worktree read failing mid-write
141
141
  // is expected, not an error; it just means "nothing to report this tick."
142
- function makeIdleDiffTick(cwd, { idleMs, minElapsedMs }) {
143
- let lastHash = null, lastChangeAtMs = 0, sawChange = false;
142
+ //
143
+ // An unchanged worktree isn't enough on its own: a frontier worker running
144
+ // a suite of thousands of tests, or reading before its next edit, changes
145
+ // no file for minutes. Three Senti jobs were cut off at 7 to 13 minutes
146
+ // while still working. With `activity` (the worker's transcript: event
147
+ // count and whether a tool call is still running), the breaker waits while
148
+ // the worker is active, and stops only when it has also gone quiet for
149
+ // idleMs, or when nothing has changed for ACTIVE_CAP times idleMs however
150
+ // busy it looks (a worker looping without progress).
151
+ function makeIdleDiffTick(cwd, { idleMs, minElapsedMs, activity = null }) {
152
+ const ACTIVE_CAP = 3;
153
+ let lastHash = null, lastChangeAtMs = 0, sawChange = false, lastEvents = null, lastActivityAtMs = 0, inFlight = false;
144
154
  return async elapsedMs => {
155
+ if (activity) {
156
+ try {
157
+ const a = await activity();
158
+ if (a && a.events !== lastEvents) { lastEvents = a.events; lastActivityAtMs = elapsedMs; }
159
+ inFlight = Boolean(a?.toolInFlight);
160
+ } catch { /* a failed read just means no activity signal this tick */ }
161
+ }
145
162
  let statusOut;
146
163
  try { statusOut = await gitRaw(["status", "--porcelain=v1", "-z", "--untracked-files=all"], cwd); }
147
164
  catch { return { stop: false }; }
@@ -182,6 +199,12 @@ export function createGitRecord({ run, git, gitRaw }) {
182
199
  if (!sawChange || elapsedMs < minElapsedMs) return { stop: false };
183
200
  const idleForMs = elapsedMs - lastChangeAtMs;
184
201
  if (idleForMs < idleMs) return { stop: false };
202
+ if (activity) {
203
+ const quietForMs = elapsedMs - lastActivityAtMs;
204
+ const active = inFlight || quietForMs < idleMs;
205
+ if (active && idleForMs < idleMs * ACTIVE_CAP) return { stop: false };
206
+ if (active) return { stop: true, reason: "idle_diff", detail: `worktree unchanged for ${Math.round(idleForMs / 1000)}s while the worker kept working (${ACTIVE_CAP}x the idle limit)` };
207
+ }
185
208
  return { stop: true, reason: "idle_diff", detail: `worktree unchanged for ${Math.round(idleForMs / 1000)}s` };
186
209
  };
187
210
  }
package/lib/health.mjs CHANGED
@@ -333,6 +333,14 @@ export async function checkAndRecordHealth({ projectDir, stateRoot, configDir, n
333
333
  const install = { ...readInstallVersions(installDir), copyHarnesses: Object.keys(loadHarnesses(path.join(installDir, "harnesses")).harnesses).length };
334
334
  const result = await runHealthChecks({ now, mode, armySummary, agentsError, jobsRoot: path.join(stateRoot, "jobs"), pidAlive,
335
335
  agents, openclawConfig: readOpenclawConfig(), vendors: SUBSCRIPTION_VENDORS, modelsInUse, autoPruned, usageSnapshots: readUsageSnapshots(stateRoot), install });
336
+ try {
337
+ const { loadJobRecords, agentLookup } = await import("./stats.mjs");
338
+ const { recentSuggestions } = await import("./suggestions.mjs");
339
+ for (const s of recentSuggestions(loadJobRecords(path.join(stateRoot, "jobs")), { projectDir, agentFor: agents ? agentLookup(agents, agentProviderId) : () => null, now })) {
340
+ if (s.level !== "warn") continue;
341
+ result.issues.push({ id: `suggestion:${s.key}`, severity: "warn", title: s.title, detail: s.evidence, fix: s.command ?? "nomarmy stats (routing suggestions)", short: "routing tip" });
342
+ }
343
+ } catch { /* suggestions never break a health check */ }
336
344
  const toNotify = recordHealth(path.join(stateRoot, "health.json"), result, { now });
337
345
  return { result, toNotify };
338
346
  }
@@ -61,6 +61,35 @@ export function makeAbandonedBackgroundProcessTick(stateDir, { idleMs, minElapse
61
61
  * tool call and files changed so far (liveProgress), and when. Never asks
62
62
  * to stop; a failed read just skips that beat.
63
63
  */
64
+ // local_worker_stop (or `nomarmy jobs --stop`) writes this into the job's
65
+ // folder, so any session can stop any job; the next tick ends the worker.
66
+ export const STOP_REQUEST_FILE = "stop-request.json";
67
+ export function makeStopRequestTick(jobDir) {
68
+ return async () => (fs.existsSync(path.join(jobDir, STOP_REQUEST_FILE)) ? { stop: true, reason: "stopped" } : { stop: false });
69
+ }
70
+ /**
71
+ * Ask a running job to stop. Refuses what isn't running; never kills
72
+ * anything itself (the job's own tick does, within one tick).
73
+ * @returns {{ ok: boolean, message: string }}
74
+ */
75
+ export function requestJobStop({ jobsRoot, jobId, reason = null, now = new Date() }) {
76
+ if (!/^[A-Za-z0-9._-]{1,120}$/.test(String(jobId ?? ""))) return { ok: false, message: "not a job id" };
77
+ const jobDir = path.join(jobsRoot, jobId);
78
+ let status = null;
79
+ try { status = JSON.parse(fs.readFileSync(path.join(jobDir, "status.json"), "utf8")); } catch { return { ok: false, message: `no job ${jobId} (see local_worker_jobs)` }; }
80
+ if (status.state !== "running") return { ok: false, message: `job ${jobId} isn't running (it's ${status.state ?? "unknown"})` };
81
+ if (status.phase && status.phase !== "worker" && status.phase !== "starting" && status.phase !== "worktree") {
82
+ return { ok: false, message: `job ${jobId} is past its worker (phase ${status.phase}); it's finishing on its own and spends no more model usage` };
83
+ }
84
+ if (fs.existsSync(path.join(jobDir, STOP_REQUEST_FILE))) return { ok: true, message: `a stop was already requested for ${jobId}` };
85
+ fs.writeFileSync(path.join(jobDir, STOP_REQUEST_FILE), JSON.stringify({ at: now.toISOString(), reason: reason ? String(reason).slice(0, 300) : null }));
86
+ return { ok: true, message: `stop requested for ${jobId}: its worker ends within about 15 seconds, verification is skipped, and the worktree is kept uncommitted, so continue_from can pick the work up (on another model too)` };
87
+ }
88
+
89
+ export function readStopRequest(jobDir) {
90
+ try { return JSON.parse(fs.readFileSync(path.join(jobDir, STOP_REQUEST_FILE), "utf8")); } catch { return null; }
91
+ }
92
+
64
93
  export function makeHeartbeatTick(jobDir, liveProgress) {
65
94
  // Never two beats at once: a slow beat used to overlap the next.
66
95
  let busy = false;
@@ -475,8 +504,9 @@ export function createOpenClawRunner(deps) {
475
504
  // The heartbeat runs for every job, so a running job is always watchable
476
505
  // (status.json used to keep its launch-time updatedAt until the end).
477
506
  const onTick = combineTicks([
507
+ makeStopRequestTick(jobDir),
478
508
  makeHeartbeatTick(jobDir),
479
- ...(idleDiff ? [makeIdleDiffTick(cwd, idleDiff), makeAbandonedBackgroundProcessTick(stateDir, idleDiff)] : []),
509
+ ...(idleDiff ? [makeIdleDiffTick(cwd, { ...idleDiff, activity: async () => { const t = await readOpenClawTranscriptTail(stateDir, { limit: 5 }); return t.available ? { events: t.events, toolInFlight: t.toolInFlight } : null; } }), makeAbandonedBackgroundProcessTick(stateDir, idleDiff)] : []),
480
510
  ]);
481
511
  const liveLogs = { stdout: path.join(jobDir, `openclaw${logSuffix}.stdout.log`), stderr: path.join(jobDir, `openclaw${logSuffix}.stderr.log`) };
482
512
  // Where this call's own transcript events will start (see salvageFinishedRun).
package/lib/outcome.mjs CHANGED
@@ -207,8 +207,17 @@ export function policyAdmissionProblems(job, policy) {
207
207
  // A refactor meets the regression requirement through its own contract
208
208
  // (applyRefactorContract), not the revert check.
209
209
  if (policy.require_regression_check && job.verify_regression === false && !job.refactor) problems.push("this repo requires the revert check (policy.require_regression_check in .nomarmy.yml): verify_regression can't be false (a behavior-preserving change can declare refactor: true instead)");
210
+ // stakes: high (security, data loss, irreversible): the checks a General
211
+ // could skip on routine work are mandatory, whatever the repo's policy.
212
+ if (job.stakes === "high") {
213
+ if (!job.verification && !job.refactor) problems.push("a stakes: high job needs a `verification` profile: its result is only as good as the tests that prove it");
214
+ if (job.verify_regression === false && !job.refactor) problems.push("a stakes: high job can't turn the revert check off (verify_regression: false): it's the proof a test catches the change");
215
+ }
210
216
  return problems;
211
217
  }
218
+
219
+ /** The review line every high-stakes job carries, whatever its outcome. */
220
+ export const HIGH_STAKES_NOTE = "HIGH STAKES: accept this only after an independent review: a scout on another vendor (army_role security-analyst or similar) with reviews: <this job id>, or a judge on another vendor. Passing its tests isn't enough on its own; most defects that pass every check are in security, data or deploy-only paths.";
212
221
  /**
213
222
  * A declared refactor commits only when verification passed and no test
214
223
  * file was added, changed or deleted. Mechanical, not the General's call:
package/lib/process.mjs CHANGED
@@ -73,7 +73,10 @@ export function createProcess(ctx) {
73
73
  if (ticker) clearInterval(ticker);
74
74
  child.kill("SIGTERM");
75
75
  const error = new Error(message);
76
- error.timedOut = true;
76
+ // A stop someone asked for (local_worker_stop) isn't a timeout: a
77
+ // timeout invites a report-recovery call, which would spend exactly
78
+ // the usage the stop was meant to save.
79
+ error.timedOut = stopReason !== "stopped";
77
80
  if (stopReason) error.stopReason = stopReason;
78
81
  reject(error);
79
82
  };
package/lib/scout.mjs CHANGED
@@ -438,9 +438,22 @@ export function isScoutReportUnusable(report) {
438
438
  // The recovery call carries the question itself: a reply cut off mid-run can
439
439
  // leave the resumed session without it, and a scout told only to "restate the
440
440
  // question" then came back with an empty report (a live PM scout, twice).
441
- export function scoutReportRecoveryPrompt({ report = { targetTokens: 600, hardCapTokens: 1024 }, question = null, acceptance = [] } = {}) {
441
+ // A recovery call can't count on seeing the first run's work: when OpenClaw's
442
+ // cleanup crashes ("state ownership retained"), the follow-up starts a fresh
443
+ // session, and a scout told "report only what you already found" came back
444
+ // with nothing after nine minutes of real research. So the recovery carries
445
+ // what the first run left: its cut-off reply and the files it read, which it
446
+ // may re-read to confirm line numbers, and nothing else.
447
+ const EARLIER_REPLY_CHARS = 12000;
448
+ export function scoutReportRecoveryPrompt({ report = { targetTokens: 600, hardCapTokens: 1024 }, question = null, acceptance = [], earlierReply = "", filesRead = [] } = {}) {
442
449
  const asked = question ? `\n\nThe question you were answering:\n${String(question).trim()}${acceptance?.length ? `\n\nA complete answer covers:\n${acceptance.map((a) => `- ${a}`).join("\n")}` : ""}` : "";
443
- return `Your previous reply ended without the required SCOUT REPORT, or was cut off before reaching END.${asked}\n\nDo not repeat, redo, retry, or explore further. Do not call any tool. Based only on what you already found, reply with ONLY the report below, nothing before it, nothing after it:\n\nSCOUT REPORT\nQUESTION: <the question restated in one line>\nCONFIDENCE: high | medium | low\nFINDING: <one sentence> [src/example.js:10-24]\nNOT_FOUND: none | <what you looked for and could not find>\nEND\n\nIf you did not actually find anything worth a FINDING, say so under NOT_FOUND rather than inventing one. Target ${report.targetTokens} tokens; ${report.hardCapTokens} is the hard cap.`;
450
+ const reply = String(earlierReply ?? "").trim();
451
+ const earlier = reply ? `\n\nYour earlier reply, as far as it got${reply.length > EARLIER_REPLY_CHARS ? " (its last part)" : ""}:\n${reply.length > EARLIER_REPLY_CHARS ? reply.slice(-EARLIER_REPLY_CHARS) : reply}` : "";
452
+ const files = filesRead?.length ? `\n\nFiles you read during that work: ${filesRead.slice(0, 60).join(", ")}` : "";
453
+ const tools = filesRead?.length
454
+ ? "Do not explore further. You may re-read the files listed above to confirm exact line numbers for what you found; call no other tool."
455
+ : "Do not repeat, redo, retry, or explore further. Do not call any tool.";
456
+ return `Your previous reply ended without the required SCOUT REPORT, or was cut off before reaching END.${asked}${earlier}${files}\n\n${tools} Based on what you already found, reply with ONLY the report below, nothing before it, nothing after it:\n\nSCOUT REPORT\nQUESTION: <the question restated in one line>\nCONFIDENCE: high | medium | low\nFINDING: <one sentence> [src/example.js:10-24]\nNOT_FOUND: none | <what you looked for and could not find>\nEND\n\nIf you did not actually find anything worth a FINDING, say so under NOT_FOUND rather than inventing one. Target ${report.targetTokens} tokens; ${report.hardCapTokens} is the hard cap.`;
444
457
  }
445
458
 
446
459
  // ---------------------------------------------------------------------------
package/lib/stats.mjs CHANGED
@@ -5,6 +5,9 @@
5
5
 
6
6
  import fs from "node:fs";
7
7
  import path from "node:path";
8
+ import { computeSuggestions, formatSuggestions, reviewOf } from "./suggestions.mjs";
9
+
10
+ const SUGGESTION_WINDOW_MS = 14 * 86400000;
8
11
 
9
12
  /** Every readable job record under jobsRoot. */
10
13
  export function loadJobRecords(jobsRoot) {
@@ -17,6 +20,15 @@ export function loadJobRecords(jobsRoot) {
17
20
  return records;
18
21
  }
19
22
 
23
+ /** An agent name from agents.yml for a provider id, when exactly one agent uses it. */
24
+ export function agentLookup(agents = {}, providerOf) {
25
+ return (provider) => {
26
+ if (!provider) return null;
27
+ const names = Object.entries(agents).filter(([, a]) => { try { return providerOf(a) === provider; } catch { return false; } }).map(([name]) => name);
28
+ return names.length === 1 ? names[0] : null;
29
+ };
30
+ }
31
+
20
32
  /** "7d", "24h", or a date; returns epoch ms or null. */
21
33
  export function parseSince(value, now = Date.now()) {
22
34
  if (!value) return null;
@@ -87,7 +99,7 @@ export function resolveRepo(records, value) {
87
99
  * @param {object[]} records
88
100
  * @param {{ repo?: string|null, sinceMs?: number|null, untilMs?: number|null, role?: string|null, model?: string|null }} filter
89
101
  */
90
- export function computeStats(records, { repo = null, sinceMs = null, untilMs = null, role = null, model = null } = {}) {
102
+ export function computeStats(records, { repo = null, sinceMs = null, untilMs = null, role = null, model = null, agentFor = () => null, now = Date.now() } = {}) {
91
103
  const inRange = records.filter((r) => {
92
104
  const at = Date.parse(r.startedAt ?? r.finishedAt ?? "");
93
105
  if (sinceMs != null && !(at >= sinceMs)) return false;
@@ -169,6 +181,13 @@ export function computeStats(records, { repo = null, sinceMs = null, untilMs = n
169
181
  notCompleted: sortDesc(count(jobs.filter((r) => !/^(WORKER_DONE|RECOVERED_SUCCESS|VERIFIED|SCOUT_DONE|DECOMPOSE_DONE|SCOUT_NOT_FOUND)$/.test(r.outcome ?? "")), (r) => r.outcome)),
170
182
  reviewers,
171
183
  signals: sortDesc(signals),
184
+ highStakes: (() => {
185
+ const high = implement.filter((r) => r.stakes === "high");
186
+ return { jobs: high.length, reviewed: high.filter((r) => reviewOf(r, records)).length };
187
+ })(),
188
+ // How you're set up now: the last 14 days unless a period was asked for.
189
+ suggestions: computeSuggestions(sinceMs == null ? jobs.filter((r) => Date.parse(r.startedAt ?? "") >= now - SUGGESTION_WINDOW_MS) : jobs, { agentFor }),
190
+ suggestionWindow: sinceMs == null ? "the last 14 days" : "this period",
172
191
  };
173
192
  }
174
193
 
@@ -183,6 +202,9 @@ export function formatStats(s) {
183
202
  const lines = [
184
203
  `nomArmy stats${s.repo ? ` for ${s.repo}` : " (all repositories)"}${s.role ? `, role ${s.role}` : ""}${s.model ? `, model ${s.model}` : ""}, ${s.period.from ? `${s.period.from.slice(0, 10)} to ${s.period.to.slice(0, 10)}` : "no jobs"}`,
185
204
  "",
205
+ `SUGGESTIONS (from ${s.suggestionWindow ?? "this period"}; never applied for you)`,
206
+ ...formatSuggestions(s.suggestions ?? []),
207
+ "",
186
208
  "VOLUME",
187
209
  ` Jobs ${s.volume.jobs}: ${list(s.volume.byMode)}${s.unplacedVerifyRuns ? ` (plus ${s.unplacedVerifyRuns} older verify run(s) that don't record their repository)` : ""}`,
188
210
  ` By role ${list(s.volume.byRole)}`,
@@ -198,6 +220,7 @@ export function formatStats(s) {
198
220
  ` Passed, but reverting still passed ${c.revertStillPassed}${pct(c.revertStillPassed, c.claimedDone)}`,
199
221
  ` Passed both ${c.passedBoth}${pct(c.passedBoth, c.claimedDone)}`,
200
222
  ` of those, flagged by another check ${c.flaggedAfterPassing} (mutants, Jev, judge, rewritten checks)`,
223
+ ` High-stakes jobs ${s.highStakes?.jobs ?? 0}, ${s.highStakes?.reviewed ?? 0} with an independent review`,
201
224
  " Defects the General found at integration aren't in the records; count them in your own review.",
202
225
  "",
203
226
  "DIDN'T COMPLETE",
@@ -0,0 +1,118 @@
1
+ // Routing suggestions from the job records: which role and model pairings
2
+ // are working, which aren't, and what a change would be. Evidence first,
3
+ // with minimum sample sizes; roles do different work, so a comparison across
4
+ // roles is worded as something to try, not a verdict. Never applied: the
5
+ // operator or the General decides, and each suggestion carries the command.
6
+
7
+ import { jobRole } from "./stats.mjs";
8
+
9
+ export const MIN_JOBS = 5;
10
+ const OK = /^(WORKER_DONE|RECOVERED_SUCCESS|SCOUT_DONE|SCOUT_NOT_FOUND|DECOMPOSE_DONE|VERIFIED)$/;
11
+
12
+ const provider = (r) => r.metrics?.worker_provider ?? r.worker?.provider ?? null;
13
+ const model = (r) => r.metrics?.worker_model ?? r.worker?.model ?? null;
14
+ const agent = (r) => r.labels?.agent ?? null;
15
+ const pct = (n, of) => Math.round((100 * n) / of);
16
+
17
+ /** Whether a high-stakes job has had an independent review: a scout, or a judge, on another vendor. */
18
+ export function reviewOf(job, records) {
19
+ const workerProvider = provider(job);
20
+ const scout = records.find((r) => r.mode === "scout" && r.reviews === job.jobId && provider(r) && provider(r) !== workerProvider);
21
+ if (scout) return { by: "scout", jobId: scout.jobId, provider: provider(scout) };
22
+ const judge = job.validators?.judge;
23
+ if (judge?.answer && judge.provider && judge.provider !== workerProvider) return { by: "judge", provider: judge.provider };
24
+ return null;
25
+ }
26
+
27
+ /**
28
+ * @param {object[]} records this repo's records, already filtered to a period
29
+ * @returns {{ level: "warn"|"info", key: string, title: string, evidence: string, command: string|null }[]}
30
+ */
31
+ export function computeSuggestions(records, { minJobs = MIN_JOBS, agentFor = () => null } = {}) {
32
+ const out = [];
33
+ const work = records.filter((r) => r.mode === "implement" || r.mode === "scout");
34
+
35
+ // Per role and model.
36
+ const groups = new Map();
37
+ for (const r of work) {
38
+ const key = `${jobRole(r) ?? ""}|${model(r) ?? ""}|${r.mode}`;
39
+ const g = groups.get(key) ?? { role: jobRole(r), model: model(r), mode: r.mode, agent: agent(r), provider: provider(r), jobs: 0, ok: 0, timeout: 0, unsupported: 0, tokens: 0 };
40
+ g.jobs++; if (OK.test(r.outcome ?? "")) g.ok++;
41
+ if (r.outcome === "WORKER_TIMEOUT") g.timeout++;
42
+ if (r.outcome === "SCOUT_UNSUPPORTED") g.unsupported++;
43
+ g.tokens += r.metrics?.worker_tokens_total ?? 0;
44
+ g.agent = g.agent ?? agent(r) ?? agentFor(provider(r));
45
+ groups.set(key, g);
46
+ }
47
+ const assign = (g, to = "<another model>") => (g.role ? `nomarmy army assign ${g.role} ${g.agent ?? "<agent>"} ${to}` : null);
48
+
49
+ for (const g of groups.values()) {
50
+ if (!g.model) continue;
51
+ const kind = g.mode === "scout" ? "scouts" : "implement jobs";
52
+ const who = g.role ? `${g.role} on ${g.model}` : `${kind} with no role on ${g.model}`;
53
+ // Scouts that come back empty.
54
+ if (g.mode === "scout" && g.jobs >= 3 && g.unsupported / g.jobs >= 0.4) {
55
+ out.push({ level: "warn", key: `scout-unsupported:${g.role}:${g.model}`, title: `${who}: ${g.unsupported} of ${g.jobs} scouts came back unsupported`,
56
+ evidence: "Their findings couldn't be tied to cited lines. A different agent, or report: full, usually fixes it.", command: assign(g) });
57
+ continue;
58
+ }
59
+ if (g.jobs < minJobs) continue;
60
+ // A pairing that rarely finishes.
61
+ if (g.ok / g.jobs < 0.5) {
62
+ const better = [...groups.values()].filter((o) => o !== g && o.mode === g.mode && o.jobs >= minJobs && o.model && o.ok / o.jobs >= g.ok / g.jobs + 0.2)
63
+ .sort((a, b) => b.ok / b.jobs - a.ok / a.jobs)[0];
64
+ out.push({ level: "warn", key: `low-success:${g.role}:${g.model}:${g.mode}`, title: `${who} finished ${g.ok} of ${g.jobs} ${g.role ? kind : ""}`.trim() + ` (${pct(g.ok, g.jobs)}%)`,
65
+ evidence: (better ? `${better.model} finished ${pct(better.ok, better.jobs)}% of its ${better.jobs} ${kind} here${better.role ? ` (as ${better.role})` : ""}.` : "No other model has enough jobs here to compare.") + (g.role ? "" : " These ran with no army role: send this kind of work to a role on a stronger agent instead."),
66
+ command: g.role ? assign(g, better?.model ?? "<another model>") : null });
67
+ }
68
+ // Timeouts.
69
+ if (g.timeout >= 3 && g.timeout / g.jobs >= 0.25) {
70
+ out.push({ level: "info", key: `timeouts:${g.role}:${g.model}`, title: `${who} timed out on ${g.timeout} of ${g.jobs} jobs`,
71
+ evidence: "Smaller briefs (one outcome each), a longer timeout_seconds, or a faster model would help.", command: null });
72
+ }
73
+ }
74
+
75
+ // A lighter model doing as well on implement work, for much less.
76
+ const impl = [...groups.values()].filter((g) => g.mode === "implement" && g.model && g.jobs >= minJobs && g.tokens > 0);
77
+ for (const heavy of impl) {
78
+ for (const light of impl) {
79
+ if (light === heavy || light.model === heavy.model) continue;
80
+ const lightRate = light.ok / light.jobs, heavyRate = heavy.ok / heavy.jobs;
81
+ if (lightRate >= heavyRate - 0.05 && light.tokens / light.jobs <= 0.5 * (heavy.tokens / heavy.jobs)) {
82
+ out.push({ level: "info", key: `lighter:${heavy.role}:${heavy.model}:${light.model}`,
83
+ title: `${light.model} finished ${pct(light.ok, light.jobs)}% of its jobs${light.role ? ` (as ${light.role})` : ""} on ${Math.round((light.tokens / light.jobs) / 1000)}k tokens a job; ${heavy.model}${heavy.role ? ` (as ${heavy.role})` : ""} finished ${pct(heavy.ok, heavy.jobs)}% on ${Math.round((heavy.tokens / heavy.jobs) / 1000)}k`,
84
+ evidence: "They did different work, so try it rather than switch outright: put the role on auto so the General picks per job, or send its routine pieces to the lighter model.",
85
+ command: heavy.role ? `nomarmy army assign ${heavy.role} ${heavy.agent ?? "<agent>"} auto` : null });
86
+ }
87
+ }
88
+ }
89
+
90
+ // Where the money goes.
91
+ const spend = new Map();
92
+ for (const r of work) if (Number.isFinite(r.metrics?.worker_cost_usd) && r.metrics.worker_cost_usd > 0) spend.set(model(r), (spend.get(model(r)) ?? 0) + r.metrics.worker_cost_usd);
93
+ const total = [...spend.values()].reduce((a, b) => a + b, 0);
94
+ const [top, topUsd] = [...spend.entries()].sort((a, b) => b[1] - a[1])[0] ?? [];
95
+ if (top && topUsd >= 5 && topUsd / total >= 0.5) {
96
+ out.push({ level: "info", key: `spend:${top}`, title: `${top} is ${pct(topUsd, total)}% of API spend ($${topUsd.toFixed(2)} of $${total.toFixed(2)})`,
97
+ evidence: "Worth knowing rather than changing if it's catching real problems; check its reviews' findings before moving it.", command: null });
98
+ }
99
+
100
+ // High-stakes work without an independent review.
101
+ const unreviewed = records.filter((r) => r.mode === "implement" && r.stakes === "high" && !reviewOf(r, records));
102
+ if (unreviewed.length) {
103
+ out.push({ level: "warn", key: `unreviewed:${unreviewed.map((r) => r.jobId).sort().join(",")}`, title: `${unreviewed.length} high-stakes job(s) without an independent review: ${unreviewed.slice(0, 5).map((r) => r.jobId).join(", ")}`,
104
+ evidence: "Send a scout on another vendor with reviews: <job id> before accepting them (or configure a judge on another vendor).", command: null });
105
+ }
106
+ return out;
107
+ }
108
+
109
+ /** This repository's suggestions from its last 14 days of jobs. */
110
+ export function recentSuggestions(records, { projectDir, agentFor = () => null, now = Date.now(), days = 14 } = {}) {
111
+ const since = now - days * 86400000;
112
+ return computeSuggestions(records.filter((r) => r.projectDir === projectDir && Date.parse(r.startedAt ?? "") >= since), { agentFor });
113
+ }
114
+
115
+ export function formatSuggestions(list) {
116
+ if (!list.length) return [" none: nothing in the records suggests a routing change"];
117
+ return list.flatMap((s) => [` ${s.level === "warn" ? "!" : "-"} ${s.title}`, ` ${s.evidence}`, ...(s.command ? [` ${s.command}`] : [])]);
118
+ }
@@ -99,6 +99,9 @@ export function summarizeTranscriptEvents(events) {
99
99
  }
100
100
  }
101
101
  out.filesRead = [...new Set(out.filesRead)];
102
+ // A tool call with no result yet: the worker is waiting on it (a long test
103
+ // run writes nothing to the transcript until it returns).
104
+ out.toolInFlight = pending.length > 0;
102
105
  return out;
103
106
  }
104
107
 
package/mcp/server.mjs CHANGED
@@ -36,7 +36,9 @@ import { readUsageSnapshots, usageStatus } from "../lib/usage-limits.mjs";
36
36
  import { modelRefusals } from "../lib/health.mjs";
37
37
  import { podmanProblem, podmanVmStartedAt } from "../lib/podman-health.mjs";
38
38
  import { restartNotice } from "../lib/install-freshness.mjs";
39
- import { loadJobRecords, computeStats, formatStats, parseSince, resolveRepo } from "../lib/stats.mjs";
39
+ import { requestJobStop } from "../lib/openclaw-run.mjs";
40
+ import { loadJobRecords, computeStats, formatStats, parseSince, resolveRepo, agentLookup } from "../lib/stats.mjs";
41
+ import { recentSuggestions } from "../lib/suggestions.mjs";
40
42
  import { probeModel } from "../lib/model-probe.mjs";
41
43
  import { jevSettings, judgeSettings } from "../lib/validators.mjs";
42
44
  import { agentRunsToolsOnHost } from "../lib/dispatch-schema.mjs";
@@ -301,6 +303,8 @@ export const jobSchema = z.object({
301
303
  model: z.string().regex(/^\S{1,200}$/).optional().describe("The model to run on the job's agent (an api or subscription agent), e.g. \"gpt-6-sol\". Overrides the role's model and the agent's default. Required when the role's model is \"auto\" or the agent has no default. The `army` tool lists each agent's models. Refused on the local agent, whose model `nomarmy model` sets."),
302
304
  run_id: z.string().regex(/^run-[a-z0-9-]{1,80}$/).optional().describe("The /feature run this job belongs to (from run_start). Admission then enforces the run's limits (jobs, api spend, hours) and refuses an agent the run has paused after a vendor usage-limit error; the finished job is recorded into the run."),
303
305
  report: z.enum(["brief", "standard", "full"]).optional().describe("How much the worker may report back, capped by its agent's tier: brief (today's local-sized report), standard (the default), full (the frontier ceiling: about 2k tokens for implement, 4k for a scout). An api or subscription scout defaults to full because its findings are the point; other jobs default to standard. The report lands in your own context and is re-read every later turn. No effect on the local model, whose caps are calibrated."),
306
+ stakes: z.enum(["normal", "high"]).optional().describe("implement: how much a mistake would cost, separate from how hard the work is. high for anything touching security or access control, personal or tenant data, data loss, money, or changes that can't be undone: a verification profile is then required, the revert check can't be turned off, and the job always comes back needing review until an independent review (a scout on another vendor with reviews: <job id>, or a judge on another vendor) has looked at it. A one-line auth change is simple and high-stakes."),
307
+ reviews: z.string().regex(/^[A-Za-z0-9._-]{1,120}$/).optional().describe("scout: the job id this scout independently reviews, so the review is recorded against that job (nomarmy stats shows high-stakes jobs with and without one). Use a different vendor than the job's worker."),
304
308
  commit_subject: z.string().max(200).optional().describe("implement: the subject line of the commit nomArmy makes on the worker branch, e.g. \"Keep held-back tables in the list_tables cache\". Defaults to the task's first sentence; the body is the worker's NOTE, and the job id is a trailer."),
305
309
  army_role: z.string().regex(/^[a-z][a-z0-9-]{0,63}$/).optional().describe("Dispatch by army role (e.g. \"sr-dev\", \"security-analyst\"): nomArmy runs it on the agent this repo assigns to that role and puts the role's description at the top of the brief. Call the `army` tool first to see this repo's roles. Mutually exclusive with agent. Add on_behalf_of in case the role's agent is a subscription; it's ignored otherwise."),
306
310
  confirm_over_limit: z.boolean().optional().describe("Override a reached usage limit: the General must ask the operator before resubmitting with confirm_over_limit: true, or send the job to another agent. nomArmy never sets it itself."),
@@ -352,7 +356,7 @@ function jobArgs(args, workerId) {
352
356
  return { task: args.task, acceptance: args.acceptance, verification: args.verification, mode: args.mode, baseRef: args.base_ref,
353
357
  timeoutSeconds: args.timeout_seconds, profile: args.profile, reasoning: args.reasoning, pool: args.pool,
354
358
  subscriptionWorker, onBehalfOf: args.on_behalf_of, model: args.model ?? null, reportSize: args.report ?? null, evidence: args.evidence,
355
- verifyRegression: resolveVerifyRegression(args), commitSubject: args.commit_subject ?? null, refactor: Boolean(args.refactor), continueFrom: args.continue_from ?? null, workerId };
359
+ verifyRegression: resolveVerifyRegression(args), commitSubject: args.commit_subject ?? null, refactor: Boolean(args.refactor), continueFrom: args.continue_from ?? null, stakes: args.stakes ?? null, reviews: args.reviews ?? null, workerId };
356
360
  }
357
361
  server.tool("local_worker", "Run one isolated local worker and wait for it. mode=implement edits in its own worktree and the coordinator commits only on a valid done report (or a recovered job that passed independent verification); failed or incomplete worktrees are retained. mode=scout answers a question from a read-only snapshot with mandatory [path:line] citations that nomArmy verifies and expands. mode=decompose (also read-only) proposes 2+ independent subtasks for a broad objective instead of one worker turn trying to do too much; the proposal is never auto-dispatched, review it and make a separate call with the subtasks you choose. Refuses under memory pressure or over capacity; use local_worker_start + local_worker_status to avoid blocking.", jobSchema.shape,
358
362
  async rawArgs => {
@@ -438,11 +442,21 @@ server.tool("stats", "What nomArmy's own job records show for this repository (o
438
442
  }, async ({ since, until, all_repos, repo, role, model, format }) => {
439
443
  try {
440
444
  const records = loadJobRecords(jobsRoot);
441
- const stats = computeStats(records, { repo: repo ? resolveRepo(records, repo) : all_repos ? null : projectDir, sinceMs: parseSince(since), untilMs: parseSince(until), role: role ?? null, model: model ?? null });
445
+ let agentFor = () => null;
446
+ try { agentFor = agentLookup(agentsConfig().agents, agentProviderId); } catch { /* commands name <agent> */ }
447
+ const stats = computeStats(records, { repo: repo ? resolveRepo(records, repo) : all_repos ? null : projectDir, sinceMs: parseSince(since), untilMs: parseSince(until), role: role ?? null, model: model ?? null, agentFor });
442
448
  return toolText(format === "json" ? JSON.stringify(stats, null, 2) : formatStats(stats));
443
449
  } catch (error) { return toolText(error.message, true); }
444
450
  });
445
451
 
452
+ server.tool("local_worker_stop", "Stop a running job's worker, for example one burning a frontier model's usage on the wrong track. Its worker ends within about 15 seconds, without the report-recovery call a timeout gets and without running verification; its worktree is kept uncommitted, so a new job with continue_from: <job_id> (and a cheaper model if you like) can finish the work. Works for a job started by any session. Refuses a job that isn't running or is already past its worker.", {
453
+ job_id: z.string().regex(/^[A-Za-z0-9._-]{1,120}$/),
454
+ reason: z.string().max(300).optional().describe("Why it's being stopped; recorded in the job's issues."),
455
+ }, async ({ job_id, reason }) => {
456
+ const r = requestJobStop({ jobsRoot, jobId: job_id, reason: reason ?? null });
457
+ return toolText(r.message, !r.ok);
458
+ });
459
+
446
460
  server.tool("local_worker_capacity", "What this host can take right now: context per nom and the brief/report budgets derived from it, memory pressure and whether another job would be admitted, and the jobs currently running. Read-only.", {}, async () => {
447
461
  await budgetState.refresh();
448
462
  return toolText(JSON.stringify(withRestartNotice(capacitySnapshot()), null, 2));
@@ -552,6 +566,12 @@ server.tool("army", "Who you, the General, are and who you call for what in this
552
566
  role.modelNote = `${role.model} isn't in OpenClaw's catalog for ${role.agent}; \`army assign\` checked it with a real test call when it was set, and the catalog can lag new models. Use it as assigned; if a job reports "Unknown model", reassign.`;
553
567
  }
554
568
  }
569
+ // Routing suggestions from this repo's recent jobs (lib/suggestions.mjs):
570
+ // tell the operator about them; never apply one without their say-so.
571
+ try {
572
+ const list = recentSuggestions(loadJobRecords(jobsRoot), { projectDir, agentFor: agentLookup(agents, agentProviderId) });
573
+ if (list.length) summary.suggestions = { note: "From this repo's last 14 days of jobs. Tell the operator; change routing only with their say-so (nomarmy army assign).", items: list.map(({ level, title, evidence, command }) => ({ level, title, evidence, command })) };
574
+ } catch { /* suggestions are a bonus; the army summary stands without them */ }
555
575
  return toolText(JSON.stringify(withRestartNotice(summary), null, 2));
556
576
  } catch (error) {
557
577
  return toolText(error.message, true);
package/package.json CHANGED
@@ -3,7 +3,7 @@
3
3
  "description": "Every byte verified: a harness for AI coding workers whose claims are never trusted. Your coding assistant stays in charge while workers implement and test in sandboxes, and nomArmy checks every change before it is committed.",
4
4
  "author": "Rayson Technologies",
5
5
  "license": "Apache-2.0",
6
- "version": "0.1.0-alpha.12",
6
+ "version": "0.1.0-alpha.14",
7
7
  "private": false,
8
8
  "type": "module",
9
9
  "engines": {
@@ -12,9 +12,10 @@ You are the General. Build this feature end to end with nomArmy's army and come
12
12
 
13
13
  Follow the army's workflow, calling only the roles the work needs:
14
14
 
15
+ 0. **Routing.** The `army` tool's `suggestions` come from this repo's recent jobs (a role that keeps failing, a cheaper model doing as well, unreviewed high-stakes work). Tell the operator about any at the start and in your final summary; change routing only if they say so.
15
16
  1. **Plan.** Scout the repo as needed (`repo_evidence` first; a scout only for research that would pull many files into your context). Write the plan into the run log: the outcome, acceptance criteria, the pieces, and which role gets each.
16
- 2. **Build.** Dispatch with `army_role` (and `on_behalf_of` when the role's agent is a subscription). The Sr Dev takes the core and harder work; the Jr Dev takes simple, fully specified pieces; UI/UX takes UI. For a role on `auto`, pick the model from the agent's list in the `army` tool: the lighter model for routine work, the frontier one for subtle work.
17
- 3. **Review.** When the build is in, call the specialists that apply (data architect for data work, security analyst for anything touching auth, input, secrets or data exposure), then the PM against the plan. Send what they find back to the builders as new, bounded jobs. A build job that comes back partial, blocked or failing verification is finished with `continue_from: <its job id>` and a brief of just the correction, never by fixing its files yourself: only verified work lands.
17
+ 2. **Build.** Dispatch with `army_role` (and `on_behalf_of` when the role's agent is a subscription). The Sr Dev takes the core and harder work; the Jr Dev takes simple, fully specified pieces; UI/UX takes UI. For a role on `auto`, pick the model from the agent's list in the `army` tool: the lighter model for routine work, the frontier one for subtle work. Mark a job `stakes: high` when a mistake would be costly (security or access control, personal or tenant data, data loss, money, anything irreversible), however small the change: that's separate from how hard it is.
18
+ 3. **Review.** When the build is in, call the specialists that apply (data architect for data work, security analyst for anything touching auth, input, secrets or data exposure), then the PM against the plan. Every `stakes: high` build job gets an independent review before you accept it: a scout on a different vendor than its worker, with `reviews: <that job id>`. Send what they find back to the builders as new, bounded jobs. A build job that comes back partial, blocked or failing verification is finished with `continue_from: <its job id>` and a brief of just the correction, never by fixing its files yourself: only verified work lands.
18
19
  4. **Acceptance.** PO and stakeholder test end to end. Checks that only run existing tests use mode: verify; writing new e2e checks is still an implement job. Fix what they find the same way.
19
20
  5. **Integrate.** Review every diff against nomArmy's verified record -- a worker's report is a claim, not evidence -- and bring the accepted work together on one branch. **Never merge into the developer's branch, and never push.** The finished state is a branch ready for the operator to review and merge.
20
21
 
@@ -26,7 +27,7 @@ Stop and ask only for something irreversible or outside this repository: merging
26
27
 
27
28
  ## Watching jobs
28
29
 
29
- Prefer `local_worker_start`. Right after starting a job, if your coordinator can run a background command, run `nomarmy jobs --wait <job_id>` in the background so you're told the moment it finishes and can tell the operator. Otherwise poll `local_worker_status` with the longest `wait_seconds` it allows. `nomarmy jobs --events` remains available for a stream covering every job. `run_status` lists the run's running jobs as well as finished ones. The operator gets a desktop notification whenever a job finishes and whenever the run crosses a limit, so you don't need to relay each one.
30
+ Prefer `local_worker_start`. Right after starting a job, if your coordinator can run a background command, run `nomarmy jobs --wait <job_id>` in the background so you're told the moment it finishes and can tell the operator. Otherwise poll `local_worker_status` with the longest `wait_seconds` it allows. For several jobs at once, `nomarmy jobs --events --until-done` exits when every one has finished. The plain `nomarmy jobs --events` stream never exits on its own while jobs run, so don't run it as a background command (you'd only hear when it exits); use it only with a monitor that wakes on each line. `run_status` lists the run's running jobs as well as finished ones. The operator gets a desktop notification whenever a job finishes and whenever the run crosses a limit, so you don't need to relay each one.
30
31
 
31
32
  ## Limits
32
33