jules-orchestrator-kit 0.68.0 → 0.69.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/CHANGELOG.md CHANGED
@@ -5,6 +5,34 @@ All notable changes to this project will be documented in this file.
5
5
  The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
6
6
  and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
7
7
 
8
+ ## [0.69.0] - 2026-09-04
9
+ *A denominator is not evidence if the things counted in it were never read.*
10
+
11
+ A third cold-start trial against v0.68.0. Twelve findings; six reproduced, and the two most serious were graded lower by the trial than they deserved. The pattern in both: the guard recognised a line as an assertion, counted it in `assertionsSeen`, and reported `PASS` without ever reading the value being asserted. `UNREADABLE` exists precisely so that "I could not read this" and "I read this and it is fine" look different — and a line could pass the readability test while its expectation stayed opaque.
12
+
13
+ ### Fixed
14
+ - **Expected-Value-First Assertions Were Dismissed As Reworded Messages (`src/security.mjs`)**: `messageArgIndices` treated a string in argument 0 of any two-or-more argument assertion as prose for a human. JUnit and PHPUnit *document* the opposite order — `assertEquals(expected, actual)` — so no Java or PHP repository had any protection against a rewritten string expectation, and neither did Python's `assertIn(member, container)` or `assertNotIn`. Measured: `assertEquals("Hello World", out)` → `assertEquals("Hello Tampered", out)` returned `PASS` with `assertionsSeen: 2`. The one dialect that really does put prose first is JUnit 4's three-argument form, and it is distinguishable, because its trailing argument is the actual value rather than a message.
15
+ - **A Rewritten Regex Was Neither A Change Nor A Loss (`src/security.mjs`)**: `blankLiterals` blanked strings, numbers and booleans, and left regex literals alone — so `toMatch(/Hello World/)` and `toMatch(/Hello Tampered/)` normalised to two different shapes, never met in a bucket, and balanced each other out in the count: one specific assertion removed, one added, silence. Jest's `toMatch` and `toThrow` and RSpec's `match` all take the pattern as their only argument, which is where a test states what it expects. The lookbehind separating a regex from a division is what makes this safe: an operand never precedes the opening `/` of a pattern.
16
+ - **Renaming A Test Was Reported As Rewriting An Expectation (`src/security.mjs`)**: on the one-line form — `test("adds", () => { assert.strictEqual(add(2, 3), 5); });` — the name blanked to the same shape as its replacement, the two paired, and an author who renamed a test and nothing else was handed a `CRITICAL` finding quoting their whole line back at them. The multi-line form was never affected, because `NON_JOINING_CALL` already keeps a test name from joining to the assertion below it; this is the same rule for statements that fit on one line. A rename that also moves the expectation still differs after the name is blanked, so it is still reported.
17
+ - **A Verification Command With An Environment Prefix Never Started (`src/git.mjs`)**: `PYTHONPATH=src python3 -m pytest` is how a large part of the Python world writes its test command. `execFileSync` took the whole assignment as the program name and failed with `spawnSync PYTHONPATH=src ENOENT`, which the gate reported as a failed verification — telling the user their tests broke when the command had never run. Leading assignments are peeled into the child's environment, which behaves the same on every platform where handing the string to a shell would not.
18
+ - **`node -e ""` Was An Oracle (`src/stack-detector.mjs`)**: `true`, `echo ok` and `exit 0` were refused; the same no-op spelled as an interpreter with an empty program was accepted, and the gate returned `APPROVED (Exit 0)` on a change verified by nothing. Nothing can enumerate every way to write a command that runs nothing — that is what the collection floor is for, one-sidedly and by design — but the spellings that *look* like work are a small, decidable set.
19
+ - **The Timeout Message Named Neither The Limit Nor The Knob (`src/git.mjs`, `src/engine.mjs`, `src/wizard-init.mjs`, `README.md`)**: the default was 60 seconds, which an ordinary mid-sized suite exceeds on cold caches, and the explanation was suppressed by a guard testing for a token that Node's own `spawnSync sh ETIMEDOUT` already contains — so it was skipped exactly when it was needed. `verify.timeout_ms` existed and worked, and appeared in no README, no generated config, and no error message. Default is now 300000, the key is written into the config `init` generates, and the message says the command was killed rather than failed.
20
+
21
+ ### Changed
22
+ - **`init` Rejects An Oracle That Proves Nothing (`src/wizard-init.mjs`)**: `pnpm -r test` on a workspace whose packages declare no test script exits 0, prints nothing, and runs nothing — and the probe blessed it, leaving the repository configured to approve every future change against silence. Choosing a command is the right moment to be strict about this: at `init` the cost of rejecting a candidate is trying the next one, where at gate time it would be a hard red on a repository that is fine. Two passes now — prefer a command that proves it ran something, and only settle for one that merely exits 0 while saying so out loud.
23
+ - **Seven More Runners State Their Count (`src/ops/test-collection.mjs`)**: Python `unittest`, GoogleTest, Catch2, `deno test`, `bun test`, AVA, and a bare TAP plan all landed in the same "I could not tell" bucket as a command that ran nothing. The floor stays one-sided — an unrecognised runner still passes, because failing on "I could not tell" would hard-red every correct repository not on the list — so widening it only ever converts silence into an answer.
24
+
25
+ ### Also Fixed
26
+ - **The Documented Default Was Never The Loader's Default (`src/config.mjs`)**: the fallback was added to the engine, and `loadConfig` always supplied a value — so the fallback was unreachable and every repository without an explicit `timeout_ms` kept the one-minute limit while the changelog said five. A default belongs where the value is resolved, not where it is consumed. This is the third time in this project's history that a rule has been written in one place while another kept the old answer, so the regression test asserts the value a caller actually receives rather than any one site.
27
+ - **Two Tests Gated This Repository Against Itself (`test/engine.test.mjs`, `test/kit.test.mjs`)**: both called the gate with `root: process.cwd()`, so the verify stage ran `npm test` — the whole suite, from inside the suite. Neither could finish; the stage timeout was their only stopping condition, and each asserted little more than that a boolean was a boolean. They cost one full timeout per run and were invisible while that timeout was 60 seconds. Raising it to 300 turned them into a CI failure on every matrix cell, which is how a pair of tests that verified nothing finally got noticed. Both now run against a fixture with a bounded oracle and assert the verdict. The suite went from 67s to 19s.
28
+ - **The Egress Guard Could Not See `.js` (`test/egress-allowlist.test.mjs`)**: the scan collected `.mjs` only, so `bin/init.js` — published as the `jules-init` binary — sat outside the boundary entirely. Harmless as it stands, but the guard's whole purpose is that a reviewer can trust the boundary without reading every commit.
29
+
30
+ ### Not Reproduced
31
+ Six of the twelve did not hold, and five quoted terminal output that does not exist anywhere in the shipped code. `--allow-test-change expectation` propagates correctly through both phases and returns `APPROVED`; `--allow-test-change removal` allows deleting a dead test alongside its dead function, also `APPROVED` — both reported failures were a test suite genuinely broken by the reporter's own edit, which is the same fixture artefact this project has hit repeatedly. `agentctl plan approve --dry-run` errors with *"Session ID is required"* rather than printing the quoted *"Plan Approved Successfully!"*. `task create -p` did not block on a prompt. A protected `package.json` is the design, and the gate already prints `To allow protected files in this run, pass: agentctl gate --allow-protected` — the finding stated there were no flag hints in the output. The `TEST_DIALECT_UNREADABLE` remediation copy was already correct: *"Exit 6 Test Integrity Violation — no secret was found"*, not the quoted *"Secret Leak Prevented"*. And `pnpm -r test` exits 0 on pnpm 10.33.4 rather than the reported `ERR_PNPM_RECURSIVE_RUN_NO_SCRIPT` — the real defect there was worse than the one reported, and is fixed above.
32
+
33
+ ### Added
34
+ - **Eleven Cases In The Policy Contract (`src/guard-policy.mjs`)**: seven canaries for the expectation forms that were invisible — JUnit and PHPUnit expected-first, `assertIn`, `assertNotIn`, and regex patterns in `toMatch`, `toThrow` and RSpec `match` — and four innocent edits for the renames that were called tampering, plus JUnit 4's message-first form, which must stay silent. 42 canaries, 16 innocent edits.
35
+
8
36
  ## [0.68.0] - 2026-09-04
9
37
  *A check that examined nothing does not get to say APPROVED.*
10
38
 
package/README.md CHANGED
@@ -208,7 +208,7 @@ To maximize PR merge rates, dispatch tasks according to deterministic boundaries
208
208
  * **Fail-Closed Security & Secret Redaction:** Evaluates explicit Deny rules before Allow rules against canonicalized, case-folded paths. Redacts high-entropy keys and base64-encoded credentials (such as Kubernetes `Secret` manifests).
209
209
  * **Complexity & Cost Router:** Zero-dependency heuristic classifier (`src/router.mjs`) routing mechanical tasks to lightweight models while reserving primary models for complex refactors, with a `node --check` syntax-verification gate that transparently escalates a FAST-tier result to the primary provider if it left broken JS on disk.
210
210
  * **Terminal UI & Diagnostic Matrix (`agentctl doctor`):** Interactive terminal dashboard, task sidecar manager, and automated transactional self-repair.
211
- * **Verified Test Suite:** Tested with **1163 unit tests across 166 suites passing in < 15.0s**.
211
+ * **Verified Test Suite:** Tested with **1223 unit tests across 173 suites**, green on every supported platform.
212
212
 
213
213
  <br/>
214
214
 
@@ -286,6 +286,9 @@ branchPrefix: "agent/" # Prefix for task branches
286
286
  verify:
287
287
  test: "npm test"
288
288
  build: "npm run build"
289
+ timeout_ms: 300000 # Per-stage kill time in ms (default 300000)
290
+ minTests: 1 # Floor for "the suite actually ran" (0 disables)
291
+ required: true # false = this repo uses only the scope/secret phases
289
292
 
290
293
  # Scope protection rules (Deny-first evaluation)
291
294
  scope:
package/ROADMAP_V1.md CHANGED
@@ -11,15 +11,23 @@ The **jules-orchestrator-kit** is the zero-dependency safety gatekeeper and self
11
11
  ## 📌 Release Milestones Overview
12
12
 
13
13
  ```
14
- v0.68.0 (Current Stable) ──► v0.69.0 (Distributed Swarms & Leases) ──► v1.0.0 (Production Hardened Kernel)
15
- (Can It Still Go Red?) (Multi-Agent DAG & Resource Locks) (Enterprise Telemetry & SLA)
14
+ v0.69.0 (Current Stable) ──► v0.70.0 (Distributed Swarms & Leases) ──► v1.0.0 (Production Hardened Kernel)
15
+ (Read, Not Just Counted) (Multi-Agent DAG & Resource Locks) (Enterprise Telemetry & SLA)
16
16
  ```
17
17
 
18
18
  ---
19
19
 
20
- ## ✅ Shipped Milestones (v0.20.0 – v0.68.0)
20
+ ## ✅ Shipped Milestones (v0.20.0 – v0.69.0)
21
21
 
22
22
 
23
+ ### v0.69.0: Read, Not Just Counted
24
+ - [x] **Expected-Value-First Assertions (`src/security.mjs`)** — JUnit and PHPUnit document the order the guard read as prose.
25
+ - [x] **A Regex Is An Expected Value (`src/security.mjs`)** — a rewritten pattern was neither a change nor a loss.
26
+ - [x] **Renaming A Test Is Not Tampering (`src/security.mjs`)** — on the one-line form, the name blanked into the expectation.
27
+ - [x] **An Environment Prefix Runs (`src/git.mjs`)** — `PYTHONPATH=src pytest` never started, and was reported as a failure.
28
+ - [x] **`node -e ""` Is A Placeholder (`src/stack-detector.mjs`)** — the same no-op, spelled to look like work.
29
+ - [x] **`init` Rejects An Oracle That Proves Nothing (`src/wizard-init.mjs`)** — silence at setup is silence at every gate after it.
30
+
23
31
  ### v0.68.0: Not An Approval Either
24
32
  - [x] **An Unreadable Dialect Blocks (`src/security.mjs`)** — the guard said it had not checked, and approved anyway.
25
33
  - [x] **A Command That Cannot Fail Is Not An Oracle (`src/engine.mjs`)** — `task create` already refused what the gate accepted.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "jules-orchestrator-kit",
3
- "version": "0.68.0",
3
+ "version": "0.69.0",
4
4
  "description": "Zero-dependency safety gatekeeper, test oracle generator, and multi-agent coordination protocol for autonomous coding agents — Google Jules, Claude Code, Codex and Gemini CLI.",
5
5
  "repository": {
6
6
  "type": "git",
package/src/config.mjs CHANGED
@@ -626,7 +626,13 @@ export function loadConfig(root = resolveRoot(), explicitPath = null) {
626
626
  const rawTeardown = parsed.teardown_cmd ?? parsed.verify?.teardown;
627
627
  const rawBuild = parsed.build_cmd ?? parsed.verify?.build;
628
628
  const rawUnit = parsed.verify?.unit;
629
- const verifyTimeoutMs = parsed.verify?.timeoutMs ?? parsed.verify?.timeout_ms ?? 60000;
629
+ // Five minutes, matching the value `init` writes and the README documents.
630
+ // The engine carries the same number as a fallback, but this loader always
631
+ // supplies one — so the fallback never fired, and every repository without
632
+ // an explicit `timeout_ms` kept the old one-minute limit while the
633
+ // changelog said otherwise. The default has to live where the value is
634
+ // resolved, not where it is consumed.
635
+ const verifyTimeoutMs = parsed.verify?.timeoutMs ?? parsed.verify?.timeout_ms ?? 300_000;
630
636
  const autoVerify = resolveVerify(root);
631
637
 
632
638
  const rawTier = String(process.env.JULES_TIER || parsed.tier || FALLBACK_TIER).toLowerCase();
@@ -697,7 +703,7 @@ export function loadConfig(root = resolveRoot(), explicitPath = null) {
697
703
  // Named so `doctor` can tell an operator that the gate they enabled is
698
704
  // running one fewer check than they think, and why.
699
705
  profileSkipped: profilePlan?.skipped ?? [],
700
- timeoutMs: Number.isFinite(Number(verifyTimeoutMs)) ? Number(verifyTimeoutMs) : 60000,
706
+ timeoutMs: Number.isFinite(Number(verifyTimeoutMs)) ? Number(verifyTimeoutMs) : 300_000,
701
707
  },
702
708
  evidence: {
703
709
  enabled: parsed.evidence?.enabled ?? true,
package/src/engine.mjs CHANGED
@@ -393,7 +393,11 @@ export async function gate(opts = {}) {
393
393
  }
394
394
 
395
395
  let flakyVerdictResult = null;
396
- const verifyTimeout = trustedVerify.timeoutMs || 60000;
396
+ // Five minutes. A minute was too short for an ordinary suite — a mid-sized
397
+ // Node repository takes longer than that on cold caches — and the timeout
398
+ // surfaced as a plain non-zero verification, so the gate told the user their
399
+ // tests failed when it had killed them. Raise it with verify.timeout_ms.
400
+ const verifyTimeout = trustedVerify.timeoutMs || 300_000;
397
401
 
398
402
  // Monorepo scoping: run the suites the change can actually break.
399
403
  //
package/src/git.mjs CHANGED
@@ -15,6 +15,10 @@ export class GateError extends Error {
15
15
  }
16
16
  }
17
17
 
18
+ // `NAME=value` at the head of a command line: a POSIX environment prefix,
19
+ // not the name of a program.
20
+ const ENV_ASSIGNMENT = /^[A-Za-z_][A-Za-z0-9_]*=/;
21
+
18
22
  const DEFAULT_TIMEOUT = 10 * 60 * 1000; // 10 minutes
19
23
  const DEFAULT_MAX_BUFFER = 10 * 1024 * 1024; // 10 MB
20
24
 
@@ -139,6 +143,8 @@ export function runCmd(command, opts = {}) {
139
143
  let args = [];
140
144
  let useShell = false;
141
145
  let shellCmd = "";
146
+ /** @type {Record<string,string>|null} */
147
+ let envOverlay = null;
142
148
  // True when `args` came from splitting a whitespace-separated string, which
143
149
  // means no element can itself contain whitespace. See the Windows note below.
144
150
  let tokenized = false;
@@ -153,6 +159,24 @@ export function runCmd(command, opts = {}) {
153
159
  shellCmd = trimmed;
154
160
  } else {
155
161
  const tokens = trimmed.split(/\s+/).filter(Boolean);
162
+
163
+ // A POSIX command may be prefixed with environment assignments, and
164
+ // half the Python world writes its test command exactly that way:
165
+ // `PYTHONPATH=src python3 -m pytest`. execFileSync took the first token
166
+ // as the program and failed with `spawnSync PYTHONPATH=src ENOENT` —
167
+ // which the gate then reported as a failed verification, so the user
168
+ // was told their tests broke when the command had never started.
169
+ //
170
+ // Peeling the assignments into the child's environment behaves the same
171
+ // on every platform. Handing the string to a shell would not: cmd.exe
172
+ // has no such syntax, so the Windows matrix would keep failing.
173
+ while (tokens.length > 1 && ENV_ASSIGNMENT.test(tokens[0])) {
174
+ const token = tokens.shift();
175
+ const eq = token.indexOf("=");
176
+ if (!envOverlay) envOverlay = {};
177
+ envOverlay[token.slice(0, eq)] = token.slice(eq + 1);
178
+ }
179
+
156
180
  binary = tokens[0] || "";
157
181
  args = tokens.slice(1);
158
182
  tokenized = true;
@@ -177,10 +201,12 @@ export function runCmd(command, opts = {}) {
177
201
  // wrapped in cmd.exe with every element quoted per the C-runtime + cmd.exe
178
202
  // rules, which preserves argv exactly while still letting the shell resolve
179
203
  // the `.cmd` shim (P-07).
204
+ const childEnv = envOverlay ? { ...(opts.env || process.env), ...envOverlay } : opts.env || process.env;
205
+
180
206
  const winShim = tokenized && process.platform === "win32";
181
207
  const winSpawn =
182
208
  !useShell && !winShim && process.platform === "win32"
183
- ? resolveWindowsSpawn(binary, args, opts.env || process.env)
209
+ ? resolveWindowsSpawn(binary, args, childEnv)
184
210
  : null;
185
211
 
186
212
  if (!binary && !useShell) {
@@ -194,7 +220,7 @@ export function runCmd(command, opts = {}) {
194
220
  cwd,
195
221
  encoding: "utf-8",
196
222
  stdio: ["ignore", "pipe", "pipe"],
197
- env: opts.env || process.env,
223
+ env: childEnv,
198
224
  timeout,
199
225
  maxBuffer,
200
226
  })
@@ -204,7 +230,7 @@ export function runCmd(command, opts = {}) {
204
230
  shell: winShim,
205
231
  windowsVerbatimArguments: Boolean(winSpawn && winSpawn.verbatim),
206
232
  stdio: ["ignore", "pipe", "pipe"],
207
- env: opts.env || process.env,
233
+ env: childEnv,
208
234
  timeout,
209
235
  maxBuffer,
210
236
  });
@@ -218,8 +244,16 @@ export function runCmd(command, opts = {}) {
218
244
  let stdout = (err.stdout || "").toString().trim();
219
245
  let stderr = (err.stderr || err.message || "").toString().trim();
220
246
 
221
- if (isTimeout && !stderr.includes("ETIMEDOUT")) {
222
- stderr = `Command execution timed out after ${timeout}ms (ETIMEDOUT)${stderr ? "\n" + stderr : ""}`;
247
+ if (isTimeout) {
248
+ // Node's own message for this is `spawnSync sh ETIMEDOUT`, which already
249
+ // contains the token the old guard tested for — so the explanation was
250
+ // skipped exactly when it was needed, and the user was left with five
251
+ // words that name neither the limit nor the way to raise it.
252
+ stderr =
253
+ `Command execution timed out after ${timeout}ms (ETIMEDOUT). ` +
254
+ `The command was killed, not failed: raise verify.timeout_ms in ` +
255
+ `.agent/config.yml if this suite legitimately takes longer.` +
256
+ (stderr ? "\n" + stderr : "");
223
257
  }
224
258
  if (isNobufs && !stderr.includes("ENOBUFS")) {
225
259
  stderr = `Command output buffer exceeded limit of ${maxBuffer} bytes (ENOBUFS)${stderr ? "\n" + stderr : ""}`;
@@ -68,6 +68,62 @@ export const TEST_PATH_CASES = [
68
68
  * the guard the layout the implementation was written from.
69
69
  */
70
70
  export const TAMPER_CANARIES = [
71
+ // Argument order is not a detail. JUnit and PHPUnit document
72
+ // `assertEquals(expected, actual)`, and Python's containment helpers take
73
+ // the expected member first — all of which the guard used to read as "a
74
+ // human message was reworded" and dismiss. Each of these was a silent PASS.
75
+ {
76
+ id: "expectation/junit-expected-first",
77
+ file: "src/test/java/CalcTest.java",
78
+ removed: [' assertEquals("Hello World", out);'],
79
+ added: [' assertEquals("Hello Tampered", out);'],
80
+ expect: "ASSERTION_EXPECTATION_CHANGED",
81
+ },
82
+ {
83
+ id: "expectation/phpunit-expected-first",
84
+ file: "tests/CalcTest.php",
85
+ removed: [' $this->assertSame("Hello World", $out);'],
86
+ added: [' $this->assertSame("Hello Tampered", $out);'],
87
+ expect: "ASSERTION_EXPECTATION_CHANGED",
88
+ },
89
+ {
90
+ id: "expectation/unittest-assert-in",
91
+ file: "tests/test_cli.py",
92
+ removed: [' self.assertIn("Hello World", result.output)'],
93
+ added: [' self.assertIn("Hello Tampered", result.output)'],
94
+ expect: "ASSERTION_EXPECTATION_CHANGED",
95
+ },
96
+ {
97
+ id: "expectation/unittest-assert-not-in",
98
+ file: "tests/test_cli.py",
99
+ removed: [' self.assertNotIn("error", result.output)'],
100
+ added: [' self.assertNotIn("zzz", result.output)'],
101
+ expect: "ASSERTION_EXPECTATION_CHANGED",
102
+ },
103
+ // A regex literal is an expected value. It never blanked, so the rewritten
104
+ // pattern never met its original in a shape bucket and the swap was
105
+ // reported as neither a change nor a loss.
106
+ {
107
+ id: "expectation/jest-to-match-regex",
108
+ file: "test/cli.test.js",
109
+ removed: [" expect(out).toMatch(/Hello World/);"],
110
+ added: [" expect(out).toMatch(/Hello Tampered/);"],
111
+ expect: "ASSERTION_EXPECTATION_CHANGED",
112
+ },
113
+ {
114
+ id: "expectation/jest-to-throw-regex",
115
+ file: "test/cli.test.js",
116
+ removed: [" expect(fn).toThrow(/permission denied/);"],
117
+ added: [" expect(fn).toThrow(/x/);"],
118
+ expect: "ASSERTION_EXPECTATION_CHANGED",
119
+ },
120
+ {
121
+ id: "expectation/rspec-match-regex",
122
+ file: "spec/cli_spec.rb",
123
+ removed: [" expect(out).to match(/Hello World/)"],
124
+ added: [" expect(out).to match(/Hello Tampered/)"],
125
+ expect: "ASSERTION_EXPECTATION_CHANGED",
126
+ },
71
127
  {
72
128
  id: "skip-injection/node",
73
129
  file: "test/calc.test.js",
@@ -398,6 +454,41 @@ export const SCOPE_CANARIES = [
398
454
  * from never showed it.
399
455
  */
400
456
  export const INNOCENT_EDITS = [
457
+ // Renaming a test is one of the most ordinary edits there is. On the
458
+ // one-line form the name blanked to the same shape as its replacement, the
459
+ // two paired, and the pair was reported as a rewritten expectation.
460
+ {
461
+ id: "rename-test-oneline/node",
462
+ file: "test/calc.test.js",
463
+ context: "// arithmetic",
464
+ removed: ['test("adds", () => { assert.strictEqual(add(2, 3), 5); });'],
465
+ added: ['test("adds positives", () => { assert.strictEqual(add(2, 3), 5); });'],
466
+ why: "the test name is prose about the test, not a value it asserts",
467
+ },
468
+ {
469
+ id: "rename-test-oneline/jest",
470
+ file: "test/calc.spec.js",
471
+ context: "// arithmetic",
472
+ removed: ['it("adds", () => { expect(add(2, 3)).toBe(5); });'],
473
+ added: ['it("adds two positives", () => { expect(add(2, 3)).toBe(5); });'],
474
+ why: "same rename, the dialect a new user is most likely to be in",
475
+ },
476
+ {
477
+ id: "rename-test-oneline/tap",
478
+ file: "test/calc.test.js",
479
+ context: "// arithmetic",
480
+ removed: ['t.test("adds", (ct) => { ct.equal(add(2, 3), 5); ct.end(); });'],
481
+ added: ['t.test("adds positives", (ct) => { ct.equal(add(2, 3), 5); ct.end(); });'],
482
+ why: "the receiver is a sub-test callback argument, not an assertion subject",
483
+ },
484
+ {
485
+ id: "reword-junit-message-first",
486
+ file: "src/test/java/CalcTest.java",
487
+ context: "// arithmetic",
488
+ removed: [' assertEquals("why", expected, out);'],
489
+ added: [' assertEquals("why this matters", expected, out);'],
490
+ why: "JUnit 4 puts the message first — rewording it changes no expectation",
491
+ },
401
492
  {
402
493
  id: "add-assertion-beside-comment/node",
403
494
  file: "test/calc.test.js",
@@ -50,6 +50,27 @@ const COUNT_PATTERNS = [
50
50
  { name: "dotnet", re: /\bTotal(?:\s+tests)?:\s*(\d+)/i },
51
51
  // swift test / XCTest — "Executed 12 tests"
52
52
  { name: "xctest", re: /\bExecuted\s+(\d+)\s+tests?/i },
53
+ // Every runner below states a count the gate used to read as silence, and
54
+ // silence is what makes `pnpm -r test` on a workspace with no package-level
55
+ // test script indistinguishable from a suite that ran. The floor stays
56
+ // one-sided — an unrecognised runner still passes — so widening the list
57
+ // only ever converts an "I could not tell" into an answer.
58
+ // Python unittest — "Ran 12 tests in 0.003s"
59
+ { name: "unittest", re: /^Ran\s+(\d+)\s+tests?\s+in\b/m },
60
+ // GoogleTest — "[==========] 12 tests from 3 test suites ran."
61
+ { name: "gtest", re: /\[=+\]\s+(\d+)\s+tests?\s+from\b/ },
62
+ // Catch2 — "test cases: 12 | 12 passed"
63
+ { name: "catch2", re: /\btest cases:\s*(\d+)/i },
64
+ // deno test — "ok | 12 passed | 0 failed"
65
+ { name: "deno", re: /\|\s*(\d+)\s+passed\s*\|/ },
66
+ // bun test — "12 pass"
67
+ { name: "bun", re: /^\s*(\d+)\s+pass\s*$/m },
68
+ // ava — "12 tests passed"
69
+ { name: "ava", re: /^\s*(\d+)\s+tests?\s+passed\s*$/m },
70
+ // A bare TAP plan, which is how tape, bats and the TAP producers that
71
+ // print no summary state their count. Last of the count patterns: node:test
72
+ // emits TAP too, and its own line above is the more specific reading.
73
+ { name: "tap", re: /^1\.\.(\d+)\s*$/m },
53
74
  ];
54
75
 
55
76
  /** Per-test lines, which `go test` only prints under -v. */
package/src/security.mjs CHANGED
@@ -1219,6 +1219,21 @@ const isSpecificAssertion = (str) => SPECIFIC_ASSERTION.test(str);
1219
1219
  const blankLiterals = (str) =>
1220
1220
  str
1221
1221
  .replace(/(['"`])(?:\\.|(?!\1)[^\\])*\1/g, "\u0000S")
1222
+ // A regex literal is an expected value like any other. Without this,
1223
+ // `toMatch(/Hello World/)` and `toMatch(/Hello Tampered/)` normalized to
1224
+ // two different shapes, never met in a bucket, and the rewrite was
1225
+ // reported as neither a change nor a loss — one specific assertion out,
1226
+ // one in, and silence. Runs after the string pass so a `/` inside a
1227
+ // string is already gone, and before the number pass so a pattern
1228
+ // containing digits collapses whole.
1229
+ //
1230
+ // The lookbehind is what separates a regex from a division: an operand
1231
+ // never precedes `/` here, only an opening paren, a comma, or an
1232
+ // operator, which is where a test's expected pattern actually sits.
1233
+ .replace(
1234
+ /(?<=[(,=:[!&|?{;]\s{0,8})\/(?![*/])(?:\\.|\[(?:\\.|[^\]\\\n])*\]|[^/\\\n])+\/[dgimsuvy]*/g,
1235
+ "\u0000R"
1236
+ )
1222
1237
  // The sign belongs to the literal: without it `3` and `-1` normalized to
1223
1238
  // different shapes and the rewritten expectation was never paired.
1224
1239
  .replace(
@@ -1802,8 +1817,22 @@ function isPureStringLiteral(arg) {
1802
1817
  */
1803
1818
  function messageArgIndices(args) {
1804
1819
  const idx = new Set();
1805
- if (args.length >= 3 && isPureStringLiteral(args[args.length - 1])) idx.add(args.length - 1);
1806
- if (args.length >= 2 && isPureStringLiteral(args[0])) idx.add(0);
1820
+ const lastIsMessage = args.length >= 3 && isPureStringLiteral(args[args.length - 1]);
1821
+ if (lastIsMessage) idx.add(args.length - 1);
1822
+
1823
+ // JUnit 4 is the one common dialect that puts the message *first*:
1824
+ // `assertEquals("why this matters", expected, actual)`. It is also
1825
+ // distinguishable, because its trailing argument is the actual value rather
1826
+ // than prose — so a call that already carries a trailing message is not
1827
+ // that shape, whatever its first argument looks like.
1828
+ //
1829
+ // Reading argument 0 as prose whenever it happened to be a string is what
1830
+ // made a whole family of assertions invisible: `assertEquals(expected,
1831
+ // actual)` — JUnit's and PHPUnit's own two-argument order — along with
1832
+ // Python's `assertIn(member, container)` and `assertNotIn`. A rewritten
1833
+ // expectation in any of them was dismissed as a reworded message, and the
1834
+ // guard reported PASS on a check it had not performed.
1835
+ if (!lastIsMessage && args.length >= 3 && isPureStringLiteral(args[0])) idx.add(0);
1807
1836
  return idx;
1808
1837
  }
1809
1838
 
@@ -1846,6 +1875,63 @@ function splitTrailingMessage(clean) {
1846
1875
  return { head: clean.slice(0, lastComma), msg: tail };
1847
1876
  }
1848
1877
 
1878
+ // A test declaration whose first argument is the test's name. The name is
1879
+ // prose about the test, not a value the test asserts — `test("adds", ...)`
1880
+ // renamed to `test("adds positives", ...)` is the rename the diff says it is.
1881
+ const TEST_DECL_CALL =
1882
+ /\b(?:it|test|describe|context|suite|specify|bench|scenario)\s*\(|\b[a-z_$][\w$]{0,2}\.(?:test|Run|describe)\s*\(/;
1883
+
1884
+ /**
1885
+ * Replace the name argument of a test declaration with a placeholder.
1886
+ *
1887
+ * @param {string} clean - comment-stripped statement text
1888
+ * @returns {string|null} the text with the name blanked, or null when the
1889
+ * statement is not a test declaration with a literal name.
1890
+ */
1891
+ function blankTestName(clean) {
1892
+ const m = TEST_DECL_CALL.exec(clean);
1893
+ if (!m) return null;
1894
+
1895
+ let i = m.index + m[0].length;
1896
+ while (i < clean.length && /\s/.test(clean[i])) i++;
1897
+ const quote = clean[i];
1898
+ if (quote !== '"' && quote !== "'" && quote !== "`") return null;
1899
+
1900
+ let j = i + 1;
1901
+ while (j < clean.length) {
1902
+ if (clean[j] === "\\") { j += 2; continue; }
1903
+ if (clean[j] === quote) break;
1904
+ j += 1;
1905
+ }
1906
+ if (j >= clean.length) return null;
1907
+
1908
+ return clean.slice(0, i) + "\u0000T" + clean.slice(j + 1);
1909
+ }
1910
+
1911
+ /**
1912
+ * True when the only thing that changed was the test's name.
1913
+ *
1914
+ * A one-line `test("adds", () => { assert.strictEqual(add(2, 3), 5); });`
1915
+ * blanks to the same shape as its renamed copy, so the two pair — and the
1916
+ * pair was then reported as a rewritten expectation, quoting the whole line
1917
+ * back at an author who had renamed a test and nothing else. Renaming a test
1918
+ * is one of the most ordinary edits there is, and a gate that calls it
1919
+ * tampering is a gate that gets switched off.
1920
+ *
1921
+ * The multi-line form was never affected: NON_JOINING_CALL already keeps a
1922
+ * test name from joining to the assertion below it. This is the same rule for
1923
+ * the statements that fit on one line.
1924
+ *
1925
+ * A rename that also moves the expectation still differs after the name is
1926
+ * blanked, so it is still reported.
1927
+ */
1928
+ function differsOnlyInTestName(cleanRemoved, cleanAdded) {
1929
+ const a = blankTestName(cleanRemoved);
1930
+ const b = blankTestName(cleanAdded);
1931
+ if (a === null || b === null) return false;
1932
+ return a.replace(/\s+/g, "") === b.replace(/\s+/g, "");
1933
+ }
1934
+
1849
1935
  function differsOnlyInMessage(cleanRemoved, cleanAdded, lang) {
1850
1936
  const ta = splitTrailingMessage(cleanRemoved);
1851
1937
  const tb = splitTrailingMessage(cleanAdded);
@@ -1874,6 +1960,17 @@ function differsOnlyInMessage(cleanRemoved, cleanAdded, lang) {
1874
1960
  return sawDifference;
1875
1961
  }
1876
1962
 
1963
+ /**
1964
+ * True when a paired difference is something other than a rewritten
1965
+ * expectation — a reworded message, or a renamed test.
1966
+ */
1967
+ function isNonExpectationDifference(cleanRemoved, cleanAdded, lang) {
1968
+ return (
1969
+ differsOnlyInMessage(cleanRemoved, cleanAdded, lang) ||
1970
+ differsOnlyInTestName(cleanRemoved, cleanAdded)
1971
+ );
1972
+ }
1973
+
1877
1974
  /**
1878
1975
  * Pair rewritten expectations across the removed and added images of every
1879
1976
  * hunk of one file, and report each pair.
@@ -1988,7 +2085,7 @@ function detectExpectationRewrites(file, hunks, stats, violations) {
1988
2085
  pairedOld.add(r);
1989
2086
  pairedNew.add(a);
1990
2087
  if (r.canon === a.canon) continue;
1991
- if (differsOnlyInMessage(r.clean, a.clean, lang)) continue;
2088
+ if (isNonExpectationDifference(r.clean, a.clean, lang)) continue;
1992
2089
  pairs.push({ r: r.s, a: a.s });
1993
2090
  }
1994
2091
  }
@@ -2018,7 +2115,7 @@ function detectExpectationRewrites(file, hunks, stats, violations) {
2018
2115
  // different assertions that merely resemble each other.
2019
2116
  if (!same(0)) continue;
2020
2117
  if (ra.every((_, i) => same(i))) continue;
2021
- if (differsOnlyInMessage(r.clean, a.clean, lang)) continue;
2118
+ if (isNonExpectationDifference(r.clean, a.clean, lang)) continue;
2022
2119
  pairs.push({ r: r.s, a: a.s });
2023
2120
  pairedOld.add(r);
2024
2121
  pairedNew.add(a);
@@ -2039,7 +2136,7 @@ function detectExpectationRewrites(file, hunks, stats, violations) {
2039
2136
  if (sr === sa && hasLiteralPlaceholder(sr)) {
2040
2137
  const clr = stripComments(r.text, lang);
2041
2138
  const cla = stripComments(a.text, lang);
2042
- if (clr.replace(/\s+/g, "") !== cla.replace(/\s+/g, "") && !differsOnlyInMessage(clr, cla, lang)) {
2139
+ if (clr.replace(/\s+/g, "") !== cla.replace(/\s+/g, "") && !isNonExpectationDifference(clr, cla, lang)) {
2043
2140
  pairs.push({ r, a });
2044
2141
  }
2045
2142
  }
@@ -66,10 +66,20 @@ function makefileHasTestTarget(root) {
66
66
  * placeholder by this definition, and correctly so: it exits non-zero, which
67
67
  * fails loudly rather than certifying nothing.
68
68
  */
69
+ // An interpreter handed an empty program: `node -e ""`, `python3 -c ''`,
70
+ // `sh -c ""`. It reads as a test command at a glance and runs nothing.
71
+ //
72
+ // This does not enumerate every way to write a no-op — nothing can, and the
73
+ // collection floor is what reports the general case, one-sidedly and by
74
+ // design. It closes the spellings that look like work.
75
+ const EMPTY_PROGRAM =
76
+ /^(?:\S*[/\\])?(?:node|nodejs|deno|bun|python[\d.]*|ruby|perl|php|sh|bash|zsh|dash)(?:\s+-{1,2}[\w-]+)*\s+(?:-e|-c|--eval|--execute)(?:\s*(?:""|''|``))?\s*$/;
77
+
69
78
  export function isPlaceholderTestScript(cmd) {
70
79
  if (typeof cmd !== "string") return false;
71
80
  const trimmed = cmd.trim();
72
81
  if (!trimmed) return true;
82
+ if (EMPTY_PROGRAM.test(trimmed)) return true;
73
83
  // Drop the announcements; what matters is what the shell is left doing.
74
84
  const remainder = trimmed
75
85
  .split(/&&|;/)
@@ -6,6 +6,7 @@ import { detectDefaultBranch } from "./git.mjs";
6
6
  import { resolveWorkspaceBoundary, oracleCandidates } from "./stack-detector.mjs";
7
7
  import { PROFILE_NAMES, PROFILE_DESCRIPTIONS } from "./profiles.mjs";
8
8
  import { detectStackOracles, runVerificationProbe } from "./wizard-oracle.mjs";
9
+ import { parseCollectedTests } from "./ops/test-collection.mjs";
9
10
  import { select, multiSelect, input, confirm, spinner, isTTY } from "./tui.mjs";
10
11
  import { KIT_VERSION } from "./version.mjs";
11
12
 
@@ -215,6 +216,9 @@ verify:
215
216
  # to their sub-projects and runs only those suites (monorepos)
216
217
  scope: ${verifyScope}
217
218
  test: "${verify.test}"
219
+ # How long a verification stage may run before the gate kills it (default
220
+ # 300000). Raise it for a suite that legitimately takes longer.
221
+ timeout_ms: 300000
218
222
  build: "${verify.build}"
219
223
  lint: "${verify.lint}"
220
224
  typecheck: "${verify.typecheck}"
@@ -314,25 +318,77 @@ export function loadPresets(root = process.cwd()) {
314
318
  *
315
319
  * @returns {Promise<string>} the command to save
316
320
  */
321
+ /**
322
+ * How much a probe actually proved.
323
+ *
324
+ * Exit 0 is the weakest of the three answers. `pnpm -r test` on a workspace
325
+ * whose packages declare no test script exits 0, prints nothing, and runs
326
+ * nothing — and it was accepted here as a verified oracle, which left the
327
+ * repository configured to approve every future change against silence.
328
+ *
329
+ * Choosing a command is the right moment to be strict about this: at init the
330
+ * cost of rejecting a candidate is trying the next one, where at gate time it
331
+ * would be a hard red on a repository that is fine.
332
+ */
333
+ function probeVerdict(probeRes) {
334
+ if (!probeRes.ok) return "failed";
335
+ const { count } = parseCollectedTests(probeRes.stdout, probeRes.stderr);
336
+ if (count === null) return "silent";
337
+ return count > 0 ? "ran" : "empty";
338
+ }
339
+
317
340
  async function resolveRunnableOracle(root, testCmd, options = {}) {
318
341
  if (!testCmd) return testCmd;
319
342
  const probeSp = spinner(`Probing oracle: ${testCmd}`, options);
320
343
  const probeRes = await runVerificationProbe(testCmd, root);
321
- if (probeRes.ok) {
344
+ const verdict = probeVerdict(probeRes);
345
+ if (verdict === "ran") {
322
346
  probeSp.stop(`Oracle verified successfully (${probeRes.durationMs}ms)`);
323
347
  return testCmd;
324
348
  }
325
- probeSp.fail(`Oracle verification probe failed (Exit ${probeRes.code})`);
349
+ if (verdict === "failed") {
350
+ probeSp.fail(`Oracle verification probe failed (Exit ${probeRes.code})`);
351
+ } else {
352
+ probeSp.fail(
353
+ verdict === "empty"
354
+ ? `"${testCmd}" exited 0 but ran no tests — looking for a command that does`
355
+ : `"${testCmd}" exited 0 without stating how many tests it ran — looking for a better one`
356
+ );
357
+ }
326
358
 
327
359
  const alternates = oracleCandidates(root, testCmd).filter((c) => c !== testCmd).slice(0, 3);
360
+ // Two passes: prefer a command that proves it ran something, and only then
361
+ // settle for one that merely exits 0. Falling back on the first exit-0
362
+ // candidate would reintroduce exactly the silence this rejects.
363
+ const settled = [];
328
364
  for (const cand of alternates) {
329
365
  const altSp = spinner(`Trying ${cand}`, options);
330
366
  const altRes = await runVerificationProbe(cand, root);
331
- if (altRes.ok) {
367
+ const altVerdict = probeVerdict(altRes);
368
+ if (altVerdict === "ran") {
332
369
  altSp.stop(`${cand} runs here (${altRes.durationMs}ms) — using it instead`);
333
370
  return cand;
334
371
  }
335
- altSp.fail(`${cand} also failed (Exit ${altRes.code})`);
372
+ if (altVerdict === "silent") {
373
+ altSp.fail(`${cand} exited 0 without stating a test count`);
374
+ settled.push(cand);
375
+ continue;
376
+ }
377
+ altSp.fail(altVerdict === "empty" ? `${cand} ran no tests` : `${cand} also failed (Exit ${altRes.code})`);
378
+ }
379
+
380
+ // Nothing proved it ran a suite. A command that at least starts and exits 0
381
+ // still beats one that does not run at all, so it is used — and said out
382
+ // loud, because the gate will only be able to check it by exit code.
383
+ if (settled.length > 0 || verdict === "silent") {
384
+ const chosen = verdict === "silent" ? testCmd : settled[0];
385
+ const out = options.stdout || process.stdout;
386
+ out.write("\n");
387
+ out.write(` \u26a0\ufe0f "${chosen}" exits 0 but states no test count.\n`);
388
+ out.write(" The gate can verify it by exit code alone, which cannot tell a\n");
389
+ out.write(" full suite from a command that ran nothing. If this repository\n");
390
+ out.write(" has a suite, point verify.test at it in .agent/config.yml.\n\n");
391
+ return chosen;
336
392
  }
337
393
 
338
394
  // Nothing runs. Say so in terms the user can act on, rather than leaving a