jules-orchestrator-kit 0.70.0 → 0.71.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +36 -0
- package/README.md +1 -1
- package/ROADMAP_V1.md +12 -3
- package/bin/agentctl.mjs +54 -0
- package/package.json +1 -1
- package/scripts/guard-reach-check.mjs +46 -3
- package/src/config.mjs +88 -3
- package/src/engine.mjs +22 -1
- package/src/guard-policy.mjs +174 -0
- package/src/ops/test-collection.mjs +74 -0
- package/src/security.mjs +113 -2
- package/src/stack-detector.mjs +52 -14
- package/src/wizard-init.mjs +38 -15
- package/src/wizard-task.mjs +15 -1
package/CHANGELOG.md
CHANGED
|
@@ -5,6 +5,40 @@ All notable changes to this project will be documented in this file.
|
|
|
5
5
|
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
|
|
6
6
|
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
7
7
|
|
|
8
|
+
## [0.71.0] - 2026-09-05
|
|
9
|
+
*A blanket is not a check, and silence is not a suite.*
|
|
10
|
+
|
|
11
|
+
A fourth cold-start trial, against v0.70.0 on four repositories nobody here chose. Seven findings, and this is the first trial where every one of them held: no quoted output was fabricated, and all four repository SHAs matched. Two of them prove entries in this file wrong, which is recorded below rather than quietly amended.
|
|
12
|
+
|
|
13
|
+
The two worst were the two halves of the same sentence — a gate that approves against nothing, and a gate that rejects what is fine.
|
|
14
|
+
|
|
15
|
+
### Fixed
|
|
16
|
+
- **A Command That Printed Nothing Was An Oracle (`src/ops/test-collection.mjs`, `src/wizard-init.mjs`)**: `pnpm -r test` on a workspace whose packages declare no test script exits 0 and writes **zero bytes** on stdout and stderr. Measured on `unjs/unimport`: `init` chose it, and `agentctl check` returned `APPROVED (Exit 0)` with an advisory that it could not tell a full suite from a command that ran nothing. The collection floor is one-sided on purpose — an unrecognised runner states no count and passes, because hard-redding every runner not on the list would be worse than the hole it closes — but an unrecognised runner still *prints*. Zero bytes is not a dialect the list has yet to learn; it is a command that ran nothing, and it is now the one absence that decides. v0.69.0 claimed `init` already rejected this candidate: it did not, because `probeVerdict` sorted no-output into `silent` (the forgiving bucket) rather than `empty`, so the exact command that verdict was introduced to reject was the one case it waved through. The rule reads the command as well as the output, and both halves are required: `tsc --noEmit`, `node --check`, `go vet` and `python3 -m compileall` also exit 0 in silence, and two of those are commands this kit writes itself for a repository with no suite yet — keying on silence alone hard-redded every one of them, which is the same first-run rejection of correct code the floor is otherwise careful to avoid. So a command that is *recognisably* a suite invocation and printed nothing ran no suite, while a checker that printed nothing did what it promised. Resolved once, in `producedNoOutput` and `looksLikeTestSuiteCommand`, and consumed by both the floor and the probe.
|
|
17
|
+
- **Python Was Verified Against A Different Copy Of Itself (`src/stack-detector.mjs`)**: `python3 -m pytest` puts the *working directory* on `sys.path`, which is the fix for a module at the root and no help at all when the package lives under `src/` — there, `import iniconfig` finds nothing locally and falls through to site-packages. Measured on `pytest-dev/iniconfig` with `_parse.py` gutted to `return False`: 49 tests passed and the gate returned `APPROVED (Exit 0)`. The suite really did run, so no amount of counting collected tests can see it; only the import path can. A `src/` layout is now detected by the `__init__.py` under it and the oracle is written as `PYTHONPATH=src python3 -m pytest`, which puts the working tree first whether or not the package is also installed.
|
|
18
|
+
- **Ordinary Work On A Test File Was A CRITICAL Block (`src/security.mjs`)**: `assertionsSeen === 0 && (unreadable.length > 0 || examined > 0)` decided `UNREADABLE`, and that `|| examined > 0` threw away the distinction the whole apparatus exists to draw. `ASSERTION_SHAPED` and `unreadable[]` were built to separate *"assertion-shaped lines were present and none parsed"* from *"there were no assertions in these lines at all"*. Collapsed together, **any** changed substantive line in a test file with no recognised assertion became a `CRITICAL` rejection: renaming a test function did it, and so did adding `import os`. The tell was in the finding itself — `file: null`, `line: null`, no sample, because `unreadable` was empty — so the guard blocked while holding no evidence of anything, and advised a **pytest** repository that its assertion library might be unsupported, from a list that names pytest. Worse, the remedy it printed was `verify.tamperGuard: "warn"`, which switches the real checks off along with the blanket. `UNREADABLE` now requires its evidence. Nothing is weakened: a removed or rewritten assertion is a recognised assertion line, raises `assertionsSeen`, and never reached this branch.
|
|
19
|
+
- **The Scaffold Failed The Repository's Own Linter (`src/config.mjs`, `src/wizard-init.mjs`, `src/stack-detector.mjs`)**: `init` wrapped every generated string in double quotes — valid YAML, idiomatic nowhere. Measured on `unjs/unimport`: following the README verbatim produced **40 eslint errors** (`yaml/quotes`, `yaml/plain-scalar`) across `.agent/config.yml` and `.agent/jules.yml`, so a newcomer's first `agentctl check` was a hard red on files `agentctl init` had written a minute earlier, in a repository that was green before. Two causes, and the second was the better one. `yamlScalar` now emits a plain scalar where YAML permits one and single quotes where it does not, which is simply how YAML is written — a glob beginning with `*` stays quoted, because a leading star is an alias reference. That took it from 40 errors to 16. The remaining 16 came from a single generated comment beginning with the word `global`, which ESLint reads as the `/* global a, b */` directive in any language it has a parser for, YAML included — so it declared globals named after each English word of the sentence and then reported them as unused. The sentence leads with the key it explains instead. Nothing in either fix is specific to eslint or to that repository, and the emitter rule lives in one place that both writers use. Measured again on the same clone: 40 errors to 0.
|
|
20
|
+
- **A Test-Integrity Failure Told The Operator To Rotate Their Keys (`bin/agentctl.mjs`)**: `TEST_DIALECT_UNREADABLE` was missing from `INTEGRITY_TYPES`, so the one finding that is most obviously not a secret fell through to the exit-6 secret branch. The screen printed `Phase [SECRETS (test integrity — no secret found)] : ❌ FAIL` and then `💡 Remediation Hint (Exit 6 Secret Leak Prevented) • High-entropy credential or secret detected in patch • rotate any exposed keys immediately` — two statements about the same run, contradicting each other in the same output. v0.67.0 fixed the phase label and left the hint; v0.69.0 recorded the hint as already correct without measuring it. It has its own branch now, and it is derived from `onlyIntegrityFindings` rather than tested independently, so a run carrying both an unreadable file and a real tampering finding gets the tampering advice instead of falling past both into the secret branch — which is the shape of the original hole.
|
|
21
|
+
- **`task create -p` Blocked Forever In A Real Terminal (`src/wizard-task.mjs`)**: the README advertises `task create -p "..."` as the way to *"skip straight to review"*. Interactivity was decided by `isTTY` alone and `-p` was consulted only by the TODO-import branch, so in a terminal the advertised quickstart stopped at `? Task Title` and waited for a keypress that never came — then asked for the instructions it had already been handed. v0.69.0 recorded this as not reproduced, correctly, from a non-TTY run: the headless path never asks, so the defect was invisible exactly where it was measured. `-p` now means the headless path everywhere, and the regression test supplies no answers at all, because only an assertion that nothing is asked could have caught a hang.
|
|
22
|
+
|
|
23
|
+
### Added
|
|
24
|
+
- **A Test Renamed Out Of Its Runner's Discovery Is Its Own Finding (`src/security.mjs`)**: pytest collects `test_*` and Go collects `Test*`, so `def test_totals` becoming `def totals` deletes the test from every future run while leaving it in the file, fully written, assertions intact — and every count in the guard stays level. This was previously caught only as a side effect of the blanket removed above. `TEST_DEREGISTERED` names it, explains that the runner collects by name, and is waivable as `--allow-test-change deregistration` like every other kind. Exact on the remainder: `test_totals` → `totals` is a de-registration, `test_totals` → `test_totals_rounded` is a rename and stays silent.
|
|
25
|
+
- **A Loosened Run No Longer Looks Like A Strict One (`bin/agentctl.mjs`)**: every escape hatch turns a check off and the report said so nowhere — `agentctl check --allow-test-change expectation` over a rewritten expectation printed a phase list byte-identical to a clean, unloosened `APPROVED`. A reviewer, a CI log, or the author a week later could not tell that a check had been waived. An `⚠️ OVERRIDES ACTIVE` block now names each active flag, plus `verify.required: false` and `verify.tamperGuard: "warn"`, and it prints on approval as well as rejection, because approval is the case where it is load-bearing.
|
|
26
|
+
- **Eleven More Cases In The Policy Contract (`src/guard-policy.mjs`)**: five innocent edits for the work that was being rejected — adding an import, a multi-line test rename, a docstring, a type annotation, a Go fixture — four de-registration canaries across both name-driven runners, six silent-run canaries and five silent *static gates* that must keep passing — the counterweight, without which the silence rule would hard-red every honest checker. 42 canaries, 21 innocent edits, 4 de-registrations, 6 silent runs, 5 static gates.
|
|
27
|
+
- **The Innocent-Edit Contract Measures The Verdict (`test/assertion-dialects.test.mjs`, `scripts/guard-reach-check.mjs`)**: it asserted only that no *violation* was produced, and `UNREADABLE` is a status rather than a violation — so the whole list stayed green for two releases while `scanDiff` returned `CRITICAL` and the gate rejected the change. Every innocent edit is now also asserted through `scanDiff`. Anything that can end a run has to be measured where the run ends.
|
|
28
|
+
|
|
29
|
+
### Also Fixed
|
|
30
|
+
- **The Loosening Flags Reached One Phase And Not The Next (`src/engine.mjs`)**: `scanDiff` is given `allowTestChanges` and honours it; the `assert:test-integrity` verify stage received only its own stage object and re-ran the same guard with none of them. So an override was accepted by one phase and ignored by the next — `--allow-test-change deregistration` turned `SECRETS` green and then failed the run at anti-tamper, having already told the operator to pass the flag they had just passed. Found while testing the new kind, not reported by the trial. This is the eighth time in this project's history that one rule has been written in one place while another kept the old answer.
|
|
31
|
+
|
|
32
|
+
- **The Manifest Parser Truncated Any Value Containing A Hash (`src/config.mjs`)**: `parseYaml` opened a comment at the first `#` on a line, quoted or not — so `verify.test: 'pytest -k "not #slow"'` was read as `pytest -k "not`, and the gate ran a command the operator never wrote while reporting on it as if they had. A URL fragment did the same. Comment-stripping is now quote-aware and follows the YAML rule that a `#` opens a comment only at the start of a line or after whitespace, and `coerce` undoes the `''` escape that `yamlScalar` emits. Pre-existing, and it mattered more the moment this release added an emitter: a writer and a reader in one module that disagree about the same file is this project's recurring defect with both halves in view. `YAML_ROUNDTRIP_CASES` pins 19 values — hashes, colons, leading stars, apostrophes, the empty string, and the words YAML would otherwise read as booleans — through both a key and a list item.
|
|
33
|
+
|
|
34
|
+
### Corrections To Earlier Entries
|
|
35
|
+
Two `### Not Reproduced` claims in **[0.69.0]** were wrong, and both were wrong in the same way — measured somewhere the defect could not appear.
|
|
36
|
+
|
|
37
|
+
- *"The `TEST_DIALECT_UNREADABLE` remediation copy was already correct."* It was not. v0.67.0 had fixed the phase label, not the hint, and the two were on screen together. Reproduced verbatim on `pytest-dev/iniconfig` and fixed above.
|
|
38
|
+
- *"`task create -p` did not block on a prompt."* True without a TTY and false with one. Reproduced under `script(1)` and fixed above.
|
|
39
|
+
|
|
40
|
+
One finding in the fourth trial quoted output that does not reproduce: its `--allow-test-change expectation` case shows `APPROVED (Exit 0)`, where the run rejects at `Exit 4` because the reporter's own `sed` genuinely breaks the dotenv suite — the same fixture artefact this project has hit in every trial. The finding itself, that no override leaves a trace in the report, held and is fixed above.
|
|
41
|
+
|
|
8
42
|
## [0.70.0] - 2026-09-05
|
|
9
43
|
*A session that has not finished is not a session that passed.*
|
|
10
44
|
|
|
@@ -45,6 +79,8 @@ A third cold-start trial against v0.68.0. Twelve findings; six reproduced, and t
|
|
|
45
79
|
- **The Egress Guard Could Not See `.js` (`test/egress-allowlist.test.mjs`)**: the scan collected `.mjs` only, so `bin/init.js` — published as the `jules-init` binary — sat outside the boundary entirely. Harmless as it stands, but the guard's whole purpose is that a reviewer can trust the boundary without reading every commit.
|
|
46
80
|
|
|
47
81
|
### Not Reproduced
|
|
82
|
+
> **Corrected in [0.71.0]:** two of the claims below were wrong — the `TEST_DIALECT_UNREADABLE` remediation copy and `task create -p`. Both were measured where the defect could not appear. See the corrections section of that release.
|
|
83
|
+
|
|
48
84
|
Six of the twelve did not hold, and five quoted terminal output that does not exist anywhere in the shipped code. `--allow-test-change expectation` propagates correctly through both phases and returns `APPROVED`; `--allow-test-change removal` allows deleting a dead test alongside its dead function, also `APPROVED` — both reported failures were a test suite genuinely broken by the reporter's own edit, which is the same fixture artefact this project has hit repeatedly. `agentctl plan approve --dry-run` errors with *"Session ID is required"* rather than printing the quoted *"Plan Approved Successfully!"*. `task create -p` did not block on a prompt. A protected `package.json` is the design, and the gate already prints `To allow protected files in this run, pass: agentctl gate --allow-protected` — the finding stated there were no flag hints in the output. The `TEST_DIALECT_UNREADABLE` remediation copy was already correct: *"Exit 6 Test Integrity Violation — no secret was found"*, not the quoted *"Secret Leak Prevented"*. And `pnpm -r test` exits 0 on pnpm 10.33.4 rather than the reported `ERR_PNPM_RECURSIVE_RUN_NO_SCRIPT` — the real defect there was worse than the one reported, and is fixed above.
|
|
49
85
|
|
|
50
86
|
### Added
|
package/README.md
CHANGED
|
@@ -208,7 +208,7 @@ To maximize PR merge rates, dispatch tasks according to deterministic boundaries
|
|
|
208
208
|
* **Fail-Closed Security & Secret Redaction:** Evaluates explicit Deny rules before Allow rules against canonicalized, case-folded paths. Redacts high-entropy keys and base64-encoded credentials (such as Kubernetes `Secret` manifests).
|
|
209
209
|
* **Complexity & Cost Router:** Zero-dependency heuristic classifier (`src/router.mjs`) routing mechanical tasks to lightweight models while reserving primary models for complex refactors, with a `node --check` syntax-verification gate that transparently escalates a FAST-tier result to the primary provider if it left broken JS on disk.
|
|
210
210
|
* **Terminal UI & Diagnostic Matrix (`agentctl doctor`):** Interactive terminal dashboard, task sidecar manager, and automated transactional self-repair.
|
|
211
|
-
* **Verified Test Suite:** Tested with **
|
|
211
|
+
* **Verified Test Suite:** Tested with **1299 unit tests across 174 suites**, green on every supported platform.
|
|
212
212
|
|
|
213
213
|
<br/>
|
|
214
214
|
|
package/ROADMAP_V1.md
CHANGED
|
@@ -11,15 +11,24 @@ The **jules-orchestrator-kit** is the zero-dependency safety gatekeeper and self
|
|
|
11
11
|
## 📌 Release Milestones Overview
|
|
12
12
|
|
|
13
13
|
```
|
|
14
|
-
v0.
|
|
15
|
-
(
|
|
14
|
+
v0.71.0 (Current Stable) ──► v0.72.0 (Distributed Swarms & Leases) ──► v1.0.0 (Production Hardened Kernel)
|
|
15
|
+
(A Blanket Is Not A Check) (Multi-Agent DAG & Resource Locks) (Enterprise Telemetry & SLA)
|
|
16
16
|
```
|
|
17
17
|
|
|
18
18
|
---
|
|
19
19
|
|
|
20
|
-
## ✅ Shipped Milestones (v0.20.0 – v0.
|
|
20
|
+
## ✅ Shipped Milestones (v0.20.0 – v0.71.0)
|
|
21
21
|
|
|
22
22
|
|
|
23
|
+
### v0.71.0: A Blanket Is Not A Check
|
|
24
|
+
- [x] **Silence Is Not A Suite (`src/ops/test-collection.mjs`, `src/wizard-init.mjs`)** — a command that claims to run tests and prints nothing ran none; a static gate that prints nothing did its job.
|
|
25
|
+
- [x] **Python Verified Against The Working Tree (`src/stack-detector.mjs`)** — a `src/` layout resolved its imports to site-packages, and broken code passed 49 tests.
|
|
26
|
+
- [x] **`UNREADABLE` Requires Its Evidence (`src/security.mjs`)** — adding an import to a test file was a CRITICAL block.
|
|
27
|
+
- [x] **A Test Renamed Out Of Discovery Is Its Own Finding (`src/security.mjs`)** — `TEST_DEREGISTERED`, waivable like every other kind.
|
|
28
|
+
- [x] **The Scaffold Passes A Repository's Own YAML Linter (`src/config.mjs`)** — 40 eslint errors on files `init` had just written.
|
|
29
|
+
- [x] **A Loosened Run Says So (`bin/agentctl.mjs`)** — an override left no trace in the report at all.
|
|
30
|
+
- [x] **`task create -p` Skips The Questions (`src/wizard-task.mjs`)** — the advertised quickstart blocked forever in a real terminal.
|
|
31
|
+
|
|
23
32
|
### v0.70.0: Not Finished Is Not Passed
|
|
24
33
|
- [x] **An Unfinished Session Is Not COMPLETED (`src/engine.mjs`)** — terminal, blocked and timed-out are three verdicts, not one.
|
|
25
34
|
- [x] **The Retry Carries The Failure (`src/session-ops.mjs`)** — it was reading four fields the API does not return.
|
package/bin/agentctl.mjs
CHANGED
|
@@ -425,6 +425,31 @@ async function main() {
|
|
|
425
425
|
} else {
|
|
426
426
|
console.log(`\n🛡️ agentctl Safety Gate Audit Results (Base: ${values.base}, Mode: ${selectedMode})`);
|
|
427
427
|
console.log(`-----------------------------------------------------`);
|
|
428
|
+
// A loosened run must not be able to pass for a strict one.
|
|
429
|
+
//
|
|
430
|
+
// Every escape hatch here turns a check off, and the report said so
|
|
431
|
+
// nowhere: `agentctl check --allow-test-change expectation` over a
|
|
432
|
+
// rewritten expectation printed a phase list byte-identical to a clean,
|
|
433
|
+
// unloosened APPROVED. Anyone reading the output afterwards — a
|
|
434
|
+
// reviewer, a CI log, the author a week later — could not tell that a
|
|
435
|
+
// check had been waived, which makes the waiver invisible exactly where
|
|
436
|
+
// it matters most. Printed before the phases, and on approval as well
|
|
437
|
+
// as rejection, because an approval is the case where it is load-bearing.
|
|
438
|
+
const overrides = [];
|
|
439
|
+
if (values["allow-protected"]) overrides.push("--allow-protected (protected paths permitted)");
|
|
440
|
+
if (values["allow-unreadable-tests"]) overrides.push("--allow-unreadable-tests (unreadable dialect permitted)");
|
|
441
|
+
if (values["allow-test-modifications"]) overrides.push("--allow-test-modifications (every tamper check waived)");
|
|
442
|
+
if (values["allow-test-change"]) {
|
|
443
|
+
const kinds = [].concat(values["allow-test-change"]).join(", ");
|
|
444
|
+
overrides.push(`--allow-test-change ${kinds} (tamper check waived: ${kinds})`);
|
|
445
|
+
}
|
|
446
|
+
if (config.verify?.required === false) overrides.push("verify.required: false (nothing is executed)");
|
|
447
|
+
if (config.verify?.tamperGuard === "warn") overrides.push('verify.tamperGuard: "warn" (unreadable dialects report only)');
|
|
448
|
+
if (overrides.length > 0) {
|
|
449
|
+
console.log(` ⚠️ OVERRIDES ACTIVE — this run is not a strict pass:`);
|
|
450
|
+
for (const o of overrides) console.log(` • ${o}`);
|
|
451
|
+
console.log(`-----------------------------------------------------`);
|
|
452
|
+
}
|
|
428
453
|
for (const p of res.phases) {
|
|
429
454
|
const status = p.ok ? "✅ PASS" : "❌ FAIL";
|
|
430
455
|
// Test tampering shares this phase, and its exit code, with the
|
|
@@ -472,13 +497,32 @@ async function main() {
|
|
|
472
497
|
// arrive under the same phase and the same code.
|
|
473
498
|
const secretsPhase = res.phases.find((p) => p.phase === "secrets" && !p.ok);
|
|
474
499
|
const findingTypes = new Set((secretsPhase?.findings || []).map((f) => f.type));
|
|
500
|
+
// `TEST_DIALECT_UNREADABLE` belongs in this set and was missing from
|
|
501
|
+
// it, so the one finding that is *most* obviously not a secret fell
|
|
502
|
+
// through to the secret hint: the report printed
|
|
503
|
+
// "Phase [SECRETS (test integrity — no secret found)] : ❌ FAIL" and
|
|
504
|
+
// then told the operator a high-entropy credential had leaked and to
|
|
505
|
+
// rotate their keys. v0.67.0 fixed the phase label and left the hint;
|
|
506
|
+
// v0.69.0 recorded the hint as already correct without measuring it.
|
|
507
|
+
// Both statements were about the same screen, and they contradicted
|
|
508
|
+
// each other on it.
|
|
475
509
|
const INTEGRITY_TYPES = new Set([
|
|
476
510
|
"TEST_TAMPERING_DETECTED",
|
|
511
|
+
"TEST_DIALECT_UNREADABLE",
|
|
477
512
|
"EDGE_RUNTIME_VIOLATION",
|
|
478
513
|
"CROSS_PACKAGE_BOUNDARY_VIOLATION",
|
|
479
514
|
]);
|
|
480
515
|
const onlyIntegrityFindings =
|
|
481
516
|
findingTypes.size > 0 && [...findingTypes].every((t) => INTEGRITY_TYPES.has(t));
|
|
517
|
+
// An unreadable dialect is not a weakened assertion either, and the
|
|
518
|
+
// integrity hint tells the operator to restore an assertion that no
|
|
519
|
+
// finding names. Split by what is actually present — and derived from
|
|
520
|
+
// `onlyIntegrityFindings` rather than tested independently, so a run
|
|
521
|
+
// carrying both an unreadable file and a real tampering finding gets
|
|
522
|
+
// the tampering advice instead of falling past both into the secret
|
|
523
|
+
// branch, which is how this hole was shaped the first time.
|
|
524
|
+
const onlyUnreadableDialect =
|
|
525
|
+
onlyIntegrityFindings && [...findingTypes].every((t) => t === "TEST_DIALECT_UNREADABLE");
|
|
482
526
|
|
|
483
527
|
// Exit 3 is also what a strictTestLock tamper verdict returns, so the
|
|
484
528
|
// code alone cannot pick the hint — a scope remediation for a rewritten
|
|
@@ -504,6 +548,16 @@ async function main() {
|
|
|
504
548
|
if (firstError) console.log(` • Last attempt error: ${firstError}`);
|
|
505
549
|
console.log(` • Review error fingerprints via: agentctl doctor\n`);
|
|
506
550
|
}
|
|
551
|
+
} else if (onlyUnreadableDialect) {
|
|
552
|
+
console.log(`💡 Remediation Hint (Exit ${res.code} Unreadable Test Dialect — no secret was found):`);
|
|
553
|
+
console.log(` • Nothing leaked. The guard could not recognise the assertions in a changed test`);
|
|
554
|
+
console.log(` file, so it could not check that file for tampering — and a check that examined`);
|
|
555
|
+
console.log(` nothing does not get to say APPROVED.`);
|
|
556
|
+
console.log(` • Best: report the dialect so it gets covered. The guard knows Node, pytest, Go,`);
|
|
557
|
+
console.log(` Rust, JUnit, RSpec, PHPUnit, Minitest, XCTest, chai and node-tap.`);
|
|
558
|
+
console.log(` • Or say so on the record: verify.tamperGuard: "warn" in .agent/config.yml`);
|
|
559
|
+
console.log(` • Or allow this one run: agentctl check --allow-unreadable-tests`);
|
|
560
|
+
console.log(` • Nothing needs rotating: this exit code is shared with the secret scanner.\n`);
|
|
507
561
|
} else if (onlyIntegrityFindings) {
|
|
508
562
|
console.log(`💡 Remediation Hint (Exit ${res.code} Test Integrity Violation — no secret was found):`);
|
|
509
563
|
console.log(` • The diff weakens or removes verification rather than leaking a credential.`);
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "jules-orchestrator-kit",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.71.0",
|
|
4
4
|
"description": "Zero-dependency safety gatekeeper, test oracle generator, and multi-agent coordination protocol for autonomous coding agents — Google Jules, Claude Code, Codex and Gemini CLI.",
|
|
5
5
|
"repository": {
|
|
6
6
|
"type": "git",
|
|
@@ -42,10 +42,10 @@
|
|
|
42
42
|
* Exit codes: 0 = every guard reachable, 1 = a guard has gone silent.
|
|
43
43
|
*/
|
|
44
44
|
|
|
45
|
-
import { checkTestTampering, checkScope } from "../src/security.mjs";
|
|
45
|
+
import { checkTestTampering, checkScope, scanDiff } from "../src/security.mjs";
|
|
46
46
|
import { isTestPath } from "../src/test-paths.mjs";
|
|
47
47
|
import { normalizeScope } from "../src/config.mjs";
|
|
48
|
-
import { parseCollectedTests } from "../src/ops/test-collection.mjs";
|
|
48
|
+
import { parseCollectedTests, checkCollectionFloor } from "../src/ops/test-collection.mjs";
|
|
49
49
|
import {
|
|
50
50
|
TEST_PATH_CASES,
|
|
51
51
|
TAMPER_CANARIES,
|
|
@@ -53,6 +53,9 @@ import {
|
|
|
53
53
|
PREDICATE_MUTANTS,
|
|
54
54
|
EMPTY_RUN_CANARIES,
|
|
55
55
|
COUNTED_RUN_CANARIES,
|
|
56
|
+
SILENT_RUN_CANARIES,
|
|
57
|
+
SILENT_STATIC_GATES,
|
|
58
|
+
DEREGISTRATION_CANARIES,
|
|
56
59
|
SCOPE_CANARIES,
|
|
57
60
|
INNOCENT_EDITS,
|
|
58
61
|
UNREADABLE_DIALECTS,
|
|
@@ -154,13 +157,53 @@ const canaryResults = new Map();
|
|
|
154
157
|
add("canaries: assertion rules parsed an assertion", noAssertions.length === 0, noAssertions.length ? noAssertions.join(", ") : "assertionsSeen > 0 on every assertion finding");
|
|
155
158
|
}
|
|
156
159
|
|
|
160
|
+
{
|
|
161
|
+
const missed = DEREGISTRATION_CANARIES.filter(
|
|
162
|
+
(c) => !(checkTestTampering(canaryDiff(c)).violations || []).some((v) => v.type === "TEST_DEREGISTERED")
|
|
163
|
+
);
|
|
164
|
+
add(
|
|
165
|
+
"canaries: a test renamed out of discovery is caught",
|
|
166
|
+
missed.length === 0,
|
|
167
|
+
missed.length ? missed.map((c) => `${c.id} (${c.why})`).join("; ") : `${DEREGISTRATION_CANARIES.length} de-registrations reported`
|
|
168
|
+
);
|
|
169
|
+
}
|
|
170
|
+
|
|
171
|
+
{
|
|
172
|
+
// Zero bytes on both streams is the one absence that decides. Measured
|
|
173
|
+
// through the floor rather than the parser: the parser answers "no count
|
|
174
|
+
// stated", and it is the floor that has to refuse to certify against it.
|
|
175
|
+
const passed = SILENT_RUN_CANARIES.filter(
|
|
176
|
+
(c) => checkCollectionFloor({ ok: true, stdout: c.stdout, stderr: c.stderr, command: c.command }).ok !== false
|
|
177
|
+
);
|
|
178
|
+
add(
|
|
179
|
+
"policy: a command that printed nothing verified nothing",
|
|
180
|
+
passed.length === 0,
|
|
181
|
+
passed.length ? passed.map((c) => c.id).join(", ") : `${SILENT_RUN_CANARIES.length} silent runs refused`
|
|
182
|
+
);
|
|
183
|
+
|
|
184
|
+
// And the counterweight: a checker that succeeds in silence is doing its job.
|
|
185
|
+
const redded = SILENT_STATIC_GATES.filter(
|
|
186
|
+
(c) => checkCollectionFloor({ ok: true, stdout: "", stderr: "", command: c.command }).ok === false
|
|
187
|
+
);
|
|
188
|
+
add(
|
|
189
|
+
"policy: a silent static gate is not a silent suite",
|
|
190
|
+
redded.length === 0,
|
|
191
|
+
redded.length ? `${redded.map((c) => c.id).join(", ")} hard-redded` : `${SILENT_STATIC_GATES.length} static gates still pass`
|
|
192
|
+
);
|
|
193
|
+
}
|
|
194
|
+
|
|
157
195
|
// --- 3. The opposite failure: flagging what is innocent ---------------------
|
|
158
196
|
{
|
|
159
197
|
const noisy = [];
|
|
160
198
|
for (const e of INNOCENT_EDITS) {
|
|
161
199
|
const res = checkTestTampering(canaryDiff(e));
|
|
162
200
|
const types = (res.violations || []).map((v) => v.type);
|
|
163
|
-
if (types.length > 0) noisy.push(`${e.id}
|
|
201
|
+
if (types.length > 0) noisy.push(`${e.id} \u2192 ${JSON.stringify(types)} (${e.why})`);
|
|
202
|
+
// The verdict, not just the violation list. `UNREADABLE` is a status, so
|
|
203
|
+
// this check stayed green for two releases while the gate rejected every
|
|
204
|
+
// edit below on a CRITICAL finding.
|
|
205
|
+
const scan = scanDiff(canaryDiff(e));
|
|
206
|
+
if (!scan.ok) noisy.push(`${e.id} \u2192 gate rejected: ${JSON.stringify((scan.findings || []).map((f) => f.type))}`);
|
|
164
207
|
}
|
|
165
208
|
add(
|
|
166
209
|
"innocent edits stay silent",
|
package/src/config.mjs
CHANGED
|
@@ -205,7 +205,13 @@ const BLOCKED_KEYS = new Set(["__proto__", "constructor", "prototype"]);
|
|
|
205
205
|
|
|
206
206
|
function coerce(val) {
|
|
207
207
|
if (!val) return "";
|
|
208
|
-
if (
|
|
208
|
+
if (val.startsWith("'") && val.endsWith("'") && val.length >= 2) {
|
|
209
|
+
// `''` is how a single-quoted YAML scalar escapes a quote, and it is what
|
|
210
|
+
// `yamlScalar` emits. Reading the quotes off without undoing the escape
|
|
211
|
+
// meant the emitter and the parser disagreed about the same file.
|
|
212
|
+
return val.slice(1, -1).replace(/''/g, "'");
|
|
213
|
+
}
|
|
214
|
+
if (val.startsWith('"') && val.endsWith('"') && val.length >= 2) {
|
|
209
215
|
return val.slice(1, -1);
|
|
210
216
|
}
|
|
211
217
|
if (val === "true") return true;
|
|
@@ -315,6 +321,86 @@ export function normalizeScope(parsed = {}) {
|
|
|
315
321
|
};
|
|
316
322
|
}
|
|
317
323
|
|
|
324
|
+
/**
|
|
325
|
+
* A YAML scalar quoted only as much as YAML requires.
|
|
326
|
+
*
|
|
327
|
+
* The scaffold wrapped every generated string in double quotes, which is valid
|
|
328
|
+
* YAML and idiomatic nowhere. On a repository that lints its own YAML that is a
|
|
329
|
+
* hard red on the kit's own files: measured on `unjs/unimport`, following the
|
|
330
|
+
* README verbatim produced 40 eslint errors — `yaml/quotes` and
|
|
331
|
+
* `yaml/plain-scalar` — across `.agent/config.yml` and `.agent/jules.yml`, so a
|
|
332
|
+
* newcomer's very first `agentctl check` failed on output `agentctl init` had
|
|
333
|
+
* just written, in a repository that was green a minute earlier. Teaching a new
|
|
334
|
+
* user that the gate is broken is the most expensive thing this tool can do.
|
|
335
|
+
*
|
|
336
|
+
* Nothing here is specific to eslint or to that repository. Emitting a plain
|
|
337
|
+
* scalar where YAML permits one, and single quotes where it does not, is simply
|
|
338
|
+
* how YAML is written; the linters that check it are agreeing with the spec.
|
|
339
|
+
*
|
|
340
|
+
* The quoting test is deliberately conservative — anything that could change
|
|
341
|
+
* meaning unquoted gets quotes. A glob beginning with a star stays quoted,
|
|
342
|
+
* because a leading `*` is an alias reference; `pnpm -r test` does not, because
|
|
343
|
+
* there is nothing in it to misread.
|
|
344
|
+
*/
|
|
345
|
+
const YAML_NEEDS_QUOTES =
|
|
346
|
+
/^$|^[-?:,[\]{}#&*!|>'"%@`]|^\s|\s$|:\s|\s#|:$|[\n\r\t]|^(?:true|false|yes|no|on|off|null|~)$/i;
|
|
347
|
+
|
|
348
|
+
export function yamlScalar(value) {
|
|
349
|
+
const s = value === null || value === undefined ? "" : String(value);
|
|
350
|
+
if (!YAML_NEEDS_QUOTES.test(s) && !/^[+-]?(?:\d[\d_]*)(?:\.[\d_]*)?(?:[eE][+-]?\d+)?$/.test(s)) return s;
|
|
351
|
+
return `'${s.replace(/'/g, "''")}'`;
|
|
352
|
+
}
|
|
353
|
+
|
|
354
|
+
/**
|
|
355
|
+
* Can a quoted scalar begin at this offset?
|
|
356
|
+
*
|
|
357
|
+
* A quote character is only an opening quote where a value may start — after
|
|
358
|
+
* `key:`, after a `-` list marker, at the head of a flow collection, or at the
|
|
359
|
+
* start of the line. Anywhere else it is an ordinary apostrophe, and treating
|
|
360
|
+
* `test: it's fine` as an unterminated string would swallow the rest of it.
|
|
361
|
+
*/
|
|
362
|
+
function opensQuotedScalar(line, i) {
|
|
363
|
+
return /(?:^\s*|:\s+|-\s+|[[{,]\s*)$/.test(line.slice(0, i));
|
|
364
|
+
}
|
|
365
|
+
|
|
366
|
+
/**
|
|
367
|
+
* Strip a trailing `#` comment without cutting into a quoted scalar.
|
|
368
|
+
*
|
|
369
|
+
* `rawLine.indexOf("#")` stood here, which is not the YAML rule twice over: a
|
|
370
|
+
* `#` opens a comment only when it begins the line or follows whitespace, and
|
|
371
|
+
* never inside quotes. So any configured value containing a hash was silently
|
|
372
|
+
* truncated on the way in — `test: 'pytest -k "not #slow"'` became
|
|
373
|
+
* `test: 'pytest -k "not`, and the gate then ran a command the user never
|
|
374
|
+
* wrote. Silent truncation of a verification command is the worst shape this
|
|
375
|
+
* file can produce, because every later phase reports on whatever did run.
|
|
376
|
+
*
|
|
377
|
+
* It matters more now that `yamlScalar` exists: an emitter and a parser that
|
|
378
|
+
* disagree about the same file is this project's recurring defect with both
|
|
379
|
+
* halves in one module.
|
|
380
|
+
*/
|
|
381
|
+
function stripYamlComment(rawLine) {
|
|
382
|
+
let quote = null;
|
|
383
|
+
for (let i = 0; i < rawLine.length; i++) {
|
|
384
|
+
const ch = rawLine[i];
|
|
385
|
+
if (quote) {
|
|
386
|
+
if (quote === '"' && ch === "\\") {
|
|
387
|
+
i++;
|
|
388
|
+
} else if (ch === quote) {
|
|
389
|
+
// `''` is an escaped quote inside a single-quoted scalar, not the end.
|
|
390
|
+
if (quote === "'" && rawLine[i + 1] === "'") i++;
|
|
391
|
+
else quote = null;
|
|
392
|
+
}
|
|
393
|
+
continue;
|
|
394
|
+
}
|
|
395
|
+
if ((ch === '"' || ch === "'") && opensQuotedScalar(rawLine, i)) {
|
|
396
|
+
quote = ch;
|
|
397
|
+
continue;
|
|
398
|
+
}
|
|
399
|
+
if (ch === "#" && (i === 0 || /\s/.test(rawLine[i - 1]))) return rawLine.slice(0, i);
|
|
400
|
+
}
|
|
401
|
+
return rawLine;
|
|
402
|
+
}
|
|
403
|
+
|
|
318
404
|
/**
|
|
319
405
|
* Indent-stack zero-dependency YAML parser with prototype pollution protection.
|
|
320
406
|
*/
|
|
@@ -331,8 +417,7 @@ export function parseYaml(src) {
|
|
|
331
417
|
const lines = src.split("\n");
|
|
332
418
|
|
|
333
419
|
for (let rawLine of lines) {
|
|
334
|
-
const
|
|
335
|
-
const line = commentIdx !== -1 ? rawLine.slice(0, commentIdx) : rawLine;
|
|
420
|
+
const line = stripYamlComment(rawLine);
|
|
336
421
|
const trimmed = line.trim();
|
|
337
422
|
if (!trimmed) continue;
|
|
338
423
|
|
package/src/engine.mjs
CHANGED
|
@@ -459,7 +459,28 @@ export async function gate(opts = {}) {
|
|
|
459
459
|
let assertMetrics = {};
|
|
460
460
|
|
|
461
461
|
if (isAssertion) {
|
|
462
|
-
|
|
462
|
+
// The loosening flags have to reach here too.
|
|
463
|
+
//
|
|
464
|
+
// `scanDiff` above is given `allowTestChanges` and honours it, while
|
|
465
|
+
// the `assert:test-integrity` stage received only its own stage object
|
|
466
|
+
// and re-ran the same guard with none of them. So an override was
|
|
467
|
+
// accepted by one phase and ignored by the next: `--allow-test-change
|
|
468
|
+
// deregistration` turned the SECRETS phase green and then failed the
|
|
469
|
+
// run at the anti-tamper stage, with a flag hint the operator had
|
|
470
|
+
// already followed. One rule, two places, and the second kept the old
|
|
471
|
+
// answer — for the eighth time in this project's history, which is why
|
|
472
|
+
// the regression test asserts the verdict a caller receives rather
|
|
473
|
+
// than the behaviour of either site.
|
|
474
|
+
const assertRes = runAssertion(
|
|
475
|
+
{
|
|
476
|
+
...stage,
|
|
477
|
+
allowTestModifications: opts.allowTestModifications === true,
|
|
478
|
+
allowTestChanges: opts.allowTestChanges,
|
|
479
|
+
tamperGuard: trustedVerify.tamperGuard,
|
|
480
|
+
allowUnreadableTests: opts.allowUnreadableTests === true,
|
|
481
|
+
},
|
|
482
|
+
root
|
|
483
|
+
);
|
|
463
484
|
durationMs = assertRes.metrics?.durationMs ?? (Date.now() - startTime);
|
|
464
485
|
stdoutRedacted = assertRes.stdout || "";
|
|
465
486
|
stderrRedacted = assertRes.stderr || "";
|
package/src/guard-policy.mjs
CHANGED
|
@@ -454,6 +454,57 @@ export const SCOPE_CANARIES = [
|
|
|
454
454
|
* from never showed it.
|
|
455
455
|
*/
|
|
456
456
|
export const INNOCENT_EDITS = [
|
|
457
|
+
// Everything below this comment was a CRITICAL rejection until v0.71.0.
|
|
458
|
+
//
|
|
459
|
+
// Not through any rule that examined them: `UNREADABLE` was reached whenever
|
|
460
|
+
// a test file had changed lines and none of them parsed as an assertion, so
|
|
461
|
+
// the verdict was the same for a repository speaking an unsupported dialect
|
|
462
|
+
// and for one where the edit simply was not an assertion. Adding an import
|
|
463
|
+
// was a CRITICAL block. The finding carried no file, no line and no sample,
|
|
464
|
+
// because there was nothing to name — and it advised a pytest repository
|
|
465
|
+
// that its assertion library might be unsupported, from a list naming pytest.
|
|
466
|
+
//
|
|
467
|
+
// These are ordinary work. Every one of them has to be silent.
|
|
468
|
+
{
|
|
469
|
+
id: "add-import/pytest",
|
|
470
|
+
file: "testing/test_iniconfig.py",
|
|
471
|
+
context: "# fixtures",
|
|
472
|
+
removed: [],
|
|
473
|
+
added: ["import os"],
|
|
474
|
+
why: "adding an import to a test file is not an assertion and not tampering",
|
|
475
|
+
},
|
|
476
|
+
{
|
|
477
|
+
id: "rename-test-multiline/pytest",
|
|
478
|
+
file: "testing/test_iniconfig.py",
|
|
479
|
+
context: "# parsing",
|
|
480
|
+
removed: ["def test_parse_strips_inline_comments() -> None:"],
|
|
481
|
+
added: ["def test_parse_strips_inline_comments_from_continuations() -> None:"],
|
|
482
|
+
why: "still collected by pytest, so nothing left the run",
|
|
483
|
+
},
|
|
484
|
+
{
|
|
485
|
+
id: "docstring/pytest",
|
|
486
|
+
file: "testing/test_iniconfig.py",
|
|
487
|
+
context: "# parsing",
|
|
488
|
+
removed: [' """Check comments."""'],
|
|
489
|
+
added: [' """Check that inline comments are stripped."""'],
|
|
490
|
+
why: "prose about the test, not a value it asserts",
|
|
491
|
+
},
|
|
492
|
+
{
|
|
493
|
+
id: "type-annotation/pytest",
|
|
494
|
+
file: "testing/test_iniconfig.py",
|
|
495
|
+
context: "# fixtures",
|
|
496
|
+
removed: ["def make_config(data):"],
|
|
497
|
+
added: ["def make_config(data: str) -> IniConfig:"],
|
|
498
|
+
why: "a helper's signature, typed; no claim changed",
|
|
499
|
+
},
|
|
500
|
+
{
|
|
501
|
+
id: "add-fixture/go",
|
|
502
|
+
file: "calc_test.go",
|
|
503
|
+
context: "// helpers",
|
|
504
|
+
removed: [],
|
|
505
|
+
added: ["var cases = []int{1, 2, 3}"],
|
|
506
|
+
why: "test data added, nothing asserted or unasserted",
|
|
507
|
+
},
|
|
457
508
|
// Renaming a test is one of the most ordinary edits there is. On the
|
|
458
509
|
// one-line form the name blanked to the same shape as its replacement, the
|
|
459
510
|
// two paired, and the pair was reported as a rewritten expectation.
|
|
@@ -598,6 +649,129 @@ export const INNOCENT_EDITS = [
|
|
|
598
649
|
* will always end somewhere; what must never happen again is that the edge
|
|
599
650
|
* is silent.
|
|
600
651
|
*/
|
|
652
|
+
/**
|
|
653
|
+
* Renames that delete a test from the run without deleting a line of it.
|
|
654
|
+
*
|
|
655
|
+
* pytest collects `test_*`, Go collects `Test*`. For those runners the name is
|
|
656
|
+
* the registration, so `def test_totals` → `def totals` removes the test as
|
|
657
|
+
* completely as deleting the file, and every count in the tamper guard stays
|
|
658
|
+
* level: nothing removed, nothing weakened, nothing rewritten.
|
|
659
|
+
*
|
|
660
|
+
* These were caught only as a side effect of the blanket that rejected every
|
|
661
|
+
* unrecognised edit to a test file — which also rejected adding an import, and
|
|
662
|
+
* whose printed remedy switched the real checks off. They are their own
|
|
663
|
+
* finding now, so narrowing that blanket costs nothing.
|
|
664
|
+
*/
|
|
665
|
+
export const DEREGISTRATION_CANARIES = [
|
|
666
|
+
{
|
|
667
|
+
id: "pytest/underscore",
|
|
668
|
+
file: "tests/test_totals.py",
|
|
669
|
+
context: "# totals",
|
|
670
|
+
removed: ["def test_totals_round_to_cents():"],
|
|
671
|
+
added: ["def totals_round_to_cents():"],
|
|
672
|
+
why: "pytest collects by the test_ prefix, so this test no longer runs",
|
|
673
|
+
},
|
|
674
|
+
{
|
|
675
|
+
id: "pytest/camel",
|
|
676
|
+
file: "tests/test_totals.py",
|
|
677
|
+
context: "# totals",
|
|
678
|
+
removed: ["def testTotals():"],
|
|
679
|
+
added: ["def Totals():"],
|
|
680
|
+
why: "the prefix is what registers it, in either spelling",
|
|
681
|
+
},
|
|
682
|
+
{
|
|
683
|
+
id: "go/test",
|
|
684
|
+
file: "calc_test.go",
|
|
685
|
+
context: "// totals",
|
|
686
|
+
removed: ["func TestTotals(t *testing.T) {"],
|
|
687
|
+
added: ["func Totals(t *testing.T) {"],
|
|
688
|
+
why: "go test collects by the Test prefix",
|
|
689
|
+
},
|
|
690
|
+
{
|
|
691
|
+
id: "go/benchmark",
|
|
692
|
+
file: "calc_test.go",
|
|
693
|
+
context: "// totals",
|
|
694
|
+
removed: ["func BenchmarkTotals(b *testing.B) {"],
|
|
695
|
+
added: ["func Totals(b *testing.B) {"],
|
|
696
|
+
why: "the same rule for the other collected prefixes",
|
|
697
|
+
},
|
|
698
|
+
];
|
|
699
|
+
|
|
700
|
+
/**
|
|
701
|
+
* A verification command that exits 0 having written nothing at all.
|
|
702
|
+
*
|
|
703
|
+
* The one absence that is evidence. The floor is otherwise deliberately
|
|
704
|
+
* one-sided — an unrecognised runner states no count and passes, because
|
|
705
|
+
* hard-redding every runner not on the list would be worse than the hole it
|
|
706
|
+
* closes — but an unrecognised runner still prints something. Zero bytes on
|
|
707
|
+
* both streams is a command that ran nothing, and `pnpm -r test` on a
|
|
708
|
+
* workspace whose packages declare no test script is exactly that: it was
|
|
709
|
+
* indistinguishable from a full suite by every signal the gate had.
|
|
710
|
+
*/
|
|
711
|
+
export const SILENT_RUN_CANARIES = [
|
|
712
|
+
{ id: "pnpm -r test", command: "pnpm -r test", stdout: "", stderr: "" },
|
|
713
|
+
{ id: "npm test", command: "npm test", stdout: "", stderr: "" },
|
|
714
|
+
{ id: "yarn workspaces test", command: "yarn workspaces foreach run test", stdout: "", stderr: "" },
|
|
715
|
+
{ id: "pytest", command: "python3 -m pytest", stdout: "", stderr: "" },
|
|
716
|
+
{ id: "go test", command: "go test ./...", stdout: "", stderr: "" },
|
|
717
|
+
{ id: "whitespace is not output", command: "npm test", stdout: " \n", stderr: "\n" },
|
|
718
|
+
];
|
|
719
|
+
|
|
720
|
+
/**
|
|
721
|
+
* Honest static gates that print nothing, and must keep passing.
|
|
722
|
+
*
|
|
723
|
+
* The counterweight, and the reason the rule above reads the command as well
|
|
724
|
+
* as the output. `tsc --noEmit`, `node --check`, `go vet` and `compileall`
|
|
725
|
+
* all exit 0 in silence when they succeed — that is what success looks like
|
|
726
|
+
* for a checker — and two of them are commands this kit writes itself for a
|
|
727
|
+
* repository that has no suite yet. A rule keyed on silence alone hard-redded
|
|
728
|
+
* every one of them, which is exactly the first-run rejection of correct code
|
|
729
|
+
* that the collection floor is otherwise so careful to avoid.
|
|
730
|
+
*/
|
|
731
|
+
export const SILENT_STATIC_GATES = [
|
|
732
|
+
{ id: "tsc", command: "tsc --noEmit" },
|
|
733
|
+
{ id: "node --check", command: "node --check index.js" },
|
|
734
|
+
{ id: "go vet", command: "go vet ./..." },
|
|
735
|
+
{ id: "compileall", command: "python3 -m compileall -q ." },
|
|
736
|
+
{ id: "generated parse gate", command: "node --check src/index.mjs" },
|
|
737
|
+
];
|
|
738
|
+
|
|
739
|
+
/**
|
|
740
|
+
* Values that must survive a trip through the emitter and back.
|
|
741
|
+
*
|
|
742
|
+
* `yamlScalar` writes the manifests and `parseYaml` reads them, and a pair
|
|
743
|
+
* like that disagreeing about one character is this project's recurring
|
|
744
|
+
* defect with both halves in a single module. The parser opened a comment at
|
|
745
|
+
* the first `#` on a line, quoted or not, so `verify.test` containing a hash
|
|
746
|
+
* was silently truncated on the way in and the gate ran a command the user
|
|
747
|
+
* never wrote — reporting on it as if it were theirs.
|
|
748
|
+
*
|
|
749
|
+
* The corpus is the hard cases on purpose: hashes, colons, leading stars,
|
|
750
|
+
* apostrophes, the empty string, and the words YAML would otherwise read as
|
|
751
|
+
* booleans.
|
|
752
|
+
*/
|
|
753
|
+
export const YAML_ROUNDTRIP_CASES = [
|
|
754
|
+
"pnpm -r test",
|
|
755
|
+
"PYTHONPATH=src python3 -m pytest",
|
|
756
|
+
'pytest -k "not #slow"',
|
|
757
|
+
"has #hash",
|
|
758
|
+
"# leading hash",
|
|
759
|
+
"a: b",
|
|
760
|
+
"**/*.pem",
|
|
761
|
+
"**/.env.*",
|
|
762
|
+
".github/**",
|
|
763
|
+
"",
|
|
764
|
+
"true",
|
|
765
|
+
"yes",
|
|
766
|
+
"1.5",
|
|
767
|
+
"it's fine",
|
|
768
|
+
"it's #1",
|
|
769
|
+
"O'Brien",
|
|
770
|
+
"agent/",
|
|
771
|
+
" leading space",
|
|
772
|
+
"trailing space ",
|
|
773
|
+
];
|
|
774
|
+
|
|
601
775
|
export const UNREADABLE_DIALECTS = [
|
|
602
776
|
{
|
|
603
777
|
id: "hspec",
|
|
@@ -111,6 +111,56 @@ const GO_RAN_SOMETHING = /^(?:(?:ok|FAIL)\s+(?!\d+\s)\S+|---\s+(?:PASS|FAIL|SKIP
|
|
|
111
111
|
* `count` is null when no recognised runner stated one — which is not a
|
|
112
112
|
* finding, only an absence of evidence.
|
|
113
113
|
*/
|
|
114
|
+
/**
|
|
115
|
+
* Did the command write nothing at all, on either stream?
|
|
116
|
+
*
|
|
117
|
+
* This is the one absence that is evidence rather than the lack of it. The
|
|
118
|
+
* one-sided floor exists because an unrecognised runner states no count, and
|
|
119
|
+
* hard-redding every runner not on the list would be worse than the hole it
|
|
120
|
+
* closes. But an unrecognised runner still *prints*: dots, a summary line,
|
|
121
|
+
* a package name, something. Zero bytes on both streams is not a dialect the
|
|
122
|
+
* list has yet to learn — it is a command that ran nothing.
|
|
123
|
+
*
|
|
124
|
+
* `pnpm -r test` on a workspace whose packages declare no test script is the
|
|
125
|
+
* shape that made this necessary: it exits 0, writes nothing anywhere, and
|
|
126
|
+
* was indistinguishable from a full suite by every signal the gate had.
|
|
127
|
+
*
|
|
128
|
+
* Resolved here and consumed in two places — the gate's floor and `init`'s
|
|
129
|
+
* oracle probe — because writing the rule once in each is how this project
|
|
130
|
+
* has repeatedly ended up with two answers to one question.
|
|
131
|
+
*/
|
|
132
|
+
export function producedNoOutput(stdout = "", stderr = "") {
|
|
133
|
+
return `${stdout || ""}${stderr || ""}`.trim() === "";
|
|
134
|
+
}
|
|
135
|
+
|
|
136
|
+
/**
|
|
137
|
+
* Does this command claim to run a test suite?
|
|
138
|
+
*
|
|
139
|
+
* Silence alone cannot carry the verdict, and the first version of this rule
|
|
140
|
+
* assumed it could. `node --check index.js`, `tsc --noEmit`, `go vet ./...`
|
|
141
|
+
* and `python3 -m compileall -q .` all exit 0 having printed nothing — and
|
|
142
|
+
* they are honest static gates, two of which this kit writes itself for
|
|
143
|
+
* repositories that have no suite yet. Failing on silence alone hard-redded
|
|
144
|
+
* every one of them, which is the same first-run rejection of correct code
|
|
145
|
+
* that the whole collection floor is careful to avoid.
|
|
146
|
+
*
|
|
147
|
+
* Nothing in the *output* separates `pnpm -r test` from `tsc --noEmit`; both
|
|
148
|
+
* are empty. The difference is in what the command says it is. So this reads
|
|
149
|
+
* the command, the same way `isPlaceholderTestScript` does: a command that is
|
|
150
|
+
* recognisably a suite invocation and printed nothing ran no suite, while a
|
|
151
|
+
* static checker that printed nothing did exactly what it promised.
|
|
152
|
+
*
|
|
153
|
+
* One-sided in the safe direction, like everything else here. An unrecognised
|
|
154
|
+
* command is not treated as a suite, so an unusual runner keeps its advisory
|
|
155
|
+
* pass rather than becoming a hard red.
|
|
156
|
+
*/
|
|
157
|
+
const TEST_SUITE_COMMAND =
|
|
158
|
+
/(?:^|\s|\/)(?:pytest|jest|vitest|mocha|ava|karma|jasmine|nyc|c8|tap|tape|rspec|minitest|phpunit|behave|nose2?|ginkgo|gotestsum|nextest)\b|\b(?:go|cargo|swift|dart|flutter|deno|bun|dotnet|mix|lein|sbt|gradlew?|mvn)\s+test\b|\bunittest\b|-m\s+(?:pytest|unittest)\b|\bnode\s+--test\b|(?:^|&&|;|\|)\s*(?:npm|pnpm|yarn|bun|npx)\b[^&;|]*?\btest\b/;
|
|
159
|
+
|
|
160
|
+
export function looksLikeTestSuiteCommand(cmd) {
|
|
161
|
+
return typeof cmd === "string" && TEST_SUITE_COMMAND.test(cmd);
|
|
162
|
+
}
|
|
163
|
+
|
|
114
164
|
export function parseCollectedTests(stdout = "", stderr = "") {
|
|
115
165
|
const text = `${stdout || ""}\n${stderr || ""}`;
|
|
116
166
|
if (!text.trim()) return { count: null, runner: null };
|
|
@@ -168,6 +218,30 @@ export function checkCollectionFloor(testResult, opts = {}) {
|
|
|
168
218
|
return { ok: true, count: null, runner: null, reason: null };
|
|
169
219
|
}
|
|
170
220
|
|
|
221
|
+
// A command that says it runs a suite, and printed nothing, ran no suite.
|
|
222
|
+
//
|
|
223
|
+
// Both halves are required. Silence alone would hard-red `tsc --noEmit` and
|
|
224
|
+
// `python3 -m compileall`, which are honest static gates this kit generates
|
|
225
|
+
// itself; the command shape alone would say nothing, because a real suite
|
|
226
|
+
// prints. Together they are decidable, and they are exactly `pnpm -r test`
|
|
227
|
+
// on a workspace whose packages declare no test script.
|
|
228
|
+
if (looksLikeTestSuiteCommand(testResult.command) && producedNoOutput(testResult.stdout, testResult.stderr)) {
|
|
229
|
+
return {
|
|
230
|
+
ok: false,
|
|
231
|
+
count: 0,
|
|
232
|
+
runner: null,
|
|
233
|
+
silent: true,
|
|
234
|
+
reason:
|
|
235
|
+
`The verification command ${testResult.command ? `${JSON.stringify(testResult.command)} ` : ""}` +
|
|
236
|
+
`exited 0 and wrote nothing at all — no test names, no summary, no count. ` +
|
|
237
|
+
`Every test runner prints something, so this command ran no suite, and approving this change ` +
|
|
238
|
+
`would certify nothing. A workspace command such as \`pnpm -r test\` does this when no package ` +
|
|
239
|
+
`declares a test script. Point verify.test at the suite that covers this repository ` +
|
|
240
|
+
`(often the root script rather than the recursive one), or — if this repository intentionally ` +
|
|
241
|
+
`uses only the scope and secret phases — set verify.required: false, which says so on the record.`,
|
|
242
|
+
};
|
|
243
|
+
}
|
|
244
|
+
|
|
171
245
|
const { count, runner } = parseCollectedTests(testResult.stdout, testResult.stderr);
|
|
172
246
|
|
|
173
247
|
// Deliberately one-sided: only a *stated* zero fails, because failing on
|
package/src/security.mjs
CHANGED
|
@@ -1875,6 +1875,51 @@ function splitTrailingMessage(clean) {
|
|
|
1875
1875
|
return { head: clean.slice(0, lastComma), msg: tail };
|
|
1876
1876
|
}
|
|
1877
1877
|
|
|
1878
|
+
/**
|
|
1879
|
+
* Test declarations that a runner finds by the *name* of the function.
|
|
1880
|
+
*
|
|
1881
|
+
* pytest collects `def test_*`, Go collects `func Test*`, and unittest and
|
|
1882
|
+
* Minitest collect `def test_*` off the case class. For those runners the
|
|
1883
|
+
* name is not prose — it is the registration. Renaming `test_totals` to
|
|
1884
|
+
* `totals` deletes the test from the run as completely as removing the file,
|
|
1885
|
+
* and the diff shows a rename.
|
|
1886
|
+
*
|
|
1887
|
+
* Only these name-driven runners are listed. `it("...")`, `#[test]` and
|
|
1888
|
+
* `@Test` register by call, attribute or annotation, so renaming what they
|
|
1889
|
+
* declare removes nothing, and the ordinary rename rules already cover them.
|
|
1890
|
+
*/
|
|
1891
|
+
const NAME_REGISTERED_DECLS = [
|
|
1892
|
+
{ lang: "python", re: /^\s*(?:async\s+)?def\s+([A-Za-z_]\w*)\s*\(/, discovered: /^test/i },
|
|
1893
|
+
{ lang: "go", re: /^\s*func\s+([A-Za-z_]\w*)\s*\(/, discovered: /^(?:Test|Benchmark|Fuzz|Example)/ },
|
|
1894
|
+
];
|
|
1895
|
+
|
|
1896
|
+
/**
|
|
1897
|
+
* The declared name on this line, and whether the runner would collect it.
|
|
1898
|
+
*
|
|
1899
|
+
* @returns {{ name: string, collected: boolean }|null}
|
|
1900
|
+
*/
|
|
1901
|
+
function declaredTestName(text) {
|
|
1902
|
+
for (const rule of NAME_REGISTERED_DECLS) {
|
|
1903
|
+
const m = rule.re.exec(text);
|
|
1904
|
+
if (m) return { name: m[1], collected: rule.discovered.test(m[1]) };
|
|
1905
|
+
}
|
|
1906
|
+
return null;
|
|
1907
|
+
}
|
|
1908
|
+
|
|
1909
|
+
/**
|
|
1910
|
+
* Is `after` the same declaration as `before` with its discovery prefix gone?
|
|
1911
|
+
*
|
|
1912
|
+
* Exact on the remainder, deliberately. `test_totals` → `totals` is a
|
|
1913
|
+
* de-registration; `test_totals` → `test_totals_rounded` is a rename and must
|
|
1914
|
+
* stay silent, which is the false red this check exists alongside rather than
|
|
1915
|
+
* instead of.
|
|
1916
|
+
*/
|
|
1917
|
+
function isDeregistration(before, after) {
|
|
1918
|
+
if (!before.collected || after.collected) return false;
|
|
1919
|
+
const stripped = before.name.replace(/^test[_-]?/i, "").replace(/^(?:Test|Benchmark|Fuzz|Example)/, "");
|
|
1920
|
+
return stripped.length > 0 && stripped === after.name;
|
|
1921
|
+
}
|
|
1922
|
+
|
|
1878
1923
|
// A test declaration whose first argument is the test's name. The name is
|
|
1879
1924
|
// prose about the test, not a value the test asserts — `test("adds", ...)`
|
|
1880
1925
|
// renamed to `test("adds positives", ...)` is the rename the diff says it is.
|
|
@@ -2211,6 +2256,7 @@ export const TAMPER_KINDS = new Map([
|
|
|
2211
2256
|
["ASSERTION_REMOVAL", "removal"],
|
|
2212
2257
|
["ASSERTION_WEAKENED", "weakening"],
|
|
2213
2258
|
["ASSERTION_EXPECTATION_CHANGED", "expectation"],
|
|
2259
|
+
["TEST_DEREGISTERED", "deregistration"],
|
|
2214
2260
|
]);
|
|
2215
2261
|
|
|
2216
2262
|
/** Every kind name, for CLI validation and help text. */
|
|
@@ -2406,7 +2452,7 @@ export function checkTestTampering(diffOrText = "", options = {}) {
|
|
|
2406
2452
|
}
|
|
2407
2453
|
|
|
2408
2454
|
if (!fileAssertions.has(currentFile)) {
|
|
2409
|
-
fileAssertions.set(currentFile, { removed: [], added: 0, addedTexts: [], removedSpecific: [], addedSpecific: 0, hunks: [], examined: 0, recognised: 0, unreadable: [] });
|
|
2455
|
+
fileAssertions.set(currentFile, { removed: [], added: 0, addedTexts: [], removedSpecific: [], addedSpecific: 0, hunks: [], examined: 0, recognised: 0, unreadable: [], declRemoved: [], declAdded: [] });
|
|
2410
2456
|
}
|
|
2411
2457
|
const fileStats = fileAssertions.get(currentFile);
|
|
2412
2458
|
if (pendingHunk) {
|
|
@@ -2419,6 +2465,10 @@ export function checkTestTampering(diffOrText = "", options = {}) {
|
|
|
2419
2465
|
const deletedText = line.slice(1);
|
|
2420
2466
|
if (hunk) hunk.lines.push({ kind: "-", text: deletedText, oldNo: currentOldLineNo, newNo: null });
|
|
2421
2467
|
countExamined(fileStats, deletedText);
|
|
2468
|
+
if (!isCommentLine(deletedText)) {
|
|
2469
|
+
const decl = declaredTestName(deletedText);
|
|
2470
|
+
if (decl) fileStats.declRemoved.push({ ...decl, line: currentOldLineNo, text: deletedText });
|
|
2471
|
+
}
|
|
2422
2472
|
if (!isCommentLine(deletedText) && ASSERTION_PATTERN.test(deletedText)) {
|
|
2423
2473
|
fileStats.removed.push({ line: currentOldLineNo, text: deletedText });
|
|
2424
2474
|
if (isSpecificAssertion(deletedText)) {
|
|
@@ -2430,6 +2480,10 @@ export function checkTestTampering(diffOrText = "", options = {}) {
|
|
|
2430
2480
|
const addedText = line.slice(1);
|
|
2431
2481
|
if (hunk) hunk.lines.push({ kind: "+", text: addedText, oldNo: null, newNo: currentNewLineNo });
|
|
2432
2482
|
countExamined(fileStats, addedText);
|
|
2483
|
+
if (!isCommentLine(addedText)) {
|
|
2484
|
+
const decl = declaredTestName(addedText);
|
|
2485
|
+
if (decl) fileStats.declAdded.push({ ...decl, line: currentNewLineNo, text: addedText });
|
|
2486
|
+
}
|
|
2433
2487
|
let isVacuous = false;
|
|
2434
2488
|
|
|
2435
2489
|
// Check skip injections
|
|
@@ -2563,6 +2617,40 @@ export function checkTestTampering(diffOrText = "", options = {}) {
|
|
|
2563
2617
|
}
|
|
2564
2618
|
}
|
|
2565
2619
|
|
|
2620
|
+
// A test renamed out of its runner's discovery convention.
|
|
2621
|
+
//
|
|
2622
|
+
// pytest collects `test_*` and nothing else, so `def test_totals` becoming
|
|
2623
|
+
// `def totals` deletes the test from every future run while leaving it in
|
|
2624
|
+
// the file, fully written, with all its assertions intact. Every count in
|
|
2625
|
+
// this guard stays level: nothing was removed, weakened or rewritten.
|
|
2626
|
+
//
|
|
2627
|
+
// Until now this was caught only by accident, as a side effect of the
|
|
2628
|
+
// blanket that blocked every unrecognised edit to a test file — which also
|
|
2629
|
+
// blocked adding an import, and whose printed remedy (`tamperGuard: "warn"`)
|
|
2630
|
+
// switched off the real checks along with the blanket. Narrowing that blanket
|
|
2631
|
+
// is what makes this its own finding, with its own name and its own remedy.
|
|
2632
|
+
for (const [file, stats] of fileAssertions.entries()) {
|
|
2633
|
+
const takenAdds = new Set();
|
|
2634
|
+
for (const before of stats.declRemoved || []) {
|
|
2635
|
+
if (!before.collected) continue;
|
|
2636
|
+
const idx = (stats.declAdded || []).findIndex((after, i) => !takenAdds.has(i) && isDeregistration(before, after));
|
|
2637
|
+
if (idx === -1) continue;
|
|
2638
|
+
takenAdds.add(idx);
|
|
2639
|
+
const after = stats.declAdded[idx];
|
|
2640
|
+
violations.push({
|
|
2641
|
+
file,
|
|
2642
|
+
line: after.line ?? before.line,
|
|
2643
|
+
type: "TEST_DEREGISTERED",
|
|
2644
|
+
reason:
|
|
2645
|
+
`Test Tamper Guard: ${JSON.stringify(before.name)} was renamed to ${JSON.stringify(after.name)} in ${file}` +
|
|
2646
|
+
`${after.line ? `:${after.line}` : ""}. The runner collects tests by name, so the test still exists in ` +
|
|
2647
|
+
`the file and no longer runs — the same effect as deleting it, with none of the signs. ` +
|
|
2648
|
+
`If the test is genuinely obsolete, delete it; if it is being turned into a helper, say so with ` +
|
|
2649
|
+
`--allow-test-change deregistration.`,
|
|
2650
|
+
});
|
|
2651
|
+
}
|
|
2652
|
+
}
|
|
2653
|
+
|
|
2566
2654
|
for (const [file, stats] of fileAssertions.entries()) {
|
|
2567
2655
|
// An expectation that was rewritten rather than removed.
|
|
2568
2656
|
//
|
|
@@ -2668,10 +2756,33 @@ export function checkTestTampering(diffOrText = "", options = {}) {
|
|
|
2668
2756
|
// assertion, is not the same as "checked and clean" — it is the state where
|
|
2669
2757
|
// this guard has nothing to say. Saying nothing and saying "approved" have
|
|
2670
2758
|
// to look different, which is the whole reason `status` exists.
|
|
2759
|
+
// `unreadable` is the evidence, and it is required.
|
|
2760
|
+
//
|
|
2761
|
+
// `|| examined > 0` used to stand here, and it threw away the distinction
|
|
2762
|
+
// this whole apparatus exists to draw. `ASSERTION_SHAPED` and `unreadable[]`
|
|
2763
|
+
// were built to separate "assertion-shaped lines were present and none of
|
|
2764
|
+
// them parsed" — a dialect the guard cannot read — from "there were no
|
|
2765
|
+
// assertions in these lines at all", which is most ordinary work on a test
|
|
2766
|
+
// file. That clause collapsed the two, so *any* changed substantive line in
|
|
2767
|
+
// a test file with no recognised assertion became a CRITICAL block:
|
|
2768
|
+
// measured on `pytest-dev/iniconfig`, renaming a test function did it, and
|
|
2769
|
+
// so did adding `import os`.
|
|
2770
|
+
//
|
|
2771
|
+
// The tell was in the finding itself: it carried `file: null`, `line: null`
|
|
2772
|
+
// and no sample, because `unreadable` was empty — the guard blocked while
|
|
2773
|
+
// holding no evidence of anything, and advised a pytest repository that its
|
|
2774
|
+
// assertion library might be unsupported, from a list that names pytest.
|
|
2775
|
+
//
|
|
2776
|
+
// Nothing is weakened by requiring the evidence. A removed or rewritten
|
|
2777
|
+
// assertion is a recognised assertion line, so it raises `assertionsSeen`
|
|
2778
|
+
// and goes to the ordinary removal and weakening checks; it never reached
|
|
2779
|
+
// this branch. What is lost is only the blanket, and a blanket that fires
|
|
2780
|
+
// on `import os` teaches its way around itself: the remedy it printed was
|
|
2781
|
+
// `tamperGuard: "warn"`, which switches the real guard off too.
|
|
2671
2782
|
const status =
|
|
2672
2783
|
reported.length > 0
|
|
2673
2784
|
? "FAIL"
|
|
2674
|
-
: assertionsSeen === 0 &&
|
|
2785
|
+
: assertionsSeen === 0 && unreadable.length > 0
|
|
2675
2786
|
? "UNREADABLE"
|
|
2676
2787
|
: examined > 0
|
|
2677
2788
|
? "PASS"
|
package/src/stack-detector.mjs
CHANGED
|
@@ -1,6 +1,7 @@
|
|
|
1
1
|
import { readFileSync, writeFileSync, existsSync, readdirSync, mkdirSync } from "node:fs";
|
|
2
2
|
import { join, relative } from "node:path";
|
|
3
3
|
import { whichBinary } from "./provider-readiness.mjs";
|
|
4
|
+
import { yamlScalar } from "./config.mjs";
|
|
4
5
|
|
|
5
6
|
/**
|
|
6
7
|
* The Python interpreter to invoke, by whatever name this machine has it under.
|
|
@@ -31,11 +32,48 @@ function pythonBin(env = process.env) {
|
|
|
31
32
|
* Falls back to the bare console script only when no interpreter can be found
|
|
32
33
|
* to host the module.
|
|
33
34
|
*/
|
|
34
|
-
export function pytestCmd(env = process.env) {
|
|
35
|
+
export function pytestCmd(env = process.env, root = null) {
|
|
36
|
+
const prefix = root && isSrcLayout(root) ? "PYTHONPATH=src " : "";
|
|
35
37
|
for (const name of ["python3", "python", "py"]) {
|
|
36
|
-
if (whichBinary(name, env)) return `${name} -m pytest`;
|
|
38
|
+
if (whichBinary(name, env)) return `${prefix}${name} -m pytest`;
|
|
39
|
+
}
|
|
40
|
+
return `${prefix}pytest`;
|
|
41
|
+
}
|
|
42
|
+
|
|
43
|
+
/**
|
|
44
|
+
* Does this repository keep its package under `src/` rather than at the root?
|
|
45
|
+
*
|
|
46
|
+
* The sibling of the bug above, and a worse one. `-m` puts the *working
|
|
47
|
+
* directory* on `sys.path` — which is the fix for a module at the root, and
|
|
48
|
+
* no help at all when the package lives in `src/`. There, `import iniconfig`
|
|
49
|
+
* finds nothing in the working directory and falls through to whatever is
|
|
50
|
+
* installed in site-packages. The suite then runs green against a *different
|
|
51
|
+
* copy of the library than the one in the diff*: measured on `pytest-dev/
|
|
52
|
+
* iniconfig`, `_parse.py` gutted to `return False`, 49 tests passed, and the
|
|
53
|
+
* gate returned APPROVED (Exit 0).
|
|
54
|
+
*
|
|
55
|
+
* That is this project's worst failure shape — a check that examined
|
|
56
|
+
* something other than the thing under review, reporting a pass — and no
|
|
57
|
+
* amount of counting collected tests can see it, because the tests really
|
|
58
|
+
* did run. Only the import path can.
|
|
59
|
+
*
|
|
60
|
+
* `PYTHONPATH=src` is the ordinary spelling and puts the working tree first
|
|
61
|
+
* whether or not the package is also installed. Leading assignments are
|
|
62
|
+
* peeled into the child's environment by `runCommand`, so this needs no shell.
|
|
63
|
+
*
|
|
64
|
+
* Keyed on `__init__.py` under `src/`, which is the marker of a Python
|
|
65
|
+
* package and not of a Rust, C or JavaScript `src/` directory.
|
|
66
|
+
*/
|
|
67
|
+
export function isSrcLayout(root) {
|
|
68
|
+
const src = join(root, "src");
|
|
69
|
+
if (!existsSync(src)) return false;
|
|
70
|
+
try {
|
|
71
|
+
return readdirSync(src, { withFileTypes: true }).some(
|
|
72
|
+
(e) => e.isDirectory() && existsSync(join(src, e.name, "__init__.py"))
|
|
73
|
+
);
|
|
74
|
+
} catch (_) {
|
|
75
|
+
return false;
|
|
37
76
|
}
|
|
38
|
-
return "pytest";
|
|
39
77
|
}
|
|
40
78
|
|
|
41
79
|
/**
|
|
@@ -234,7 +272,7 @@ export function oracleCandidates(root = process.cwd(), detected = "") {
|
|
|
234
272
|
} catch (_) {}
|
|
235
273
|
}
|
|
236
274
|
if (has("pytest.ini") || has("pyproject.toml") || has("setup.py") || has("tox.ini") || has("setup.cfg")) {
|
|
237
|
-
push(pytestCmd());
|
|
275
|
+
push(pytestCmd(process.env, root));
|
|
238
276
|
}
|
|
239
277
|
if (has("Cargo.toml")) push("cargo test");
|
|
240
278
|
if (has("go.mod")) push("go test ./...");
|
|
@@ -386,7 +424,7 @@ export function detectPolyglotStack(projectRoot = process.cwd()) {
|
|
|
386
424
|
}
|
|
387
425
|
if (existsSync(join(projectRoot, "pyproject.toml")) || existsSync(join(projectRoot, "requirements.txt")) || existsSync(join(projectRoot, "setup.py"))) {
|
|
388
426
|
const triggerFile = existsSync(join(projectRoot, "pyproject.toml")) ? "pyproject.toml" : existsSync(join(projectRoot, "requirements.txt")) ? "requirements.txt" : "setup.py";
|
|
389
|
-
return { ...container, stack: "python", testCmd: pytestCmd(), buildCmd: `${pythonBin()} -m compileall -q .`, triggerFile };
|
|
427
|
+
return { ...container, stack: "python", testCmd: pytestCmd(process.env, projectRoot), buildCmd: `${pythonBin()} -m compileall -q .`, triggerFile };
|
|
390
428
|
}
|
|
391
429
|
if (existsSync(join(projectRoot, "mix.exs"))) {
|
|
392
430
|
return { ...container, stack: "mix", testCmd: "mix test", buildCmd: "mix compile", triggerFile: "mix.exs" };
|
|
@@ -470,7 +508,7 @@ export function detectPolyglotStack(projectRoot = process.cwd()) {
|
|
|
470
508
|
}
|
|
471
509
|
const pyFile = rootFiles.find((f) => f.endsWith(".py"));
|
|
472
510
|
if (pyFile) {
|
|
473
|
-
return { ...container, stack: "python", testCmd: pytestCmd(), buildCmd: `${pythonBin()} -m compileall -q .`, triggerFile: pyFile };
|
|
511
|
+
return { ...container, stack: "python", testCmd: pytestCmd(process.env, projectRoot), buildCmd: `${pythonBin()} -m compileall -q .`, triggerFile: pyFile };
|
|
474
512
|
}
|
|
475
513
|
} catch (_) {}
|
|
476
514
|
|
|
@@ -966,19 +1004,19 @@ export function bootstrapZeroTestRepo(root = process.cwd(), options = {}) {
|
|
|
966
1004
|
try {
|
|
967
1005
|
let rawConfig = readFileSync(configPath, "utf-8");
|
|
968
1006
|
if (/^\s*test:\s*.*$/m.test(rawConfig)) {
|
|
969
|
-
rawConfig = rawConfig.replace(/^\s*test:\s*.*$/m, ` test:
|
|
1007
|
+
rawConfig = rawConfig.replace(/^\s*test:\s*.*$/m, () => ` test: ${yamlScalar(testCmd)}`);
|
|
970
1008
|
} else if (/^\s*verify:\s*$/m.test(rawConfig)) {
|
|
971
|
-
rawConfig = rawConfig.replace(/^\s*verify:\s*$/m, `verify:\n test:
|
|
1009
|
+
rawConfig = rawConfig.replace(/^\s*verify:\s*$/m, () => `verify:\n test: ${yamlScalar(testCmd)}`);
|
|
972
1010
|
} else {
|
|
973
|
-
rawConfig += `\nverify:\n test:
|
|
1011
|
+
rawConfig += `\nverify:\n test: ${yamlScalar(testCmd)}\n`;
|
|
974
1012
|
}
|
|
975
1013
|
if (detected.buildCmd && /^\s*build:\s*["']?["']?\s*$/m.test(rawConfig)) {
|
|
976
|
-
rawConfig = rawConfig.replace(/^\s*build:\s*.*$/m, ` build:
|
|
1014
|
+
rawConfig = rawConfig.replace(/^\s*build:\s*.*$/m, () => ` build: ${yamlScalar(detected.buildCmd)}`);
|
|
977
1015
|
}
|
|
978
1016
|
writeFileSync(configPath, rawConfig, "utf-8");
|
|
979
1017
|
} catch (_) {}
|
|
980
1018
|
} else {
|
|
981
|
-
const cfg = `version: 1\nprovider: jules\ntier: free\nverify:\n test:
|
|
1019
|
+
const cfg = `version: 1\nprovider: jules\ntier: free\nverify:\n test: ${yamlScalar(testCmd)}\n build: ${yamlScalar(detected.buildCmd || "")}\nlimits:\n diff_kb: 75\n daily_tasks: 15\n repair_attempts: 3\nbranch_prefix: agent/\nbase_branch: main\n`;
|
|
982
1020
|
writeFileSync(configPath, cfg, "utf-8");
|
|
983
1021
|
}
|
|
984
1022
|
|
|
@@ -987,12 +1025,12 @@ export function bootstrapZeroTestRepo(root = process.cwd(), options = {}) {
|
|
|
987
1025
|
try {
|
|
988
1026
|
let rawJules = readFileSync(julesPath, "utf-8");
|
|
989
1027
|
if (/^\s*test_cmd:\s*.*$/m.test(rawJules)) {
|
|
990
|
-
rawJules = rawJules.replace(/^\s*test_cmd:\s*.*$/m, `test_cmd:
|
|
1028
|
+
rawJules = rawJules.replace(/^\s*test_cmd:\s*.*$/m, () => `test_cmd: ${yamlScalar(testCmd)}`);
|
|
991
1029
|
} else {
|
|
992
|
-
rawJules += `\ntest_cmd:
|
|
1030
|
+
rawJules += `\ntest_cmd: ${yamlScalar(testCmd)}\n`;
|
|
993
1031
|
}
|
|
994
1032
|
if (detected.buildCmd && /^\s*build_cmd:\s*["']?["']?\s*$/m.test(rawJules)) {
|
|
995
|
-
rawJules = rawJules.replace(/^\s*build_cmd:\s*.*$/m, `build_cmd:
|
|
1033
|
+
rawJules = rawJules.replace(/^\s*build_cmd:\s*.*$/m, () => `build_cmd: ${yamlScalar(detected.buildCmd)}`);
|
|
996
1034
|
}
|
|
997
1035
|
writeFileSync(julesPath, rawJules, "utf-8");
|
|
998
1036
|
} catch (_) {}
|
package/src/wizard-init.mjs
CHANGED
|
@@ -1,12 +1,12 @@
|
|
|
1
1
|
import { existsSync, readFileSync, writeFileSync, openSync, fsyncSync, closeSync, renameSync, mkdirSync, readdirSync } from "node:fs";
|
|
2
2
|
import { join } from "node:path";
|
|
3
|
-
import { parseYaml, TIER_PRESETS, VENDOR_TIERS, FALLBACK_TIER } from "./config.mjs";
|
|
3
|
+
import { parseYaml, yamlScalar, TIER_PRESETS, VENDOR_TIERS, FALLBACK_TIER } from "./config.mjs";
|
|
4
4
|
import { suggestProvider, detectAvailableProviders } from "./provider-readiness.mjs";
|
|
5
5
|
import { detectDefaultBranch } from "./git.mjs";
|
|
6
6
|
import { resolveWorkspaceBoundary, oracleCandidates } from "./stack-detector.mjs";
|
|
7
7
|
import { PROFILE_NAMES, PROFILE_DESCRIPTIONS } from "./profiles.mjs";
|
|
8
8
|
import { detectStackOracles, runVerificationProbe } from "./wizard-oracle.mjs";
|
|
9
|
-
import { parseCollectedTests } from "./ops/test-collection.mjs";
|
|
9
|
+
import { parseCollectedTests, producedNoOutput, looksLikeTestSuiteCommand } from "./ops/test-collection.mjs";
|
|
10
10
|
import { select, multiSelect, input, confirm, spinner, isTTY } from "./tui.mjs";
|
|
11
11
|
import { KIT_VERSION } from "./version.mjs";
|
|
12
12
|
|
|
@@ -201,6 +201,19 @@ export function planInit(root = process.cwd(), options = {}) {
|
|
|
201
201
|
? `\nlimits:\n concurrency: ${limits.concurrency}\n daily_tasks: ${limits.daily_tasks}\n stagger_ms: ${limits.stagger_ms}\n diff_kb: ${limits.diff_kb}\n`
|
|
202
202
|
: "";
|
|
203
203
|
|
|
204
|
+
// A generated comment must not begin with an ESLint directive keyword.
|
|
205
|
+
//
|
|
206
|
+
// `global`, `globals`, `exported`, `eslint`, `eslint-disable` and friends are
|
|
207
|
+
// configuration when they open a comment — in any language ESLint has a
|
|
208
|
+
// parser for, YAML included. This template began a line with "global runs
|
|
209
|
+
// the ...", which ESLint read as `/* global runs, the, ... */`: a declaration
|
|
210
|
+
// of globals named after each word of the sentence. Measured on
|
|
211
|
+
// `unjs/unimport`, that produced 18 `no-unused-vars` errors quoting
|
|
212
|
+
// individual English words back at the user, on a file `init` had written
|
|
213
|
+
// thirty seconds earlier.
|
|
214
|
+
//
|
|
215
|
+
// The word is unavoidable — `global` is the name of the setting being
|
|
216
|
+
// explained — so the sentence leads with the key instead.
|
|
204
217
|
const configYaml = `# Agent Orchestrator Kit Config (v${KIT_VERSION})
|
|
205
218
|
# provider: jules | claude-code | codex | gemini-flash (agentctl providers)
|
|
206
219
|
version: 1
|
|
@@ -212,16 +225,16 @@ ${limitsBlock}
|
|
|
212
225
|
verify:
|
|
213
226
|
# minimal | standard | max — see: agentctl profile --list
|
|
214
227
|
profile: ${profile}
|
|
215
|
-
# global runs the repository
|
|
216
|
-
# to their sub-projects
|
|
228
|
+
# scope: global runs the commands this repository declares; affected
|
|
229
|
+
# resolves changed files to their sub-projects, running only those suites
|
|
217
230
|
scope: ${verifyScope}
|
|
218
|
-
test:
|
|
231
|
+
test: ${yamlScalar(verify.test)}
|
|
219
232
|
# How long a verification stage may run before the gate kills it (default
|
|
220
233
|
# 300000). Raise it for a suite that legitimately takes longer.
|
|
221
234
|
timeout_ms: 300000
|
|
222
|
-
build:
|
|
223
|
-
lint:
|
|
224
|
-
typecheck:
|
|
235
|
+
build: ${yamlScalar(verify.build)}
|
|
236
|
+
lint: ${yamlScalar(verify.lint)}
|
|
237
|
+
typecheck: ${yamlScalar(verify.typecheck)}
|
|
225
238
|
|
|
226
239
|
presets:
|
|
227
240
|
${selectedPresets.map((p) => ` - ${p}`).join("\n")}
|
|
@@ -246,11 +259,11 @@ ${selectedPresets.map((p) => ` - ${p}`).join("\n")}
|
|
|
246
259
|
|
|
247
260
|
const julesYaml = `# Google Jules Repository Configuration (Version 2)
|
|
248
261
|
version: 2
|
|
249
|
-
test_cmd:
|
|
250
|
-
build_cmd:
|
|
262
|
+
test_cmd: ${yamlScalar(verify.test)}
|
|
263
|
+
build_cmd: ${yamlScalar(verify.build)}
|
|
251
264
|
forbidden_paths:
|
|
252
|
-
${forbiddenPaths.map((p) => ` -
|
|
253
|
-
allow_paths: ${allowPaths.length > 0 ? "\n" + allowPaths.map((p) => ` -
|
|
265
|
+
${forbiddenPaths.map((p) => ` - ${yamlScalar(p)}`).join("\n")}
|
|
266
|
+
allow_paths: ${allowPaths.length > 0 ? "\n" + allowPaths.map((p) => ` - ${yamlScalar(p)}`).join("\n") : "[]"}
|
|
254
267
|
`;
|
|
255
268
|
|
|
256
269
|
return {
|
|
@@ -330,8 +343,18 @@ export function loadPresets(root = process.cwd()) {
|
|
|
330
343
|
* cost of rejecting a candidate is trying the next one, where at gate time it
|
|
331
344
|
* would be a hard red on a repository that is fine.
|
|
332
345
|
*/
|
|
333
|
-
function probeVerdict(probeRes) {
|
|
346
|
+
function probeVerdict(probeRes, cmd) {
|
|
334
347
|
if (!probeRes.ok) return "failed";
|
|
348
|
+
// Writing nothing at all is `empty`, not `silent`. The distinction is the
|
|
349
|
+
// whole point: `silent` is the forgiving bucket that keeps a command the
|
|
350
|
+
// guard could not read, and `pnpm -r test` landed in it because it prints
|
|
351
|
+
// no output to be unreadable. So the candidate this verdict was introduced
|
|
352
|
+
// to reject was the one case it waved through, and every repository
|
|
353
|
+
// scaffolded on such a workspace kept it.
|
|
354
|
+
// Same rule as the gate's floor, from the same predicate: a command that
|
|
355
|
+
// claims to run a suite and printed nothing ran none. A static gate that
|
|
356
|
+
// printed nothing did what it promised, so it keeps the forgiving verdict.
|
|
357
|
+
if (looksLikeTestSuiteCommand(cmd) && producedNoOutput(probeRes.stdout, probeRes.stderr)) return "empty";
|
|
335
358
|
const { count } = parseCollectedTests(probeRes.stdout, probeRes.stderr);
|
|
336
359
|
if (count === null) return "silent";
|
|
337
360
|
return count > 0 ? "ran" : "empty";
|
|
@@ -341,7 +364,7 @@ async function resolveRunnableOracle(root, testCmd, options = {}) {
|
|
|
341
364
|
if (!testCmd) return testCmd;
|
|
342
365
|
const probeSp = spinner(`Probing oracle: ${testCmd}`, options);
|
|
343
366
|
const probeRes = await runVerificationProbe(testCmd, root);
|
|
344
|
-
const verdict = probeVerdict(probeRes);
|
|
367
|
+
const verdict = probeVerdict(probeRes, testCmd);
|
|
345
368
|
if (verdict === "ran") {
|
|
346
369
|
probeSp.stop(`Oracle verified successfully (${probeRes.durationMs}ms)`);
|
|
347
370
|
return testCmd;
|
|
@@ -364,7 +387,7 @@ async function resolveRunnableOracle(root, testCmd, options = {}) {
|
|
|
364
387
|
for (const cand of alternates) {
|
|
365
388
|
const altSp = spinner(`Trying ${cand}`, options);
|
|
366
389
|
const altRes = await runVerificationProbe(cand, root);
|
|
367
|
-
const altVerdict = probeVerdict(altRes);
|
|
390
|
+
const altVerdict = probeVerdict(altRes, cand);
|
|
368
391
|
if (altVerdict === "ran") {
|
|
369
392
|
altSp.stop(`${cand} runs here (${altRes.durationMs}ms) — using it instead`);
|
|
370
393
|
return cand;
|
package/src/wizard-task.mjs
CHANGED
|
@@ -261,7 +261,21 @@ ${fullPrompt}
|
|
|
261
261
|
* @returns {Promise<{ ok: boolean, dryRun: boolean, taskFile: string, written: boolean, plan: object }>}
|
|
262
262
|
*/
|
|
263
263
|
export async function runTaskCreateWizard(root = process.cwd(), options = {}) {
|
|
264
|
-
|
|
264
|
+
// `-p` is the documented way to skip the questions, so it has to skip them.
|
|
265
|
+
//
|
|
266
|
+
// README: "Pass the prompt to skip straight to review:
|
|
267
|
+
// npx jules-orchestrator-kit task create -p 'Refactor the invoice module'".
|
|
268
|
+
// Interactivity was decided by `isTTY` alone, and `-p` was consulted only by
|
|
269
|
+
// the TODO-import branch below — so in a real terminal the advertised
|
|
270
|
+
// quickstart stopped at "? Task Title" and waited for a keypress forever,
|
|
271
|
+
// then asked for the instructions it had already been handed.
|
|
272
|
+
//
|
|
273
|
+
// It looked fine under test because a non-TTY run takes the headless path and
|
|
274
|
+
// never asks. That is the same defect the flag itself has: one rule, two
|
|
275
|
+
// paths, and only the path nobody was watching kept the old answer. Making
|
|
276
|
+
// `-p` mean the headless path everywhere makes the two agree by construction.
|
|
277
|
+
const promptSupplied = typeof options.prompt === "string" && options.prompt.trim() !== "";
|
|
278
|
+
const interactive = options.interactive !== false && !promptSupplied && isTTY(options.stdin || process.stdin);
|
|
265
279
|
|
|
266
280
|
let title = options.title;
|
|
267
281
|
let promptText = options.prompt;
|