jules-orchestrator-kit 0.69.0 → 0.71.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +53 -0
- package/README.md +1 -1
- package/ROADMAP_V1.md +18 -3
- package/bin/agentctl.mjs +61 -1
- package/package.json +1 -1
- package/scripts/guard-reach-check.mjs +46 -3
- package/src/config.mjs +88 -3
- package/src/engine.mjs +139 -24
- package/src/guard-policy.mjs +174 -0
- package/src/ops/test-collection.mjs +74 -0
- package/src/security.mjs +113 -2
- package/src/session-ops.mjs +117 -10
- package/src/stack-detector.mjs +52 -14
- package/src/wizard-init.mjs +38 -15
- package/src/wizard-task.mjs +15 -1
package/CHANGELOG.md
CHANGED
|
@@ -5,6 +5,57 @@ All notable changes to this project will be documented in this file.
|
|
|
5
5
|
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
|
|
6
6
|
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
7
7
|
|
|
8
|
+
## [0.71.0] - 2026-09-05
|
|
9
|
+
*A blanket is not a check, and silence is not a suite.*
|
|
10
|
+
|
|
11
|
+
A fourth cold-start trial, against v0.70.0 on four repositories nobody here chose. Seven findings, and this is the first trial where every one of them held: no quoted output was fabricated, and all four repository SHAs matched. Two of them prove entries in this file wrong, which is recorded below rather than quietly amended.
|
|
12
|
+
|
|
13
|
+
The two worst were the two halves of the same sentence — a gate that approves against nothing, and a gate that rejects what is fine.
|
|
14
|
+
|
|
15
|
+
### Fixed
|
|
16
|
+
- **A Command That Printed Nothing Was An Oracle (`src/ops/test-collection.mjs`, `src/wizard-init.mjs`)**: `pnpm -r test` on a workspace whose packages declare no test script exits 0 and writes **zero bytes** on stdout and stderr. Measured on `unjs/unimport`: `init` chose it, and `agentctl check` returned `APPROVED (Exit 0)` with an advisory that it could not tell a full suite from a command that ran nothing. The collection floor is one-sided on purpose — an unrecognised runner states no count and passes, because hard-redding every runner not on the list would be worse than the hole it closes — but an unrecognised runner still *prints*. Zero bytes is not a dialect the list has yet to learn; it is a command that ran nothing, and it is now the one absence that decides. v0.69.0 claimed `init` already rejected this candidate: it did not, because `probeVerdict` sorted no-output into `silent` (the forgiving bucket) rather than `empty`, so the exact command that verdict was introduced to reject was the one case it waved through. The rule reads the command as well as the output, and both halves are required: `tsc --noEmit`, `node --check`, `go vet` and `python3 -m compileall` also exit 0 in silence, and two of those are commands this kit writes itself for a repository with no suite yet — keying on silence alone hard-redded every one of them, which is the same first-run rejection of correct code the floor is otherwise careful to avoid. So a command that is *recognisably* a suite invocation and printed nothing ran no suite, while a checker that printed nothing did what it promised. Resolved once, in `producedNoOutput` and `looksLikeTestSuiteCommand`, and consumed by both the floor and the probe.
|
|
17
|
+
- **Python Was Verified Against A Different Copy Of Itself (`src/stack-detector.mjs`)**: `python3 -m pytest` puts the *working directory* on `sys.path`, which is the fix for a module at the root and no help at all when the package lives under `src/` — there, `import iniconfig` finds nothing locally and falls through to site-packages. Measured on `pytest-dev/iniconfig` with `_parse.py` gutted to `return False`: 49 tests passed and the gate returned `APPROVED (Exit 0)`. The suite really did run, so no amount of counting collected tests can see it; only the import path can. A `src/` layout is now detected by the `__init__.py` under it and the oracle is written as `PYTHONPATH=src python3 -m pytest`, which puts the working tree first whether or not the package is also installed.
|
|
18
|
+
- **Ordinary Work On A Test File Was A CRITICAL Block (`src/security.mjs`)**: `assertionsSeen === 0 && (unreadable.length > 0 || examined > 0)` decided `UNREADABLE`, and that `|| examined > 0` threw away the distinction the whole apparatus exists to draw. `ASSERTION_SHAPED` and `unreadable[]` were built to separate *"assertion-shaped lines were present and none parsed"* from *"there were no assertions in these lines at all"*. Collapsed together, **any** changed substantive line in a test file with no recognised assertion became a `CRITICAL` rejection: renaming a test function did it, and so did adding `import os`. The tell was in the finding itself — `file: null`, `line: null`, no sample, because `unreadable` was empty — so the guard blocked while holding no evidence of anything, and advised a **pytest** repository that its assertion library might be unsupported, from a list that names pytest. Worse, the remedy it printed was `verify.tamperGuard: "warn"`, which switches the real checks off along with the blanket. `UNREADABLE` now requires its evidence. Nothing is weakened: a removed or rewritten assertion is a recognised assertion line, raises `assertionsSeen`, and never reached this branch.
|
|
19
|
+
- **The Scaffold Failed The Repository's Own Linter (`src/config.mjs`, `src/wizard-init.mjs`, `src/stack-detector.mjs`)**: `init` wrapped every generated string in double quotes — valid YAML, idiomatic nowhere. Measured on `unjs/unimport`: following the README verbatim produced **40 eslint errors** (`yaml/quotes`, `yaml/plain-scalar`) across `.agent/config.yml` and `.agent/jules.yml`, so a newcomer's first `agentctl check` was a hard red on files `agentctl init` had written a minute earlier, in a repository that was green before. Two causes, and the second was the better one. `yamlScalar` now emits a plain scalar where YAML permits one and single quotes where it does not, which is simply how YAML is written — a glob beginning with `*` stays quoted, because a leading star is an alias reference. That took it from 40 errors to 16. The remaining 16 came from a single generated comment beginning with the word `global`, which ESLint reads as the `/* global a, b */` directive in any language it has a parser for, YAML included — so it declared globals named after each English word of the sentence and then reported them as unused. The sentence leads with the key it explains instead. Nothing in either fix is specific to eslint or to that repository, and the emitter rule lives in one place that both writers use. Measured again on the same clone: 40 errors to 0.
|
|
20
|
+
- **A Test-Integrity Failure Told The Operator To Rotate Their Keys (`bin/agentctl.mjs`)**: `TEST_DIALECT_UNREADABLE` was missing from `INTEGRITY_TYPES`, so the one finding that is most obviously not a secret fell through to the exit-6 secret branch. The screen printed `Phase [SECRETS (test integrity — no secret found)] : ❌ FAIL` and then `💡 Remediation Hint (Exit 6 Secret Leak Prevented) • High-entropy credential or secret detected in patch • rotate any exposed keys immediately` — two statements about the same run, contradicting each other in the same output. v0.67.0 fixed the phase label and left the hint; v0.69.0 recorded the hint as already correct without measuring it. It has its own branch now, and it is derived from `onlyIntegrityFindings` rather than tested independently, so a run carrying both an unreadable file and a real tampering finding gets the tampering advice instead of falling past both into the secret branch — which is the shape of the original hole.
|
|
21
|
+
- **`task create -p` Blocked Forever In A Real Terminal (`src/wizard-task.mjs`)**: the README advertises `task create -p "..."` as the way to *"skip straight to review"*. Interactivity was decided by `isTTY` alone and `-p` was consulted only by the TODO-import branch, so in a terminal the advertised quickstart stopped at `? Task Title` and waited for a keypress that never came — then asked for the instructions it had already been handed. v0.69.0 recorded this as not reproduced, correctly, from a non-TTY run: the headless path never asks, so the defect was invisible exactly where it was measured. `-p` now means the headless path everywhere, and the regression test supplies no answers at all, because only an assertion that nothing is asked could have caught a hang.
|
|
22
|
+
|
|
23
|
+
### Added
|
|
24
|
+
- **A Test Renamed Out Of Its Runner's Discovery Is Its Own Finding (`src/security.mjs`)**: pytest collects `test_*` and Go collects `Test*`, so `def test_totals` becoming `def totals` deletes the test from every future run while leaving it in the file, fully written, assertions intact — and every count in the guard stays level. This was previously caught only as a side effect of the blanket removed above. `TEST_DEREGISTERED` names it, explains that the runner collects by name, and is waivable as `--allow-test-change deregistration` like every other kind. Exact on the remainder: `test_totals` → `totals` is a de-registration, `test_totals` → `test_totals_rounded` is a rename and stays silent.
|
|
25
|
+
- **A Loosened Run No Longer Looks Like A Strict One (`bin/agentctl.mjs`)**: every escape hatch turns a check off and the report said so nowhere — `agentctl check --allow-test-change expectation` over a rewritten expectation printed a phase list byte-identical to a clean, unloosened `APPROVED`. A reviewer, a CI log, or the author a week later could not tell that a check had been waived. An `⚠️ OVERRIDES ACTIVE` block now names each active flag, plus `verify.required: false` and `verify.tamperGuard: "warn"`, and it prints on approval as well as rejection, because approval is the case where it is load-bearing.
|
|
26
|
+
- **Eleven More Cases In The Policy Contract (`src/guard-policy.mjs`)**: five innocent edits for the work that was being rejected — adding an import, a multi-line test rename, a docstring, a type annotation, a Go fixture — four de-registration canaries across both name-driven runners, six silent-run canaries and five silent *static gates* that must keep passing — the counterweight, without which the silence rule would hard-red every honest checker. 42 canaries, 21 innocent edits, 4 de-registrations, 6 silent runs, 5 static gates.
|
|
27
|
+
- **The Innocent-Edit Contract Measures The Verdict (`test/assertion-dialects.test.mjs`, `scripts/guard-reach-check.mjs`)**: it asserted only that no *violation* was produced, and `UNREADABLE` is a status rather than a violation — so the whole list stayed green for two releases while `scanDiff` returned `CRITICAL` and the gate rejected the change. Every innocent edit is now also asserted through `scanDiff`. Anything that can end a run has to be measured where the run ends.
|
|
28
|
+
|
|
29
|
+
### Also Fixed
|
|
30
|
+
- **The Loosening Flags Reached One Phase And Not The Next (`src/engine.mjs`)**: `scanDiff` is given `allowTestChanges` and honours it; the `assert:test-integrity` verify stage received only its own stage object and re-ran the same guard with none of them. So an override was accepted by one phase and ignored by the next — `--allow-test-change deregistration` turned `SECRETS` green and then failed the run at anti-tamper, having already told the operator to pass the flag they had just passed. Found while testing the new kind, not reported by the trial. This is the eighth time in this project's history that one rule has been written in one place while another kept the old answer.
|
|
31
|
+
|
|
32
|
+
- **The Manifest Parser Truncated Any Value Containing A Hash (`src/config.mjs`)**: `parseYaml` opened a comment at the first `#` on a line, quoted or not — so `verify.test: 'pytest -k "not #slow"'` was read as `pytest -k "not`, and the gate ran a command the operator never wrote while reporting on it as if they had. A URL fragment did the same. Comment-stripping is now quote-aware and follows the YAML rule that a `#` opens a comment only at the start of a line or after whitespace, and `coerce` undoes the `''` escape that `yamlScalar` emits. Pre-existing, and it mattered more the moment this release added an emitter: a writer and a reader in one module that disagree about the same file is this project's recurring defect with both halves in view. `YAML_ROUNDTRIP_CASES` pins 19 values — hashes, colons, leading stars, apostrophes, the empty string, and the words YAML would otherwise read as booleans — through both a key and a list item.
|
|
33
|
+
|
|
34
|
+
### Corrections To Earlier Entries
|
|
35
|
+
Two `### Not Reproduced` claims in **[0.69.0]** were wrong, and both were wrong in the same way — measured somewhere the defect could not appear.
|
|
36
|
+
|
|
37
|
+
- *"The `TEST_DIALECT_UNREADABLE` remediation copy was already correct."* It was not. v0.67.0 had fixed the phase label, not the hint, and the two were on screen together. Reproduced verbatim on `pytest-dev/iniconfig` and fixed above.
|
|
38
|
+
- *"`task create -p` did not block on a prompt."* True without a TTY and false with one. Reproduced under `script(1)` and fixed above.
|
|
39
|
+
|
|
40
|
+
One finding in the fourth trial quoted output that does not reproduce: its `--allow-test-change expectation` case shows `APPROVED (Exit 0)`, where the run rejects at `Exit 4` because the reporter's own `sed` genuinely breaks the dotenv suite — the same fixture artefact this project has hit in every trial. The finding itself, that no override leaves a trace in the report, held and is fixed above.
|
|
41
|
+
|
|
42
|
+
## [0.70.0] - 2026-09-05
|
|
43
|
+
*A session that has not finished is not a session that passed.*
|
|
44
|
+
|
|
45
|
+
An audit of the Jules session layer against the API it talks to. Twelve findings, each traced to a file and line in `docs/jules-quality-plan.md`; the three below are the ones that let the kit believe something about a session that was not true.
|
|
46
|
+
|
|
47
|
+
### Fixed
|
|
48
|
+
- **The Retry Was Dispatched Without The Failure It Existed To Fix (`src/session-ops.mjs`)**: `retrySession` collected its diagnostics from `act.error`, `act.executionOutput`, `act.exitCode` and `act.status`. None of those four fields exists in the documented `Activity` type — what the API returns is `artifacts[].bashOutput.{command,output,exitCode}` and `sessionFailed.reason`. Measured against a response shaped exactly like the documentation, a `FAILED` session whose bash artifact carries `not ok 1 - invoice totals must round to cents … AssertionError: expected 10.01 to equal 10.00` with `exitCode: 1`, the retry went out with `[PREVIOUS_ATTEMPT_FAILURE_DIAGNOSTIC]` set to *"Previous session did not complete cleanly."* It was handed a sentence when it needed the assertion, the file and the line. `extractFailureDiagnostics` reads the documented fields, returns them highest-signal first so the 4000-character cut drops the least useful evidence rather than the one that failed, deduplicates repeats, and still reads the legacy spellings for provider shapes that have not been observed.
|
|
49
|
+
- **An Unfinished Session Was Reported As COMPLETED (`src/engine.mjs`)**: `pollSessionState` recognised two of the nine documented `SessionState` values and returned `String(session.status || "COMPLETED")` for every other exit — and `session.status` is never set on a fresh dispatch. Measured: a session in `AWAITING_USER_FEEDBACK` and one still `IN_PROGRESS` when the poll budget expired both came back `COMPLETED`, so `QUEUED`, `PLANNING`, `PAUSED` and a session waiting on a human were indistinguishable from success. Terminal, blocked and timed-out are now three distinct verdicts, a provider that stops answering is `unreachable` rather than a timeout describing the wrong thing, and nothing synthesises `COMPLETED` any more. The one path that still returns it is the dry-run simulation, and it is now flagged `simulated`.
|
|
50
|
+
|
|
51
|
+
### Added
|
|
52
|
+
- **A Contract For The Poll (`test/session-poll.test.mjs`)**: 22 cases across all nine `SessionState` values, the timeout and wall-clock paths, the `approvePlan` side effect granted and refused, the unreachable paths, and the dry-run short-circuit. `pollSessionState` decides whether an agent session is believed to have finished writing — every later gate builds on that — and `grep -rn "pollSession" test/` returned nothing. That is the same class of hole `scripts/guard-reach-check.mjs` exists to close, in the one function that most needed it.
|
|
53
|
+
- **A Green Run Is Not Evidence (`test/session-ops.test.mjs`)**: the opposite failure of a diagnostic collector is one that flags everything, which buries the command that failed under a hundred that passed. `pytest` printing `13 passed`, node:test printing `# fail 0` and `go test` printing `ok` on exit 0 each yield no diagnostics.
|
|
54
|
+
|
|
55
|
+
### Changed
|
|
56
|
+
- **`agentctl retry` Says When The Trace Is The Fallback (`bin/agentctl.mjs`, `src/session-ops.mjs`)**: `failureReason` alone cannot tell a real trace from the generic sentence — both are non-empty strings. `retrySession` now returns `diagnosticsFound` and `diagnosticSources`, and the CLI prints the count and where the evidence came from, or says plainly that the retry is going out with nothing but the generic line.
|
|
57
|
+
- **A Non-Terminal Session Is Announced Before The Gate Runs (`src/engine.mjs`)**: the repair loop polled the session and discarded the answer, then ran re-verification against a tree the agent might not have finished writing. The verdict is now read: a non-terminal session prints `[SESSION_NOT_TERMINAL]` naming what it is waiting on and appends a `session_not_terminal` telemetry event. The gate still runs either way — it is the authority on whether the change works — but it no longer runs silently on a half-applied patch.
|
|
58
|
+
|
|
8
59
|
## [0.69.0] - 2026-09-04
|
|
9
60
|
*A denominator is not evidence if the things counted in it were never read.*
|
|
10
61
|
|
|
@@ -28,6 +79,8 @@ A third cold-start trial against v0.68.0. Twelve findings; six reproduced, and t
|
|
|
28
79
|
- **The Egress Guard Could Not See `.js` (`test/egress-allowlist.test.mjs`)**: the scan collected `.mjs` only, so `bin/init.js` — published as the `jules-init` binary — sat outside the boundary entirely. Harmless as it stands, but the guard's whole purpose is that a reviewer can trust the boundary without reading every commit.
|
|
29
80
|
|
|
30
81
|
### Not Reproduced
|
|
82
|
+
> **Corrected in [0.71.0]:** two of the claims below were wrong — the `TEST_DIALECT_UNREADABLE` remediation copy and `task create -p`. Both were measured where the defect could not appear. See the corrections section of that release.
|
|
83
|
+
|
|
31
84
|
Six of the twelve did not hold, and five quoted terminal output that does not exist anywhere in the shipped code. `--allow-test-change expectation` propagates correctly through both phases and returns `APPROVED`; `--allow-test-change removal` allows deleting a dead test alongside its dead function, also `APPROVED` — both reported failures were a test suite genuinely broken by the reporter's own edit, which is the same fixture artefact this project has hit repeatedly. `agentctl plan approve --dry-run` errors with *"Session ID is required"* rather than printing the quoted *"Plan Approved Successfully!"*. `task create -p` did not block on a prompt. A protected `package.json` is the design, and the gate already prints `To allow protected files in this run, pass: agentctl gate --allow-protected` — the finding stated there were no flag hints in the output. The `TEST_DIALECT_UNREADABLE` remediation copy was already correct: *"Exit 6 Test Integrity Violation — no secret was found"*, not the quoted *"Secret Leak Prevented"*. And `pnpm -r test` exits 0 on pnpm 10.33.4 rather than the reported `ERR_PNPM_RECURSIVE_RUN_NO_SCRIPT` — the real defect there was worse than the one reported, and is fixed above.
|
|
32
85
|
|
|
33
86
|
### Added
|
package/README.md
CHANGED
|
@@ -208,7 +208,7 @@ To maximize PR merge rates, dispatch tasks according to deterministic boundaries
|
|
|
208
208
|
* **Fail-Closed Security & Secret Redaction:** Evaluates explicit Deny rules before Allow rules against canonicalized, case-folded paths. Redacts high-entropy keys and base64-encoded credentials (such as Kubernetes `Secret` manifests).
|
|
209
209
|
* **Complexity & Cost Router:** Zero-dependency heuristic classifier (`src/router.mjs`) routing mechanical tasks to lightweight models while reserving primary models for complex refactors, with a `node --check` syntax-verification gate that transparently escalates a FAST-tier result to the primary provider if it left broken JS on disk.
|
|
210
210
|
* **Terminal UI & Diagnostic Matrix (`agentctl doctor`):** Interactive terminal dashboard, task sidecar manager, and automated transactional self-repair.
|
|
211
|
-
* **Verified Test Suite:** Tested with **
|
|
211
|
+
* **Verified Test Suite:** Tested with **1299 unit tests across 174 suites**, green on every supported platform.
|
|
212
212
|
|
|
213
213
|
<br/>
|
|
214
214
|
|
package/ROADMAP_V1.md
CHANGED
|
@@ -11,15 +11,30 @@ The **jules-orchestrator-kit** is the zero-dependency safety gatekeeper and self
|
|
|
11
11
|
## 📌 Release Milestones Overview
|
|
12
12
|
|
|
13
13
|
```
|
|
14
|
-
v0.
|
|
15
|
-
(
|
|
14
|
+
v0.71.0 (Current Stable) ──► v0.72.0 (Distributed Swarms & Leases) ──► v1.0.0 (Production Hardened Kernel)
|
|
15
|
+
(A Blanket Is Not A Check) (Multi-Agent DAG & Resource Locks) (Enterprise Telemetry & SLA)
|
|
16
16
|
```
|
|
17
17
|
|
|
18
18
|
---
|
|
19
19
|
|
|
20
|
-
## ✅ Shipped Milestones (v0.20.0 – v0.
|
|
20
|
+
## ✅ Shipped Milestones (v0.20.0 – v0.71.0)
|
|
21
21
|
|
|
22
22
|
|
|
23
|
+
### v0.71.0: A Blanket Is Not A Check
|
|
24
|
+
- [x] **Silence Is Not A Suite (`src/ops/test-collection.mjs`, `src/wizard-init.mjs`)** — a command that claims to run tests and prints nothing ran none; a static gate that prints nothing did its job.
|
|
25
|
+
- [x] **Python Verified Against The Working Tree (`src/stack-detector.mjs`)** — a `src/` layout resolved its imports to site-packages, and broken code passed 49 tests.
|
|
26
|
+
- [x] **`UNREADABLE` Requires Its Evidence (`src/security.mjs`)** — adding an import to a test file was a CRITICAL block.
|
|
27
|
+
- [x] **A Test Renamed Out Of Discovery Is Its Own Finding (`src/security.mjs`)** — `TEST_DEREGISTERED`, waivable like every other kind.
|
|
28
|
+
- [x] **The Scaffold Passes A Repository's Own YAML Linter (`src/config.mjs`)** — 40 eslint errors on files `init` had just written.
|
|
29
|
+
- [x] **A Loosened Run Says So (`bin/agentctl.mjs`)** — an override left no trace in the report at all.
|
|
30
|
+
- [x] **`task create -p` Skips The Questions (`src/wizard-task.mjs`)** — the advertised quickstart blocked forever in a real terminal.
|
|
31
|
+
|
|
32
|
+
### v0.70.0: Not Finished Is Not Passed
|
|
33
|
+
- [x] **An Unfinished Session Is Not COMPLETED (`src/engine.mjs`)** — terminal, blocked and timed-out are three verdicts, not one.
|
|
34
|
+
- [x] **The Retry Carries The Failure (`src/session-ops.mjs`)** — it was reading four fields the API does not return.
|
|
35
|
+
- [x] **`agentctl retry` Says When The Trace Is The Fallback (`bin/agentctl.mjs`)** — a generic sentence must not look like evidence.
|
|
36
|
+
- [x] **A Contract For The Poll (`test/session-poll.test.mjs`)** — 22 cases over all nine documented session states.
|
|
37
|
+
|
|
23
38
|
### v0.69.0: Read, Not Just Counted
|
|
24
39
|
- [x] **Expected-Value-First Assertions (`src/security.mjs`)** — JUnit and PHPUnit document the order the guard read as prose.
|
|
25
40
|
- [x] **A Regex Is An Expected Value (`src/security.mjs`)** — a rewritten pattern was neither a change nor a loss.
|
package/bin/agentctl.mjs
CHANGED
|
@@ -425,6 +425,31 @@ async function main() {
|
|
|
425
425
|
} else {
|
|
426
426
|
console.log(`\n🛡️ agentctl Safety Gate Audit Results (Base: ${values.base}, Mode: ${selectedMode})`);
|
|
427
427
|
console.log(`-----------------------------------------------------`);
|
|
428
|
+
// A loosened run must not be able to pass for a strict one.
|
|
429
|
+
//
|
|
430
|
+
// Every escape hatch here turns a check off, and the report said so
|
|
431
|
+
// nowhere: `agentctl check --allow-test-change expectation` over a
|
|
432
|
+
// rewritten expectation printed a phase list byte-identical to a clean,
|
|
433
|
+
// unloosened APPROVED. Anyone reading the output afterwards — a
|
|
434
|
+
// reviewer, a CI log, the author a week later — could not tell that a
|
|
435
|
+
// check had been waived, which makes the waiver invisible exactly where
|
|
436
|
+
// it matters most. Printed before the phases, and on approval as well
|
|
437
|
+
// as rejection, because an approval is the case where it is load-bearing.
|
|
438
|
+
const overrides = [];
|
|
439
|
+
if (values["allow-protected"]) overrides.push("--allow-protected (protected paths permitted)");
|
|
440
|
+
if (values["allow-unreadable-tests"]) overrides.push("--allow-unreadable-tests (unreadable dialect permitted)");
|
|
441
|
+
if (values["allow-test-modifications"]) overrides.push("--allow-test-modifications (every tamper check waived)");
|
|
442
|
+
if (values["allow-test-change"]) {
|
|
443
|
+
const kinds = [].concat(values["allow-test-change"]).join(", ");
|
|
444
|
+
overrides.push(`--allow-test-change ${kinds} (tamper check waived: ${kinds})`);
|
|
445
|
+
}
|
|
446
|
+
if (config.verify?.required === false) overrides.push("verify.required: false (nothing is executed)");
|
|
447
|
+
if (config.verify?.tamperGuard === "warn") overrides.push('verify.tamperGuard: "warn" (unreadable dialects report only)');
|
|
448
|
+
if (overrides.length > 0) {
|
|
449
|
+
console.log(` ⚠️ OVERRIDES ACTIVE — this run is not a strict pass:`);
|
|
450
|
+
for (const o of overrides) console.log(` • ${o}`);
|
|
451
|
+
console.log(`-----------------------------------------------------`);
|
|
452
|
+
}
|
|
428
453
|
for (const p of res.phases) {
|
|
429
454
|
const status = p.ok ? "✅ PASS" : "❌ FAIL";
|
|
430
455
|
// Test tampering shares this phase, and its exit code, with the
|
|
@@ -472,13 +497,32 @@ async function main() {
|
|
|
472
497
|
// arrive under the same phase and the same code.
|
|
473
498
|
const secretsPhase = res.phases.find((p) => p.phase === "secrets" && !p.ok);
|
|
474
499
|
const findingTypes = new Set((secretsPhase?.findings || []).map((f) => f.type));
|
|
500
|
+
// `TEST_DIALECT_UNREADABLE` belongs in this set and was missing from
|
|
501
|
+
// it, so the one finding that is *most* obviously not a secret fell
|
|
502
|
+
// through to the secret hint: the report printed
|
|
503
|
+
// "Phase [SECRETS (test integrity — no secret found)] : ❌ FAIL" and
|
|
504
|
+
// then told the operator a high-entropy credential had leaked and to
|
|
505
|
+
// rotate their keys. v0.67.0 fixed the phase label and left the hint;
|
|
506
|
+
// v0.69.0 recorded the hint as already correct without measuring it.
|
|
507
|
+
// Both statements were about the same screen, and they contradicted
|
|
508
|
+
// each other on it.
|
|
475
509
|
const INTEGRITY_TYPES = new Set([
|
|
476
510
|
"TEST_TAMPERING_DETECTED",
|
|
511
|
+
"TEST_DIALECT_UNREADABLE",
|
|
477
512
|
"EDGE_RUNTIME_VIOLATION",
|
|
478
513
|
"CROSS_PACKAGE_BOUNDARY_VIOLATION",
|
|
479
514
|
]);
|
|
480
515
|
const onlyIntegrityFindings =
|
|
481
516
|
findingTypes.size > 0 && [...findingTypes].every((t) => INTEGRITY_TYPES.has(t));
|
|
517
|
+
// An unreadable dialect is not a weakened assertion either, and the
|
|
518
|
+
// integrity hint tells the operator to restore an assertion that no
|
|
519
|
+
// finding names. Split by what is actually present — and derived from
|
|
520
|
+
// `onlyIntegrityFindings` rather than tested independently, so a run
|
|
521
|
+
// carrying both an unreadable file and a real tampering finding gets
|
|
522
|
+
// the tampering advice instead of falling past both into the secret
|
|
523
|
+
// branch, which is how this hole was shaped the first time.
|
|
524
|
+
const onlyUnreadableDialect =
|
|
525
|
+
onlyIntegrityFindings && [...findingTypes].every((t) => t === "TEST_DIALECT_UNREADABLE");
|
|
482
526
|
|
|
483
527
|
// Exit 3 is also what a strictTestLock tamper verdict returns, so the
|
|
484
528
|
// code alone cannot pick the hint — a scope remediation for a rewritten
|
|
@@ -504,6 +548,16 @@ async function main() {
|
|
|
504
548
|
if (firstError) console.log(` • Last attempt error: ${firstError}`);
|
|
505
549
|
console.log(` • Review error fingerprints via: agentctl doctor\n`);
|
|
506
550
|
}
|
|
551
|
+
} else if (onlyUnreadableDialect) {
|
|
552
|
+
console.log(`💡 Remediation Hint (Exit ${res.code} Unreadable Test Dialect — no secret was found):`);
|
|
553
|
+
console.log(` • Nothing leaked. The guard could not recognise the assertions in a changed test`);
|
|
554
|
+
console.log(` file, so it could not check that file for tampering — and a check that examined`);
|
|
555
|
+
console.log(` nothing does not get to say APPROVED.`);
|
|
556
|
+
console.log(` • Best: report the dialect so it gets covered. The guard knows Node, pytest, Go,`);
|
|
557
|
+
console.log(` Rust, JUnit, RSpec, PHPUnit, Minitest, XCTest, chai and node-tap.`);
|
|
558
|
+
console.log(` • Or say so on the record: verify.tamperGuard: "warn" in .agent/config.yml`);
|
|
559
|
+
console.log(` • Or allow this one run: agentctl check --allow-unreadable-tests`);
|
|
560
|
+
console.log(` • Nothing needs rotating: this exit code is shared with the secret scanner.\n`);
|
|
507
561
|
} else if (onlyIntegrityFindings) {
|
|
508
562
|
console.log(`💡 Remediation Hint (Exit ${res.code} Test Integrity Violation — no secret was found):`);
|
|
509
563
|
console.log(` • The diff weakens or removes verification rather than leaking a credential.`);
|
|
@@ -959,7 +1013,13 @@ async function main() {
|
|
|
959
1013
|
console.log(`------------------------------------------------------------------`);
|
|
960
1014
|
console.log(` Original Session : ${res.originalSessionId}`);
|
|
961
1015
|
console.log(` New Session ID : ${res.newSession?.id || "N/A"}`);
|
|
962
|
-
console.log(
|
|
1016
|
+
console.log(
|
|
1017
|
+
` Failure Trace Added : ${
|
|
1018
|
+
res.diagnosticsFound > 0
|
|
1019
|
+
? `YES (${res.diagnosticsFound} block${res.diagnosticsFound === 1 ? "" : "s"} from ${[...new Set(res.diagnosticSources || [])].join(", ")})`
|
|
1020
|
+
: "NO — the session carried no readable diagnostics; the retry got the generic sentence"
|
|
1021
|
+
}`
|
|
1022
|
+
);
|
|
963
1023
|
console.log(`------------------------------------------------------------------\n`);
|
|
964
1024
|
}
|
|
965
1025
|
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "jules-orchestrator-kit",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.71.0",
|
|
4
4
|
"description": "Zero-dependency safety gatekeeper, test oracle generator, and multi-agent coordination protocol for autonomous coding agents — Google Jules, Claude Code, Codex and Gemini CLI.",
|
|
5
5
|
"repository": {
|
|
6
6
|
"type": "git",
|
|
@@ -42,10 +42,10 @@
|
|
|
42
42
|
* Exit codes: 0 = every guard reachable, 1 = a guard has gone silent.
|
|
43
43
|
*/
|
|
44
44
|
|
|
45
|
-
import { checkTestTampering, checkScope } from "../src/security.mjs";
|
|
45
|
+
import { checkTestTampering, checkScope, scanDiff } from "../src/security.mjs";
|
|
46
46
|
import { isTestPath } from "../src/test-paths.mjs";
|
|
47
47
|
import { normalizeScope } from "../src/config.mjs";
|
|
48
|
-
import { parseCollectedTests } from "../src/ops/test-collection.mjs";
|
|
48
|
+
import { parseCollectedTests, checkCollectionFloor } from "../src/ops/test-collection.mjs";
|
|
49
49
|
import {
|
|
50
50
|
TEST_PATH_CASES,
|
|
51
51
|
TAMPER_CANARIES,
|
|
@@ -53,6 +53,9 @@ import {
|
|
|
53
53
|
PREDICATE_MUTANTS,
|
|
54
54
|
EMPTY_RUN_CANARIES,
|
|
55
55
|
COUNTED_RUN_CANARIES,
|
|
56
|
+
SILENT_RUN_CANARIES,
|
|
57
|
+
SILENT_STATIC_GATES,
|
|
58
|
+
DEREGISTRATION_CANARIES,
|
|
56
59
|
SCOPE_CANARIES,
|
|
57
60
|
INNOCENT_EDITS,
|
|
58
61
|
UNREADABLE_DIALECTS,
|
|
@@ -154,13 +157,53 @@ const canaryResults = new Map();
|
|
|
154
157
|
add("canaries: assertion rules parsed an assertion", noAssertions.length === 0, noAssertions.length ? noAssertions.join(", ") : "assertionsSeen > 0 on every assertion finding");
|
|
155
158
|
}
|
|
156
159
|
|
|
160
|
+
{
|
|
161
|
+
const missed = DEREGISTRATION_CANARIES.filter(
|
|
162
|
+
(c) => !(checkTestTampering(canaryDiff(c)).violations || []).some((v) => v.type === "TEST_DEREGISTERED")
|
|
163
|
+
);
|
|
164
|
+
add(
|
|
165
|
+
"canaries: a test renamed out of discovery is caught",
|
|
166
|
+
missed.length === 0,
|
|
167
|
+
missed.length ? missed.map((c) => `${c.id} (${c.why})`).join("; ") : `${DEREGISTRATION_CANARIES.length} de-registrations reported`
|
|
168
|
+
);
|
|
169
|
+
}
|
|
170
|
+
|
|
171
|
+
{
|
|
172
|
+
// Zero bytes on both streams is the one absence that decides. Measured
|
|
173
|
+
// through the floor rather than the parser: the parser answers "no count
|
|
174
|
+
// stated", and it is the floor that has to refuse to certify against it.
|
|
175
|
+
const passed = SILENT_RUN_CANARIES.filter(
|
|
176
|
+
(c) => checkCollectionFloor({ ok: true, stdout: c.stdout, stderr: c.stderr, command: c.command }).ok !== false
|
|
177
|
+
);
|
|
178
|
+
add(
|
|
179
|
+
"policy: a command that printed nothing verified nothing",
|
|
180
|
+
passed.length === 0,
|
|
181
|
+
passed.length ? passed.map((c) => c.id).join(", ") : `${SILENT_RUN_CANARIES.length} silent runs refused`
|
|
182
|
+
);
|
|
183
|
+
|
|
184
|
+
// And the counterweight: a checker that succeeds in silence is doing its job.
|
|
185
|
+
const redded = SILENT_STATIC_GATES.filter(
|
|
186
|
+
(c) => checkCollectionFloor({ ok: true, stdout: "", stderr: "", command: c.command }).ok === false
|
|
187
|
+
);
|
|
188
|
+
add(
|
|
189
|
+
"policy: a silent static gate is not a silent suite",
|
|
190
|
+
redded.length === 0,
|
|
191
|
+
redded.length ? `${redded.map((c) => c.id).join(", ")} hard-redded` : `${SILENT_STATIC_GATES.length} static gates still pass`
|
|
192
|
+
);
|
|
193
|
+
}
|
|
194
|
+
|
|
157
195
|
// --- 3. The opposite failure: flagging what is innocent ---------------------
|
|
158
196
|
{
|
|
159
197
|
const noisy = [];
|
|
160
198
|
for (const e of INNOCENT_EDITS) {
|
|
161
199
|
const res = checkTestTampering(canaryDiff(e));
|
|
162
200
|
const types = (res.violations || []).map((v) => v.type);
|
|
163
|
-
if (types.length > 0) noisy.push(`${e.id}
|
|
201
|
+
if (types.length > 0) noisy.push(`${e.id} \u2192 ${JSON.stringify(types)} (${e.why})`);
|
|
202
|
+
// The verdict, not just the violation list. `UNREADABLE` is a status, so
|
|
203
|
+
// this check stayed green for two releases while the gate rejected every
|
|
204
|
+
// edit below on a CRITICAL finding.
|
|
205
|
+
const scan = scanDiff(canaryDiff(e));
|
|
206
|
+
if (!scan.ok) noisy.push(`${e.id} \u2192 gate rejected: ${JSON.stringify((scan.findings || []).map((f) => f.type))}`);
|
|
164
207
|
}
|
|
165
208
|
add(
|
|
166
209
|
"innocent edits stay silent",
|
package/src/config.mjs
CHANGED
|
@@ -205,7 +205,13 @@ const BLOCKED_KEYS = new Set(["__proto__", "constructor", "prototype"]);
|
|
|
205
205
|
|
|
206
206
|
function coerce(val) {
|
|
207
207
|
if (!val) return "";
|
|
208
|
-
if (
|
|
208
|
+
if (val.startsWith("'") && val.endsWith("'") && val.length >= 2) {
|
|
209
|
+
// `''` is how a single-quoted YAML scalar escapes a quote, and it is what
|
|
210
|
+
// `yamlScalar` emits. Reading the quotes off without undoing the escape
|
|
211
|
+
// meant the emitter and the parser disagreed about the same file.
|
|
212
|
+
return val.slice(1, -1).replace(/''/g, "'");
|
|
213
|
+
}
|
|
214
|
+
if (val.startsWith('"') && val.endsWith('"') && val.length >= 2) {
|
|
209
215
|
return val.slice(1, -1);
|
|
210
216
|
}
|
|
211
217
|
if (val === "true") return true;
|
|
@@ -315,6 +321,86 @@ export function normalizeScope(parsed = {}) {
|
|
|
315
321
|
};
|
|
316
322
|
}
|
|
317
323
|
|
|
324
|
+
/**
|
|
325
|
+
* A YAML scalar quoted only as much as YAML requires.
|
|
326
|
+
*
|
|
327
|
+
* The scaffold wrapped every generated string in double quotes, which is valid
|
|
328
|
+
* YAML and idiomatic nowhere. On a repository that lints its own YAML that is a
|
|
329
|
+
* hard red on the kit's own files: measured on `unjs/unimport`, following the
|
|
330
|
+
* README verbatim produced 40 eslint errors — `yaml/quotes` and
|
|
331
|
+
* `yaml/plain-scalar` — across `.agent/config.yml` and `.agent/jules.yml`, so a
|
|
332
|
+
* newcomer's very first `agentctl check` failed on output `agentctl init` had
|
|
333
|
+
* just written, in a repository that was green a minute earlier. Teaching a new
|
|
334
|
+
* user that the gate is broken is the most expensive thing this tool can do.
|
|
335
|
+
*
|
|
336
|
+
* Nothing here is specific to eslint or to that repository. Emitting a plain
|
|
337
|
+
* scalar where YAML permits one, and single quotes where it does not, is simply
|
|
338
|
+
* how YAML is written; the linters that check it are agreeing with the spec.
|
|
339
|
+
*
|
|
340
|
+
* The quoting test is deliberately conservative — anything that could change
|
|
341
|
+
* meaning unquoted gets quotes. A glob beginning with a star stays quoted,
|
|
342
|
+
* because a leading `*` is an alias reference; `pnpm -r test` does not, because
|
|
343
|
+
* there is nothing in it to misread.
|
|
344
|
+
*/
|
|
345
|
+
const YAML_NEEDS_QUOTES =
|
|
346
|
+
/^$|^[-?:,[\]{}#&*!|>'"%@`]|^\s|\s$|:\s|\s#|:$|[\n\r\t]|^(?:true|false|yes|no|on|off|null|~)$/i;
|
|
347
|
+
|
|
348
|
+
export function yamlScalar(value) {
|
|
349
|
+
const s = value === null || value === undefined ? "" : String(value);
|
|
350
|
+
if (!YAML_NEEDS_QUOTES.test(s) && !/^[+-]?(?:\d[\d_]*)(?:\.[\d_]*)?(?:[eE][+-]?\d+)?$/.test(s)) return s;
|
|
351
|
+
return `'${s.replace(/'/g, "''")}'`;
|
|
352
|
+
}
|
|
353
|
+
|
|
354
|
+
/**
|
|
355
|
+
* Can a quoted scalar begin at this offset?
|
|
356
|
+
*
|
|
357
|
+
* A quote character is only an opening quote where a value may start — after
|
|
358
|
+
* `key:`, after a `-` list marker, at the head of a flow collection, or at the
|
|
359
|
+
* start of the line. Anywhere else it is an ordinary apostrophe, and treating
|
|
360
|
+
* `test: it's fine` as an unterminated string would swallow the rest of it.
|
|
361
|
+
*/
|
|
362
|
+
function opensQuotedScalar(line, i) {
|
|
363
|
+
return /(?:^\s*|:\s+|-\s+|[[{,]\s*)$/.test(line.slice(0, i));
|
|
364
|
+
}
|
|
365
|
+
|
|
366
|
+
/**
|
|
367
|
+
* Strip a trailing `#` comment without cutting into a quoted scalar.
|
|
368
|
+
*
|
|
369
|
+
* `rawLine.indexOf("#")` stood here, which is not the YAML rule twice over: a
|
|
370
|
+
* `#` opens a comment only when it begins the line or follows whitespace, and
|
|
371
|
+
* never inside quotes. So any configured value containing a hash was silently
|
|
372
|
+
* truncated on the way in — `test: 'pytest -k "not #slow"'` became
|
|
373
|
+
* `test: 'pytest -k "not`, and the gate then ran a command the user never
|
|
374
|
+
* wrote. Silent truncation of a verification command is the worst shape this
|
|
375
|
+
* file can produce, because every later phase reports on whatever did run.
|
|
376
|
+
*
|
|
377
|
+
* It matters more now that `yamlScalar` exists: an emitter and a parser that
|
|
378
|
+
* disagree about the same file is this project's recurring defect with both
|
|
379
|
+
* halves in one module.
|
|
380
|
+
*/
|
|
381
|
+
function stripYamlComment(rawLine) {
|
|
382
|
+
let quote = null;
|
|
383
|
+
for (let i = 0; i < rawLine.length; i++) {
|
|
384
|
+
const ch = rawLine[i];
|
|
385
|
+
if (quote) {
|
|
386
|
+
if (quote === '"' && ch === "\\") {
|
|
387
|
+
i++;
|
|
388
|
+
} else if (ch === quote) {
|
|
389
|
+
// `''` is an escaped quote inside a single-quoted scalar, not the end.
|
|
390
|
+
if (quote === "'" && rawLine[i + 1] === "'") i++;
|
|
391
|
+
else quote = null;
|
|
392
|
+
}
|
|
393
|
+
continue;
|
|
394
|
+
}
|
|
395
|
+
if ((ch === '"' || ch === "'") && opensQuotedScalar(rawLine, i)) {
|
|
396
|
+
quote = ch;
|
|
397
|
+
continue;
|
|
398
|
+
}
|
|
399
|
+
if (ch === "#" && (i === 0 || /\s/.test(rawLine[i - 1]))) return rawLine.slice(0, i);
|
|
400
|
+
}
|
|
401
|
+
return rawLine;
|
|
402
|
+
}
|
|
403
|
+
|
|
318
404
|
/**
|
|
319
405
|
* Indent-stack zero-dependency YAML parser with prototype pollution protection.
|
|
320
406
|
*/
|
|
@@ -331,8 +417,7 @@ export function parseYaml(src) {
|
|
|
331
417
|
const lines = src.split("\n");
|
|
332
418
|
|
|
333
419
|
for (let rawLine of lines) {
|
|
334
|
-
const
|
|
335
|
-
const line = commentIdx !== -1 ? rawLine.slice(0, commentIdx) : rawLine;
|
|
420
|
+
const line = stripYamlComment(rawLine);
|
|
336
421
|
const trimmed = line.trim();
|
|
337
422
|
if (!trimmed) continue;
|
|
338
423
|
|