jules-orchestrator-kit 0.70.0 → 0.72.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/.agent/rules/jules-protocol.md +0 -1
- package/CHANGELOG.md +66 -0
- package/JULES_RULES_TEMPLATE.md +0 -1
- package/README.md +56 -1
- package/ROADMAP_V1.md +31 -6
- package/bin/agentctl.mjs +101 -8
- package/package.json +1 -1
- package/scripts/guard-reach-check.mjs +74 -8
- package/src/config.mjs +434 -5
- package/src/coverage.mjs +85 -13
- package/src/engine.mjs +106 -130
- package/src/git.mjs +104 -1
- package/src/guard-policy.mjs +581 -0
- package/src/ops/command-registry.mjs +109 -0
- package/src/ops/test-collection.mjs +103 -3
- package/src/security.mjs +554 -4
- package/src/stack-detector.mjs +73 -14
- package/src/test-paths.mjs +18 -1
- package/src/wizard-init.mjs +54 -27
- package/src/wizard-oracle.mjs +1 -1
- package/src/wizard-task.mjs +15 -1
|
@@ -36,4 +36,3 @@ To maximize the ratio of mergeable PRs vs. failed or hallucinated sessions:
|
|
|
36
36
|
15. **Web Excellence & Frontend Guardrails:** Enforce quantitative Core Web Vitals (LCP < 1.2s, CLS < 0.05), WCAG 2.2 AA/AAA semantic accessibility, Schema.org JSON-LD compliance, and Playwright multi-viewport responsive testing.
|
|
37
37
|
16. **Airtight Positive Enclosures ("Pink Elephant" Rule):** Replace long negative constraint lists with strict positive operational perimeters (`ONLY modify [Target/Module]`) to prevent attention-drift in deep context windows.
|
|
38
38
|
17. **Sterile / Clinical Vocabulary Mandate:** Eradicate aggressive verbs (`kill`, `amputate`, `sabotage`, `destroy`) from prompts and policies. Use clinical equivalents (`terminate PID`, `prune code`, `mutate test logic`, `purge cache`) to prevent false-positive safety classifier trips in Google Cloud VMs.
|
|
39
|
-
|
package/CHANGELOG.md
CHANGED
|
@@ -5,6 +5,70 @@ All notable changes to this project will be documented in this file.
|
|
|
5
5
|
The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
|
|
6
6
|
and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
|
|
7
7
|
|
|
8
|
+
## [0.72.0] - 2026-09-08
|
|
9
|
+
*A guard is not a policy if the diff under review can rewrite the rules.*
|
|
10
|
+
|
|
11
|
+
A fifth cold-start audit, against published v0.71.0 on four polyglot repositories (`p-limit`, `itsdangerous`, `is`, `mini_markdown`). Twenty-two findings across three tiers:
|
|
12
|
+
- Tier 1 (Security & Isolation): uncommitted scaffold tampering, untrusted working-tree configs in staged/committed evaluation, and unisolated snapshot leaks.
|
|
13
|
+
- Tier 2 (Tamper Guard Hardening): root test file canonicalization, conditional assertion bypasses, de-registration detection, impossible preconditions, and uncovered new modules.
|
|
14
|
+
- Tier 3 (False Reds Elimination & CLI Polish): EOF markdown blank lines, cargo clippy warning defaults, multi-target Cargo aggregation, lockfile supply chain remediation diagnostics, `--strict-locks` CLI parsing, truthful waiver telemetry, dry-run evidence suppression, focused subcommand help routing, and documented repository uninstall.
|
|
15
|
+
|
|
16
|
+
### Fixed
|
|
17
|
+
- **Authoritative Base Policy Resolution (`src/config.mjs`, `src/engine.mjs`)**: Staged and committed gate modes evaluated `.agent/config.yml` from the dirty working copy on disk rather than the base commit under review. An agent diff could raise `diffKb` above 75 KB or replace verification commands with no-ops. `resolveTrustedPolicy()` now extracts authoritative verification stages and security policy from `git show <base>:.agent/config.yml` (or `.agent/jules.yml`), rejecting untrusted disk edits and committed `base: HEAD` self-comparisons (F06, F07, F08).
|
|
18
|
+
- **Uncommitted Scaffold Integrity (`src/config.mjs`)**: `checkBootstrapPolicyIntegrity()` inspects uncommitted scaffolds in bootstrap mode, refusing configurations that replace verification with no-ops (`process.exit(0)`, `sh -c :`), lower verification profiles to `minimal`, or disable `strictTestLock` / `tamperGuard` prior to the initial commit (F06).
|
|
19
|
+
- **Snapshot Materialization & Ephemeral Sandbox Isolation (`src/git.mjs`, `src/engine.mjs`)**: Phase 4 verification ran directly in the dirty repository tree, allowing uncommitted edits outside the evaluated diff to leak into verification runs. `materializeSnapshot()` now extracts the staged index (`git checkout-index`) or committed revision into temporary detached worktrees with symlinked dependency trees (`node_modules`, `.venv`), ensuring hermetic execution (F10).
|
|
20
|
+
- **Python src-layout Invariant (`src/stack-detector.mjs`, `src/engine.mjs`)**: Packages structured under `src/` without root packaging could resolve imports to stale site-packages or fail discovery. `isSrcLayout()` now detects `src/<pkg>/` topologies and automatically prepends `src/` to `PYTHONPATH` during verification (F11).
|
|
21
|
+
- **Empty Test Collection Canaries (`src/ops/test-collection.mjs`, `src/guard-policy.mjs`)**: Go's `ok ... [no tests to run]` and pytest's `--collect-only` exited 0 while verifying nothing. Both are now classified as `count: 0` empty collections, activating empty-run canaries (F09).
|
|
22
|
+
- **Cargo Multi-Target Aggregation (`src/ops/test-collection.mjs`)**: Multi-target Cargo suites (lib unit tests + integration tests) print `running N tests` per target. `parseCollectedTests` now aggregates all target summaries via `matchAll`, preventing a 0-test unit target from hiding 58 passing integration tests (F15).
|
|
23
|
+
- **Cargo Clippy Oracle Defaults (`src/wizard-oracle.mjs`)**: Removed forced `-D warnings` from default `cargo clippy` candidates, preventing green external repositories with benign compiler warnings from failing verification gates (F14).
|
|
24
|
+
- **Truthful Waiver & Override Telemetry (`bin/agentctl.mjs`)**: Added active waiver auditing for `JULES_ALLOW_COMMAND_FILE_CHANGES`, `minTests: 0`, and non-required stage failures, and corrected `verify.required: false` explanation from `(nothing is executed)` to `(verification failures and empty test suites permitted)` (F18, F19).
|
|
25
|
+
- **Dry-Run Evidence Suppression (`src/engine.mjs`, `bin/agentctl.mjs`)**: Suppressed `.agent/evidence/` disk writes when `--dry-run` or `JULES_DRY_RUN=1` is active, and added visible `[DRY-RUN]` markers to `agentctl plan approve` and `agentctl session get` (F20).
|
|
26
|
+
- **Subcommand Help Routing (`bin/agentctl.mjs`, `src/ops/command-registry.mjs`)**: Registered descriptors for `pr harvest`, `session get`, and `plan approve`, and routed `agentctl help <subcommand>` to focused subcommand usage rather than dumping all 50 commands (F21).
|
|
27
|
+
- **Markdown EOF Hygiene (`JULES_RULES_TEMPLATE.md`, `.agent/rules/jules-protocol.md`)**: Stripped trailing double newlines (F13).
|
|
28
|
+
|
|
29
|
+
### Added
|
|
30
|
+
- **Canonical Root Test Guard (`src/test-paths.mjs`, `src/security.mjs`)**: Root test files like `test.js` are now classified as canonical test files under `BUILTIN_PROTECT` and anti-tamper auditing (F01).
|
|
31
|
+
- **Conditional Assertion Guard (`src/security.mjs`)**: Ternary and conditional wrapper expressions cannot mask broken logic by rewriting assertions into conditional skips (F03).
|
|
32
|
+
- **De-registration & Skip Detection (`src/security.mjs`)**: Detects removed `#[test]` attributes in Rust, Go build tags, xfail decorators in Python, and body-first early returns (F04).
|
|
33
|
+
- **Reachable Preconditions Guard (`src/security.mjs`)**: Catches impossible preconditions like `if len(x) < 0:` placed before assertions (F05).
|
|
34
|
+
- **Zero-Coverage Added Module Detection (`src/coverage.mjs`)**: Brand-new code modules that receive zero test execution fail coverage thresholds rather than being ignored (F12).
|
|
35
|
+
- **Complete Uninstall & Undo-Init Documentation (`README.md`)**: Added full removal documentation and clean commands (`git rm -rf --ignore-unmatch ... && rm -rf .agent .agentctl`) preserving pre-existing user files (F22).
|
|
36
|
+
- **Cold-Start Trial Regression Suites**: Added 52 new regression tests across `test/cold-start-trial-f01-f12.test.mjs`, `test/cold-start-trial-f06-f11.test.mjs`, and `test/cold-start-trial-f13-f22.test.mjs`.
|
|
37
|
+
|
|
38
|
+
## [0.71.0] - 2026-09-05
|
|
39
|
+
*A blanket is not a check, and silence is not a suite.*
|
|
40
|
+
|
|
41
|
+
A fourth cold-start trial, against v0.70.0 on four repositories nobody here chose. Seven findings, and this is the first trial where every one of them held: no quoted output was fabricated, and all four repository SHAs matched. Two of them prove entries in this file wrong, which is recorded below rather than quietly amended.
|
|
42
|
+
|
|
43
|
+
The two worst were the two halves of the same sentence — a gate that approves against nothing, and a gate that rejects what is fine.
|
|
44
|
+
|
|
45
|
+
### Fixed
|
|
46
|
+
- **A Command That Printed Nothing Was An Oracle (`src/ops/test-collection.mjs`, `src/wizard-init.mjs`)**: `pnpm -r test` on a workspace whose packages declare no test script exits 0 and writes **zero bytes** on stdout and stderr. Measured on `unjs/unimport`: `init` chose it, and `agentctl check` returned `APPROVED (Exit 0)` with an advisory that it could not tell a full suite from a command that ran nothing. The collection floor is one-sided on purpose — an unrecognised runner states no count and passes, because hard-redding every runner not on the list would be worse than the hole it closes — but an unrecognised runner still *prints*. Zero bytes is not a dialect the list has yet to learn; it is a command that ran nothing, and it is now the one absence that decides. v0.69.0 claimed `init` already rejected this candidate: it did not, because `probeVerdict` sorted no-output into `silent` (the forgiving bucket) rather than `empty`, so the exact command that verdict was introduced to reject was the one case it waved through. The rule reads the command as well as the output, and both halves are required: `tsc --noEmit`, `node --check`, `go vet` and `python3 -m compileall` also exit 0 in silence, and two of those are commands this kit writes itself for a repository with no suite yet — keying on silence alone hard-redded every one of them, which is the same first-run rejection of correct code the floor is otherwise careful to avoid. So a command that is *recognisably* a suite invocation and printed nothing ran no suite, while a checker that printed nothing did what it promised. Resolved once, in `producedNoOutput` and `looksLikeTestSuiteCommand`, and consumed by both the floor and the probe.
|
|
47
|
+
- **Python Was Verified Against A Different Copy Of Itself (`src/stack-detector.mjs`)**: `python3 -m pytest` puts the *working directory* on `sys.path`, which is the fix for a module at the root and no help at all when the package lives under `src/` — there, `import iniconfig` finds nothing locally and falls through to site-packages. Measured on `pytest-dev/iniconfig` with `_parse.py` gutted to `return False`: 49 tests passed and the gate returned `APPROVED (Exit 0)`. The suite really did run, so no amount of counting collected tests can see it; only the import path can. A `src/` layout is now detected by the `__init__.py` under it and the oracle is written as `PYTHONPATH=src python3 -m pytest`, which puts the working tree first whether or not the package is also installed.
|
|
48
|
+
- **Ordinary Work On A Test File Was A CRITICAL Block (`src/security.mjs`)**: `assertionsSeen === 0 && (unreadable.length > 0 || examined > 0)` decided `UNREADABLE`, and that `|| examined > 0` threw away the distinction the whole apparatus exists to draw. `ASSERTION_SHAPED` and `unreadable[]` were built to separate *"assertion-shaped lines were present and none parsed"* from *"there were no assertions in these lines at all"*. Collapsed together, **any** changed substantive line in a test file with no recognised assertion became a `CRITICAL` rejection: renaming a test function did it, and so did adding `import os`. The tell was in the finding itself — `file: null`, `line: null`, no sample, because `unreadable` was empty — so the guard blocked while holding no evidence of anything, and advised a **pytest** repository that its assertion library might be unsupported, from a list that names pytest. Worse, the remedy it printed was `verify.tamperGuard: "warn"`, which switches the real checks off along with the blanket. `UNREADABLE` now requires its evidence. Nothing is weakened: a removed or rewritten assertion is a recognised assertion line, raises `assertionsSeen`, and never reached this branch.
|
|
49
|
+
- **The Scaffold Failed The Repository's Own Linter (`src/config.mjs`, `src/wizard-init.mjs`, `src/stack-detector.mjs`)**: `init` wrapped every generated string in double quotes — valid YAML, idiomatic nowhere. Measured on `unjs/unimport`: following the README verbatim produced **40 eslint errors** (`yaml/quotes`, `yaml/plain-scalar`) across `.agent/config.yml` and `.agent/jules.yml`, so a newcomer's first `agentctl check` was a hard red on files `agentctl init` had written a minute earlier, in a repository that was green before. Two causes, and the second was the better one. `yamlScalar` now emits a plain scalar where YAML permits one and single quotes where it does not, which is simply how YAML is written — a glob beginning with `*` stays quoted, because a leading star is an alias reference. That took it from 40 errors to 16. The remaining 16 came from a single generated comment beginning with the word `global`, which ESLint reads as the `/* global a, b */` directive in any language it has a parser for, YAML included — so it declared globals named after each English word of the sentence and then reported them as unused. The sentence leads with the key it explains instead. Nothing in either fix is specific to eslint or to that repository, and the emitter rule lives in one place that both writers use. Measured again on the same clone: 40 errors to 0.
|
|
50
|
+
- **A Test-Integrity Failure Told The Operator To Rotate Their Keys (`bin/agentctl.mjs`)**: `TEST_DIALECT_UNREADABLE` was missing from `INTEGRITY_TYPES`, so the one finding that is most obviously not a secret fell through to the exit-6 secret branch. The screen printed `Phase [SECRETS (test integrity — no secret found)] : ❌ FAIL` and then `💡 Remediation Hint (Exit 6 Secret Leak Prevented) • High-entropy credential or secret detected in patch • rotate any exposed keys immediately` — two statements about the same run, contradicting each other in the same output. v0.67.0 fixed the phase label and left the hint; v0.69.0 recorded the hint as already correct without measuring it. It has its own branch now, and it is derived from `onlyIntegrityFindings` rather than tested independently, so a run carrying both an unreadable file and a real tampering finding gets the tampering advice instead of falling past both into the secret branch — which is the shape of the original hole.
|
|
51
|
+
- **`task create -p` Blocked Forever In A Real Terminal (`src/wizard-task.mjs`)**: the README advertises `task create -p "..."` as the way to *"skip straight to review"*. Interactivity was decided by `isTTY` alone and `-p` was consulted only by the TODO-import branch, so in a terminal the advertised quickstart stopped at `? Task Title` and waited for a keypress that never came — then asked for the instructions it had already been handed. v0.69.0 recorded this as not reproduced, correctly, from a non-TTY run: the headless path never asks, so the defect was invisible exactly where it was measured. `-p` now means the headless path everywhere, and the regression test supplies no answers at all, because only an assertion that nothing is asked could have caught a hang.
|
|
52
|
+
|
|
53
|
+
### Added
|
|
54
|
+
- **A Test Renamed Out Of Its Runner's Discovery Is Its Own Finding (`src/security.mjs`)**: pytest collects `test_*` and Go collects `Test*`, so `def test_totals` becoming `def totals` deletes the test from every future run while leaving it in the file, fully written, assertions intact — and every count in the guard stays level. This was previously caught only as a side effect of the blanket removed above. `TEST_DEREGISTERED` names it, explains that the runner collects by name, and is waivable as `--allow-test-change deregistration` like every other kind. Exact on the remainder: `test_totals` → `totals` is a de-registration, `test_totals` → `test_totals_rounded` is a rename and stays silent.
|
|
55
|
+
- **A Loosened Run No Longer Looks Like A Strict One (`bin/agentctl.mjs`)**: every escape hatch turns a check off and the report said so nowhere — `agentctl check --allow-test-change expectation` over a rewritten expectation printed a phase list byte-identical to a clean, unloosened `APPROVED`. A reviewer, a CI log, or the author a week later could not tell that a check had been waived. An `⚠️ OVERRIDES ACTIVE` block now names each active flag, plus `verify.required: false` and `verify.tamperGuard: "warn"`, and it prints on approval as well as rejection, because approval is the case where it is load-bearing.
|
|
56
|
+
- **Eleven More Cases In The Policy Contract (`src/guard-policy.mjs`)**: five innocent edits for the work that was being rejected — adding an import, a multi-line test rename, a docstring, a type annotation, a Go fixture — four de-registration canaries across both name-driven runners, six silent-run canaries and five silent *static gates* that must keep passing — the counterweight, without which the silence rule would hard-red every honest checker. 42 canaries, 21 innocent edits, 4 de-registrations, 6 silent runs, 5 static gates.
|
|
57
|
+
- **The Innocent-Edit Contract Measures The Verdict (`test/assertion-dialects.test.mjs`, `scripts/guard-reach-check.mjs`)**: it asserted only that no *violation* was produced, and `UNREADABLE` is a status rather than a violation — so the whole list stayed green for two releases while `scanDiff` returned `CRITICAL` and the gate rejected the change. Every innocent edit is now also asserted through `scanDiff`. Anything that can end a run has to be measured where the run ends.
|
|
58
|
+
|
|
59
|
+
### Also Fixed
|
|
60
|
+
- **The Loosening Flags Reached One Phase And Not The Next (`src/engine.mjs`)**: `scanDiff` is given `allowTestChanges` and honours it; the `assert:test-integrity` verify stage received only its own stage object and re-ran the same guard with none of them. So an override was accepted by one phase and ignored by the next — `--allow-test-change deregistration` turned `SECRETS` green and then failed the run at anti-tamper, having already told the operator to pass the flag they had just passed. Found while testing the new kind, not reported by the trial. This is the eighth time in this project's history that one rule has been written in one place while another kept the old answer.
|
|
61
|
+
|
|
62
|
+
- **The Manifest Parser Truncated Any Value Containing A Hash (`src/config.mjs`)**: `parseYaml` opened a comment at the first `#` on a line, quoted or not — so `verify.test: 'pytest -k "not #slow"'` was read as `pytest -k "not`, and the gate ran a command the operator never wrote while reporting on it as if they had. A URL fragment did the same. Comment-stripping is now quote-aware and follows the YAML rule that a `#` opens a comment only at the start of a line or after whitespace, and `coerce` undoes the `''` escape that `yamlScalar` emits. Pre-existing, and it mattered more the moment this release added an emitter: a writer and a reader in one module that disagree about the same file is this project's recurring defect with both halves in view. `YAML_ROUNDTRIP_CASES` pins 19 values — hashes, colons, leading stars, apostrophes, the empty string, and the words YAML would otherwise read as booleans — through both a key and a list item.
|
|
63
|
+
|
|
64
|
+
### Corrections To Earlier Entries
|
|
65
|
+
Two `### Not Reproduced` claims in **[0.69.0]** were wrong, and both were wrong in the same way — measured somewhere the defect could not appear.
|
|
66
|
+
|
|
67
|
+
- *"The `TEST_DIALECT_UNREADABLE` remediation copy was already correct."* It was not. v0.67.0 had fixed the phase label, not the hint, and the two were on screen together. Reproduced verbatim on `pytest-dev/iniconfig` and fixed above.
|
|
68
|
+
- *"`task create -p` did not block on a prompt."* True without a TTY and false with one. Reproduced under `script(1)` and fixed above.
|
|
69
|
+
|
|
70
|
+
One finding in the fourth trial quoted output that does not reproduce: its `--allow-test-change expectation` case shows `APPROVED (Exit 0)`, where the run rejects at `Exit 4` because the reporter's own `sed` genuinely breaks the dotenv suite — the same fixture artefact this project has hit in every trial. The finding itself, that no override leaves a trace in the report, held and is fixed above.
|
|
71
|
+
|
|
8
72
|
## [0.70.0] - 2026-09-05
|
|
9
73
|
*A session that has not finished is not a session that passed.*
|
|
10
74
|
|
|
@@ -45,6 +109,8 @@ A third cold-start trial against v0.68.0. Twelve findings; six reproduced, and t
|
|
|
45
109
|
- **The Egress Guard Could Not See `.js` (`test/egress-allowlist.test.mjs`)**: the scan collected `.mjs` only, so `bin/init.js` — published as the `jules-init` binary — sat outside the boundary entirely. Harmless as it stands, but the guard's whole purpose is that a reviewer can trust the boundary without reading every commit.
|
|
46
110
|
|
|
47
111
|
### Not Reproduced
|
|
112
|
+
> **Corrected in [0.71.0]:** two of the claims below were wrong — the `TEST_DIALECT_UNREADABLE` remediation copy and `task create -p`. Both were measured where the defect could not appear. See the corrections section of that release.
|
|
113
|
+
|
|
48
114
|
Six of the twelve did not hold, and five quoted terminal output that does not exist anywhere in the shipped code. `--allow-test-change expectation` propagates correctly through both phases and returns `APPROVED`; `--allow-test-change removal` allows deleting a dead test alongside its dead function, also `APPROVED` — both reported failures were a test suite genuinely broken by the reporter's own edit, which is the same fixture artefact this project has hit repeatedly. `agentctl plan approve --dry-run` errors with *"Session ID is required"* rather than printing the quoted *"Plan Approved Successfully!"*. `task create -p` did not block on a prompt. A protected `package.json` is the design, and the gate already prints `To allow protected files in this run, pass: agentctl gate --allow-protected` — the finding stated there were no flag hints in the output. The `TEST_DIALECT_UNREADABLE` remediation copy was already correct: *"Exit 6 Test Integrity Violation — no secret was found"*, not the quoted *"Secret Leak Prevented"*. And `pnpm -r test` exits 0 on pnpm 10.33.4 rather than the reported `ERR_PNPM_RECURSIVE_RUN_NO_SCRIPT` — the real defect there was worse than the one reported, and is fixed above.
|
|
49
115
|
|
|
50
116
|
### Added
|
package/JULES_RULES_TEMPLATE.md
CHANGED
|
@@ -124,4 +124,3 @@ HARD CONSTRAINTS:
|
|
|
124
124
|
- BEFORE opening the PR: Run `git fetch origin <base> && git rebase origin/<base>`, then re-verify. If the rebase leaves an empty diff, the work already landed — do NOT submit.
|
|
125
125
|
- Remove any scratch files you created for debugging before submitting. Do not delete files that are part of the project.
|
|
126
126
|
```
|
|
127
|
-
|
package/README.md
CHANGED
|
@@ -208,7 +208,7 @@ To maximize PR merge rates, dispatch tasks according to deterministic boundaries
|
|
|
208
208
|
* **Fail-Closed Security & Secret Redaction:** Evaluates explicit Deny rules before Allow rules against canonicalized, case-folded paths. Redacts high-entropy keys and base64-encoded credentials (such as Kubernetes `Secret` manifests).
|
|
209
209
|
* **Complexity & Cost Router:** Zero-dependency heuristic classifier (`src/router.mjs`) routing mechanical tasks to lightweight models while reserving primary models for complex refactors, with a `node --check` syntax-verification gate that transparently escalates a FAST-tier result to the primary provider if it left broken JS on disk.
|
|
210
210
|
* **Terminal UI & Diagnostic Matrix (`agentctl doctor`):** Interactive terminal dashboard, task sidecar manager, and automated transactional self-repair.
|
|
211
|
-
* **Verified Test Suite:** Tested with **
|
|
211
|
+
* **Verified Test Suite:** Tested with **1399 unit tests across 195 suites**, green on every supported platform.
|
|
212
212
|
|
|
213
213
|
<br/>
|
|
214
214
|
|
|
@@ -448,6 +448,8 @@ const result = await fast.dispatch({ prompt: "Fix a typo." }, { root: process.cw
|
|
|
448
448
|
|
|
449
449
|
| Feature | Module / Command | Architectural Description | Status |
|
|
450
450
|
| :--- | :--- | :--- | :---: |
|
|
451
|
+
| **Cold-Start Hardened Kernel & Tamper Defense** | `src/config.mjs`, `src/engine.mjs`, `src/git.mjs`, `src/security.mjs` | Full remediation of 22 cold-start audit findings (F01–F22): authoritative base policy resolution, ephemeral snapshot worktree isolation, canonical root test tamper guard, conditional assertion defense, multi-target Cargo test aggregation, Python src-layout injection, and complete repository uninstall documentation. | **v0.72.0** *(Shipped)* |
|
|
452
|
+
| **Silence Is Not A Suite & Scaffolding Linter Fixes** | `src/ops/test-collection.mjs`, `src/wizard-init.mjs`, `src/config.mjs` | Reject zero-output test suite commands, quote-aware YAML parser with scalar emission, test de-registration detection (`TEST_DEREGISTERED`), and active waiver telemetry banner. | **v0.71.0** *(Shipped)* |
|
|
451
453
|
| **Diagnostics That Reach the Operator** | `src/security.mjs`, `src/engine.mjs`, `bin/agentctl.mjs` | Secret findings name the file and line, a failed verify stage reports its command, exit code and output, and `queue`/`swarm` name each failed task and exit `1` rather than reporting success for a run that dispatched nothing. | **v0.41.1** *(Shipped)* |
|
|
452
454
|
| **One Scaffolding Path & First-Install Fixes** | `src/scaffold.mjs`, `src/security.mjs` | `agentctl init` and `jules-init` scaffold from one source and write the runtime `.gitignore` entries, so the kit's own bookkeeping no longer reaches its own gate; a lockfile bump no longer fails closed as a secret leak. | **v0.41.1** *(Shipped)* |
|
|
453
455
|
| **Queue Runner Fidelity** | `src/dag-engine.mjs`, `src/engine.mjs` | Queue selection is by task shape rather than file extension, so manifests and READMEs are skipped instead of dispatched, and `--dry-run` leaves the queue untouched. | **v0.38.2** *(Shipped)* |
|
|
@@ -473,6 +475,59 @@ const result = await fast.dispatch({ prompt: "Fix a typo." }, { root: process.cw
|
|
|
473
475
|
|
|
474
476
|
<br/>
|
|
475
477
|
|
|
478
|
+
## 🧹 Complete Uninstall / Removing the Kit (Undo Init)
|
|
479
|
+
|
|
480
|
+
If you need to completely remove `jules-orchestrator-kit` from a repository after running `agentctl init`, follow the procedure below. Note that `agentctl clean` is an operational maintenance command (cleaning ephemeral locks, temporary worktrees, and evidence caches), not an uninstaller.
|
|
481
|
+
|
|
482
|
+
### 1. Generated Assets & Manifest
|
|
483
|
+
|
|
484
|
+
`agentctl init` / `scaffoldRepoAssets()` writes the following project files and directories:
|
|
485
|
+
- **Core configuration and rules:** `.agent/config.yml` (or `.agent/jules.yml`), `.agent/rules/`, `.agent/prompts/`, `.agent/workflows/`, and `AGENTS.md`.
|
|
486
|
+
- **System contracts:** `SPEC.md`, `CONSTRAINTS.md` (and optional `DESIGN.md`).
|
|
487
|
+
- **Queue runtime stub:** `.agent/jules-queue/README.md`.
|
|
488
|
+
- **Optional IDE & CI integrations:** `.github/workflows/agent-gate.yml`, `.gitlab-ci.agent-gate.yml`, and `.cursor/rules/jules.mdc`.
|
|
489
|
+
|
|
490
|
+
### 2. Runtime State & Working Trees
|
|
491
|
+
|
|
492
|
+
During execution, the kit produces untracked runtime artifacts in:
|
|
493
|
+
- `.agent/evidence/` — Cryptographic evidence manifests and stage run recordings.
|
|
494
|
+
- `.agent/state/` — Flaky test ledgers, budget trackers, and escalation queues.
|
|
495
|
+
- `.agent/worktrees/` — Isolated snapshot worktrees used by the verification sandbox.
|
|
496
|
+
- `.agent/history/` and `.agent/handovers/` — Local agent session memories.
|
|
497
|
+
|
|
498
|
+
### 3. Removal Procedure (Preserving Pre-Existing User Files)
|
|
499
|
+
|
|
500
|
+
To completely undo `init` and restore your working tree to its exact original state:
|
|
501
|
+
|
|
502
|
+
```bash
|
|
503
|
+
# 1. Remove tracked orchestrator assets (skips any files that were not scaffolded)
|
|
504
|
+
git rm -rf --ignore-unmatch \
|
|
505
|
+
.agent \
|
|
506
|
+
AGENTS.md \
|
|
507
|
+
SPEC.md \
|
|
508
|
+
CONSTRAINTS.md \
|
|
509
|
+
DESIGN.md \
|
|
510
|
+
.github/workflows/agent-gate.yml \
|
|
511
|
+
.gitlab-ci.agent-gate.yml \
|
|
512
|
+
.cursor/rules/jules.mdc
|
|
513
|
+
|
|
514
|
+
# 2. Remove untracked runtime directories and temporary caches
|
|
515
|
+
rm -rf .agent .agentctl
|
|
516
|
+
|
|
517
|
+
# 3. Clean up .gitignore additions
|
|
518
|
+
# Revert the appended "# Jules Orchestrator runtime state & credentials" block from .gitignore
|
|
519
|
+
git checkout .gitignore # If .gitignore had no other unstaged changes, or edit by hand
|
|
520
|
+
|
|
521
|
+
# 4. Optional: Uninstall global CLI package
|
|
522
|
+
npm uninstall -g jules-orchestrator-kit
|
|
523
|
+
```
|
|
524
|
+
|
|
525
|
+
<br/>
|
|
526
|
+
|
|
527
|
+
---
|
|
528
|
+
|
|
529
|
+
<br/>
|
|
530
|
+
|
|
476
531
|
## 📖 Documentation & External References
|
|
477
532
|
|
|
478
533
|
- [**System Architecture & Pipeline Overview**](./docs/architecture.md) — Comprehensive technical sequence diagrams and control plane specifications.
|
package/ROADMAP_V1.md
CHANGED
|
@@ -11,14 +11,39 @@ The **jules-orchestrator-kit** is the zero-dependency safety gatekeeper and self
|
|
|
11
11
|
## 📌 Release Milestones Overview
|
|
12
12
|
|
|
13
13
|
```
|
|
14
|
-
v0.
|
|
15
|
-
(
|
|
14
|
+
v0.72.0 (Current Stable) ──► v0.73.0 (Distributed Swarms & Leases) ──► v1.0.0 (Production Hardened Kernel)
|
|
15
|
+
(Cold-Start Hardened Kernel) (Multi-Agent DAG & Resource Locks) (Enterprise Telemetry & SLA)
|
|
16
16
|
```
|
|
17
17
|
|
|
18
18
|
---
|
|
19
19
|
|
|
20
|
-
## ✅ Shipped Milestones (v0.20.0 – v0.
|
|
21
|
-
|
|
20
|
+
## ✅ Shipped Milestones (v0.20.0 – v0.72.0)
|
|
21
|
+
|
|
22
|
+
|
|
23
|
+
### v0.72.0: Cold-Start Hardened Kernel & Tamper Defense
|
|
24
|
+
- [x] **Canonical Root Test Guard (F01)** — `test.js` at repository root is inside the tamper guard.
|
|
25
|
+
- [x] **Conditional Expectation Guard (F03)** — ternary and conditional assertions cannot mask broken logic.
|
|
26
|
+
- [x] **De-registration & Skip Detection (F04)** — removed `#[test]`, build tags, xfail decorators, and body-first early returns are caught.
|
|
27
|
+
- [x] **Reachable Preconditions (F05)** — impossible guard conditions (`len < 0`) cannot neutralise assertions.
|
|
28
|
+
- [x] **Uncommitted Scaffold Integrity (F06)** — rejects uncommitted scaffolds that disable verification or lower profiles.
|
|
29
|
+
- [x] **Trusted Base Policy Resolution (F07)** — authoritative verification stages resolved from base commit (`git show <base>:.agent/config.yml`), never trusting uncommitted edits under review.
|
|
30
|
+
- [x] **Committed Base Branch Integrity (F08)** — rejects `--base HEAD` in committed mode.
|
|
31
|
+
- [x] **Empty Test Collection Canaries (F09)** — empty collections (Go `[no tests to run]`, pytest `--collect-only`, no-op scripts) recognized as 0 tests.
|
|
32
|
+
- [x] **Snapshot Worktree Isolation (F10)** — isolates staged index and committed revisions in ephemeral worktrees with symlinked dependencies.
|
|
33
|
+
- [x] **Python src-layout Invariant (F11)** — automatically injects `PYTHONPATH=src` for package layouts.
|
|
34
|
+
- [x] **Zero-Coverage Added Module Detection (F12)** — untracked/unexecuted new files fail coverage.
|
|
35
|
+
- [x] **Clean Scaffold & Oracle Tuning (F13–F15)** — markdown newline hygiene, cargo clippy without `-D warnings`, and multi-target Cargo aggregation.
|
|
36
|
+
- [x] **Supply Chain Diagnostics & Waiver Telemetry (F16–F20)** — lockfile tamper hints, `--strict-locks` flag, waiver auditing, and dry-run evidence suppression.
|
|
37
|
+
- [x] **Targeted Help & Complete Uninstall (F21–F22)** — targeted subcommand help routing and documented full removal procedure.
|
|
38
|
+
|
|
39
|
+
### v0.71.0: A Blanket Is Not A Check
|
|
40
|
+
- [x] **Silence Is Not A Suite (`src/ops/test-collection.mjs`, `src/wizard-init.mjs`)** — a command that claims to run tests and prints nothing ran none; a static gate that prints nothing did its job.
|
|
41
|
+
- [x] **Python Verified Against The Working Tree (`src/stack-detector.mjs`)** — a `src/` layout resolved its imports to site-packages, and broken code passed 49 tests.
|
|
42
|
+
- [x] **`UNREADABLE` Requires Its Evidence (`src/security.mjs`)** — adding an import to a test file was a CRITICAL block.
|
|
43
|
+
- [x] **A Test Renamed Out Of Discovery Is Its Own Finding (`src/security.mjs`)** — `TEST_DEREGISTERED`, waivable like every other kind.
|
|
44
|
+
- [x] **The Scaffold Passes A Repository's Own YAML Linter (`src/config.mjs`)** — 40 eslint errors on files `init` had just written.
|
|
45
|
+
- [x] **A Loosened Run Says So (`bin/agentctl.mjs`)** — an override left no trace in the report at all.
|
|
46
|
+
- [x] **`task create -p` Skips The Questions (`src/wizard-task.mjs`)** — the advertised quickstart blocked forever in a real terminal.
|
|
22
47
|
|
|
23
48
|
### v0.70.0: Not Finished Is Not Passed
|
|
24
49
|
- [x] **An Unfinished Session Is Not COMPLETED (`src/engine.mjs`)** — terminal, blocked and timed-out are three verdicts, not one.
|
|
@@ -181,9 +206,9 @@ The **jules-orchestrator-kit** is the zero-dependency safety gatekeeper and self
|
|
|
181
206
|
|
|
182
207
|
---
|
|
183
208
|
|
|
184
|
-
## 🎯 Target Milestones (v0.
|
|
209
|
+
## 🎯 Target Milestones (v0.73.0 & v1.0.0)
|
|
185
210
|
|
|
186
|
-
### v0.
|
|
211
|
+
### v0.73.0: Distributed File Leases & Preemptive DAG Scheduling
|
|
187
212
|
- [ ] **Atomic Filesystem Lease & Heartbeat Protocol (`src/engine.mjs`, `src/flaky-ledger.mjs`)** — Directory-mutex file leasing with heartbeat timestamps, stale-lock detection via PID liveness inspection, and tombstone rotation without third-party daemons or Redis.
|
|
188
213
|
- [ ] **Preemptive Task Cancellation & Interface Fingerprints (`src/dag-engine.mjs`)** — Automatically aborts and yields downstream swarm tasks when upstream exported symbol interfaces diverge from their cryptographic SHA-256 fingerprints.
|
|
189
214
|
- [ ] **POSIX/Win32 Process Group Guillotine (`src/engine.mjs`)** — Tree teardown via `process.kill(-pid, 'SIGKILL')` on POSIX and `taskkill /T /F /PID` on Windows to eliminate orphaned dev-servers and background watchers.
|
package/bin/agentctl.mjs
CHANGED
|
@@ -211,6 +211,23 @@ async function main() {
|
|
|
211
211
|
process.exit(0);
|
|
212
212
|
}
|
|
213
213
|
|
|
214
|
+
if (command === "help") {
|
|
215
|
+
const target = args[1];
|
|
216
|
+
if (!target || target === "--help" || target === "-h") {
|
|
217
|
+
printHelp();
|
|
218
|
+
process.exit(0);
|
|
219
|
+
}
|
|
220
|
+
const { getCommandDescriptor, formatCommandHelp } = await import("../src/ops/command-registry.mjs");
|
|
221
|
+
const subSub = args[2] && !args[2].startsWith("-") ? `${target} ${args[2]}` : target;
|
|
222
|
+
const desc = getCommandDescriptor(subSub) || getCommandDescriptor(target);
|
|
223
|
+
if (desc) {
|
|
224
|
+
console.log(formatCommandHelp(desc));
|
|
225
|
+
process.exit(0);
|
|
226
|
+
}
|
|
227
|
+
console.log(`\nUsage: agentctl ${args.slice(1).join(" ")} [options]\n\nFor general help, run: agentctl --help\n`);
|
|
228
|
+
process.exit(0);
|
|
229
|
+
}
|
|
230
|
+
|
|
214
231
|
// Bare `agentctl` answers "what do I do next" rather than dumping thirty
|
|
215
232
|
// commands. The help text is a reference for people who already know the
|
|
216
233
|
// tool; a newcomer cannot tell which entry is step one, and guessing wrong
|
|
@@ -248,7 +265,7 @@ async function main() {
|
|
|
248
265
|
console.log(formatCommandHelp(desc));
|
|
249
266
|
process.exit(0);
|
|
250
267
|
}
|
|
251
|
-
|
|
268
|
+
console.log(`\nUsage: agentctl ${command}${subArgs[0] && !subArgs[0].startsWith("-") ? " " + subArgs[0] : ""} [options]\n\nFor general help, run: agentctl --help\n`);
|
|
252
269
|
process.exit(0);
|
|
253
270
|
}
|
|
254
271
|
|
|
@@ -376,7 +393,7 @@ async function main() {
|
|
|
376
393
|
const { values } = parseArgs({
|
|
377
394
|
args: args.slice(1),
|
|
378
395
|
options: {
|
|
379
|
-
base: { type: "string", short: "b"
|
|
396
|
+
base: { type: "string", short: "b" },
|
|
380
397
|
mode: { type: "string", short: "m", default: "working-tree" },
|
|
381
398
|
"working-tree": { type: "boolean" },
|
|
382
399
|
staged: { type: "boolean" },
|
|
@@ -384,6 +401,7 @@ async function main() {
|
|
|
384
401
|
fix: { type: "boolean" },
|
|
385
402
|
"allow-protected": { type: "boolean" },
|
|
386
403
|
"allow-unreadable-tests": { type: "boolean" },
|
|
404
|
+
"strict-locks": { type: "boolean" },
|
|
387
405
|
// The tamper guard has always had an override — `allowTestModifications`
|
|
388
406
|
// — and it was reachable only from JavaScript. So a legitimate change
|
|
389
407
|
// of spec, which necessarily rewrites what a test expects, hit a
|
|
@@ -417,14 +435,54 @@ async function main() {
|
|
|
417
435
|
allowUnreadableTests: values["allow-unreadable-tests"],
|
|
418
436
|
allowTestModifications: values["allow-test-modifications"],
|
|
419
437
|
allowTestChanges: values["allow-test-change"],
|
|
438
|
+
strictLocks: values["strict-locks"],
|
|
439
|
+
dryRun: values["dry-run"],
|
|
420
440
|
jsonReport: values["json-report"],
|
|
421
441
|
});
|
|
422
442
|
|
|
443
|
+
const overrides = [];
|
|
444
|
+
const julesCmdAllow = process.env.JULES_ALLOW_COMMAND_FILE_CHANGES === "true" || process.env.JULES_ALLOW_COMMAND_FILE_CHANGES === "1";
|
|
445
|
+
const agentCmdAllow = process.env.AGENT_ALLOW_COMMAND_FILE_CHANGES === "true" || process.env.AGENT_ALLOW_COMMAND_FILE_CHANGES === "1";
|
|
446
|
+
if (values["allow-protected"] || julesCmdAllow || agentCmdAllow) {
|
|
447
|
+
const via = values["allow-protected"] ? "--allow-protected" : (julesCmdAllow ? "JULES_ALLOW_COMMAND_FILE_CHANGES" : "AGENT_ALLOW_COMMAND_FILE_CHANGES");
|
|
448
|
+
overrides.push(`${via} (protected paths permitted)`);
|
|
449
|
+
}
|
|
450
|
+
if (values["allow-unreadable-tests"]) overrides.push("--allow-unreadable-tests (unreadable dialect permitted)");
|
|
451
|
+
if (values["allow-test-modifications"]) overrides.push("--allow-test-modifications (every tamper check waived)");
|
|
452
|
+
if (values["allow-test-change"]) {
|
|
453
|
+
const kinds = [].concat(values["allow-test-change"]).join(", ");
|
|
454
|
+
overrides.push(`--allow-test-change ${kinds} (tamper check waived: ${kinds})`);
|
|
455
|
+
}
|
|
456
|
+
if (config?.verify?.required === false) overrides.push("verify.required: false (verification failures and empty test suites permitted)");
|
|
457
|
+
if (config?.verify?.minTests === 0 || config?.verify?.min_tests === 0 || config?.minTests === 0 || config?.min_tests === 0) overrides.push("verify.minTests: 0 (test collection floor disabled)");
|
|
458
|
+
if (config?.verify?.tamperGuard === "warn") overrides.push('verify.tamperGuard: "warn" (unreadable dialects report only)');
|
|
459
|
+
const optionalFails = (res.phases?.find((p) => p.phase === "verify")?.executionRecords || []).filter((r) => r.required === false && !r.ok);
|
|
460
|
+
for (const r of optionalFails) {
|
|
461
|
+
overrides.push(`stage.${r.id || r.name}: required: false (failing optional stage permitted)`);
|
|
462
|
+
}
|
|
463
|
+
|
|
423
464
|
if (values.json) {
|
|
465
|
+
res.overrides = overrides;
|
|
424
466
|
console.log(JSON.stringify(res, null, 2));
|
|
425
467
|
} else {
|
|
426
|
-
|
|
468
|
+
const reportedBase = res.base || values.base || "main";
|
|
469
|
+
console.log(`\n🛡️ agentctl Safety Gate Audit Results (Base: ${reportedBase}, Mode: ${selectedMode})`);
|
|
427
470
|
console.log(`-----------------------------------------------------`);
|
|
471
|
+
// A loosened run must not be able to pass for a strict one.
|
|
472
|
+
//
|
|
473
|
+
// Every escape hatch here turns a check off, and the report said so
|
|
474
|
+
// nowhere: `agentctl check --allow-test-change expectation` over a
|
|
475
|
+
// rewritten expectation printed a phase list byte-identical to a clean,
|
|
476
|
+
// unloosened APPROVED. Anyone reading the output afterwards — a
|
|
477
|
+
// reviewer, a CI log, the author a week later — could not tell that a
|
|
478
|
+
// check had been waived, which makes the waiver invisible exactly where
|
|
479
|
+
// it matters most. Printed before the phases, and on approval as well
|
|
480
|
+
// as rejection, because an approval is the case where it is load-bearing.
|
|
481
|
+
if (overrides.length > 0) {
|
|
482
|
+
console.log(` ⚠️ OVERRIDES ACTIVE — this run is not a strict pass:`);
|
|
483
|
+
for (const o of overrides) console.log(` • ${o}`);
|
|
484
|
+
console.log(`-----------------------------------------------------`);
|
|
485
|
+
}
|
|
428
486
|
for (const p of res.phases) {
|
|
429
487
|
const status = p.ok ? "✅ PASS" : "❌ FAIL";
|
|
430
488
|
// Test tampering shares this phase, and its exit code, with the
|
|
@@ -472,13 +530,32 @@ async function main() {
|
|
|
472
530
|
// arrive under the same phase and the same code.
|
|
473
531
|
const secretsPhase = res.phases.find((p) => p.phase === "secrets" && !p.ok);
|
|
474
532
|
const findingTypes = new Set((secretsPhase?.findings || []).map((f) => f.type));
|
|
533
|
+
// `TEST_DIALECT_UNREADABLE` belongs in this set and was missing from
|
|
534
|
+
// it, so the one finding that is *most* obviously not a secret fell
|
|
535
|
+
// through to the secret hint: the report printed
|
|
536
|
+
// "Phase [SECRETS (test integrity — no secret found)] : ❌ FAIL" and
|
|
537
|
+
// then told the operator a high-entropy credential had leaked and to
|
|
538
|
+
// rotate their keys. v0.67.0 fixed the phase label and left the hint;
|
|
539
|
+
// v0.69.0 recorded the hint as already correct without measuring it.
|
|
540
|
+
// Both statements were about the same screen, and they contradicted
|
|
541
|
+
// each other on it.
|
|
475
542
|
const INTEGRITY_TYPES = new Set([
|
|
476
543
|
"TEST_TAMPERING_DETECTED",
|
|
544
|
+
"TEST_DIALECT_UNREADABLE",
|
|
477
545
|
"EDGE_RUNTIME_VIOLATION",
|
|
478
546
|
"CROSS_PACKAGE_BOUNDARY_VIOLATION",
|
|
479
547
|
]);
|
|
480
548
|
const onlyIntegrityFindings =
|
|
481
549
|
findingTypes.size > 0 && [...findingTypes].every((t) => INTEGRITY_TYPES.has(t));
|
|
550
|
+
// An unreadable dialect is not a weakened assertion either, and the
|
|
551
|
+
// integrity hint tells the operator to restore an assertion that no
|
|
552
|
+
// finding names. Split by what is actually present — and derived from
|
|
553
|
+
// `onlyIntegrityFindings` rather than tested independently, so a run
|
|
554
|
+
// carrying both an unreadable file and a real tampering finding gets
|
|
555
|
+
// the tampering advice instead of falling past both into the secret
|
|
556
|
+
// branch, which is how this hole was shaped the first time.
|
|
557
|
+
const onlyUnreadableDialect =
|
|
558
|
+
onlyIntegrityFindings && [...findingTypes].every((t) => t === "TEST_DIALECT_UNREADABLE");
|
|
482
559
|
|
|
483
560
|
// Exit 3 is also what a strictTestLock tamper verdict returns, so the
|
|
484
561
|
// code alone cannot pick the hint — a scope remediation for a rewritten
|
|
@@ -504,6 +581,16 @@ async function main() {
|
|
|
504
581
|
if (firstError) console.log(` • Last attempt error: ${firstError}`);
|
|
505
582
|
console.log(` • Review error fingerprints via: agentctl doctor\n`);
|
|
506
583
|
}
|
|
584
|
+
} else if (onlyUnreadableDialect) {
|
|
585
|
+
console.log(`💡 Remediation Hint (Exit ${res.code} Unreadable Test Dialect — no secret was found):`);
|
|
586
|
+
console.log(` • Nothing leaked. The guard could not recognise the assertions in a changed test`);
|
|
587
|
+
console.log(` file, so it could not check that file for tampering — and a check that examined`);
|
|
588
|
+
console.log(` nothing does not get to say APPROVED.`);
|
|
589
|
+
console.log(` • Best: report the dialect so it gets covered. The guard knows Node, pytest, Go,`);
|
|
590
|
+
console.log(` Rust, JUnit, RSpec, PHPUnit, Minitest, XCTest, chai and node-tap.`);
|
|
591
|
+
console.log(` • Or say so on the record: verify.tamperGuard: "warn" in .agent/config.yml`);
|
|
592
|
+
console.log(` • Or allow this one run: agentctl check --allow-unreadable-tests`);
|
|
593
|
+
console.log(` • Nothing needs rotating: this exit code is shared with the secret scanner.\n`);
|
|
507
594
|
} else if (onlyIntegrityFindings) {
|
|
508
595
|
console.log(`💡 Remediation Hint (Exit ${res.code} Test Integrity Violation — no secret was found):`);
|
|
509
596
|
console.log(` • The diff weakens or removes verification rather than leaking a credential.`);
|
|
@@ -553,9 +640,15 @@ async function main() {
|
|
|
553
640
|
console.log(` • Commit them once and the gate goes green:`);
|
|
554
641
|
console.log(` git add ${untracked.join(" ")} && git commit -m "chore: add agent config"\n`);
|
|
555
642
|
} else {
|
|
643
|
+
const hasLockfileViolation = scopeFiles.some((f) => /(?:^|\/)(?:package-lock\.json|yarn\.lock|pnpm-lock\.yaml|bun\.lockb?|Cargo\.lock|go\.sum|uv\.lock|poetry\.lock|Pipfile\.lock|composer\.lock)$/.test(f));
|
|
556
644
|
console.log(`💡 Remediation Hint (Exit ${res.code} Scope Violation):`);
|
|
557
|
-
|
|
558
|
-
|
|
645
|
+
if (hasLockfileViolation) {
|
|
646
|
+
console.log(` • Lockfiles are protected against supply-chain tampering. Routine dependency`);
|
|
647
|
+
console.log(` updates are permitted using the maintainer waiver: agentctl gate --allow-protected\n`);
|
|
648
|
+
} else {
|
|
649
|
+
console.log(` • To allow protected files in this run, pass: agentctl gate --allow-protected`);
|
|
650
|
+
console.log(` • Or remove protected/denied paths from the diff before dispatching.\n`);
|
|
651
|
+
}
|
|
559
652
|
}
|
|
560
653
|
} else if (failedPhase === "git_resolution") {
|
|
561
654
|
console.log(`💡 Remediation Hint (Exit ${res.code} Base Branch Unresolvable):`);
|
|
@@ -2198,7 +2291,7 @@ async function main() {
|
|
|
2198
2291
|
if (values.json) {
|
|
2199
2292
|
console.log(JSON.stringify(res, null, 2));
|
|
2200
2293
|
} else {
|
|
2201
|
-
console.log(`\n✅ Plan Approved Successfully!`);
|
|
2294
|
+
console.log(values["dry-run"] ? `\n[DRY-RUN] Plan Approved Successfully (Simulation)!` : `\n✅ Plan Approved Successfully!`);
|
|
2202
2295
|
console.log(` Session ID : ${res.id}`);
|
|
2203
2296
|
console.log(` Status : ${res.status}\n`);
|
|
2204
2297
|
}
|
|
@@ -2235,7 +2328,7 @@ async function main() {
|
|
|
2235
2328
|
if (values.json) {
|
|
2236
2329
|
console.log(JSON.stringify(res, null, 2));
|
|
2237
2330
|
} else {
|
|
2238
|
-
console.log(`\n✅ Plan Approved Successfully!`);
|
|
2331
|
+
console.log(values["dry-run"] ? `\n[DRY-RUN] Plan Approved Successfully (Simulation)!` : `\n✅ Plan Approved Successfully!`);
|
|
2239
2332
|
console.log(` Session ID : ${res.id}`);
|
|
2240
2333
|
console.log(` Status : ${res.status}\n`);
|
|
2241
2334
|
}
|
|
@@ -2271,7 +2364,7 @@ async function main() {
|
|
|
2271
2364
|
if (values.json) {
|
|
2272
2365
|
console.log(JSON.stringify(res, null, 2));
|
|
2273
2366
|
} else {
|
|
2274
|
-
console.log(`\n📋 Remote Session Status:`);
|
|
2367
|
+
console.log(values["dry-run"] ? `\n📋 [DRY-RUN] Remote Session Status (Simulation):` : `\n📋 Remote Session Status:`);
|
|
2275
2368
|
console.log(` Session ID : ${res.id}`);
|
|
2276
2369
|
console.log(` Status : ${res.status}\n`);
|
|
2277
2370
|
}
|
package/package.json
CHANGED
|
@@ -1,6 +1,6 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "jules-orchestrator-kit",
|
|
3
|
-
"version": "0.
|
|
3
|
+
"version": "0.72.0",
|
|
4
4
|
"description": "Zero-dependency safety gatekeeper, test oracle generator, and multi-agent coordination protocol for autonomous coding agents — Google Jules, Claude Code, Codex and Gemini CLI.",
|
|
5
5
|
"repository": {
|
|
6
6
|
"type": "git",
|
|
@@ -42,10 +42,10 @@
|
|
|
42
42
|
* Exit codes: 0 = every guard reachable, 1 = a guard has gone silent.
|
|
43
43
|
*/
|
|
44
44
|
|
|
45
|
-
import { checkTestTampering, checkScope } from "../src/security.mjs";
|
|
45
|
+
import { checkTestTampering, checkScope, scanDiff } from "../src/security.mjs";
|
|
46
46
|
import { isTestPath } from "../src/test-paths.mjs";
|
|
47
47
|
import { normalizeScope } from "../src/config.mjs";
|
|
48
|
-
import { parseCollectedTests } from "../src/ops/test-collection.mjs";
|
|
48
|
+
import { parseCollectedTests, checkCollectionFloor } from "../src/ops/test-collection.mjs";
|
|
49
49
|
import {
|
|
50
50
|
TEST_PATH_CASES,
|
|
51
51
|
TAMPER_CANARIES,
|
|
@@ -53,6 +53,9 @@ import {
|
|
|
53
53
|
PREDICATE_MUTANTS,
|
|
54
54
|
EMPTY_RUN_CANARIES,
|
|
55
55
|
COUNTED_RUN_CANARIES,
|
|
56
|
+
SILENT_RUN_CANARIES,
|
|
57
|
+
SILENT_STATIC_GATES,
|
|
58
|
+
DEREGISTRATION_CANARIES,
|
|
56
59
|
SCOPE_CANARIES,
|
|
57
60
|
INNOCENT_EDITS,
|
|
58
61
|
UNREADABLE_DIALECTS,
|
|
@@ -72,9 +75,14 @@ function canaryDiff(c) {
|
|
|
72
75
|
const lines = [`--- a/${c.file}`, `+++ b/${c.file}`, "@@ -1,20 +1,20 @@", ` ${ctx}`];
|
|
73
76
|
// Unchanged lines the change sits inside. An assertion whose keyword is on
|
|
74
77
|
// one of these is the shape the line-level denominator could not see.
|
|
75
|
-
|
|
76
|
-
|
|
77
|
-
|
|
78
|
+
// Entries may contain newlines (a multi-line fixture); each physical line
|
|
79
|
+
// carries its own diff prefix, exactly as git emits it.
|
|
80
|
+
const each = (arr, prefix) => {
|
|
81
|
+
for (const l of arr) for (const part of String(l).split("\n")) lines.push(`${prefix}${part}`);
|
|
82
|
+
};
|
|
83
|
+
each(c.lead || [], " ");
|
|
84
|
+
each(c.removed, "-");
|
|
85
|
+
each(c.added, "+");
|
|
78
86
|
lines.push(` ${ctx}`);
|
|
79
87
|
return lines.join("\n");
|
|
80
88
|
}
|
|
@@ -135,17 +143,35 @@ const canaryResults = new Map();
|
|
|
135
143
|
const silent = [];
|
|
136
144
|
const noDenominator = [];
|
|
137
145
|
const noAssertions = [];
|
|
146
|
+
// Findings whose own evidence is a comment line: Go build constraints are
|
|
147
|
+
// comments to the compiler, so a constraint-only diff contributes no
|
|
148
|
+
// *examined* code lines — the finding's denominator is the file it sits in,
|
|
149
|
+
// and `filesSeen` carries that. Requiring `inputsSeen` here would ask the
|
|
150
|
+
// counter to count comments, which is exactly what it must skip.
|
|
151
|
+
const COMMENT_LINE_FINDING = /^skip-injection\/go-build/;
|
|
138
152
|
for (const c of [...TAMPER_CANARIES, ...MULTILINE_CANARIES]) {
|
|
139
153
|
const res = checkTestTampering(canaryDiff(c));
|
|
140
154
|
const hit = (res.violations || []).some((v) => v.type === c.expect);
|
|
141
155
|
canaryResults.set(c.id, hit);
|
|
142
156
|
if (!hit) silent.push(`${c.id} expected ${c.expect}, got ${JSON.stringify((res.violations || []).map((v) => v.type))}`);
|
|
143
157
|
// A finding with no denominator is the shape this script exists to reject.
|
|
144
|
-
if (hit && !(res.inputsSeen > 0)) noDenominator.push(c.id);
|
|
158
|
+
if (hit && !(res.inputsSeen > 0) && !COMMENT_LINE_FINDING.test(c.id)) noDenominator.push(c.id);
|
|
159
|
+
if (hit && COMMENT_LINE_FINDING.test(c.id) && !(res.filesSeen > 0)) noDenominator.push(c.id);
|
|
145
160
|
// Counting lines was not enough: a JUnit diff reported one input examined
|
|
146
161
|
// and a clean PASS while every assertion in it went unrecognised. A rule
|
|
147
162
|
// about assertions has to say how many assertions it actually read.
|
|
148
|
-
|
|
163
|
+
//
|
|
164
|
+
// A finding about test *execution* rather than assertion content —
|
|
165
|
+
// skips, xfail/cfg/build-constraint exclusion, de-registration by rename
|
|
166
|
+
// or attribute removal — may legitimately be the only change in the
|
|
167
|
+
// diff, with no assertion line on either side. Asserting an
|
|
168
|
+
// `assertionsSeen` there is a denominator the finding does not have; the
|
|
169
|
+
// assertions a dead-tagged test holds are in context, not in the edit.
|
|
170
|
+
const assertionFinding =
|
|
171
|
+
c.expect !== "TEST_SKIP_INJECTION" &&
|
|
172
|
+
c.expect !== "TEST_DEREGISTERED" &&
|
|
173
|
+
!/^(skip-injection|deregistration)\//.test(c.id);
|
|
174
|
+
if (hit && assertionFinding && !(res.assertionsSeen > 0)) {
|
|
149
175
|
noAssertions.push(`${c.id} (${res.assertionsSeen} assertions parsed)`);
|
|
150
176
|
}
|
|
151
177
|
}
|
|
@@ -154,13 +180,53 @@ const canaryResults = new Map();
|
|
|
154
180
|
add("canaries: assertion rules parsed an assertion", noAssertions.length === 0, noAssertions.length ? noAssertions.join(", ") : "assertionsSeen > 0 on every assertion finding");
|
|
155
181
|
}
|
|
156
182
|
|
|
183
|
+
{
|
|
184
|
+
const missed = DEREGISTRATION_CANARIES.filter(
|
|
185
|
+
(c) => !(checkTestTampering(canaryDiff(c)).violations || []).some((v) => v.type === "TEST_DEREGISTERED")
|
|
186
|
+
);
|
|
187
|
+
add(
|
|
188
|
+
"canaries: a test renamed out of discovery is caught",
|
|
189
|
+
missed.length === 0,
|
|
190
|
+
missed.length ? missed.map((c) => `${c.id} (${c.why})`).join("; ") : `${DEREGISTRATION_CANARIES.length} de-registrations reported`
|
|
191
|
+
);
|
|
192
|
+
}
|
|
193
|
+
|
|
194
|
+
{
|
|
195
|
+
// Zero bytes on both streams is the one absence that decides. Measured
|
|
196
|
+
// through the floor rather than the parser: the parser answers "no count
|
|
197
|
+
// stated", and it is the floor that has to refuse to certify against it.
|
|
198
|
+
const passed = SILENT_RUN_CANARIES.filter(
|
|
199
|
+
(c) => checkCollectionFloor({ ok: true, stdout: c.stdout, stderr: c.stderr, command: c.command }).ok !== false
|
|
200
|
+
);
|
|
201
|
+
add(
|
|
202
|
+
"policy: a command that printed nothing verified nothing",
|
|
203
|
+
passed.length === 0,
|
|
204
|
+
passed.length ? passed.map((c) => c.id).join(", ") : `${SILENT_RUN_CANARIES.length} silent runs refused`
|
|
205
|
+
);
|
|
206
|
+
|
|
207
|
+
// And the counterweight: a checker that succeeds in silence is doing its job.
|
|
208
|
+
const redded = SILENT_STATIC_GATES.filter(
|
|
209
|
+
(c) => checkCollectionFloor({ ok: true, stdout: "", stderr: "", command: c.command }).ok === false
|
|
210
|
+
);
|
|
211
|
+
add(
|
|
212
|
+
"policy: a silent static gate is not a silent suite",
|
|
213
|
+
redded.length === 0,
|
|
214
|
+
redded.length ? `${redded.map((c) => c.id).join(", ")} hard-redded` : `${SILENT_STATIC_GATES.length} static gates still pass`
|
|
215
|
+
);
|
|
216
|
+
}
|
|
217
|
+
|
|
157
218
|
// --- 3. The opposite failure: flagging what is innocent ---------------------
|
|
158
219
|
{
|
|
159
220
|
const noisy = [];
|
|
160
221
|
for (const e of INNOCENT_EDITS) {
|
|
161
222
|
const res = checkTestTampering(canaryDiff(e));
|
|
162
223
|
const types = (res.violations || []).map((v) => v.type);
|
|
163
|
-
if (types.length > 0) noisy.push(`${e.id}
|
|
224
|
+
if (types.length > 0) noisy.push(`${e.id} \u2192 ${JSON.stringify(types)} (${e.why})`);
|
|
225
|
+
// The verdict, not just the violation list. `UNREADABLE` is a status, so
|
|
226
|
+
// this check stayed green for two releases while the gate rejected every
|
|
227
|
+
// edit below on a CRITICAL finding.
|
|
228
|
+
const scan = scanDiff(canaryDiff(e));
|
|
229
|
+
if (!scan.ok) noisy.push(`${e.id} \u2192 gate rejected: ${JSON.stringify((scan.findings || []).map((f) => f.type))}`);
|
|
164
230
|
}
|
|
165
231
|
add(
|
|
166
232
|
"innocent edits stay silent",
|