jules-orchestrator-kit 0.72.2 → 0.72.3

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -13,7 +13,8 @@ This document outlines the hard constraints and system prompting best practices
13
13
  - **Git Base Drift (Stale Merge-Base Reverts)**: If `main` advances during a session, `git diff main pr-N` shows branch *divergence*, not the applied patch. Merging blindly can silently revert unrelated files updated on `main`.
14
14
  - **Edge Isolates (Runtime Boundary Breaches)**: Edge environments (e.g. Cloudflare `workerd`) enforce strict limits (128 MB RAM, 10 MiB bundle cap). Jules may import heavy native libraries (`sharp`, `canvas`) that pass Node tests in the VM but crash worker deployment.
15
15
  - **CMS/DB Credentials (Visual & E2E Test Failures)**: Cloud VMs lack live CMS API keys, DB credentials, and display servers. Headful E2E or visual screenshot tests (e.g. Playwright) fail with 500 errors.
16
- - **CLI Dry-Run Drift (Misleading Pull Diffs)**: Running `jules remote pull --session <id>` (dry-run without `--apply`) on a session that crashed or made no commits can output cached or unrelated diffs.
16
+ - **Runtime Environment & Resource Gates (KVM / 4 vCPU / 8 GiB RAM / No Swap)**: The execution sandbox runs in KVM on Ubuntu 24.04.2 LTS (`x86_64`, kernel 6.8.0, hostname `devbox`, user `jules`, uid=1001, passwordless `sudo`) with the target repo mounted at `/workspace`. Memory is capped via cgroup v2 `memory.max` at 8 GiB with swap completely disabled (`memory.swap.max = 0`); exceeding resident memory triggers the Linux kernel OOM killer (Exit code 137) immediately. Writable OverlayFS quota is 20–30 GiB. PID 1 is a headless supervisor without `systemd` (`systemctl` and `service` commands fail), so background services must be launched directly via bash (`cmd &`).
17
+ - **Bootstrap Ingestion Hierarchy**: The startup harness parses only the Task Prompt / Issue Context, root `AGENTS.md`, and root `README.md`. Sub-package `AGENTS.md` manifests are discovered dynamically during file traversal (not during bootstrap planning), and IDE/agent configs (`.agent/rules/`, `.cursorrules`, `CLAUDE.md`, `.jules.yml`) are ignored by the harness. Critical invariants must be stated in root `AGENTS.md` or injected into the task envelope.
17
18
 
18
19
  ## 2. System Prompting & Guardrail Best Practices
19
20
 
@@ -22,7 +23,7 @@ To maximize the ratio of mergeable PRs vs. failed or hallucinated sessions:
22
23
  1. **Strict File Scoping:** Constrain file I/O using explicit glob patterns in session prompts.
23
24
  2. **Immutable Boundary Directives:** Explicitly forbid modification of `*.lock` files, database migration histories, and core configuration files.
24
25
  3. **Deterministic Test Verification Mandate:** Require explicit verification commands with zero-exit-code constraints before PR generation is permitted.
25
- 4. **Sub-Package `AGENTS.md` Hierarchy:** Place localized `AGENTS.md` files at sub-package boundaries in monorepos to restrict dependency resolution graphs and operational blast radius.
26
+ 4. **Sub-Package `AGENTS.md` Hierarchy:** Place localized `AGENTS.md` files at sub-package boundaries in monorepos to restrict dependency resolution graphs and operational blast radius. Note: Sub-package manifests are discovered during file traversal, not during initial bootstrap planning, so critical repo-wide invariants must remain in root `AGENTS.md`.
26
27
  5. **Evidence-Based PR Requirement:** Require every PR to include commands run, exit codes, coverage/performance deltas, and risk assessments.
27
28
  6. **No-Weakening Rule:** Explicitly forbid deleting tests, reducing assertion strength, disabling lint/type checks, or ignoring security warnings.
28
29
  7. **Benchmark Threshold Rule:** Performance changes must include multiple benchmark runs, median comparison, and a minimum improvement threshold (e.g. ≥ 5%).
@@ -31,7 +32,7 @@ To maximize the ratio of mergeable PRs vs. failed or hallucinated sessions:
31
32
  10. **Stop-on-Uncertainty Rule:** If the task cannot be completed safely within scope, the agent must stop without opening a PR rather than guessing.
32
33
  11. **Pre-Dispatch Grounding Mandate:** Verify all file paths, script names, and exported symbols against the live repository tree before writing them into a prompt.
33
34
  12. **Programmatic CI Scope Guarding:** Enforce prompt constraints at the CI level using an unbypassable `Agent Scope Guard` workflow that evaluates diffs against a protected paths manifest (`.agent/protected-paths.json`).
34
- 13. **Google Labs Exploration Budget Protocol:** Execute tasks across 3 discrete phases: (1) Discovery & Symbol Tracing (silent inspection, write NO code), (2) Oracle & Test Formulation, and (3) Surgical Implementation & Verification.
35
+ 13. **Google Labs Exploration Budget Protocol ("Use deep planning mode"):** Execute tasks across 3 discrete phases: (1) Discovery & Symbol Tracing (silent inspection, write NO code), (2) Oracle & Test Formulation, and (3) Surgical Implementation & Verification. Incorporate "Use deep planning mode" in task envelopes to steer model attention toward methodical verification before code modification.
35
36
  14. **Critic Agent Steering (Adversarial Pre-Review):** Jules' internal Critic Agent must evaluate proposed patches for edge-case regressions, $O(n^2)$ bottlenecks, unhandled parameters, and layout shifts (CLS) prior to PR submission. When modifying test suites or error handling, the agent must deliberately mutate production code and induce real failure conditions to prove that tests and catch blocks actually fail as intended (preventing tautological tests).
36
37
  15. **Web Excellence & Frontend Guardrails:** Enforce quantitative Core Web Vitals (LCP < 1.2s, CLS < 0.05), WCAG 2.2 AA/AAA semantic accessibility, Schema.org JSON-LD compliance, and Playwright multi-viewport responsive testing.
37
38
  16. **Airtight Positive Enclosures ("Pink Elephant" Rule):** Replace long negative constraint lists with strict positive operational perimeters (`ONLY modify [Target/Module]`) to prevent attention-drift in deep context windows.
package/CHANGELOG.md CHANGED
@@ -5,6 +5,15 @@ All notable changes to this project will be documented in this file.
5
5
  The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.0.0/),
6
6
  and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html).
7
7
 
8
+ ## [0.72.3] - 2026-09-08
9
+ *A directive not understood is silence; a tag not stripped is a backdoor.*
10
+
11
+ Hardens prompt boundaries against Unicode Tag ASCII smuggling, expands clinical vocabulary to mitigate Vertex AI false-positive aborts, injects deep planning directives, and grounds runtime VM/cgroup constraints:
12
+ - **Unicode Tag Plane & ASCII Smuggling Defense (`src/prompt-guard.mjs`, `src/security.mjs`)**: Gemini Pro tokenizes invisible Unicode Tag characters (`U+E0000` to `U+E007F`) as operational ASCII tokens. Both `ZERO_WIDTH_AND_BIDI_REGEX` and `INVISIBLE_CHARS` now strip and scan Plane 14 tag characters (`[\u{E0000}-\u{E007F}]`), preventing hidden prompt injections in GitHub issues/PR comments and credential hiding in diffs.
13
+ - **Vertex AI Safety Classifier Mitigations (`src/prompt-guard.mjs`)**: Upstream Vertex AI content moderation (`HARM_CATEGORY_DANGEROUS_CONTENT`) frequently aborts sessions upon encountering aggressive process management syntax. `STERILE_VOCABULARY_MAP` now clinicalizes `kill -9` $\rightarrow$ `terminate with SIGTERM`, `SIGKILL` $\rightarrow$ `SIGTERM`, `reap zombie processes` $\rightarrow$ `collect child processes`, and `exploit_payload` $\rightarrow$ `test_payload`.
14
+ - **Deep Planning Mode Directive Injection (`src/task-optimizer.mjs`, `src/web-templates.mjs`)**: Integrated the model-steering phrase `"Use deep planning mode."` into the Exploration Budget Protocol envelopes, reinforcing phased reasoning prior to code mutation.
15
+ - **Grounded VM & cgroup Invariant Documentation (`.agent/rules/jules-protocol.md`)**: Formally documented KVM sandbox boundaries: Ubuntu 24.04.2 LTS (`x86_64`), cgroup v2 hard limit 8 GiB RAM with zero swap (`memory.swap.max = 0`), 20–30 GiB OverlayFS writable quota, `/workspace` repo mount, headless supervisor init (no systemd), and the three-source startup ingestion hierarchy (Task Prompt, root `AGENTS.md`, and root `README.md`).
16
+
8
17
  ## [0.72.2] - 2026-09-08
9
18
  *A test suite that writes to stderr did not write silence.*
10
19
 
package/README.md CHANGED
@@ -49,32 +49,28 @@
49
49
  <a id="quickstart"></a>
50
50
  ## Quickstart
51
51
 
52
- Get running in any repository in 3 commands. `init` asks seven questions and
53
- fills in a sensible answer for each; `--yes` accepts all of them, detects the
54
- stack, and probes the test command it picked before writing it down.
52
+ Configure any repository in three steps. `init` inspects project manifests, detects the stack, probes the test runner, and scaffolds repository guardrails:
55
53
 
56
54
  ```bash
57
- # 1. Scaffold config, AGENTS.md, role prompts and guardrails
58
- # (auto-detects Python, Rust, Go, Node, PHP, etc.)
59
- # Drop --yes to choose provider, plan, profile and workflows yourself.
55
+ # 1. Scaffold configuration, AGENTS.md, role prompts, and guardrails
56
+ # Auto-detects Python, Rust, Go, Bun, Deno, Node, PHP, .NET, etc.
57
+ # Omit --yes to select provider, plan tier, and verification profile interactively.
60
58
  npx jules-orchestrator-kit init --yes
61
59
  ```
62
60
 
63
61
  ```bash
64
- # 2. Commit what init wrote — .agent/config.yml is protected by BUILTIN_PROTECT,
65
- # so leaving it uncommitted makes the first gate reject your tree
62
+ # 2. Commit the scaffolded configuration
63
+ # .agent/config.yml is protected by scope guards; committing establishes the trusted base policy.
66
64
  git add .agent AGENTS.md SPEC.md CONSTRAINTS.md .gitignore && git commit -m "chore: add agent config"
67
65
  ```
68
66
 
69
67
  ```bash
70
- # 3. Author a scoped, verified task envelope with guardrails & secret scrubbing
71
- # Interactive by default. Pass the prompt to skip straight to review:
72
- npx jules-orchestrator-kit task create -p "Refactor the invoice module"
68
+ # 3. Author a scoped, verified task envelope
69
+ # Interactive by default. Pass --prompt and --verify to define requirements directly:
70
+ npx jules-orchestrator-kit task create -p "Refactor invoice calculation" --verify "npm test"
73
71
  ```
74
72
 
75
- `init` reads the repository, not a template: it detects the stack, picks a
76
- provider this machine can actually reach, and generates a CI workflow for the
77
- toolchain the project uses. Nothing about your setup is assumed.
73
+ `init` derives configuration directly from repository manifests, connects reachable agent providers, and generates CI workflows matching the project toolchain.
78
74
 
79
75
  ```bash
80
76
  # Which agents can this machine dispatch to, and what is missing for the rest?
@@ -87,14 +83,12 @@ npx jules-orchestrator-kit profile --set max
87
83
  ```
88
84
 
89
85
  > [!TIP]
90
- > **Not sure what to run next?**
91
- > `agentctl` with no arguments reads the repository state and prints the single
92
- > next step — missing git repo, missing API key, empty queue, tasks ready to
93
- > dispatch — instead of a wall of commands.
86
+ > **Context-Aware Next Step:**
87
+ > Running `agentctl` without arguments inspects the local repository state (git status, active API keys, queued tasks) and prints the immediate next action.
94
88
 
95
89
  > [!TIP]
96
- > **Prefer a global CLI?**
97
- > Install globally to access `agentctl` directly:
90
+ > **Global Installation:**
91
+ > Install globally for direct command access:
98
92
  > ```bash
99
93
  > npm install -g jules-orchestrator-kit
100
94
  > agentctl init && agentctl task create && agentctl queue
@@ -109,35 +103,30 @@ npx jules-orchestrator-kit profile --set max
109
103
  <a id="any-repository"></a>
110
104
  ## Using It In Any Repository
111
105
 
112
- Four things differ between projects, and the kit resolves each one from the
113
- repository rather than from a template.
106
+ The kit derives configuration directly from repository manifests across five core dimensions:
114
107
 
115
- | What differs | How it is resolved | Inspect / override |
108
+ | Dimension | Resolution Mechanism | Inspect / Override |
116
109
  | :--- | :--- | :--- |
117
- | **Which suites to run** | A monorepo change resolves to the sub-projects it touches (`verify.scope: affected`), widening back to the root command as soon as it reaches a shared file. Off by default, on for repositories `init` detects as monorepos. | `agentctl check --json` · `verify.scope` in `.agent/config.yml` |
118
- | **The stack** | `detectStack()` recognises 24+ ecosystems (Cargo, Go, Python/Django, Maven/Gradle, .NET, PHP/Laravel, Ruby, Elixir, Swift, Flutter/Dart, CMake, Bun, Deno, Node + Turbo/pnpm/Nx workspaces) and derives the setup, lint, test and build commands from the manifest it finds. | `agentctl doctor` · `verify:` in `.agent/config.yml` |
119
- | **The agent** | `provider:` selects Google Jules (hosted REST), the Claude Code CLI, the Codex CLI or the Gemini CLI. Readiness means a credential for the hosted one and a binary on `PATH` for the local ones — never both. | `agentctl providers` · `agentctl init --provider <name>` |
120
- | **How hard to verify** | `verify.profile` expands at load time into a stage pipeline that skips gates the runtime cannot support, and says which and why. | `agentctl profile` · `agentctl profile --set max` |
121
- | **Where CI runs** | A workflow is *generated* for the detected stack — the project's toolchain plus Node for the CLI — not copied from this repository. | `agentctl ci init [--target github\|gitlab]` |
110
+ | **Monorepo Scope** | Monorepo diffs resolve to affected sub-projects (`verify.scope: affected`), widening to root commands when shared files change. Activated automatically when monorepo manifests are detected. | `agentctl check --json`<br/>`verify.scope` in `.agent/config.yml` |
111
+ | **Stack & Tooling** | `detectPolyglotStack()` inspects 26+ ecosystems (Cargo, Go, Python, Bun, Deno, Maven, Gradle, .NET, PHP, Ruby, Elixir, Swift, Flutter, CMake, Make, Turbo/pnpm/Nx) and extracts native test and build commands. | `agentctl doctor`<br/>`verify:` in `.agent/config.yml` |
112
+ | **Agent Provider** | Supports Google Jules (hosted REST), Claude Code CLI, OpenAI Codex CLI, and Gemini CLI. Validates environment credentials for hosted APIs and `PATH` binaries for local agents. | `agentctl providers`<br/>`agentctl init --provider <name>` |
113
+ | **Verification Depth** | `verify.profile` (`minimal`, `standard`, `max`) expands dynamically into stack-compatible verification stages, reporting explicit skip reasons for unsupported platform checks. | `agentctl profile`<br/>`agentctl profile --set max` |
114
+ | **CI Generation** | Generates tailored CI workflows containing the project's native runtime and toolchain rather than copying a fixed template. | `agentctl ci init [--target github\|gitlab]` |
122
115
 
123
- ### Verification profiles
116
+ ### Verification Profiles
124
117
 
125
- | Profile | Runs | Use when |
118
+ | Profile | Stages | Recommended Use |
126
119
  | :--- | :--- | :--- |
127
- | `minimal` | setup → tests | The suite is slow, the stack is unfamiliar, or it is day one. |
128
- | `standard` | setup → lint → tests → build → anti-tamper on the diff | The everyday gate. Scaffolded by default. |
129
- | `max` | everything above → mutation scoring → V8 diff coverage *(Node runtimes only)* → 3-pass flakiness probe | The change is consequential, or an agent has been getting green too easily. |
120
+ | `minimal` | Setup → Tests | Large/slow test suites or initial project onboarding. |
121
+ | `standard` | Setup → Lint → Tests → Build → Diff Anti-Tamper | Default gate for routine pull requests. |
122
+ | `max` | All stages above → AST Mutation Scoring → V8 Diff Coverage *(Node)* → 3-Pass Flakiness Probe | High-risk refactors or critical infrastructure changes. |
130
123
 
131
- Nothing in a profile is Node-specific by assumption. Gates a runtime cannot
132
- support are skipped with a stated reason rather than failing the diff — a Cargo
133
- repository on `max` runs mutation and stability probing and is never asked for
134
- `NODE_V8_COVERAGE`.
124
+ Verification profiles evaluate gates dynamically per runtime. Unsupported platform checks (such as V8 coverage on Cargo or Go projects) are bypassed with explicit diagnostic logs rather than failing the gate.
135
125
 
136
- ### No provider? Still useful
126
+ ### Standalone Local Verification
137
127
 
138
- Every gate below runs locally with no API key, no CLI and no network:
139
- `agentctl check`, `mutate`, `coverage`, `probe`, `assert`, `evidence`, `rules`,
140
- `doctor`. The provider is only needed to *dispatch* work, not to verify it.
128
+ All security, integrity, and test gates execute locally without external network access or API keys:
129
+ `agentctl check`, `agentctl gate`, `agentctl mutate`, `agentctl coverage`, `agentctl probe`, `agentctl evidence`, `agentctl doctor`. Agent providers are required only for dispatching autonomous tasks.
141
130
 
142
131
  <br/>
143
132
 
@@ -195,20 +184,20 @@ To maximize PR merge rates, dispatch tasks according to deterministic boundaries
195
184
 
196
185
  ## Core Capabilities
197
186
 
198
- * **Provider-Agnostic:** Dispatches to Google Jules (hosted REST), the Claude Code CLI, the OpenAI Codex CLI or the Gemini CLI. `agentctl providers` probes each one — a credential for the hosted provider, a binary on `PATH` for the local ones — and every verification gate works with no provider configured at all.
199
- * **Vendor-Neutral Configuration:** Every `JULES_*` environment variable also answers to an `AGENT_*` spelling (`AGENT_API_KEY`, `AGENT_REPO`, `AGENT_SWARM_CONCURRENCY`). The legacy name always wins where both are set, so adding an alias cannot change a working setup.
200
- * **One-Word Verification Depth:** `verify.profile: minimal | standard | max` expands at load time into a stack-aware pipeline — `max` adds mutation scoring, flakiness probing and, where the runtime emits it, V8 diff coverage. A Cargo repository is never asked for `NODE_V8_COVERAGE`.
201
- * **Generated, Not Copied, CI:** `agentctl ci init` writes a GitHub Actions or GitLab job carrying the toolchain the detected stack needs (`setup-python`, `setup-go`, `setup-java`, …) plus Node for the CLI itself.
202
- * **Zero Runtime Dependencies:** Built exclusively on Node.js 20+ built-in modules (`node:fs`, `node:child_process`, `node:crypto`, `node:path`, `node:http`, `node:tty`, `node:test`).
187
+ * **Multi-Provider Dispatch:** Dispatches to Google Jules (hosted REST API), Claude Code CLI, OpenAI Codex CLI, and Gemini CLI. `agentctl providers` inspects environment credentials and binary availability across providers.
188
+ * **Vendor-Neutral Configuration:** Supports both `JULES_*` and `AGENT_*` environment variables (`AGENT_API_KEY`, `AGENT_REPO`, `AGENT_SWARM_CONCURRENCY`), with legacy `JULES_*` variables taking precedence.
189
+ * **Dynamic Verification Profiles:** Configured via `verify.profile: minimal | standard | max`. Automatically schedules linting, unit testing, build stages, AST mutation testing, and stability probing suited to the project toolchain.
190
+ * **Stack-Native Generated CI:** `agentctl ci init` generates GitHub Actions and GitLab CI configurations containing the project's exact toolchain (`setup-python`, `setup-go`, `setup-bun`, etc.) alongside Node.js for CLI execution.
191
+ * **Zero Runtime Dependencies:** Implemented strictly using native Node.js 20+ standard modules (`node:fs`, `node:child_process`, `node:crypto`, `node:path`, `node:http`, `node:readline`, `node:test`).
203
192
  * **Cross-Platform Parity:** Verified 100% green across Linux, macOS (Darwin), and Windows on Node 20, 22, and 24.
204
- * **Autonomous Self-Healing Loop:** Captures test stderr/stdout, fingerprints error traces, and feeds structured context back into automated repair turns (up to 3 attempts) before human escalation.
205
- * **Fail-Closed Verification:** A change that ran no verification command at all is rejected, not approved — "nothing to run" is not a pass. Repositories using only the scope and secret phases opt out explicitly with `verify.required: false`.
206
- * **Anti-Tamper That Reads Semantics:** Counting assertions cannot see a value check swapped for a truthiness check. The guard tracks assertions that name an expected value, so weakening a test is a violation even when the line count is unchanged.
207
- * **Binary-Aware Scanning:** Files git renders as `Binary files ... differ` are read directly for structured credentials, and their real size is charged against the diff ceiling, so a leading NUL byte cannot hide a token and a committed blob cannot walk past the payload governor.
208
- * **Fail-Closed Security & Secret Redaction:** Evaluates explicit Deny rules before Allow rules against canonicalized, case-folded paths. Redacts high-entropy keys and base64-encoded credentials (such as Kubernetes `Secret` manifests).
209
- * **Complexity & Cost Router:** Zero-dependency heuristic classifier (`src/router.mjs`) routing mechanical tasks to lightweight models while reserving primary models for complex refactors, with a `node --check` syntax-verification gate that transparently escalates a FAST-tier result to the primary provider if it left broken JS on disk.
210
- * **Terminal UI & Diagnostic Matrix (`agentctl doctor`):** Interactive terminal dashboard, task sidecar manager, and automated transactional self-repair.
211
- * **Verified Test Suite:** Tested with **1403 unit tests across 196 suites**, green on every supported platform.
193
+ * **Autonomous OODA Repair Loop:** Captures test stdout/stderr traces, fingerprints failure patterns, and executes automated repair cycles (up to 3 turns) before requesting human intervention.
194
+ * **Fail-Closed Verification:** Rejects diffs that execute zero verification commands unless explicitly waived with `verify.required: false`.
195
+ * **Semantic Anti-Tamper Guard:** Detects test tampering across languages: weakened assertions, removed assertions, vacuous tautologies (`expect(true).toBe(true)`), and assertions nested inside dead conditions (`if False:`, `if (false)`, `if 0:`).
196
+ * **Binary & Symlink Payload Inspection:** Inspects binary diffs and symlink targets directly, charging real byte sizes against the diff ceiling to prevent payload governor bypasses.
197
+ * **Fail-Closed Security & Secret Scrubbing:** Evaluates Deny-before-Allow rules against canonicalized paths. Detects high-entropy strings and base64-encoded credentials (e.g. Kubernetes manifests).
198
+ * **Complexity & Cost Router:** Zero-dependency heuristic classifier (`src/router.mjs`) routing mechanical tasks to lightweight models while reserving primary models for complex refactors, backed by syntax-check fallback recovery.
199
+ * **Terminal UI & Diagnostics (`agentctl doctor`):** Interactive terminal dashboard, VFS lock management, and automated system diagnostics.
200
+ * **Mechanically Verified:** Comprehensive test suite of **1410 unit tests across 196 suites**, with 59 activation-coverage canaries and 100% pass rate.
212
201
 
213
202
  <br/>
214
203
 
@@ -452,22 +441,14 @@ const result = await fast.dispatch({ prompt: "Fix a typo." }, { root: process.cw
452
441
  | **Staged Diff Fidelity & Indentation Dead Guards** | `src/git.mjs`, `src/security.mjs` | Query cached index in staged mode (`git diff --cached <base>`), detect literal falsity dead guards (`if False:`, `if (false)`, `if 0:`), and support indentation-aware block traversal for Python test suites. | **v0.72.1** *(Shipped)* |
453
442
  | **Cold-Start Hardened Kernel & Tamper Defense** | `src/config.mjs`, `src/engine.mjs`, `src/git.mjs`, `src/security.mjs` | Full remediation of 22 cold-start audit findings (F01–F22): authoritative base policy resolution, ephemeral snapshot worktree isolation, canonical root test tamper guard, conditional assertion defense, multi-target Cargo test aggregation, Python src-layout injection, and complete repository uninstall documentation. | **v0.72.0** *(Shipped)* |
454
443
  | **Silence Is Not A Suite & Scaffolding Linter Fixes** | `src/ops/test-collection.mjs`, `src/wizard-init.mjs`, `src/config.mjs` | Reject zero-output test suite commands, quote-aware YAML parser with scalar emission, test de-registration detection (`TEST_DEREGISTERED`), and active waiver telemetry banner. | **v0.71.0** *(Shipped)* |
455
- | **Diagnostics That Reach the Operator** | `src/security.mjs`, `src/engine.mjs`, `bin/agentctl.mjs` | Secret findings name the file and line, a failed verify stage reports its command, exit code and output, and `queue`/`swarm` name each failed task and exit `1` rather than reporting success for a run that dispatched nothing. | **v0.41.1** *(Shipped)* |
456
- | **One Scaffolding Path & First-Install Fixes** | `src/scaffold.mjs`, `src/security.mjs` | `agentctl init` and `jules-init` scaffold from one source and write the runtime `.gitignore` entries, so the kit's own bookkeeping no longer reaches its own gate; a lockfile bump no longer fails closed as a secret leak. | **v0.41.1** *(Shipped)* |
457
- | **Queue Runner Fidelity** | `src/dag-engine.mjs`, `src/engine.mjs` | Queue selection is by task shape rather than file extension, so manifests and READMEs are skipped instead of dispatched, and `--dry-run` leaves the queue untouched. | **v0.38.2** *(Shipped)* |
458
- | **Release Gate Enforcement & Wizard Smoke Test** | `.github/workflows/jules-audit.yml`, `scripts/release.mjs`, `test/wizard-smoke.test.mjs` | Doc-sync gate runs in CI rather than by hand, releases block on a green CI matrix for `HEAD`, per-test deadlines turn a hang into a failure, and the real `init` wizard is driven end to end over a fake TTY. | **v0.38.1** *(Shipped)* |
459
- | **Multi-OS CI Matrix & TUI Hardening** | `scripts/run-tests.mjs`, `src/state.mjs`, `src/git.mjs` | Automated 9-job CI matrix across Linux, macOS, and Windows on Node 20/22/24 with raw-mode TUI resilience and native Windows command quoting. | **v0.38.0** *(Shipped)* |
460
- | **Base64 Secret Detection & Budget Fix** | `src/security.mjs`, `src/budget.mjs` | Secret scanner decodes base64 before matching structured patterns (K8s secrets), and `budget reset` preserves confirmed provider sessions. | **v0.37.0** *(Shipped)* |
461
- | **Universal AI Crawler Policy & llms.txt** | `src/web-templates.mjs` (`web-ai-access`) | Cross-surface consistency for crawler directives (`robots.txt`, meta tags, `X-Robots-Tag`) and `llms.txt` local route integrity. | **v0.36.0** *(Shipped)* |
462
- | **Silence Governor & Flaky Test Swarm** | `src/webhook.mjs`, `src/flaky-ledger.mjs` | Notification alert throttling with interruption budgeting, and automated anti-flakiness swarm coordinator. | **v0.35.0** *(Shipped)* |
463
- | **Rolling 24h Quota & Plan Concurrency** | `src/state.mjs`, `src/config.mjs` | Rolling 24-hour quota accounting matching vendor reset windows and true concurrency limits (3/15/60). | **v0.34.0** *(Shipped)* |
464
- | **Cost Router & Guided First Run** | `src/router.mjs`, `src/ops/next-step.mjs` | Heuristic task classifier routing trivial tasks to fast models, and guided single-command first run workflow. | **v0.33.0** *(Shipped)* |
465
- | **DAG Task Queue & Specialist Roles** | `src/dag-engine.mjs`, `src/evidence.mjs` | Kahn's-algorithm dependency queue execution (`queue --dag`), specialist role prompts (`overseer`, `bolt`, `sentinel`, `janitor`), and SHA-256 evidence manifests. | **v0.32.5** *(Shipped)* |
466
- | **Warm Session Resumption & PR Bundler** | `src/provider.mjs`, `src/engine.mjs` | Multi-turn warm session context streaming via `POST /v1alpha/sessions/{id}:sendMessage` & evidence PR descriptions. | **v0.31.0** *(Shipped)* |
467
- | **TDD Harness & Prompt Falsifiability Linter** | `agentctl test-gen`, `agentctl task optimize` | Automated RED-state test generator, `scope.deny` test locking, and prompt testability linter with fuzzy path resolution. | **v0.31.0** *(Shipped)* |
468
- | **Atomic Git Checkpoint & Rollback** | `agentctl rollback` (`src/ops/checkpoint.mjs`) | Pre-flight git HEAD/stash snapshotting, atomic rollback restoration, and 10-session pruning rotation. | **v0.31.0** *(Shipped)* |
469
- | **Terminal UI Engine** | `src/tui.mjs`, `src/key-decoder.mjs` | Zero-dependency terminal capabilities detector, sequence key decoder, and interactive prompt widgets. | **v0.30.0** *(Shipped)* |
470
- | **PR Review Auto-Remediation Loop** | `agentctl review-repair` (`src/review-repair.mjs`) | Ingests GitHub PR review comments (`CHANGES_REQUESTED`), extracts line/file context, and dispatches automated repair turns. | **v0.27.0** *(Shipped)* |
444
+ | **Terminal State Classification & Error Diagnostics** | `src/engine.mjs`, `src/session-ops.mjs` | Triple-verdict session resolution (terminal, blocked, timed-out), diagnostic extraction for failed session retries (`AssertionError`, tracebacks), and 22-case session polling contract. | **v0.70.0** *(Shipped)* |
445
+ | **Expected-Value Scanning & Safe Scaffolding** | `src/security.mjs`, `src/wizard-init.mjs` | Multi-language expected-value first assertions (JUnit/PHPUnit), prefix-aware test rename verification, and rejection of empty/trivial verification oracles. | **v0.69.0** *(Shipped)* |
446
+ | **Dialect Hardening & Scope Protection** | `src/security.mjs`, `src/config.mjs`, `scripts/guard-reach-check.mjs` | Unreadable test dialects fail closed, lockfiles and toolchain pins guarded against silent tampering, and 59-canary activation coverage gate in CI. | **v0.63.0** *(Shipped)* |
447
+ | **Process-Group Reaping & Subprocess Lifecycles** | `scripts/run-tests.mjs`, `src/git.mjs` | Tree-wide process group cleanup on interruption and zero-output test collection guards. | **v0.60.0** *(Shipped)* |
448
+
449
+ <br/>
450
+
451
+ > For the complete history of all shipped milestones (v0.20.0 – v0.72.2), see [ROADMAP_V1.md](ROADMAP_V1.md).
471
452
 
472
453
  </details>
473
454
 
@@ -479,7 +460,7 @@ const result = await fast.dispatch({ prompt: "Fix a typo." }, { root: process.cw
479
460
 
480
461
  ## 🧹 Complete Uninstall / Removing the Kit (Undo Init)
481
462
 
482
- If you need to completely remove `jules-orchestrator-kit` from a repository after running `agentctl init`, follow the procedure below. Note that `agentctl clean` is an operational maintenance command (cleaning ephemeral locks, temporary worktrees, and evidence caches), not an uninstaller.
463
+ To completely remove `jules-orchestrator-kit` from a repository after running `agentctl init`, follow the procedure below. Note that `agentctl clean` performs operational maintenance (clearing ephemeral locks, temporary worktrees, and evidence caches), not an uninstaller.
483
464
 
484
465
  ### 1. Generated Assets & Manifest
485
466
 
@@ -554,6 +535,6 @@ npm uninstall -g jules-orchestrator-kit
554
535
  <br/>
555
536
 
556
537
  <div align="center">
557
- <p><b>jules-orchestrator-kit</b> • Built with zero external dependencies for Google Jules and autonomous agent workflows.</p>
538
+ <p><b>jules-orchestrator-kit</b> • Zero runtime dependencies • MIT License • Universal safety and verification for autonomous coding agents.</p>
558
539
  </div>
559
540
 
package/ROADMAP_V1.md CHANGED
@@ -11,13 +11,19 @@ The **jules-orchestrator-kit** is the zero-dependency safety gatekeeper and self
11
11
  ## 📌 Release Milestones Overview
12
12
 
13
13
  ```
14
- v0.72.2 (Current Stable) ──► v0.73.0 (Distributed Swarms & Leases) ──► v1.0.0 (Production Hardened Kernel)
15
- (Child Streams & Bun/Deno) (Multi-Agent DAG & Resource Locks) (Enterprise Telemetry & SLA)
14
+ v0.72.3 (Current Stable) ──► v0.73.0 (Distributed Swarms & Leases) ──► v1.0.0 (Production Hardened Kernel)
15
+ (ASCII Smuggle & Safety) (Multi-Agent DAG & Resource Locks) (Enterprise Telemetry & SLA)
16
16
  ```
17
17
 
18
18
  ---
19
19
 
20
- ## ✅ Shipped Milestones (v0.20.0 – v0.72.2)
20
+ ## ✅ Shipped Milestones (v0.20.0 – v0.72.3)
21
+
22
+ ### v0.72.3: ASCII Smuggling Defense, Safety Filter Mitigations & Deep Planning Envelopes
23
+ - [x] **Unicode Tag ASCII Smuggling Defense (`src/prompt-guard.mjs`, `src/security.mjs`)** — strips and detects Plane 14 Unicode Tag characters (`U+E0000`–`U+E007F`) across untrusted inputs and secret diffs.
24
+ - [x] **Vertex AI Safety Moderation Mitigations (`src/prompt-guard.mjs`)** — clinicalizes `kill -9`, `SIGKILL`, zombie reaping, and exploit terms to prevent upstream `HARM_CATEGORY_DANGEROUS_CONTENT` aborts.
25
+ - [x] **Deep Planning Mode Steering Directive (`src/task-optimizer.mjs`, `src/web-templates.mjs`)** — injects `"Use deep planning mode."` into exploration budget task envelopes.
26
+ - [x] **Runtime Environment & Ingestion Directives (`.agent/rules/jules-protocol.md`)** — documented KVM 8 GiB swap=0 OOM limits, 20-30 GiB OverlayFS quota, `/workspace` mount, non-systemd supervisor, and startup ingestion hierarchy.
21
27
 
22
28
  ### v0.72.2: Child Process Stream Fidelity & Polyglot Build Detection
23
29
  - [x] **Child Process Stream Fidelity (`src/git.mjs`)** — `runCmd()` invokes native `spawnSync`, preserving both `stdout` and `stderr` streams on exit 0 so test runners emitting summaries to stderr (`bun test`) are fully recognized by `parseCollectedTests`.
@@ -221,8 +227,8 @@ The **jules-orchestrator-kit** is the zero-dependency safety gatekeeper and self
221
227
  ### v0.73.0: Distributed File Leases & Preemptive DAG Scheduling
222
228
  - [ ] **Atomic Filesystem Lease & Heartbeat Protocol (`src/engine.mjs`, `src/flaky-ledger.mjs`)** — Directory-mutex file leasing with heartbeat timestamps, stale-lock detection via PID liveness inspection, and tombstone rotation without third-party daemons or Redis.
223
229
  - [ ] **Preemptive Task Cancellation & Interface Fingerprints (`src/dag-engine.mjs`)** — Automatically aborts and yields downstream swarm tasks when upstream exported symbol interfaces diverge from their cryptographic SHA-256 fingerprints.
224
- - [ ] **POSIX/Win32 Process Group Guillotine (`src/engine.mjs`)** — Tree teardown via `process.kill(-pid, 'SIGKILL')` on POSIX and `taskkill /T /F /PID` on Windows to eliminate orphaned dev-servers and background watchers.
225
- - [ ] **Unicode Trojan Source & Homoglyph Fencing (`src/security.mjs`)** — Deterministic token scanner using V8 Unicode Property Escapes (`\p{Script=...}`) and NFKC normalization to block invisible Bidi overrides (CVE-2021-42574) and mixed-script homoglyphs.
230
+ - [ ] **POSIX/Win32 Process Group Guillotine (`src/git.mjs:runCmd`, `src/engine.mjs`)** — Tree teardown via `process.kill(-pid, 'SIGKILL')` on POSIX and `taskkill /T /F /PID` on Windows in `runCmd` to eliminate orphaned test runners, dev-servers and background watchers on timeout (`ETIMEDOUT`).
231
+ - [ ] **Unicode Trojan Source & Homoglyph Fencing (`src/security.mjs`)** — Deterministic token scanner using V8 Unicode Property Escapes (`\p{Script=...}`) and NFKC normalization to block invisible Bidi overrides (CVE-2021-42574, isolates partially patched in PR #21) and mixed-script homoglyphs.
226
232
 
227
233
  ---
228
234
 
package/bin/agentctl.mjs CHANGED
@@ -293,6 +293,8 @@ async function main() {
293
293
  "check-premise": { type: "boolean" },
294
294
  idempotent: { type: "boolean" },
295
295
  author: { type: "string" },
296
+ "verify-cmd": { type: "string", short: "v" },
297
+ verify: { type: "string" },
296
298
  "dry-run": { type: "boolean", short: "d" },
297
299
  json: { type: "boolean", short: "j" },
298
300
  },
@@ -333,6 +335,7 @@ async function main() {
333
335
  requirePlanApproval: values["require-plan-approval"],
334
336
  checkPremise: values["check-premise"] || values.idempotent,
335
337
  author: values.author,
338
+ verifyCmd: values["verify-cmd"] || values.verify,
336
339
  };
337
340
 
338
341
  try {
@@ -1498,7 +1501,9 @@ async function main() {
1498
1501
  console.log(`--------------------------------------------------`);
1499
1502
  console.log(` Project Root : ${root}`);
1500
1503
  console.log(` Config File : ${config._file || "None (Using defaults)"}`);
1501
- console.log(` Detected Stack : ${detectStack(root).stack}`);
1504
+ const stackInfo = detectStack(root);
1505
+ const edgeSuffix = stackInfo.isEdgeRuntime ? ` (Edge: ${stackInfo.edgePlatform || "generic"})` : "";
1506
+ console.log(` Detected Stack : ${stackInfo.stack}${edgeSuffix}`);
1502
1507
  console.log(` Test Command : ${config.verify.test || "(None)"}`);
1503
1508
  console.log(` Build Command : ${config.verify.build || "(None)"}`);
1504
1509
  console.log(` Daily Budget : ${formatBudgetLine(budgetStatus(config, root))}`);
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "jules-orchestrator-kit",
3
- "version": "0.72.2",
3
+ "version": "0.72.3",
4
4
  "description": "Zero-dependency safety gatekeeper, test oracle generator, and multi-agent coordination protocol for autonomous coding agents — Google Jules, Claude Code, Codex and Gemini CLI.",
5
5
  "repository": {
6
6
  "type": "git",
@@ -317,6 +317,21 @@ export function checkDocSync(root = process.cwd(), opts = {}) {
317
317
  }
318
318
  }
319
319
 
320
+ // 6. SECURITY.md — supported version table must name the current version series.
321
+ const security = readIfExists(join(root, "SECURITY.md"));
322
+ if (security === null) {
323
+ add("SECURITY present", false, "SECURITY.md not found");
324
+ } else {
325
+ const sem = version.split(".");
326
+ const expectedSeries = `v${sem[0]}.${sem[1]}.x`;
327
+ const secMatch = security.match(/\|\s*`v(\d+\.\d+)\.x`\s*\(Latest\)\s*\|/);
328
+ add(
329
+ "SECURITY supported version",
330
+ secMatch ? `v${secMatch[1]}.x` === expectedSeries : false,
331
+ secMatch ? `found v${secMatch[1]}.x, expected ${expectedSeries}` : "no `vX.Y.x (Latest)` marker found"
332
+ );
333
+ }
334
+
320
335
  return { ok: checks.every((c) => c.ok), version, checks };
321
336
  }
322
337
 
package/src/mcp.mjs CHANGED
@@ -563,9 +563,11 @@ export async function handleMcpRequest(request, opts = {}) {
563
563
  };
564
564
  }
565
565
 
566
+ const isDryRun = Boolean(args.dryRun ?? opts?.dryRun ?? opts?.config?.dryRun);
567
+
566
568
  if (toolName === "jules_list_sessions") {
567
569
  const provider = createProvider(config.provider || "jules", config);
568
- const res = await provider.listSessions({ root, pageSize: args.pageSize, pageToken: args.pageToken });
570
+ const res = await provider.listSessions({ root, pageSize: args.pageSize, pageToken: args.pageToken, dryRun: isDryRun });
569
571
  return {
570
572
  jsonrpc: "2.0",
571
573
  id,
@@ -584,7 +586,7 @@ export async function handleMcpRequest(request, opts = {}) {
584
586
  };
585
587
  }
586
588
  const provider = createProvider(config.provider || "jules", config);
587
- const res = await provider.listActivities(args.sessionId, { root, pageSize: args.pageSize, pageToken: args.pageToken });
589
+ const res = await provider.listActivities(args.sessionId, { root, pageSize: args.pageSize, pageToken: args.pageToken, dryRun: isDryRun });
588
590
  return {
589
591
  jsonrpc: "2.0",
590
592
  id,
@@ -621,7 +623,7 @@ export async function handleMcpRequest(request, opts = {}) {
621
623
  };
622
624
  }
623
625
  const provider = createProvider(config.provider || "jules", config);
624
- const res = await provider.archiveSession(args.sessionId, { root });
626
+ const res = await provider.archiveSession(args.sessionId, { root, dryRun: isDryRun });
625
627
  return {
626
628
  jsonrpc: "2.0",
627
629
  id,
@@ -640,7 +642,7 @@ export async function handleMcpRequest(request, opts = {}) {
640
642
  };
641
643
  }
642
644
  const provider = createProvider(config.provider || "jules", config);
643
- const res = await provider.deleteSession(args.sessionId, { root });
645
+ const res = await provider.deleteSession(args.sessionId, { root, dryRun: isDryRun });
644
646
  return {
645
647
  jsonrpc: "2.0",
646
648
  id,
package/src/memory.mjs CHANGED
Binary file
@@ -6,7 +6,7 @@
6
6
  * neutralizes LLM control role markers and prompt injection patterns, and wraps inputs in strict tags.
7
7
  */
8
8
 
9
- const ZERO_WIDTH_AND_BIDI_REGEX = /[\u200B-\u200F\u202A-\u202E\u2060-\u206F\uFEFF]/g;
9
+ const ZERO_WIDTH_AND_BIDI_REGEX = /[\u200B-\u200F\u202A-\u202E\u2060-\u206F\uFEFF]|[\u{E0000}-\u{E007F}]/gu;
10
10
  const ANSI_ESCAPE_REGEX = /\u001B(?:[@-Z\\-_]|\[[0-?]*[ -/]*[@-~])/g;
11
11
 
12
12
  const CONTROL_TAGS_REGEX = /<\|im_start\|>|<\|im_end\|>|<\|system\|>|<\|assistant\|>|<\|user\|>|<\|endoftext\|>|\[INST\]|\[\/INST\]|<<SYS>>|<\/SYS>>/gi;
@@ -27,15 +27,21 @@ const INJECTION_PATTERNS = [
27
27
  * to prevent false-positive safety classifier tripwires in Google VM environments.
28
28
  */
29
29
  export const STERILE_VOCABULARY_MAP = [
30
+ { pattern: /\bkill\s+-9\b/gi, replacement: "terminate with SIGTERM" },
31
+ { pattern: /\bSIGKILL\b/g, replacement: "SIGTERM" },
30
32
  { pattern: /\bkill\s+(?:the\s+)?processes\b/gi, replacement: "terminate the processes" },
31
33
  { pattern: /\bkill\s+(?:the\s+)?process\b/gi, replacement: "terminate the process" },
32
34
  { pattern: /\bkill\s+(?:the\s+)?pid\b/gi, replacement: "terminate the PID" },
35
+ { pattern: /\breap\s+(?:the\s+)?(?:zombie\s+)?processes\b/gi, replacement: "collect child processes" },
36
+ { pattern: /\breap\s+(?:the\s+)?(?:zombie\s+)?process\b/gi, replacement: "collect child process" },
37
+ { pattern: /\breap\s+(?:the\s+)?zombies\b/gi, replacement: "collect child processes" },
33
38
  { pattern: /\bamputate\s+(?:dead\s+|unused\s+)?code\b/gi, replacement: "prune unused code" },
34
39
  { pattern: /\bamputate\b/gi, replacement: "prune" },
35
40
  { pattern: /\bsabotage\s+(?:the\s+)?tests?\b/gi, replacement: "mutate test logic" },
36
41
  { pattern: /\bsabotage\b/gi, replacement: "mutate" },
37
42
  { pattern: /\bdestroy\s+(?:the\s+)?(state|cache|file|data|directory)\b/gi, replacement: "purge the $1" },
38
43
  { pattern: /\bwipe\s+(?:the\s+)?(state|cache|directory|disk|table)\b/gi, replacement: "clear the $1" },
44
+ { pattern: /\bexploit_payload\b/gi, replacement: "test_payload" },
39
45
  ];
40
46
 
41
47
  /**
package/src/security.mjs CHANGED
@@ -3,7 +3,7 @@ import { isTestPath } from "./test-paths.mjs";
3
3
  import { dirname, join, basename } from "node:path";
4
4
  import { randomBytes } from "node:crypto";
5
5
  import { canonicalizePath, isWindowsAbsolutePath } from "./config.mjs";
6
- import { detectCrossPackageBoundaryViolations } from "./stack-detector.mjs";
6
+ import { detectCrossPackageBoundaryViolations, detectEdgeRuntime } from "./stack-detector.mjs";
7
7
 
8
8
  export const HIGH_CONFIDENCE_PATTERNS = [
9
9
  /\bghp_[A-Za-z0-9_]{36,255}\b/g,
@@ -448,7 +448,8 @@ export function checkEdgeRuntimeImports(diffOrText = "", options = {}) {
448
448
 
449
449
  const isEdgeExplicit = options.isEdgeRuntime === true;
450
450
  const hasEdgeExport = /export\s+const\s+runtime\s*=\s*['"]edge['"]/i.test(diffOrText);
451
- const isEdgeContext = isEdgeExplicit || hasEdgeExport;
451
+ const isEdgeDetected = options.root ? detectEdgeRuntime(options.root).isEdgeRuntime : false;
452
+ const isEdgeContext = isEdgeExplicit || hasEdgeExport || isEdgeDetected;
452
453
 
453
454
  if (!isEdgeContext) {
454
455
  return { ok: true, violations: [] };
@@ -526,9 +527,9 @@ export function checkCrossPackageImports(diffOrText = "", root = process.cwd(),
526
527
  };
527
528
  }
528
529
 
529
- // Zero-width and bidi-control characters. Inserting one mid-token defeats a
530
- // regex without changing how the value renders, copies, or authenticates.
531
- const INVISIBLE_CHARS = /[\u00AD\u200B-\u200F\u2028\u2029\u202A-\u202E\u2060-\u2064\uFEFF]/g;
530
+ // Zero-width, bidi-control characters, and Unicode tag plane (U+E0000..U+E007F).
531
+ // Inserting one mid-token defeats a regex without changing how the value renders, copies, or authenticates.
532
+ const INVISIBLE_CHARS = /[\u00AD\u200B-\u200F\u2028\u2029\u202A-\u202E\u2060-\u2064\u2066-\u2069\uFEFF]|[\u{E0000}-\u{E007F}]/gu;
532
533
 
533
534
  // Unicode lookalikes that NFKD does NOT decompose. Full-width and other
534
535
  // compatibility forms are handled by String#normalize("NFKD") below; these are
@@ -552,6 +553,8 @@ const CONFUSABLE_TO_ASCII = new Map([
552
553
  ["K", "K"], // U+212A KELVIN SIGN
553
554
  ]);
554
555
 
556
+ const CONFUSABLE_REGEX = new RegExp([...CONFUSABLE_TO_ASCII.keys()].join("|"), "g");
557
+
555
558
  /**
556
559
  * Reduces the confusable spellings a credential can hide behind to plain
557
560
  * ASCII before the secret patterns run (`SEC-04`).
@@ -577,9 +580,7 @@ function normalizeSecretText(str) {
577
580
  out = out.normalize("NFKD");
578
581
  } catch (_) {}
579
582
  out = out.replace(/[\u0300-\u036f]/g, "");
580
- for (const [from, to] of CONFUSABLE_TO_ASCII) {
581
- out = out.split(from).join(to);
582
- }
583
+ out = out.replace(CONFUSABLE_REGEX, (m) => CONFUSABLE_TO_ASCII.get(m));
583
584
  return out;
584
585
  }
585
586
 
@@ -221,7 +221,10 @@ test("every source file parses", () => {
221
221
  }
222
222
 
223
223
  export function detectEdgeRuntime(projectRoot = process.cwd()) {
224
- const isCloudflare = existsSync(join(projectRoot, "wrangler.toml")) || existsSync(join(projectRoot, "wrangler.json"));
224
+ const isCloudflare =
225
+ existsSync(join(projectRoot, "wrangler.toml")) ||
226
+ existsSync(join(projectRoot, "wrangler.json")) ||
227
+ existsSync(join(projectRoot, "wrangler.jsonc"));
225
228
  const isDeno = existsSync(join(projectRoot, "deno.json")) || existsSync(join(projectRoot, "deno.jsonc"));
226
229
  const isNetlify = existsSync(join(projectRoot, "netlify.toml"));
227
230
 
@@ -441,7 +444,7 @@ export function detectPolyglotStack(projectRoot = process.cwd()) {
441
444
 
442
445
  // 5. Python / Django / Elixir / Ruby / Java
443
446
  if (existsSync(join(projectRoot, "manage.py"))) {
444
- return { ...container, stack: "django", testCmd: "python manage.py test --keepdb", buildCmd: "python manage.py check", triggerFile: "manage.py" };
447
+ return { ...container, stack: "django", testCmd: `${pythonBin()} manage.py test --keepdb`, buildCmd: `${pythonBin()} manage.py check`, triggerFile: "manage.py" };
445
448
  }
446
449
  if (existsSync(join(projectRoot, "pyproject.toml")) || existsSync(join(projectRoot, "requirements.txt")) || existsSync(join(projectRoot, "setup.py"))) {
447
450
  const triggerFile = existsSync(join(projectRoot, "pyproject.toml")) ? "pyproject.toml" : existsSync(join(projectRoot, "requirements.txt")) ? "requirements.txt" : "setup.py";
@@ -4,6 +4,7 @@ import { checkScope } from "./security.mjs";
4
4
  import { detectStackOracles } from "./wizard-oracle.mjs";
5
5
  import { loadConfig } from "./config.mjs";
6
6
  import { sanitizePromptVocabulary } from "./prompt-guard.mjs";
7
+ import { buildGuardrailFooter } from "./wizard-task.mjs";
7
8
 
8
9
  /**
9
10
  * Calculates Levenshtein distance between two strings.
@@ -350,7 +351,7 @@ export function optimizeTaskPrompt(promptText, options = {}) {
350
351
  const isWeb = Boolean(options.web || (analysis.webIntent && analysis.webIntent.isWeb));
351
352
  const includeExplorationBudget = options.explorationBudget !== false;
352
353
  const includeCriticGuidance = options.criticGuidance !== false;
353
- const verifyCmd = analysis.oracle.command || options.verifyCmd || "npm test";
354
+ const verifyCmd = analysis.oracle.command || options.verifyCmd || null;
354
355
 
355
356
  // Construct structured Markdown envelope
356
357
  const lines = [];
@@ -359,16 +360,24 @@ export function optimizeTaskPrompt(promptText, options = {}) {
359
360
 
360
361
  if (includeExplorationBudget) {
361
362
  lines.push("## Google Labs Exploration Budget Protocol (3-Phase Discovery)");
362
- lines.push("To maximize diagnostic accuracy (Hit@5 57%), execute this task in 3 distinct phases:");
363
+ lines.push("Use deep planning mode. To maximize diagnostic accuracy (Hit@5 57%), execute this task in 3 distinct phases:");
363
364
  lines.push("1. **PHASE 1: DISCOVERY & SYMBOL TRACING (Stay Silent, Write NO Code)**");
364
365
  lines.push(" - Read target source files, definitions, and dependent call sites.");
365
366
  lines.push(" - Formulate diagnostic hypothesis and verify exact symbol signatures before making edits.");
366
367
  lines.push("2. **PHASE 2: ORACLE FORMULATION**");
367
- lines.push(` - Execute baseline verification: \`${verifyCmd}\`.`);
368
+ if (verifyCmd) {
369
+ lines.push(` - Execute baseline verification: \`${verifyCmd}\`.`);
370
+ } else {
371
+ lines.push(" - Execute baseline verification: (None specified - define verification command).");
372
+ }
368
373
  lines.push(" - Identify specific test assertions, benchmarks, or status codes to satisfy.");
369
374
  lines.push("3. **PHASE 3: SURGICAL IMPLEMENTATION & VERIFICATION**");
370
375
  lines.push(" - Apply minimal, zero-bloat code modifications.");
371
- lines.push(` - Execute \`${verifyCmd}\` and verify 100% clean exit code 0.`);
376
+ if (verifyCmd) {
377
+ lines.push(` - Execute \`${verifyCmd}\` and verify 100% clean exit code 0.`);
378
+ } else {
379
+ lines.push(" - Execute automated verification suite and verify 100% clean exit code 0.");
380
+ }
372
381
  lines.push("");
373
382
  }
374
383
 
@@ -376,6 +385,8 @@ export function optimizeTaskPrompt(promptText, options = {}) {
376
385
  lines.push(`- **Goal**: ${promptBody}`);
377
386
  if (analysis.oracle.command) {
378
387
  lines.push(`- **Verification Command**: \`${analysis.oracle.command}\` (Must pass cleanly with exit code 0)`);
388
+ } else {
389
+ lines.push("- **Verification Command**: (None) - WARNING: Missing test oracle.");
379
390
  }
380
391
  lines.push("- **Falsifiability Criteria**: Zero deleted tests, zero weakened assertions, zero lint errors.");
381
392
  lines.push("");
@@ -401,10 +412,26 @@ export function optimizeTaskPrompt(promptText, options = {}) {
401
412
  lines.push("");
402
413
  }
403
414
 
415
+ const root = options.rootDir || process.cwd();
416
+ let cfg = options.config;
417
+ if (!cfg) {
418
+ try {
419
+ cfg = loadConfig(root);
420
+ } catch (_) {
421
+ cfg = {};
422
+ }
423
+ }
424
+
404
425
  lines.push("## Standard Guardrails");
405
- lines.push("- Do NOT modify package.json, lockfiles, or .github/ infrastructure files.");
406
- lines.push("- Diff Payload Governor: Keep total diff payload under 75 KB (\`git diff | wc -c\`).");
407
- lines.push(`- Verify before finishing: Execute \`${verifyCmd}\` and confirm zero errors.`);
426
+ const rawFooter = buildGuardrailFooter(cfg, {
427
+ baseBranch: options.baseBranch,
428
+ diffKb: options.diffKb,
429
+ }).trim();
430
+ const cleanedFooter = rawFooter.replace(/^---\s*\n?HARD CONSTRAINTS:\s*\n?/, "").trim();
431
+ lines.push(cleanedFooter);
432
+ if (verifyCmd) {
433
+ lines.push(`- Verify before finishing: Execute \`${verifyCmd}\` and confirm zero errors.`);
434
+ }
408
435
 
409
436
  const optimizedPrompt = lines.join("\n");
410
437
 
@@ -920,7 +920,7 @@ export function synthesizeWebEnvelope(templateId, userParams = {}, options = {})
920
920
 
921
921
  if (explorationBudget) {
922
922
  lines.push("## Google Labs Exploration Budget Protocol (3-Phase Discovery)");
923
- lines.push("To maximize diagnostic accuracy (Hit@5 57%), execute this task in 3 distinct phases:");
923
+ lines.push("Use deep planning mode. To maximize diagnostic accuracy (Hit@5 57%), execute this task in 3 distinct phases:");
924
924
  lines.push("1. **PHASE 1: DISCOVERY & SYMBOL TRACING (Stay Silent, Write NO Code)**");
925
925
  lines.push(" - Read target component files, CSS definitions, imports, and existing test specs.");
926
926
  lines.push(" - Formulate diagnostic hypothesis and verify symbol signatures before planning edits.");
@@ -3,6 +3,7 @@ import { join } from "node:path";
3
3
  import { spawnSync } from "node:child_process";
4
4
  import { detectStack, parseYaml } from "./config.mjs";
5
5
  import { runCmd } from "./git.mjs";
6
+ import { pytestCmd } from "./stack-detector.mjs";
6
7
 
7
8
  /**
8
9
  * Check if a command-line binary is available on the system PATH.
@@ -92,7 +93,7 @@ export function detectStackOracles(root = process.cwd()) {
92
93
  }
93
94
 
94
95
  // Check Go module
95
- if (stack.includes("go") || existsSync(join(root, "go.mod"))) {
96
+ if (stack === "go" || existsSync(join(root, "go.mod"))) {
96
97
  candidates.testCmd = candidates.testCmd || "go test ./...";
97
98
  candidates.buildCmd = candidates.buildCmd || "go build ./...";
98
99
  if (hasBinary("golangci-lint")) {
@@ -102,8 +103,8 @@ export function detectStackOracles(root = process.cwd()) {
102
103
  }
103
104
 
104
105
  // Check Python
105
- if (stack.includes("pytest") || existsSync(join(root, "pyproject.toml")) || existsSync(join(root, "requirements.txt"))) {
106
- candidates.testCmd = candidates.testCmd || "pytest";
106
+ if (stack === "python" || stack === "django" || existsSync(join(root, "pyproject.toml")) || existsSync(join(root, "requirements.txt"))) {
107
+ candidates.testCmd = candidates.testCmd || pytestCmd(process.env, root);
107
108
  candidates.buildCmd = candidates.buildCmd || "";
108
109
  if (hasBinary("flake8")) {
109
110
  candidates.lintCmd = candidates.lintCmd || "flake8 .";