jules-orchestrator-kit 0.29.1 β†’ 0.31.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/README.md CHANGED
@@ -23,12 +23,15 @@
23
23
 
24
24
  <p align="center">
25
25
  <a href="#what-is-kit">πŸ’‘&nbsp;What&nbsp;is&nbsp;Kit?</a> &nbsp;β€’&nbsp;
26
+ <a href="#who-is-it-for">πŸ‘₯&nbsp;Who&nbsp;Is&nbsp;It&nbsp;For?</a> &nbsp;β€’&nbsp;
26
27
  <a href="#quickstart">⚑&nbsp;Quickstart</a> &nbsp;β€’&nbsp;
27
- <a href="#triage-guidelines">🎯&nbsp;Triage&nbsp;Guidelines</a> &nbsp;β€’&nbsp;
28
+ <a href="#triage-guidelines">🎯&nbsp;Triage</a> &nbsp;β€’&nbsp;
28
29
  <a href="#matrix">πŸ“Š&nbsp;Matrix</a>
29
30
  <br/>
31
+ <a href="#configuration">βš™οΈ&nbsp;Configuration</a> &nbsp;β€’&nbsp;
30
32
  <a href="#architecture">πŸ›οΈ&nbsp;Architecture</a> &nbsp;β€’&nbsp;
31
33
  <a href="#cli-docs">πŸ› οΈ&nbsp;CLI&nbsp;Docs</a> &nbsp;β€’&nbsp;
34
+ <a href="#providers">πŸ”Œ&nbsp;Providers</a> &nbsp;β€’&nbsp;
32
35
  <a href="#roadmap">πŸ—ΊοΈ&nbsp;Roadmap</a>
33
36
  </p>
34
37
 
@@ -59,25 +62,59 @@
59
62
 
60
63
  <br/>
61
64
 
65
+ <a id="who-is-it-for"></a>
66
+ ## πŸ‘₯ Who Is It For?
67
+
68
+ Whether you are trying your first AI coding session or running enterprise monorepo swarms, `jules-orchestrator-kit` scales with your workflow:
69
+
70
+ <br/>
71
+
72
+ | Role | Primary Value Proposition | Key Commands |
73
+ | :--- | :--- | :--- |
74
+ | **🌱 Beginners & Solo Developers** | Safely experiment with AI agents without risking broken code, leaked API keys, or ruined git history. | `agentctl init`<br/>`agentctl task create` |
75
+ | **πŸ“¦ Single-Repo Maintainers** | Automate bug fixes, dependency updates, and PR reviews with automated OODA test verification. | `agentctl gate`<br/>`agentctl queue` |
76
+ | **πŸ—οΈ Monorepo Engineering Teams** | Isolate subproject verification (`backend/`, `frontend/`, `cli/`) so agent edits never thrash global test suites. | `agentctl swarm`<br/>`agentctl lock` |
77
+ | **πŸ›‘οΈ Platform & Security Engineers** | Enforce zero-trust security policies, pre-commit secret scrubbing, and strict 75 KB diff payload limits. | `agentctl doctor`<br/>`agentctl dashboard` |
78
+
79
+ <br/>
80
+
81
+ ---
82
+
83
+ <br/>
84
+
62
85
  ## 🎯 Why `jules-orchestrator-kit`?
63
86
 
64
87
  Autonomous coding agents can write software at 100Γ— human speedβ€”but unconstrained agents introduce silent regressions, leak API keys, hallucinate test assertions, and thrash shared monorepos.
65
88
 
66
89
  `jules-orchestrator-kit` provides the missing **Safety, Orchestration, and Verification Kernel** for high-reliability AI agent deployments:
67
90
 
68
- * **πŸ”’ Zero Runtime Dependencies:** Built exclusively on Node.js 20+ built-ins (`node:fs`, `node:child_process`, `node:crypto`, `node:path`, `node:http`, `node:test`). Zero third-party npm packages mean zero supply-chain CVE risk.
91
+ * **πŸ”₯ Warm Multi-Turn Session Resumption (`v0.31.0`):** Streams OODA repair prompts directly into active Google Jules session streams via `POST /v1alpha/sessions/{id}:sendMessage`, saving 60–80% context tokens while preserving reasoning context.
92
+
93
+ * **πŸ§ͺ Automated TDD Red-to-Green Harness (`agentctl test-gen`):** Scaffolds falsifiable unit tests from bug specs, verifies **RED** failure state, locks the test file in `scope.deny`, and tasks Jules with making it pass (**GREEN** state).
94
+
95
+ * **πŸ›‘οΈ 1-Click Atomic Git Checkpoint & Rollback (`agentctl rollback`):** Snapshots working tree state, git diffs, and stashes before every session, enabling instant 1-command git restoration.
96
+
97
+ * **🌐 Verification Sandbox & SSR Hydration Prober (`verify.server`):** Executes deterministic `setup`/`teardown` hooks for databases and boots dev servers to intercept React/Next.js SSR hydration panics before approving PRs.
98
+
99
+ * **⚑ AST Blast-Radius Selective Testing:** Traverses file import dependency graphs to execute only affected downstream test suites, cutting monorepo test latency from minutes to milliseconds.
100
+
101
+ * **🚨 Asynchronous HITL Escalation Bridge (`agentctl escalate`):** Dispatches Slack & Discord webhook alerts when Jules needs feedback, allowing engineers to unblock agents asynchronously via `agentctl resume <id> --response "<reply>"`.
102
+
103
+ * **πŸ”’ Zero Runtime Dependencies:** Built exclusively on Node.js 20+ built-ins (`node:fs`, `node:child_process`, `node:crypto`, `node:path`, `node:http`, `node:tty`, `node:test`). Zero third-party npm packages mean zero supply-chain CVE risk.
69
104
 
70
105
  * **πŸ›‘οΈ Fail-Closed Security Gatekeeper:** Unconditionally evaluates explicit Deny rules *before* Allow rules, redacts high-entropy secrets and PII from dry-runs and git diffs, and rejects PRs exceeding the 75 KB Diff Payload governor.
71
106
 
72
107
  * **πŸ”„ Autonomous OODA Self-Healing:** Captures test stderr/stdout, normalizes failure fingerprints, and feeds structured error contexts back into repair iterations (up to 3 automatic attempts) before human escalation.
73
108
 
109
+ * **πŸ’» Native Interactive UX & Command Palette (`v0.30.0`):** Features a zero-dependency full-screen Terminal Engine (`capabilities`, `key-decoder`, `renderer`, `layout`, `widgets`), interactive diagnostic matrix (`agentctl doctor`), task queue/swarm managers (`agentctl queue`, `agentctl swarm`), and a searchable Command Palette.
110
+
74
111
  * **🌐 Universal Polyglot Spine:** Natively auto-detects 24+ tech stacks (PHP/Laravel/WordPress, .NET/C#, Python, Go, Rust, C/C++, Flutter/Swift, Node/Deno/Bun) and transparently wraps verification suites in Docker Compose or Devcontainer sandboxes.
75
112
 
76
113
  * **πŸ“‚ Scoped Monorepo Boundary Resolver:** Statically maps changed files up directory ancestry to invoke isolated subshell test suites (`(cd backend && pytest) && (cd cli && cargo test)`), eliminating global test thrashing.
77
114
 
78
115
  * **πŸš€ Zero-Test Bootstrapping (`agentctl bootstrap`):** Synthesizes deterministic syntax-check and smoke-test verification oracles for untested legacy repositories so agents always operate against a falsifiable feedback loop.
79
116
 
80
- * **πŸ“ˆ Proven Scale & Reliability:** Empirically tested with **224 unit tests across 55 suites passing in < 1.2s**, supporting 300+ daily agent sessions per repository.
117
+ * **πŸ“ˆ Proven Scale & Reliability:** Empirically tested with **368 unit tests across 52 suites passing in < 3.0s**, supporting 300+ daily agent sessions per repository.
81
118
 
82
119
  <br/>
83
120
 
@@ -124,11 +161,12 @@ To ensure maximum merge success, dispatch tasks according to our deterministic t
124
161
  <a id="matrix"></a>
125
162
  ## πŸ“Š Feature Comparison Matrix
126
163
 
127
- | Dimension | Raw Agent Execution (No Orchestrator) | Standard CI/CD Pipelines | `jules-orchestrator-kit` (v0.27+) |
164
+ | Dimension | Raw Agent Execution (No Orchestrator) | Standard CI/CD Pipelines | `jules-orchestrator-kit` (v0.30.0) |
128
165
  | :--- | :--- | :--- | :--- |
129
166
  | **Self-Healing Loop** | ❌ None (Crashes on test error) | ❌ None (Fails build; notifies human) | βœ… **Autonomous OODA Loop** (Max 3 repair turns with error fingerprinting) |
167
+ | **Interactive UX Engine**| ❌ Raw unformatted CLI dumps | ❌ Non-interactive log outputs | βœ… **Native TUI Engine & Command Palette** (Zero-dependency alternate-screen TUI) |
130
168
  | **Scope Isolation** | ❌ None (Can modify CI files or lockfiles) | 🟑 Post-commit branch rules only | βœ… **Fail-Closed Scope Guard** (Deny-first evaluation; blocks protected paths) |
131
- | **Polyglot Stack Detection**| ❌ Manual prompt instructions | 🟑 Hardcoded YAML workflow steps | βœ… **Universal 24+ Stack Detector** (`src/stack-detector.mjs`) |
169
+ | **Polyglot Stack Detection**| ❌ Manual prompt instructions | 🟑 Hardcoded YAML workflow steps | βœ… **Universal 24+ Stack Detector** (`src/config.mjs`) |
132
170
  | **Flaky Test Quarantine** | ❌ Fails session randomly | ❌ Breaks CI pipeline randomly | βœ… **Wilson-Score Statistical Quarantine** (Oscillation β‰₯ 0.40 quarantined automatically) |
133
171
  | **Monorepo Scoping** | ❌ Runs full global test suite | 🟑 Requires custom Nx/Turbo scripting | βœ… **Scoped Subshell Boundary Resolver** (`resolveWorkspaceBoundary`) |
134
172
  | **Zero-Test Bootstrapping**| ❌ Halts without verification oracle | ❌ Fails build if no tests exist | βœ… **Instant Oracle Synthesis** (`php -l`, `compileall`, `dotnet build`, `tsc`, `smoke`) |
@@ -148,62 +186,42 @@ To ensure maximum merge success, dispatch tasks according to our deterministic t
148
186
  <br/>
149
187
 
150
188
  <a id="quickstart"></a>
151
- ## ⚑ Universal 30-Second Quickstart (Zero to Verified PR)
189
+ ## ⚑ Guided Quickstart (Audit β†’ Author β†’ Verify)
152
190
 
153
- Get from zero to an autonomously verified GitHub Pull Request across any software ecosystem in 30 seconds.
191
+ Get started with a safe 3-step workflow across any repository:
154
192
 
155
193
  <br/>
156
194
 
157
- ### 1️⃣ Node.js / TypeScript (npm, pnpm, yarn, bun, deno)
195
+ ### Step 1: Security & Scope Gate Audit
196
+ Audit your current working tree or branch for secret leaks, protected path violations, and verification readiness:
158
197
  ```bash
159
- # Dispatch a scoped task; auto-detects package.json / tsconfig.json and runs type-checked tests
160
- npx jules-orchestrator-kit dispatch --title "Add rate limiting to API router" \
161
- --prompt "Implement IP-based token-bucket rate limiting in src/router.ts with unit tests."
198
+ # Run security, secret scanning, and scope gate audit without modifying files
199
+ npx jules-orchestrator-kit gate --mode working-tree
162
200
  ```
163
201
 
164
202
  <br/>
165
203
 
166
- ### 2️⃣ Python / FastAPI / Django (pytest, pyproject.toml)
204
+ ### Step 2: Author a Scoped Task
205
+ Launch the interactive task authoring wizard or create a task via CLI:
167
206
  ```bash
168
- # Bootstrap zero-test or legacy Python repo, then dispatch task
169
- npx jules-orchestrator-kit bootstrap --force
170
- npx jules-orchestrator-kit dispatch --title "Add OAuth2 JWT validation" \
171
- --prompt "Add JWT bearer authentication middleware to backend/api/auth.py and verify via pytest."
172
- ```
173
-
174
- <br/>
207
+ # Interactive authoring wizard with secret scrubbing & verification probes
208
+ npx jules-orchestrator-kit task create
175
209
 
176
- ### 3️⃣ PHP / Laravel / WordPress (Docker Compose + PHPUnit/Pest)
177
- ```bash
178
- # Auto-detects docker-compose.yml and wraps test commands in `docker compose exec -T app ...`
179
- npx jules-orchestrator-kit dispatch --title "Upgrade PHP 8.3 type annotations" \
180
- --prompt "Add strict type hints to all repository classes in app/Repositories/."
181
- ```
182
-
183
- <br/>
184
-
185
- ### 4️⃣ .NET / C# Enterprise (*.sln, *.csproj)
186
- ```bash
187
- # Auto-detects .sln / .csproj and runs `dotnet test --no-restore --nologo`
188
- npx jules-orchestrator-kit dispatch --title "Implement OrderService caching" \
189
- --prompt "Add IMemoryCache caching to OrderService.cs with xUnit coverage."
210
+ # Or dispatch directly with explicit flags
211
+ npx jules-orchestrator-kit task create --title "Fix authentication token expiration" \
212
+ --prompt "Fix JWT expiration check in src/auth.mjs and verify with npm test."
190
213
  ```
191
214
 
192
215
  <br/>
193
216
 
194
- ### 5️⃣ UI / Frontend E2E (Playwright)
217
+ ### Step 3: Local Verification & Queue Execution
218
+ Inspect diagnostics and execute queued tasks:
195
219
  ```bash
196
- # Dispatch frontend task verified via headless Playwright E2E tests
197
- npx jules-orchestrator-kit dispatch --title "Add Dark Mode Toggle Component" \
198
- --prompt "Create ThemeToggle component in src/components/ThemeToggle.tsx and verify via npx playwright test."
199
- ```
200
-
201
- <br/>
220
+ # Run interactive diagnostic matrix
221
+ npx jules-orchestrator-kit doctor --interactive
202
222
 
203
- ### 6️⃣ Polyglot Monorepo (FastAPI + React + Rust CLI)
204
- ```bash
205
- # Run a parallel worktree swarm; changed files automatically route to scoped subproject tests
206
- npx jules-orchestrator-kit swarm
223
+ # Process pending task queue in isolated worktrees
224
+ npx jules-orchestrator-kit queue --interactive
207
225
  ```
208
226
 
209
227
  <br/>
@@ -214,7 +232,7 @@ npx jules-orchestrator-kit swarm
214
232
  <br/>
215
233
 
216
234
  ```
217
- Ecosystems Natively Supported by src/stack-detector.mjs:
235
+ Ecosystems Natively Supported by src/config.mjs:
218
236
  β”œβ”€β”€ PHP / Laravel / WordPress (composer.json, phpunit.xml, pest.php, artisan, wp-cli.yml)
219
237
  β”œβ”€β”€ .NET / C# / F# (*.sln, *.csproj, *.fsproj, global.json)
220
238
  β”œβ”€β”€ Mobile / Dart / Flutter (pubspec.yaml)
@@ -242,6 +260,46 @@ Ecosystems Natively Supported by src/stack-detector.mjs:
242
260
 
243
261
  <br/>
244
262
 
263
+ <a id="configuration"></a>
264
+ ## βš™οΈ Configuration Reference (`.agent/config.yml`)
265
+
266
+ `jules-orchestrator-kit` auto-detects stack defaults, but allows explicit overrides through `.agent/config.yml`:
267
+
268
+ ```yaml
269
+ # .agent/config.yml β€” Universal Orchestrator Configuration
270
+
271
+ version: 1
272
+ provider: "jules" # Provider key ("jules" | "claude-code" | "local")
273
+ baseBranch: "main" # Default target base branch
274
+ branchPrefix: "agent/" # Prefix for task branches
275
+
276
+ # Verification commands (auto-detected by Stack Oracle if omitted)
277
+ verify:
278
+ test: "npm test"
279
+ build: "npm run build"
280
+
281
+ # Scope protection rules (Deny-first evaluation)
282
+ scope:
283
+ deny:
284
+ - ".github/**"
285
+ - ".agent/config.yml"
286
+ - "keys/**"
287
+
288
+ # Operational limits & governors
289
+ limits:
290
+ diffKb: 75 # 75 KB Diff Payload Governor limit
291
+ promptKb: 50 # Maximum prompt payload size
292
+ dailyTasks: 300 # Daily task session quota limit
293
+ repairAttempts: 3 # Maximum OODA repair iterations
294
+ concurrency: 1 # Worker slot concurrency limit
295
+ ```
296
+
297
+ <br/>
298
+
299
+ ---
300
+
301
+ <br/>
302
+
245
303
  <a id="architecture"></a>
246
304
  ## πŸ›οΈ System Architecture & Visual Diagrams
247
305
 
@@ -264,36 +322,14 @@ Every task dispatched to `jules-orchestrator-kit` executes within an immutable,
264
322
 
265
323
  <br/>
266
324
 
267
- ```mermaid
268
- flowchart TD
269
- A["πŸ“© Task Envelope"] --> B["1. Validate Scope & Base Freshness"]
270
- B --> C["2. Create Isolated Git Worktree & VFS Lock"]
271
- C --> D["3. Dispatch Task to Google Jules / LLM"]
272
- D --> E["4. Execute Scoped Verification Gate<br/><code>detectPolyglotStack().testCmd</code>"]
273
-
274
- E -->|PASS| F["5. Security Audit<br/><i>Redact Secrets, Diff < 75KB</i>"]
275
- F --> G["7. Rebase & Open PR<br/><code>git rebase main && gh pr create</code>"]
276
-
277
- E -->|FAIL| H["6. Fingerprint Stderr & Flaky Verdict"]
278
- H -->|Oscillation >= 0.40| I["🚨 Quarantined Test<br/><i>Exit Code 8</i>"]
279
- H -->|Normal Failure| J["πŸ”„ Attempt OODA Repair Turn<br/><i>Max 3 Retries; Exit 4 on Exhaust</i>"]
280
- J --> D
281
- ```
282
-
283
- <br/>
284
-
285
325
  ### 2. Polyglot Monorepo Scoped Execution Engine
286
326
  In monorepos containing multiple languages, `resolveWorkspaceBoundary(changedFiles)` traverses directory ancestry to isolate verification to affected subprojects:
287
327
 
288
- ```mermaid
289
- flowchart TD
290
- A["πŸ“ Changed Files<br/><code>['backend/api/main.py', 'cli/src/main.rs']</code>"] --> B{"Check Shared Triggers?<br/><i>docker-compose.yml, openapi.yaml</i>"}
291
- B -->|None Changed| C["Traverse Directory Ancestry"]
292
- C --> D1["<code>backend/api/main.py</code> β†’ <code>backend/pyproject.toml</code><br/><i>(Python Stack)</i>"]
293
- C --> D2["<code>cli/src/main.rs</code> β†’ <code>cli/Cargo.toml</code><br/><i>(Rust Stack)</i>"]
294
- D1 --> E["Synthesize POSIX Subshell Verification Plan<br/><code>(cd backend && pytest) && (cd cli && cargo test)</code>"]
295
- D2 --> E
296
- ```
328
+ <br/>
329
+
330
+ <p align="center">
331
+ <img src="docs/assets/monorepo-resolver.svg" alt="Polyglot Monorepo Scoped Boundary Resolver" width="100%" />
332
+ </p>
297
333
 
298
334
  <br/>
299
335
 
@@ -334,19 +370,59 @@ Native stdio server exposing task dispatch, gate verification, and risk auditing
334
370
  | :--- | :--- | :--- | :--- |
335
371
  | `init` | `agentctl init [--interactive] [--tier pro]` | Interactive onboarding wizard & stack oracle inspector generating `.agent/config.yml`. | `0` (Created) |
336
372
  | `task create` | `agentctl task create [--title <t>] [--prompt <p>]` | Interactively authors & scopes falsifiable task envelopes with secret scrubbing & preflight gate checks. | `0` (Queued), `1` (Unfalsifiable / Secret leak) |
373
+ | `task optimize` | `agentctl task optimize "<prompt>" [--fix] [--json]` | Linter & optimizer scoring prompt falsifiability (0–100), fixing typos, and checking scope violations. | `0` (Scored/Fixed) |
374
+ | `test-gen` | `agentctl test-gen --title <t> --spec <s> [--run]` | Scaffolds falsifiable unit tests, verifies **RED** failure state, and locks test in `scope.deny`. | `0` (Scaffolded/Red) |
375
+ | `rollback` | `agentctl rollback [sessionId \| --latest]` | Restores exact commit, uncommitted files, and cleans orphan task worktrees from pre-flight checkpoints. | `0` (Restored), `1` (Error) |
376
+ | `resume` | `agentctl resume <sessionId> --response "<reply>"` | Streams engineer response back into active Google Jules warm session context window. | `0` (Resumed), `1` (Error) |
337
377
  | `dispatch` | `agentctl dispatch --title <t> --prompt <p>` | Dispatches a single task to an AI agent in an isolated worktree. | `0` (Success), `1` (Arg error), `2` (429 Rate limit), `3` (Scope deny), `4` (OODA exhausted), `5` (Diff > 75KB), `6` (Secret leak) |
378
+ | `doctor` | `agentctl doctor [--interactive] [--fix safe]` | Diagnostic DAG check runner & automated transactional repair planner. | `0` (Healthy) |
379
+ | `queue` | `agentctl queue [--interactive] [--json]` | Consumes, inspects, and executes task envelopes in `.agent/jules-queue/` (supports `--json`). | `0` (Complete) |
380
+ | `swarm` | `agentctl swarm [--interactive] [--json]` | Runs parallel multi-agent swarm across worker slots with process PID liveness detection (supports `--json`). | `0` (Complete) |
338
381
  | `scan` | `agentctl scan` | Scans codebase for TODO/FIXME annotations to seed task authoring. | `0` (Scanned) |
339
382
  | `review-repair`| `agentctl review-repair <pr-comments.json>`| Parses GitHub PR review comments and synthesizes actionable OODA repair tasks. | `0` (Parsed), `1` (Missing file) |
340
383
  | `dashboard` | `agentctl dashboard [port]` | Starts zero-dependency local HTTP telemetry and audit visualizer dashboard. | `0` (Running) |
341
- | `gate` / `audit`| `agentctl gate --mode working-tree --json` | Runs security, secret scanning, and verification gate against working tree or branch. | `0` (Approved), `3` (Scope violation), `5` (Diff limit), `6` (Secret leak) |
384
+ | `gate` / `audit`| `agentctl gate --mode working-tree [--json]` | Runs security, secret scanning, and verification gate against working tree or branch (supports `--json`). | `0` (Approved), `3` (Scope violation), `5` (Diff limit), `6` (Secret leak) |
342
385
  | `bootstrap` | `agentctl bootstrap [--force] [--json]` | Inspects an untested repository and synthesizes `.agent/config.yml` with a zero-test verification oracle (`php -l`, `compileall`, `dotnet build`, `tsc`, `smoke`). | `0` (Bootstrapped / Existing) |
343
- | `queue` | `agentctl queue` | Consumes and executes pending markdown task envelopes in `.agent/queue/`. | `0` (Complete) |
344
- | `swarm` | `agentctl swarm` | Runs parallel multi-agent swarm across queued tasks with token-bucket concurrency. | `0` (Complete) |
345
- | `doctor` | `agentctl doctor` | Diagnostic inspect: displays detected stack, container wrapper, test command, and daily session budget. | `0` (Healthy) |
346
386
  | `lock` | `agentctl lock <acquire\|release\|status>`| Manages VFS mutex locks for multi-agent non-overlapping file ownership. | `0` (Locked/Released), `1` (Conflict) |
347
387
  | `clean` | `agentctl clean` | Prunes stale git worktrees, lockfiles, and temporary ledgers. | `0` (Clean) |
348
388
  | `mcp` | `agentctl mcp` | Starts stdio Model Context Protocol (MCP) server for tool integration. | `0` / Stdio stream |
349
- | `version` | `agentctl version` | Outputs orchestrator kit semantic version (`v0.29.0`). | `0` |
389
+ | `mcp init` | `agentctl mcp init [--target cursor\|vscode\|claude\|all]` | 1-click scaffolding for Cursor (`.cursor/mcp.json`), VS Code tasks (`tasks.json`), and Claude Desktop. | `0` (Scaffolded) |
390
+ | `version` | `agentctl version` | Outputs orchestrator kit semantic version (`v0.30.0`). | `0` |
391
+
392
+ <br/>
393
+
394
+ ---
395
+
396
+ <br/>
397
+
398
+ <a id="providers"></a>
399
+ ## πŸ”Œ Provider Integration & SDK Usage
400
+
401
+ `jules-orchestrator-kit` supports standard AI agent platforms and programmatic failover routing:
402
+
403
+ ### 1. Google Jules Native Integration
404
+ Dispatch tasks using canonical task envelopes or the Google Jules REST v1alpha API.
405
+
406
+ ### 2. Multi-Provider Failover SDK (`createFailoverProvider`)
407
+ Programmatically configure ordered provider failover (e.g. falling back to secondary providers on HTTP 429 rate limits):
408
+
409
+ ```javascript
410
+ import { createFailoverProvider, loadConfig } from "jules-orchestrator-kit";
411
+
412
+ const config = loadConfig(process.cwd());
413
+ const provider = createFailoverProvider(["jules", "claude-code"], config);
414
+
415
+ const result = await provider.dispatch(
416
+ { title: "Repair failing tests", prompt: "Fix the failing test suite." },
417
+ { root: process.cwd() }
418
+ );
419
+ ```
420
+
421
+ ### 3. Model Context Protocol (MCP) Server
422
+ Expose orchestrator gates and queue controls over stdio to client tools (Antigravity, Claude, Cursor):
423
+ ```bash
424
+ npx jules-orchestrator-kit mcp
425
+ ```
350
426
 
351
427
  <br/>
352
428
 
@@ -355,15 +431,24 @@ Native stdio server exposing task dispatch, gate verification, and risk auditing
355
431
  <br/>
356
432
 
357
433
  <a id="roadmap"></a>
358
- ## πŸ—ΊοΈ v0.29+ Feature Roadmap
434
+ ## πŸ—ΊοΈ Feature Roadmap & Release History
359
435
 
360
- | Feature | Module / Command | Architectural Blueprint | Target Release |
436
+ | Feature | Module / Command | Architectural Description | Target Release |
361
437
  | :--- | :--- | :--- | :---: |
438
+ | **Warm Session Resumption & PR Bundler** | `src/provider.mjs`, `src/engine.mjs` | Multi-turn warm session context streaming via `POST /v1alpha/sessions/{id}:sendMessage` & evidence PR descriptions. | **v0.31.0** *(Shipped)* |
439
+ | **TDD Harness & Prompt Falsifiability Linter** | `agentctl test-gen`, `agentctl task optimize` | Automated RED-state test generator, `scope.deny` test locking, and prompt testability linter with fuzzy path resolution. | **v0.31.0** *(Shipped)* |
440
+ | **Atomic Git Checkpoint & Rollback** | `agentctl rollback` (`src/ops/checkpoint.mjs`) | Pre-flight git HEAD/stash snapshotting, atomic rollback restoration, and 10-session pruning rotation. | **v0.31.0** *(Shipped)* |
441
+ | **Verification Sandbox & SSR Hydration Prober** | `verify.server`, `verify.setup`/`teardown` | Isolated process group dev server probing, Next.js/React SSR panic detection, and deterministic DB hooks. | **v0.31.0** *(Shipped)* |
442
+ | **AST Selective Testing & Escalation Bridge** | `src/dag-engine.mjs`, `agentctl escalate` | Downstream import test resolution, Slack/Discord webhook alerts, and async `agentctl resume` unblocking. | **v0.31.0** *(Shipped)* |
443
+ | **IDE Native MCP Config Scaffolder** | `agentctl mcp init` (`src/ops/ide-scaffold.mjs`) | 1-click scaffolding for Cursor (`.cursor/mcp.json`), VS Code tasks (`tasks.json`), and Claude Desktop. | **v0.31.0** *(Shipped)* |
444
+ | **Interactive UX Engine & TUI Engine** | `src/ux/` (`capabilities`, `key-decoder`, `renderer`, `layout`, `widgets`) | Zero-dependency terminal capabilities detector, sequence key decoder, virtual frame renderer, and widgets. | **v0.30.0** *(Shipped)* |
445
+ | **Guided Diagnostics & Transactional Core** | `src/ops/` (`doctor-registry`, `doctor-planner`, `transaction`, `receipts`) | Diagnostic check DAG (`runDoctorChecks`), pure fix planner (`planDiagnosticFixes`), and transactional executor with rollback. | **v0.30.0** *(Shipped)* |
446
+ | **Interactive Queue & Swarm Manager** | `src/ux/`, `src/ops/` (`queue-model`, `swarm-model`, `task-actions`, `swarm-actions`) | Task sidecar state machine, queue snapshot builder, PID liveness reconciler, task actions, and swarm actions. | **v0.30.0** *(Shipped)* |
447
+ | **Command Registry & Command Palette** | `src/ops/command-registry.mjs`, `src/ux/palette.mjs` | Single-source command descriptor registry (`COMMAND_REGISTRY`), `--help` string formatter, fuzzy search filter, and command palette. | **v0.30.0** *(Shipped)* |
362
448
  | **Onboarding & Stack Oracle Wizard** | `agentctl init --interactive` (`src/wizard-init.mjs`) | Zero-dependency interactive CLI wizard auto-detecting verification oracles, quota tiers, and preset workflows. | **v0.29.0** *(Shipped)* |
363
449
  | **Guided Task Authoring Subsystem** | `agentctl task create` (`src/wizard-task.mjs`) | Guided task authoring with TODO candidate harvesting, Shannon entropy secret scrubbing, and guardrail footer synthesis. | **v0.29.0** *(Shipped)* |
364
- | **PR Review Auto-Remediation Loop** | `agentctl review-repair` (`src/review-repair.mjs`) | Ingests GitHub PR review comments (`CHANGES_REQUESTED`), extracts line/file context, and dispatches automated OODA repair turns until reviewer comments are resolved. | **v0.27.0** *(Shipped)* |
365
- | **Multi-Provider Failover Router** | `createFailoverProvider` (`src/provider.mjs`) | Ordered router (`["jules", "claude-code", "local-mcp"]`) that seamlessly falls back to secondary LLMs on HTTP 429 rate limits or 5xx service unavailability. | **v0.27.0** *(Shipped)* |
366
- | **Telemetry & Audit Web Dashboard**| `agentctl dashboard` (`src/dashboard.mjs`) | Zero-dependency local HTTP server displaying real-time DAG execution graphs, Wilson-Score flaky test ledgers, and SHA-256 telemetry chains. | **v0.27.0** *(Shipped)* |
450
+ | **P0 Remediation & Safety Alignment** | Queue, Task Envelope & Secrets (`src/wizard-task.mjs`) | Canonical queue path alignment (`.agent/jules-queue/`), path traversal guards, atomic writes, multiline secret scans, and JSON headers. | **v0.29.1** *(Shipped)* |
451
+ | **PR Review Auto-Remediation Loop** | `agentctl review-repair` (`src/review-repair.mjs`) | Ingests GitHub PR review comments (`CHANGES_REQUESTED`), extracts line/file context, and dispatches automated OODA repair turns. | **v0.27.0** *(Shipped)* |
367
452
 
368
453
  <br/>
369
454
 
@@ -371,12 +456,10 @@ Native stdio server exposing task dispatch, gate verification, and risk auditing
371
456
 
372
457
  <br/>
373
458
 
374
- ## πŸ“– Recipes, Documentation & Prior Art
459
+ ## πŸ“– Documentation & Architecture
375
460
 
461
+ - [**System Architecture & Pipeline Overview**](./docs/architecture.md) β€” Comprehensive technical sequence diagram and control plane architecture.
376
462
  - [**Google Jules Official Documentation**](https://jules.google) β€” Official platform overview and API specifications for Google Jules.
377
- - [**Onboarding & Task Wizard Specification**](./docs/ONBOARDING_AND_TASK_WIZARD_SPEC.md) β€” Normative architecture specification for zero-dependency TUI, stack oracle, and guided task authoring.
378
- - [**Universal Polyglot Architecture & Zero-Test Specification**](./docs/UNIVERSAL_POLYGLOT_ARCHITECTURE.md) β€” Comprehensive technical report on boundary resolution, OODA math, and B2B workflows.
379
- - [**v0.27.0 Architectural Audit & Platform Evolution**](./docs/V0.27_ARCHITECTURAL_AUDIT_AND_EVOLUTION.md) β€” PR review remediation, failover router, and local dashboard specs.
380
463
  - [**Examples & Task Envelope Recipes**](./EXAMPLES.md) β€” Production YAML and Markdown task envelopes.
381
464
  - [**Changelog**](./CHANGELOG.md) β€” Full release history and migration guides.
382
465
 
package/bin/agentctl.mjs CHANGED
@@ -2,7 +2,7 @@
2
2
 
3
3
  import { parseArgs } from "node:util";
4
4
  import { readFileSync, existsSync, readdirSync } from "node:fs";
5
- import { join } from "node:path";
5
+ import { join, resolve } from "node:path";
6
6
  import { loadConfig, resolveRoot, detectStack, bootstrapZeroTestRepo } from "../src/config.mjs";
7
7
  import { gate, dispatch, run, isTaskFile } from "../src/engine.mjs";
8
8
  import { acquireLock, releaseLock, lockStatus, checkDailyBudget, getQueueDir } from "../src/state.mjs";
@@ -14,7 +14,7 @@ const command = args[0];
14
14
 
15
15
  function printHelp() {
16
16
  console.log(`
17
- πŸš€ agentctl v0.29.1 β€” Universal Agent Orchestrator & Safety Gatekeeper
17
+ πŸš€ agentctl v0.30.0 β€” Universal Agent Orchestrator & Safety Gatekeeper
18
18
 
19
19
  Usage: agentctl <command> [options]
20
20
 
@@ -32,6 +32,11 @@ Commands:
32
32
  dashboard Start local HTTP telemetry and audit dashboard
33
33
  init Scaffold .agent/ config and run onboarding wizard
34
34
  task create Interactively author and scope a Jules task envelope
35
+ task optimize Linter & optimizer for Jules task prompts (--fix, --json)
36
+ test-gen Scaffold & run automated TDD Red-to-Green test cycle (--run)
37
+ mcp init Scaffold IDE integration config (cursor | vscode | claude | all)
38
+ rollback Restore git state & working tree to atomic pre-flight checkpoint
39
+ resume Resume warm session with human response (--response "<text>")
35
40
  status Display queue and system status summary
36
41
  scan Scan codebase for TODO/FIXME task candidates
37
42
  version Output agentctl version
@@ -391,8 +396,66 @@ async function main() {
391
396
  console.log(` Auto-PR : ${res.plan.flags.autoPr}`);
392
397
  }
393
398
  process.exit(0);
399
+ } else if (subCommand === "optimize") {
400
+ const { values, positionals } = parseArgs({
401
+ args: args.slice(2),
402
+ options: {
403
+ fix: { type: "boolean", short: "f" },
404
+ file: { type: "string" },
405
+ dir: { type: "string", short: "d" },
406
+ json: { type: "boolean", short: "j" },
407
+ "verify-cmd": { type: "string", short: "v" },
408
+ },
409
+ allowPositionals: true,
410
+ });
411
+
412
+ const { scorePromptFalsifiability, optimizeTaskPrompt } = await import("../src/task-optimizer.mjs");
413
+ const targetDir = values.dir ? resolve(values.dir) : root;
414
+ let promptText = positionals.join(" ");
415
+
416
+ if (values.file) {
417
+ if (existsSync(values.file)) {
418
+ promptText = readFileSync(values.file, "utf-8");
419
+ } else {
420
+ console.error(`Error: File '${values.file}' does not exist.`);
421
+ process.exit(1);
422
+ }
423
+ }
424
+
425
+ if (values.fix) {
426
+ const opt = optimizeTaskPrompt(promptText, { rootDir: targetDir, verifyCmd: values["verify-cmd"] });
427
+ if (values.json) {
428
+ console.log(JSON.stringify(opt, null, 2));
429
+ } else {
430
+ console.log(opt.optimizedPrompt);
431
+ }
432
+ process.exit(0);
433
+ }
434
+
435
+ const analysis = scorePromptFalsifiability(promptText, { rootDir: targetDir, verifyCmd: values["verify-cmd"] });
436
+ if (values.json) {
437
+ console.log(JSON.stringify(analysis, null, 2));
438
+ } else {
439
+ console.log(`\n🎯 Task Prompt Falsifiability Analysis`);
440
+ console.log(`--------------------------------------------------`);
441
+ console.log(` Score / Grade : ${analysis.score} / 100 (${analysis.grade})`);
442
+ console.log(` Falsifiable : ${analysis.isFalsifiable ? "βœ… YES" : "❌ NO"}`);
443
+ if (analysis.oracle.command) {
444
+ console.log(` Oracle Command : "${analysis.oracle.command}" (${analysis.oracle.autoDetected ? "Auto-detected" : "User-supplied"})`);
445
+ }
446
+ if (analysis.issues.length > 0) {
447
+ console.log(`\n Issues Identified (${analysis.issues.length}):`);
448
+ analysis.issues.forEach((i) => console.log(` - [${i.type}] ${i.message} (-${i.penalty} pts)`));
449
+ }
450
+ if (analysis.suggestions.length > 0) {
451
+ console.log(`\n Suggestions for Improvement:`);
452
+ analysis.suggestions.forEach((s) => console.log(` - ${s}`));
453
+ }
454
+ console.log(`--------------------------------------------------\n`);
455
+ }
456
+ process.exit(analysis.isFalsifiable ? 0 : 1);
394
457
  } else {
395
- console.error(`Unknown task subcommand '${subCommand}'. Supported: agentctl task create`);
458
+ console.error(`Unknown task subcommand '${subCommand}'. Supported: agentctl task create, agentctl task optimize`);
396
459
  process.exit(1);
397
460
  }
398
461
  break;
@@ -427,7 +490,143 @@ async function main() {
427
490
  break;
428
491
  }
429
492
 
493
+ case "rollback": {
494
+ const { restoreCheckpoint } = await import("../src/ops/checkpoint.mjs");
495
+ const targetId = args[1] || "--latest";
496
+ try {
497
+ const res = restoreCheckpoint(targetId, { root });
498
+ console.log(`\nβœ… Git Checkpoint Restored Successfully!`);
499
+ console.log(` Session ID : ${res.id}`);
500
+ console.log(` HEAD SHA : ${res.headSha || "N/A"}`);
501
+ console.log(` RestoredAt : ${res.restoredAt}\n`);
502
+ process.exit(0);
503
+ } catch (err) {
504
+ console.error(`❌ Rollback Failed: ${err.message}`);
505
+ process.exit(1);
506
+ }
507
+ break;
508
+ }
509
+
510
+ case "resume": {
511
+ const { createProvider } = await import("../src/provider.mjs");
512
+ const sessionId = args[1];
513
+ const { values } = parseArgs({
514
+ args: args.slice(1),
515
+ options: {
516
+ response: { type: "string", short: "r" },
517
+ "dry-run": { type: "boolean", short: "d" },
518
+ json: { type: "boolean", short: "j" },
519
+ },
520
+ allowPositionals: true,
521
+ });
522
+
523
+ if (!sessionId || sessionId.startsWith("-")) {
524
+ console.error("Error: Session ID is required for agentctl resume <sessionId>.");
525
+ process.exit(1);
526
+ }
527
+
528
+ const responseText = values.response || args.slice(2).join(" ");
529
+ if (!responseText) {
530
+ console.error("Error: --response text is required to resume warm session.");
531
+ process.exit(1);
532
+ }
533
+
534
+ const provider = createProvider(config.provider || "jules", config);
535
+ try {
536
+ const res = await provider.resume(sessionId, responseText, { root, dryRun: values["dry-run"] });
537
+ if (values.json) {
538
+ console.log(JSON.stringify(res, null, 2));
539
+ } else {
540
+ console.log(`\nβœ… Warm Session Resumed Successfully!`);
541
+ console.log(` Session ID : ${res.id}`);
542
+ console.log(` Status : ${res.status}\n`);
543
+ }
544
+ process.exit(0);
545
+ } catch (err) {
546
+ console.error(`❌ Resume Failed: ${err.message}`);
547
+ process.exit(1);
548
+ }
549
+ break;
550
+ }
551
+
552
+ case "test-gen": {
553
+ const { scaffoldTddTest, runTddCycle } = await import("../src/ops/tdd-generator.mjs");
554
+ const { values } = parseArgs({
555
+ args: args.slice(1),
556
+ options: {
557
+ title: { type: "string", short: "t" },
558
+ spec: { type: "string", short: "s" },
559
+ run: { type: "boolean", short: "r" },
560
+ json: { type: "boolean", short: "j" },
561
+ },
562
+ allowPositionals: true,
563
+ });
564
+
565
+ const title = values.title || args[1] || "feature-spec";
566
+ const specText = values.spec || args.slice(2).join(" ") || "TDD requirement specification.";
567
+
568
+ try {
569
+ if (values.run) {
570
+ const res = await runTddCycle({ title, spec: specText }, { root });
571
+ if (values.json) {
572
+ console.log(JSON.stringify(res, null, 2));
573
+ } else {
574
+ console.log(`\nπŸ”΄ TDD RED Check Verified!`);
575
+ console.log(` Test File : ${res.testFile}`);
576
+ console.log(` Scope Lock : Locked into scope.deny`);
577
+ console.log(` Status : Ready for green implementation dispatch\n`);
578
+ }
579
+ } else {
580
+ const scaffold = scaffoldTddTest({ title, spec: specText }, { root });
581
+ if (values.json) {
582
+ console.log(JSON.stringify(scaffold, null, 2));
583
+ } else {
584
+ console.log(`\nπŸ§ͺ TDD Test Scaffolded Successfully!`);
585
+ console.log(` File Path : ${scaffold.relativePath}`);
586
+ console.log(` Command : ${scaffold.testCmd}\n`);
587
+ }
588
+ }
589
+ process.exit(0);
590
+ } catch (err) {
591
+ console.error(`❌ TDD Generation Failed: ${err.message}`);
592
+ process.exit(1);
593
+ }
594
+ break;
595
+ }
596
+
430
597
  case "mcp": {
598
+ const subAction = args[1];
599
+ if (subAction === "init") {
600
+ const { scaffoldIdeConfig } = await import("../src/ops/ide-scaffold.mjs");
601
+ const { values } = parseArgs({
602
+ args: args.slice(2),
603
+ options: {
604
+ target: { type: "string", short: "t", default: "all" },
605
+ json: { type: "boolean", short: "j" },
606
+ },
607
+ allowPositionals: true,
608
+ });
609
+
610
+ const target = values.target || args[2] || "all";
611
+ try {
612
+ const res = scaffoldIdeConfig(target, { root });
613
+ if (values.json) {
614
+ console.log(JSON.stringify(res, null, 2));
615
+ } else {
616
+ console.log(`\nπŸ”Œ IDE MCP Config Scaffolded Successfully!`);
617
+ console.log(` Target : ${res.target}`);
618
+ for (const item of res.results) {
619
+ console.log(` - ${item.target.toUpperCase()} : ${item.file}`);
620
+ }
621
+ console.log("");
622
+ }
623
+ process.exit(0);
624
+ } catch (err) {
625
+ console.error(`❌ IDE MCP Scaffold Failed: ${err.message}`);
626
+ process.exit(1);
627
+ }
628
+ }
629
+
431
630
  const { startMcpServer } = await import("../src/mcp.mjs");
432
631
  startMcpServer();
433
632
  break;
package/index.mjs CHANGED
@@ -85,4 +85,24 @@ export { planInit, loadPresets, runInitWizard, TIER_PROFILES, BUILTIN_PRESETS }
85
85
  // Guided Task Authoring Subsystem
86
86
  export { planTaskCreate, runTaskCreateWizard, GUARDRAIL_FOOTER } from "./src/wizard-task.mjs";
87
87
 
88
+ // Prompt Falsifiability & Task Optimizer Engine
89
+ export { scorePromptFalsifiability, optimizeTaskPrompt, levenshteinDistance, extractPathTokens } from "./src/task-optimizer.mjs";
90
+
91
+ // Atomic Git Checkpoints & Rollback
92
+ export { createCheckpoint, restoreCheckpoint, listCheckpoints, pruneCheckpoints, CheckpointError } from "./src/ops/checkpoint.mjs";
93
+
94
+ // Webhook & HITL Escalation Bridge
95
+ export { dispatchEscalation, verifySignature, parseWebhookPayload, routeWebhookEvent, createWebhookServer } from "./src/webhook.mjs";
96
+
97
+ // PR Review Evidence Bundler & Dev Server Probe
98
+ export { synthesizePrDescription, probeDevServer } from "./src/engine.mjs";
99
+
100
+ // Automated TDD Red-to-Green Harness
101
+ export { scaffoldTddTest, runTddCycle, TddError } from "./src/ops/tdd-generator.mjs";
102
+
103
+ // IDE Native MCP Scaffolder
104
+ export { scaffoldIdeConfig, IdeScaffoldError } from "./src/ops/ide-scaffold.mjs";
105
+
106
+
107
+
88
108