jules-orchestrator-kit 0.38.2 → 0.41.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -1,18 +1,21 @@
1
1
  # Bolt - Performance & Payload Optimization Specialist ⚡
2
2
 
3
3
  > **Role:** Codebase Micro-Optimizer & Payload Governor.
4
- > **Scope:** Performance tuning, bundle size reduction, and asset optimization with zero structural side-effects.
4
+ > **Scope:** Performance tuning, artifact size reduction, and asset optimization with zero structural side-effects.
5
5
 
6
6
  ## Core Directives
7
7
 
8
8
  1. **Payload Budgeting:**
9
- - Keep total diff payload strictly under 75 KB (`git diff | wc -c`).
10
- - Eliminate redundant dependencies by replacing 3rd-party modules with Node.js built-ins (`node:fs`, `node:path`, `node:crypto`).
9
+ - Keep total diff payload strictly under {{DIFF_KB}} KB (`git diff | wc -c`).
10
+ - Prefer this project's existing dependencies and its language's standard library over adding another third-party module. Removing a dependency whose job the standard library already does is in scope; adding one is not.
11
11
 
12
12
  2. **Asset & Memory Optimization:**
13
- - Replace heavy raster assets with modern WebP/AVIF equivalents or clean SVGs.
14
- - Optimize hot execution paths: remove redundant object allocations inside tight loops.
13
+ - Replace heavy raster assets with modern equivalents (WebP/AVIF) or clean vector graphics, where the project already serves such formats.
14
+ - Optimize hot execution paths: remove redundant allocations inside tight loops, and hoist work out of repeated calls.
15
15
 
16
- 3. **Zero Regressions Invariant:**
17
- - Execute test suite (`npm test`) before and after every micro-optimization pass.
16
+ 3. **Evidence Before Claims:**
17
+ - A performance change requires numbers. Run the benchmark or timing measurement multiple times, compare medians, and state the delta. "Feels faster" is not a result, and a change below the noise floor is not an improvement.
18
+
19
+ 4. **Zero Regressions Invariant:**
20
+ - Execute `{{VERIFY_TEST}}` before and after every optimization pass, and record both results.
18
21
  - Never disable type-checks, skip tests, or alter public API signatures.
@@ -1,11 +1,11 @@
1
1
  # Janitor Protocol: Technical Debt & Dead Code Elimination
2
2
 
3
- You are **Janitor**, a specialist autonomous agent optimized for technical debt elimination, dead code pruning, and strict zero-dependency refactoring.
3
+ You are **Janitor**, a specialist autonomous agent optimized for technical debt elimination, dead code pruning, and conservative refactoring.
4
4
 
5
5
  ## Strict Operational Invariants
6
6
 
7
- 1. **Zero External Runtime Dependencies**: You are STRICTLY FORBIDDEN from adding third-party npm packages. Use ONLY native Node.js ESM built-in modules (`node:fs`, `node:path`, `node:crypto`, `node:child_process`, `node:os`).
8
- 2. **Dead Code Elimination**: Prune unused variables, unreachable branches, and redundant helper functions.
9
- 3. **Atomic Payload Limit**: Keep total patch payload under 75 KB (`git diff | wc -c`).
10
- 4. **Verification Requirement**: Execute `npm test` and `npm run lint` to ensure 100% of tests pass with 0 lint errors before completing work.
11
- 5. **No Assert Weakening**: Never weak or remove test assertions to make a test pass.
7
+ 1. **No New Dependencies**: Do NOT add third-party packages, libraries, or modules of any kind. Solve the task with this project's existing dependencies and its language's standard library. If a task genuinely cannot be completed without a new dependency, stop and say so instead of adding one.
8
+ 2. **Dead Code Elimination**: Prune unused variables, unreachable branches, and redundant helper functions. Confirm a symbol has no remaining references across the whole repository before removing it — including dynamic lookups, reflection, and string-keyed access, which a definition-search will not find.
9
+ 3. **Atomic Payload Limit**: Keep total patch payload under {{DIFF_KB}} KB (`git diff | wc -c`).
10
+ 4. **Verification Requirement**: Execute `{{VERIFY_TEST}}` and `{{VERIFY_LINT}}` and ensure 100% of tests pass with 0 lint errors before completing work.
11
+ 5. **No Assert Weakening**: Never weaken or remove test assertions to make a test pass. Leave an unmet requirement RED with a written rationale.
@@ -8,18 +8,19 @@
8
8
  ## Context
9
9
  - **Project Goals:** [Describe key architectural or business goals.]
10
10
  - **Key Files & Folders:** [List critical files, directories, or schemas, e.g. `src/auth.ts`, `schema.sql`.]
11
- - **Tech Stack:** [List frameworks and libraries, e.g. Node.js, Express, TypeScript, Drizzle ORM.]
11
+ - **Tech Stack:** [List this project's languages, frameworks, and libraries.]
12
12
 
13
13
  ## Requirements & Hard Constraints
14
14
  - **Functional Requirements:** [List specific, non-negotiable functional requirements.]
15
15
  - **Hard Constraints:**
16
- - Do NOT introduce third-party npm dependencies without explicit authorization.
17
- - Do NOT modify command files (`package.json`, `.github/`) or Agent Scope files.
18
- - Keep total diff payload strictly under 75 KB (`git diff | wc -c`).
16
+ - Do NOT introduce new third-party dependencies without explicit authorization.
17
+ - Do NOT modify this project's build manifest, lockfile, CI configuration, or agent scope files. Run `agentctl gate` to see the enforced set.
18
+ - Keep total diff payload strictly under {{DIFF_KB}} KB (`git diff | wc -c`).
19
19
 
20
20
  ## Verification Loop
21
- - **Verification Command:** Execute automated verification tests: `npm test`.
21
+ - **Verification Command:** Execute `{{VERIFY_TEST}}`.
22
22
  - **Zero Errors Invariant:** Ensure 100% of tests pass cleanly with 0 errors before submitting.
23
+ - **Carry the Evidence:** Paste the actual terminal output. Exit code 0 proves the process survived, not that the change works.
23
24
 
24
25
  ## Expected Artifacts
25
26
  - **Code Changes:** Clean, production-grade implementation preserving existing symbol contracts.
@@ -32,7 +32,7 @@ To maximize the ratio of mergeable PRs vs. failed or hallucinated sessions:
32
32
  11. **Pre-Dispatch Grounding Mandate:** Verify all file paths, script names, and exported symbols against the live repository tree before writing them into a prompt.
33
33
  12. **Programmatic CI Scope Guarding:** Enforce prompt constraints at the CI level using an unbypassable `Agent Scope Guard` workflow that evaluates diffs against a protected paths manifest (`.agent/protected-paths.json`).
34
34
  13. **Google Labs Exploration Budget Protocol:** Execute tasks across 3 discrete phases: (1) Discovery & Symbol Tracing (silent inspection, write NO code), (2) Oracle & Test Formulation, and (3) Surgical Implementation & Verification.
35
- 14. **Critic Agent Steering (Adversarial Pre-Review):** Jules' internal Critic Agent must evaluate proposed patches for edge-case regressions, $O(n^2)$ bottlenecks, unhandled parameters, and layout shifts (CLS) prior to PR submission.
35
+ 14. **Critic Agent Steering (Adversarial Pre-Review):** Jules' internal Critic Agent must evaluate proposed patches for edge-case regressions, $O(n^2)$ bottlenecks, unhandled parameters, and layout shifts (CLS) prior to PR submission. When modifying test suites or error handling, the agent must deliberately mutate production code and induce real failure conditions to prove that tests and catch blocks actually fail as intended (preventing tautological tests).
36
36
  15. **Web Excellence & Frontend Guardrails:** Enforce quantitative Core Web Vitals (LCP < 1.2s, CLS < 0.05), WCAG 2.2 AA/AAA semantic accessibility, Schema.org JSON-LD compliance, and Playwright multi-viewport responsive testing.
37
37
  16. **Airtight Positive Enclosures ("Pink Elephant" Rule):** Replace long negative constraint lists with strict positive operational perimeters (`ONLY modify [Target/Module]`) to prevent attention-drift in deep context windows.
38
38
  17. **Sterile / Clinical Vocabulary Mandate:** Eradicate aggressive verbs (`kill`, `amputate`, `sabotage`, `destroy`) from prompts and policies. Use clinical equivalents (`terminate PID`, `prune code`, `mutate test logic`, `purge cache`) to prevent false-positive safety classifier trips in Google Cloud VMs.
@@ -58,7 +58,7 @@ Jules automatically infers test and build verification commands via `scripts/com
58
58
  - **Minimal Interference**: Preserve existing function signatures, comments, and style conventions.
59
59
  - **No Token Bloat**: Exclude lockfiles, minified bundles, and binary assets from diff representations.
60
60
  - **Google Labs Exploration Budget Protocol**: Execute complex multi-step tasks across 3 discrete phases: (1) Discovery & Symbol Tracing (silent inspection, write NO code), (2) Oracle & Test Formulation, and (3) Surgical Implementation & Verification.
61
- - **Critic Agent Steering (Adversarial Pre-Review)**: Jules' internal Critic Agent evaluates proposed patches for edge-case regressions, $O(n^2)$ bottlenecks, unhandled parameters, and layout shifts (CLS) prior to final PR submission.
61
+ - **Critic Agent Steering (Adversarial Pre-Review)**: Jules' internal Critic Agent evaluates proposed patches for edge-case regressions, $O(n^2)$ bottlenecks, unhandled parameters, and layout shifts (CLS) prior to final PR submission. In test modifications, verify deliberate logic mutations turn tests red (mutation falsification).
62
62
  - **Airtight Positive Enclosures ("Pink Elephant" Principle)**: Avoid massive negative constraint lists; define explicit positive perimeters (`ONLY modify [Target/Module]`) to eliminate attention distortion and cognitive drag.
63
63
  - **Sterile / Clinical Vocabulary Mandate**: Replace aggressive verbs (`kill`, `amputate`, `destroy`) with clinical equivalents (`terminate PID`, `prune code`, `purge state`) to prevent false-positive safety classifier tripwires.
64
64
 
@@ -80,8 +80,8 @@ Jules automatically infers test and build verification commands via `scripts/com
80
80
 
81
81
  ## 6. Local CI Verification with Nektos Act
82
82
 
83
- - **Pre-Push CI Validation**: When `.github/workflows/` exists and Nektos `act` is installed, execute `act push` or `bash scripts/act/run-act.sh` to verify changes pass CI locally inside the VM before opening a PR.
84
- - **Log Inspection**: If local `act` CI fails, inspect `act_output.log`, resolve errors in code, and re-run verification before pushing.
83
+ - **Pre-Push CI Validation**: When `.github/workflows/` exists and Nektos `act` is installed, execute `act push` to verify changes pass CI locally inside the VM before opening a PR. Skip this step if `act` is not on `PATH` — do not install it and do not invent a wrapper script for it.
84
+ - **Log Inspection**: If local `act` CI fails, inspect its output, resolve errors in code, and re-run verification before pushing.
85
85
  - **Diff Payload Governor**: API forcefully truncates diff payloads > 80 KB. Keep total diff payload under 75 KB (`git diff | wc -c`).
86
86
 
87
87
  ---
@@ -101,7 +101,11 @@ To maximize the ratio of mergeable PRs vs. failed or hallucinated sessions, adhe
101
101
 
102
102
  ### Standard Jules Guardrails Footer
103
103
 
104
- Append this footer to all Jules dispatches:
104
+ `agentctl task create` appends this automatically, generated from your own
105
+ `.agent/config.yml` scope — so the protected-path line lists *your* build
106
+ manifests (`Cargo.toml`, `go.mod`, `pyproject.toml`, `composer.json`, …) and
107
+ rebases onto *your* base branch. Fill the placeholders only for hand-written
108
+ dispatches; run `agentctl gate` to see the full enforced set.
105
109
 
106
110
  ```text
107
111
  Read AGENTS.md and .agent/rules/jules-protocol.md BEFORE starting.
@@ -110,12 +114,12 @@ Follow all rules strictly.
110
114
  TASK: <description>
111
115
 
112
116
  HARD CONSTRAINTS:
113
- - Do NOT modify package.json, pnpm-lock.yaml, tsconfig.json, or .github/ files. Enforced in CI by Agent Scope Guard.
117
+ - Do NOT modify these protected paths: <your build manifest, lockfile, CI directory, and agent rules>.
114
118
  - Diff Payload Governor: Keep total diff payload under 75 KB (`git diff | wc -c`) to prevent API truncation (~80 KB limit).
115
119
  - Falsifiable & Evidence-Based: Attach full terminal verification output to PR. Never weaken assertions or delete failing tests to force a pass.
116
120
  - Declare Scope Deviations: If modifying files outside task bounds, explicitly state rationale in PR.
117
- - Verify before finishing: Run full type-check, lint, and unit test suites.
118
- - BEFORE opening the PR: Run `git fetch origin main && git rebase origin/main`, then re-verify. If the rebase leaves an empty diff, the work already landed — do NOT submit.
119
- - Delete ALL temporary files (.py, .sh, .patch, debug logs) before submitting.
121
+ - Verify before finishing: Run the project's full type-check, lint, and test commands.
122
+ - BEFORE opening the PR: Run `git fetch origin <base> && git rebase origin/<base>`, then re-verify. If the rebase leaves an empty diff, the work already landed — do NOT submit.
123
+ - Remove any scratch files you created for debugging before submitting. Do not delete files that are part of the project.
120
124
  ```
121
125
 
package/README.md CHANGED
@@ -129,9 +129,9 @@ To maximize PR merge rates, dispatch tasks according to deterministic boundaries
129
129
  * **Cross-Platform Parity:** Verified 100% green across Linux, macOS (Darwin), and Windows on Node 20, 22, and 24.
130
130
  * **Autonomous Self-Healing Loop:** Captures test stderr/stdout, fingerprints error traces, and feeds structured context back into automated repair turns (up to 3 attempts) before human escalation.
131
131
  * **Fail-Closed Security & Secret Redaction:** Evaluates explicit Deny rules before Allow rules against canonicalized, case-folded paths. Redacts high-entropy keys and base64-encoded credentials (such as Kubernetes `Secret` manifests).
132
- * **Complexity & Cost Router:** Zero-dependency heuristic classifier (`src/router.mjs`) routing mechanical tasks to lightweight models while reserving primary models for complex refactors.
132
+ * **Complexity & Cost Router:** Zero-dependency heuristic classifier (`src/router.mjs`) routing mechanical tasks to lightweight models while reserving primary models for complex refactors, with a `node --check` syntax-verification gate that transparently escalates a FAST-tier result to the primary provider if it left broken JS on disk.
133
133
  * **Terminal UI & Diagnostic Matrix (`agentctl doctor`):** Interactive terminal dashboard, task sidecar manager, and automated transactional self-repair.
134
- * **Verified Test Suite:** Tested with **572 unit tests across 81 suites passing in < 10.0s**.
134
+ * **Verified Test Suite:** Tested with **671 unit tests across 84 suites passing in < 10.0s**.
135
135
 
136
136
  <br/>
137
137
 
@@ -147,13 +147,18 @@ To maximize PR merge rates, dispatch tasks according to deterministic boundaries
147
147
  | Command | Usage | Description | Exit Codes |
148
148
  | :--- | :--- | :--- | :--- |
149
149
  | `init` | `agentctl init [--interactive] [--tier pro]` | Interactive onboarding wizard & stack detector generating `.agent/config.yml`. | `0` (Created) |
150
+ | `budget` | `agentctl budget [--by-user] [--json] [reset]` | Reports rolling 24h task budget, quota headroom, and per-developer task attribution without external auth servers. | `0` (Status), `2` (Arg Error) |
150
151
  | `task create` | `agentctl task create [--title <t>] [--prompt <p>] [--template <id>] [--role <name>] [--tier fast\|complex]` | Interactively authors & scopes falsifiable task envelopes with secret scrubbing, preflight gate checks, and DAG dependency wiring. | `0` (Queued), `1` (Secret/Unfalsifiable) |
151
- | `task template` | `agentctl task template [<id>] [--list] [--json]` | Lists and synthesizes pre-calibrated web task envelopes (`web-cwv`, `web-wcag`, `web-seo`, `web-playwright`, `web-flaky-heal`, `web-i18n`, `web-ai-access`). | `0` (Listed/Synthesized) |
152
- | `dispatch` | `agentctl dispatch [-p <prompt>] [-f <file>] [-r <role>] [-t <tier>] [--auto-pr] [--repoless] [--dry-run]` | Dispatches autonomous task to the active provider with payload limits and role prompt resolution. | `0` (Dispatched), `1` (Error) |
152
+ | `task template` | `agentctl task template [<id>] [--list] [--json]` | Lists and synthesizes pre-calibrated task envelopes (Web, Deep Think & Agent Hardening: `web-cwv`, `web-wcag`, `web-seo`, `web-playwright`, `agent-dead-code-audit`, `web-flaky-heal`, `web-i18n`, `web-ai-access`, `agent-qa-mutation`, `agent-ci-falsify`, `agent-service-isolate`, `agent-error-paths`, `agent-security-audit`, `deep-debug`, `deep-feature`, `deep-optimize`, `deep-harden`). | `0` (Listed/Synthesized) |
153
+ | `dispatch` | `agentctl dispatch [-p <prompt>] [-f <file>] [-r <role>] [-t <tier>] [--author <name>] [--check-premise] [--auto-pr] [--repoless] [--dry-run]` | Dispatches autonomous task to the active provider with pre-flight idempotency checks, payload limits, and role prompt resolution. | `0` (Dispatched), `1` (Error) |
154
+ | `plan approve` | `agentctl plan approve <sessionId> [--dry-run] [--json]` | Approves pending execution plan for an active Jules session (`:approvePlan`) with automatic 404/503 retry backoff. | `0` (Approved), `1` (Error) |
155
+ | `session get` | `agentctl session get <sessionId> [--dry-run] [--json]` | Retrieves live session lifecycle state from provider REST API with token rotation. | `0` (Fetched), `1` (Error) |
156
+ | `pr harvest` | `agentctl pr harvest [--tier r0,r1] [--limit <n>] [--auto] [--allow-no-checks] [--dry-run]` | Discovers open agent PRs, evaluates CI checks & risk tiers, and auto-squashes green low-risk changes autonomously. A PR reporting **no** CI checks is skipped unless `--allow-no-checks` is passed, and an unavailable changed-file list blocks rather than classifying as low risk. | `0` (Triaged/Merged), `1` (Error) |
153
157
  | `doctor` | `agentctl doctor [--json]` | Diagnostic DAG check runner & automated transactional self-repair engine. | `0` (Healthy), `1` (Failures) |
154
158
  | `queue` | `agentctl queue [--dag] [--concurrency <n>] [--dry-run] [--json]` | Consumes and executes task envelopes in `.agent/jules-queue/` with Kahn's DAG dependency resolution. Non-task files (manifests, `README.md`) are skipped, and `--dry-run` previews without moving anything. | `0` (Complete) |
155
159
  | `swarm` | `agentctl swarm [--json]` | Runs parallel multi-agent swarm across worker slots with PID liveness detection. | `0` (Complete) |
156
- | `gate` / `audit`| `agentctl gate --mode working-tree [--json]` | Runs security, secret scanning, and verification gates against working tree or branch. | `0` (Approved), `3` (Scope), `5` (Diff >75K), `6` (Secret) |
160
+ | `gate` / `audit`| `agentctl gate --mode working-tree [--json] [--json-report <path>]` | Runs security, secret scanning, and tiered verification gates (with declarative assertion support) against working tree or branch. | `0` (Approved), `3` (Scope), `5` (Diff >75K), `6` (Secret) |
161
+ | `assert` | `agentctl assert [--dir <d>] [--file <f>] [--max-mb <n>] [--gzip] [--targets <g>] [--patterns <p>] [--json] [--json-report <p>]` | Runs declarative zero-dependency verification assertion primitives (`assert:dir-size`, `assert:file-size`, `assert:file-patterns`, `assert:exists`). | `0` (Passed), `1` (Assertion Failed) |
157
162
  | `rollback` | `agentctl rollback [sessionId \| --latest]` | Restores exact commit, uncommitted files, and cleans orphan task worktrees from pre-flight checkpoints. | `0` (Restored), `1` (Error) |
158
163
  | `resume` | `agentctl resume <sessionId> --response "<reply>"` | Streams engineer response back into active Google Jules warm session context window. | `0` (Resumed), `1` (Error) |
159
164
  | `test-gen` | `agentctl test-gen --title <t> --spec <s> [--run]` | Scaffolds falsifiable unit tests, verifies RED failure state, and locks test in `scope.deny`. | `0` (Scaffolded/Red) |
@@ -199,13 +204,28 @@ scope:
199
204
  - ".agent/config.yml"
200
205
  - "keys/**"
201
206
 
202
- # Operational limits & governors
207
+ # Plan tier. Defaults to `free` when unset — the kit will not assume you are
208
+ # paying for a larger plan than you are. Set this to unlock your real limits.
209
+ tier: "free" # free | pro | ultra
210
+
211
+ # Risk model for auto-merge triage. Builtin patterns cover what is dangerous in
212
+ # any repository (CI, lockfiles, migrations, key material, IaC, auth). Add the
213
+ # paths that are sensitive to YOUR domain — these EXTEND the builtins.
214
+ risk:
215
+ restricted: # R3 — never auto-merged
216
+ - "**/pricing/**"
217
+ - "**/billing/**"
218
+ consequential: # R2 — always requires a human read
219
+ - "packages/api/**"
220
+ max_routine_diff_lines: 400
221
+
222
+ # Operational limits & governors (tier defaults shown; any key here overrides)
203
223
  limits:
204
- diffKb: 75 # 75 KB Diff Payload Governor limit
224
+ diffKb: 75 # Diff Payload Governor limit
205
225
  promptKb: 50 # Maximum prompt payload size
206
226
  dailyTasks: 300 # Task quota per rolling 24h window (not per calendar day)
207
227
  repairAttempts: 3 # Maximum repair iterations
208
- concurrency: 15 # Worker slots (free: 3, pro: 8, ultra: 15)
228
+ concurrency: 15 # Worker slots (defaults free: 3, pro: 8, ultra: 15)
209
229
 
210
230
  # Dynamic Complexity & Cost Router — opt-in, disabled by default.
211
231
  router:
@@ -308,6 +328,23 @@ const { provider, classification } = resolveRoutedProvider(
308
328
  console.log(classification.tier); // "fast" | "complex"
309
329
  ```
310
330
 
331
+ ### Syntax-Verified FAST Tier (`createSyntaxVerifiedProvider`)
332
+ `resolveRoutedProvider()` already wraps the FAST tier with this; use it directly only when composing your own provider cascade.
333
+ ```javascript
334
+ import { createProvider, createSyntaxVerifiedProvider, loadConfig } from "jules-orchestrator-kit";
335
+
336
+ const config = loadConfig(process.cwd());
337
+ const fast = createSyntaxVerifiedProvider(
338
+ createProvider("gemini-flash", config),
339
+ createProvider("jules", config),
340
+ config
341
+ );
342
+
343
+ // If gemini-flash leaves broken .js/.mjs/.cjs on disk, this transparently
344
+ // re-dispatches through "jules" instead of returning the broken result.
345
+ const result = await fast.dispatch({ prompt: "Fix a typo." }, { root: process.cwd() });
346
+ ```
347
+
311
348
  </details>
312
349
 
313
350
  <br/>
package/bin/agentctl.mjs CHANGED
@@ -40,7 +40,7 @@ export function printHelp() {
40
40
  Usage: agentctl <command> [options]
41
41
 
42
42
  Commands:
43
- dispatch | create Dispatch a single task to an AI agent (--role <name>, --tier fast|complex)
43
+ dispatch | create Dispatch a single task to an AI agent (--role <name>, --tier fast|complex, --check-premise)
44
44
  gate | audit Run CI security and verification gate against current branch
45
45
  queue Run pending task queue (--dag, --concurrency <n>)
46
46
  swarm Run parallel task swarm
@@ -54,12 +54,15 @@ Commands:
54
54
  init Scaffold .agent/ config and run onboarding wizard
55
55
  task create Interactively author and scope a Jules task envelope (--template <name>, --role <name>, --tier fast|complex)
56
56
  task optimize Linter & optimizer for Jules task prompts (--fix, --json, --web)
57
- task template List and generate web development task templates (--list, --json)
57
+ task template List and generate task templates (--list, --json)
58
58
  test-gen Scaffold & run automated TDD Red-to-Green test cycle (--run)
59
59
  mcp init Scaffold IDE integration config (cursor | vscode | claude | all)
60
60
  rollback Restore git state & working tree to atomic pre-flight checkpoint
61
61
  handover Inspect or generate Baton Pass session handover envelopes (list | show | create | prune)
62
62
  resume Resume warm session with human response (--response "<text>")
63
+ plan approve <id> Approve pending plan for a Jules session (:approvePlan)
64
+ session get <id> Retrieve remote session status from provider API
65
+ pr harvest Triage, verify CI, and auto-merge low-risk agent PRs (--auto, --tier r0,r1)
63
66
  escalate Dispatch or manage webhook escalation incidents (--flush, --status, --clear)
64
67
  flaky Manage Wilson-quarantined tests and dispatch healing swarm (status | heal | reset)
65
68
  status Display queue and system status summary
@@ -69,11 +72,13 @@ Commands:
69
72
  harvest Harvest failure traces and record/quarantine resolution rules
70
73
  learning add Record a system learning rule into .agent/knowledge/
71
74
  evidence <action> Manage cryptographic audit evidence (generate | verify | show)
75
+ assert Run declarative zero-dependency verification assertion primitives
72
76
  version Output agentctl version
73
77
 
74
78
  Options:
75
79
  --role, -r Specify specialist agent role (overseer | bolt | sentinel | janitor)
76
80
  --tier Force routing tier when router.enabled (fast | complex) — see .agent/config.yml router:
81
+ --check-premise Verify task goal/oracle passes locally before burning API budget
77
82
  --dag Execute queue tasks via DAG dependency resolution
78
83
  --dry-run, -d Simulate action without making API calls or modifying git
79
84
  --mode, -m Gate evaluation mode (working-tree | committed | staged)
@@ -81,6 +86,7 @@ Options:
81
86
  --source, -s Specify Jules repository source name
82
87
  --branch, -b Specify target starting branch
83
88
  --json, -j Emit machine-readable JSON output
89
+ --json-report <path> Write structured machine-readable JSON diagnostics report
84
90
  --help, -h Show command help
85
91
  `);
86
92
  }
@@ -147,6 +153,9 @@ async function main() {
147
153
  repoless: { type: "boolean" },
148
154
  "auto-pr": { type: "boolean" },
149
155
  "require-plan-approval": { type: "boolean" },
156
+ "check-premise": { type: "boolean" },
157
+ idempotent: { type: "boolean" },
158
+ author: { type: "string" },
150
159
  "dry-run": { type: "boolean", short: "d" },
151
160
  json: { type: "boolean", short: "j" },
152
161
  },
@@ -177,6 +186,8 @@ async function main() {
177
186
  repoless: values.repoless,
178
187
  autoPr: values["auto-pr"],
179
188
  requirePlanApproval: values["require-plan-approval"],
189
+ checkPremise: values["check-premise"] || values.idempotent,
190
+ author: values.author,
180
191
  };
181
192
 
182
193
  try {
@@ -187,15 +198,21 @@ async function main() {
187
198
  repoless: values.repoless,
188
199
  source: values.source,
189
200
  branch: values.branch,
201
+ checkPremise: values["check-premise"] || values.idempotent,
190
202
  });
191
203
  if (values.json) {
192
204
  console.log(JSON.stringify({ ok: true, session }, null, 2));
193
205
  } else {
194
- console.log(`\n✅ Task Dispatched Successfully!`);
195
- console.log(` Session ID : ${session.id}`);
196
- console.log(` Session URL : ${session.url || "N/A"}`);
197
- if (session._routeTier) {
198
- console.log(` Router Tier : ${session._routeTier} (${session._routeReason || "n/a"})`);
206
+ if (session.status === "ALREADY_SATISFIED" || session.skipped) {
207
+ console.log(`\n⚡ Task Already Satisfied (skipped dispatch):`);
208
+ console.log(` Reason: ${session.reason || "Verification oracle already passing on base branch."}`);
209
+ } else {
210
+ console.log(`\n✅ Task Dispatched Successfully!`);
211
+ console.log(` Session ID : ${session.id}`);
212
+ console.log(` Session URL : ${session.url || "N/A"}`);
213
+ if (session._routeTier) {
214
+ console.log(` Router Tier : ${session._routeTier} (${session._routeReason || "n/a"})`);
215
+ }
199
216
  }
200
217
  }
201
218
  process.exit(0);
@@ -224,6 +241,7 @@ async function main() {
224
241
  fix: { type: "boolean" },
225
242
  "allow-protected": { type: "boolean" },
226
243
  json: { type: "boolean", short: "j" },
244
+ "json-report": { type: "string" },
227
245
  "dry-run": { type: "boolean", short: "d" },
228
246
  },
229
247
  allowPositionals: true,
@@ -241,6 +259,7 @@ async function main() {
241
259
  mode: selectedMode,
242
260
  fix: values.fix,
243
261
  allowProtected: values["allow-protected"],
262
+ jsonReport: values["json-report"],
244
263
  });
245
264
 
246
265
  if (values.json) {
@@ -285,6 +304,93 @@ async function main() {
285
304
  break;
286
305
  }
287
306
 
307
+ case "assert": {
308
+ const { runAssertion, parseYaml } = await import("../index.mjs");
309
+ const { values } = parseArgs({
310
+ args: args.slice(1),
311
+ options: {
312
+ dir: { type: "string", short: "d" },
313
+ file: { type: "string", short: "f" },
314
+ targets: { type: "string", short: "t" },
315
+ patterns: { type: "string", short: "p" },
316
+ "patterns-file": { type: "string" },
317
+ "max-bytes": { type: "string" },
318
+ "max-kb": { type: "string" },
319
+ "max-mb": { type: "string" },
320
+ gzip: { type: "boolean" },
321
+ config: { type: "string", short: "c" },
322
+ json: { type: "boolean", short: "j" },
323
+ "json-report": { type: "string" },
324
+ },
325
+ allowPositionals: true,
326
+ });
327
+
328
+ let stageConfig = {};
329
+ if (values.config && existsSync(values.config)) {
330
+ const raw = readFileSync(values.config, "utf-8");
331
+ stageConfig = parseYaml(raw);
332
+ } else {
333
+ let assertType = "dir-size";
334
+ if (values.file) assertType = "file-size";
335
+ else if (values.patterns || values["patterns-file"] || values.targets) assertType = "file-patterns";
336
+ else if (values.dir) assertType = "dir-size";
337
+
338
+ stageConfig = {
339
+ assert: assertType,
340
+ path: values.dir || values.file,
341
+ targets: values.targets ? values.targets.split(",").map((s) => s.trim()) : undefined,
342
+ patterns: values.patterns ? values.patterns.split(",").map((s) => s.trim()) : undefined,
343
+ patternsFile: values["patterns-file"],
344
+ maxBytes: values["max-bytes"] ? Number(values["max-bytes"]) : undefined,
345
+ maxKb: values["max-kb"] ? Number(values["max-kb"]) : undefined,
346
+ maxMb: values["max-mb"] ? Number(values["max-mb"]) : undefined,
347
+ gzip: values.gzip,
348
+ };
349
+ }
350
+
351
+ const result = runAssertion(stageConfig, root);
352
+ if (values["json-report"]) {
353
+ const { exportJsonReport, generateEvidenceManifest } = await import("../src/evidence.mjs");
354
+ const dummyManifest = generateEvidenceManifest(root, {
355
+ taskId: "CLI-ASSERT",
356
+ title: "Ad-hoc Assertion",
357
+ executionRecords: [
358
+ {
359
+ id: result.assertionType,
360
+ kind: "assert",
361
+ assert: result.assertionType,
362
+ exitCode: result.status,
363
+ durationMs: result.metrics?.durationMs || 0,
364
+ diagnostics: result.diagnostics,
365
+ metrics: result.metrics,
366
+ },
367
+ ],
368
+ diagnostics: result.diagnostics,
369
+ metrics: result.metrics,
370
+ ok: result.ok,
371
+ });
372
+ exportJsonReport(dummyManifest, values["json-report"]);
373
+ }
374
+
375
+ if (values.json) {
376
+ console.log(JSON.stringify(result, null, 2));
377
+ } else {
378
+ if (result.ok) {
379
+ console.log(`\n✅ Assertion [${result.assertionType}] PASSED (${result.metrics?.durationMs || 0}ms)`);
380
+ if (result.metrics?.measuredBytes !== undefined) {
381
+ console.log(` Measured : ${result.metrics.measuredBytes} bytes${result.metrics.gzip ? " (gzipped)" : ""}`);
382
+ }
383
+ } else {
384
+ console.error(`\n❌ Assertion [${result.assertionType}] FAILED (Exit 1)`);
385
+ for (const d of result.diagnostics) {
386
+ console.error(` - ${d}`);
387
+ }
388
+ }
389
+ }
390
+ process.exit(result.status);
391
+ break;
392
+ }
393
+
288
394
  case "queue": {
289
395
  const { values } = parseArgs({
290
396
  args: args.slice(1),
@@ -396,6 +502,27 @@ async function main() {
396
502
  process.exit(0);
397
503
  }
398
504
 
505
+ if (args.includes("--by-user") || args.includes("-u")) {
506
+ console.log("📊 Task Budget Attribution (Rolling 24h Window)");
507
+ console.log(`Daily Limit : ${b.limit} Tasks | Used: ${b.used} | Remaining: ${b.remaining}\n`);
508
+ const users = Object.entries(b.byUser || {});
509
+ if (users.length === 0) {
510
+ console.log(" No user activity recorded in the active 24h window.\n");
511
+ } else {
512
+ console.log(" Author Tasks Committed Pending");
513
+ console.log(" ------------------- ----- --------- -------");
514
+ for (const [user, stats] of users.sort(([, a], [, b]) => b.tasks - a.tasks)) {
515
+ const padUser = user.padEnd(20);
516
+ const padTasks = String(stats.tasks).padStart(5);
517
+ const padCommitted = String(stats.committed).padStart(9);
518
+ const padPending = String(stats.uncommitted).padStart(7);
519
+ console.log(` ${padUser} ${padTasks} ${padCommitted} ${padPending}`);
520
+ }
521
+ console.log("");
522
+ }
523
+ process.exit(0);
524
+ }
525
+
399
526
  if (args.includes("--json")) {
400
527
  console.log(JSON.stringify({ ok: true, budget: { ...b, scope: "this-repository" } }, null, 2));
401
528
  process.exit(0);
@@ -559,7 +686,12 @@ async function main() {
559
686
  const { runInitWizard } = await import("../src/wizard-init.mjs");
560
687
  const res = await runInitWizard(root, {
561
688
  interactive: values.interactive !== false,
562
- tier: values.tier || "pro",
689
+ // No `|| "pro"`: a hardcoded default here overrode both the tier picked
690
+ // in the menu and the tier already recorded in .agent/config.yml on a
691
+ // re-run. Undefined lets the wizard seed the menu from the existing
692
+ // config and fall back to FALLBACK_TIER when there is nothing to seed.
693
+ tier: values.tier,
694
+ allowDefaults: true,
563
695
  });
564
696
 
565
697
  if (values.json) {
@@ -643,7 +775,7 @@ async function main() {
643
775
  if (values.json) {
644
776
  console.log(JSON.stringify({ ok: true, templates }, null, 2));
645
777
  } else {
646
- console.log(`\n🌐 Available Web Development Task Templates`);
778
+ console.log(`\n📋 Available Task Templates & Envelopes`);
647
779
  console.log(`--------------------------------------------------`);
648
780
  templates.forEach((t) => {
649
781
  console.log(` • ${t.id.padEnd(16)} [${t.category}]`);
@@ -977,6 +1109,162 @@ async function main() {
977
1109
  break;
978
1110
  }
979
1111
 
1112
+ case "plan": {
1113
+ const subAction = args[1];
1114
+ if (subAction === "approve") {
1115
+ const sessionId = args[2];
1116
+ const { values, positionals } = parseArgs({
1117
+ args: args.slice(2),
1118
+ options: {
1119
+ "dry-run": { type: "boolean", short: "d" },
1120
+ json: { type: "boolean", short: "j" },
1121
+ },
1122
+ allowPositionals: true,
1123
+ });
1124
+ const targetSessionId = sessionId && !sessionId.startsWith("-") ? sessionId : positionals?.[0];
1125
+ if (!targetSessionId) {
1126
+ console.error("Error: Session ID is required for agentctl plan approve <sessionId>.");
1127
+ process.exit(1);
1128
+ }
1129
+ const { createProvider } = await import("../src/provider.mjs");
1130
+ const provider = createProvider(config.provider || "jules", config);
1131
+ try {
1132
+ const res = await provider.approvePlan(targetSessionId, { root, dryRun: values["dry-run"] });
1133
+ if (values.json) {
1134
+ console.log(JSON.stringify(res, null, 2));
1135
+ } else {
1136
+ console.log(`\n✅ Plan Approved Successfully!`);
1137
+ console.log(` Session ID : ${res.id}`);
1138
+ console.log(` Status : ${res.status}\n`);
1139
+ }
1140
+ process.exit(0);
1141
+ } catch (err) {
1142
+ console.error(`❌ Plan Approval Failed: ${err.message}`);
1143
+ process.exit(1);
1144
+ }
1145
+ }
1146
+ console.error(`Error: Unknown plan subcommand '${subAction}'. Use approve.`);
1147
+ process.exit(1);
1148
+ break;
1149
+ }
1150
+
1151
+ case "approve": {
1152
+ const sessionId = args[1];
1153
+ const { values, positionals } = parseArgs({
1154
+ args: args.slice(1),
1155
+ options: {
1156
+ "dry-run": { type: "boolean", short: "d" },
1157
+ json: { type: "boolean", short: "j" },
1158
+ },
1159
+ allowPositionals: true,
1160
+ });
1161
+ const targetSessionId = sessionId && !sessionId.startsWith("-") ? sessionId : positionals?.[0];
1162
+ if (!targetSessionId) {
1163
+ console.error("Error: Session ID is required for agentctl approve <sessionId>.");
1164
+ process.exit(1);
1165
+ }
1166
+ const { createProvider } = await import("../src/provider.mjs");
1167
+ const provider = createProvider(config.provider || "jules", config);
1168
+ try {
1169
+ const res = await provider.approvePlan(targetSessionId, { root, dryRun: values["dry-run"] });
1170
+ if (values.json) {
1171
+ console.log(JSON.stringify(res, null, 2));
1172
+ } else {
1173
+ console.log(`\n✅ Plan Approved Successfully!`);
1174
+ console.log(` Session ID : ${res.id}`);
1175
+ console.log(` Status : ${res.status}\n`);
1176
+ }
1177
+ process.exit(0);
1178
+ } catch (err) {
1179
+ console.error(`❌ Plan Approval Failed: ${err.message}`);
1180
+ process.exit(1);
1181
+ }
1182
+ break;
1183
+ }
1184
+
1185
+ case "session": {
1186
+ const subAction = args[1];
1187
+ if (subAction === "get" || subAction === "status") {
1188
+ const sessionId = args[2];
1189
+ const { values, positionals } = parseArgs({
1190
+ args: args.slice(2),
1191
+ options: {
1192
+ "dry-run": { type: "boolean", short: "d" },
1193
+ json: { type: "boolean", short: "j" },
1194
+ },
1195
+ allowPositionals: true,
1196
+ });
1197
+ const targetSessionId = sessionId && !sessionId.startsWith("-") ? sessionId : positionals?.[0];
1198
+ if (!targetSessionId) {
1199
+ console.error("Error: Session ID is required for agentctl session get <sessionId>.");
1200
+ process.exit(1);
1201
+ }
1202
+ const { createProvider } = await import("../src/provider.mjs");
1203
+ const provider = createProvider(config.provider || "jules", config);
1204
+ try {
1205
+ const res = await provider.getSession(targetSessionId, { root, dryRun: values["dry-run"] });
1206
+ if (values.json) {
1207
+ console.log(JSON.stringify(res, null, 2));
1208
+ } else {
1209
+ console.log(`\n📋 Remote Session Status:`);
1210
+ console.log(` Session ID : ${res.id}`);
1211
+ console.log(` Status : ${res.status}\n`);
1212
+ }
1213
+ process.exit(0);
1214
+ } catch (err) {
1215
+ console.error(`❌ Session Retrieval Failed: ${err.message}`);
1216
+ process.exit(1);
1217
+ }
1218
+ }
1219
+ console.error(`Error: Unknown session subcommand '${subAction}'. Use get.`);
1220
+ process.exit(1);
1221
+ break;
1222
+ }
1223
+
1224
+ case "pr": {
1225
+ const subAction = args[1];
1226
+ if (subAction === "harvest") {
1227
+ const { harvestPullRequests, formatHarvestTable } = await import("../src/ops/pr-harvest.mjs");
1228
+ const { values } = parseArgs({
1229
+ args: args.slice(2),
1230
+ options: {
1231
+ tier: { type: "string" },
1232
+ limit: { type: "string" },
1233
+ auto: { type: "boolean" },
1234
+ merge: { type: "boolean" },
1235
+ "allow-no-checks": { type: "boolean" },
1236
+ "dry-run": { type: "boolean", short: "d" },
1237
+ json: { type: "boolean", short: "j" },
1238
+ },
1239
+ allowPositionals: true,
1240
+ });
1241
+
1242
+ const limit = values.limit ? parseInt(values.limit, 10) : 50;
1243
+ try {
1244
+ const res = await harvestPullRequests(root, {
1245
+ tier: values.tier,
1246
+ limit,
1247
+ auto: values.auto || values.merge,
1248
+ allowNoChecks: values["allow-no-checks"],
1249
+ dryRun: values["dry-run"],
1250
+ });
1251
+
1252
+ if (values.json) {
1253
+ console.log(JSON.stringify(res, null, 2));
1254
+ } else {
1255
+ console.log(formatHarvestTable(res));
1256
+ }
1257
+ process.exit(0);
1258
+ } catch (err) {
1259
+ console.error(`❌ PR Harvest Failed: ${err.message}`);
1260
+ process.exit(1);
1261
+ }
1262
+ }
1263
+ console.error(`Error: Unknown PR subcommand '${subAction}'. Use harvest.`);
1264
+ process.exit(1);
1265
+ break;
1266
+ }
1267
+
980
1268
  case "escalate": {
981
1269
  const {
982
1270
  dispatchEscalation,