jonah-fleet 1.11.0 → 1.13.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "jonah-fleet",
3
- "version": "1.11.0",
3
+ "version": "1.13.0",
4
4
  "description": "Standalone autonomous agent fleet with Symphony orchestration, claim protocols, and continuous improvement loops",
5
5
  "main": "dist/index.js",
6
6
  "types": "dist/index.d.ts",
package/schema.json CHANGED
@@ -163,6 +163,32 @@
163
163
  },
164
164
  "additionalProperties": false,
165
165
  "description": "Configuration for label management and protection shields"
166
+ },
167
+ "lessons": {
168
+ "oneOf": [
169
+ {
170
+ "type": "boolean",
171
+ "description": "Enable or disable opt-in LESSONS.md operational memory tier"
172
+ },
173
+ {
174
+ "type": "object",
175
+ "properties": {
176
+ "enabled": {
177
+ "type": "boolean",
178
+ "default": true,
179
+ "description": "Enable opt-in LESSONS.md operational memory tier"
180
+ },
181
+ "maxEntries": {
182
+ "type": "number",
183
+ "default": 25,
184
+ "description": "Maximum number of active operational lesson entries (hard cap: 25)"
185
+ }
186
+ },
187
+ "additionalProperties": false,
188
+ "description": "Configuration for opt-in operational memory tier (LESSONS.md)"
189
+ }
190
+ ],
191
+ "description": "Configuration for opt-in operational memory tier (LESSONS.md)"
166
192
  }
167
193
  },
168
194
  "required": ["version", "preset", "routines", "skills"],
@@ -0,0 +1,7 @@
1
+ # Operational Lessons & Repository Heuristics
2
+ <!-- Invariants: Max 25 entries. Hard cap. Older entries graduate or demote to LESSONS_ARCHIVE.md -->
3
+
4
+ ### [subsystem] Title
5
+ - **Symptom:** <Symptom description or error message>
6
+ - **Root Cause:** <Brief explanation of the underlying cause>
7
+ - **Rule:** <Actionable rule or constraint to follow>
@@ -210,11 +210,11 @@ How the fleet guarantees continuous review throughput, recovers from transient A
210
210
 
211
211
  ---
212
212
 
213
- ## Upstream Symphony & Funes Intel & Architectural Evaluation Framework
213
+ ## Upstream Symphony, Funes & Orbital Intel & Architectural Evaluation Framework
214
214
 
215
- How changes and innovations from upstream ecosystems—[openai/symphony](https://github.com/openai/symphony) for issue-tracker orchestration and [huggingface/funes](https://github.com/huggingface/funes) for agent memory & session indexing—are systematically audited and evaluated for incorporation into Jonah Fleet:
215
+ How changes and innovations from upstream ecosystems—[openai/symphony](https://github.com/openai/symphony) for issue-tracker orchestration, [huggingface/funes](https://github.com/huggingface/funes) for agent memory & session indexing, and [zqiren/Orbital](https://github.com/zqiren/Orbital) for project agents & worker transports—are systematically audited and evaluated for incorporation into Jonah Fleet:
216
216
 
217
- 1. **Automated Ecosystem Radar (`symphony-radar.yml`)**: A weekly scheduled workflow runs `.github/scripts/fetch-symphony-radar.js` to inspect upstream commits, specification updates (`SPEC.md`), and releases across `openai/symphony` (orchestration) and `huggingface/funes` (memory tooling), generating an actionable digest issue in Jonah Fleet.
217
+ 1. **Automated Ecosystem Radar (`symphony-radar.yml`)**: A weekly scheduled workflow runs `.github/scripts/fetch-symphony-radar.js` to inspect upstream commits, specification updates (`SPEC.md`), releases, and pull requests across `openai/symphony` (orchestration), `huggingface/funes` (memory tooling), and `zqiren/Orbital` (project agents & worker transports), generating an actionable digest issue in Jonah Fleet.
218
218
  2. **The 4 Evaluation Layers**:
219
219
  - **Layer 1 (Zero-Daemon Invariant)**: Can the enhancement execute in ephemeral GitHub Actions and `agy` CLI sessions without requiring a 24/7 background server or persistent WebSocket?
220
220
  - **Layer 2 (Issue Tracker Abstraction)**: Does the pattern map cleanly to native GitHub Issues, labels, and PR checks without proprietary tracker dependencies?
@@ -225,9 +225,14 @@ How changes and innovations from upstream ecosystems—[openai/symphony](https:/
225
225
  - **Pull-Based Memory Delivery**: Memory served strictly on demand via MCP (`recall`, `get`) to prevent prompt context bloat.
226
226
  - **Cross-Session Provenance**: Verbatim turns and provenance retention instead of lossy summary drift.
227
227
  - **Multi-Agent Portability**: Standardized trace ingestion across Antigravity CLI (`agy`), Claude Code, and Codex.
228
- 4. **Classification & Action Protocol**:
229
- - **🟢 Category A (Adopt Directly)**: Security guardrails, claim lock invariants, reader/writer rules, prompt engineering optimizations, deterministic zero-LLM indexing.
230
- - **🟡 Category B (Adapt to Actions/CLI)**: Dynamic orchestrator pacing, backpressure controls, multi-stage review checks, pull-based memory MCP integrations.
228
+ 4. **Project Agent & Worker Transports Evaluation Dimensions (Orbital Watch)**:
229
+ - **Layer-1 Context Memory Files**: In-repo memory files (`LESSONS.md`, `CONTEXT.md`) maintained directly by agents on PR branches to avoid uncommitted disk drift or git merge collisions.
230
+ - **ACP/PTY Worker Transports**: Agent Client Protocol (ACP) and pseudo-terminal delegation to sub-agents, adaptable for local CLI runner execution.
231
+ - **Prompt Prefix Caching Benchmarks**: Partitioning prompt structures into Static $\rightarrow$ Semi-Stable $\rightarrow$ Dynamic tiers to maximize prefix cache hit rates (~95%).
232
+ - **Fail-Closed Safety Guards**: Deterministic action-hash repetition guards, cycle detection, and circuit breakers that halt runaway execution loops before token budgets are breached.
233
+ 5. **Classification & Action Protocol**:
234
+ - **🟢 Category A (Adopt Directly)**: Security guardrails, claim lock invariants, reader/writer rules, prompt engineering & prefix caching optimizations, fail-closed safety guards, deterministic zero-LLM indexing.
235
+ - **🟡 Category B (Adapt to Actions/CLI)**: Dynamic orchestrator pacing, backpressure controls, multi-stage review checks, pull-based memory MCP integrations, ACP/PTY worker transports.
231
236
  - **🔴 Category C (Skip)**: Elixir/OTP supervision trees, BEAM memory tuning, proprietary runtime internals, always-loaded memory context dumps.
232
237
 
233
238
  ---
@@ -257,3 +262,39 @@ How agent routines are partitioned between cloud GitHub Actions (24/7 cloud runn
257
262
  - Local agent PR convergence claims post: `🔒 Addressing review findings by local autowork session (host: <hostname>) <timestamp>`.
258
263
  - Local processes trap `SIGINT`/`SIGTERM` to unassign claims and remove worktrees cleanly on exit.
259
264
  - Standard stale-claim rules (6h for issues, 2h for PRs) safely reclaim orphaned local claims if a machine powers down unexpectedly.
265
+
266
+ ---
267
+
268
+ ## Headless Execution & Asynchronous Non-Yielding Guardrail
269
+
270
+ In headless CLI environments (`agy -p` / GitHub Actions), agent sessions terminate immediately whenever the model yields a turn without active tool calls. Therefore:
271
+
272
+ 1. **Zero-Yield Waiting Invariant**: Agents MUST NEVER call `schedule` or emit a terminal turn with plain text to "wait" for background commands, timers, or long-running checks. In headless mode, yielding the turn halts the process immediately with exit code 0 before reaching the Definition of Done.
273
+ 2. **Active Task Supervision**: If a verification command (`npm test`, `npm run type-check`) is sent to the background by `run_command`, the agent must actively poll `manage_task(Action='status')` or inspect code while waiting within the continuous tool-calling loop.
274
+ 3. **CI Trust Bar & Test Discipline**: Peer review routines should trust green passing remote CI checks (GitHub Actions or Vercel preview deployments) on the PR's head commit rather than initiating slow, background-prone full test runs. Run repository verification locally ONLY if CI status is unconfirmed, missing, or failing.
275
+
276
+ ---
277
+
278
+ ## Structured Human Escalation Card Protocol ("Why I believe this")
279
+
280
+ How autonomous routines escalate decisions, ambiguities, and blockers to human maintainers without unbounded back-and-forth or vague questions:
281
+
282
+ 1. **Mandatory 4-Part Escalation Schema**: Whenever a routine cannot proceed autonomously due to ambiguity, conflicting requirements, unobservable acceptance criteria, or repeated review ping-pong—and applies `needs-human` or `needs-info`—it MUST post a comment structured as the mandatory 4-part escalation card (inspired by Orbital's Workbench Provenance format):
283
+ ```markdown
284
+ ## 🛑 Escalation: Human Decision Required
285
+ - **Decision Needed**: [1 focused question or choice]
286
+ - **Evidence ("Why I believe this")**: [Specific files, lines, test outputs, or conflicting docs]
287
+ - **Evaluated Options & Trade-offs**:
288
+ - *Option A*: [Pros / Cons]
289
+ - *Option B*: [Pros / Cons]
290
+ - **Recommended Path**: [Agent recommendation]
291
+ ```
292
+ 2. **Card Invariants**:
293
+ - **Decision Needed**: Exactly 1 high-leverage question or choice required from the maintainer or reporter. Prohibit question dumps or vague "please provide more details".
294
+ - **Evidence ("Why I believe this")**: Concrete artifacts, specific file paths, line numbers, test outputs, or contradicting specification documents justifying why the routine cannot proceed without human guidance.
295
+ - **Evaluated Options & Trade-offs**: At least two distinct, viable options with concrete pros and cons. Never ask maintainers to solve problems from scratch without agent-evaluated trade-offs.
296
+ - **Recommended Path**: The agent's recommended decision and reasoning, allowing maintainers to unblock execution with a simple confirmation.
297
+ 3. **Cross-Routine Enforcement**:
298
+ - `autowork.md`: Required when tripping the Ambiguity Gate (Step 12), encountering a 2nd-strike permanent blocker (`needs-human`), or hitting the review Ping-Pong Cap (Step 3b).
299
+ - `triage/SKILL.md`: Required when transitioning issues or PRs to `needs-info` or `ready-for-human`.
300
+ - `issues-housekeeping.md`: Required when auditing and escalating ambiguous, stale, or infeasible issues with `needs-human` or `needs-info`.
@@ -29,6 +29,7 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
29
29
  - **Single-flight per issue & PR**: multiple autowork runs can execute concurrently. Both issues and pull requests are shared resources — never begin implementing an issue without first claiming it (see Claim protocol in Phase 2), and never begin addressing findings on an open PR without first claiming it (see PR Claim protocol in Phase 1).
30
30
  - **Language Requirement**: All GitHub issue titles, descriptions, task checklists, and comments MUST be written in **English**.
31
31
  - **Session link footer**: sign every GitHub post (issue comments, PR comments, PR descriptions) with the Antigravity run footer (`_Generated by [Antigravity](${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}/actions/runs/${GITHUB_RUN_ID})_`). Inline review-thread line comments are exempt.
32
+ - **Headless Execution & Asynchronous Non-Yielding Guardrail**: You are running in a headless autonomous session. You MUST NEVER call `schedule` or yield the turn with plain text while waiting for background tasks or verification checks. If a verification command runs in the background, actively inspect its completion with `manage_task` or run with sufficient `WaitMsBeforeAsync` (up to `10000` ms). Never stop calling tools or yield your turn until the routine's terminal Definition of Done is fully reached.
32
33
 
33
34
  ## Negative examples (DO NOT do these)
34
35
 
@@ -44,7 +45,8 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
44
45
  - Do not start implementing an issue before claiming it (both assignment AND claim comment).
45
46
  - Do not mark a PR ready while its `mergeable_state` is `dirty` — resolve merge conflicts first.
46
47
  - Do not fall into the **Telemetry Rabbit Hole**: do not spend cycles instrumenting elaborate fallback telemetry or defensive error handling for features that suffer from lack of user intent rather than software bugs.
47
- - Do not guess or invent arbitrary specifications for ambiguous issues — post clarifying questions, label `needs-info`, and release the claim instead of blindly writing code.
48
+ - Do not guess or invent arbitrary specifications for ambiguous issues — post the 4-part escalation card ('Why I believe this'), label `needs-info`, and release the claim instead of blindly writing code.
49
+ - Do not call `schedule` or yield the turn with plain text while waiting for background verification tasks — stay in the tool loop until the Definition of Done is met.
48
50
 
49
51
  ## Instructions
50
52
 
@@ -88,8 +90,17 @@ a. **Read the target issue and check eligibility.** Eligible = open, unassigned
88
90
  - **Build & type-check verification**: run the repository's test, type-check, and lint commands from `AGENTS.md` (e.g. `npm test`, `npm run type-check`, `npm run lint`, `pytest`, `cargo test`). Confirm zero errors and zero test failures.
89
91
  - **Active Origin Sync & Clean-Merge Gate**: Run `git fetch origin main && git merge origin/main --no-edit` to absorb any newly merged pull requests and resolve any conflicts locally. Verify `git merge-tree origin/main HEAD` reports no conflicts before marking ready.
90
92
  - **Release claim on ready**: mark the PR ready (`gh pr ready <PR>`) and unassign yourself (`gh pr edit <PR> --remove-assignee <login>`) so Peer Review can evaluate without holding stale agent reservation locks.
91
- Only mark the PR ready after passing every check above.
92
- 3b. **Ping-pong cap**: If this same PR has bounced between draft and ready 3 or more times over the same substantive finding, stop re-marking it ready. Post a comment summarizing the disagreement for human resolution and leave the PR in draft.
93
+ Only mark the PR ready after passing every check above.
94
+ 3b. **Ping-pong cap**: If this same PR has bounced between draft and ready 3 or more times over the same substantive finding, stop re-marking it ready. Post the mandatory 4-part escalation card summarizing the disagreement for human resolution and leave the PR in draft:
95
+ ```markdown
96
+ ## 🛑 Escalation: Human Decision Required
97
+ - **Decision Needed**: [1 focused question or choice]
98
+ - **Evidence ("Why I believe this")**: [Specific files, lines, test outputs, or conflicting docs]
99
+ - **Evaluated Options & Trade-offs**:
100
+ - *Option A*: [Pros / Cons]
101
+ - *Option B*: [Pros / Cons]
102
+ - **Recommended Path**: [Agent recommendation]
103
+ ```
93
104
  3c. **Orphaned Ready PR Recovery**: If an open PR authored by this routine is `ready_for_review`, has passing CI, no unaddressed review comments, and has received no review activity for over 2 hours (e.g. because peer review crashed or encountered quota limits), kickstart the review routine by posting `/review` comment or toggling draft and ready (`gh pr ready <PR> --undo && gh pr ready <PR>`).
94
105
  - **Passing CI Verification Gate**: Verify via `gh pr view <PR> --json statusCheckRollup,mergeStateStatus` that all required and existing checks have completed with `conclusion: "SUCCESS"` and `mergeStateStatus` is `CLEAN` (neither `UNSTABLE`, `BLOCKED`, nor `DIRTY`).
95
106
  - **Unapproved/Pending Workflow Invariant**: NEVER post `/review` or toggle draft state if checks are in-progress, failing, or awaiting approval (`conclusion: "ACTION_REQUIRED"`). Doing so creates an infinite comment storm while workflows remain paused awaiting human permissions.
@@ -128,13 +139,23 @@ a. **Read the target issue and check eligibility.** Eligible = open, unassigned
128
139
  - Read the issue description, linked code, and comment thread.
129
140
  - If bug: use `/diagnosing-bugs` to establish reproduction test before fixing.
130
141
  - If large/complex: use `/domain-modeling` and `/codebase-design`.
142
+ - **Pre-Flight Memory Scan**: If `LESSONS.md` exists, grep matching subsystem tags (`grep -E "^### \[(subsystem)\]" LESSONS.md -A 4`) to incorporate known landmines into implementation plans before writing code.
131
143
  - **Ambiguity & Missing Acceptance Criteria Gate**: Challenge underspecified or incomplete requests before writing any code. If the issue lacks observable acceptance criteria, relies on unverified assumptions, or leaves critical technical/UX decisions ambiguous:
132
144
  - Do NOT guess or invent arbitrary requirements to force completion.
133
- - Post a comment on the issue posing 1–3 focused clarifying questions that identify the exact decisions or trade-offs needed.
145
+ - Post the mandatory 4-part escalation card on the issue:
146
+ ```markdown
147
+ ## 🛑 Escalation: Human Decision Required
148
+ - **Decision Needed**: [1 focused question or choice]
149
+ - **Evidence ("Why I believe this")**: [Specific files, lines, test outputs, or conflicting docs]
150
+ - **Evaluated Options & Trade-offs**:
151
+ - *Option A*: [Pros / Cons]
152
+ - *Option B*: [Pros / Cons]
153
+ - **Recommended Path**: [Agent recommendation]
154
+ ```
134
155
  - Apply the `needs-info` label and release the claim (unassign).
135
156
  - Select the next candidate (evaluating ambiguous issues counts toward step 12's infeasible-continuation cap).
136
157
  - **Intent vs. Defect Guardrail**: When investigating issues related to low conversion, zero-click events, or underperforming features: verify whether the issue is a software defect or a lack of user intent. If data indicates the root cause is **lack of user intent** (e.g. button is rendered above fold and functions correctly when clicked, but user interaction rate is <2%) rather than a software defect, do NOT fall into the **telemetry rabbit hole** (adding elaborate fallback telemetry, downstream error handling, or defensive rendering). Categorize the issue as a **product/UX question** (`needs-design` / `roadmap/*`), comment explaining the lack of user intent, release the claim (unassign), and select the next candidate.
137
- - If infeasible: comment explaining blocker, release claim (unassign), and select next candidate (up to 3 infeasible evaluations per run). If permanent blocker on 2nd strike, apply `needs-human` label and tag repo owner.
158
+ - If infeasible: comment explaining blocker, release claim (unassign), and select next candidate (up to 3 infeasible evaluations per run). If permanent blocker on 2nd strike, post the mandatory 4-part escalation card (`## 🛑 Escalation: Human Decision Required`), apply `needs-human` label, and tag repo owner.
138
159
  12a. **Umbrella-issue handoff + batching:** If candidate is an umbrella epic:
139
160
  - Read `🧭 Decomposition plan` comment (or create if first run).
140
161
  - Pick next slice(s), batching up to 3 same-recipe slices into one child issue + PR.
@@ -151,7 +172,9 @@ a. **Read the target issue and check eligibility.** Eligible = open, unassigned
151
172
  13. **Implementation & PR creation:**
152
173
  - Branch from freshly fetched `origin/main` with descriptive name (e.g. `feat/...` or `fix/...`).
153
174
  - Drive implementation via `/tdd` (red-green-refactor).
175
+ - **Diagnostic Reflex**: On unexpected test/build failure during TDD, grep symptom text in `LESSONS.md` before making speculative code edits.
154
176
  - Run repository tests and verification.
177
+ - **Pre-PR Lessons Capture Gate**: If solving the issue required overcoming a non-obvious quirk not caught by tests/linters, append a 3-line structured entry on the active PR branch adhering to the 25-entry hard cap. Never write to `LESSONS.md` directly on `main`.
155
178
  - **Milestone 2 (Verification & Tests)**: Once implementation passes tests and type checks, emit milestone card if `$ROUTINE_ISSUE_NUMBER` is set:
156
179
  ```bash
157
180
  gh issue comment "$ROUTINE_ISSUE_NUMBER" --body "### 🧪 Milestone: Verification & Tests
@@ -39,7 +39,17 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
39
39
  3. **Priority review**: Check open P1/P2/P3 issues. Promote critical bugs or unblocked items; demote items that lack immediate priority.
40
40
  4. **Duplicate & consolidation check**: Identify duplicate issues; close duplicates with cross-references. Consolidate small, related micro-tasks into batch issues.
41
41
  5. **Premise-obsolete & stale check**: If an issue's premise was resolved by already-merged PRs or recent refactors, close as completed with evidence.
42
- 6. **Label audit & safe prune**: Ensure open issues carry standard role labels (`needs-triage`, `ready-for-agent`, `needs-human`, etc.). Use `/triage` if classifying incoming issues. Run `npx --yes jonah-fleet labels prune --yes` (or `jonah-fleet labels prune --yes`) to safely prune strictly unused boilerplate labels (`issues: 0`, `pullRequests: 0`, non-protected taxonomy) without deleting historical or fleet taxonomy labels.
42
+ 6. **Label audit & safe prune**: Ensure open issues carry standard role labels (`needs-triage`, `ready-for-agent`, `needs-human`, etc.). Use `/triage` if classifying incoming issues. Whenever applying `needs-human` or `needs-info` to escalate an ambiguous, stale, or infeasible issue, mandate formatting the escalation comment with the 4-part card:
43
+ ```markdown
44
+ ## 🛑 Escalation: Human Decision Required
45
+ - **Decision Needed**: [1 focused question or choice]
46
+ - **Evidence ("Why I believe this")**: [Specific files, lines, test outputs, or conflicting docs]
47
+ - **Evaluated Options & Trade-offs**:
48
+ - *Option A*: [Pros / Cons]
49
+ - *Option B*: [Pros / Cons]
50
+ - **Recommended Path**: [Agent recommendation]
51
+ ```
52
+ Run `npx --yes jonah-fleet labels prune --yes` (or `jonah-fleet labels prune --yes`) to safely prune strictly unused boilerplate labels (`issues: 0`, `pullRequests: 0`, non-protected taxonomy) without deleting historical or fleet taxonomy labels.
43
53
  7. **Closed-loop verification check**: For projects running impact or verification loops, audit recently closed roadmap/feature issues against tracking issues to ensure shipped levers do not remain untracked.
44
54
 
45
55
  ### Phase 3: Summary
@@ -71,6 +71,8 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
71
71
  - **Passive Order-Taking Anomaly ("Yes-Man Blindspot")**: The Ambiguity Gate trigger rate across intake runs in `autowork` or `triage` is <5% despite elevated PR review bounces ($\ge 2$) or high iteration usage ($\ge 35$), indicating agents are silently guessing requirements and building flawed implementations rather than interrogating underspecified issues.
72
72
  - **Speculative Runaway Waste**: An agent run consumed >50k tokens on an underspecified issue with 0 clarifying questions asked, and subsequently failed, bounced, or required post-merge rework.
73
73
  6. **Analyze resolved bugs & review comments**: Examine closed bug issues, merged bug-fix PRs, and review feedback for missing checks in authoring (`autowork.md`) or review (`peer-review.md`).
74
+ 7. **Scan & Audit Operational Lessons (`LESSONS.md`)**:
75
+ - Scan `LESSONS.md` during routine optimization sweeps, graduating stable rules to automated linter/CI checks or archiving stale entries to `LESSONS_ARCHIVE.md`.
74
76
 
75
77
  ### 2. Formulate preventative improvements
76
78
 
@@ -81,6 +83,7 @@ Translate findings into concrete preventative improvements and remediation trigg
81
83
  - **Ping-Pong Convergence**: For Review Loop Burn, tighten reviewer trust & noise filtering, enforce clean-merge gates, and apply ping-pong caps to prevent endless bounce cycles.
82
84
  - **Loop Discovery Mechanical Audits**: For Feedback Loop Stagnation, tighten discovery sweeps by mandating deterministic per-issue matching tables and itemized reconciliation against upstream closed issues/PRs rather than allowing un-itemized generic summary assertions.
83
85
  - **Ambiguity Gate & Benchmark Eval Feeding**: For Passive Order-Taking and Speculative Runaway Waste, tighten Step 12 criteria in `autowork.md` and `triage.md` to mandate clarifying questions, and automatically extract the problem issue into a `BenchmarkIssue` test case to feed the automated ambiguity benchmark eval suite (`tests/evals.test.ts`), ensuring future agent prompts are continuously tested against real failure cases.
86
+ - **Operational Memory Graduation & Archiving**: Scan `LESSONS.md` to identify recurring, stable rules for graduation into automated linter rules or CI workflow checks. Move obsolete or overflow entries (>25 cap) to `LESSONS_ARCHIVE.md`.
84
87
  - **Verification & Invariant Tests**: Add automated test cases in `tests/` verifying prompt invariant preservation and schema conformity.
85
88
 
86
89
  ### 3. Open Fix PR (Local or Upstream Bridge)
@@ -33,6 +33,8 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
33
33
  - **No speculative work**: review only the diff in the PR.
34
34
  - **Language Requirement**: All GitHub comments, review notes, and follow-up issues MUST be written in **English**.
35
35
  - **Session link footer**: sign every GitHub post with the Antigravity run footer (`_Generated by [Antigravity](${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}/actions/runs/${GITHUB_RUN_ID})_`).
36
+ - **Headless Execution & Asynchronous Non-Yielding Guardrail**: You are running in a headless autonomous CLI session. You MUST NEVER call `schedule` or yield the turn with plain text to "wait" for background tasks, timers, or long-running commands. In headless CLI mode, yielding the turn terminates the session immediately before reaching the Definition of Done. When invoking `run_command` for test runs or type checks, ALWAYS set `WaitMsBeforeAsync` high enough (up to the maximum `10000` ms) or if a command moves to the background, keep your execution turn active by checking task status (`manage_task(Action='status')`) or inspecting files; NEVER call `schedule` or emit a terminal turn with plain text before taking the final action.
37
+ - **Verification Command & CI Discipline**: When remote CI checks (e.g. GitHub Actions, Vercel preview deployments) are already present and passed on the PR's head commit, treat CI as verified and green. Avoid executing slow, redundant local full test suites or full type-checks when CI is already green. Run local verification commands (`npm test`, `npm run type-check`) ONLY if CI is missing, unconfirmed, or failing.
36
38
 
37
39
  ## Final action: merge or bounce to draft
38
40
 
@@ -58,6 +60,8 @@ Every review ends in exactly one of two states:
58
60
  - On re-review, do not raise new findings in code that was unchanged since the prior review — only inspect the delta commits.
59
61
  - Do not bounce a PR for non-blocking style/preference findings when all correctness checks pass.
60
62
  - Do not bounce an external contributor PR purely for missing `Closes #N` when the PR description provides a clear specification.
63
+ - Do not call `schedule` or yield the turn with plain text while waiting for background verification tasks — stay in the tool loop until the final action is taken.
64
+ - Do not rerun full local test suites when remote CI checks on the head commit are already green and passing.
61
65
 
62
66
  ## Instructions
63
67
 
@@ -101,6 +105,7 @@ Check if `$PR_NUMBER` is set:
101
105
  1. Run `/code-review` over the diff (or delta commits if re-review) evaluating:
102
106
  - **Standards**: Conformance to `AGENTS.md` (or `CLAUDE.md`/`GEMINI.md`), conventions, and architecture.
103
107
  - **Spec Compliance**: Verification against the linked issue's deliverables (`## Tasks`), or against the PR description's summary/changes if no tracking issue is linked.
108
+ - **Operational Memory & Lessons Invariant**: Inspect `LESSONS.md` diffs in PRs to verify the lesson is accurate, non-trivial, follows the 3-line structured schema, and the file adheres to the 25-entry hard cap.
104
109
  2. Run Security Pass: auth gates, permission checks, injection risks, sensitive credentials.
105
110
  3. **Design System & Viewport Density Pass** (if PR modifies frontend/rendered UI):
106
111
  - **Token Purity**: Check for arbitrary CSS/Tailwind sizing overrides (e.g. `text-[...px]`, `w-[...px]`) or bespoke button styling bypassing standard design system tokens.
@@ -108,7 +113,7 @@ Check if `$PR_NUMBER` is set:
108
113
  - **CTA Hierarchy**: Verify at most one primary forward CTA (`.btn-primary`) per active screen/tab.
109
114
  - **Mobile Viewport Budget**: Verify that new banners, nudges, or sticky elements do not stack concurrently above the fold on mobile viewports (~390px).
110
115
  4. If PR modifies rendered UI, verify screenshots or visual components if tooling/scripts are available.
111
- 5. Run repository verification commands (tests, type-check) if CI status is unconfirmed.
116
+ 5. Run repository verification commands (tests, type-check) ONLY if CI status is unconfirmed or failing. If remote CI (such as GitHub Actions or Vercel preview) is already green and passing on the head commit, treat CI as verified. If verification commands must be run locally, specify `WaitMsBeforeAsync: 10000` and NEVER call `schedule` or yield the turn with plain text while waiting for background tasks.
112
117
 
113
118
  ### Step 5: Classify Findings & Make Decision
114
119