jonah-fleet 1.5.0 → 1.8.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +136 -0
- package/README.md +25 -2
- package/dist/commands/daemon.d.ts.map +1 -1
- package/dist/commands/init.d.ts +3 -0
- package/dist/commands/init.d.ts.map +1 -1
- package/dist/commands/labels.d.ts +11 -0
- package/dist/commands/labels.d.ts.map +1 -0
- package/dist/commands/status.d.ts.map +1 -1
- package/dist/index.js +2292 -476
- package/dist/lib/daemon-keys.d.ts +103 -0
- package/dist/lib/daemon-keys.d.ts.map +1 -0
- package/dist/lib/daemon.d.ts +41 -2
- package/dist/lib/daemon.d.ts.map +1 -1
- package/dist/lib/evals.d.ts +56 -1
- package/dist/lib/evals.d.ts.map +1 -1
- package/dist/lib/labels.d.ts +48 -0
- package/dist/lib/labels.d.ts.map +1 -0
- package/dist/lib/manifest.d.ts +6 -0
- package/dist/lib/manifest.d.ts.map +1 -1
- package/dist/lib/presets.d.ts +47 -1
- package/dist/lib/presets.d.ts.map +1 -1
- package/dist/lib/runner.d.ts +66 -0
- package/dist/lib/runner.d.ts.map +1 -1
- package/dist/lib/telemetry.d.ts +13 -0
- package/dist/lib/telemetry.d.ts.map +1 -1
- package/dist/lib/terminal-card.d.ts +23 -0
- package/dist/lib/terminal-card.d.ts.map +1 -1
- package/package.json +4 -4
- package/schema.json +68 -0
- package/templates/evals/ambiguity-benchmark.json +92 -0
- package/templates/prompts/ORCHESTRATION.md +57 -36
- package/templates/prompts/_prompt-template.md +13 -1
- package/templates/prompts/autowork.md +63 -27
- package/templates/prompts/issues-housekeeping.md +1 -1
- package/templates/prompts/optimizer.md +3 -0
- package/templates/prompts/peer-review.md +13 -3
- package/templates/workflows/autowork-cron.yml +74 -2
- package/templates/workflows/dependency-check-cron.yml +41 -2
- package/templates/workflows/issues-housekeeping-cron.yml +41 -2
- package/templates/workflows/prompt-optimizer-cron.yml +41 -2
- package/templates/workflows/trigger-autowork-manual.yml +43 -2
- package/templates/workflows/trigger-autowork-on-bug.yml +47 -7
- package/templates/workflows/trigger-autowork-on-merge.yml +48 -7
- package/templates/workflows/trigger-review-routine.yml +46 -7
|
@@ -2,20 +2,20 @@
|
|
|
2
2
|
|
|
3
3
|
## Objective
|
|
4
4
|
|
|
5
|
-
This routine runs in two modes, decided in Step 0. In **Scan mode** (a scheduled run, no issue named): converge on existing open work before starting anything new — priority order (1) address review comments on open PRs, (2) close issues whose PRs are merged, (3) only then pick a new issue. In **Targeted mode** (fired with a specific issue in the payload): work
|
|
5
|
+
This routine runs in two modes, decided in Step 0. In **Scan mode** (a scheduled run, no issue named): converge on existing open work before starting anything new — priority order (1) address review comments on open PRs, (2) close issues whose PRs are merged, (3) only then pick a new issue. In **Targeted mode** (fired with a specific issue in the payload): work _that_ issue as the run's objective, **ahead of** the convergence steps above — the fire exists to start its issue immediately, so an unrelated pending PR does not preempt it (Step 0.5); fall back to the Scan flow only if the target is ineligible. Either way, at most one issue may be **implemented** (code written, branch pushed) per run — the sole exception is batching up to 3 same-recipe slices of a single _umbrella_ issue into one child issue + PR (step 12a); that batch is still one concern, not a second issue. Evaluating a candidate and finding it infeasible does not count as "working" it: in Scan mode, step 12's infeasible-continuation cap lets a run evaluate up to 3 candidates for feasibility before it must stop, so a single blocked issue can't consume an entire run without any other progress being attempted.
|
|
6
6
|
|
|
7
7
|
## Definition of Done
|
|
8
8
|
|
|
9
|
-
This routine runs in two modes (Step 0): **Targeted** (a fire named an issue) and **Scan** (scheduled / no issue named). In **Targeted mode**, the run is SUCCESS if you claimed and implemented the target issue to a pushed draft PR (or documented why it is infeasible and released the claim, or ended via step 13's **collision bail** — a competing open PR discovered immediately before opening yours: branch pushed, claim comment annotated, no unassign, no second PR) — or, when the target was
|
|
9
|
+
This routine runs in two modes (Step 0): **Targeted** (a fire named an issue) and **Scan** (scheduled / no issue named). In **Targeted mode**, the run is SUCCESS if you claimed and implemented the target issue to a pushed draft PR (or documented why it is infeasible and released the claim, or ended via step 13's **collision bail** — a competing open PR discovered immediately before opening yours: branch pushed, claim comment annotated, no unassign, no second PR) — or, when the target was _ineligible_ (closed / has an open PR / claimed by a live run), you fell back to the Scan flow and met the Scan criteria below.
|
|
10
10
|
|
|
11
11
|
In **Scan mode**, the run is SUCCESS only if ALL of these are true:
|
|
12
12
|
|
|
13
|
-
- [ ] Checked all open PRs for unresolved or unaddressed review comments
|
|
13
|
+
- [ ] Checked all open PRs for unresolved or unaddressed review comments; if actionable findings were found, claimed the PR before writing fixes (see the Phase 1 PR Claim protocol), confirmed sole ownership, and pushed fixes
|
|
14
14
|
- [ ] Closed any issues whose corresponding PRs are all merged
|
|
15
|
-
- [ ] If no open PRs needed attention: picked the highest-priority
|
|
15
|
+
- [ ] If no open PRs needed attention: picked the highest-priority _unclaimed_ issue (P1 > P2 > P3); if it turned out to be already-done (step 10 — all its PRs merged), closed it and moved on to the next-priority candidate rather than stopping there; otherwise **claimed it before starting work** (see the Claim protocol), and either implemented a fix and **opened a draft pull request on GitHub via `gh pr create --draft`** or left a comment explaining why autonomous completion is blocked and released the claim, then repeated candidate selection for the next-priority issue per step 12's infeasible-continuation cap
|
|
16
16
|
- [ ] If an issue was implemented: successfully opened a draft pull request on GitHub via `gh pr create --draft` referencing the issue (`Closes #N`) in its body, verified the PR exists (a returned PR URL is mandatory), updated the issue's `## Tasks` checkboxes (`- [ ]` → `- [x]`) for every deliverable the PR ships, marked the PR ready for review (`gh pr ready <PR>`), and executed step 15's in-session review wait (never stop at merely pushing the branch or editing the issue; a pushed branch with no open PR on GitHub is a fatal invariant violation and must be logged as FAILURE)
|
|
17
|
-
- [ ] Did not open a new PR while any existing PR by this routine has unaddressed review comments (a finding you have replied to with a rationale counts as
|
|
18
|
-
- [ ] Never worked an issue that was already actively claimed (assigned) by another live run, and confirmed sole ownership of the claim before writing any code — reclaiming a
|
|
17
|
+
- [ ] Did not open a new PR while any existing PR by this routine has unaddressed review comments (a finding you have replied to with a rationale counts as _addressed_, even if the thread is still technically unresolved) — **Targeted mode is exempt**
|
|
18
|
+
- [ ] Never worked an issue or PR that was already actively claimed (assigned) by another live run, and confirmed sole ownership of the claim before writing any code — reclaiming a _stale_ claim (a dead run's orphaned assignment, per step 10a / `ORCHESTRATION.md`) is permitted
|
|
19
19
|
- [ ] If work was an umbrella slice/batch (step 12a): the child issue created for it carries the required Summary/Tasks/Why/Complexity template and exactly one type/size/priority label, and the `🧭 Decomposition plan` markers were left `🔍 in review` until the child PR merges
|
|
20
20
|
|
|
21
21
|
If any criterion cannot be met, stop immediately and log FAILURE with the reason.
|
|
@@ -26,18 +26,19 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
|
|
|
26
26
|
- **Max scope**: at most one issue may be **implemented** per run. Do not pick up a second, unrelated issue to implement after finishing (or abandoning mid-implementation) the first. (Umbrella slice batching up to 3 slices per step 12a is allowed).
|
|
27
27
|
- **No mid-run context switch**: priorities are evaluated ONCE, at the start of the run (Phase 1 → Phase 2). Finish the issue you started; the next run's Phase 1 will pick up newly surfaced work.
|
|
28
28
|
- **No speculative work**: only take actions directly required by the Definition of Done. Do not refactor adjacent code, open bonus issues, or add improvements not requested by the issue.
|
|
29
|
-
- **Single-flight per issue**: multiple autowork runs can execute concurrently.
|
|
29
|
+
- **Single-flight per issue & PR**: multiple autowork runs can execute concurrently. Both issues and pull requests are shared resources — never begin implementing an issue without first claiming it (see Claim protocol in Phase 2), and never begin addressing findings on an open PR without first claiming it (see PR Claim protocol in Phase 1).
|
|
30
30
|
- **Language Requirement**: All GitHub issue titles, descriptions, task checklists, and comments MUST be written in **English**.
|
|
31
31
|
- **Session link footer**: sign every GitHub post (issue comments, PR comments, PR descriptions) with the Antigravity run footer (`_Generated by [Antigravity](${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}/actions/runs/${GITHUB_RUN_ID})_`). Inline review-thread line comments are exempt.
|
|
32
32
|
|
|
33
33
|
## Negative examples (DO NOT do these)
|
|
34
34
|
|
|
35
35
|
- Do not address a subset of review findings and stop — handle ALL findings in one run before stopping or marking the PR as ready.
|
|
36
|
-
- Do not mark a PR as ready for review without first re-fetching review threads
|
|
36
|
+
- Do not mark a PR as ready for review without first re-fetching review threads _and PR-level comments_ and confirming every finding matching step 2's trust & noise rules has been handled.
|
|
37
|
+
- Do not start addressing review findings on a PR before claiming it (both PR assignment AND claim comment).
|
|
37
38
|
- Do not strand a PR in draft because you disagree with a finding — reply with your rationale, mark the PR ready, and let peer-review re-evaluate.
|
|
38
39
|
- Do not push a branch and stop without running `gh pr create --draft` to actually open the pull request — a pushed branch with no open PR cannot be picked up by the peer-review routine.
|
|
39
40
|
- Do not open a new PR if you already have 3+ open PRs — converge before creating more (**Scan mode only**; Targeted mode is exempt).
|
|
40
|
-
- Do not
|
|
41
|
+
- Do not _implement_ multiple _unrelated_ issues in a single run.
|
|
41
42
|
- Do not close an issue just because it is old — only close if the work is done and PRs are merged.
|
|
42
43
|
- Do not attempt an issue that requires environment secrets, manual testing, or external service setup — mark it as infeasible with a comment.
|
|
43
44
|
- Do not start implementing an issue before claiming it (both assignment AND claim comment).
|
|
@@ -50,35 +51,55 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
|
|
|
50
51
|
### Step 0: Determine Targeted vs Scan mode (before Phase 1)
|
|
51
52
|
|
|
52
53
|
This routine runs in two modes, decided here before any other work:
|
|
53
|
-
|
|
54
|
+
|
|
55
|
+
- **Targeted mode** — an environment variable `$ISSUE_NUMBER` (or explicit issue payload) names a target issue. Work _that_ issue as the run's objective, **ahead of** Phase 1 convergence and the normal priority scan.
|
|
54
56
|
- **Scan mode** — no issue is named (a scheduled cron run). Run Phase 1, then select an issue by priority in Phase 2.
|
|
55
57
|
|
|
56
58
|
Check if `$ISSUE_NUMBER` environment variable is set (or scan invocation text for `#<number>` / `/issues/<number>`):
|
|
59
|
+
|
|
57
60
|
- **Found one → Targeted mode.** Record it as this run's target issue and proceed to Step 0.5.
|
|
58
61
|
- **None found → Scan mode.** Proceed as a scheduled run: Phase 1, then normal priority selection in Phase 2.
|
|
59
62
|
|
|
60
63
|
### Step 0.5: Targeted mode — work the target issue first (only when Step 0 found one)
|
|
61
64
|
|
|
62
65
|
a. **Read the target issue and check eligibility.** Eligible = open, unassigned or reclaimable stale claim, no open PR (`Closes #N`), no unclosed inward blocking dependencies (`Blocked by #N` or `Depends on #N` where `#N` is open), and not carrying `needs-human` or `needs-design`.
|
|
63
|
-
|
|
64
|
-
|
|
66
|
+
|
|
67
|
+
- **Ineligible** → fall back to Phase 1 and run the normal Scan flow.
|
|
68
|
+
b. **Eligible → claim, then implement.** Call `get_me` once to learn your own login (step 7), reclaim stale claim if applicable (step 10a), run Claim protocol (step 11), evaluate and implement per steps 12–13, open draft PR, and run in-session review wait (step 15).
|
|
65
69
|
|
|
66
70
|
### Phase 1: Converge on open work (Scan mode; skipped in Targeted mode)
|
|
67
71
|
|
|
68
72
|
1. List all open PRs authored by this routine.
|
|
69
|
-
|
|
73
|
+
1a. **PR Claim Protocol (Single-Flight PR Convergence)**:
|
|
74
|
+
- For each open PR needing attention (bounced to draft with actionable findings, or open with unaddressed review comments):
|
|
75
|
+
- **Check eligibility**: re-read candidate PR (`gh pr view <PR> --json assignees,comments`). Skip PRs currently assigned to another live runner or carrying an active claim comment posted within the last 2 hours (unless stale claim per `ORCHESTRATION.md`).
|
|
76
|
+
- Call `get_me` once to learn your own login if not already known.
|
|
77
|
+
- **Claim atomically**: assign yourself to the PR (`gh pr edit <PR> --add-assignee <login>`) AND post a claim comment on the PR:
|
|
78
|
+
- In **Cloud Actions**: `🔒 Addressing review findings by autowork run {run_url} {timestamp}`
|
|
79
|
+
- In **Local Agent**: `🔒 Addressing review findings by local autowork session (host: {hostname}) {timestamp}`
|
|
80
|
+
- **Confirm sole ownership**: fetch comments on the PR. Earliest `created_at` among competing claim comments in this round wins. If you lost the race, unassign yourself (`gh pr edit <PR> --remove-assignee <login>`), annotate your comment, and evaluate the next PR (or proceed to Phase 2 if none left).
|
|
81
|
+
2. For the claimed PR, fetch ALL review threads and PR-level comments. Filter for actionable findings matching trust & noise rules (author login matching `get_me`, non-noise). Address every single one in this run — push fixes for actionable findings, reply to clarifying questions, and comment on deferred items.
|
|
70
82
|
3. For PRs where you have addressed all findings:
|
|
71
|
-
3a. **Pre-ready self-audit** (run before marking ready):
|
|
83
|
+
3a. **Pre-ready self-audit** (run before marking ready):
|
|
72
84
|
- **Automated review passes**: run `/code-review` (evaluating along Standards in `AGENTS.md` and Spec in the issue's `## Tasks`) and security review over the diff. Fix what they flag.
|
|
73
85
|
- **Repository conventions scan**: read and verify all rules and conventions specified in `AGENTS.md` (or `CLAUDE.md`/`GEMINI.md`), project-level skills in `.agents/skills/`, and project documentation.
|
|
74
86
|
- **Design System & Viewport Pre-flight** (for frontend/UI diffs): self-audit diffs against design tokens (no arbitrary class overrides), WCAG AA 4.5:1 contrast ratios on dark/light surfaces, single primary CTA hierarchy per screen, and mobile viewport crowding (avoid stacked nudges/banners above the fold at ~390px).
|
|
75
87
|
- **Documentation accuracy**: update relevant docs (`ARCHITECTURE.md`, `CODEMAP.md`, `API.md`, `CHANGELOG.md` if maintained by repo).
|
|
76
88
|
- **Build & type-check verification**: run the repository's test, type-check, and lint commands from `AGENTS.md` (e.g. `npm test`, `npm run type-check`, `npm run lint`, `pytest`, `cargo test`). Confirm zero errors and zero test failures.
|
|
77
89
|
- **Clean-merge gate**: verify `git merge-tree origin/main HEAD` reports no conflicts.
|
|
78
|
-
|
|
79
|
-
|
|
80
|
-
|
|
81
|
-
|
|
90
|
+
- **Release claim on ready**: mark the PR ready (`gh pr ready <PR>`) and unassign yourself (`gh pr edit <PR> --remove-assignee <login>`) so Peer Review can evaluate without holding stale agent reservation locks.
|
|
91
|
+
Only mark the PR ready after passing every check above.
|
|
92
|
+
3b. **Ping-pong cap**: If this same PR has bounced between draft and ready 3 or more times over the same substantive finding, stop re-marking it ready. Post a comment summarizing the disagreement for human resolution and leave the PR in draft.
|
|
93
|
+
3c. **Orphaned Ready PR Recovery**: If an open PR authored by this routine is `ready_for_review`, has passing CI, no unaddressed review comments, and has received no review activity for over 2 hours (e.g. because peer review crashed or encountered quota limits), kickstart the review routine by posting `/review` comment or toggling draft and ready (`gh pr ready <PR> --undo && gh pr ready <PR>`).
|
|
94
|
+
- **Passing CI Verification Gate**: Verify via `gh pr view <PR> --json statusCheckRollup,mergeStateStatus` that all required and existing checks have completed with `conclusion: "SUCCESS"` and `mergeStateStatus` is `CLEAN` (neither `UNSTABLE`, `BLOCKED`, nor `DIRTY`).
|
|
95
|
+
- **Unapproved/Pending Workflow Invariant**: NEVER post `/review` or toggle draft state if checks are in-progress, failing, or awaiting approval (`conclusion: "ACTION_REQUIRED"`). Doing so creates an infinite comment storm while workflows remain paused awaiting human permissions.
|
|
96
|
+
4. **Check open issues with merged PRs / work done — close them**:
|
|
97
|
+
- For open issues, check if recent git history or merged PRs reference the issue (e.g. `git log -n 50 --grep="#<ISSUE_NUMBER>"` or `gh pr list --state merged --search "<ISSUE_NUMBER>" --limit 5`).
|
|
98
|
+
- If a merged PR or commit on `main` resolved the issue, close it immediately:
|
|
99
|
+
```bash
|
|
100
|
+
gh issue close <ISSUE_NUMBER> --comment "Closed via merged PR #<PR_NUMBER> (found in git history)."
|
|
101
|
+
```
|
|
102
|
+
|
|
82
103
|
5. If any PR was updated in this phase, STOP — run is SUCCESS.
|
|
83
104
|
|
|
84
105
|
### Phase 2: New work (only if Phase 1 had nothing to do)
|
|
@@ -94,14 +115,14 @@ b. **Eligible → claim, then implement.** Call `get_me` once to learn your own
|
|
|
94
115
|
- If running in **Local Agent** (`$LOCAL_AGENT`): Scan mode selects across all priorities (P0 → P1 → P2 → P3) without time gating, prioritizing active local backlog consumption.
|
|
95
116
|
- Skip assigned issues (unless stale claim per `ORCHESTRATION.md`), issues with open PRs, issues with unclosed blocking dependencies (`Blocked by #N` / `Depends on #N`), issues labeled `needs-human`, `needs-design`, or `needs-info`, and issues under cross-run cooldown.
|
|
96
117
|
9. If no eligible candidate exists, STOP — run is SUCCESS with "No unclaimed work available".
|
|
97
|
-
10. If the candidate should be closed already (work done,
|
|
98
|
-
10a. **Stale-claim reclamation:** If candidate is a stale claim per `ORCHESTRATION.md`, re-read immediately before writing, unassign the dead owner, post reclamation comment, and proceed to claim.
|
|
118
|
+
10. If the candidate should be closed already (work done, merged PR found in git history or via `gh pr list --state merged --search "<ISSUE_NUMBER>"`), close it immediately (`gh issue close <ISSUE_NUMBER> --comment "Closed: work already merged in PR #<PR_NUMBER>."`) and return to step 8.
|
|
119
|
+
10a. **Stale-claim reclamation:** If candidate is a stale claim per `ORCHESTRATION.md`, re-read immediately before writing, unassign the dead owner, post reclamation comment, and proceed to claim.
|
|
99
120
|
11. **Claim protocol:**
|
|
100
121
|
a. Re-read candidate issue immediately before claiming (`issue_read`). If assigned, abort and pick next candidate.
|
|
101
122
|
b. Claim atomically: assign yourself (`login` from step 7) AND post claim comment:
|
|
102
|
-
|
|
103
|
-
|
|
104
|
-
|
|
123
|
+
- In **Cloud Actions**: `🔒 Claimed by autowork run {run_url} {timestamp}`
|
|
124
|
+
- In **Local Agent**: `🔒 Claimed by local autowork session (host: {hostname}) {timestamp}`
|
|
125
|
+
c. Confirm sole ownership by counting `🔒 Claimed by autowork run` / `🔒 Claimed by local autowork session` comments. Earliest `created_at` wins. If you lost the race, leave assignee as is, annotate your comment, and pick next candidate.
|
|
105
126
|
12. Evaluate whether the claimed issue can be completed autonomously:
|
|
106
127
|
- Confirm your `🔒` claim comment is present on the issue.
|
|
107
128
|
- Read the issue description, linked code, and comment thread.
|
|
@@ -114,7 +135,7 @@ b. **Eligible → claim, then implement.** Call `get_me` once to learn your own
|
|
|
114
135
|
- Select the next candidate (evaluating ambiguous issues counts toward step 12's infeasible-continuation cap).
|
|
115
136
|
- **Intent vs. Defect Guardrail**: When investigating issues related to low conversion, zero-click events, or underperforming features: verify whether the issue is a software defect or a lack of user intent. If data indicates the root cause is **lack of user intent** (e.g. button is rendered above fold and functions correctly when clicked, but user interaction rate is <2%) rather than a software defect, do NOT fall into the **telemetry rabbit hole** (adding elaborate fallback telemetry, downstream error handling, or defensive rendering). Categorize the issue as a **product/UX question** (`needs-design` / `roadmap/*`), comment explaining the lack of user intent, release the claim (unassign), and select the next candidate.
|
|
116
137
|
- If infeasible: comment explaining blocker, release claim (unassign), and select next candidate (up to 3 infeasible evaluations per run). If permanent blocker on 2nd strike, apply `needs-human` label and tag repo owner.
|
|
117
|
-
12a. **Umbrella-issue handoff + batching:** If candidate is an umbrella epic:
|
|
138
|
+
12a. **Umbrella-issue handoff + batching:** If candidate is an umbrella epic:
|
|
118
139
|
- Read `🧭 Decomposition plan` comment (or create if first run).
|
|
119
140
|
- Pick next slice(s), batching up to 3 same-recipe slices into one child issue + PR.
|
|
120
141
|
- Create and claim child issue first, then update plan marker to `🚧 in progress — child #M`, release umbrella claim, and implement against child.
|
|
@@ -126,20 +147,35 @@ b. **Eligible → claim, then implement.** Call `get_me` once to learn your own
|
|
|
126
147
|
- Open draft PR via `gh pr create --draft --head <branch> --base main --title "<title>" --body "<body referencing Closes #N>"`.
|
|
127
148
|
- **PR Priority Label Mirroring**: If the issue carried a priority label (`priority/P0`, `priority/P1`, `priority/P2`, `priority/P3`), add the identical priority label to the PR (`gh pr edit <PR> --add-label "<label>"` or via `--label` in create) so downstream review workflows can filter triggers immediately.
|
|
128
149
|
- Verify PR URL returned. Update issue `## Tasks` checkboxes.
|
|
150
|
+
- **Issue Cross-Reference Comment Guardrail**: Post an explicit comment on the tracking issue referencing the newly created PR (`gh issue comment <ISSUE_NUMBER> --body "Work in progress in PR #<PR_NUMBER>."`). This guarantees an unambiguous, permanent link on the issue timeline even when GitHub's native UI link is suppressed for bot draft PRs.
|
|
151
|
+
- **Warm Context Assignment**: Assign yourself to the newly opened PR (`gh pr edit <PR> --add-assignee <login>`) to hold the reservation across Step 15's in-session review wait so parallel Scan routines recognise the PR as actively held by a live session.
|
|
129
152
|
- Mark PR ready (`gh pr ready <PR>`).
|
|
130
153
|
14. If run aborts before opening PR, release claim (unassign).
|
|
131
154
|
15. **In-Session Peer Review Wait & Immediate Convergence (Warm Context):**
|
|
132
155
|
- Poll PR status for up to 10–12 minutes.
|
|
133
|
-
- If merged: record terminal SUCCESS and exit cleanly.
|
|
156
|
+
- If merged: unassign yourself, record terminal SUCCESS and exit cleanly.
|
|
134
157
|
- If bounced to draft with findings: fetch review comments, apply fixes in active worktree, run tests, push fix commit, re-mark ready (`gh pr ready <PR>`), and complete run.
|
|
135
|
-
- If timeout (>12m): exit cleanly; scheduled cron will handle subsequent rounds.
|
|
158
|
+
- If timeout (>12m): unassign yourself from the PR (`gh pr edit <PR> --remove-assignee <login>`) and exit cleanly; scheduled cron will handle subsequent rounds.
|
|
136
159
|
|
|
137
160
|
## Logging
|
|
138
161
|
|
|
139
162
|
After completing (SUCCESS or FAILURE), write a log file to `.github/prompts/logs/autowork/{timestamp}.md` following the schema in `.github/prompts/logs/_template.md`. Include:
|
|
163
|
+
|
|
140
164
|
- The prompt SHA (run `git rev-parse --short HEAD:.github/prompts/autowork.md`)
|
|
141
165
|
- Every Definition of Done criterion with YES/NO and evidence
|
|
142
166
|
- Full execution trace with tool calls
|
|
143
167
|
- If FAILURE: root cause, category, and suggested fix
|
|
144
168
|
|
|
145
|
-
**
|
|
169
|
+
**Log Delivery Protocol & Invariants**:
|
|
170
|
+
|
|
171
|
+
- **Negative Rule**: NEVER commit or push run logs to a feature branch or open PR branch. Doing so emits a `pull_request: synchronize` event under bot credentials, triggering GitHub Actions workflow approval gates (`action_required`) that stall CI.
|
|
172
|
+
- **Mandatory `[skip ci]`**: Always append `[skip ci]` to any log commit message.
|
|
173
|
+
- **Direct Push to `main`**: Commit the log file directly to `main` and push — explicitly permitted for files under `.github/prompts/logs/**`:
|
|
174
|
+
```bash
|
|
175
|
+
git checkout main
|
|
176
|
+
git pull origin main
|
|
177
|
+
git add .github/prompts/logs/autowork/{timestamp}.md
|
|
178
|
+
git commit -m "docs(log): record autowork run {timestamp} [skip ci]"
|
|
179
|
+
git push origin main
|
|
180
|
+
```
|
|
181
|
+
Follow the Log delivery fallback in `ORCHESTRATION.md` if direct push fails.
|
|
@@ -39,7 +39,7 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
|
|
|
39
39
|
3. **Priority review**: Check open P1/P2/P3 issues. Promote critical bugs or unblocked items; demote items that lack immediate priority.
|
|
40
40
|
4. **Duplicate & consolidation check**: Identify duplicate issues; close duplicates with cross-references. Consolidate small, related micro-tasks into batch issues.
|
|
41
41
|
5. **Premise-obsolete & stale check**: If an issue's premise was resolved by already-merged PRs or recent refactors, close as completed with evidence.
|
|
42
|
-
6. **Label audit**: Ensure open issues carry standard role labels (`needs-triage`, `ready-for-agent`, `needs-human`, etc.). Use `/triage` if classifying incoming issues.
|
|
42
|
+
6. **Label audit & safe prune**: Ensure open issues carry standard role labels (`needs-triage`, `ready-for-agent`, `needs-human`, etc.). Use `/triage` if classifying incoming issues. Run `npx --yes jonah-fleet labels prune --yes` (or `jonah-fleet labels prune --yes`) to safely prune strictly unused boilerplate labels (`issues: 0`, `pullRequests: 0`, non-protected taxonomy) without deleting historical or fleet taxonomy labels.
|
|
43
43
|
7. **Closed-loop verification check**: For projects running impact or verification loops, audit recently closed roadmap/feature issues against tracking issues to ensure shipped levers do not remain untracked.
|
|
44
44
|
|
|
45
45
|
### Phase 3: Summary
|
|
@@ -59,6 +59,8 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
|
|
|
59
59
|
- **Iteration Ceiling Exhaustion**: >20% of runs in a routine terminate at the `token_limit` / max iteration cap.
|
|
60
60
|
- **Review Loop Burn**: Pull requests experiencing >= 3 bounce rounds between autowork and peer-review over unresolved or recurring findings.
|
|
61
61
|
- **Feedback Loop Stagnation**: A downstream processing routine (e.g. impact measurement, verification, or triage) records 0 intake (`filed: 0` or 0 new items processed) across $\ge 2$ consecutive runs while upstream PRs merge or roadmap/feature issues close in the same window. Flags that discovery sweeps have stalled or become overly coarse.
|
|
62
|
+
- **Passive Order-Taking Anomaly ("Yes-Man Blindspot")**: The Ambiguity Gate trigger rate across intake runs in `autowork` or `triage` is <5% despite elevated PR review bounces ($\ge 2$) or high iteration usage ($\ge 35$), indicating agents are silently guessing requirements and building flawed implementations rather than interrogating underspecified issues.
|
|
63
|
+
- **Speculative Runaway Waste**: An agent run consumed >50k tokens on an underspecified issue with 0 clarifying questions asked, and subsequently failed, bounced, or required post-merge rework.
|
|
62
64
|
6. **Analyze resolved bugs & review comments**: Examine closed bug issues, merged bug-fix PRs, and review feedback for missing checks in authoring (`autowork.md`) or review (`peer-review.md`).
|
|
63
65
|
|
|
64
66
|
### 2. Formulate preventative improvements
|
|
@@ -69,6 +71,7 @@ Translate findings into concrete preventative improvements and remediation trigg
|
|
|
69
71
|
- **Iteration Ceiling & Self-Audit Tuning**: For Iteration Ceiling Exhaustion, adjust max iteration bounds or tighten pre-ready self-audits in `autowork.md` to catch defects before review cycles start.
|
|
70
72
|
- **Ping-Pong Convergence**: For Review Loop Burn, tighten reviewer trust & noise filtering, enforce clean-merge gates, and apply ping-pong caps to prevent endless bounce cycles.
|
|
71
73
|
- **Loop Discovery Mechanical Audits**: For Feedback Loop Stagnation, tighten discovery sweeps by mandating deterministic per-issue matching tables and itemized reconciliation against upstream closed issues/PRs rather than allowing un-itemized generic summary assertions.
|
|
74
|
+
- **Ambiguity Gate & Benchmark Eval Feeding**: For Passive Order-Taking and Speculative Runaway Waste, tighten Step 12 criteria in `autowork.md` and `triage.md` to mandate clarifying questions, and automatically extract the problem issue into a `BenchmarkIssue` test case to feed the automated ambiguity benchmark eval suite (`tests/evals.test.ts`), ensuring future agent prompts are continuously tested against real failure cases.
|
|
72
75
|
- **Verification & Invariant Tests**: Add automated test cases in `tests/` verifying prompt invariant preservation and schema conformity.
|
|
73
76
|
|
|
74
77
|
### 3. Open Fix PR (Local or Upstream Bridge)
|
|
@@ -8,7 +8,7 @@ Review a pull request (in Targeted mode for a specific `$PR_NUMBER`, or in Scan
|
|
|
8
8
|
|
|
9
9
|
Before reading further, before any tool call, and before deciding the mode, check the environment variable `$PR_NUMBER`.
|
|
10
10
|
|
|
11
|
-
- **$PR_NUMBER is set → Targeted mode.** The value of `$PR_NUMBER`
|
|
11
|
+
- **$PR_NUMBER is set → Targeted mode.** The value of `$PR_NUMBER`is your target PR. You may also check`$PR_URL`. Skip all selection logic.
|
|
12
12
|
- **$PR_NUMBER is not set → Scan mode.** Only then select a PR by priority.
|
|
13
13
|
|
|
14
14
|
**Targeted mode is sticky: it can never fall back to Scan mode.** Once the invocation contains a PR reference, you must review exactly that PR. If you cannot act on it (closed/merged/missing), STOP and log FAILURE.
|
|
@@ -19,7 +19,7 @@ The run is SUCCESS only if ALL of these are true:
|
|
|
19
19
|
|
|
20
20
|
- [ ] Identified the target PR: if one was named in the invocation, reviewed exactly that PR; otherwise listed open PRs and selected one by priority
|
|
21
21
|
- [ ] Ran the code-review pass (`/code-review` and security pass), and posted findings as inline review comments
|
|
22
|
-
- [ ] Took exactly one final action: squash-merged (if PR is good, CI green and present; executed Autonomous Issue Synthesis if unlinked) OR posted findings and **converted the PR back to draft** (`gh pr ready <N> --undo`) for author/autowork in-session fixes OR, if round cap reached at round 5 with blocking findings, converted to draft and escalated to human
|
|
22
|
+
- [ ] Took exactly one final action: squash-merged (if PR is good, CI green and present; executed Autonomous Issue Synthesis if unlinked; closed tracking issue explicitly if referenced) OR posted findings and **converted the PR back to draft** (`gh pr ready <N> --undo`) for author/autowork in-session fixes OR, if round cap reached at round 5 with blocking findings, converted to draft and escalated to human
|
|
23
23
|
- [ ] If merging: captured deferred non-blocking findings per materiality bar (filed follow-up issues for material ones, batched or dropped immaterial ones)
|
|
24
24
|
- [ ] If in Scan mode and no eligible PRs exist, logged SUCCESS with "No PRs to review"
|
|
25
25
|
|
|
@@ -37,8 +37,9 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
|
|
|
37
37
|
## Final action: merge or bounce to draft
|
|
38
38
|
|
|
39
39
|
Every review ends in exactly one of two states:
|
|
40
|
+
|
|
40
41
|
- **Merge** — only if PR is good, CI is green and verified on the head commit. If the PR does not reference a tracked issue (`Closes #N`), execute Autonomous Issue Synthesis prior to merge.
|
|
41
|
-
- Sequence: (1) if unlinked, synthesize tracking issue (`gh issue create`) and link to PR (`gh pr edit`), (2) squash-merge, (3) submit held review comments, (
|
|
42
|
+
- Sequence: (1) if unlinked, synthesize tracking issue (`gh issue create`) and link to PR (`gh pr edit`), (2) squash-merge (`gh pr merge <N> --squash --delete-branch`), (3) explicitly close tracking issue if referenced (`gh issue close <ISSUE_NUMBER>`), (4) submit held review comments, (5) file follow-up issues for deferred material findings.
|
|
42
43
|
- Immaterial findings (style/preference) default to dying in the review thread or getting batched.
|
|
43
44
|
- Mechanical doc fixes (missing changelog line, doc typo in diff) can be committed directly to `main` after squash-merge.
|
|
44
45
|
- **Bounce to draft** — if any **blocking** finding remains (correctness bug, security flaw, failing/missing CI, broken contract):
|
|
@@ -63,6 +64,7 @@ Every review ends in exactly one of two states:
|
|
|
63
64
|
### Step 0: Determine the target PR (do this FIRST)
|
|
64
65
|
|
|
65
66
|
Check if `$PR_NUMBER` is set:
|
|
67
|
+
|
|
66
68
|
- **$PR_NUMBER is set → Targeted mode.** Review that exact PR. Skip selection steps 1–2.
|
|
67
69
|
- **$PR_NUMBER is not set → Scan mode.** Proceed to steps 1–2.
|
|
68
70
|
|
|
@@ -98,12 +100,14 @@ Check if `$PR_NUMBER` is set:
|
|
|
98
100
|
### Step 5: Classify Findings & Make Decision
|
|
99
101
|
|
|
100
102
|
Classify each finding:
|
|
103
|
+
|
|
101
104
|
- **Blocking**: Broken logic, security hole, data loss, regression, broken tests, missing deliverable from the issue/PR specification.
|
|
102
105
|
- **Non-blocking**: Minor refactor, style preference, performance micro-optimization, missing `Closes #N` on contributor PRs with self-contained descriptions.
|
|
103
106
|
|
|
104
107
|
### Step 5.5: Autonomous Issue Synthesis (for unlinked PRs)
|
|
105
108
|
|
|
106
109
|
If the PR is clean and approved for merge, but lacks a `Closes #N` tracking link:
|
|
110
|
+
|
|
107
111
|
1. Synthesize a retroactive tracking issue on GitHub:
|
|
108
112
|
```bash
|
|
109
113
|
gh issue create --title "<PR Title>" --body "Tracked retroactively from external pull request #<PR_NUMBER>.\n\n## Deliverables & Context\n<PR Description>\n\n_Synthesized autonomously by Jonah Fleet Peer Review_"
|
|
@@ -121,7 +125,12 @@ If the PR is clean and approved for merge, but lacks a `Closes #N` tracking link
|
|
|
121
125
|
- If `N >= 5`: Convert PR to draft, post summary comment escalating to repo maintainer, and apply `needs-human` label.
|
|
122
126
|
- **If Clean (or only Non-blocking findings)**:
|
|
123
127
|
- If PR lacks `Closes #N`, execute Autonomous Issue Synthesis (Step 5.5).
|
|
128
|
+
- Extract the tracking issue number `$ISSUE_NUMBER` from the PR description or title (e.g. `Closes #<N>`, `Fixes #<N>`, `Resolves #<N>`).
|
|
124
129
|
- Squash-merge the PR: `gh pr merge <N> --squash --delete-branch`.
|
|
130
|
+
- **Explicit Tracking Issue Closure Guardrail**: If a tracking issue was referenced (`$ISSUE_NUMBER`), explicitly close it immediately after merge rather than relying solely on GitHub's native keyword parser (which often fails to trigger on bot-merged squash commits or draft PRs):
|
|
131
|
+
```bash
|
|
132
|
+
gh issue close "$ISSUE_NUMBER" --comment "Closed via PR #<N> (merged into main)."
|
|
133
|
+
```
|
|
125
134
|
- Submit held review comments.
|
|
126
135
|
- File follow-up issues for material non-blocking findings.
|
|
127
136
|
- If mechanical doc fixes are needed, commit directly to `main`.
|
|
@@ -129,6 +138,7 @@ If the PR is clean and approved for merge, but lacks a `Closes #N` tracking link
|
|
|
129
138
|
## Logging
|
|
130
139
|
|
|
131
140
|
After completing (SUCCESS or FAILURE), write a log file to `.github/prompts/logs/peer-review/{timestamp}.md` following the schema in `.github/prompts/logs/_template.md`. Include:
|
|
141
|
+
|
|
132
142
|
- Prompt SHA
|
|
133
143
|
- Target PR number and decision (MERGE / BOUNCE / ESCALATE)
|
|
134
144
|
- Execution trace and findings summary
|
|
@@ -36,6 +36,8 @@ jobs:
|
|
|
36
36
|
uses: actions/checkout@v4
|
|
37
37
|
with:
|
|
38
38
|
fetch-depth: 0
|
|
39
|
+
token: ${{ secrets.GH_PAT || github.token }}
|
|
40
|
+
|
|
39
41
|
|
|
40
42
|
- name: Setup Node.js
|
|
41
43
|
uses: actions/setup-node@v4
|
|
@@ -103,6 +105,45 @@ jobs:
|
|
|
103
105
|
fi
|
|
104
106
|
echo "skills_prompt=$SKILLS_PROMPT" >> "$GITHUB_OUTPUT"
|
|
105
107
|
|
|
108
|
+
- name: Resolve routine configuration
|
|
109
|
+
id: routine-config
|
|
110
|
+
run: |
|
|
111
|
+
ROUTINE="autowork"
|
|
112
|
+
DEFAULT_MODEL="gemini-3.7-flash-high"
|
|
113
|
+
DEFAULT_TIMEOUT="60"
|
|
114
|
+
|
|
115
|
+
MODEL="$DEFAULT_MODEL"
|
|
116
|
+
TIMEOUT="$DEFAULT_TIMEOUT"
|
|
117
|
+
|
|
118
|
+
if [ -f "agents-manifest.json" ]; then
|
|
119
|
+
CONFIG_JSON=$(node -e '
|
|
120
|
+
try {
|
|
121
|
+
const fs = require("fs");
|
|
122
|
+
const m = JSON.parse(fs.readFileSync("agents-manifest.json", "utf8"));
|
|
123
|
+
const r = process.argv[1];
|
|
124
|
+
const model = m.models?.[r] || m.models?.default;
|
|
125
|
+
const timeout = m.budgets?.timeoutMinutes?.[r] || m.budgets?.timeoutMinutes?.default;
|
|
126
|
+
console.log(JSON.stringify({ model: model || "", timeout: timeout ? String(timeout) : "" }));
|
|
127
|
+
} catch (e) {
|
|
128
|
+
console.log(JSON.stringify({ model: "", timeout: "" }));
|
|
129
|
+
}
|
|
130
|
+
' "$ROUTINE")
|
|
131
|
+
|
|
132
|
+
RESOLVED_MODEL=$(node -e "try { console.log(JSON.parse(process.argv[1]).model || ''); } catch {}" "$CONFIG_JSON")
|
|
133
|
+
RESOLVED_TIMEOUT=$(node -e "try { console.log(JSON.parse(process.argv[1]).timeout || ''); } catch {}" "$CONFIG_JSON")
|
|
134
|
+
|
|
135
|
+
if [ -n "$RESOLVED_MODEL" ]; then
|
|
136
|
+
MODEL="$RESOLVED_MODEL"
|
|
137
|
+
fi
|
|
138
|
+
if [ -n "$RESOLVED_TIMEOUT" ]; then
|
|
139
|
+
TIMEOUT="$RESOLVED_TIMEOUT"
|
|
140
|
+
fi
|
|
141
|
+
fi
|
|
142
|
+
|
|
143
|
+
echo "model=$MODEL" >> "$GITHUB_OUTPUT"
|
|
144
|
+
echo "print_timeout=${TIMEOUT}m" >> "$GITHUB_OUTPUT"
|
|
145
|
+
echo "Resolved routine configuration: model=$MODEL, timeout=${TIMEOUT}m"
|
|
146
|
+
|
|
106
147
|
- name: Run Autowork
|
|
107
148
|
timeout-minutes: 60
|
|
108
149
|
env:
|
|
@@ -123,9 +164,9 @@ jobs:
|
|
|
123
164
|
fi
|
|
124
165
|
|
|
125
166
|
agy -p "$PROMPT" \
|
|
126
|
-
--model gemini-3.7-flash-high \
|
|
167
|
+
--model "${{ steps.routine-config.outputs.model || 'gemini-3.7-flash-high' }}" \
|
|
127
168
|
--output-format text \
|
|
128
|
-
--print-timeout 60m \
|
|
169
|
+
--print-timeout "${{ steps.routine-config.outputs.print_timeout || '60m' }}" \
|
|
129
170
|
--dangerously-skip-permissions
|
|
130
171
|
|
|
131
172
|
- name: Emit run summary
|
|
@@ -170,3 +211,34 @@ jobs:
|
|
|
170
211
|
fi
|
|
171
212
|
fi
|
|
172
213
|
fi
|
|
214
|
+
|
|
215
|
+
- name: Commit and push run log directly to main
|
|
216
|
+
if: always()
|
|
217
|
+
env:
|
|
218
|
+
GH_TOKEN: ${{ secrets.GH_PAT || github.token }}
|
|
219
|
+
run: |
|
|
220
|
+
LATEST_LOG=$(ls -t .github/prompts/logs/autowork/*.md 2>/dev/null | head -n 1)
|
|
221
|
+
if [ -n "$LATEST_LOG" ] && [ -f "$LATEST_LOG" ]; then
|
|
222
|
+
TMP_LOG_FILE=$(mktemp)
|
|
223
|
+
cp "$LATEST_LOG" "$TMP_LOG_FILE"
|
|
224
|
+
LOG_REL_PATH="$LATEST_LOG"
|
|
225
|
+
LOG_NAME=$(basename "$LATEST_LOG" .md)
|
|
226
|
+
|
|
227
|
+
git fetch origin main:main || git fetch origin main || true
|
|
228
|
+
git checkout main || git checkout -B main origin/main || true
|
|
229
|
+
|
|
230
|
+
mkdir -p "$(dirname "$LOG_REL_PATH")"
|
|
231
|
+
cp "$TMP_LOG_FILE" "$LOG_REL_PATH"
|
|
232
|
+
rm -f "$TMP_LOG_FILE"
|
|
233
|
+
|
|
234
|
+
git add "$LOG_REL_PATH"
|
|
235
|
+
if ! git diff --cached --quiet; then
|
|
236
|
+
git commit -m "docs(log): record autowork run ${LOG_NAME} [skip ci]" || true
|
|
237
|
+
if [ -n "$GH_TOKEN" ]; then
|
|
238
|
+
git push "https://x-access-token:${GH_TOKEN}@github.com/${GITHUB_REPOSITORY}.git" main || git push origin main || true
|
|
239
|
+
else
|
|
240
|
+
git push origin main || true
|
|
241
|
+
fi
|
|
242
|
+
fi
|
|
243
|
+
fi
|
|
244
|
+
|
|
@@ -65,6 +65,45 @@ jobs:
|
|
|
65
65
|
printf '%s' "$ANTIGRAVITY_OAUTH_TOKEN" > "$HOME/.gemini/antigravity-cli/antigravity-oauth-token"
|
|
66
66
|
chmod 600 "$HOME/.gemini/antigravity-cli/antigravity-oauth-token"
|
|
67
67
|
|
|
68
|
+
- name: Resolve routine configuration
|
|
69
|
+
id: routine-config
|
|
70
|
+
run: |
|
|
71
|
+
ROUTINE="dependency-update-security-check"
|
|
72
|
+
DEFAULT_MODEL="gemini-3.7-flash"
|
|
73
|
+
DEFAULT_TIMEOUT="25"
|
|
74
|
+
|
|
75
|
+
MODEL="$DEFAULT_MODEL"
|
|
76
|
+
TIMEOUT="$DEFAULT_TIMEOUT"
|
|
77
|
+
|
|
78
|
+
if [ -f "agents-manifest.json" ]; then
|
|
79
|
+
CONFIG_JSON=$(node -e '
|
|
80
|
+
try {
|
|
81
|
+
const fs = require("fs");
|
|
82
|
+
const m = JSON.parse(fs.readFileSync("agents-manifest.json", "utf8"));
|
|
83
|
+
const r = process.argv[1];
|
|
84
|
+
const model = m.models?.[r] || m.models?.default;
|
|
85
|
+
const timeout = m.budgets?.timeoutMinutes?.[r] || m.budgets?.timeoutMinutes?.default;
|
|
86
|
+
console.log(JSON.stringify({ model: model || "", timeout: timeout ? String(timeout) : "" }));
|
|
87
|
+
} catch (e) {
|
|
88
|
+
console.log(JSON.stringify({ model: "", timeout: "" }));
|
|
89
|
+
}
|
|
90
|
+
' "$ROUTINE")
|
|
91
|
+
|
|
92
|
+
RESOLVED_MODEL=$(node -e "try { console.log(JSON.parse(process.argv[1]).model || ''); } catch {}" "$CONFIG_JSON")
|
|
93
|
+
RESOLVED_TIMEOUT=$(node -e "try { console.log(JSON.parse(process.argv[1]).timeout || ''); } catch {}" "$CONFIG_JSON")
|
|
94
|
+
|
|
95
|
+
if [ -n "$RESOLVED_MODEL" ]; then
|
|
96
|
+
MODEL="$RESOLVED_MODEL"
|
|
97
|
+
fi
|
|
98
|
+
if [ -n "$RESOLVED_TIMEOUT" ]; then
|
|
99
|
+
TIMEOUT="$RESOLVED_TIMEOUT"
|
|
100
|
+
fi
|
|
101
|
+
fi
|
|
102
|
+
|
|
103
|
+
echo "model=$MODEL" >> "$GITHUB_OUTPUT"
|
|
104
|
+
echo "print_timeout=${TIMEOUT}m" >> "$GITHUB_OUTPUT"
|
|
105
|
+
echo "Resolved routine configuration: model=$MODEL, timeout=${TIMEOUT}m"
|
|
106
|
+
|
|
68
107
|
- name: Run Dependency Update Check
|
|
69
108
|
timeout-minutes: 25
|
|
70
109
|
env:
|
|
@@ -79,7 +118,7 @@ jobs:
|
|
|
79
118
|
PROMPT="You are the Dependency Update & Security Check routine. Read and follow .github/prompts/dependency-update-security-check.md exactly. Check dependencies for updates and vulnerabilities, and create or update tracking issues."
|
|
80
119
|
|
|
81
120
|
agy -p "$PROMPT" \
|
|
82
|
-
--model gemini-3.7-flash
|
|
121
|
+
--model "${{ steps.routine-config.outputs.model || 'gemini-3.7-flash' }}" \
|
|
83
122
|
--output-format text \
|
|
84
|
-
--print-timeout 25m \
|
|
123
|
+
--print-timeout "${{ steps.routine-config.outputs.print_timeout || '25m' }}" \
|
|
85
124
|
--dangerously-skip-permissions
|
|
@@ -65,6 +65,45 @@ jobs:
|
|
|
65
65
|
printf '%s' "$ANTIGRAVITY_OAUTH_TOKEN" > "$HOME/.gemini/antigravity-cli/antigravity-oauth-token"
|
|
66
66
|
chmod 600 "$HOME/.gemini/antigravity-cli/antigravity-oauth-token"
|
|
67
67
|
|
|
68
|
+
- name: Resolve routine configuration
|
|
69
|
+
id: routine-config
|
|
70
|
+
run: |
|
|
71
|
+
ROUTINE="issues-housekeeping"
|
|
72
|
+
DEFAULT_MODEL="gemini-3.7-flash"
|
|
73
|
+
DEFAULT_TIMEOUT="40"
|
|
74
|
+
|
|
75
|
+
MODEL="$DEFAULT_MODEL"
|
|
76
|
+
TIMEOUT="$DEFAULT_TIMEOUT"
|
|
77
|
+
|
|
78
|
+
if [ -f "agents-manifest.json" ]; then
|
|
79
|
+
CONFIG_JSON=$(node -e '
|
|
80
|
+
try {
|
|
81
|
+
const fs = require("fs");
|
|
82
|
+
const m = JSON.parse(fs.readFileSync("agents-manifest.json", "utf8"));
|
|
83
|
+
const r = process.argv[1];
|
|
84
|
+
const model = m.models?.[r] || m.models?.default;
|
|
85
|
+
const timeout = m.budgets?.timeoutMinutes?.[r] || m.budgets?.timeoutMinutes?.default;
|
|
86
|
+
console.log(JSON.stringify({ model: model || "", timeout: timeout ? String(timeout) : "" }));
|
|
87
|
+
} catch (e) {
|
|
88
|
+
console.log(JSON.stringify({ model: "", timeout: "" }));
|
|
89
|
+
}
|
|
90
|
+
' "$ROUTINE")
|
|
91
|
+
|
|
92
|
+
RESOLVED_MODEL=$(node -e "try { console.log(JSON.parse(process.argv[1]).model || ''); } catch {}" "$CONFIG_JSON")
|
|
93
|
+
RESOLVED_TIMEOUT=$(node -e "try { console.log(JSON.parse(process.argv[1]).timeout || ''); } catch {}" "$CONFIG_JSON")
|
|
94
|
+
|
|
95
|
+
if [ -n "$RESOLVED_MODEL" ]; then
|
|
96
|
+
MODEL="$RESOLVED_MODEL"
|
|
97
|
+
fi
|
|
98
|
+
if [ -n "$RESOLVED_TIMEOUT" ]; then
|
|
99
|
+
TIMEOUT="$RESOLVED_TIMEOUT"
|
|
100
|
+
fi
|
|
101
|
+
fi
|
|
102
|
+
|
|
103
|
+
echo "model=$MODEL" >> "$GITHUB_OUTPUT"
|
|
104
|
+
echo "print_timeout=${TIMEOUT}m" >> "$GITHUB_OUTPUT"
|
|
105
|
+
echo "Resolved routine configuration: model=$MODEL, timeout=${TIMEOUT}m"
|
|
106
|
+
|
|
68
107
|
- name: Run Issues Housekeeping
|
|
69
108
|
timeout-minutes: 40
|
|
70
109
|
env:
|
|
@@ -79,7 +118,7 @@ jobs:
|
|
|
79
118
|
PROMPT="You are the Issues Housekeeping routine for this repository. Read and follow .github/prompts/issues-housekeeping.md exactly. Sweep open issues for staleness, duplicates, priority accuracy, and orphaned claims."
|
|
80
119
|
|
|
81
120
|
agy -p "$PROMPT" \
|
|
82
|
-
--model gemini-3.7-flash
|
|
121
|
+
--model "${{ steps.routine-config.outputs.model || 'gemini-3.7-flash' }}" \
|
|
83
122
|
--output-format text \
|
|
84
|
-
--print-timeout 40m \
|
|
123
|
+
--print-timeout "${{ steps.routine-config.outputs.print_timeout || '40m' }}" \
|
|
85
124
|
--dangerously-skip-permissions
|
|
@@ -65,6 +65,45 @@ jobs:
|
|
|
65
65
|
printf '%s' "$ANTIGRAVITY_OAUTH_TOKEN" > "$HOME/.gemini/antigravity-cli/antigravity-oauth-token"
|
|
66
66
|
chmod 600 "$HOME/.gemini/antigravity-cli/antigravity-oauth-token"
|
|
67
67
|
|
|
68
|
+
- name: Resolve routine configuration
|
|
69
|
+
id: routine-config
|
|
70
|
+
run: |
|
|
71
|
+
ROUTINE="optimizer"
|
|
72
|
+
DEFAULT_MODEL="gemini-3.7-flash-high"
|
|
73
|
+
DEFAULT_TIMEOUT="35"
|
|
74
|
+
|
|
75
|
+
MODEL="$DEFAULT_MODEL"
|
|
76
|
+
TIMEOUT="$DEFAULT_TIMEOUT"
|
|
77
|
+
|
|
78
|
+
if [ -f "agents-manifest.json" ]; then
|
|
79
|
+
CONFIG_JSON=$(node -e '
|
|
80
|
+
try {
|
|
81
|
+
const fs = require("fs");
|
|
82
|
+
const m = JSON.parse(fs.readFileSync("agents-manifest.json", "utf8"));
|
|
83
|
+
const r = process.argv[1];
|
|
84
|
+
const model = m.models?.[r] || m.models?.default;
|
|
85
|
+
const timeout = m.budgets?.timeoutMinutes?.[r] || m.budgets?.timeoutMinutes?.default;
|
|
86
|
+
console.log(JSON.stringify({ model: model || "", timeout: timeout ? String(timeout) : "" }));
|
|
87
|
+
} catch (e) {
|
|
88
|
+
console.log(JSON.stringify({ model: "", timeout: "" }));
|
|
89
|
+
}
|
|
90
|
+
' "$ROUTINE")
|
|
91
|
+
|
|
92
|
+
RESOLVED_MODEL=$(node -e "try { console.log(JSON.parse(process.argv[1]).model || ''); } catch {}" "$CONFIG_JSON")
|
|
93
|
+
RESOLVED_TIMEOUT=$(node -e "try { console.log(JSON.parse(process.argv[1]).timeout || ''); } catch {}" "$CONFIG_JSON")
|
|
94
|
+
|
|
95
|
+
if [ -n "$RESOLVED_MODEL" ]; then
|
|
96
|
+
MODEL="$RESOLVED_MODEL"
|
|
97
|
+
fi
|
|
98
|
+
if [ -n "$RESOLVED_TIMEOUT" ]; then
|
|
99
|
+
TIMEOUT="$RESOLVED_TIMEOUT"
|
|
100
|
+
fi
|
|
101
|
+
fi
|
|
102
|
+
|
|
103
|
+
echo "model=$MODEL" >> "$GITHUB_OUTPUT"
|
|
104
|
+
echo "print_timeout=${TIMEOUT}m" >> "$GITHUB_OUTPUT"
|
|
105
|
+
echo "Resolved routine configuration: model=$MODEL, timeout=${TIMEOUT}m"
|
|
106
|
+
|
|
68
107
|
- name: Run Prompt Optimizer
|
|
69
108
|
timeout-minutes: 35
|
|
70
109
|
env:
|
|
@@ -79,9 +118,9 @@ jobs:
|
|
|
79
118
|
PROMPT="You are the Prompt Optimizer routine for this repository. Read and follow .github/prompts/optimizer.md exactly. Diagnose failures, token anomalies, and review loops, and propose preventative prompt and test fixes."
|
|
80
119
|
|
|
81
120
|
agy -p "$PROMPT" \
|
|
82
|
-
--model gemini-3.7-flash-high \
|
|
121
|
+
--model "${{ steps.routine-config.outputs.model || 'gemini-3.7-flash-high' }}" \
|
|
83
122
|
--output-format text \
|
|
84
|
-
--print-timeout 35m \
|
|
123
|
+
--print-timeout "${{ steps.routine-config.outputs.print_timeout || '35m' }}" \
|
|
85
124
|
--dangerously-skip-permissions
|
|
86
125
|
|
|
87
126
|
- name: Emit run summary
|