jonah-fleet 1.8.0 β†’ 1.9.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (43) hide show
  1. package/CHANGELOG.md +52 -0
  2. package/README.md +3 -3
  3. package/dist/commands/labels.d.ts +1 -1
  4. package/dist/commands/labels.d.ts.map +1 -1
  5. package/dist/commands/status.d.ts.map +1 -1
  6. package/dist/commands/sync.d.ts.map +1 -1
  7. package/dist/commands/telemetry.d.ts.map +1 -1
  8. package/dist/index.js +666 -151
  9. package/dist/lib/daemon-keys.d.ts.map +1 -1
  10. package/dist/lib/daemon.d.ts.map +1 -1
  11. package/dist/lib/fleet-query.d.ts.map +1 -1
  12. package/dist/lib/labels.d.ts +20 -0
  13. package/dist/lib/labels.d.ts.map +1 -1
  14. package/dist/lib/manifest.d.ts.map +1 -1
  15. package/dist/lib/presets.d.ts +6 -1
  16. package/dist/lib/presets.d.ts.map +1 -1
  17. package/dist/lib/runner.d.ts +48 -0
  18. package/dist/lib/runner.d.ts.map +1 -1
  19. package/dist/lib/telemetry.d.ts.map +1 -1
  20. package/dist/lib/terminal-card.d.ts +15 -1
  21. package/dist/lib/terminal-card.d.ts.map +1 -1
  22. package/package.json +1 -1
  23. package/schema.json +10 -5
  24. package/templates/prompts/ORCHESTRATION.md +49 -11
  25. package/templates/prompts/_prompt-template.md +10 -13
  26. package/templates/prompts/analytics-review.md +4 -2
  27. package/templates/prompts/autowork.md +43 -12
  28. package/templates/prompts/dependency-update-security-check.md +4 -2
  29. package/templates/prompts/design-review.md +125 -0
  30. package/templates/prompts/issues-housekeeping.md +9 -7
  31. package/templates/prompts/optimizer.md +31 -12
  32. package/templates/prompts/peer-review.md +29 -4
  33. package/templates/prompts/product-planning.md +14 -10
  34. package/templates/workflows/autowork-cron.yml +99 -39
  35. package/templates/workflows/dependency-check-cron.yml +89 -3
  36. package/templates/workflows/design-review-cron.yml +221 -0
  37. package/templates/workflows/issues-housekeeping-cron.yml +89 -3
  38. package/templates/workflows/prompt-optimizer-cron.yml +96 -10
  39. package/templates/workflows/sync-fleet.yml +4 -0
  40. package/templates/workflows/trigger-autowork-manual.yml +116 -12
  41. package/templates/workflows/trigger-autowork-on-bug.yml +135 -5
  42. package/templates/workflows/trigger-autowork-on-merge.yml +135 -5
  43. package/templates/workflows/trigger-review-routine.yml +119 -18
@@ -86,7 +86,7 @@ a. **Read the target issue and check eligibility.** Eligible = open, unassigned
86
86
  - **Design System & Viewport Pre-flight** (for frontend/UI diffs): self-audit diffs against design tokens (no arbitrary class overrides), WCAG AA 4.5:1 contrast ratios on dark/light surfaces, single primary CTA hierarchy per screen, and mobile viewport crowding (avoid stacked nudges/banners above the fold at ~390px).
87
87
  - **Documentation accuracy**: update relevant docs (`ARCHITECTURE.md`, `CODEMAP.md`, `API.md`, `CHANGELOG.md` if maintained by repo).
88
88
  - **Build & type-check verification**: run the repository's test, type-check, and lint commands from `AGENTS.md` (e.g. `npm test`, `npm run type-check`, `npm run lint`, `pytest`, `cargo test`). Confirm zero errors and zero test failures.
89
- - **Clean-merge gate**: verify `git merge-tree origin/main HEAD` reports no conflicts.
89
+ - **Active Origin Sync & Clean-Merge Gate**: Run `git fetch origin main && git merge origin/main --no-edit` to absorb any newly merged pull requests and resolve any conflicts locally. Verify `git merge-tree origin/main HEAD` reports no conflicts before marking ready.
90
90
  - **Release claim on ready**: mark the PR ready (`gh pr ready <PR>`) and unassign yourself (`gh pr edit <PR> --remove-assignee <login>`) so Peer Review can evaluate without holding stale agent reservation locks.
91
91
  Only mark the PR ready after passing every check above.
92
92
  3b. **Ping-pong cap**: If this same PR has bounced between draft and ready 3 or more times over the same substantive finding, stop re-marking it ready. Post a comment summarizing the disagreement for human resolution and leave the PR in draft.
@@ -139,10 +139,28 @@ a. **Read the target issue and check eligibility.** Eligible = open, unassigned
139
139
  - Read `🧭 Decomposition plan` comment (or create if first run).
140
140
  - Pick next slice(s), batching up to 3 same-recipe slices into one child issue + PR.
141
141
  - Create and claim child issue first, then update plan marker to `🚧 in progress β€” child #M`, release umbrella claim, and implement against child.
142
+ - **Milestone 1 (Intake & Strategy)**: If `$ROUTINE_ISSUE_NUMBER` is set, emit milestone card to the routine issue thread:
143
+ ```bash
144
+ gh issue comment "$ROUTINE_ISSUE_NUMBER" --body "### 🧭 Milestone: Intake & Strategy
145
+ - **Phase**: \`Phase 1 Β· Target Selection & Planning\`
146
+ - **Status**: ⏳ In Progress
147
+ - **Target / Context**: \`#<TARGET_ISSUE> (<Title>)\`
148
+ - **Key Decision / Finding**: <Summary of plan, root cause, or reproduction strategy>
149
+ - **Next**: Implementation & Test Verification" || true
150
+ ```
142
151
  13. **Implementation & PR creation:**
143
152
  - Branch from freshly fetched `origin/main` with descriptive name (e.g. `feat/...` or `fix/...`).
144
153
  - Drive implementation via `/tdd` (red-green-refactor).
145
154
  - Run repository tests and verification.
155
+ - **Milestone 2 (Verification & Tests)**: Once implementation passes tests and type checks, emit milestone card if `$ROUTINE_ISSUE_NUMBER` is set:
156
+ ```bash
157
+ gh issue comment "$ROUTINE_ISSUE_NUMBER" --body "### πŸ§ͺ Milestone: Verification & Tests
158
+ - **Phase**: \`Phase 2 Β· Implementation & Verification\`
159
+ - **Status**: βœ… Tests Passing
160
+ - **Target / Context**: \`#<TARGET_ISSUE>\`
161
+ - **Key Decision / Finding**: Tests verified passing with zero failures and clean type-check.
162
+ - **Next**: PR Creation & Autonomous Handoff" || true
163
+ ```
146
164
  - Check for competing open PRs immediately before creating PR. If collision, bail cleanly.
147
165
  - Open draft PR via `gh pr create --draft --head <branch> --base main --title "<title>" --body "<body referencing Closes #N>"`.
148
166
  - **PR Priority Label Mirroring**: If the issue carried a priority label (`priority/P0`, `priority/P1`, `priority/P2`, `priority/P3`), add the identical priority label to the PR (`gh pr edit <PR> --add-label "<label>"` or via `--label` in create) so downstream review workflows can filter triggers immediately.
@@ -150,6 +168,15 @@ a. **Read the target issue and check eligibility.** Eligible = open, unassigned
150
168
  - **Issue Cross-Reference Comment Guardrail**: Post an explicit comment on the tracking issue referencing the newly created PR (`gh issue comment <ISSUE_NUMBER> --body "Work in progress in PR #<PR_NUMBER>."`). This guarantees an unambiguous, permanent link on the issue timeline even when GitHub's native UI link is suppressed for bot draft PRs.
151
169
  - **Warm Context Assignment**: Assign yourself to the newly opened PR (`gh pr edit <PR> --add-assignee <login>`) to hold the reservation across Step 15's in-session review wait so parallel Scan routines recognise the PR as actively held by a live session.
152
170
  - Mark PR ready (`gh pr ready <PR>`).
171
+ - **Milestone 3 (Autonomous Handoff)**: Once PR is marked ready, emit milestone card if `$ROUTINE_ISSUE_NUMBER` is set:
172
+ ```bash
173
+ gh issue comment "$ROUTINE_ISSUE_NUMBER" --body "### πŸš€ Milestone: Autonomous Handoff
174
+ - **Phase**: \`Phase 3 Β· Autonomous Handoff\`
175
+ - **Status**: βœ… PR Marked Ready
176
+ - **Target / Context**: \`PR #<PR_NUMBER> (closes #<TARGET_ISSUE>)\`
177
+ - **Key Decision / Finding**: Pushed branch and marked PR ready for review.
178
+ - **Next**: In-session review wait or run completion" || true
179
+ ```
153
180
  14. If run aborts before opening PR, release claim (unassign).
154
181
  15. **In-Session Peer Review Wait & Immediate Convergence (Warm Context):**
155
182
  - Poll PR status for up to 10–12 minutes.
@@ -159,23 +186,27 @@ a. **Read the target issue and check eligibility.** Eligible = open, unassigned
159
186
 
160
187
  ## Logging
161
188
 
162
- After completing (SUCCESS or FAILURE), write a log file to `.github/prompts/logs/autowork/{timestamp}.md` following the schema in `.github/prompts/logs/_template.md`. Include:
189
+ After completing (SUCCESS or FAILURE), record run execution details to `.jonah-fleet/run-report.md`. Include:
163
190
 
164
191
  - The prompt SHA (run `git rev-parse --short HEAD:.github/prompts/autowork.md`)
165
192
  - Every Definition of Done criterion with YES/NO and evidence
166
193
  - Full execution trace with tool calls
167
194
  - If FAILURE: root cause, category, and suggested fix
168
195
 
169
- **Log Delivery Protocol & Invariants**:
196
+ **Issue Logging Protocol & Invariants**:
170
197
 
171
- - **Negative Rule**: NEVER commit or push run logs to a feature branch or open PR branch. Doing so emits a `pull_request: synchronize` event under bot credentials, triggering GitHub Actions workflow approval gates (`action_required`) that stall CI.
172
- - **Mandatory `[skip ci]`**: Always append `[skip ci]` to any log commit message.
173
- - **Direct Push to `main`**: Commit the log file directly to `main` and push β€” explicitly permitted for files under `.github/prompts/logs/**`:
198
+ - **Negative Rule**: NEVER commit or push run logs to any git branch (`main` or feature branches). Run logs are recorded exclusively as GitHub Issues and local `.jonah-fleet/runs/*.json` artifacts.
199
+ - **Milestone 4 (Run Completed)**: If `$ROUTINE_ISSUE_NUMBER` is set in the environment, append the concluding compact milestone card to close out the comment stream:
174
200
  ```bash
175
- git checkout main
176
- git pull origin main
177
- git add .github/prompts/logs/autowork/{timestamp}.md
178
- git commit -m "docs(log): record autowork run {timestamp} [skip ci]"
179
- git push origin main
201
+ gh issue comment "$ROUTINE_ISSUE_NUMBER" --body "### 🏁 Milestone: Run Completed
202
+ - **Phase**: \`Phase 4 Β· Reconciliation\`
203
+ - **Status**: βœ… SUCCESS
204
+ - **Target / Context**: \`#<TARGET_ISSUE>\` Β· \`PR #<PR_NUMBER>\`
205
+ - **Key Decision / Finding**: All DoD criteria satisfied and verified.
206
+ - **Next**: Routine finished; issue closed by harness" || true
180
207
  ```
181
- Follow the Log delivery fallback in `ORCHESTRATION.md` if direct push fails.
208
+ - **Harness Reconciliation**: When running under GitHub Actions or `jonah-fleet daemon`, the harness initializes the run issue (`status:running`), replaces the top-level issue body with `.jonah-fleet/run-report.md`, and reconciles the final status upon completion:
209
+ - On **SUCCESS**: Labels `status:success`, removes `status:running`, and closes the issue.
210
+ - On **FAILURE**: If interrupted or failed, the harness posts the Interruption Card, edits the body, and marks `status:failure,needs-attention`.
211
+ - On clean local runs outside a harness, write `.jonah-fleet/runs/{timestamp}.json`.
212
+ - Follow the Routine Issue Logging & Telemetry Protocol in `ORCHESTRATION.md`.
@@ -36,9 +36,11 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
36
36
 
37
37
  ## Logging
38
38
 
39
- After completing (SUCCESS or FAILURE), write a log file to `.github/prompts/logs/dependency-update-security-check/{timestamp}.md` following the schema in `.github/prompts/logs/_template.md`. Include:
39
+ After completing (SUCCESS or FAILURE), record run execution details to `.jonah-fleet/run-report.md`. Include:
40
40
  - Prompt SHA
41
41
  - Tally of audited dependencies and security findings
42
42
  - List of created or updated issues
43
43
 
44
- **Important**: Commit the log file directly to `main` and push. Follow the Log delivery fallback in `ORCHESTRATION.md` if direct push fails.
44
+ **Issue Logging Protocol**:
45
+ - Record run execution details to `.jonah-fleet/run-report.md` (or update `$ROUTINE_ISSUE_NUMBER`).
46
+ - Follow the Routine Issue Logging & Telemetry Protocol in `ORCHESTRATION.md`. Never commit run logs to git branches.
@@ -0,0 +1,125 @@
1
+ # Design Review
2
+
3
+ ## Objective
4
+
5
+ Audit player-facing surfaces for design system token deviations, visual clutter, candidates for feature pruning, and high-impact UX improvements. Provide a continuous, budget-capped quality loop across the app by combining dynamic route discovery, static code analysis, and headless visual critiques. Directly file concrete, self-contained design fixes for Autowork while staging structural UX opportunities and pruning proposals into a dedicated Design Review issue for Product Planning to ingest.
6
+
7
+ ## Definition of Done
8
+
9
+ The run is SUCCESS only if ALL of these are true:
10
+
11
+ - [ ] Evaluated target surfaces: if `TARGET_SURFACE` was provided via manual dispatch, targeted that specific surface; otherwise performed dynamic route discovery across `src/app/[locale]/**/page.tsx` (excluding `admin/**`), prioritized any `[NEW SURFACE]` at Priority 0, and proceeded through the rotation sequence within the 30-iteration budget
12
+ - [ ] Performed Tier 1 Static Code Audit on target surfaces: checked against `DESIGN_SYSTEM.md` foundations, flagging arbitrary Tailwind values (`text-[...]`, `w-[...]`, arbitrary colors), non-standard button styling, multiple `.btn-primary` actions per view, and contrast token regressions
13
+ - [ ] Performed Tier 2 Headless Browser & Visual Critique: launched the dev environment and captured viewports at mobile (390px) and desktop (1280px) via `run` and `design-critique` skills, auditing whitespace consistency, visual hierarchy, touch targets, and viewport density (enforcing the mobile fold budget: max 1 banner/nudge)
14
+ - [ ] Cross-referenced UI friction and telemetry signals: checked for recurring rage clicks (`$rageclick`) and identified low-interaction (<2% CTR) secondary elements or persistent nudges suitable for pruning
15
+ - [ ] Executed Hybrid Handoff:
16
+ - If self-contained token, class, or contrast violations were found, filed or updated exactly one grouped punchlist issue titled `🎨 Design Polish & Token Cleanup β€” {YYYY-MM-DD}` labeled `design/polish` for Autowork
17
+ - Created or updated exactly one tracking issue titled `🎨 Design Review β€” {YYYY-MM-DD}` documenting: Surfaces Reviewed (with any `[NEW SURFACE]` highlighted), Next Surfaces in Rotation, Visual Critique & Clutter Analysis, Feature Pruning & Deprecation Recommendations, and Experience Enhancement Opportunities
18
+ - [ ] Emitted structured design directives in the run report (`DESIGN_DIRECTIVE: [PRUNE | REDESIGN | POLISH]`) and recorded the next queued surfaces for the subsequent run
19
+ - [ ] Recorded the final execution report to `.jonah-fleet/run-report.md` following `ORCHESTRATION.md`
20
+
21
+ If any criterion cannot be met, stop immediately and log FAILURE with the reason.
22
+
23
+ ## Constraints
24
+
25
+ - **Max iterations**: 30 β€” after 30 tool call rounds without completing Definition of Done, STOP. Log FAILURE with category `token_limit`.
26
+ - **Max scope**: auditing player-facing UI surfaces only (`src/app/[locale]/**/page.tsx`, excluding `admin/**`). Do not implement feature code or open feature PRs directly.
27
+ - **Budget-capped rotation**: review as many prioritized surfaces as possible within the iteration budget; do not attempt an exhaustive sweep if approaching the tool call ceiling. Stamp remaining surfaces as `Next Surfaces in Rotation`.
28
+ - **Language Requirement**: All GitHub issue titles, descriptions, task checklists, and comments MUST be written in **English**.
29
+ - **Session link footer**: sign every GitHub post with the Antigravity run footer (`_Generated by [Antigravity](${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}/actions/runs/${GITHUB_RUN_ID})_`).
30
+
31
+ ## Negative examples (DO NOT do these)
32
+
33
+ - Do not open multiple separate issues for individual token mismatches β€” group all simple token and contrast cleanup items into a single grouped punchlist issue for Autowork.
34
+ - Do not file unapproved structural redesigns or feature removals as direct Autowork issues β€” stage them in the `🎨 Design Review β€” {date}` tracking issue for Product Planning to review.
35
+ - Do not review the admin back-office (`src/app/[locale]/admin/**`) β€” it is deliberately utilitarian and exempt from the player-facing design system polish bar.
36
+ - Do not blindly repeat the same surface review order every run β€” dynamically discover routes and check previous review issues to resume rotation and fast-track unreviewed new surfaces.
37
+
38
+ ## Instructions
39
+
40
+ ### Step 1: Dynamic Route Discovery & Target Prioritization
41
+
42
+ 1. **Check for Manual Target Override**:
43
+ - If the environment variable `TARGET_SURFACE` is provided and non-empty, audit only that specified surface or route and skip rotation determination.
44
+ 2. **Dynamic Route Discovery**:
45
+ - Scan the filesystem for all active player-facing page routes: `src/app/[locale]/**/page.tsx` (ignoring any `admin/**` directories).
46
+ - Normalize the discovered paths into canonical route identifiers (e.g. `Home` (`/`), `Predictions` (`/pronostics`), `Episode Predictions` (`/pronostics/episode/[id]`), `Groups` (`/groupes`), `Group Detail` (`/groupes/[id]`), `Presenter` (`/groupes/[id]/presentateur`), `Leaderboard` (`/classements`), `Profile` (`/profil`), `Account` (`/compte`), `Auth` (`/auth/*`), `Season Overview` (`/saisons/[slug]`), `Share` (`/partage`)).
47
+ 3. **Reconcile with Previous Rotation State**:
48
+ - Locate the most recent `🎨 Design Review` issue or run issue labeled `routine:design-review`.
49
+ - Read the `Surfaces Reviewed` and `Next Surfaces in Rotation` fields.
50
+ - Any discovered route on disk that has never appeared in prior review logs is flagged as **`[NEW SURFACE]`** and placed at **Priority 0** (Fast-Track).
51
+ - Remaining surfaces are sequenced based on `Next Surfaces in Rotation` (resuming rotation where the previous run stopped).
52
+ - Any route present in the previous rotation list that no longer exists on disk is automatically pruned.
53
+ 4. **Set Run Queue**:
54
+ - Build the active review queue: `[NEW SURFACE]` items first, followed by the queued rotation surfaces.
55
+ - Keep a budget awareness counter: as iterations proceed, stop adding new surfaces when remaining tool calls are needed for issues and logging.
56
+
57
+ ### Step 2: Static Code & Token Audit
58
+
59
+ For each surface in the active queue:
60
+ 1. Locate the page component and its imported presentation components under `src/components/**`.
61
+ 2. Inspect the styling using `.agents/skills/design-system/SKILL.md`:
62
+ - **Token Purity**: flag hardcoded hex codes (`#...`), rgb values, or arbitrary Tailwind brackets (e.g. `text-[13px]`, `w-[310px]`).
63
+ - **Surface Hierarchy**: verify background hierarchy adheres to `bg-gray-950` (base) β†’ `bg-gray-900` / `.card` (card) β†’ `bg-gray-800` / `.input-field` (input).
64
+ - **Contrast Tokens**: ensure readable secondary text uses `text-gray-400` minimum (never `text-gray-500`, which fails WCAG AA on dark backgrounds).
65
+ - **CTA Discipline**: verify each view/tab contains at most **one** primary action button (`.btn-primary`). Flag competing primary buttons.
66
+ - **Button & Pill Standards**: verify buttons use established classes (`.btn-primary`, `.btn-secondary`, `.btn-ghost`) and status pills use standard semantic colors.
67
+
68
+ ### Step 3: Headless Browser & Visual Critique
69
+
70
+ For the surfaces being audited:
71
+ 1. Start the local Next.js dev server with mock credentials following `.agents/skills/run/SKILL.md`:
72
+ ```bash
73
+ NEXT_PUBLIC_SUPABASE_URL=http://localhost:54321 NEXT_PUBLIC_SUPABASE_ANON_KEY=placeholder npm run dev
74
+ ```
75
+ 2. Using Playwright, capture screenshots at two standard viewports:
76
+ - **Mobile**: 390px width
77
+ - **Desktop**: 1280px width
78
+ 3. Apply `.agents/skills/design-critique/SKILL.md` to evaluate visual evidence:
79
+ - **Viewport Density & Nudge Budget**: On mobile (390px), do sticky headers, promotional banners, and teasers crowd the viewport above the fold? Enforce: max 1 banner/nudge visible concurrently.
80
+ - **Visual Hierarchy & Clutter**: Are cards nested redundantly? Is there visual noise from excessive borders, conflicting accents, or dense text blocks?
81
+ - **Touch Targets & Spacing**: Are tap targets at least 44x44px? Is whitespace rhythmic and intentional?
82
+ 4. Shut down the dev server cleanly after capture.
83
+
84
+ ### Step 4: Telemetry & Friction Cross-Reference
85
+
86
+ 1. If telemetry endpoints or PostHog logs are available:
87
+ - Audit `$rageclick` events associated with URLs and selectors of the reviewed surfaces. Pinpoint confusing or non-interactive elements masquerading as buttons.
88
+ - Review click-through rates (CTR) on secondary promotional cards or banner elements.
89
+ 2. Formulate **Pruning Recommendations**:
90
+ - Any persistent card, teaser, or secondary control with near-zero interaction, high rage-click density, or severe visual clutter is recommended for deprecation (`RECOMMENDATION: DEPRECATE`) or simplification (`RECOMMENDATION: PRUNE`).
91
+
92
+ ### Step 5: Hybrid Handoff & Issue Generation
93
+
94
+ 1. **Direct Autowork Punchlist Issue**:
95
+ - If concrete, self-contained token violations, contrast bugs, or duplicate `.btn-primary` classes are found, compile them into a single grouped punchlist issue:
96
+ - Title: `🎨 Design Polish & Token Cleanup β€” {YYYY-MM-DD}`
97
+ - Labels: `design/polish`, `autowork-candidate`
98
+ - Content: Structured task list with target files, line references, current violation, and suggested replacement token/class.
99
+ 2. **Overarching Design Review Tracking Issue**:
100
+ - Create or update a single tracking issue:
101
+ - Title: `🎨 Design Review β€” {YYYY-MM-DD}`
102
+ - Labels: `area/design`, `review`
103
+ - Content:
104
+ - **Surfaces Reviewed**: list audited routes, noting any `[NEW SURFACE]`
105
+ - **Next Surfaces in Rotation**: list unreviewed routes queued for the subsequent run
106
+ - **Visual Critique & Clutter Analysis**: mobile fold density, hierarchy, and layout observations
107
+ - **Pruning & Deprecation Candidates**: specific features/elements recommended to be pruned or consolidated
108
+ - **Experience Enhancement Opportunities**: structural UX improvements staged for Product Planning
109
+ 3. **Structured Directives**:
110
+ - Output clear directives: `DESIGN_DIRECTIVE: [PRUNE | REDESIGN | POLISH]` to be consumed downstream by `product-planning.md`.
111
+
112
+ ## Logging
113
+
114
+ Follow the Routine Issue Logging Protocol in `ORCHESTRATION.md`:
115
+ 1. Write the final run report to `.jonah-fleet/run-report.md`.
116
+ 2. Include the Run Summary table with:
117
+ - Routine: `design-review`
118
+ - Prompt SHA (run `git rev-parse --short HEAD:.github/prompts/design-review.md` or template)
119
+ - Result: `SUCCESS` or `FAILURE`
120
+ - Every Definition of Done criterion with YES/NO and evidence
121
+ - Surfaces Reviewed and Next Surfaces in Rotation
122
+ - Created issues (`🎨 Design Polish & Token Cleanup`, `🎨 Design Review`)
123
+ - If FAILURE: root cause, category, and suggested fix
124
+ 3. The surrounding execution harness will reconcile the corresponding GitHub issue.
125
+
@@ -15,7 +15,7 @@ The run is SUCCESS only if ALL of these are true:
15
15
  - [ ] Label audit completed: every open issue has consistent type, size, and priority labels
16
16
  - [ ] Dependency check completed: issues with `## Dependencies` verified against blocker status
17
17
  - [ ] Orphaned-claim sweep completed: stale autowork claims (per `ORCHESTRATION.md`) released back to the unclaimed pool
18
- - [ ] Draft log-only PR sweep completed: open draft PRs whose changed files are entirely under `.github/prompts/logs/**` undrafted and squash-merged
18
+ - [ ] Stalled routine run audit completed: routine log issues in `status:running` older than 6 hours marked as `status:failure` (timed_out)
19
19
  - [ ] Summary posted listing all changes made
20
20
 
21
21
  If any criterion cannot be met, stop immediately and log FAILURE with the reason.
@@ -24,14 +24,14 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
24
24
 
25
25
  - **Max iterations**: 40 β€” after 40 tool call rounds without completing Definition of Done, STOP. Log FAILURE with category `token_limit`.
26
26
  - **Max scope**: housekeeping only. Do not implement code fixes or open feature PRs.
27
- - **No speculative work**: only modify issue metadata (labels, status, comments, releasing stale assignees) and land draft log-only PRs.
27
+ - **No speculative work**: only modify issue metadata (labels, status, comments, releasing stale assignees, reconciling stalled routine runs).
28
28
  - **Language Requirement**: All GitHub issue titles, descriptions, task checklists, and comments MUST be written in **English**.
29
29
 
30
30
  ## Instructions
31
31
 
32
32
  ### Phase 1: Quick Recovery & Clearing
33
33
 
34
- 1. **Draft log-only PR sweep**: Undraft and squash-merge open draft PRs whose diffs are entirely under `.github/prompts/logs/**`.
34
+ 1. **Stalled routine run sweep**: Check open issues with labels `routine-log` and `status:running`. If an issue has been in `status:running` for > 6 hours without updates, add label `status:failure`, remove `status:running`, add label `needs-attention`, and comment noting the runner timeout or crash.
35
35
  2. **Orphaned-claim sweep**: Sweep assigned issues. If an issue meets the 3 stale-claim conditions in `ORCHESTRATION.md` (autowork claim comment, no open PR, comment > 6 hours old), re-read immediately before writing, unassign the dead owner, and post a release comment.
36
36
 
37
37
  ### Phase 2: Backlog Hygiene
@@ -44,13 +44,15 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
44
44
 
45
45
  ### Phase 3: Summary
46
46
 
47
- 8. Post a summary comment or log recording all actions taken (priority shifts, closed duplicates, released claims, landed log PRs).
47
+ 8. Post a summary comment or log recording all actions taken (priority shifts, closed duplicates, released claims, reconciled stalled runs).
48
48
 
49
49
  ## Logging
50
50
 
51
- After completing (SUCCESS or FAILURE), write a log file to `.github/prompts/logs/issues-housekeeping/{timestamp}.md` following the schema in `.github/prompts/logs/_template.md`. Include:
51
+ After completing (SUCCESS or FAILURE), record run execution details to `.jonah-fleet/run-report.md`. Include:
52
52
  - Prompt SHA
53
- - Tally of issues audited, claims released, PRs merged
53
+ - Tally of issues audited, claims released, stalled runs reconciled
54
54
  - List of closed or modified issues
55
55
 
56
- **Important**: Commit the log file directly to `main` and push. Follow the Log delivery fallback in `ORCHESTRATION.md` if direct push fails.
56
+ **Issue Logging Protocol**:
57
+ - Record run execution details to `.jonah-fleet/run-report.md` (or update `$ROUTINE_ISSUE_NUMBER`).
58
+ - Follow the Routine Issue Logging & Telemetry Protocol in `ORCHESTRATION.md`. Never commit run logs to git branches.
@@ -10,10 +10,10 @@ Additionally, this routine acts as the **Upstream Evolution Bridge**: when a pro
10
10
 
11
11
  The run is SUCCESS if ALL of these are true:
12
12
 
13
- - [ ] All log files from the incremental window in `.github/prompts/logs/` have been scanned
14
- - [ ] Every FAILURE log has been categorized and analyzed
15
- - [ ] Inefficiency and review loops per PR have been computed across SUCCESS logs
16
- - [ ] Per-agent token and cost consumption metrics have been aggregated across in-window logs and evaluated against the 70% weekly budget ceiling in `ORCHESTRATION.md`
13
+ - [ ] All routine run issues from the incremental window (labeled `routine-log`) have been scanned
14
+ - [ ] Every FAILURE run has been categorized and analyzed
15
+ - [ ] Inefficiency and review loops per PR have been computed across SUCCESS runs
16
+ - [ ] Per-agent token and cost consumption metrics have been aggregated across in-window routine runs and evaluated against the 70% weekly budget ceiling in `ORCHESTRATION.md`
17
17
  - [ ] Closed bug issues and merged bug-fix PRs in the window have been analyzed for systemic root causes
18
18
  - [ ] For each fixable pattern:
19
19
  - If project-specific: opened a local PR with a prompt, template, or test fix and marked ready for review
@@ -27,23 +27,32 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
27
27
 
28
28
  - **Max iterations**: 30 β€” stop after 30 tool call rounds.
29
29
  - **Max scope**: one PR per identified problem. Do not bundle unrelated fixes.
30
- - **No speculative work**: only fix patterns evidenced by logs, closed bug issues, or reviewer findings.
30
+ - **No speculative work**: only fix patterns evidenced by routine issues, closed bug issues, or reviewer findings.
31
31
  - **Language Requirement**: All GitHub issue titles, descriptions, task checklists, and comments MUST be written in **English**.
32
32
 
33
33
  ## Instructions
34
34
 
35
35
  ### 0. Establish the incremental scan boundary
36
36
 
37
- 1. Check `.github/prompts/logs/optimizer/` for the most recent optimizer log.
37
+ 1. Check the most recent routine issue with labels `routine-log,routine:optimizer` via `gh issue list --label routine-log --label routine:optimizer --state all --limit 1`.
38
38
  2. Extract the timestamp as the scan boundary (or last 7 days if first run).
39
+ 3. **Milestone 1 (Intake & Fleet Telemetry Scan)**: If `$ROUTINE_ISSUE_NUMBER` is set, emit milestone card to the routine issue thread:
40
+ ```bash
41
+ gh issue comment "$ROUTINE_ISSUE_NUMBER" --body "### 🧭 Milestone: Intake & Fleet Telemetry Scan
42
+ - **Phase**: \`Phase 1 Β· Telemetry Intake & Boundary Scan\`
43
+ - **Status**: ⏳ In Progress
44
+ - **Target / Context**: \`Fleet Telemetry Window\`
45
+ - **Key Decision / Finding**: Established scan boundary; analyzing in-window routine issues and token pacing.
46
+ - **Next**: Anomaly Diagnosis & Proposals" || true
47
+ ```
39
48
 
40
49
  ### 1. Collect signals & analyze logs
41
50
 
42
- 1. **Scan in-window log files**: Read all log files in `.github/prompts/logs/*/` within the incremental scan window.
43
- 2. **Extract failure categories**: Categorize runs logging `FAILURE` (`prompt_unclear`, `data_issue`, `token_limit`, `infeasible_task`).
51
+ 1. **Query in-window routine issues**: List routine run issues within the window using `gh issue list --label routine-log --state all --limit 100 --json number,title,body,state,labels,createdAt`.
52
+ 2. **Extract failure categories**: Categorize runs labeled `status:failure` or logging `FAILURE` (`prompt_unclear`, `data_issue`, `token_limit`, `infeasible_task`).
44
53
  3. **Compute efficiency metrics**: Identify PRs experiencing $\ge 3$ review bounce rounds and runs with high iteration usage relative to limits.
45
54
  4. **Aggregate per-agent token & cost consumption**:
46
- - Parse the metadata table from each in-window log: `Routine`, `Input tokens`, `Output tokens`, `Estimated cost`, `Iterations used` (e.g. `26 / 65`), and `Result` (`SUCCESS` or `FAILURE`).
55
+ - Parse the metadata table from each in-window routine issue body: `Routine`, `Input tokens`, `Output tokens`, `Estimated cost`, `Iterations used` (e.g. `26 / 65`), and `Result` (`SUCCESS` or `FAILURE`).
47
56
  - Group logs by `Routine` (`autowork`, `peer-review`, `issues-housekeeping`, `dependency-update-security-check`, `optimizer`, `product-planning`).
48
57
  - For each routine, compute:
49
58
  - **Run count**: total completed runs.
@@ -88,9 +97,9 @@ Translate findings into concrete preventative improvements and remediation trigg
88
97
 
89
98
  ## Logging
90
99
 
91
- After completing (SUCCESS or FAILURE), write a log file to `.github/prompts/logs/optimizer/{timestamp}.md` following the schema in `.github/prompts/logs/_template.md`. Include:
100
+ After completing (SUCCESS or FAILURE), record run execution details to `.jonah-fleet/run-report.md`. Include:
92
101
  - Prompt SHA
93
- - Analyzed logs count and identified patterns
102
+ - Analyzed routine issues count and identified patterns
94
103
  - **Token & Cost Consumption by Agent** scorecard table:
95
104
 
96
105
  ```markdown
@@ -109,5 +118,15 @@ After completing (SUCCESS or FAILURE), write a log file to `.github/prompts/logs
109
118
  - Weekly token budget pacing evaluation (pacing vs 70% ceiling in `ORCHESTRATION.md`)
110
119
  - PRs opened (local or upstream)
111
120
 
112
- **Important**: Commit the log file directly to `main` and push. Follow the Log delivery fallback in `ORCHESTRATION.md` if direct push fails.
121
+ **Issue Logging Protocol**:
122
+ - **Milestone 4 (Run Completed)**: If `$ROUTINE_ISSUE_NUMBER` is set, emit compact milestone card to close out the comment stream:
123
+ ```bash
124
+ gh issue comment "$ROUTINE_ISSUE_NUMBER" --body "### 🏁 Milestone: Run Completed
125
+ - **Phase**: \`Phase 4 Β· Reconciliation\`
126
+ - **Status**: βœ… SUCCESS
127
+ - **Target / Context**: \`Fleet Optimization Sweep\`
128
+ - **Key Decision / Finding**: Fleet telemetry sweep completed and scored against token ceiling.
129
+ - **Next**: Routine finished; issue closed by harness" || true
130
+ ```
131
+ - Follow the Routine Issue Logging & Telemetry Protocol in `ORCHESTRATION.md`. Never commit run logs to git branches.
113
132
 
@@ -70,7 +70,7 @@ Check if `$PR_NUMBER` is set:
70
70
 
71
71
  ### Steps 1–2: Select a PR (Scan mode only)
72
72
 
73
- 1. List all open PRs, excluding drafts, pure log PRs (`.github/prompts/logs/**`), and automated release PRs (`release-please--*` / `chore(main): release*`).
73
+ 1. List all open PRs, excluding drafts and automated release PRs (`release-please--*` / `chore(main): release*`).
74
74
  2. **Dual Execution Priority Routing**:
75
75
  - If running in **Cloud Actions** (`$GITHUB_ACTIONS` / `$CI`): Scan mode prioritizes PRs labeled `priority/P0` or `priority/P1` (or untagged PRs). Lower-priority PRs (`priority/P2`, `priority/P3`) are eligible for cloud review ONLY if they have remained ready and unreviewed for more than 48 hours (`created_at` older than 48h, acting as a cloud catchup sweep).
76
76
  - If running in **Local Agent** (`$LOCAL_AGENT`): Scan mode reviews any ready PR across all priorities (P0 β†’ P1 β†’ P2 β†’ P3) with no age gating.
@@ -80,8 +80,21 @@ Check if `$PR_NUMBER` is set:
80
80
 
81
81
  ### Step 3: Round tracking & Starting Review marker
82
82
 
83
- - Post a "Starting review (round N)" comment on the target PR to claim the review window.
83
+ - Post a "Starting review (round N)" comment on the target PR to claim the review window:
84
+ - If `$ROUTINE_ISSUE_NUMBER` is set in the environment or routine prompt, include a link to the tracking log issue:
85
+ `Starting review (round N) Β· [Run Log #$ROUTINE_ISSUE_NUMBER](${GITHUB_SERVER_URL:-https://github.com}/${GITHUB_REPOSITORY}/issues/${ROUTINE_ISSUE_NUMBER})`
86
+ (or `Starting review (round N) Β· Run log: #$ROUTINE_ISSUE_NUMBER`).
87
+ - If `$ROUTINE_ISSUE_NUMBER` is not set, post `Starting review (round N)`.
84
88
  - Check round count `N`. If `N >= 5` and blocking findings persist, prepare to escalate.
89
+ - **Milestone 1 (Intake & Review Scope)**: If `$ROUTINE_ISSUE_NUMBER` is set, emit milestone card to the routine issue thread:
90
+ ```bash
91
+ gh issue comment "$ROUTINE_ISSUE_NUMBER" --body "### 🧭 Milestone: Intake & Review Scope
92
+ - **Phase**: \`Phase 1 Β· Review Target & Scope\`
93
+ - **Status**: ⏳ In Progress
94
+ - **Target / Context**: \`PR #<PR_NUMBER> (Round N)\`
95
+ - **Key Decision / Finding**: Claimed review window; starting multi-angle evaluation passes.
96
+ - **Next**: Multi-Angle Code Review Pass" || true
97
+ ```
85
98
 
86
99
  ### Step 4: Multi-Angle Code Review Pass
87
100
 
@@ -124,6 +137,7 @@ If the PR is clean and approved for merge, but lacks a `Closes #N` tracking link
124
137
  - If `N < 5`: Post inline comments, submit review as `COMMENT`, and convert PR to draft (`gh pr ready <N> --undo`).
125
138
  - If `N >= 5`: Convert PR to draft, post summary comment escalating to repo maintainer, and apply `needs-human` label.
126
139
  - **If Clean (or only Non-blocking findings)**:
140
+ - If the PR branch has minor mechanical merge conflicts against `origin/main` (e.g. adjacent `CHANGELOG.md` entries from concurrent merges) while code and tests are sound, resolve the conflict mechanically via `/resolving-merge-conflicts` before squash-merging rather than bouncing the PR back to draft.
127
141
  - If PR lacks `Closes #N`, execute Autonomous Issue Synthesis (Step 5.5).
128
142
  - Extract the tracking issue number `$ISSUE_NUMBER` from the PR description or title (e.g. `Closes #<N>`, `Fixes #<N>`, `Resolves #<N>`).
129
143
  - Squash-merge the PR: `gh pr merge <N> --squash --delete-branch`.
@@ -132,15 +146,26 @@ If the PR is clean and approved for merge, but lacks a `Closes #N` tracking link
132
146
  gh issue close "$ISSUE_NUMBER" --comment "Closed via PR #<N> (merged into main)."
133
147
  ```
134
148
  - Submit held review comments.
149
+ - In any review summary, decision comment, or escalation comment posted to the PR, include the run log reference if `$ROUTINE_ISSUE_NUMBER` is set (e.g. `- **Run Log**: [#$ROUTINE_ISSUE_NUMBER](${GITHUB_SERVER_URL:-https://github.com}/${GITHUB_REPOSITORY}/issues/${ROUTINE_ISSUE_NUMBER})`).
135
150
  - File follow-up issues for material non-blocking findings.
136
151
  - If mechanical doc fixes are needed, commit directly to `main`.
137
152
 
138
153
  ## Logging
139
154
 
140
- After completing (SUCCESS or FAILURE), write a log file to `.github/prompts/logs/peer-review/{timestamp}.md` following the schema in `.github/prompts/logs/_template.md`. Include:
155
+ After completing (SUCCESS or FAILURE), record run execution details to `.jonah-fleet/run-report.md`. Include:
141
156
 
142
157
  - Prompt SHA
143
158
  - Target PR number and decision (MERGE / BOUNCE / ESCALATE)
144
159
  - Execution trace and findings summary
145
160
 
146
- **Important**: Commit the log file directly to `main` and push. Follow the Log delivery fallback in `ORCHESTRATION.md` if direct push fails.
161
+ **Issue Logging Protocol**:
162
+ - **Milestone 4 (Run Completed)**: If `$ROUTINE_ISSUE_NUMBER` is set, emit compact milestone card to close out the comment stream:
163
+ ```bash
164
+ gh issue comment "$ROUTINE_ISSUE_NUMBER" --body "### 🏁 Milestone: Run Completed
165
+ - **Phase**: \`Phase 4 Β· Reconciliation\`
166
+ - **Status**: βœ… SUCCESS
167
+ - **Target / Context**: \`PR #<PR_NUMBER>\`
168
+ - **Key Decision / Finding**: Review completed with decision \`<MERGE | BOUNCE | ESCALATE>\`.
169
+ - **Next**: Routine finished; issue closed by harness" || true
170
+ ```
171
+ - Follow the Routine Issue Logging & Telemetry Protocol in `ORCHESTRATION.md`. Never commit run logs to git branches.
@@ -11,8 +11,8 @@ This routine runs behind a **human approval gate**: it **never files autowork-re
11
11
  This routine runs in two modes: **Propose** (scheduled cron sweep / unapproved fire) and **Promote** (operator-approved fire).
12
12
 
13
13
  In **Propose mode**, SUCCESS requires:
14
- - [ ] Read current roadmap, domain documentation, closed measurement trackers, and recent feedback/analytics findings
15
- - [ ] Performed Feature Pruning & Deprecation Audit: evaluated shipped features and measurement outcomes for low-ROI (<2% user adoption or >50% failure rate) features, drafting deprecation, removal, or simplification proposals
14
+ - [ ] Read current roadmap, domain documentation, closed measurement trackers, recent feedback/analytics findings, and the latest open `🎨 Design Review` issue / recent design-review routine run issues
15
+ - [ ] Performed Feature Pruning & Deprecation Audit: evaluated shipped features, measurement outcomes (<2% user adoption or >50% failure rate), and Design Review pruning/clutter directives, drafting deprecation, removal, or simplification proposals
16
16
  - [ ] Created or updated exactly one dated staging issue (`πŸ—ΊοΈ Product Plan β€” {date}`) containing:
17
17
  - Up to 3 well-scoped proposals (Summary/Tasks/Why/Complexity), covering additions, pivots, or deprecations
18
18
  - Backlog re-ranking recommendations
@@ -51,10 +51,10 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
51
51
 
52
52
  ### Steps 1–3: Propose Mode (Staging Proposals)
53
53
 
54
- 1. Read `ROADMAP.md`, `AGENTS.md`, closed measurement trackers with `RECOMMENDATION: [PIVOT | DEPRECATE | ITERATE]`, and open issues.
54
+ 1. Read `ROADMAP.md`, `AGENTS.md`, closed measurement trackers with `RECOMMENDATION: [PIVOT | DEPRECATE | ITERATE]`, the latest open `🎨 Design Review` issue (and recent run issues with `gh issue list --label "routine:design-review"`), and open issues.
55
55
  2. **Feature Pruning & Deprecation Audit**:
56
- - Audit shipped features and closed measurement tracker verdicts.
57
- - For any feature with <2% user adoption, sub-threshold CTR, or >50% failure rate, draft explicit deprecation, removal, or pivot proposals to keep the codebase lean and eliminate maintenance waste.
56
+ - Audit shipped features, closed measurement tracker verdicts, and `🎨 Design Review` clutter/pruning findings.
57
+ - For any feature with <2% user adoption, sub-threshold CTR, >50% failure rate, or persistent UI clutter flagged by Design Review, draft explicit deprecation, removal, or pivot proposals to keep the codebase lean and eliminate maintenance waste.
58
58
  3. Draft up to 3 high-impact proposals (including additions, pivots, or deprecations) based on roadmap priorities, measurement outcomes, and user feedback.
59
59
  4. For proposals sized `size/M` or above, draft a formal specification using `/to-spec`.
60
60
  5. Stage all proposals in a dedicated staging issue: `πŸ—ΊοΈ Product Plan β€” {YYYY-MM-DD}` assigned to the repo maintainer.
@@ -70,9 +70,13 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
70
70
 
71
71
  ## Logging
72
72
 
73
- After completing (SUCCESS or FAILURE), write a log file to `.github/prompts/logs/product-planning/{timestamp}.md` following the schema in `.github/prompts/logs/_template.md`. Include:
74
- - Prompt SHA
75
- - Mode (Propose or Promote)
76
- - Staged or promoted proposals tally
73
+ Follow the Routine Issue Logging Protocol in `ORCHESTRATION.md`:
74
+ 1. Write the final run report to `.jonah-fleet/run-report.md`.
75
+ 2. Include the Run Summary table with:
76
+ - Routine: `product-planning`
77
+ - Mode: `Propose` or `Promote`
78
+ - Result: `SUCCESS` or `FAILURE`
79
+ - Prompt SHA: (run `git rev-parse --short HEAD:.github/prompts/product-planning.md` or template)
80
+ - Staged or promoted proposals tally
81
+ 3. The surrounding execution harness will reconcile the corresponding GitHub issue.
77
82
 
78
- **Important**: Commit the log file directly to `main` and push. Follow the Log delivery fallback in `ORCHESTRATION.md` if direct push fails.