jonah-fleet 1.8.0 β 1.9.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +52 -0
- package/README.md +3 -3
- package/dist/commands/labels.d.ts +1 -1
- package/dist/commands/labels.d.ts.map +1 -1
- package/dist/commands/status.d.ts.map +1 -1
- package/dist/commands/sync.d.ts.map +1 -1
- package/dist/commands/telemetry.d.ts.map +1 -1
- package/dist/index.js +666 -151
- package/dist/lib/daemon-keys.d.ts.map +1 -1
- package/dist/lib/daemon.d.ts.map +1 -1
- package/dist/lib/fleet-query.d.ts.map +1 -1
- package/dist/lib/labels.d.ts +20 -0
- package/dist/lib/labels.d.ts.map +1 -1
- package/dist/lib/manifest.d.ts.map +1 -1
- package/dist/lib/presets.d.ts +6 -1
- package/dist/lib/presets.d.ts.map +1 -1
- package/dist/lib/runner.d.ts +48 -0
- package/dist/lib/runner.d.ts.map +1 -1
- package/dist/lib/telemetry.d.ts.map +1 -1
- package/dist/lib/terminal-card.d.ts +15 -1
- package/dist/lib/terminal-card.d.ts.map +1 -1
- package/package.json +1 -1
- package/schema.json +10 -5
- package/templates/prompts/ORCHESTRATION.md +49 -11
- package/templates/prompts/_prompt-template.md +10 -13
- package/templates/prompts/analytics-review.md +4 -2
- package/templates/prompts/autowork.md +43 -12
- package/templates/prompts/dependency-update-security-check.md +4 -2
- package/templates/prompts/design-review.md +125 -0
- package/templates/prompts/issues-housekeeping.md +9 -7
- package/templates/prompts/optimizer.md +31 -12
- package/templates/prompts/peer-review.md +29 -4
- package/templates/prompts/product-planning.md +14 -10
- package/templates/workflows/autowork-cron.yml +99 -39
- package/templates/workflows/dependency-check-cron.yml +89 -3
- package/templates/workflows/design-review-cron.yml +221 -0
- package/templates/workflows/issues-housekeeping-cron.yml +89 -3
- package/templates/workflows/prompt-optimizer-cron.yml +96 -10
- package/templates/workflows/sync-fleet.yml +4 -0
- package/templates/workflows/trigger-autowork-manual.yml +116 -12
- package/templates/workflows/trigger-autowork-on-bug.yml +135 -5
- package/templates/workflows/trigger-autowork-on-merge.yml +135 -5
- package/templates/workflows/trigger-review-routine.yml +119 -18
|
@@ -86,7 +86,7 @@ a. **Read the target issue and check eligibility.** Eligible = open, unassigned
|
|
|
86
86
|
- **Design System & Viewport Pre-flight** (for frontend/UI diffs): self-audit diffs against design tokens (no arbitrary class overrides), WCAG AA 4.5:1 contrast ratios on dark/light surfaces, single primary CTA hierarchy per screen, and mobile viewport crowding (avoid stacked nudges/banners above the fold at ~390px).
|
|
87
87
|
- **Documentation accuracy**: update relevant docs (`ARCHITECTURE.md`, `CODEMAP.md`, `API.md`, `CHANGELOG.md` if maintained by repo).
|
|
88
88
|
- **Build & type-check verification**: run the repository's test, type-check, and lint commands from `AGENTS.md` (e.g. `npm test`, `npm run type-check`, `npm run lint`, `pytest`, `cargo test`). Confirm zero errors and zero test failures.
|
|
89
|
-
- **Clean-
|
|
89
|
+
- **Active Origin Sync & Clean-Merge Gate**: Run `git fetch origin main && git merge origin/main --no-edit` to absorb any newly merged pull requests and resolve any conflicts locally. Verify `git merge-tree origin/main HEAD` reports no conflicts before marking ready.
|
|
90
90
|
- **Release claim on ready**: mark the PR ready (`gh pr ready <PR>`) and unassign yourself (`gh pr edit <PR> --remove-assignee <login>`) so Peer Review can evaluate without holding stale agent reservation locks.
|
|
91
91
|
Only mark the PR ready after passing every check above.
|
|
92
92
|
3b. **Ping-pong cap**: If this same PR has bounced between draft and ready 3 or more times over the same substantive finding, stop re-marking it ready. Post a comment summarizing the disagreement for human resolution and leave the PR in draft.
|
|
@@ -139,10 +139,28 @@ a. **Read the target issue and check eligibility.** Eligible = open, unassigned
|
|
|
139
139
|
- Read `π§ Decomposition plan` comment (or create if first run).
|
|
140
140
|
- Pick next slice(s), batching up to 3 same-recipe slices into one child issue + PR.
|
|
141
141
|
- Create and claim child issue first, then update plan marker to `π§ in progress β child #M`, release umbrella claim, and implement against child.
|
|
142
|
+
- **Milestone 1 (Intake & Strategy)**: If `$ROUTINE_ISSUE_NUMBER` is set, emit milestone card to the routine issue thread:
|
|
143
|
+
```bash
|
|
144
|
+
gh issue comment "$ROUTINE_ISSUE_NUMBER" --body "### π§ Milestone: Intake & Strategy
|
|
145
|
+
- **Phase**: \`Phase 1 Β· Target Selection & Planning\`
|
|
146
|
+
- **Status**: β³ In Progress
|
|
147
|
+
- **Target / Context**: \`#<TARGET_ISSUE> (<Title>)\`
|
|
148
|
+
- **Key Decision / Finding**: <Summary of plan, root cause, or reproduction strategy>
|
|
149
|
+
- **Next**: Implementation & Test Verification" || true
|
|
150
|
+
```
|
|
142
151
|
13. **Implementation & PR creation:**
|
|
143
152
|
- Branch from freshly fetched `origin/main` with descriptive name (e.g. `feat/...` or `fix/...`).
|
|
144
153
|
- Drive implementation via `/tdd` (red-green-refactor).
|
|
145
154
|
- Run repository tests and verification.
|
|
155
|
+
- **Milestone 2 (Verification & Tests)**: Once implementation passes tests and type checks, emit milestone card if `$ROUTINE_ISSUE_NUMBER` is set:
|
|
156
|
+
```bash
|
|
157
|
+
gh issue comment "$ROUTINE_ISSUE_NUMBER" --body "### π§ͺ Milestone: Verification & Tests
|
|
158
|
+
- **Phase**: \`Phase 2 Β· Implementation & Verification\`
|
|
159
|
+
- **Status**: β
Tests Passing
|
|
160
|
+
- **Target / Context**: \`#<TARGET_ISSUE>\`
|
|
161
|
+
- **Key Decision / Finding**: Tests verified passing with zero failures and clean type-check.
|
|
162
|
+
- **Next**: PR Creation & Autonomous Handoff" || true
|
|
163
|
+
```
|
|
146
164
|
- Check for competing open PRs immediately before creating PR. If collision, bail cleanly.
|
|
147
165
|
- Open draft PR via `gh pr create --draft --head <branch> --base main --title "<title>" --body "<body referencing Closes #N>"`.
|
|
148
166
|
- **PR Priority Label Mirroring**: If the issue carried a priority label (`priority/P0`, `priority/P1`, `priority/P2`, `priority/P3`), add the identical priority label to the PR (`gh pr edit <PR> --add-label "<label>"` or via `--label` in create) so downstream review workflows can filter triggers immediately.
|
|
@@ -150,6 +168,15 @@ a. **Read the target issue and check eligibility.** Eligible = open, unassigned
|
|
|
150
168
|
- **Issue Cross-Reference Comment Guardrail**: Post an explicit comment on the tracking issue referencing the newly created PR (`gh issue comment <ISSUE_NUMBER> --body "Work in progress in PR #<PR_NUMBER>."`). This guarantees an unambiguous, permanent link on the issue timeline even when GitHub's native UI link is suppressed for bot draft PRs.
|
|
151
169
|
- **Warm Context Assignment**: Assign yourself to the newly opened PR (`gh pr edit <PR> --add-assignee <login>`) to hold the reservation across Step 15's in-session review wait so parallel Scan routines recognise the PR as actively held by a live session.
|
|
152
170
|
- Mark PR ready (`gh pr ready <PR>`).
|
|
171
|
+
- **Milestone 3 (Autonomous Handoff)**: Once PR is marked ready, emit milestone card if `$ROUTINE_ISSUE_NUMBER` is set:
|
|
172
|
+
```bash
|
|
173
|
+
gh issue comment "$ROUTINE_ISSUE_NUMBER" --body "### π Milestone: Autonomous Handoff
|
|
174
|
+
- **Phase**: \`Phase 3 Β· Autonomous Handoff\`
|
|
175
|
+
- **Status**: β
PR Marked Ready
|
|
176
|
+
- **Target / Context**: \`PR #<PR_NUMBER> (closes #<TARGET_ISSUE>)\`
|
|
177
|
+
- **Key Decision / Finding**: Pushed branch and marked PR ready for review.
|
|
178
|
+
- **Next**: In-session review wait or run completion" || true
|
|
179
|
+
```
|
|
153
180
|
14. If run aborts before opening PR, release claim (unassign).
|
|
154
181
|
15. **In-Session Peer Review Wait & Immediate Convergence (Warm Context):**
|
|
155
182
|
- Poll PR status for up to 10β12 minutes.
|
|
@@ -159,23 +186,27 @@ a. **Read the target issue and check eligibility.** Eligible = open, unassigned
|
|
|
159
186
|
|
|
160
187
|
## Logging
|
|
161
188
|
|
|
162
|
-
After completing (SUCCESS or FAILURE),
|
|
189
|
+
After completing (SUCCESS or FAILURE), record run execution details to `.jonah-fleet/run-report.md`. Include:
|
|
163
190
|
|
|
164
191
|
- The prompt SHA (run `git rev-parse --short HEAD:.github/prompts/autowork.md`)
|
|
165
192
|
- Every Definition of Done criterion with YES/NO and evidence
|
|
166
193
|
- Full execution trace with tool calls
|
|
167
194
|
- If FAILURE: root cause, category, and suggested fix
|
|
168
195
|
|
|
169
|
-
**
|
|
196
|
+
**Issue Logging Protocol & Invariants**:
|
|
170
197
|
|
|
171
|
-
- **Negative Rule**: NEVER commit or push run logs to
|
|
172
|
-
- **
|
|
173
|
-
- **Direct Push to `main`**: Commit the log file directly to `main` and push β explicitly permitted for files under `.github/prompts/logs/**`:
|
|
198
|
+
- **Negative Rule**: NEVER commit or push run logs to any git branch (`main` or feature branches). Run logs are recorded exclusively as GitHub Issues and local `.jonah-fleet/runs/*.json` artifacts.
|
|
199
|
+
- **Milestone 4 (Run Completed)**: If `$ROUTINE_ISSUE_NUMBER` is set in the environment, append the concluding compact milestone card to close out the comment stream:
|
|
174
200
|
```bash
|
|
175
|
-
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
|
|
179
|
-
|
|
201
|
+
gh issue comment "$ROUTINE_ISSUE_NUMBER" --body "### π Milestone: Run Completed
|
|
202
|
+
- **Phase**: \`Phase 4 Β· Reconciliation\`
|
|
203
|
+
- **Status**: β
SUCCESS
|
|
204
|
+
- **Target / Context**: \`#<TARGET_ISSUE>\` Β· \`PR #<PR_NUMBER>\`
|
|
205
|
+
- **Key Decision / Finding**: All DoD criteria satisfied and verified.
|
|
206
|
+
- **Next**: Routine finished; issue closed by harness" || true
|
|
180
207
|
```
|
|
181
|
-
|
|
208
|
+
- **Harness Reconciliation**: When running under GitHub Actions or `jonah-fleet daemon`, the harness initializes the run issue (`status:running`), replaces the top-level issue body with `.jonah-fleet/run-report.md`, and reconciles the final status upon completion:
|
|
209
|
+
- On **SUCCESS**: Labels `status:success`, removes `status:running`, and closes the issue.
|
|
210
|
+
- On **FAILURE**: If interrupted or failed, the harness posts the Interruption Card, edits the body, and marks `status:failure,needs-attention`.
|
|
211
|
+
- On clean local runs outside a harness, write `.jonah-fleet/runs/{timestamp}.json`.
|
|
212
|
+
- Follow the Routine Issue Logging & Telemetry Protocol in `ORCHESTRATION.md`.
|
|
@@ -36,9 +36,11 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
|
|
|
36
36
|
|
|
37
37
|
## Logging
|
|
38
38
|
|
|
39
|
-
After completing (SUCCESS or FAILURE),
|
|
39
|
+
After completing (SUCCESS or FAILURE), record run execution details to `.jonah-fleet/run-report.md`. Include:
|
|
40
40
|
- Prompt SHA
|
|
41
41
|
- Tally of audited dependencies and security findings
|
|
42
42
|
- List of created or updated issues
|
|
43
43
|
|
|
44
|
-
**
|
|
44
|
+
**Issue Logging Protocol**:
|
|
45
|
+
- Record run execution details to `.jonah-fleet/run-report.md` (or update `$ROUTINE_ISSUE_NUMBER`).
|
|
46
|
+
- Follow the Routine Issue Logging & Telemetry Protocol in `ORCHESTRATION.md`. Never commit run logs to git branches.
|
|
@@ -0,0 +1,125 @@
|
|
|
1
|
+
# Design Review
|
|
2
|
+
|
|
3
|
+
## Objective
|
|
4
|
+
|
|
5
|
+
Audit player-facing surfaces for design system token deviations, visual clutter, candidates for feature pruning, and high-impact UX improvements. Provide a continuous, budget-capped quality loop across the app by combining dynamic route discovery, static code analysis, and headless visual critiques. Directly file concrete, self-contained design fixes for Autowork while staging structural UX opportunities and pruning proposals into a dedicated Design Review issue for Product Planning to ingest.
|
|
6
|
+
|
|
7
|
+
## Definition of Done
|
|
8
|
+
|
|
9
|
+
The run is SUCCESS only if ALL of these are true:
|
|
10
|
+
|
|
11
|
+
- [ ] Evaluated target surfaces: if `TARGET_SURFACE` was provided via manual dispatch, targeted that specific surface; otherwise performed dynamic route discovery across `src/app/[locale]/**/page.tsx` (excluding `admin/**`), prioritized any `[NEW SURFACE]` at Priority 0, and proceeded through the rotation sequence within the 30-iteration budget
|
|
12
|
+
- [ ] Performed Tier 1 Static Code Audit on target surfaces: checked against `DESIGN_SYSTEM.md` foundations, flagging arbitrary Tailwind values (`text-[...]`, `w-[...]`, arbitrary colors), non-standard button styling, multiple `.btn-primary` actions per view, and contrast token regressions
|
|
13
|
+
- [ ] Performed Tier 2 Headless Browser & Visual Critique: launched the dev environment and captured viewports at mobile (390px) and desktop (1280px) via `run` and `design-critique` skills, auditing whitespace consistency, visual hierarchy, touch targets, and viewport density (enforcing the mobile fold budget: max 1 banner/nudge)
|
|
14
|
+
- [ ] Cross-referenced UI friction and telemetry signals: checked for recurring rage clicks (`$rageclick`) and identified low-interaction (<2% CTR) secondary elements or persistent nudges suitable for pruning
|
|
15
|
+
- [ ] Executed Hybrid Handoff:
|
|
16
|
+
- If self-contained token, class, or contrast violations were found, filed or updated exactly one grouped punchlist issue titled `π¨ Design Polish & Token Cleanup β {YYYY-MM-DD}` labeled `design/polish` for Autowork
|
|
17
|
+
- Created or updated exactly one tracking issue titled `π¨ Design Review β {YYYY-MM-DD}` documenting: Surfaces Reviewed (with any `[NEW SURFACE]` highlighted), Next Surfaces in Rotation, Visual Critique & Clutter Analysis, Feature Pruning & Deprecation Recommendations, and Experience Enhancement Opportunities
|
|
18
|
+
- [ ] Emitted structured design directives in the run report (`DESIGN_DIRECTIVE: [PRUNE | REDESIGN | POLISH]`) and recorded the next queued surfaces for the subsequent run
|
|
19
|
+
- [ ] Recorded the final execution report to `.jonah-fleet/run-report.md` following `ORCHESTRATION.md`
|
|
20
|
+
|
|
21
|
+
If any criterion cannot be met, stop immediately and log FAILURE with the reason.
|
|
22
|
+
|
|
23
|
+
## Constraints
|
|
24
|
+
|
|
25
|
+
- **Max iterations**: 30 β after 30 tool call rounds without completing Definition of Done, STOP. Log FAILURE with category `token_limit`.
|
|
26
|
+
- **Max scope**: auditing player-facing UI surfaces only (`src/app/[locale]/**/page.tsx`, excluding `admin/**`). Do not implement feature code or open feature PRs directly.
|
|
27
|
+
- **Budget-capped rotation**: review as many prioritized surfaces as possible within the iteration budget; do not attempt an exhaustive sweep if approaching the tool call ceiling. Stamp remaining surfaces as `Next Surfaces in Rotation`.
|
|
28
|
+
- **Language Requirement**: All GitHub issue titles, descriptions, task checklists, and comments MUST be written in **English**.
|
|
29
|
+
- **Session link footer**: sign every GitHub post with the Antigravity run footer (`_Generated by [Antigravity](${GITHUB_SERVER_URL}/${GITHUB_REPOSITORY}/actions/runs/${GITHUB_RUN_ID})_`).
|
|
30
|
+
|
|
31
|
+
## Negative examples (DO NOT do these)
|
|
32
|
+
|
|
33
|
+
- Do not open multiple separate issues for individual token mismatches β group all simple token and contrast cleanup items into a single grouped punchlist issue for Autowork.
|
|
34
|
+
- Do not file unapproved structural redesigns or feature removals as direct Autowork issues β stage them in the `π¨ Design Review β {date}` tracking issue for Product Planning to review.
|
|
35
|
+
- Do not review the admin back-office (`src/app/[locale]/admin/**`) β it is deliberately utilitarian and exempt from the player-facing design system polish bar.
|
|
36
|
+
- Do not blindly repeat the same surface review order every run β dynamically discover routes and check previous review issues to resume rotation and fast-track unreviewed new surfaces.
|
|
37
|
+
|
|
38
|
+
## Instructions
|
|
39
|
+
|
|
40
|
+
### Step 1: Dynamic Route Discovery & Target Prioritization
|
|
41
|
+
|
|
42
|
+
1. **Check for Manual Target Override**:
|
|
43
|
+
- If the environment variable `TARGET_SURFACE` is provided and non-empty, audit only that specified surface or route and skip rotation determination.
|
|
44
|
+
2. **Dynamic Route Discovery**:
|
|
45
|
+
- Scan the filesystem for all active player-facing page routes: `src/app/[locale]/**/page.tsx` (ignoring any `admin/**` directories).
|
|
46
|
+
- Normalize the discovered paths into canonical route identifiers (e.g. `Home` (`/`), `Predictions` (`/pronostics`), `Episode Predictions` (`/pronostics/episode/[id]`), `Groups` (`/groupes`), `Group Detail` (`/groupes/[id]`), `Presenter` (`/groupes/[id]/presentateur`), `Leaderboard` (`/classements`), `Profile` (`/profil`), `Account` (`/compte`), `Auth` (`/auth/*`), `Season Overview` (`/saisons/[slug]`), `Share` (`/partage`)).
|
|
47
|
+
3. **Reconcile with Previous Rotation State**:
|
|
48
|
+
- Locate the most recent `π¨ Design Review` issue or run issue labeled `routine:design-review`.
|
|
49
|
+
- Read the `Surfaces Reviewed` and `Next Surfaces in Rotation` fields.
|
|
50
|
+
- Any discovered route on disk that has never appeared in prior review logs is flagged as **`[NEW SURFACE]`** and placed at **Priority 0** (Fast-Track).
|
|
51
|
+
- Remaining surfaces are sequenced based on `Next Surfaces in Rotation` (resuming rotation where the previous run stopped).
|
|
52
|
+
- Any route present in the previous rotation list that no longer exists on disk is automatically pruned.
|
|
53
|
+
4. **Set Run Queue**:
|
|
54
|
+
- Build the active review queue: `[NEW SURFACE]` items first, followed by the queued rotation surfaces.
|
|
55
|
+
- Keep a budget awareness counter: as iterations proceed, stop adding new surfaces when remaining tool calls are needed for issues and logging.
|
|
56
|
+
|
|
57
|
+
### Step 2: Static Code & Token Audit
|
|
58
|
+
|
|
59
|
+
For each surface in the active queue:
|
|
60
|
+
1. Locate the page component and its imported presentation components under `src/components/**`.
|
|
61
|
+
2. Inspect the styling using `.agents/skills/design-system/SKILL.md`:
|
|
62
|
+
- **Token Purity**: flag hardcoded hex codes (`#...`), rgb values, or arbitrary Tailwind brackets (e.g. `text-[13px]`, `w-[310px]`).
|
|
63
|
+
- **Surface Hierarchy**: verify background hierarchy adheres to `bg-gray-950` (base) β `bg-gray-900` / `.card` (card) β `bg-gray-800` / `.input-field` (input).
|
|
64
|
+
- **Contrast Tokens**: ensure readable secondary text uses `text-gray-400` minimum (never `text-gray-500`, which fails WCAG AA on dark backgrounds).
|
|
65
|
+
- **CTA Discipline**: verify each view/tab contains at most **one** primary action button (`.btn-primary`). Flag competing primary buttons.
|
|
66
|
+
- **Button & Pill Standards**: verify buttons use established classes (`.btn-primary`, `.btn-secondary`, `.btn-ghost`) and status pills use standard semantic colors.
|
|
67
|
+
|
|
68
|
+
### Step 3: Headless Browser & Visual Critique
|
|
69
|
+
|
|
70
|
+
For the surfaces being audited:
|
|
71
|
+
1. Start the local Next.js dev server with mock credentials following `.agents/skills/run/SKILL.md`:
|
|
72
|
+
```bash
|
|
73
|
+
NEXT_PUBLIC_SUPABASE_URL=http://localhost:54321 NEXT_PUBLIC_SUPABASE_ANON_KEY=placeholder npm run dev
|
|
74
|
+
```
|
|
75
|
+
2. Using Playwright, capture screenshots at two standard viewports:
|
|
76
|
+
- **Mobile**: 390px width
|
|
77
|
+
- **Desktop**: 1280px width
|
|
78
|
+
3. Apply `.agents/skills/design-critique/SKILL.md` to evaluate visual evidence:
|
|
79
|
+
- **Viewport Density & Nudge Budget**: On mobile (390px), do sticky headers, promotional banners, and teasers crowd the viewport above the fold? Enforce: max 1 banner/nudge visible concurrently.
|
|
80
|
+
- **Visual Hierarchy & Clutter**: Are cards nested redundantly? Is there visual noise from excessive borders, conflicting accents, or dense text blocks?
|
|
81
|
+
- **Touch Targets & Spacing**: Are tap targets at least 44x44px? Is whitespace rhythmic and intentional?
|
|
82
|
+
4. Shut down the dev server cleanly after capture.
|
|
83
|
+
|
|
84
|
+
### Step 4: Telemetry & Friction Cross-Reference
|
|
85
|
+
|
|
86
|
+
1. If telemetry endpoints or PostHog logs are available:
|
|
87
|
+
- Audit `$rageclick` events associated with URLs and selectors of the reviewed surfaces. Pinpoint confusing or non-interactive elements masquerading as buttons.
|
|
88
|
+
- Review click-through rates (CTR) on secondary promotional cards or banner elements.
|
|
89
|
+
2. Formulate **Pruning Recommendations**:
|
|
90
|
+
- Any persistent card, teaser, or secondary control with near-zero interaction, high rage-click density, or severe visual clutter is recommended for deprecation (`RECOMMENDATION: DEPRECATE`) or simplification (`RECOMMENDATION: PRUNE`).
|
|
91
|
+
|
|
92
|
+
### Step 5: Hybrid Handoff & Issue Generation
|
|
93
|
+
|
|
94
|
+
1. **Direct Autowork Punchlist Issue**:
|
|
95
|
+
- If concrete, self-contained token violations, contrast bugs, or duplicate `.btn-primary` classes are found, compile them into a single grouped punchlist issue:
|
|
96
|
+
- Title: `π¨ Design Polish & Token Cleanup β {YYYY-MM-DD}`
|
|
97
|
+
- Labels: `design/polish`, `autowork-candidate`
|
|
98
|
+
- Content: Structured task list with target files, line references, current violation, and suggested replacement token/class.
|
|
99
|
+
2. **Overarching Design Review Tracking Issue**:
|
|
100
|
+
- Create or update a single tracking issue:
|
|
101
|
+
- Title: `π¨ Design Review β {YYYY-MM-DD}`
|
|
102
|
+
- Labels: `area/design`, `review`
|
|
103
|
+
- Content:
|
|
104
|
+
- **Surfaces Reviewed**: list audited routes, noting any `[NEW SURFACE]`
|
|
105
|
+
- **Next Surfaces in Rotation**: list unreviewed routes queued for the subsequent run
|
|
106
|
+
- **Visual Critique & Clutter Analysis**: mobile fold density, hierarchy, and layout observations
|
|
107
|
+
- **Pruning & Deprecation Candidates**: specific features/elements recommended to be pruned or consolidated
|
|
108
|
+
- **Experience Enhancement Opportunities**: structural UX improvements staged for Product Planning
|
|
109
|
+
3. **Structured Directives**:
|
|
110
|
+
- Output clear directives: `DESIGN_DIRECTIVE: [PRUNE | REDESIGN | POLISH]` to be consumed downstream by `product-planning.md`.
|
|
111
|
+
|
|
112
|
+
## Logging
|
|
113
|
+
|
|
114
|
+
Follow the Routine Issue Logging Protocol in `ORCHESTRATION.md`:
|
|
115
|
+
1. Write the final run report to `.jonah-fleet/run-report.md`.
|
|
116
|
+
2. Include the Run Summary table with:
|
|
117
|
+
- Routine: `design-review`
|
|
118
|
+
- Prompt SHA (run `git rev-parse --short HEAD:.github/prompts/design-review.md` or template)
|
|
119
|
+
- Result: `SUCCESS` or `FAILURE`
|
|
120
|
+
- Every Definition of Done criterion with YES/NO and evidence
|
|
121
|
+
- Surfaces Reviewed and Next Surfaces in Rotation
|
|
122
|
+
- Created issues (`π¨ Design Polish & Token Cleanup`, `π¨ Design Review`)
|
|
123
|
+
- If FAILURE: root cause, category, and suggested fix
|
|
124
|
+
3. The surrounding execution harness will reconcile the corresponding GitHub issue.
|
|
125
|
+
|
|
@@ -15,7 +15,7 @@ The run is SUCCESS only if ALL of these are true:
|
|
|
15
15
|
- [ ] Label audit completed: every open issue has consistent type, size, and priority labels
|
|
16
16
|
- [ ] Dependency check completed: issues with `## Dependencies` verified against blocker status
|
|
17
17
|
- [ ] Orphaned-claim sweep completed: stale autowork claims (per `ORCHESTRATION.md`) released back to the unclaimed pool
|
|
18
|
-
- [ ]
|
|
18
|
+
- [ ] Stalled routine run audit completed: routine log issues in `status:running` older than 6 hours marked as `status:failure` (timed_out)
|
|
19
19
|
- [ ] Summary posted listing all changes made
|
|
20
20
|
|
|
21
21
|
If any criterion cannot be met, stop immediately and log FAILURE with the reason.
|
|
@@ -24,14 +24,14 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
|
|
|
24
24
|
|
|
25
25
|
- **Max iterations**: 40 β after 40 tool call rounds without completing Definition of Done, STOP. Log FAILURE with category `token_limit`.
|
|
26
26
|
- **Max scope**: housekeeping only. Do not implement code fixes or open feature PRs.
|
|
27
|
-
- **No speculative work**: only modify issue metadata (labels, status, comments, releasing stale assignees
|
|
27
|
+
- **No speculative work**: only modify issue metadata (labels, status, comments, releasing stale assignees, reconciling stalled routine runs).
|
|
28
28
|
- **Language Requirement**: All GitHub issue titles, descriptions, task checklists, and comments MUST be written in **English**.
|
|
29
29
|
|
|
30
30
|
## Instructions
|
|
31
31
|
|
|
32
32
|
### Phase 1: Quick Recovery & Clearing
|
|
33
33
|
|
|
34
|
-
1. **
|
|
34
|
+
1. **Stalled routine run sweep**: Check open issues with labels `routine-log` and `status:running`. If an issue has been in `status:running` for > 6 hours without updates, add label `status:failure`, remove `status:running`, add label `needs-attention`, and comment noting the runner timeout or crash.
|
|
35
35
|
2. **Orphaned-claim sweep**: Sweep assigned issues. If an issue meets the 3 stale-claim conditions in `ORCHESTRATION.md` (autowork claim comment, no open PR, comment > 6 hours old), re-read immediately before writing, unassign the dead owner, and post a release comment.
|
|
36
36
|
|
|
37
37
|
### Phase 2: Backlog Hygiene
|
|
@@ -44,13 +44,15 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
|
|
|
44
44
|
|
|
45
45
|
### Phase 3: Summary
|
|
46
46
|
|
|
47
|
-
8. Post a summary comment or log recording all actions taken (priority shifts, closed duplicates, released claims,
|
|
47
|
+
8. Post a summary comment or log recording all actions taken (priority shifts, closed duplicates, released claims, reconciled stalled runs).
|
|
48
48
|
|
|
49
49
|
## Logging
|
|
50
50
|
|
|
51
|
-
After completing (SUCCESS or FAILURE),
|
|
51
|
+
After completing (SUCCESS or FAILURE), record run execution details to `.jonah-fleet/run-report.md`. Include:
|
|
52
52
|
- Prompt SHA
|
|
53
|
-
- Tally of issues audited, claims released,
|
|
53
|
+
- Tally of issues audited, claims released, stalled runs reconciled
|
|
54
54
|
- List of closed or modified issues
|
|
55
55
|
|
|
56
|
-
**
|
|
56
|
+
**Issue Logging Protocol**:
|
|
57
|
+
- Record run execution details to `.jonah-fleet/run-report.md` (or update `$ROUTINE_ISSUE_NUMBER`).
|
|
58
|
+
- Follow the Routine Issue Logging & Telemetry Protocol in `ORCHESTRATION.md`. Never commit run logs to git branches.
|
|
@@ -10,10 +10,10 @@ Additionally, this routine acts as the **Upstream Evolution Bridge**: when a pro
|
|
|
10
10
|
|
|
11
11
|
The run is SUCCESS if ALL of these are true:
|
|
12
12
|
|
|
13
|
-
- [ ] All
|
|
14
|
-
- [ ] Every FAILURE
|
|
15
|
-
- [ ] Inefficiency and review loops per PR have been computed across SUCCESS
|
|
16
|
-
- [ ] Per-agent token and cost consumption metrics have been aggregated across in-window
|
|
13
|
+
- [ ] All routine run issues from the incremental window (labeled `routine-log`) have been scanned
|
|
14
|
+
- [ ] Every FAILURE run has been categorized and analyzed
|
|
15
|
+
- [ ] Inefficiency and review loops per PR have been computed across SUCCESS runs
|
|
16
|
+
- [ ] Per-agent token and cost consumption metrics have been aggregated across in-window routine runs and evaluated against the 70% weekly budget ceiling in `ORCHESTRATION.md`
|
|
17
17
|
- [ ] Closed bug issues and merged bug-fix PRs in the window have been analyzed for systemic root causes
|
|
18
18
|
- [ ] For each fixable pattern:
|
|
19
19
|
- If project-specific: opened a local PR with a prompt, template, or test fix and marked ready for review
|
|
@@ -27,23 +27,32 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
|
|
|
27
27
|
|
|
28
28
|
- **Max iterations**: 30 β stop after 30 tool call rounds.
|
|
29
29
|
- **Max scope**: one PR per identified problem. Do not bundle unrelated fixes.
|
|
30
|
-
- **No speculative work**: only fix patterns evidenced by
|
|
30
|
+
- **No speculative work**: only fix patterns evidenced by routine issues, closed bug issues, or reviewer findings.
|
|
31
31
|
- **Language Requirement**: All GitHub issue titles, descriptions, task checklists, and comments MUST be written in **English**.
|
|
32
32
|
|
|
33
33
|
## Instructions
|
|
34
34
|
|
|
35
35
|
### 0. Establish the incremental scan boundary
|
|
36
36
|
|
|
37
|
-
1. Check
|
|
37
|
+
1. Check the most recent routine issue with labels `routine-log,routine:optimizer` via `gh issue list --label routine-log --label routine:optimizer --state all --limit 1`.
|
|
38
38
|
2. Extract the timestamp as the scan boundary (or last 7 days if first run).
|
|
39
|
+
3. **Milestone 1 (Intake & Fleet Telemetry Scan)**: If `$ROUTINE_ISSUE_NUMBER` is set, emit milestone card to the routine issue thread:
|
|
40
|
+
```bash
|
|
41
|
+
gh issue comment "$ROUTINE_ISSUE_NUMBER" --body "### π§ Milestone: Intake & Fleet Telemetry Scan
|
|
42
|
+
- **Phase**: \`Phase 1 Β· Telemetry Intake & Boundary Scan\`
|
|
43
|
+
- **Status**: β³ In Progress
|
|
44
|
+
- **Target / Context**: \`Fleet Telemetry Window\`
|
|
45
|
+
- **Key Decision / Finding**: Established scan boundary; analyzing in-window routine issues and token pacing.
|
|
46
|
+
- **Next**: Anomaly Diagnosis & Proposals" || true
|
|
47
|
+
```
|
|
39
48
|
|
|
40
49
|
### 1. Collect signals & analyze logs
|
|
41
50
|
|
|
42
|
-
1. **
|
|
43
|
-
2. **Extract failure categories**: Categorize runs logging `FAILURE` (`prompt_unclear`, `data_issue`, `token_limit`, `infeasible_task`).
|
|
51
|
+
1. **Query in-window routine issues**: List routine run issues within the window using `gh issue list --label routine-log --state all --limit 100 --json number,title,body,state,labels,createdAt`.
|
|
52
|
+
2. **Extract failure categories**: Categorize runs labeled `status:failure` or logging `FAILURE` (`prompt_unclear`, `data_issue`, `token_limit`, `infeasible_task`).
|
|
44
53
|
3. **Compute efficiency metrics**: Identify PRs experiencing $\ge 3$ review bounce rounds and runs with high iteration usage relative to limits.
|
|
45
54
|
4. **Aggregate per-agent token & cost consumption**:
|
|
46
|
-
- Parse the metadata table from each in-window
|
|
55
|
+
- Parse the metadata table from each in-window routine issue body: `Routine`, `Input tokens`, `Output tokens`, `Estimated cost`, `Iterations used` (e.g. `26 / 65`), and `Result` (`SUCCESS` or `FAILURE`).
|
|
47
56
|
- Group logs by `Routine` (`autowork`, `peer-review`, `issues-housekeeping`, `dependency-update-security-check`, `optimizer`, `product-planning`).
|
|
48
57
|
- For each routine, compute:
|
|
49
58
|
- **Run count**: total completed runs.
|
|
@@ -88,9 +97,9 @@ Translate findings into concrete preventative improvements and remediation trigg
|
|
|
88
97
|
|
|
89
98
|
## Logging
|
|
90
99
|
|
|
91
|
-
After completing (SUCCESS or FAILURE),
|
|
100
|
+
After completing (SUCCESS or FAILURE), record run execution details to `.jonah-fleet/run-report.md`. Include:
|
|
92
101
|
- Prompt SHA
|
|
93
|
-
- Analyzed
|
|
102
|
+
- Analyzed routine issues count and identified patterns
|
|
94
103
|
- **Token & Cost Consumption by Agent** scorecard table:
|
|
95
104
|
|
|
96
105
|
```markdown
|
|
@@ -109,5 +118,15 @@ After completing (SUCCESS or FAILURE), write a log file to `.github/prompts/logs
|
|
|
109
118
|
- Weekly token budget pacing evaluation (pacing vs 70% ceiling in `ORCHESTRATION.md`)
|
|
110
119
|
- PRs opened (local or upstream)
|
|
111
120
|
|
|
112
|
-
**
|
|
121
|
+
**Issue Logging Protocol**:
|
|
122
|
+
- **Milestone 4 (Run Completed)**: If `$ROUTINE_ISSUE_NUMBER` is set, emit compact milestone card to close out the comment stream:
|
|
123
|
+
```bash
|
|
124
|
+
gh issue comment "$ROUTINE_ISSUE_NUMBER" --body "### π Milestone: Run Completed
|
|
125
|
+
- **Phase**: \`Phase 4 Β· Reconciliation\`
|
|
126
|
+
- **Status**: β
SUCCESS
|
|
127
|
+
- **Target / Context**: \`Fleet Optimization Sweep\`
|
|
128
|
+
- **Key Decision / Finding**: Fleet telemetry sweep completed and scored against token ceiling.
|
|
129
|
+
- **Next**: Routine finished; issue closed by harness" || true
|
|
130
|
+
```
|
|
131
|
+
- Follow the Routine Issue Logging & Telemetry Protocol in `ORCHESTRATION.md`. Never commit run logs to git branches.
|
|
113
132
|
|
|
@@ -70,7 +70,7 @@ Check if `$PR_NUMBER` is set:
|
|
|
70
70
|
|
|
71
71
|
### Steps 1β2: Select a PR (Scan mode only)
|
|
72
72
|
|
|
73
|
-
1. List all open PRs, excluding drafts
|
|
73
|
+
1. List all open PRs, excluding drafts and automated release PRs (`release-please--*` / `chore(main): release*`).
|
|
74
74
|
2. **Dual Execution Priority Routing**:
|
|
75
75
|
- If running in **Cloud Actions** (`$GITHUB_ACTIONS` / `$CI`): Scan mode prioritizes PRs labeled `priority/P0` or `priority/P1` (or untagged PRs). Lower-priority PRs (`priority/P2`, `priority/P3`) are eligible for cloud review ONLY if they have remained ready and unreviewed for more than 48 hours (`created_at` older than 48h, acting as a cloud catchup sweep).
|
|
76
76
|
- If running in **Local Agent** (`$LOCAL_AGENT`): Scan mode reviews any ready PR across all priorities (P0 β P1 β P2 β P3) with no age gating.
|
|
@@ -80,8 +80,21 @@ Check if `$PR_NUMBER` is set:
|
|
|
80
80
|
|
|
81
81
|
### Step 3: Round tracking & Starting Review marker
|
|
82
82
|
|
|
83
|
-
- Post a "Starting review (round N)" comment on the target PR to claim the review window
|
|
83
|
+
- Post a "Starting review (round N)" comment on the target PR to claim the review window:
|
|
84
|
+
- If `$ROUTINE_ISSUE_NUMBER` is set in the environment or routine prompt, include a link to the tracking log issue:
|
|
85
|
+
`Starting review (round N) Β· [Run Log #$ROUTINE_ISSUE_NUMBER](${GITHUB_SERVER_URL:-https://github.com}/${GITHUB_REPOSITORY}/issues/${ROUTINE_ISSUE_NUMBER})`
|
|
86
|
+
(or `Starting review (round N) Β· Run log: #$ROUTINE_ISSUE_NUMBER`).
|
|
87
|
+
- If `$ROUTINE_ISSUE_NUMBER` is not set, post `Starting review (round N)`.
|
|
84
88
|
- Check round count `N`. If `N >= 5` and blocking findings persist, prepare to escalate.
|
|
89
|
+
- **Milestone 1 (Intake & Review Scope)**: If `$ROUTINE_ISSUE_NUMBER` is set, emit milestone card to the routine issue thread:
|
|
90
|
+
```bash
|
|
91
|
+
gh issue comment "$ROUTINE_ISSUE_NUMBER" --body "### π§ Milestone: Intake & Review Scope
|
|
92
|
+
- **Phase**: \`Phase 1 Β· Review Target & Scope\`
|
|
93
|
+
- **Status**: β³ In Progress
|
|
94
|
+
- **Target / Context**: \`PR #<PR_NUMBER> (Round N)\`
|
|
95
|
+
- **Key Decision / Finding**: Claimed review window; starting multi-angle evaluation passes.
|
|
96
|
+
- **Next**: Multi-Angle Code Review Pass" || true
|
|
97
|
+
```
|
|
85
98
|
|
|
86
99
|
### Step 4: Multi-Angle Code Review Pass
|
|
87
100
|
|
|
@@ -124,6 +137,7 @@ If the PR is clean and approved for merge, but lacks a `Closes #N` tracking link
|
|
|
124
137
|
- If `N < 5`: Post inline comments, submit review as `COMMENT`, and convert PR to draft (`gh pr ready <N> --undo`).
|
|
125
138
|
- If `N >= 5`: Convert PR to draft, post summary comment escalating to repo maintainer, and apply `needs-human` label.
|
|
126
139
|
- **If Clean (or only Non-blocking findings)**:
|
|
140
|
+
- If the PR branch has minor mechanical merge conflicts against `origin/main` (e.g. adjacent `CHANGELOG.md` entries from concurrent merges) while code and tests are sound, resolve the conflict mechanically via `/resolving-merge-conflicts` before squash-merging rather than bouncing the PR back to draft.
|
|
127
141
|
- If PR lacks `Closes #N`, execute Autonomous Issue Synthesis (Step 5.5).
|
|
128
142
|
- Extract the tracking issue number `$ISSUE_NUMBER` from the PR description or title (e.g. `Closes #<N>`, `Fixes #<N>`, `Resolves #<N>`).
|
|
129
143
|
- Squash-merge the PR: `gh pr merge <N> --squash --delete-branch`.
|
|
@@ -132,15 +146,26 @@ If the PR is clean and approved for merge, but lacks a `Closes #N` tracking link
|
|
|
132
146
|
gh issue close "$ISSUE_NUMBER" --comment "Closed via PR #<N> (merged into main)."
|
|
133
147
|
```
|
|
134
148
|
- Submit held review comments.
|
|
149
|
+
- In any review summary, decision comment, or escalation comment posted to the PR, include the run log reference if `$ROUTINE_ISSUE_NUMBER` is set (e.g. `- **Run Log**: [#$ROUTINE_ISSUE_NUMBER](${GITHUB_SERVER_URL:-https://github.com}/${GITHUB_REPOSITORY}/issues/${ROUTINE_ISSUE_NUMBER})`).
|
|
135
150
|
- File follow-up issues for material non-blocking findings.
|
|
136
151
|
- If mechanical doc fixes are needed, commit directly to `main`.
|
|
137
152
|
|
|
138
153
|
## Logging
|
|
139
154
|
|
|
140
|
-
After completing (SUCCESS or FAILURE),
|
|
155
|
+
After completing (SUCCESS or FAILURE), record run execution details to `.jonah-fleet/run-report.md`. Include:
|
|
141
156
|
|
|
142
157
|
- Prompt SHA
|
|
143
158
|
- Target PR number and decision (MERGE / BOUNCE / ESCALATE)
|
|
144
159
|
- Execution trace and findings summary
|
|
145
160
|
|
|
146
|
-
**
|
|
161
|
+
**Issue Logging Protocol**:
|
|
162
|
+
- **Milestone 4 (Run Completed)**: If `$ROUTINE_ISSUE_NUMBER` is set, emit compact milestone card to close out the comment stream:
|
|
163
|
+
```bash
|
|
164
|
+
gh issue comment "$ROUTINE_ISSUE_NUMBER" --body "### π Milestone: Run Completed
|
|
165
|
+
- **Phase**: \`Phase 4 Β· Reconciliation\`
|
|
166
|
+
- **Status**: β
SUCCESS
|
|
167
|
+
- **Target / Context**: \`PR #<PR_NUMBER>\`
|
|
168
|
+
- **Key Decision / Finding**: Review completed with decision \`<MERGE | BOUNCE | ESCALATE>\`.
|
|
169
|
+
- **Next**: Routine finished; issue closed by harness" || true
|
|
170
|
+
```
|
|
171
|
+
- Follow the Routine Issue Logging & Telemetry Protocol in `ORCHESTRATION.md`. Never commit run logs to git branches.
|
|
@@ -11,8 +11,8 @@ This routine runs behind a **human approval gate**: it **never files autowork-re
|
|
|
11
11
|
This routine runs in two modes: **Propose** (scheduled cron sweep / unapproved fire) and **Promote** (operator-approved fire).
|
|
12
12
|
|
|
13
13
|
In **Propose mode**, SUCCESS requires:
|
|
14
|
-
- [ ] Read current roadmap, domain documentation, closed measurement trackers,
|
|
15
|
-
- [ ] Performed Feature Pruning & Deprecation Audit: evaluated shipped features
|
|
14
|
+
- [ ] Read current roadmap, domain documentation, closed measurement trackers, recent feedback/analytics findings, and the latest open `π¨ Design Review` issue / recent design-review routine run issues
|
|
15
|
+
- [ ] Performed Feature Pruning & Deprecation Audit: evaluated shipped features, measurement outcomes (<2% user adoption or >50% failure rate), and Design Review pruning/clutter directives, drafting deprecation, removal, or simplification proposals
|
|
16
16
|
- [ ] Created or updated exactly one dated staging issue (`πΊοΈ Product Plan β {date}`) containing:
|
|
17
17
|
- Up to 3 well-scoped proposals (Summary/Tasks/Why/Complexity), covering additions, pivots, or deprecations
|
|
18
18
|
- Backlog re-ranking recommendations
|
|
@@ -51,10 +51,10 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
|
|
|
51
51
|
|
|
52
52
|
### Steps 1β3: Propose Mode (Staging Proposals)
|
|
53
53
|
|
|
54
|
-
1. Read `ROADMAP.md`, `AGENTS.md`, closed measurement trackers with `RECOMMENDATION: [PIVOT | DEPRECATE | ITERATE]`, and open issues.
|
|
54
|
+
1. Read `ROADMAP.md`, `AGENTS.md`, closed measurement trackers with `RECOMMENDATION: [PIVOT | DEPRECATE | ITERATE]`, the latest open `π¨ Design Review` issue (and recent run issues with `gh issue list --label "routine:design-review"`), and open issues.
|
|
55
55
|
2. **Feature Pruning & Deprecation Audit**:
|
|
56
|
-
- Audit shipped features
|
|
57
|
-
- For any feature with <2% user adoption, sub-threshold CTR,
|
|
56
|
+
- Audit shipped features, closed measurement tracker verdicts, and `π¨ Design Review` clutter/pruning findings.
|
|
57
|
+
- For any feature with <2% user adoption, sub-threshold CTR, >50% failure rate, or persistent UI clutter flagged by Design Review, draft explicit deprecation, removal, or pivot proposals to keep the codebase lean and eliminate maintenance waste.
|
|
58
58
|
3. Draft up to 3 high-impact proposals (including additions, pivots, or deprecations) based on roadmap priorities, measurement outcomes, and user feedback.
|
|
59
59
|
4. For proposals sized `size/M` or above, draft a formal specification using `/to-spec`.
|
|
60
60
|
5. Stage all proposals in a dedicated staging issue: `πΊοΈ Product Plan β {YYYY-MM-DD}` assigned to the repo maintainer.
|
|
@@ -70,9 +70,13 @@ If any criterion cannot be met, stop immediately and log FAILURE with the reason
|
|
|
70
70
|
|
|
71
71
|
## Logging
|
|
72
72
|
|
|
73
|
-
|
|
74
|
-
|
|
75
|
-
|
|
76
|
-
-
|
|
73
|
+
Follow the Routine Issue Logging Protocol in `ORCHESTRATION.md`:
|
|
74
|
+
1. Write the final run report to `.jonah-fleet/run-report.md`.
|
|
75
|
+
2. Include the Run Summary table with:
|
|
76
|
+
- Routine: `product-planning`
|
|
77
|
+
- Mode: `Propose` or `Promote`
|
|
78
|
+
- Result: `SUCCESS` or `FAILURE`
|
|
79
|
+
- Prompt SHA: (run `git rev-parse --short HEAD:.github/prompts/product-planning.md` or template)
|
|
80
|
+
- Staged or promoted proposals tally
|
|
81
|
+
3. The surrounding execution harness will reconcile the corresponding GitHub issue.
|
|
77
82
|
|
|
78
|
-
**Important**: Commit the log file directly to `main` and push. Follow the Log delivery fallback in `ORCHESTRATION.md` if direct push fails.
|