@beremaran/ralphie 0.0.0-stage → 0.2.1
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/CHANGELOG.md +843 -0
- package/LICENSE +21 -0
- package/README.md +53 -2
- package/dist/ralphie.js +33689 -0
- package/docs/README.md +71 -0
- package/docs/architecture.md +159 -0
- package/docs/cli-reference.md +126 -0
- package/docs/configuration.md +269 -0
- package/docs/development.md +249 -0
- package/docs/getting-started.md +134 -0
- package/docs/operations-and-recovery.md +328 -0
- package/docs/safety.md +189 -0
- package/docs/workflows.md +411 -0
- package/package.json +83 -3
- package/vendor/mattpocock-skills/LICENSE +21 -0
- package/vendor/mattpocock-skills/code-review/SKILL.md +87 -0
- package/vendor/mattpocock-skills/code-review/agents/openai.yaml +3 -0
- package/vendor/mattpocock-skills/codebase-design/DEEPENING.md +37 -0
- package/vendor/mattpocock-skills/codebase-design/DESIGN-IT-TWICE.md +44 -0
- package/vendor/mattpocock-skills/codebase-design/SKILL.md +114 -0
- package/vendor/mattpocock-skills/codebase-design/agents/openai.yaml +3 -0
- package/vendor/mattpocock-skills/diagnosing-bugs/SKILL.md +138 -0
- package/vendor/mattpocock-skills/diagnosing-bugs/agents/openai.yaml +3 -0
- package/vendor/mattpocock-skills/diagnosing-bugs/scripts/hitl-loop.template.sh +44 -0
- package/vendor/mattpocock-skills/implement/SKILL.md +15 -0
- package/vendor/mattpocock-skills/implement/agents/openai.yaml +5 -0
- package/vendor/mattpocock-skills/lock.json +37 -0
- package/vendor/mattpocock-skills/tdd/SKILL.md +38 -0
- package/vendor/mattpocock-skills/tdd/agents/openai.yaml +3 -0
- package/vendor/mattpocock-skills/tdd/mocking.md +59 -0
- package/vendor/mattpocock-skills/tdd/tests.md +77 -0
- package/vendor/mattpocock-skills/to-tickets/SKILL.md +105 -0
- package/vendor/mattpocock-skills/to-tickets/agents/openai.yaml +5 -0
- package/vendor/mattpocock-skills/triage/AGENT-BRIEF.md +207 -0
- package/vendor/mattpocock-skills/triage/OUT-OF-SCOPE.md +105 -0
- package/vendor/mattpocock-skills/triage/SKILL.md +112 -0
- package/vendor/mattpocock-skills/triage/agents/openai.yaml +5 -0
|
@@ -0,0 +1,411 @@
|
|
|
1
|
+
# Workflows
|
|
2
|
+
|
|
3
|
+
This page is for operators and contributors who need to understand how Ralphie
|
|
4
|
+
routes issues, performs implementation and decomposition, and delivers the
|
|
5
|
+
result. It is the authoritative description of workflow semantics and diagrams;
|
|
6
|
+
see the [documentation index](README.md) for setup, CLI, safety, and recovery
|
|
7
|
+
references.
|
|
8
|
+
|
|
9
|
+
> [!CAUTION]
|
|
10
|
+
> Ralphie commits and pushes directly to the selected branch. Read the
|
|
11
|
+
> [safety model](safety.md) and validate against a repository you control before
|
|
12
|
+
> enabling delivery mutations.
|
|
13
|
+
|
|
14
|
+
## Routing overview
|
|
15
|
+
|
|
16
|
+
### Intake
|
|
17
|
+
|
|
18
|
+
Ralphie processes open issues that carry the label `labels.ready-for-agent`
|
|
19
|
+
maps to (default `ready-for-agent`) and every label listed in
|
|
20
|
+
`intake.requireLabels`. Issues without them are never read.
|
|
21
|
+
|
|
22
|
+
### AFK triage
|
|
23
|
+
|
|
24
|
+
Off unless `triage.enabled` is `true`. Before the queue starts, Ralphie runs one
|
|
25
|
+
read-only `triager` session per issue that is not agent-ready yet, in three
|
|
26
|
+
buckets only: issues with no triage state label, issues labelled
|
|
27
|
+
`labels.needs-triage`, and `labels.needs-info` issues whose reporter commented
|
|
28
|
+
after the last `## Triage Notes` comment. Issues with several state labels are
|
|
29
|
+
left to a human, and `intake.requireLabels` still narrows the set.
|
|
30
|
+
|
|
31
|
+
The session runs the vendored `/triage` skill under an overlay that skips the
|
|
32
|
+
maintainer steps and grilling. It returns one outcome, and Ralphie applies it:
|
|
33
|
+
|
|
34
|
+
- `promote`: Ralphie posts the Agent Brief (after the AI disclaimer), moves
|
|
35
|
+
the issue to `labels.ready-for-agent`, and the issue is implemented later in
|
|
36
|
+
the same run like any agent-ready issue.
|
|
37
|
+
- `needs_info`: handed off to `labels.needs-info` with the Triage Notes
|
|
38
|
+
template (see [Hand-offs](#hand-offs)).
|
|
39
|
+
- `ready_for_human`: handed off to `labels.ready-for-human`.
|
|
40
|
+
- `already_implemented`: a fresh resolution verifier must prove it. Only then
|
|
41
|
+
does Ralphie comment where the behavior lives and close the issue as
|
|
42
|
+
completed. If the verifier disagrees or fails, the issue is handed off to
|
|
43
|
+
`labels.ready-for-human` instead.
|
|
44
|
+
|
|
45
|
+
Triage never rejects a request: it cannot apply `wontfix` and never writes
|
|
46
|
+
`.out-of-scope/`. A triage session that fails marks only that issue failed.
|
|
47
|
+
|
|
48
|
+
### Pre-flight
|
|
49
|
+
|
|
50
|
+
Each issue gets exactly one read-only, schema-validated `preflight` session.
|
|
51
|
+
It returns one disposition:
|
|
52
|
+
|
|
53
|
+
- `actionable` with `fitsOneSession`: `true` routes to implementation, `false`
|
|
54
|
+
routes to decomposition.
|
|
55
|
+
- `already_resolved`: tentative; see below.
|
|
56
|
+
- `blocked` with the numbers of the open issues it waits on: the issue is
|
|
57
|
+
skipped for this run with no label or other GitHub change, and is picked up
|
|
58
|
+
again on a later run once the blockers close.
|
|
59
|
+
- `hand_off` with a reason, evidence and questions: the issue is handed
|
|
60
|
+
off (see [Hand-offs](#hand-offs)) while Ralphie continues with the next
|
|
61
|
+
queue item.
|
|
62
|
+
|
|
63
|
+
The prompt pins the exact checked-out commit so evidence is never mistaken for
|
|
64
|
+
a newer revision. An actionable result is retained as an artifact, so a
|
|
65
|
+
restart does not repeat the session.
|
|
66
|
+
|
|
67
|
+
```mermaid
|
|
68
|
+
flowchart TD
|
|
69
|
+
A[Open GitHub issue] --> Z[Pre-flight session]
|
|
70
|
+
Z -->|Hand-off| Y[Swap triage label, comment, continue queue]
|
|
71
|
+
Z -->|Blocked by open issue| W[Skip, no label change, continue queue]
|
|
72
|
+
Z -->|Actionable or apparently resolved| B{fitsOneSession}
|
|
73
|
+
B -->|true| C[Implementation session]
|
|
74
|
+
C --> D[Deterministically stage changes]
|
|
75
|
+
D -->|Changes present| V[Configured verification]
|
|
76
|
+
V -->|Passed| E[Candidate commit, standards and spec reviews]
|
|
77
|
+
V -->|Command failed| R[Resume implementer with /diagnosing-bugs]
|
|
78
|
+
R --> D
|
|
79
|
+
D -->|No changes| N[Fresh structured resolution verification]
|
|
80
|
+
N -->|Resolved with evidence| O[Close issue as completed]
|
|
81
|
+
N -->|Unresolved or uncertain| P[Fail and leave issue open]
|
|
82
|
+
E -->|Approved| F[Squash candidates and reverify]
|
|
83
|
+
E -->|Changes requested| G[Resume implementer with /implement]
|
|
84
|
+
G --> D
|
|
85
|
+
E -->|Review rounds exhausted| H[Preserve diagnostics and restore checkout]
|
|
86
|
+
B -->|false| I[Structured decomposition]
|
|
87
|
+
H --> I
|
|
88
|
+
I --> J[Create and cross-link child issues]
|
|
89
|
+
J --> K[Leave original issue untouched and open]
|
|
90
|
+
K --> L[Refresh issue queue]
|
|
91
|
+
F --> M[Commit and non-force push]
|
|
92
|
+
M --> O
|
|
93
|
+
```
|
|
94
|
+
|
|
95
|
+
## Implementation workflow
|
|
96
|
+
|
|
97
|
+
1. Capture the exact clean branch and commit as an issue checkpoint.
|
|
98
|
+
2. Ask a fresh `implementer` session to run the vendored `/implement` skill
|
|
99
|
+
(which drives `/tdd`) with an overlay: do not commit, leave changes in the
|
|
100
|
+
working tree, skip the closing code review. The session must return
|
|
101
|
+
`{status: done | needs_attention, summary, commitMessage, needsAttention?}`;
|
|
102
|
+
prose or premature model termination is not completion. The literal
|
|
103
|
+
`needs_attention` status is a hand-off request, which Ralphie routes through
|
|
104
|
+
hand-off verification rather than treating it as a failure. The latest issue comment that starts
|
|
105
|
+
with `## Agent Brief` is the contract: it is included in full and is exempt
|
|
106
|
+
from comment trimming, with the body and other comments as background.
|
|
107
|
+
Without a brief the issue body is the contract.
|
|
108
|
+
3. Stage every change deterministically and capture the exact staged diff.
|
|
109
|
+
4. Run the configured deterministic verification commands, when any. If a
|
|
110
|
+
command exits non-zero, resume the implementer's session with
|
|
111
|
+
`/diagnosing-bugs` and the bounded output, then restage and retry up to
|
|
112
|
+
`limits.verificationFixes` times (default five). See
|
|
113
|
+
[Fix sessions](#fix-sessions). If the commands still fail after the last
|
|
114
|
+
repair, the issue is not retried: it ends in a `ready-for-human` hand-off
|
|
115
|
+
with reason `implementation_exhausted` (see [Hand-offs](#hand-offs)).
|
|
116
|
+
5. After verification passes or is skipped (no `verify` commands configured),
|
|
117
|
+
create a local candidate commit on top of the checkpoint. Candidate commits
|
|
118
|
+
are never pushed.
|
|
119
|
+
6. Run the review gate: a standards reviewer and a spec reviewer start in
|
|
120
|
+
parallel as read-only sessions on the checkpoint-to-candidate range, using
|
|
121
|
+
the vendored `/code-review` axes. Reviewers cannot run Git (see
|
|
122
|
+
[Safety](safety.md#agent-and-mutation-boundaries)), so Ralphie puts the
|
|
123
|
+
fixed point, commit list and range diff in each prompt. A range diff over
|
|
124
|
+
100,000 characters is truncated in the prompt, so Ralphie also writes the
|
|
125
|
+
full commit list and diff to `.ralphie-review/candidate-diff.txt` in the
|
|
126
|
+
checkout (listed in `.git/info/exclude`, so it can never be staged), names
|
|
127
|
+
that path in both prompts and tells the reviewers to read it whole with their
|
|
128
|
+
Read tool before judging; the file is removed when the reviews end. The
|
|
129
|
+
standards reviewer gets the repository's standards sources (`AGENTS.md`, `CLAUDE.md`,
|
|
130
|
+
`CONTRIBUTING.md`, `CODING_STANDARDS.md`, `GLOSSARY.md`, `docs/adr`,
|
|
131
|
+
`docs/agents`) and the smell baseline; the spec reviewer gets the Agent
|
|
132
|
+
Brief, or the issue body when there is none. Each returns structured
|
|
133
|
+
findings and Ralphie computes the verdict. A hard documented-standard
|
|
134
|
+
violation and any missing, partial, wrong or scope-creep spec finding
|
|
135
|
+
block. Smells never block.
|
|
136
|
+
7. If the gate blocks, resume the implementer's session through `/implement`
|
|
137
|
+
with the findings, restage, reverify and add another candidate commit, then
|
|
138
|
+
review again.
|
|
139
|
+
8. Stop after approval or `limits.reviewRounds` review attempts (default five).
|
|
140
|
+
On approval, squash all candidate commits (soft reset to the checkpoint)
|
|
141
|
+
and reverify the final tree. If repair changes the approved tree, review the
|
|
142
|
+
repaired tree again.
|
|
143
|
+
9. Create the single issue commit with the implementer's commit message,
|
|
144
|
+
validated when the result is received (subject non-empty and at most 72
|
|
145
|
+
characters, optional body). No separate commit-message session runs. Exactly
|
|
146
|
+
one created commit is delivered per issue.
|
|
147
|
+
10. Recheck the remote and push the commit without force, then close the
|
|
148
|
+
source issue after the push is verified.
|
|
149
|
+
|
|
150
|
+
When implementation produces no changes, a fresh read-only session must prove
|
|
151
|
+
that the current checkout already resolves the issue and return concrete
|
|
152
|
+
evidence. A proven resolution is completed and closed. An unresolved result is
|
|
153
|
+
fed back to a fresh implementation session for up to
|
|
154
|
+
`limits.implementationAttempts` attempts. Every other terminal failure of
|
|
155
|
+
an implementation (an exhausted retry budget, a review loop that repeats the
|
|
156
|
+
same blocking findings, a failed session, a timed-out fixer session, a review
|
|
157
|
+
fix that changes nothing) ends in an `implementation_exhausted` hand-off, so a failing
|
|
158
|
+
issue never returns to the queue unchanged. A timed-out implementer session counts as a
|
|
159
|
+
failed implementation attempt: a fresh session retries, told that its
|
|
160
|
+
predecessor timed out and may have left edits, while
|
|
161
|
+
`limits.implementationAttempts` remain, and the last timeout hands off
|
|
162
|
+
`ready-for-human`. A fixer timeout hands off directly without a retry. A run
|
|
163
|
+
interrupted by a stop request is not a failure and hands nothing off. If the review budget is
|
|
164
|
+
exhausted, Ralphie preserves the patch and review diagnostics, restores the
|
|
165
|
+
clean checkpoint, and sends the issue through decomposition.
|
|
166
|
+
|
|
167
|
+
Pre-flight's `already_resolved` disposition is tentative. A fresh verifier must
|
|
168
|
+
confirm it before completion; an `unresolved` result corrects the route to
|
|
169
|
+
actionable, proceeds to implementation, and supplies its summary
|
|
170
|
+
and evidence to the first implementation session.
|
|
171
|
+
|
|
172
|
+
Agent session failures (a harness error, timeout, or result that never
|
|
173
|
+
validates) in pre-flight, the resolution verifier, hand-off verification or
|
|
174
|
+
decomposition are handled like implementation failures: the issue becomes a
|
|
175
|
+
`ready-for-human` hand-off (reason `needs_human_judgment`) with preserved
|
|
176
|
+
diagnostics, so it does not re-enter the queue unchanged. Checkout, repository
|
|
177
|
+
invariant and GitHub errors are not session failures; they fail the issue and
|
|
178
|
+
the next run retries it, as does any failure after a stop request. A hand-off
|
|
179
|
+
whose verifier session failed stays pending and is resumed by the next run.
|
|
180
|
+
|
|
181
|
+
A session failure that is environmental is not a hand-off either. A rate,
|
|
182
|
+
usage, session or quota limit (for example Claude's "You've hit your session
|
|
183
|
+
limit"), an overloaded or unavailable provider (429, 5xx), a network error,
|
|
184
|
+
exhausted credits, or an expired login says nothing about the issue, and every
|
|
185
|
+
further session would fail the same way. Ralphie restores the checkout,
|
|
186
|
+
leaves the issue's labels and comments untouched, records a `deferred`
|
|
187
|
+
outcome and stops the rest of the queue (exit status in
|
|
188
|
+
[Operations and recovery](operations-and-recovery.md#failure-cancellation-and-exit-status)),
|
|
189
|
+
with a message naming the failure and, when the harness reports it, the reset
|
|
190
|
+
time. Rerun once the limit
|
|
191
|
+
clears or the login is renewed. Definite failures (an invalid or never
|
|
192
|
+
validating result, an unknown model, refused access, a session timeout, real
|
|
193
|
+
work that did not succeed) still hand off.
|
|
194
|
+
|
|
195
|
+
### Fix sessions
|
|
196
|
+
|
|
197
|
+
Fixes continue the implementer's own session, so the agent keeps the context it
|
|
198
|
+
built. The `fixer` role resumes it on the same harness. A fresh `fixer` session
|
|
199
|
+
starts from the issue, the current diff and the findings instead when:
|
|
200
|
+
|
|
201
|
+
- resuming fails (the harness error is reported as an info event, not a failed
|
|
202
|
+
stage);
|
|
203
|
+
- the implementer reported no session id;
|
|
204
|
+
- the `fixer` role runs on a different harness than the implementer;
|
|
205
|
+
- the session is near its context limit. Harnesses do not report context use,
|
|
206
|
+
so Ralphie estimates it from the prompt and reply sizes and switches at about
|
|
207
|
+
400,000 characters.
|
|
208
|
+
|
|
209
|
+
The session that did the last fix is the one the next fix continues. The
|
|
210
|
+
`limits.verificationFixes` and `limits.reviewRounds` budgets count fixes the
|
|
211
|
+
same way in both modes. Exhausting `limits.verificationFixes` ends in the
|
|
212
|
+
`implementation_exhausted` hand-off; exhausting `limits.reviewRounds` sends the
|
|
213
|
+
issue through decomposition.
|
|
214
|
+
|
|
215
|
+
The boundary between agent work and deterministic operations stays explicit
|
|
216
|
+
throughout the loop:
|
|
217
|
+
|
|
218
|
+
```mermaid
|
|
219
|
+
sequenceDiagram
|
|
220
|
+
participant R as Ralphie
|
|
221
|
+
participant GH as GitHub
|
|
222
|
+
participant G as Git
|
|
223
|
+
participant P as Harness
|
|
224
|
+
|
|
225
|
+
R->>G: Capture clean branch checkpoint
|
|
226
|
+
R->>G: Verify destination and remote base
|
|
227
|
+
R->>P: Start fresh implementation session
|
|
228
|
+
P-->>R: Edit the checkout
|
|
229
|
+
R->>G: Stage all changes and read exact diff
|
|
230
|
+
|
|
231
|
+
alt Changes present
|
|
232
|
+
loop Until approved or review rounds exhausted
|
|
233
|
+
R->>G: Run configured verification commands (when any)
|
|
234
|
+
opt Verification command fails and repair budget remains
|
|
235
|
+
R->>P: Resume implementer with /diagnosing-bugs
|
|
236
|
+
P-->>R: Update the checkout
|
|
237
|
+
R->>G: Restage and rerun verification
|
|
238
|
+
end
|
|
239
|
+
R->>G: Create local candidate commit
|
|
240
|
+
R->>P: Start standards and spec review sessions
|
|
241
|
+
P-->>R: Return findings; Ralphie computes the verdict
|
|
242
|
+
opt Changes requested and budget remains
|
|
243
|
+
R->>P: Resume implementer with /implement
|
|
244
|
+
P-->>R: Update the checkout
|
|
245
|
+
R->>G: Restage changes and read exact diff
|
|
246
|
+
end
|
|
247
|
+
end
|
|
248
|
+
R->>G: Reverify the exact approved staged tree
|
|
249
|
+
alt Review approved
|
|
250
|
+
R->>G: Squash candidates, commit with the implementer's message
|
|
251
|
+
R->>G: Revalidate destination, HEAD, and remote base
|
|
252
|
+
R->>G: Push selected branch without force
|
|
253
|
+
G->>GH: Send branch update
|
|
254
|
+
GH-->>G: Accept or return authoritative policy rejection
|
|
255
|
+
R->>GH: Close issue as completed
|
|
256
|
+
else Review budget exhausted
|
|
257
|
+
R->>G: Preserve patch and restore checkpoint
|
|
258
|
+
R->>GH: Continue through decomposition
|
|
259
|
+
end
|
|
260
|
+
else No changes
|
|
261
|
+
R->>P: Start fresh structured resolution verification
|
|
262
|
+
P-->>R: Return status and concrete evidence
|
|
263
|
+
opt Resolved
|
|
264
|
+
R->>GH: Close issue as completed
|
|
265
|
+
end
|
|
266
|
+
end
|
|
267
|
+
```
|
|
268
|
+
|
|
269
|
+
## Hand-offs
|
|
270
|
+
|
|
271
|
+
Anything that needs a human becomes a hand-off, and hand-offs are always on.
|
|
272
|
+
Ralphie replaces the issue's triage state label (the issue ends with exactly
|
|
273
|
+
one of the five, named by the `labels` mapping), posts one comment that
|
|
274
|
+
starts with the AI disclaimer `> *This was generated by AI during triage.*`,
|
|
275
|
+
and records a `hand-off` outcome. Because the issue no longer carries the
|
|
276
|
+
agent-ready label, the next run's intake skips it until a human relabels it.
|
|
277
|
+
|
|
278
|
+
| Reason | State label | Comment |
|
|
279
|
+
| --- | --- | --- |
|
|
280
|
+
| `missing_information`, `conflicting_requirements`, `cannot_reproduce`, `outdated_premise` | `needs-info` | Triage Notes: what was established, what is still needed from the reporter. |
|
|
281
|
+
| `implementation_exhausted` (attempts or verification repairs ran out, or the implementation failed terminally) | `ready-for-human` | `## Hand-off` write-up: reason, summary, what was tried, what a human must decide, and the diagnostics location. |
|
|
282
|
+
| `decomposition_limit_reached` (review never converged at the depth limit) | `ready-for-human` | Same write-up. |
|
|
283
|
+
| `external_dependency` reported by a session | `ready-for-human` | Same write-up. |
|
|
284
|
+
|
|
285
|
+
Exhausted attempts preserve a diagnostic patch and restore the clean checkout
|
|
286
|
+
before the hand-off, like the other recovery paths. An open-blocker skip (from
|
|
287
|
+
pre-flight or from queue order) is not a hand-off and changes nothing on
|
|
288
|
+
GitHub, and neither is a deferral caused by a limit, outage or expired login. The write-up deliberately avoids the `## Agent Brief` heading, which
|
|
289
|
+
marks the contract an implementer works from. The recovery details are in
|
|
290
|
+
[Operations and recovery](operations-and-recovery.md#hand-off-handling).
|
|
291
|
+
|
|
292
|
+
## Decomposition workflow
|
|
293
|
+
|
|
294
|
+
1. Ask a `decomposer` session to run the vendored `/to-tickets` skill with an
|
|
295
|
+
overlay: the quiz is skipped and the breakdown comes back as structured
|
|
296
|
+
output instead of being published.
|
|
297
|
+
2. Create child issues, blockers first, in the to-tickets issue template
|
|
298
|
+
(`## Parent`, `## What to build`, `## Acceptance criteria`, `## Blocked by`)
|
|
299
|
+
behind Ralphie's hidden stable marker, each with the agent-ready label
|
|
300
|
+
(`labels.ready-for-agent`) plus every label of the parent that appears in
|
|
301
|
+
`intake.requireLabels`, so a run scoped by those labels picks the children up.
|
|
302
|
+
3. Attach each created or recovered child to the original issue as a **native
|
|
303
|
+
GitHub sub-issue**, reconciling against GitHub's reported hierarchy.
|
|
304
|
+
4. Represent each declared `dependsOn` edge as a **native GitHub
|
|
305
|
+
`blocked_by` dependency** and persist the dependency mapping artifact.
|
|
306
|
+
5. Leave the original issue **open and its body untouched**. It is the
|
|
307
|
+
tracking parent; it is never closed as a duplicate merely because it was
|
|
308
|
+
decomposed.
|
|
309
|
+
|
|
310
|
+
Stable markers and persisted child mappings make the workflow retry-safe: a
|
|
311
|
+
retry discovers previously created children instead of duplicating them,
|
|
312
|
+
and native relationships are reconciled idempotently. Eligible children can
|
|
313
|
+
enter the main implementation loop during the same run; the decomposed parent
|
|
314
|
+
stays out of the queue because it is a tracking issue, not executable work.
|
|
315
|
+
|
|
316
|
+
```mermaid
|
|
317
|
+
flowchart LR
|
|
318
|
+
A[Original issue] --> B[Structured task breakdown]
|
|
319
|
+
B --> C{Existing child marker?}
|
|
320
|
+
C -->|Yes| D[Reuse child issue]
|
|
321
|
+
C -->|No| E[Create child issue]
|
|
322
|
+
D --> F[Reconcile native sub-issues]
|
|
323
|
+
E --> F
|
|
324
|
+
F --> G[Create native blocked_by dependencies]
|
|
325
|
+
G --> H[Keep parent open and untouched]
|
|
326
|
+
H --> I[Refresh open-issue queue]
|
|
327
|
+
```
|
|
328
|
+
|
|
329
|
+
The decomposition session is read-only and returns an
|
|
330
|
+
`issueBreakdownDecisionSchema` result containing at least two children, each
|
|
331
|
+
with a stable key, a title, what to build, acceptance criteria, and its
|
|
332
|
+
blockers. Each child must fit one session, and the blocking graph must be
|
|
333
|
+
acyclic. The breakdown is persisted before the first GitHub mutation.
|
|
334
|
+
|
|
335
|
+
Each child receives a stable marker containing root, parent, key, and depth.
|
|
336
|
+
The positive `limits.maxDecompositionDepth` setting (default `3`) bounds recursive
|
|
337
|
+
splitting and is persisted in run state. If direct pre-flight routing or review
|
|
338
|
+
exhaustion would exceed it, Ralphie does not attempt another breakdown: it
|
|
339
|
+
hands the issue off as `ready-for-human` with reason
|
|
340
|
+
`decomposition_limit_reached` and continues independent queued work. The
|
|
341
|
+
issue is not marked complete, so its dependents remain blocked.
|
|
342
|
+
Ralphie discovers those markers and reconciles them with any persisted mapping
|
|
343
|
+
before creating anything. Thus a lost create response or a partial linking
|
|
344
|
+
failure does not blindly duplicate children. Creation,
|
|
345
|
+
number recording, native sub-issue attachment, and dependency creation
|
|
346
|
+
are separate mutations; a child already
|
|
347
|
+
attached to the wrong parent, or a native relationship that disagrees with a
|
|
348
|
+
child's marker, halts with a recovery diagnostic instead of silently
|
|
349
|
+
reparenting or duplicating issues.
|
|
350
|
+
|
|
351
|
+
The decomposed parent remains open as the native tracking issue and exposes
|
|
352
|
+
GitHub's completion progress for its sub-issues. Ralphie never edits its body;
|
|
353
|
+
it recognises the parent by its native sub-issues (and by open children whose
|
|
354
|
+
marker names it), so it is not queued again. It is closed as `completed`,
|
|
355
|
+
with one explanatory comment, only when its child work is finished: completing the
|
|
356
|
+
final child reconciles its parent immediately, and every run also
|
|
357
|
+
reconciles decomposed parents it discovers or refreshes, so a parent whose
|
|
358
|
+
final child closed in a previous run is completed on a later run. The open-issue
|
|
359
|
+
queue is refreshed after decomposition; newly eligible children can run during
|
|
360
|
+
the same invocation. If dependencies remain open after the queue is exhausted,
|
|
361
|
+
Ralphie records each blocked issue as a skipped outcome naming its open
|
|
362
|
+
blockers and leaves it pending instead of handing it to an agent, then drains
|
|
363
|
+
later work and completes the run. Blocked issues remain open and keep their
|
|
364
|
+
labels.
|
|
365
|
+
|
|
366
|
+
A pre-flight `fitsOneSession: false` route returns `decomposed`. Review exhaustion returns an
|
|
367
|
+
`escalated` outcome containing the recovery diagnostic path and, after
|
|
368
|
+
successful decomposition, the created child numbers. Both transitions refresh
|
|
369
|
+
the queue.
|
|
370
|
+
|
|
371
|
+
### Platform support for native sub-issues and dependencies
|
|
372
|
+
|
|
373
|
+
Native sub-issues and `blocked_by` dependencies are GitHub REST features required
|
|
374
|
+
for decomposition. Ralphie's current GitHub client targets `github.com` only;
|
|
375
|
+
GitHub Enterprise Server is not supported. There is **no body-link fallback**:
|
|
376
|
+
Ralphie never silently degrades to body-only hierarchy semantics.
|
|
377
|
+
|
|
378
|
+
- Creating, recovering, or linking children fails with an actionable error
|
|
379
|
+
naming the missing platform capability when an endpoint is unavailable or the
|
|
380
|
+
token lacks issue write permission.
|
|
381
|
+
- The compatibility check is per live operation: the first relationship read or
|
|
382
|
+
write against an unsupported endpoint surfaces the error.
|
|
383
|
+
- Recovery metadata (stable markers and the persisted key/dependency mappings)
|
|
384
|
+
remains the idempotency record, so a run can continue after a recoverable
|
|
385
|
+
relationship failure without blindly duplicating children.
|
|
386
|
+
- To verify the required `github.com` endpoints before a live run:
|
|
387
|
+
`gh api repos/{owner}/{repo}/issues/1/sub_issues` and
|
|
388
|
+
`gh api repos/{owner}/{repo}/issues/1/dependencies/blocked_by` should return
|
|
389
|
+
`200` (an empty list) rather than `404`.
|
|
390
|
+
|
|
391
|
+
## Delivery
|
|
392
|
+
|
|
393
|
+
| Issue checkout | Delivery | Source issue closure |
|
|
394
|
+
| --- | --- | --- |
|
|
395
|
+
| Selected base branch | Commit and non-force push directly to that branch; verify remote SHA and clean checkout | Close directly as `completed` after verified delivery. |
|
|
396
|
+
|
|
397
|
+
The direct-push path never uses force. A push rejection is authoritative: the
|
|
398
|
+
created commit and artifacts are retained, the run halts, and a later run
|
|
399
|
+
re-evaluates the still-open issue from a fresh checkout. Inspect the retained
|
|
400
|
+
workspace before the next run removes it.
|
|
401
|
+
|
|
402
|
+
## Queue behavior
|
|
403
|
+
|
|
404
|
+
Issue work is sequential. With the default `created:asc` sort, issues are
|
|
405
|
+
processed oldest-first. When no branch is configured, Ralphie uses `main` when
|
|
406
|
+
it exists and otherwise `master`.
|
|
407
|
+
|
|
408
|
+
For command syntax and all defaults, see the [CLI reference](cli-reference.md).
|
|
409
|
+
For workspace, Git, and GitHub guardrails, see [Safety](safety.md). For
|
|
410
|
+
interruption and failure boundaries, see
|
|
411
|
+
[Operations and recovery](operations-and-recovery.md).
|
package/package.json
CHANGED
|
@@ -1,6 +1,86 @@
|
|
|
1
1
|
{
|
|
2
2
|
"name": "@beremaran/ralphie",
|
|
3
|
-
"version": "0.
|
|
4
|
-
"
|
|
5
|
-
"
|
|
3
|
+
"version": "0.2.1",
|
|
4
|
+
"description": "Turn a GitHub issue queue into reviewed commits with coding-agent harnesses.",
|
|
5
|
+
"module": "./dist/ralphie.js",
|
|
6
|
+
"main": "./dist/ralphie.js",
|
|
7
|
+
"type": "module",
|
|
8
|
+
"bin": {
|
|
9
|
+
"ralphie": "./dist/ralphie.js"
|
|
10
|
+
},
|
|
11
|
+
"exports": {
|
|
12
|
+
".": "./dist/ralphie.js"
|
|
13
|
+
},
|
|
14
|
+
"files": [
|
|
15
|
+
"dist/ralphie.js",
|
|
16
|
+
"README.md",
|
|
17
|
+
"CHANGELOG.md",
|
|
18
|
+
"LICENSE",
|
|
19
|
+
"docs/architecture.md",
|
|
20
|
+
"docs/cli-reference.md",
|
|
21
|
+
"docs/configuration.md",
|
|
22
|
+
"docs/development.md",
|
|
23
|
+
"docs/getting-started.md",
|
|
24
|
+
"docs/operations-and-recovery.md",
|
|
25
|
+
"docs/README.md",
|
|
26
|
+
"docs/safety.md",
|
|
27
|
+
"docs/workflows.md",
|
|
28
|
+
"vendor/mattpocock-skills"
|
|
29
|
+
],
|
|
30
|
+
"repository": {
|
|
31
|
+
"type": "git",
|
|
32
|
+
"url": "https://github.com/beremaran/ralphie"
|
|
33
|
+
},
|
|
34
|
+
"homepage": "https://github.com/beremaran/ralphie",
|
|
35
|
+
"bugs": {
|
|
36
|
+
"url": "https://github.com/beremaran/ralphie/issues"
|
|
37
|
+
},
|
|
38
|
+
"license": "MIT",
|
|
39
|
+
"keywords": [
|
|
40
|
+
"bun",
|
|
41
|
+
"cli",
|
|
42
|
+
"github",
|
|
43
|
+
"github-issues",
|
|
44
|
+
"automation",
|
|
45
|
+
"ai",
|
|
46
|
+
"claude-code",
|
|
47
|
+
"agent",
|
|
48
|
+
"npm"
|
|
49
|
+
],
|
|
50
|
+
"engines": {
|
|
51
|
+
"bun": ">=1.3.0"
|
|
52
|
+
},
|
|
53
|
+
"publishConfig": {
|
|
54
|
+
"access": "public"
|
|
55
|
+
},
|
|
56
|
+
"devDependencies": {
|
|
57
|
+
"@biomejs/biome": "^2.5.10",
|
|
58
|
+
"@types/bun": "latest",
|
|
59
|
+
"@types/proper-lockfile": "^4.1.4",
|
|
60
|
+
"typescript": "^5"
|
|
61
|
+
},
|
|
62
|
+
"dependencies": {
|
|
63
|
+
"@opentui/core": "^0.5.11",
|
|
64
|
+
"octokit": "^5.0.5",
|
|
65
|
+
"proper-lockfile": "^4.1.2",
|
|
66
|
+
"zod": "^4.4.3"
|
|
67
|
+
},
|
|
68
|
+
"scripts": {
|
|
69
|
+
"start": "bun run index.ts",
|
|
70
|
+
"dev": "bun --watch ./index.ts",
|
|
71
|
+
"build": "bun run scripts/build.ts",
|
|
72
|
+
"build:package": "bun run scripts/build.ts",
|
|
73
|
+
"check": "bun run format:check && bun run lint && bun run typecheck && bun run source:audit && bun run test && bun run build",
|
|
74
|
+
"format": "biome format --write .",
|
|
75
|
+
"format:check": "biome format .",
|
|
76
|
+
"lint": "biome lint .",
|
|
77
|
+
"package:check": "bun run scripts/package-smoke.ts",
|
|
78
|
+
"package:inspect": "bun run scripts/package-smoke.ts --dry-run",
|
|
79
|
+
"prepack": "bun run build:package",
|
|
80
|
+
"source:audit": "bun run scripts/source-reachability.ts --json",
|
|
81
|
+
"smoke:live": "bun run scripts/live-smoke.ts",
|
|
82
|
+
"skills:sync": "bun run scripts/skills-sync.ts",
|
|
83
|
+
"test": "bun test tests",
|
|
84
|
+
"typecheck": "bunx tsc --noEmit"
|
|
85
|
+
}
|
|
6
86
|
}
|
|
@@ -0,0 +1,21 @@
|
|
|
1
|
+
MIT License
|
|
2
|
+
|
|
3
|
+
Copyright (c) 2026 Matt Pocock
|
|
4
|
+
|
|
5
|
+
Permission is hereby granted, free of charge, to any person obtaining a copy
|
|
6
|
+
of this software and associated documentation files (the "Software"), to deal
|
|
7
|
+
in the Software without restriction, including without limitation the rights
|
|
8
|
+
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
|
|
9
|
+
copies of the Software, and to permit persons to whom the Software is
|
|
10
|
+
furnished to do so, subject to the following conditions:
|
|
11
|
+
|
|
12
|
+
The above copyright notice and this permission notice shall be included in all
|
|
13
|
+
copies or substantial portions of the Software.
|
|
14
|
+
|
|
15
|
+
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
|
|
16
|
+
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
|
|
17
|
+
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
|
|
18
|
+
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
|
|
19
|
+
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
|
|
20
|
+
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
|
|
21
|
+
SOFTWARE.
|
|
@@ -0,0 +1,87 @@
|
|
|
1
|
+
---
|
|
2
|
+
name: code-review
|
|
3
|
+
description: "Review the changes since a fixed point (commit, branch, tag, or merge-base) along two axes: Standards (does the code follow this repo's documented coding standards?) and Spec (does the code match what the originating issue/spec asked for?). Runs both reviews in parallel sub-agents and reports them side by side. Use when the user wants to review a branch, a PR, work-in-progress changes, or asks to \"review since X\"."
|
|
4
|
+
---
|
|
5
|
+
|
|
6
|
+
Two-axis review of the diff between `HEAD` and a fixed point the user supplies:
|
|
7
|
+
|
|
8
|
+
- **Standards**: does the code conform to this repo's documented coding standards?
|
|
9
|
+
- **Spec**: does the code faithfully implement the originating issue / spec?
|
|
10
|
+
|
|
11
|
+
Both axes run as **parallel sub-agents** so they don't pollute each other's context, then this skill aggregates their findings.
|
|
12
|
+
|
|
13
|
+
The issue tracker should have been provided to you. If `docs/agents/issue-tracker.md` is missing, tell the user to run `/setup-matt-pocock-skills`.
|
|
14
|
+
|
|
15
|
+
## Process
|
|
16
|
+
|
|
17
|
+
### 1. Pin the fixed point
|
|
18
|
+
|
|
19
|
+
Whatever the user said is the fixed point (a commit SHA, branch name, tag, `main`, `HEAD~5`, etc.). If they didn't specify one, ask for it.
|
|
20
|
+
|
|
21
|
+
Capture the diff command once: `git diff <fixed-point>...HEAD` (three-dot, so the comparison is against the merge-base). Also note the list of commits via `git log <fixed-point>..HEAD --oneline`.
|
|
22
|
+
|
|
23
|
+
Before going further, confirm the fixed point resolves (`git rev-parse <fixed-point>`) and the diff is non-empty. A bad ref or empty diff should fail here, not inside two parallel sub-agents.
|
|
24
|
+
|
|
25
|
+
### 2. Identify the spec source
|
|
26
|
+
|
|
27
|
+
Look for the originating spec, in this order:
|
|
28
|
+
|
|
29
|
+
1. Issue references in the commit messages (`#123`, `Closes #45`, GitLab `!67`, etc.), fetched via the workflow in `docs/agents/issue-tracker.md`.
|
|
30
|
+
2. A path the user passed as an argument.
|
|
31
|
+
3. A spec file under `docs/`, `specs/`, or `.scratch/` matching the branch name or feature.
|
|
32
|
+
4. If nothing is found, ask the user where the spec is. If they say there isn't one, the **Spec** sub-agent will skip and report "no spec available".
|
|
33
|
+
|
|
34
|
+
### 3. Identify the standards sources
|
|
35
|
+
|
|
36
|
+
Anything in the repo that documents how code should be written, such as `CODING_STANDARDS.md` or `CONTRIBUTING.md`.
|
|
37
|
+
|
|
38
|
+
On top of whatever the repo documents, the Standards axis always carries the **smell baseline** below: a fixed set of Fowler code smells (_Refactoring_, ch.3) that applies even when a repo documents nothing. Two rules bind it:
|
|
39
|
+
|
|
40
|
+
- **The repo overrides.** A documented repo standard always wins; where it endorses something the baseline would flag, suppress the smell.
|
|
41
|
+
- **Always a judgement call.** Each smell is a labelled heuristic ("possible Feature Envy"), never a hard violation. Like any standard here, skip anything tooling already enforces.
|
|
42
|
+
|
|
43
|
+
Each smell reads *what it is* → *how to fix*; match it against the diff:
|
|
44
|
+
|
|
45
|
+
- **Mysterious Name**: a function, variable, or type whose name doesn't reveal what it does or holds. → rename it; if no honest name comes, the design's murky.
|
|
46
|
+
- **Duplicated Code**: the same logic shape appears in more than one hunk or file in the change. → extract the shared shape, call it from both.
|
|
47
|
+
- **Feature Envy**: a method that reaches into another object's data more than its own. → move the method onto the data it envies.
|
|
48
|
+
- **Data Clumps**: the same few fields or params keep travelling together (a type wanting to be born). → bundle them into one type, pass that.
|
|
49
|
+
- **Primitive Obsession**: a primitive or string standing in for a domain concept that deserves its own type. → give the concept its own small type.
|
|
50
|
+
- **Repeated Switches**: the same `switch`/`if`-cascade on the same type recurs across the change. → replace with polymorphism, or one map both sites share.
|
|
51
|
+
- **Shotgun Surgery**: one logical change forces scattered edits across many files in the diff. → gather what changes together into one module.
|
|
52
|
+
- **Divergent Change**: one file or module is edited for several unrelated reasons. → split so each module changes for one reason.
|
|
53
|
+
- **Speculative Generality**: abstraction, parameters, or hooks added for needs the spec doesn't have. → delete it; inline back until a real need shows.
|
|
54
|
+
- **Message Chains**: long `a.b().c().d()` navigation the caller shouldn't depend on. → hide the walk behind one method on the first object.
|
|
55
|
+
- **Middle Man**: a class or function that mostly just delegates onward. → cut it, call the real target direct.
|
|
56
|
+
- **Refused Bequest**: a subclass or implementer that ignores or overrides most of what it inherits. → drop the inheritance, use composition.
|
|
57
|
+
|
|
58
|
+
### 4. Spawn both sub-agents in parallel
|
|
59
|
+
|
|
60
|
+
**Standards sub-agent prompt** should include:
|
|
61
|
+
|
|
62
|
+
- The full diff command and commit list.
|
|
63
|
+
- The list of standards-source files you found in step 3, **plus the smell baseline from step 3** pasted in full (the sub-agent has no other access to it).
|
|
64
|
+
- The brief: "Report, per file/hunk where relevant, (a) every place the diff violates a documented standard: cite the standard (file + the rule); and (b) any baseline smell you spot: name it and quote the hunk. Distinguish hard violations from judgement calls: documented-standard breaches can be hard, but baseline smells are always judgement calls, and a documented repo standard overrides the baseline. Skip anything tooling enforces. Under 400 words."
|
|
65
|
+
|
|
66
|
+
**Spec sub-agent prompt** should include:
|
|
67
|
+
|
|
68
|
+
- The diff command and commit list.
|
|
69
|
+
- The path or fetched contents of the spec.
|
|
70
|
+
- The brief: "Report: (a) requirements the spec asked for that are missing or partial; (b) behaviour in the diff that wasn't asked for (scope creep); (c) requirements that look implemented but where the implementation looks wrong. Quote the spec line for each finding. Under 400 words."
|
|
71
|
+
|
|
72
|
+
If the spec is missing, skip the Spec sub-agent and note this in the final report.
|
|
73
|
+
|
|
74
|
+
### 5. Aggregate
|
|
75
|
+
|
|
76
|
+
Present the two reports under `## Standards` and `## Spec` headings, verbatim or lightly cleaned. Do **not** merge or rerank findings, because the two axes are deliberately separate (see _Why two axes_).
|
|
77
|
+
|
|
78
|
+
End with a one-line summary: total findings per axis, and the worst issue _within each axis_ (if any). Don't pick a single winner across axes: that's the reranking the separation exists to prevent.
|
|
79
|
+
|
|
80
|
+
## Why two axes
|
|
81
|
+
|
|
82
|
+
A change can pass one axis and fail the other:
|
|
83
|
+
|
|
84
|
+
- Code that follows every standard but implements the wrong thing → **Standards pass, Spec fail.**
|
|
85
|
+
- Code that does exactly what the issue asked but breaks the project's conventions → **Spec pass, Standards fail.**
|
|
86
|
+
|
|
87
|
+
Reporting them separately stops one axis from masking the other.
|
|
@@ -0,0 +1,37 @@
|
|
|
1
|
+
# Deepening
|
|
2
|
+
|
|
3
|
+
How to deepen a cluster of shallow modules safely, given its dependencies. Assumes the vocabulary in [SKILL.md](SKILL.md): **module**, **interface**, **seam**, **adapter**.
|
|
4
|
+
|
|
5
|
+
## Dependency categories
|
|
6
|
+
|
|
7
|
+
When assessing a candidate for deepening, classify its dependencies. The category determines how the deepened module is tested across its seam.
|
|
8
|
+
|
|
9
|
+
### 1. In-process
|
|
10
|
+
|
|
11
|
+
Pure computation, in-memory state, no I/O. Always deepenable: merge the modules and test through the new interface directly. No adapter needed.
|
|
12
|
+
|
|
13
|
+
### 2. Local-substitutable
|
|
14
|
+
|
|
15
|
+
Dependencies that have local test stand-ins (PGLite for Postgres, in-memory filesystem). Deepenable if the stand-in exists. The deepened module is tested with the stand-in running in the test suite. The seam is internal; no port at the module's external interface.
|
|
16
|
+
|
|
17
|
+
### 3. Remote but owned (Ports & Adapters)
|
|
18
|
+
|
|
19
|
+
Your own services across a network boundary (microservices, internal APIs). Define a **port** (interface) at the seam. The deep module owns the logic; the transport is injected as an **adapter**. Tests use an in-memory adapter. Production uses an HTTP/gRPC/queue adapter.
|
|
20
|
+
|
|
21
|
+
Recommendation shape: *"Define a port at the seam, implement an HTTP adapter for production and an in-memory adapter for testing, so the logic sits in one deep module even though it's deployed across a network."*
|
|
22
|
+
|
|
23
|
+
### 4. True external (Mock)
|
|
24
|
+
|
|
25
|
+
Third-party services (Stripe, Twilio, etc.) you don't control. The deepened module takes the external dependency as an injected port; tests provide a mock adapter.
|
|
26
|
+
|
|
27
|
+
## Seam discipline
|
|
28
|
+
|
|
29
|
+
- **One adapter means a hypothetical seam. Two adapters means a real one.** Don't introduce a port unless at least two adapters are justified (typically production + test). A single-adapter seam is just indirection.
|
|
30
|
+
- **Internal seams vs external seams.** A deep module can have internal seams (private to its implementation, used by its own tests) as well as the external seam at its interface. Don't expose internal seams through the interface just because tests use them.
|
|
31
|
+
|
|
32
|
+
## Testing strategy: replace, don't layer
|
|
33
|
+
|
|
34
|
+
- Old unit tests on shallow modules become waste once tests at the deepened module's interface exist; delete them.
|
|
35
|
+
- Write new tests at the deepened module's interface. The **interface is the test surface**.
|
|
36
|
+
- Tests assert on observable outcomes through the interface, not internal state.
|
|
37
|
+
- Tests should survive internal refactors, since they describe behaviour, not implementation. If a test has to change when the implementation changes, it's testing past the interface.
|