@beremaran/ralphie 0.0.0-stage → 0.2.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (37) hide show
  1. package/CHANGELOG.md +832 -0
  2. package/LICENSE +21 -0
  3. package/README.md +53 -2
  4. package/dist/ralphie.js +33691 -0
  5. package/docs/README.md +71 -0
  6. package/docs/architecture.md +159 -0
  7. package/docs/cli-reference.md +126 -0
  8. package/docs/configuration.md +269 -0
  9. package/docs/development.md +248 -0
  10. package/docs/getting-started.md +134 -0
  11. package/docs/operations-and-recovery.md +328 -0
  12. package/docs/safety.md +189 -0
  13. package/docs/workflows.md +411 -0
  14. package/package.json +83 -3
  15. package/vendor/mattpocock-skills/LICENSE +21 -0
  16. package/vendor/mattpocock-skills/code-review/SKILL.md +87 -0
  17. package/vendor/mattpocock-skills/code-review/agents/openai.yaml +3 -0
  18. package/vendor/mattpocock-skills/codebase-design/DEEPENING.md +37 -0
  19. package/vendor/mattpocock-skills/codebase-design/DESIGN-IT-TWICE.md +44 -0
  20. package/vendor/mattpocock-skills/codebase-design/SKILL.md +114 -0
  21. package/vendor/mattpocock-skills/codebase-design/agents/openai.yaml +3 -0
  22. package/vendor/mattpocock-skills/diagnosing-bugs/SKILL.md +138 -0
  23. package/vendor/mattpocock-skills/diagnosing-bugs/agents/openai.yaml +3 -0
  24. package/vendor/mattpocock-skills/diagnosing-bugs/scripts/hitl-loop.template.sh +44 -0
  25. package/vendor/mattpocock-skills/implement/SKILL.md +15 -0
  26. package/vendor/mattpocock-skills/implement/agents/openai.yaml +5 -0
  27. package/vendor/mattpocock-skills/lock.json +37 -0
  28. package/vendor/mattpocock-skills/tdd/SKILL.md +38 -0
  29. package/vendor/mattpocock-skills/tdd/agents/openai.yaml +3 -0
  30. package/vendor/mattpocock-skills/tdd/mocking.md +59 -0
  31. package/vendor/mattpocock-skills/tdd/tests.md +77 -0
  32. package/vendor/mattpocock-skills/to-tickets/SKILL.md +105 -0
  33. package/vendor/mattpocock-skills/to-tickets/agents/openai.yaml +5 -0
  34. package/vendor/mattpocock-skills/triage/AGENT-BRIEF.md +207 -0
  35. package/vendor/mattpocock-skills/triage/OUT-OF-SCOPE.md +105 -0
  36. package/vendor/mattpocock-skills/triage/SKILL.md +112 -0
  37. package/vendor/mattpocock-skills/triage/agents/openai.yaml +5 -0
@@ -0,0 +1,411 @@
1
+ # Workflows
2
+
3
+ This page is for operators and contributors who need to understand how Ralphie
4
+ routes issues, performs implementation and decomposition, and delivers the
5
+ result. It is the authoritative description of workflow semantics and diagrams;
6
+ see the [documentation index](README.md) for setup, CLI, safety, and recovery
7
+ references.
8
+
9
+ > [!CAUTION]
10
+ > Ralphie commits and pushes directly to the selected branch. Read the
11
+ > [safety model](safety.md) and validate against a repository you control before
12
+ > enabling delivery mutations.
13
+
14
+ ## Routing overview
15
+
16
+ ### Intake
17
+
18
+ Ralphie processes open issues that carry the label `labels.ready-for-agent`
19
+ maps to (default `ready-for-agent`) and every label listed in
20
+ `intake.requireLabels`. Issues without them are never read.
21
+
22
+ ### AFK triage
23
+
24
+ Off unless `triage.enabled` is `true`. Before the queue starts, Ralphie runs one
25
+ read-only `triager` session per issue that is not agent-ready yet, in three
26
+ buckets only: issues with no triage state label, issues labelled
27
+ `labels.needs-triage`, and `labels.needs-info` issues whose reporter commented
28
+ after the last `## Triage Notes` comment. Issues with several state labels are
29
+ left to a human, and `intake.requireLabels` still narrows the set.
30
+
31
+ The session runs the vendored `/triage` skill under an overlay that skips the
32
+ maintainer steps and grilling. It returns one outcome, and Ralphie applies it:
33
+
34
+ - `promote`: Ralphie posts the Agent Brief (after the AI disclaimer), moves
35
+ the issue to `labels.ready-for-agent`, and the issue is implemented later in
36
+ the same run like any agent-ready issue.
37
+ - `needs_info`: handed off to `labels.needs-info` with the Triage Notes
38
+ template (see [Hand-offs](#hand-offs)).
39
+ - `ready_for_human`: handed off to `labels.ready-for-human`.
40
+ - `already_implemented`: a fresh resolution verifier must prove it. Only then
41
+ does Ralphie comment where the behavior lives and close the issue as
42
+ completed. If the verifier disagrees or fails, the issue is handed off to
43
+ `labels.ready-for-human` instead.
44
+
45
+ Triage never rejects a request: it cannot apply `wontfix` and never writes
46
+ `.out-of-scope/`. A triage session that fails marks only that issue failed.
47
+
48
+ ### Pre-flight
49
+
50
+ Each issue gets exactly one read-only, schema-validated `preflight` session.
51
+ It returns one disposition:
52
+
53
+ - `actionable` with `fitsOneSession`: `true` routes to implementation, `false`
54
+ routes to decomposition.
55
+ - `already_resolved`: tentative; see below.
56
+ - `blocked` with the numbers of the open issues it waits on: the issue is
57
+ skipped for this run with no label or other GitHub change, and is picked up
58
+ again on a later run once the blockers close.
59
+ - `hand_off` with a reason, evidence and questions: the issue is handed
60
+ off (see [Hand-offs](#hand-offs)) while Ralphie continues with the next
61
+ queue item.
62
+
63
+ The prompt pins the exact checked-out commit so evidence is never mistaken for
64
+ a newer revision. An actionable result is retained as an artifact, so a
65
+ restart does not repeat the session.
66
+
67
+ ```mermaid
68
+ flowchart TD
69
+ A[Open GitHub issue] --> Z[Pre-flight session]
70
+ Z -->|Hand-off| Y[Swap triage label, comment, continue queue]
71
+ Z -->|Blocked by open issue| W[Skip, no label change, continue queue]
72
+ Z -->|Actionable or apparently resolved| B{fitsOneSession}
73
+ B -->|true| C[Implementation session]
74
+ C --> D[Deterministically stage changes]
75
+ D -->|Changes present| V[Configured verification]
76
+ V -->|Passed| E[Candidate commit, standards and spec reviews]
77
+ V -->|Command failed| R[Resume implementer with /diagnosing-bugs]
78
+ R --> D
79
+ D -->|No changes| N[Fresh structured resolution verification]
80
+ N -->|Resolved with evidence| O[Close issue as completed]
81
+ N -->|Unresolved or uncertain| P[Fail and leave issue open]
82
+ E -->|Approved| F[Squash candidates and reverify]
83
+ E -->|Changes requested| G[Resume implementer with /implement]
84
+ G --> D
85
+ E -->|Review rounds exhausted| H[Preserve diagnostics and restore checkout]
86
+ B -->|false| I[Structured decomposition]
87
+ H --> I
88
+ I --> J[Create and cross-link child issues]
89
+ J --> K[Leave original issue untouched and open]
90
+ K --> L[Refresh issue queue]
91
+ F --> M[Commit and non-force push]
92
+ M --> O
93
+ ```
94
+
95
+ ## Implementation workflow
96
+
97
+ 1. Capture the exact clean branch and commit as an issue checkpoint.
98
+ 2. Ask a fresh `implementer` session to run the vendored `/implement` skill
99
+ (which drives `/tdd`) with an overlay: do not commit, leave changes in the
100
+ working tree, skip the closing code review. The session must return
101
+ `{status: done | needs_attention, summary, commitMessage, needsAttention?}`;
102
+ prose or premature model termination is not completion. The literal
103
+ `needs_attention` status is a hand-off request, which Ralphie routes through
104
+ hand-off verification rather than treating it as a failure. The latest issue comment that starts
105
+ with `## Agent Brief` is the contract: it is included in full and is exempt
106
+ from comment trimming, with the body and other comments as background.
107
+ Without a brief the issue body is the contract.
108
+ 3. Stage every change deterministically and capture the exact staged diff.
109
+ 4. Run the configured deterministic verification commands, when any. If a
110
+ command exits non-zero, resume the implementer's session with
111
+ `/diagnosing-bugs` and the bounded output, then restage and retry up to
112
+ `limits.verificationFixes` times (default five). See
113
+ [Fix sessions](#fix-sessions). If the commands still fail after the last
114
+ repair, the issue is not retried: it ends in a `ready-for-human` hand-off
115
+ with reason `implementation_exhausted` (see [Hand-offs](#hand-offs)).
116
+ 5. After verification passes or is skipped (no `verify` commands configured),
117
+ create a local candidate commit on top of the checkpoint. Candidate commits
118
+ are never pushed.
119
+ 6. Run the review gate: a standards reviewer and a spec reviewer start in
120
+ parallel as read-only sessions on the checkpoint-to-candidate range, using
121
+ the vendored `/code-review` axes. Reviewers cannot run Git (see
122
+ [Safety](safety.md#agent-and-mutation-boundaries)), so Ralphie puts the
123
+ fixed point, commit list and range diff in each prompt. A range diff over
124
+ 100,000 characters is truncated in the prompt, so Ralphie also writes the
125
+ full commit list and diff to `.ralphie-review/candidate-diff.txt` in the
126
+ checkout (listed in `.git/info/exclude`, so it can never be staged), names
127
+ that path in both prompts and tells the reviewers to read it whole with their
128
+ Read tool before judging; the file is removed when the reviews end. The
129
+ standards reviewer gets the repository's standards sources (`AGENTS.md`, `CLAUDE.md`,
130
+ `CONTRIBUTING.md`, `CODING_STANDARDS.md`, `GLOSSARY.md`, `docs/adr`,
131
+ `docs/agents`) and the smell baseline; the spec reviewer gets the Agent
132
+ Brief, or the issue body when there is none. Each returns structured
133
+ findings and Ralphie computes the verdict. A hard documented-standard
134
+ violation and any missing, partial, wrong or scope-creep spec finding
135
+ block. Smells never block.
136
+ 7. If the gate blocks, resume the implementer's session through `/implement`
137
+ with the findings, restage, reverify and add another candidate commit, then
138
+ review again.
139
+ 8. Stop after approval or `limits.reviewRounds` review attempts (default five).
140
+ On approval, squash all candidate commits (soft reset to the checkpoint)
141
+ and reverify the final tree. If repair changes the approved tree, review the
142
+ repaired tree again.
143
+ 9. Create the single issue commit with the implementer's commit message,
144
+ validated when the result is received (subject non-empty and at most 72
145
+ characters, optional body). No separate commit-message session runs. Exactly
146
+ one created commit is delivered per issue.
147
+ 10. Recheck the remote and push the commit without force, then close the
148
+ source issue after the push is verified.
149
+
150
+ When implementation produces no changes, a fresh read-only session must prove
151
+ that the current checkout already resolves the issue and return concrete
152
+ evidence. A proven resolution is completed and closed. An unresolved result is
153
+ fed back to a fresh implementation session for up to
154
+ `limits.implementationAttempts` attempts. Every other terminal failure of
155
+ an implementation (an exhausted retry budget, a review loop that repeats the
156
+ same blocking findings, a failed session, a timed-out fixer session, a review
157
+ fix that changes nothing) ends in an `implementation_exhausted` hand-off, so a failing
158
+ issue never returns to the queue unchanged. A timed-out implementer session counts as a
159
+ failed implementation attempt: a fresh session retries, told that its
160
+ predecessor timed out and may have left edits, while
161
+ `limits.implementationAttempts` remain, and the last timeout hands off
162
+ `ready-for-human`. A fixer timeout hands off directly without a retry. A run
163
+ interrupted by a stop request is not a failure and hands nothing off. If the review budget is
164
+ exhausted, Ralphie preserves the patch and review diagnostics, restores the
165
+ clean checkpoint, and sends the issue through decomposition.
166
+
167
+ Pre-flight's `already_resolved` disposition is tentative. A fresh verifier must
168
+ confirm it before completion; an `unresolved` result corrects the route to
169
+ actionable, proceeds to implementation, and supplies its summary
170
+ and evidence to the first implementation session.
171
+
172
+ Agent session failures (a harness error, timeout, or result that never
173
+ validates) in pre-flight, the resolution verifier, hand-off verification or
174
+ decomposition are handled like implementation failures: the issue becomes a
175
+ `ready-for-human` hand-off (reason `needs_human_judgment`) with preserved
176
+ diagnostics, so it does not re-enter the queue unchanged. Checkout, repository
177
+ invariant and GitHub errors are not session failures; they fail the issue and
178
+ the next run retries it, as does any failure after a stop request. A hand-off
179
+ whose verifier session failed stays pending and is resumed by the next run.
180
+
181
+ A session failure that is environmental is not a hand-off either. A rate,
182
+ usage, session or quota limit (for example Claude's "You've hit your session
183
+ limit"), an overloaded or unavailable provider (429, 5xx), a network error,
184
+ exhausted credits, or an expired login says nothing about the issue, and every
185
+ further session would fail the same way. Ralphie restores the checkout,
186
+ leaves the issue's labels and comments untouched, records a `deferred`
187
+ outcome and stops the rest of the queue (exit status in
188
+ [Operations and recovery](operations-and-recovery.md#failure-cancellation-and-exit-status)),
189
+ with a message naming the failure and, when the harness reports it, the reset
190
+ time. Rerun once the limit
191
+ clears or the login is renewed. Definite failures (an invalid or never
192
+ validating result, an unknown model, refused access, a session timeout, real
193
+ work that did not succeed) still hand off.
194
+
195
+ ### Fix sessions
196
+
197
+ Fixes continue the implementer's own session, so the agent keeps the context it
198
+ built. The `fixer` role resumes it on the same harness. A fresh `fixer` session
199
+ starts from the issue, the current diff and the findings instead when:
200
+
201
+ - resuming fails (the harness error is reported as an info event, not a failed
202
+ stage);
203
+ - the implementer reported no session id;
204
+ - the `fixer` role runs on a different harness than the implementer;
205
+ - the session is near its context limit. Harnesses do not report context use,
206
+ so Ralphie estimates it from the prompt and reply sizes and switches at about
207
+ 400,000 characters.
208
+
209
+ The session that did the last fix is the one the next fix continues. The
210
+ `limits.verificationFixes` and `limits.reviewRounds` budgets count fixes the
211
+ same way in both modes. Exhausting `limits.verificationFixes` ends in the
212
+ `implementation_exhausted` hand-off; exhausting `limits.reviewRounds` sends the
213
+ issue through decomposition.
214
+
215
+ The boundary between agent work and deterministic operations stays explicit
216
+ throughout the loop:
217
+
218
+ ```mermaid
219
+ sequenceDiagram
220
+ participant R as Ralphie
221
+ participant GH as GitHub
222
+ participant G as Git
223
+ participant P as Harness
224
+
225
+ R->>G: Capture clean branch checkpoint
226
+ R->>G: Verify destination and remote base
227
+ R->>P: Start fresh implementation session
228
+ P-->>R: Edit the checkout
229
+ R->>G: Stage all changes and read exact diff
230
+
231
+ alt Changes present
232
+ loop Until approved or review rounds exhausted
233
+ R->>G: Run configured verification commands (when any)
234
+ opt Verification command fails and repair budget remains
235
+ R->>P: Resume implementer with /diagnosing-bugs
236
+ P-->>R: Update the checkout
237
+ R->>G: Restage and rerun verification
238
+ end
239
+ R->>G: Create local candidate commit
240
+ R->>P: Start standards and spec review sessions
241
+ P-->>R: Return findings; Ralphie computes the verdict
242
+ opt Changes requested and budget remains
243
+ R->>P: Resume implementer with /implement
244
+ P-->>R: Update the checkout
245
+ R->>G: Restage changes and read exact diff
246
+ end
247
+ end
248
+ R->>G: Reverify the exact approved staged tree
249
+ alt Review approved
250
+ R->>G: Squash candidates, commit with the implementer's message
251
+ R->>G: Revalidate destination, HEAD, and remote base
252
+ R->>G: Push selected branch without force
253
+ G->>GH: Send branch update
254
+ GH-->>G: Accept or return authoritative policy rejection
255
+ R->>GH: Close issue as completed
256
+ else Review budget exhausted
257
+ R->>G: Preserve patch and restore checkpoint
258
+ R->>GH: Continue through decomposition
259
+ end
260
+ else No changes
261
+ R->>P: Start fresh structured resolution verification
262
+ P-->>R: Return status and concrete evidence
263
+ opt Resolved
264
+ R->>GH: Close issue as completed
265
+ end
266
+ end
267
+ ```
268
+
269
+ ## Hand-offs
270
+
271
+ Anything that needs a human becomes a hand-off, and hand-offs are always on.
272
+ Ralphie replaces the issue's triage state label (the issue ends with exactly
273
+ one of the five, named by the `labels` mapping), posts one comment that
274
+ starts with the AI disclaimer `> *This was generated by AI during triage.*`,
275
+ and records a `hand-off` outcome. Because the issue no longer carries the
276
+ agent-ready label, the next run's intake skips it until a human relabels it.
277
+
278
+ | Reason | State label | Comment |
279
+ | --- | --- | --- |
280
+ | `missing_information`, `conflicting_requirements`, `cannot_reproduce`, `outdated_premise` | `needs-info` | Triage Notes: what was established, what is still needed from the reporter. |
281
+ | `implementation_exhausted` (attempts or verification repairs ran out, or the implementation failed terminally) | `ready-for-human` | `## Hand-off` write-up: reason, summary, what was tried, what a human must decide, and the diagnostics location. |
282
+ | `decomposition_limit_reached` (review never converged at the depth limit) | `ready-for-human` | Same write-up. |
283
+ | `external_dependency` reported by a session | `ready-for-human` | Same write-up. |
284
+
285
+ Exhausted attempts preserve a diagnostic patch and restore the clean checkout
286
+ before the hand-off, like the other recovery paths. An open-blocker skip (from
287
+ pre-flight or from queue order) is not a hand-off and changes nothing on
288
+ GitHub, and neither is a deferral caused by a limit, outage or expired login. The write-up deliberately avoids the `## Agent Brief` heading, which
289
+ marks the contract an implementer works from. The recovery details are in
290
+ [Operations and recovery](operations-and-recovery.md#hand-off-handling).
291
+
292
+ ## Decomposition workflow
293
+
294
+ 1. Ask a `decomposer` session to run the vendored `/to-tickets` skill with an
295
+ overlay: the quiz is skipped and the breakdown comes back as structured
296
+ output instead of being published.
297
+ 2. Create child issues, blockers first, in the to-tickets issue template
298
+ (`## Parent`, `## What to build`, `## Acceptance criteria`, `## Blocked by`)
299
+ behind Ralphie's hidden stable marker, each with the agent-ready label
300
+ (`labels.ready-for-agent`) plus every label of the parent that appears in
301
+ `intake.requireLabels`, so a run scoped by those labels picks the children up.
302
+ 3. Attach each created or recovered child to the original issue as a **native
303
+ GitHub sub-issue**, reconciling against GitHub's reported hierarchy.
304
+ 4. Represent each declared `dependsOn` edge as a **native GitHub
305
+ `blocked_by` dependency** and persist the dependency mapping artifact.
306
+ 5. Leave the original issue **open and its body untouched**. It is the
307
+ tracking parent; it is never closed as a duplicate merely because it was
308
+ decomposed.
309
+
310
+ Stable markers and persisted child mappings make the workflow retry-safe: a
311
+ retry discovers previously created children instead of duplicating them,
312
+ and native relationships are reconciled idempotently. Eligible children can
313
+ enter the main implementation loop during the same run; the decomposed parent
314
+ stays out of the queue because it is a tracking issue, not executable work.
315
+
316
+ ```mermaid
317
+ flowchart LR
318
+ A[Original issue] --> B[Structured task breakdown]
319
+ B --> C{Existing child marker?}
320
+ C -->|Yes| D[Reuse child issue]
321
+ C -->|No| E[Create child issue]
322
+ D --> F[Reconcile native sub-issues]
323
+ E --> F
324
+ F --> G[Create native blocked_by dependencies]
325
+ G --> H[Keep parent open and untouched]
326
+ H --> I[Refresh open-issue queue]
327
+ ```
328
+
329
+ The decomposition session is read-only and returns an
330
+ `issueBreakdownDecisionSchema` result containing at least two children, each
331
+ with a stable key, a title, what to build, acceptance criteria, and its
332
+ blockers. Each child must fit one session, and the blocking graph must be
333
+ acyclic. The breakdown is persisted before the first GitHub mutation.
334
+
335
+ Each child receives a stable marker containing root, parent, key, and depth.
336
+ The positive `limits.maxDecompositionDepth` setting (default `3`) bounds recursive
337
+ splitting and is persisted in run state. If direct pre-flight routing or review
338
+ exhaustion would exceed it, Ralphie does not attempt another breakdown: it
339
+ hands the issue off as `ready-for-human` with reason
340
+ `decomposition_limit_reached` and continues independent queued work. The
341
+ issue is not marked complete, so its dependents remain blocked.
342
+ Ralphie discovers those markers and reconciles them with any persisted mapping
343
+ before creating anything. Thus a lost create response or a partial linking
344
+ failure does not blindly duplicate children. Creation,
345
+ number recording, native sub-issue attachment, and dependency creation
346
+ are separate mutations; a child already
347
+ attached to the wrong parent, or a native relationship that disagrees with a
348
+ child's marker, halts with a recovery diagnostic instead of silently
349
+ reparenting or duplicating issues.
350
+
351
+ The decomposed parent remains open as the native tracking issue and exposes
352
+ GitHub's completion progress for its sub-issues. Ralphie never edits its body;
353
+ it recognises the parent by its native sub-issues (and by open children whose
354
+ marker names it), so it is not queued again. It is closed as `completed`,
355
+ with one explanatory comment, only when its child work is finished: completing the
356
+ final child reconciles its parent immediately, and every run also
357
+ reconciles decomposed parents it discovers or refreshes, so a parent whose
358
+ final child closed in a previous run is completed on a later run. The open-issue
359
+ queue is refreshed after decomposition; newly eligible children can run during
360
+ the same invocation. If dependencies remain open after the queue is exhausted,
361
+ Ralphie records each blocked issue as a skipped outcome naming its open
362
+ blockers and leaves it pending instead of handing it to an agent, then drains
363
+ later work and completes the run. Blocked issues remain open and keep their
364
+ labels.
365
+
366
+ A pre-flight `fitsOneSession: false` route returns `decomposed`. Review exhaustion returns an
367
+ `escalated` outcome containing the recovery diagnostic path and, after
368
+ successful decomposition, the created child numbers. Both transitions refresh
369
+ the queue.
370
+
371
+ ### Platform support for native sub-issues and dependencies
372
+
373
+ Native sub-issues and `blocked_by` dependencies are GitHub REST features required
374
+ for decomposition. Ralphie's current GitHub client targets `github.com` only;
375
+ GitHub Enterprise Server is not supported. There is **no body-link fallback**:
376
+ Ralphie never silently degrades to body-only hierarchy semantics.
377
+
378
+ - Creating, recovering, or linking children fails with an actionable error
379
+ naming the missing platform capability when an endpoint is unavailable or the
380
+ token lacks issue write permission.
381
+ - The compatibility check is per live operation: the first relationship read or
382
+ write against an unsupported endpoint surfaces the error.
383
+ - Recovery metadata (stable markers and the persisted key/dependency mappings)
384
+ remains the idempotency record, so a run can continue after a recoverable
385
+ relationship failure without blindly duplicating children.
386
+ - To verify the required `github.com` endpoints before a live run:
387
+ `gh api repos/{owner}/{repo}/issues/1/sub_issues` and
388
+ `gh api repos/{owner}/{repo}/issues/1/dependencies/blocked_by` should return
389
+ `200` (an empty list) rather than `404`.
390
+
391
+ ## Delivery
392
+
393
+ | Issue checkout | Delivery | Source issue closure |
394
+ | --- | --- | --- |
395
+ | Selected base branch | Commit and non-force push directly to that branch; verify remote SHA and clean checkout | Close directly as `completed` after verified delivery. |
396
+
397
+ The direct-push path never uses force. A push rejection is authoritative: the
398
+ created commit and artifacts are retained, the run halts, and a later run
399
+ re-evaluates the still-open issue from a fresh checkout. Inspect the retained
400
+ workspace before the next run removes it.
401
+
402
+ ## Queue behavior
403
+
404
+ Issue work is sequential. With the default `created:asc` sort, issues are
405
+ processed oldest-first. When no branch is configured, Ralphie uses `main` when
406
+ it exists and otherwise `master`.
407
+
408
+ For command syntax and all defaults, see the [CLI reference](cli-reference.md).
409
+ For workspace, Git, and GitHub guardrails, see [Safety](safety.md). For
410
+ interruption and failure boundaries, see
411
+ [Operations and recovery](operations-and-recovery.md).
package/package.json CHANGED
@@ -1,6 +1,86 @@
1
1
  {
2
2
  "name": "@beremaran/ralphie",
3
- "version": "0.0.0-stage",
4
- "stub": true,
5
- "description": "Temporary package placeholder for staged publishing"
3
+ "version": "0.2.0",
4
+ "description": "Turn a GitHub issue queue into reviewed commits with coding-agent harnesses.",
5
+ "module": "./dist/ralphie.js",
6
+ "main": "./dist/ralphie.js",
7
+ "type": "module",
8
+ "bin": {
9
+ "ralphie": "./dist/ralphie.js"
10
+ },
11
+ "exports": {
12
+ ".": "./dist/ralphie.js"
13
+ },
14
+ "files": [
15
+ "dist/ralphie.js",
16
+ "README.md",
17
+ "CHANGELOG.md",
18
+ "LICENSE",
19
+ "docs/architecture.md",
20
+ "docs/cli-reference.md",
21
+ "docs/configuration.md",
22
+ "docs/development.md",
23
+ "docs/getting-started.md",
24
+ "docs/operations-and-recovery.md",
25
+ "docs/README.md",
26
+ "docs/safety.md",
27
+ "docs/workflows.md",
28
+ "vendor/mattpocock-skills"
29
+ ],
30
+ "repository": {
31
+ "type": "git",
32
+ "url": "https://github.com/beremaran/ralphie"
33
+ },
34
+ "homepage": "https://github.com/beremaran/ralphie",
35
+ "bugs": {
36
+ "url": "https://github.com/beremaran/ralphie/issues"
37
+ },
38
+ "license": "MIT",
39
+ "keywords": [
40
+ "bun",
41
+ "cli",
42
+ "github",
43
+ "github-issues",
44
+ "automation",
45
+ "ai",
46
+ "claude-code",
47
+ "agent",
48
+ "npm"
49
+ ],
50
+ "engines": {
51
+ "bun": ">=1.3.0"
52
+ },
53
+ "publishConfig": {
54
+ "access": "public"
55
+ },
56
+ "devDependencies": {
57
+ "@biomejs/biome": "^2.5.10",
58
+ "@types/bun": "latest",
59
+ "@types/proper-lockfile": "^4.1.4",
60
+ "typescript": "^5"
61
+ },
62
+ "dependencies": {
63
+ "@opentui/core": "^0.5.11",
64
+ "octokit": "^5.0.5",
65
+ "proper-lockfile": "^4.1.2",
66
+ "zod": "^4.4.3"
67
+ },
68
+ "scripts": {
69
+ "start": "bun run index.ts",
70
+ "dev": "bun --watch ./index.ts",
71
+ "build": "bun run scripts/build.ts",
72
+ "build:package": "bun run scripts/build.ts",
73
+ "check": "bun run format:check && bun run lint && bun run typecheck && bun run source:audit && bun run test && bun run build",
74
+ "format": "biome format --write .",
75
+ "format:check": "biome format .",
76
+ "lint": "biome lint .",
77
+ "package:check": "bun run scripts/package-smoke.ts",
78
+ "package:inspect": "bun run scripts/package-smoke.ts --dry-run",
79
+ "prepack": "bun run build:package",
80
+ "source:audit": "bun run scripts/source-reachability.ts --json",
81
+ "smoke:live": "bun run scripts/live-smoke.ts",
82
+ "skills:sync": "bun run scripts/skills-sync.ts",
83
+ "test": "bun test tests",
84
+ "typecheck": "bunx tsc --noEmit"
85
+ }
6
86
  }
@@ -0,0 +1,21 @@
1
+ MIT License
2
+
3
+ Copyright (c) 2026 Matt Pocock
4
+
5
+ Permission is hereby granted, free of charge, to any person obtaining a copy
6
+ of this software and associated documentation files (the "Software"), to deal
7
+ in the Software without restriction, including without limitation the rights
8
+ to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
9
+ copies of the Software, and to permit persons to whom the Software is
10
+ furnished to do so, subject to the following conditions:
11
+
12
+ The above copyright notice and this permission notice shall be included in all
13
+ copies or substantial portions of the Software.
14
+
15
+ THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
16
+ IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
17
+ FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
18
+ AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
19
+ LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
20
+ OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
21
+ SOFTWARE.
@@ -0,0 +1,87 @@
1
+ ---
2
+ name: code-review
3
+ description: "Review the changes since a fixed point (commit, branch, tag, or merge-base) along two axes: Standards (does the code follow this repo's documented coding standards?) and Spec (does the code match what the originating issue/spec asked for?). Runs both reviews in parallel sub-agents and reports them side by side. Use when the user wants to review a branch, a PR, work-in-progress changes, or asks to \"review since X\"."
4
+ ---
5
+
6
+ Two-axis review of the diff between `HEAD` and a fixed point the user supplies:
7
+
8
+ - **Standards**: does the code conform to this repo's documented coding standards?
9
+ - **Spec**: does the code faithfully implement the originating issue / spec?
10
+
11
+ Both axes run as **parallel sub-agents** so they don't pollute each other's context, then this skill aggregates their findings.
12
+
13
+ The issue tracker should have been provided to you. If `docs/agents/issue-tracker.md` is missing, tell the user to run `/setup-matt-pocock-skills`.
14
+
15
+ ## Process
16
+
17
+ ### 1. Pin the fixed point
18
+
19
+ Whatever the user said is the fixed point (a commit SHA, branch name, tag, `main`, `HEAD~5`, etc.). If they didn't specify one, ask for it.
20
+
21
+ Capture the diff command once: `git diff <fixed-point>...HEAD` (three-dot, so the comparison is against the merge-base). Also note the list of commits via `git log <fixed-point>..HEAD --oneline`.
22
+
23
+ Before going further, confirm the fixed point resolves (`git rev-parse <fixed-point>`) and the diff is non-empty. A bad ref or empty diff should fail here, not inside two parallel sub-agents.
24
+
25
+ ### 2. Identify the spec source
26
+
27
+ Look for the originating spec, in this order:
28
+
29
+ 1. Issue references in the commit messages (`#123`, `Closes #45`, GitLab `!67`, etc.), fetched via the workflow in `docs/agents/issue-tracker.md`.
30
+ 2. A path the user passed as an argument.
31
+ 3. A spec file under `docs/`, `specs/`, or `.scratch/` matching the branch name or feature.
32
+ 4. If nothing is found, ask the user where the spec is. If they say there isn't one, the **Spec** sub-agent will skip and report "no spec available".
33
+
34
+ ### 3. Identify the standards sources
35
+
36
+ Anything in the repo that documents how code should be written, such as `CODING_STANDARDS.md` or `CONTRIBUTING.md`.
37
+
38
+ On top of whatever the repo documents, the Standards axis always carries the **smell baseline** below: a fixed set of Fowler code smells (_Refactoring_, ch.3) that applies even when a repo documents nothing. Two rules bind it:
39
+
40
+ - **The repo overrides.** A documented repo standard always wins; where it endorses something the baseline would flag, suppress the smell.
41
+ - **Always a judgement call.** Each smell is a labelled heuristic ("possible Feature Envy"), never a hard violation. Like any standard here, skip anything tooling already enforces.
42
+
43
+ Each smell reads *what it is* → *how to fix*; match it against the diff:
44
+
45
+ - **Mysterious Name**: a function, variable, or type whose name doesn't reveal what it does or holds. → rename it; if no honest name comes, the design's murky.
46
+ - **Duplicated Code**: the same logic shape appears in more than one hunk or file in the change. → extract the shared shape, call it from both.
47
+ - **Feature Envy**: a method that reaches into another object's data more than its own. → move the method onto the data it envies.
48
+ - **Data Clumps**: the same few fields or params keep travelling together (a type wanting to be born). → bundle them into one type, pass that.
49
+ - **Primitive Obsession**: a primitive or string standing in for a domain concept that deserves its own type. → give the concept its own small type.
50
+ - **Repeated Switches**: the same `switch`/`if`-cascade on the same type recurs across the change. → replace with polymorphism, or one map both sites share.
51
+ - **Shotgun Surgery**: one logical change forces scattered edits across many files in the diff. → gather what changes together into one module.
52
+ - **Divergent Change**: one file or module is edited for several unrelated reasons. → split so each module changes for one reason.
53
+ - **Speculative Generality**: abstraction, parameters, or hooks added for needs the spec doesn't have. → delete it; inline back until a real need shows.
54
+ - **Message Chains**: long `a.b().c().d()` navigation the caller shouldn't depend on. → hide the walk behind one method on the first object.
55
+ - **Middle Man**: a class or function that mostly just delegates onward. → cut it, call the real target direct.
56
+ - **Refused Bequest**: a subclass or implementer that ignores or overrides most of what it inherits. → drop the inheritance, use composition.
57
+
58
+ ### 4. Spawn both sub-agents in parallel
59
+
60
+ **Standards sub-agent prompt** should include:
61
+
62
+ - The full diff command and commit list.
63
+ - The list of standards-source files you found in step 3, **plus the smell baseline from step 3** pasted in full (the sub-agent has no other access to it).
64
+ - The brief: "Report, per file/hunk where relevant, (a) every place the diff violates a documented standard: cite the standard (file + the rule); and (b) any baseline smell you spot: name it and quote the hunk. Distinguish hard violations from judgement calls: documented-standard breaches can be hard, but baseline smells are always judgement calls, and a documented repo standard overrides the baseline. Skip anything tooling enforces. Under 400 words."
65
+
66
+ **Spec sub-agent prompt** should include:
67
+
68
+ - The diff command and commit list.
69
+ - The path or fetched contents of the spec.
70
+ - The brief: "Report: (a) requirements the spec asked for that are missing or partial; (b) behaviour in the diff that wasn't asked for (scope creep); (c) requirements that look implemented but where the implementation looks wrong. Quote the spec line for each finding. Under 400 words."
71
+
72
+ If the spec is missing, skip the Spec sub-agent and note this in the final report.
73
+
74
+ ### 5. Aggregate
75
+
76
+ Present the two reports under `## Standards` and `## Spec` headings, verbatim or lightly cleaned. Do **not** merge or rerank findings, because the two axes are deliberately separate (see _Why two axes_).
77
+
78
+ End with a one-line summary: total findings per axis, and the worst issue _within each axis_ (if any). Don't pick a single winner across axes: that's the reranking the separation exists to prevent.
79
+
80
+ ## Why two axes
81
+
82
+ A change can pass one axis and fail the other:
83
+
84
+ - Code that follows every standard but implements the wrong thing → **Standards pass, Spec fail.**
85
+ - Code that does exactly what the issue asked but breaks the project's conventions → **Spec pass, Standards fail.**
86
+
87
+ Reporting them separately stops one axis from masking the other.
@@ -0,0 +1,3 @@
1
+ interface:
2
+ display_name: "Code Review"
3
+ short_description: "Review a diff on standards and spec"
@@ -0,0 +1,37 @@
1
+ # Deepening
2
+
3
+ How to deepen a cluster of shallow modules safely, given its dependencies. Assumes the vocabulary in [SKILL.md](SKILL.md): **module**, **interface**, **seam**, **adapter**.
4
+
5
+ ## Dependency categories
6
+
7
+ When assessing a candidate for deepening, classify its dependencies. The category determines how the deepened module is tested across its seam.
8
+
9
+ ### 1. In-process
10
+
11
+ Pure computation, in-memory state, no I/O. Always deepenable: merge the modules and test through the new interface directly. No adapter needed.
12
+
13
+ ### 2. Local-substitutable
14
+
15
+ Dependencies that have local test stand-ins (PGLite for Postgres, in-memory filesystem). Deepenable if the stand-in exists. The deepened module is tested with the stand-in running in the test suite. The seam is internal; no port at the module's external interface.
16
+
17
+ ### 3. Remote but owned (Ports & Adapters)
18
+
19
+ Your own services across a network boundary (microservices, internal APIs). Define a **port** (interface) at the seam. The deep module owns the logic; the transport is injected as an **adapter**. Tests use an in-memory adapter. Production uses an HTTP/gRPC/queue adapter.
20
+
21
+ Recommendation shape: *"Define a port at the seam, implement an HTTP adapter for production and an in-memory adapter for testing, so the logic sits in one deep module even though it's deployed across a network."*
22
+
23
+ ### 4. True external (Mock)
24
+
25
+ Third-party services (Stripe, Twilio, etc.) you don't control. The deepened module takes the external dependency as an injected port; tests provide a mock adapter.
26
+
27
+ ## Seam discipline
28
+
29
+ - **One adapter means a hypothetical seam. Two adapters means a real one.** Don't introduce a port unless at least two adapters are justified (typically production + test). A single-adapter seam is just indirection.
30
+ - **Internal seams vs external seams.** A deep module can have internal seams (private to its implementation, used by its own tests) as well as the external seam at its interface. Don't expose internal seams through the interface just because tests use them.
31
+
32
+ ## Testing strategy: replace, don't layer
33
+
34
+ - Old unit tests on shallow modules become waste once tests at the deepened module's interface exist; delete them.
35
+ - Write new tests at the deepened module's interface. The **interface is the test surface**.
36
+ - Tests assert on observable outcomes through the interface, not internal state.
37
+ - Tests should survive internal refactors, since they describe behaviour, not implementation. If a test has to change when the implementation changes, it's testing past the interface.