@wemuda/launchrail 1.10.0 → 1.11.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
package/assets/ralph.workflow.js
CHANGED
|
@@ -11,7 +11,7 @@ export const meta = {
|
|
|
11
11
|
name: 'ralph',
|
|
12
12
|
description: 'Autonomous Ralph loop: implement ready tickets with fresh-context subagents, verification-gated',
|
|
13
13
|
whenToUse:
|
|
14
|
-
'The engine for any multi-ticket Ralph run. Scope a run via args: { only: [9, 10], width: 2 }, just [9, 10], or { max: 5 } to stop after 5 verified merges ("the next five" — the frontier picks which, in dependency order). The front door consolidates by DEFAULT (ADR-0026): it passes { target: "spec/44-mvp" } to collect the campaign onto that branch (default branch untouched; released later by one offered PR). Omitting target is the explicit trunk opt-in — each ticket merged straight into the default branch. { canary: true } holds width at 1 until the first verified merge. Args must be JSON — resolve any natural-language scope to ticket numbers, a cap, and a target before launching. For a watchable run (an explicit user ask, or a targeted intervention), use the launch-ralph skill instead — and say why.',
|
|
14
|
+
'The engine for any multi-ticket Ralph run. Scope a run via args: { only: [9, 10], width: 2 }, just [9, 10], or { max: 5 } to stop after 5 verified merges ("the next five" — the frontier picks which, in dependency order). The front door consolidates by DEFAULT (ADR-0026): it passes { target: "spec/44-mvp" } to collect the campaign onto that branch (default branch untouched; released later by one offered PR). In a session pinned to a designated working branch (hosted sessions), the front door passes that branch as the target (ADR-0028). Omitting target is the explicit trunk opt-in — each ticket merged straight into the default branch. { canary: true } holds width at 1 until the first verified merge. Args must be JSON — resolve any natural-language scope to ticket numbers, a cap, and a target before launching. For a watchable run (an explicit user ask, or a targeted intervention), use the launch-ralph skill instead — and say why.',
|
|
15
15
|
phases: [
|
|
16
16
|
{ title: 'Preflight', detail: 'read project config, resolve the integration target, run the verification gate' },
|
|
17
17
|
{ title: 'Graph', detail: 'list ready tickets and their blocking edges, verbatim' },
|
|
@@ -61,7 +61,8 @@ const POLICY = {
|
|
|
61
61
|
// Integration target. The front door consolidates by DEFAULT (ADR-0026): it resolves a
|
|
62
62
|
// scope-native branch name and passes it here, so the campaign collects on that branch and
|
|
63
63
|
// the default branch is never touched — the run ends by offering ONE release PR
|
|
64
|
-
// target -> default.
|
|
64
|
+
// target -> default. A session pinned to a designated working branch passes that
|
|
65
|
+
// branch here (ADR-0028) — the pin re-targets the run, it never changes the engine. '' is the explicit trunk opt-in: each ticket PR merged straight into
|
|
65
66
|
// the default branch; a bare launch with no target is therefore trunk mode — non-default.
|
|
66
67
|
target: A.target ?? '',
|
|
67
68
|
// Canary: hold width at 1 until the run's first verified merge proves the plumbing
|
|
@@ -70,6 +71,13 @@ const POLICY = {
|
|
|
70
71
|
// Tries per ticket: 1 attempt + 1 retry with a fresh context, then park. Deferrals
|
|
71
72
|
// (a declared blocker had not landed yet) hand their attempt back, capped separately.
|
|
72
73
|
attempts: A.attempts ?? 2,
|
|
74
|
+
// CI-wait re-polls for the merge gate — distinct from build attempts. A gate agent
|
|
75
|
+
// gets one turn and cannot idle-wait (a background sleep never resumes it), so a PR
|
|
76
|
+
// whose CI is still running comes back "ci-timeout": not done yet, not a failure.
|
|
77
|
+
// Re-poll the cheap gate this many times (reads only until it can merge) before the
|
|
78
|
+
// ticket spends a fresh implementer attempt — a slow CI must never cost a rebuild
|
|
79
|
+
// (ADR-0027). Concurrent tickets in the round supply the wall-clock CI needs.
|
|
80
|
+
gateWaits: A.gateWaits ?? 6,
|
|
73
81
|
// Backstop against a graph that never drains; deferral rounds spend from this too.
|
|
74
82
|
maxRounds: A.maxRounds ?? 25,
|
|
75
83
|
// Re-read the tracker between rounds so externally closed tickets unblock things.
|
|
@@ -264,11 +272,11 @@ ${IDEMPOTENCY}
|
|
|
264
272
|
Report honestly via the schema: "pr-open" once the PR exists against ${pre.base}; "merged" only when an adopted PR turned out to be already merged; "blocked" when a declared blocker had not landed; "verify-failed" when the verification gate would not go green; "conflict" when a conflict was too ambiguous to resolve without losing behavior (say which files and why); "failed" otherwise, with a summary a fresh retry can act on. List deliberately-out-of-scope discoveries in "punted".`
|
|
265
273
|
}
|
|
266
274
|
|
|
267
|
-
function gatePrompt(pre, ticket, build) {
|
|
275
|
+
function gatePrompt(pre, ticket, build, waited = 0) {
|
|
268
276
|
return `You are the merge gate for ticket #${ticket.number}: PR #${build.pr} is open against ${pre.base}.
|
|
269
277
|
Tracker access from this environment: ${pre.trackerAccess}
|
|
270
278
|
You own the CI wait, the squash-merge, and the tracker bookkeeping — and nothing else. You never write code, never push commits, never repair a failing branch; a failing PR is reported, not fixed here.
|
|
271
|
-
|
|
279
|
+
${waited > 0 ? `This is CI re-poll ${waited}: the CI was still running when the gate last looked, so the loop sent you back — it has very likely finished by now. Start again from step 1.\n` : ''}1. Check the PR's CI, if the repository has it. You have ONE turn and cannot idle-wait — a background sleep will not resume you, so do NOT try to sit on a long wait. Poll a handful of times, spacing the checks with the Monitor tool (never a bare sleep, never a busy loop), across the minute or two you can cover, then act on what you see: green → step 2; still in progress when your turn is ending → report status "ci-timeout" and STOP. A "ci-timeout" is not a failure — the loop simply re-polls you, cheaply, until CI lands; never merge on a run that has not finished green.
|
|
272
280
|
2. CI green (or absent): check mergeability against ${pre.base} — the base may have moved since CI started. Mergeable: squash-merge via the tracker API; if the base moves between check and merge, re-check and retry up to 3 times. A real conflict is status "not-mergeable" — name the conflicting files if the API reports them.
|
|
273
281
|
3. Merged: read issue #${ticket.number} back and close it explicitly if it is still open — "Closes #n" only auto-fires from the default branch${POLICY.target ? ', and this run does not merge there' : ', and squash-merge does not reliably fire it even there'} — then remove the ralph:building label. Report status "merged" with the merge commit sha.
|
|
274
282
|
4. CI failed on the PR: report status "ci-failed" with the failing check and a summary a fresh implementer can act on. Fix nothing.`
|
|
@@ -365,15 +373,28 @@ async function drive(pre, ticket) {
|
|
|
365
373
|
// subagent's background sleep never resumes it), and a single gate owner keeps
|
|
366
374
|
// merge ordering sane. A failing gate hands the ticket back as a failed attempt;
|
|
367
375
|
// the fresh retry adopts the PR via the idempotency clause, repairs, hands off again.
|
|
368
|
-
|
|
369
|
-
|
|
370
|
-
|
|
371
|
-
|
|
372
|
-
|
|
373
|
-
|
|
374
|
-
|
|
375
|
-
|
|
376
|
-
|
|
376
|
+
//
|
|
377
|
+
// But the gate agent cannot idle-wait either — given one turn it can only poll CI a
|
|
378
|
+
// few times, so a PR whose CI is still running comes back "ci-timeout": not done
|
|
379
|
+
// yet, not a failure of the code. Re-poll the cheap gate in place (reads + at most
|
|
380
|
+
// one merge, effort low) up to POLICY.gateWaits times before the ticket falls back
|
|
381
|
+
// to a fresh implementer attempt, so a slow CI can never cost a full rebuild
|
|
382
|
+
// (ADR-0027). Only a real verdict — merged, ci-failed, not-mergeable — leaves the
|
|
383
|
+
// loop early; a CI that never lands still parks the ticket once re-polls run out.
|
|
384
|
+
let gate
|
|
385
|
+
for (let waited = 0; ; waited += 1) {
|
|
386
|
+
gate = await agent(gatePrompt(pre, ticket, build, waited), {
|
|
387
|
+
label: `gate:#${ticket.number}${waited > 0 ? `:ci-wait${waited}` : ''}`,
|
|
388
|
+
phase: 'Gate',
|
|
389
|
+
schema: GATE_SCHEMA,
|
|
390
|
+
effort: 'low',
|
|
391
|
+
})
|
|
392
|
+
if (!gate) {
|
|
393
|
+
s.failures.push('gate agent died (infrastructure)')
|
|
394
|
+
return { ticket, ok: false, dead: true }
|
|
395
|
+
}
|
|
396
|
+
if (gate.status !== 'ci-timeout' || waited >= POLICY.gateWaits) break
|
|
397
|
+
log(`#${ticket.number}: CI still running on PR #${build.pr} — re-polling the gate (${waited + 1}/${POLICY.gateWaits})`)
|
|
377
398
|
}
|
|
378
399
|
if (gate.status !== 'merged') {
|
|
379
400
|
s.failures.push(`[attempt ${s.attempts}] PR #${build.pr} ${gate.status}: ${gate.summary}`)
|
|
@@ -18,7 +18,7 @@ From `.launchrail.yml`: `issueTracker` and the `testing` commands. From the argu
|
|
|
18
18
|
- **a count** — "the next 5", "max 5" → the loop with a merge cap. Don't hand-pick which five: the cap is a stop condition, and the frontier decides the order — the loop stops after that many *verified merges* and leaves the rest ready;
|
|
19
19
|
- **a spec, slice, or epic reference** — "spec #2's tickets", "the rest of slice 1" → resolve it to explicit numbers against the live tracker: the open tickets that belong to it (a "Part of: #n" line, the spec issue's ticket list, or a label), plus any open in-set blockers so the scope stays dependency-closed. Combinations compose: "the next 5 of spec #2" → that spec's tickets *and* a cap of 5.
|
|
20
20
|
|
|
21
|
-
**Resolve the integration target with the scope.** Every multi-ticket run **consolidates by default** (ADR-0022, ADR-0026): its per-ticket PRs collect onto one integration branch, the default branch stays untouched, and the run ends by *offering* one release PR `<target> → <default>` for the user to review and merge — never merging to the default branch on its own. You name that branch before anything launches, scope-native: the branch the user named if they gave one ("collect spec #44 on `spec/44-mvp`"); else the scope's own name when it maps to a spec, epic, or slice (`spec/<n>-<slug>`); else a generated fallback for an ad-hoc frontier (`launch/frontier-<today>`, or `launch/tickets-<n>-<m>` for a handful of loose numbers). A named target the remote lacks is created from the default branch's tip by the run's preflight. **Trunk** — each ticket merged straight to the default branch, live the moment it lands — is now the explicit opt-in: choose it only when the user asks in so many words ("land each ticket on master as it goes"), and never when the environment forbids pushing to the default branch. Whichever target, restate it in the echo — a choice the user should recognize, never silent.
|
|
21
|
+
**Resolve the integration target with the scope.** Every multi-ticket run **consolidates by default** (ADR-0022, ADR-0026): its per-ticket PRs collect onto one integration branch, the default branch stays untouched, and the run ends by *offering* one release PR `<target> → <default>` for the user to review and merge — never merging to the default branch on its own. You name that branch before anything launches, scope-native: the branch the user named if they gave one ("collect spec #44 on `spec/44-mvp`"); else **the session's designated working branch** when the environment pinned one at start — a hosted session's `claude/...` branch exists to receive exactly this run, so use it rather than minting a `spec/*` twin beside it (ADR-0028); else the scope's own name when it maps to a spec, epic, or slice (`spec/<n>-<slug>`); else a generated fallback for an ad-hoc frontier (`launch/frontier-<today>`, or `launch/tickets-<n>-<m>` for a handful of loose numbers). A named target the remote lacks is created from the default branch's tip by the run's preflight. **Trunk** — each ticket merged straight to the default branch, live the moment it lands — is now the explicit opt-in: choose it only when the user asks in so many words ("land each ticket on master as it goes"), and never when the environment forbids pushing to the default branch. A pinned-branch environment ("develop on this branch; no unprompted PRs") constrains the *deliverable*, which consolidation already honors — the default branch stays untouched, the release PR stays offered-only. It never constrains the *engine*: the user typing `/launch-implement` is the explicit go-ahead for the loop's own mechanics — per-ticket `ralph/*` branches and their PRs, every one landing on the target — and is never a reason to drop to sequential in-session building. Whichever target, restate it in the echo — a choice the user should recognize, never silent.
|
|
22
22
|
|
|
23
23
|
**Resolve prose to data before anything launches.** The loop's inputs are ticket numbers and policy values (`only`, `max`, `width`, `target`) — the workflow takes them as JSON args and refuses a natural-language string by design. Translating the user's words into that scope, against live tracker state, is *your* job, and it ends with an echo before any dispatch: "Scope: #14, #15, #19 — the remaining slice-1 tickets; #19 builds after #14. Cap: none. Target: consolidate on `spec/2-checkout` (`master` untouched; one release PR offered at the end). Engine: the `ralph` workflow." A misread scope corrected here costs a sentence; corrected after launch it costs a run. That echo is your one pre-launch report — resolve the scope, repair setup (Step 2), and route (Step 3) without narrating the checks in between; surface a step only when it fails or changes the scope.
|
|
24
24
|
|
|
@@ -30,11 +30,11 @@ If the loop's materials are missing — `modules.ralph` off in the manifest, or
|
|
|
30
30
|
|
|
31
31
|
**The frontier (or any multi-ticket scope):** the engine is the `ralph` workflow (`.claude/workflows/ralph.js`) — launch it with the resolved scope and target as JSON args, e.g. `{ only: [14, 15, 19], max: 5, target: 'spec/2-checkout' }` (`canary: true` on a project's first run), then supervise it per the `launch-ralph` skill, which owns the policies (width, attempts, cap, deferrals, the merge gate, remote-verified merges) and the supervisor's contract. Orchestrating dispatches by hand under that skill instead is the exception, chosen out loud in the echo: the user asked to watch each dispatch, the Workflow tool is unavailable here, or the run is a targeted intervention (one parked ticket). One engine, one shape — a session that invents its own fan-out is not running the loop.
|
|
32
32
|
|
|
33
|
-
**One ticket:** build it here, watchable, under the same contract a Ralph dispatch carries (kept textually parallel with `launch-ralph` — change one, change both). A single ticket has nothing to consolidate, so its base is the **default branch** and its one CI-gated PR merges there — consolidation-by-default is a multi-ticket policy (use a
|
|
33
|
+
**One ticket:** build it here, watchable, under the same contract a Ralph dispatch carries (kept textually parallel with `launch-ralph` — change one, change both). A single ticket has nothing to consolidate, so its base is the **default branch** and its one CI-gated PR merges there — consolidation-by-default is a multi-ticket policy (use a different base only if the user names one, or the session pins a designated branch — then that is the base). One deliberate divergence: you are the session, not a subagent, so you also run the merge gate yourself — waiting on CI here is fine:
|
|
34
34
|
|
|
35
35
|
1. **Dependency gate:** every ticket on the `Blocked by:` line is closed with its work merged. An open blocker stops you before any code — name it and offer to build it first.
|
|
36
36
|
2. Read the ticket and everything it links (spec sections, ADRs, journeys), plus `AGENTS.md`/`CLAUDE.md`.
|
|
37
|
-
3. Label the ticket `ralph:building`; branch `ralph/<n>-<short-slug>` from a fresh sync of the base
|
|
37
|
+
3. Label the ticket `ralph:building`; branch `ralph/<n>-<short-slug>` from a fresh sync of the base resolved above.
|
|
38
38
|
4. Implement by the **`launch-ralph-implement`** contract — TDD, the `verify` gate, browser smoke for user-facing changes, self-review, commit conventions. Name the skill; don't paraphrase it.
|
|
39
39
|
5. Pre-PR sync: merge the latest base; resolve conflicts with `launch-resolving-merge-conflicts`; re-run the gate if anything changed.
|
|
40
40
|
6. Open a PR titled from the ticket with `Closes #<n>`; adopt an existing `ralph/<n>-*` branch or PR rather than opening a second.
|
|
@@ -13,11 +13,11 @@ You are the orchestrator. **You do not write code. You do not read diffs. You do
|
|
|
13
13
|
|
|
14
14
|
One agent implementing a whole backlog in a single session degrades — context fills with diffs and half-remembered state, and quality drops with every ticket. This loop inverts that: every ticket gets a fresh-context implementer subagent that owns the build through an open PR, the loop's merge gate lands it, and nothing anyone reports is trusted until the remote confirms it.
|
|
15
15
|
|
|
16
|
-
This skill is the loop's contract and its supervisor. The loop itself runs as the deterministic `ralph` workflow (`.claude/workflows/ralph.js`, installed by init; `launchrail sync` restores it) — **the workflow is the engine for every multi-ticket run**, launched with the resolved scope and integration target as JSON args and then supervised per this skill; its script state cannot be compacted away. Orchestrating dispatches from this session instead is the exception, and it is chosen out loud — name the engine and why before anything dispatches: the user asked to watch each dispatch, the Workflow tool is unavailable in this environment, or this is a targeted intervention (one parked ticket, re-run watchably). A hand-rolled fan-out that is neither is not the loop. The two forms share one policy block: change a policy here, change it in the workflow too (ADR-0005, field-revised by ADR-0010, ADR-0022, and ADR-0026).
|
|
16
|
+
This skill is the loop's contract and its supervisor. The loop itself runs as the deterministic `ralph` workflow (`.claude/workflows/ralph.js`, installed by init; `launchrail sync` restores it) — **the workflow is the engine for every multi-ticket run**, launched with the resolved scope and integration target as JSON args and then supervised per this skill; its script state cannot be compacted away. Orchestrating dispatches from this session instead is the exception, and it is chosen out loud — name the engine and why before anything dispatches: the user asked to watch each dispatch, the Workflow tool is unavailable in this environment, or this is a targeted intervention (one parked ticket, re-run watchably). A hand-rolled fan-out that is neither is not the loop. And the exception swaps only who orchestrates, never the shape: fresh context per ticket, per-ticket PRs into the target, and the loop-owned gate survive every engine. An environment rule about branches or PRs re-targets the run (see Integration target) — it never selects this exception and never licenses sequential in-session implementation. The two forms share one policy block: change a policy here, change it in the workflow too (ADR-0005, field-revised by ADR-0010, ADR-0022, and ADR-0026).
|
|
17
17
|
|
|
18
18
|
## Policies
|
|
19
19
|
|
|
20
|
-
- **Integration target: declared, singular, restated.** Every run merges its per-ticket PRs into exactly one base, named before anything dispatches and again in the close-out. **Consolidation** (the default, ADR-0026): one integration branch (e.g. `spec/44-mvp`) collects the whole campaign and the default branch is never touched; the front door names it scope-native — the scope's own spec/epic/slice name when it maps to one, else a generated `launch/*` fallback — and passes it as the `target` arg. The run ends by *offering* one release PR `<target> → <default>` — opened only when the user says so. **Trunk** (the explicit opt-in): the repository's default branch — each verified merge is immediately on mainline, and when the run ends there is nothing left to integrate; select it only when the user asks for per-ticket merges to the default branch, and never when the environment forbids pushing there. A named target missing from the remote is created from the default branch's tip before preflight verifies it; a missing *default* branch stays a refusal. In consolidation mode `Closes #n` never auto-fires (auto-close only triggers from the default branch), so the explicit post-merge close is load-bearing, not belt-and-suspenders.
|
|
20
|
+
- **Integration target: declared, singular, restated.** Every run merges its per-ticket PRs into exactly one base, named before anything dispatches and again in the close-out. **Consolidation** (the default, ADR-0026): one integration branch (e.g. `spec/44-mvp`) collects the whole campaign and the default branch is never touched; the front door names it scope-native — the session's designated working branch when the environment pinned one at start, else the scope's own spec/epic/slice name when it maps to one, else a generated `launch/*` fallback — and passes it as the `target` arg. A pinned-branch session (ADR-0028) changes only that name: the user's start authorizes the loop's mechanics, per-ticket `ralph/*` branches and PRs included, and the pin's real demands — default branch untouched, release PR offered not opened — are exactly what consolidation already does. The run ends by *offering* one release PR `<target> → <default>` — opened only when the user says so. **Trunk** (the explicit opt-in): the repository's default branch — each verified merge is immediately on mainline, and when the run ends there is nothing left to integrate; select it only when the user asks for per-ticket merges to the default branch, and never when the environment forbids pushing there. A named target missing from the remote is created from the default branch's tip before preflight verifies it; a missing *default* branch stays a refusal. In consolidation mode `Closes #n` never auto-fires (auto-close only triggers from the default branch), so the explicit post-merge close is load-bearing, not belt-and-suspenders.
|
|
21
21
|
- **Width: 3** implementers at once. Width multiplies conflict rate and shared-machine load, not just throughput — use 1 until a run has landed tickets cleanly on this project (the workflow's `canary: true` encodes exactly that). Cut a batch below width when its tickets would obviously collide (same module, same files); when in doubt, narrow. Tickets that add DB migrations are a known collision: two parallel implementers both claim the next migration number — serialize them, or expect the second to renumber at pre-PR sync.
|
|
22
22
|
- **Cap: none** by default. The user may bound a run ("the next 5"): stop once that many merges have been *verified*, keeping every batch within the remainder so the run cannot overshoot. Failed and deferred dispatches never consume the cap — their slots go to other tickets. Hitting the cap ends the run cleanly: the rest of the frontier stays ready (reported, never parked), and close-out runs as usual.
|
|
23
23
|
- **Attempts: 2** — retry a failed ticket once with a fresh context, then park it.
|
|
@@ -90,7 +90,7 @@ When the Ralph loop runs as the `ralph` workflow instead of through this skill,
|
|
|
90
90
|
1. **Read the resolved scope back, immediately.** The first `log()` lines state it ("Scoped to #11, #12", "Stopping after 5 verified merge(s)", or "No scope — building the whole ready frontier"). An unscoped run when the user asked for three tickets is the cheapest failure to catch and the most expensive to miss — stop and relaunch if it is wrong. Scan the listed numbers for anything that is not an implementable ticket: a spec or research issue wearing `ready-for-agent` will be built as if it were work (the workflow excludes and logs obvious cases, but the label is the fix — have it corrected).
|
|
91
91
|
2. **Establish ground truth from the remote, never from the run's own reports.** On every check-in read the workflow journal (`journal.jsonl`) *and* the tracker/PRs. A merge is real only when the commit is on the base branch and the issue is closed.
|
|
92
92
|
3. **Arm check-ins across the long waits.** If the session can schedule a self-message, arm one a few minutes out (confirm scope and the first dispatches) and a longer fallback (catch completion or a stall). The workflow's completion notification is the primary signal; the check-ins are the backstop so the run survives an interruption.
|
|
93
|
-
4. **Know the healthy shapes so you don't cry wolf.** A ticket can appear twice in Build — that is the retry policy, or a *deferral* because its dependency had not landed yet (not a failure). Build ending at PR-open with a separate Gate agent doing the merge is the design, not a stall. A ticket only truly fails after two real attempts, then it parks.
|
|
93
|
+
4. **Know the healthy shapes so you don't cry wolf.** A ticket can appear twice in Build — that is the retry policy, or a *deferral* because its dependency had not landed yet (not a failure). Build ending at PR-open with a separate Gate agent doing the merge is the design, not a stall. Several Gate dispatches for one PR (`gate:#n`, then `gate:#n:ci-wait1…`) are the loop re-polling a still-running CI in place — cheap, expected on anything but the fastest CI, and *not* a build attempt: a slow CI costs re-polls, never a rebuild (ADR-0027). A ticket only truly fails after two real attempts, then it parks.
|
|
94
94
|
5. **Intervene by exception, not by reflex.** Parked ticket → dispatch a fresh scoped run for just that one. Stall (an agent stops writing, CI never returns) → diagnose from the journal. Wrong scope or wrong base → stop, fix, relaunch. Otherwise stay out of the way; the loop is built to self-correct.
|
|
95
95
|
6. **Report once at the end, concretely** — PR numbers, merge commits, issues closed, the verification outcome, anything punted — then disarm the check-ins.
|
|
96
96
|
|
package/package.json
CHANGED