@wemuda/launchrail 1.10.0 → 1.10.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -70,6 +70,13 @@ const POLICY = {
70
70
  // Tries per ticket: 1 attempt + 1 retry with a fresh context, then park. Deferrals
71
71
  // (a declared blocker had not landed yet) hand their attempt back, capped separately.
72
72
  attempts: A.attempts ?? 2,
73
+ // CI-wait re-polls for the merge gate — distinct from build attempts. A gate agent
74
+ // gets one turn and cannot idle-wait (a background sleep never resumes it), so a PR
75
+ // whose CI is still running comes back "ci-timeout": not done yet, not a failure.
76
+ // Re-poll the cheap gate this many times (reads only until it can merge) before the
77
+ // ticket spends a fresh implementer attempt — a slow CI must never cost a rebuild
78
+ // (ADR-0027). Concurrent tickets in the round supply the wall-clock CI needs.
79
+ gateWaits: A.gateWaits ?? 6,
73
80
  // Backstop against a graph that never drains; deferral rounds spend from this too.
74
81
  maxRounds: A.maxRounds ?? 25,
75
82
  // Re-read the tracker between rounds so externally closed tickets unblock things.
@@ -264,11 +271,11 @@ ${IDEMPOTENCY}
264
271
  Report honestly via the schema: "pr-open" once the PR exists against ${pre.base}; "merged" only when an adopted PR turned out to be already merged; "blocked" when a declared blocker had not landed; "verify-failed" when the verification gate would not go green; "conflict" when a conflict was too ambiguous to resolve without losing behavior (say which files and why); "failed" otherwise, with a summary a fresh retry can act on. List deliberately-out-of-scope discoveries in "punted".`
265
272
  }
266
273
 
267
- function gatePrompt(pre, ticket, build) {
274
+ function gatePrompt(pre, ticket, build, waited = 0) {
268
275
  return `You are the merge gate for ticket #${ticket.number}: PR #${build.pr} is open against ${pre.base}.
269
276
  Tracker access from this environment: ${pre.trackerAccess}
270
277
  You own the CI wait, the squash-merge, and the tracker bookkeeping — and nothing else. You never write code, never push commits, never repair a failing branch; a failing PR is reported, not fixed here.
271
- 1. Wait for CI on the PR, if the repository has it. Space checks with the Monitor tool — NEVER a bare background sleep (it will not resume you) and never a busy loop. Treat ~20 minutes as the budget; beyond it report status "ci-timeout".
278
+ ${waited > 0 ? `This is CI re-poll ${waited}: the CI was still running when the gate last looked, so the loop sent you back — it has very likely finished by now. Start again from step 1.\n` : ''}1. Check the PR's CI, if the repository has it. You have ONE turn and cannot idle-wait — a background sleep will not resume you, so do NOT try to sit on a long wait. Poll a handful of times, spacing the checks with the Monitor tool (never a bare sleep, never a busy loop), across the minute or two you can cover, then act on what you see: green → step 2; still in progress when your turn is ending → report status "ci-timeout" and STOP. A "ci-timeout" is not a failure — the loop simply re-polls you, cheaply, until CI lands; never merge on a run that has not finished green.
272
279
  2. CI green (or absent): check mergeability against ${pre.base} — the base may have moved since CI started. Mergeable: squash-merge via the tracker API; if the base moves between check and merge, re-check and retry up to 3 times. A real conflict is status "not-mergeable" — name the conflicting files if the API reports them.
273
280
  3. Merged: read issue #${ticket.number} back and close it explicitly if it is still open — "Closes #n" only auto-fires from the default branch${POLICY.target ? ', and this run does not merge there' : ', and squash-merge does not reliably fire it even there'} — then remove the ralph:building label. Report status "merged" with the merge commit sha.
274
281
  4. CI failed on the PR: report status "ci-failed" with the failing check and a summary a fresh implementer can act on. Fix nothing.`
@@ -365,15 +372,28 @@ async function drive(pre, ticket) {
365
372
  // subagent's background sleep never resumes it), and a single gate owner keeps
366
373
  // merge ordering sane. A failing gate hands the ticket back as a failed attempt;
367
374
  // the fresh retry adopts the PR via the idempotency clause, repairs, hands off again.
368
- const gate = await agent(gatePrompt(pre, ticket, build), {
369
- label: `gate:#${ticket.number}`,
370
- phase: 'Gate',
371
- schema: GATE_SCHEMA,
372
- effort: 'low',
373
- })
374
- if (!gate) {
375
- s.failures.push('gate agent died (infrastructure)')
376
- return { ticket, ok: false, dead: true }
375
+ //
376
+ // But the gate agent cannot idle-wait either — given one turn it can only poll CI a
377
+ // few times, so a PR whose CI is still running comes back "ci-timeout": not done
378
+ // yet, not a failure of the code. Re-poll the cheap gate in place (reads + at most
379
+ // one merge, effort low) up to POLICY.gateWaits times before the ticket falls back
380
+ // to a fresh implementer attempt, so a slow CI can never cost a full rebuild
381
+ // (ADR-0027). Only a real verdict — merged, ci-failed, not-mergeable — leaves the
382
+ // loop early; a CI that never lands still parks the ticket once re-polls run out.
383
+ let gate
384
+ for (let waited = 0; ; waited += 1) {
385
+ gate = await agent(gatePrompt(pre, ticket, build, waited), {
386
+ label: `gate:#${ticket.number}${waited > 0 ? `:ci-wait${waited}` : ''}`,
387
+ phase: 'Gate',
388
+ schema: GATE_SCHEMA,
389
+ effort: 'low',
390
+ })
391
+ if (!gate) {
392
+ s.failures.push('gate agent died (infrastructure)')
393
+ return { ticket, ok: false, dead: true }
394
+ }
395
+ if (gate.status !== 'ci-timeout' || waited >= POLICY.gateWaits) break
396
+ log(`#${ticket.number}: CI still running on PR #${build.pr} — re-polling the gate (${waited + 1}/${POLICY.gateWaits})`)
377
397
  }
378
398
  if (gate.status !== 'merged') {
379
399
  s.failures.push(`[attempt ${s.attempts}] PR #${build.pr} ${gate.status}: ${gate.summary}`)
@@ -90,7 +90,7 @@ When the Ralph loop runs as the `ralph` workflow instead of through this skill,
90
90
  1. **Read the resolved scope back, immediately.** The first `log()` lines state it ("Scoped to #11, #12", "Stopping after 5 verified merge(s)", or "No scope — building the whole ready frontier"). An unscoped run when the user asked for three tickets is the cheapest failure to catch and the most expensive to miss — stop and relaunch if it is wrong. Scan the listed numbers for anything that is not an implementable ticket: a spec or research issue wearing `ready-for-agent` will be built as if it were work (the workflow excludes and logs obvious cases, but the label is the fix — have it corrected).
91
91
  2. **Establish ground truth from the remote, never from the run's own reports.** On every check-in read the workflow journal (`journal.jsonl`) *and* the tracker/PRs. A merge is real only when the commit is on the base branch and the issue is closed.
92
92
  3. **Arm check-ins across the long waits.** If the session can schedule a self-message, arm one a few minutes out (confirm scope and the first dispatches) and a longer fallback (catch completion or a stall). The workflow's completion notification is the primary signal; the check-ins are the backstop so the run survives an interruption.
93
- 4. **Know the healthy shapes so you don't cry wolf.** A ticket can appear twice in Build — that is the retry policy, or a *deferral* because its dependency had not landed yet (not a failure). Build ending at PR-open with a separate Gate agent doing the merge is the design, not a stall. A ticket only truly fails after two real attempts, then it parks.
93
+ 4. **Know the healthy shapes so you don't cry wolf.** A ticket can appear twice in Build — that is the retry policy, or a *deferral* because its dependency had not landed yet (not a failure). Build ending at PR-open with a separate Gate agent doing the merge is the design, not a stall. Several Gate dispatches for one PR (`gate:#n`, then `gate:#n:ci-wait1…`) are the loop re-polling a still-running CI in place — cheap, expected on anything but the fastest CI, and *not* a build attempt: a slow CI costs re-polls, never a rebuild (ADR-0027). A ticket only truly fails after two real attempts, then it parks.
94
94
  5. **Intervene by exception, not by reflex.** Parked ticket → dispatch a fresh scoped run for just that one. Stall (an agent stops writing, CI never returns) → diagnose from the journal. Wrong scope or wrong base → stop, fix, relaunch. Otherwise stay out of the way; the loop is built to self-correct.
95
95
  6. **Report once at the end, concretely** — PR numbers, merge commits, issues closed, the verification outcome, anything punted — then disarm the check-ins.
96
96
 
package/package.json CHANGED
@@ -1,6 +1,6 @@
1
1
  {
2
2
  "name": "@wemuda/launchrail",
3
- "version": "1.10.0",
3
+ "version": "1.10.1",
4
4
  "description": "Launchrail — initialize, inspect, update, and validate repositories using the Launchrail development system.",
5
5
  "license": "MIT",
6
6
  "type": "module",