letsdo 0.4.0 → 0.5.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (49) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +39 -2
  3. data/README.md +141 -12
  4. data/bin/letsdo +5 -0
  5. data/docs/config.md +171 -0
  6. data/docs/prompts.md +209 -0
  7. data/docs/task-selection.md +239 -0
  8. data/docs/usage.md +295 -0
  9. data/letsdo.gemspec +3 -1
  10. data/lib/letsdo/agent.rb +29 -17
  11. data/lib/letsdo/agent_identity.rb +42 -0
  12. data/lib/letsdo/agent_loop/tasks.rb +91 -13
  13. data/lib/letsdo/agent_loop.rb +36 -3
  14. data/lib/letsdo/backends/backend.rb +137 -0
  15. data/lib/letsdo/backends/pi/events.rb +77 -0
  16. data/lib/letsdo/backends/pi.rb +88 -0
  17. data/lib/letsdo/cli/builder.rb +47 -95
  18. data/lib/letsdo/cli/builder_assembly.rb +151 -0
  19. data/lib/letsdo/cli/builder_metrics.rb +82 -0
  20. data/lib/letsdo/cli/doctor.rb +16 -0
  21. data/lib/letsdo/cli.rb +15 -1
  22. data/lib/letsdo/config.rb +66 -4
  23. data/lib/letsdo/control/reader.rb +131 -0
  24. data/lib/letsdo/control.rb +6 -3
  25. data/lib/letsdo/doctor/checks.rb +98 -0
  26. data/lib/letsdo/doctor.rb +42 -0
  27. data/lib/letsdo/duration.rb +23 -0
  28. data/lib/letsdo/errors.rb +7 -0
  29. data/lib/letsdo/metrics/fanout.rb +47 -0
  30. data/lib/letsdo/prompt_store.rb +48 -1
  31. data/lib/letsdo/providers/backlog.rb +124 -0
  32. data/lib/letsdo/providers/task.rb +48 -0
  33. data/lib/letsdo/retry_policy.rb +98 -0
  34. data/lib/letsdo/session_recorder/jsonl_writer.rb +73 -0
  35. data/lib/letsdo/session_recorder.rb +188 -0
  36. data/lib/letsdo/task_time_writeback.rb +106 -0
  37. data/lib/letsdo/tui/metrics.rb +7 -2
  38. data/lib/letsdo/tui/session/terminal.rb +47 -0
  39. data/lib/letsdo/tui/session/view.rb +10 -2
  40. data/lib/letsdo/tui/session.rb +27 -22
  41. data/lib/letsdo/tui/window_title.rb +133 -0
  42. data/lib/letsdo/tui.rb +4 -0
  43. data/lib/letsdo/version.rb +1 -1
  44. data/lib/letsdo.rb +16 -5
  45. metadata +25 -5
  46. data/lib/letsdo/backlog_tasks.rb +0 -59
  47. data/lib/letsdo/pi_runner/events.rb +0 -75
  48. data/lib/letsdo/pi_runner/process.rb +0 -68
  49. data/lib/letsdo/pi_runner.rb +0 -101
data/docs/usage.md ADDED
@@ -0,0 +1,295 @@
1
+ # letsdo — usage guide
2
+
3
+ This guide walks through running letsdo end to end: prerequisites, install,
4
+ the first run, what a session looks like (plain line-stream and TUI), the
5
+ orchestrator loop semantics, and how to run several agents at once.
6
+
7
+ - [Prerequisites](#prerequisites)
8
+ - [Install](#install)
9
+ - [Project layout](#project-layout)
10
+ - [First run](#first-run)
11
+ - [A session in plain mode](#a-session-in-plain-mode)
12
+ - [The interactive TUI](#the-interactive-tui)
13
+ - [How the loop behaves](#how-the-loop-behaves)
14
+ - [Running several agents](#running-several-agents)
15
+ - [Environment self-check (doctor)](#environment-self-check-doctor)
16
+ - [Exit codes](#exit-codes)
17
+ - [Next steps](#next-steps)
18
+
19
+ ## Prerequisites
20
+
21
+ - **Ruby >= 3.0**.
22
+ - The **pi agent CLI** on `PATH` — the AI backend that actually runs each
23
+ agent (`pi --mode json`). Override the command with `LETSDO_PI_COMMAND`
24
+ (see the [configuration reference](config.md)).
25
+ - The **Backlog.md CLI** (`backlog`) on `PATH` — the task provider reads the
26
+ runnable open tasks assigned to an agent via
27
+ `backlog task list --assignee <handle> --ready --sort priority`. Override
28
+ with `LETSDO_BACKLOG_COMMAND`.
29
+
30
+ Tests and the gem build use only Ruby's bundled default gems (Minitest,
31
+ Rake) — no `bundle install` needed.
32
+
33
+ ## Install
34
+
35
+ Build and install the gem from the repository:
36
+
37
+ ```sh
38
+ git clone git@github.com:sergio-fry/letsdo.git
39
+ cd letsdo
40
+ gem build letsdo.gemspec
41
+ gem install letsdo-0.3.0.gem
42
+ ```
43
+
44
+ > **Ruby 4.0.x note.** Some Ruby 4.0.x builds ship default gems out of sync —
45
+ > rdoc 8.0.0 declares `rbs >= 4.0.0` while rbs 3.x is bundled — so the
46
+ > post-install RDoc hook can raise `Gem::ConflictError` even though the gem
47
+ > files are already installed. Fix the environment by installing a matching
48
+ > rbs first (`gem install rbs -v '>= 4.0.0'`); if that is not possible,
49
+ > install the gem without documentation to skip the hook
50
+ > (`gem install letsdo-0.3.0.gem --no-document`).
51
+
52
+ Or run it straight from the checkout without installing:
53
+
54
+ ```sh
55
+ cd letsdo
56
+ ./bin/letsdo --version
57
+ ```
58
+
59
+ ## Project layout
60
+
61
+ letsdo works in a Backlog.md project root — a folder that holds your
62
+ `backlog/` tasks and your `agents/` prompts:
63
+
64
+ ```
65
+ your-backlog-project/
66
+ ├── backlog/ # the Backlog.md tasks (the single source of truth)
67
+ ├── agents/ # one prompt file per agent: agents/<name>.md
68
+ └── ... # anything else — docs, code, etc.
69
+ ```
70
+
71
+ The project root defaults to the current working directory and can be set
72
+ explicitly with `LETSDO_ROOT`.
73
+
74
+ ## First run
75
+
76
+ An agent is just a prompt file. Create one, then run the agent:
77
+
78
+ ```sh
79
+ cd your-backlog-project
80
+
81
+ # Option 1: scaffold a starter prompt, then customize it
82
+ letsdo developer --init # writes agents/developer.md, never runs the agent
83
+
84
+ # Option 2: write agents/developer.md by hand
85
+ # (see the prompt-authoring guide for what a good prompt contains)
86
+
87
+ # Run the agent: it works through all open tasks assigned to @developer
88
+ letsdo developer
89
+ ```
90
+
91
+ No prompt file yet? The agent still runs on the built-in default prompt and
92
+ letsdo tells you once how to create your own:
93
+
94
+ ```
95
+ $ letsdo newcomer
96
+ letsdo: no prompt for newcomer at /home/user/backlog-project/agents/newcomer.md
97
+ letsdo: using the built-in default prompt (create a prompt file with 'letsdo newcomer --init')
98
+ ```
99
+
100
+ `--init` never overwrites an existing prompt and never runs the agent: it
101
+ fails with exit 1 when `agents/<name>.md` already exists or the name is
102
+ unsafe (contains `/` or `\`, or is `.`/`..` — nothing is ever written
103
+ outside `agents/`). Both argument orders work: `letsdo <name> --init` and
104
+ `letsdo --init <name>`.
105
+
106
+ ## A session in plain mode
107
+
108
+ When stdout is **not** a terminal (pipes, CI, `TERM=dumb`), letsdo prints a
109
+ plain line stream:
110
+
111
+ - **stdout** — only the agent's answer text, as it is generated.
112
+ - **stderr** — service and tool lines with a shared `HH:MM:SS` prefix:
113
+
114
+ ```
115
+ HH:MM:SS ⚙ tool_name: arguments tool call (bash → the command, read/write/edit → the path)
116
+ HH:MM:SS ✓ tool_name: done (3s) tool completion, success
117
+ HH:MM:SS ✖ tool_name: error (3s) tool completion, error
118
+ ```
119
+
120
+ Tool result bodies are never printed — the call line and its one-line
121
+ completion are the whole tool story, so the stream stays readable.
122
+
123
+ - Loop service messages on stderr, e.g.:
124
+
125
+ ```
126
+ letsdo: developer has 3 open task(s)
127
+ letsdo: running developer for TASK-42
128
+ letsdo: no open tasks for developer, retrying in 10s
129
+ letsdo: backlog unavailable, retrying in 10s - run `letsdo doctor` to diagnose
130
+ letsdo: stopped
131
+ ```
132
+
133
+ When the backlog stays unreadable, only the first message of a run
134
+ carries the `letsdo doctor` hint; later retries repeat the short line.
135
+
136
+ Stop the loop with `Ctrl+C` (`SIGINT`; `SIGTERM` and `SIGHUP` — terminal
137
+ closed — work too) — a running pi child is terminated and the process
138
+ exits with code 0. With `LETSDO_DEBUG=1` the loop additionally traces
139
+ `[letsdo] loop: ...` decisions to stderr.
140
+
141
+ ### Pause and quit in plain mode
142
+
143
+ When stdin is still a terminal (for example `letsdo developer > run.log`,
144
+ where stdout is redirected but the keyboard is live), plain mode reads the
145
+ same control keys as the TUI:
146
+
147
+ | Key | Action |
148
+ | --- | --- |
149
+ | `p` | pause/resume: mid-run the running pi child is suspended at the kernel level (SIGSTOP); a second `p` resumes it (SIGCONT). Between runs the next task is held until resume. A no-op when nothing is running |
150
+ | `q` | stop — identical to `Ctrl+C`: the pi child is terminated, the process exits with code 0 |
151
+
152
+ The control reader writes nothing, so the plain stream stays
153
+ byte-identical (no escape codes, no echoed input). When stdin is **not** a
154
+ terminal (a pipe, `</dev/null`, CI), the reader is never started and the
155
+ loop is stopped only by `SIGINT`/`SIGTERM`/`SIGHUP`.
156
+
157
+ ## The interactive TUI
158
+
159
+ When stdout **and** stdin are terminals and `TERM` is not `dumb`, `letsdo
160
+ <name>` starts a full-screen interface instead of the line stream:
161
+
162
+ ```
163
+ letsdo · developer (@developer) session 00:12:34
164
+ done 3 · left 2 · task TASK-42 · 00:03:21
165
+ ├──────────────────────────────────────────────────────────┤
166
+ …scrollable combined log: agent text, tool lines and loop
167
+ service messages in arrival order, newest at the bottom…
168
+ ↑/↓ PgUp/PgDn scroll · p pause · r refresh · q quit (p resume while paused)
169
+ ```
170
+
171
+ - **Header** — agent identity (`letsdo · <name> (@handle)`) and the session
172
+ timer (monotonic, ticks every second).
173
+ - **State line** — tasks done in this session, tasks left (from the latest
174
+ backlog query), and either the running task with its elapsed time, a
175
+ waiting reason (`no open tasks (retry in 10s)` / `backlog unavailable
176
+ (retry in 10s)`), or `PAUSED` while the agent is suspended.
177
+ - **Stream** — one combined log of everything the agent produces: answer
178
+ text deltas, tool lines and loop messages, exactly as OutputStreamer
179
+ emits them. The view follows the newest line automatically (tail -f).
180
+ While paused the frame freezes but the log keeps buffering, so nothing
181
+ is lost.
182
+ - **Footer** — the key map; the `p` hint says `p pause` while the agent
183
+ runs and `p resume` while it is suspended.
184
+
185
+ Keys:
186
+
187
+ | Key | Action |
188
+ | --- | --- |
189
+ | `↑` / `↓` | scroll one line (scrolling up leaves auto-follow) |
190
+ | `PgUp` / `PgDn` | scroll one page |
191
+ | `Home` / `End` | jump to top / back to the newest line (auto-follow) |
192
+ | `p` | pause/resume the agent: mid-run the running pi is suspended at the kernel level (SIGSTOP — model generation and tool executions freeze, the frame shows `PAUSED`, the log keeps buffering); between runs the next task is held until resume. A second `p` resumes (SIGCONT). The footer flips between `p pause` and `p resume` |
193
+ | `r` | immediate backlog re-query (updates the "left" counter) |
194
+ | `q` | quit — identical to a stop signal: the pi child is terminated, the terminal is restored, exit code 0 |
195
+ | `Ctrl+C` | same as `q` inside the TUI |
196
+
197
+ Resizes (`SIGWINCH`) repaint the frame without corruption. Every exit path —
198
+ quit, signal, agent loop end — restores the terminal (alternate screen
199
+ left, cursor back). Inside tmux the window is labeled with the agent name
200
+ for the whole session, so side-by-side agent panes stay distinguishable;
201
+ the previous label (and automatic-rename) returns on exit. The TUI only
202
+ renders when it is safe to do so; in any
203
+ other context the output is byte-identical to the plain line stream, which
204
+ is what keeps CI and pipes deterministic.
205
+
206
+ ## How the loop behaves
207
+
208
+ `letsdo <name>` runs an orchestrator loop:
209
+
210
+ 1. **Query** — fetch all open tasks assigned to `@<name>` via the backlog
211
+ CLI (the assignee handle comes from `AGENT_ASSIGNEE_HANDLE`, default
212
+ `@<name>`).
213
+ 2. **Run** — take the next task and run the agent on it. **One run = one
214
+ task**; the agent must not pick up more than one task per run.
215
+ 3. **Repeat** — when a run finishes, query again.
216
+ 4. **Wait** — when no tasks are open (or the backlog is unreadable), wait
217
+ the retry interval (10 s by default, see `LETSDO_WAIT_SECONDS`) and
218
+ query again. An unreadable backlog pauses instead of crashing.
219
+ 5. **Stop** — `Ctrl+C` / `SIGTERM` / `SIGHUP` (or `q` in the TUI, and on a
220
+ terminal stdin in plain mode) stops the
221
+ loop immediately: the running pi child is terminated (even when it was
222
+ paused — SIGCONT comes before SIGTERM), the terminal is restored, exit
223
+ code 0.
224
+
225
+ The waiting is interruptible — a stop signal unwinds the loop right away
226
+ instead of waiting out the retry interval. A non-zero agent exit code is
227
+ reported but does not stop the loop.
228
+
229
+ Pause (`p` in the TUI or in plain mode on a terminal stdin) between runs
230
+ sets a gate the loop polls before starting the next run: while paused, no
231
+ new task is started even when the backlog has open ones; resume lets the
232
+ queued task run.
233
+
234
+ ## Running several agents
235
+
236
+ Agents run as separate processes, each with its own loop and its own
237
+ assignee handle:
238
+
239
+ ```sh
240
+ letsdo developer & # works on @developer tasks
241
+ letsdo analyst & # works on @analyst tasks
242
+ ```
243
+
244
+ They coordinate through the shared backlog — nothing else in common. Any
245
+ number of agents can run simultaneously; the backlog folder is the single
246
+ source of truth.
247
+
248
+ When the agents run in tmux panes, a TUI session labels its window with the
249
+ agent name (`letsdo developer` → window `developer`) and restores the
250
+ previous label on exit. tmux's automatic-rename is disabled for the session
251
+ and restored afterwards, so the label stays put instead of being overwritten
252
+ with the process name (`ruby`). Outside tmux nothing is written and the tmux
253
+ binary is never invoked.
254
+
255
+ ## Environment self-check (doctor)
256
+
257
+ `letsdo doctor` checks whether the environment can actually run the loop
258
+ and prints one line per check with a status tag and an actionable hint:
259
+
260
+ ```
261
+ $ letsdo doctor
262
+ [ OK ] ruby 4.0.2 (>= 3.3)
263
+ [ OK ] pi command found: pi
264
+ [ OK ] backlog command found: backlog
265
+ [ OK ] project root /home/user/backlog-project has backlog/tasks/
266
+ [ OK ] AGENTS.md present at /home/user/backlog-project/AGENTS.md
267
+ [ OK ] agents/ present with 1 prompt(s)
268
+ [INFO] stdout is not a TTY - plain mode
269
+ ```
270
+
271
+ Checks: Ruby version (`>= 3.3`), the `pi` command (`LETSDO_PI_COMMAND`),
272
+ the `backlog` command (`LETSDO_BACKLOG_COMMAND`), `backlog/tasks/` under
273
+ `LETSDO_ROOT`, `AGENTS.md`, a non-empty `agents/`, and whether stdout is a
274
+ TTY (TUI vs plain mode). A missing command or project layout is a `[FAIL]`;
275
+ a missing `AGENTS.md` or `agents/` is a `[WARN]`; the mode line is
276
+ `[INFO]`. Every FAIL and WARN names the fix.
277
+
278
+ It exits `0` when no check FAILs (warnings do not fail the run) and `1`
279
+ otherwise. `doctor` is a reserved agent name: it always runs the self-check
280
+ and never launches an agent.
281
+
282
+ ## Exit codes
283
+
284
+ | Code | Meaning |
285
+ | --- | --- |
286
+ | `0` | `--version` / `--help`; successful loop run and stop (incl. TUI `q`); `--init` created the prompt; `doctor` found no FAIL |
287
+ | `1` | no arguments; unknown option; `--init` on an existing/unsafe name; `doctor` found at least one FAIL |
288
+ | `2` | the AI backend command is missing (clean message, no backtrace) |
289
+
290
+ ## Next steps
291
+
292
+ - [Prompt-authoring guide](prompts.md) — what to put into `agents/<name>.md`.
293
+ - [Configuration reference](config.md) — every environment variable, its
294
+ default and where it is read.
295
+ - The README — why letsdo exists and what it does.
data/letsdo.gemspec CHANGED
@@ -40,8 +40,10 @@ Gem::Specification.new do |spec|
40
40
  # The AI backend is the external `pi` CLI (default, overridable via
41
41
  # LETSDO_PI_COMMAND) — a runtime *requirement*, not a rubygem, so it is
42
42
  # documented in the README rather than declared as a dependency.
43
+ # User-facing guides live in docs/ and are linked from the README, so
44
+ # they must ship inside the gem for installed copies to be self-contained.
43
45
  spec.files = Dir["lib/**/*.rb", "README.md", "LICENSE",
44
- "CHANGELOG.md", "letsdo.gemspec"]
46
+ "CHANGELOG.md", "docs/**/*.md", "letsdo.gemspec"]
45
47
  spec.bindir = "bin"
46
48
  spec.executables = ["letsdo"]
47
49
  spec.require_paths = ["lib"]
data/lib/letsdo/agent.rb CHANGED
@@ -2,40 +2,52 @@
2
2
 
3
3
  module Letsdo
4
4
  # A single agent run: reads the prompt from agents/<name>.md by agent name
5
- # (falling back to the built-in default prompt when the file is missing)
6
- # and runs pi with that prompt. Returns the pi exit code. There are no
7
- # unknown agents — every name runs, with the file prompt when present and
8
- # with Letsdo::DefaultPrompt::TEXT otherwise.
5
+ # (falling back to the built-in default prompt when the file is missing),
6
+ # prepends the agent's identity (name + backlog assignee handle, TASK-85)
7
+ # and delegates to the injected backend_factory. Returns the backend exit
8
+ # code. There are no unknown agents -- every name runs, with the file
9
+ # prompt when present and with Letsdo::DefaultPrompt::TEXT otherwise.
9
10
  #
10
11
  # This is the logic of one bin/agent run: a prompt store + an output
11
- # streamer + a pi runner. The orchestrator (a loop while tasks exist)
12
- # lives in Letsdo::Loop.
12
+ # streamer + a backend factory. The orchestrator (a loop while tasks
13
+ # exist) lives in Letsdo::Loop.
13
14
  class Agent
14
15
  # @param name [String] agent name (agents/<name>.md)
15
16
  # @param root [String] project root (agents/ lives there)
16
- # @param flags [Array<String>] extra pi flags
17
+ # @param backend_factory [Proc] callable(prompt:, streamer:, model:) -> backend
18
+ # The factory creates a fresh backend for each run. The prompt
19
+ # is read from PromptStore before the call; model comes from the
20
+ # agent's front-matter config (nil when absent). All flag/
21
+ # command/env wiring is the factory's job -- Agent keeps no pi
22
+ # vocabulary.
17
23
  # @param streamer [OutputStreamer] where to print output (by default
18
24
  # the real stdout/stderr)
19
- # @param command [String] the pi command (overridable for tests)
20
- def initialize(name:, root:, flags: [], streamer: nil, command: PiRunner::COMMAND)
25
+ # @param handle [String, nil] the agent's backlog assignee handle
26
+ # (Config#assignee_handle). Defaults to @<name>; the launcher
27
+ # passes the resolved handle so the injected identity always
28
+ # matches the handle the backlog tasks are assigned to.
29
+ def initialize(name:, root:, backend_factory:, streamer: nil, handle: nil)
21
30
  @name = name
22
31
  @root = root
23
- @flags = flags
32
+ @backend_factory = backend_factory
24
33
  @streamer = streamer || OutputStreamer.new
25
- @command = command
34
+ @handle = handle || "@#{name}"
26
35
  end
27
36
 
28
- # The runner of the last/current run — lets the orchestrator terminate
29
- # a running pi when the loop is stopped.
30
- attr_reader :runner
37
+ # The backend of the last/current run -- lets the orchestrator
38
+ # terminate a running backend when the loop is stopped.
39
+ attr_reader :backend
31
40
 
32
41
  # Runs the agent once.
33
42
  #
34
- # @return [Integer] pi exit code
43
+ # @return [Integer] backend exit code
35
44
  def run
36
45
  prompt = prompt_store.read(@name) || Letsdo::DefaultPrompt::TEXT
37
- @runner = PiRunner.new(prompt: prompt, flags: @flags, streamer: @streamer, command: @command)
38
- @runner.run
46
+ prompt = Letsdo::AgentIdentity.inject(prompt, name: @name, handle: @handle)
47
+ model = prompt_store.config(@name)[:model]
48
+ @backend = @backend_factory.call(prompt: prompt, streamer: @streamer,
49
+ model: model)
50
+ @backend.run
39
51
  end
40
52
 
41
53
  private
@@ -0,0 +1,42 @@
1
+ # frozen_string_literal: true
2
+
3
+ module Letsdo
4
+ # The identity preamble letsdo injects into every agent's system prompt on
5
+ # launch (TASK-85). The launcher already knows the agent name and its
6
+ # backlog assignee handle; the prompt template may not. Injecting the two
7
+ # facts unconditionally keeps every agent aware of who it is and which
8
+ # tasks are its own — independent of the template content, for the
9
+ # built-in default prompt and for every custom agents/<name>.md alike.
10
+ #
11
+ # The handle is Config#assignee_handle, so the identity an agent reads
12
+ # matches the handle its backlog tasks are assigned to.
13
+ module AgentIdentity
14
+ # The identity block prepended to a prompt. It ends with a blank line so
15
+ # the original prompt keeps its own heading structure.
16
+ #
17
+ # @param name [String] agent name (also the backlog assignee name)
18
+ # @param handle [String] backlog assignee handle (Config#assignee_handle)
19
+ # @return [String]
20
+ def self.preamble(name:, handle:)
21
+ <<~TEXT
22
+ # Your identity
23
+
24
+ You are the agent `#{name}`. Your backlog assignee handle is `#{handle}`:
25
+ the tasks assigned to `#{handle}` are yours to work. Identify yourself
26
+ as this agent and use this handle in every tracked artifact you write.
27
+
28
+ TEXT
29
+ end
30
+
31
+ # Prepends the identity block to a prompt, leaving the prompt itself
32
+ # untouched.
33
+ #
34
+ # @param prompt [String] the agent's system prompt (file or default)
35
+ # @param name [String] agent name
36
+ # @param handle [String] backlog assignee handle
37
+ # @return [String] prompt with the identity block at the very top
38
+ def self.inject(prompt, name:, handle:)
39
+ "#{preamble(name: name, handle: handle)}#{prompt}"
40
+ end
41
+ end
42
+ end
@@ -2,31 +2,62 @@
2
2
 
3
3
  module Letsdo
4
4
  # Provider reporting and per-task run wrapping for AgentLoop.
5
+ #
6
+ # Retry coordination happens here: the provider batch is reconciled
7
+ # against the previous attempt (TASK-68), tasks in retry cooldown (or
8
+ # given up) are filtered out of the batch, and each run records its
9
+ # outcome for the next reconciliation.
5
10
  module AgentLoopTasks
6
11
  private
7
12
 
8
13
  def wrapped_provider
9
14
  lambda do
10
15
  tasks = @task_provider.call
11
- @metrics&.provider_result(tasks&.length)
12
- report_provider(tasks)
13
- tasks
16
+ return provider_unavailable if tasks.nil?
17
+
18
+ reconcile_attempts(tasks)
19
+ filtered = reject_cooled_down(tasks)
20
+ @metrics&.provider_result(filtered.length)
21
+ report_provider(filtered, raw: tasks.length)
22
+ filtered
14
23
  end
15
24
  end
16
25
 
17
- def report_provider(tasks)
18
- if tasks.nil?
19
- debug('provider: backlog unavailable')
20
- @stderr.puts("letsdo: backlog unavailable, retrying in #{@wait_seconds}s")
21
- elsif tasks.empty?
26
+ def report_provider(tasks, raw: nil)
27
+ if tasks.empty?
28
+ report_empty(raw)
29
+ return
30
+ end
31
+
32
+ debug("provider: #{tasks.length} open task(s)")
33
+ @stderr.puts("letsdo: #{@name} has #{tasks.length} open task(s)")
34
+ report_backoff(raw - tasks.length) if raw && raw > tasks.length
35
+ end
36
+
37
+ def report_empty(raw)
38
+ if raw&.positive?
39
+ debug("provider: #{raw} open task(s) in retry backoff")
40
+ @stderr.puts("letsdo: no runnable tasks for #{@name} (retry backoff), " \
41
+ "retrying in #{@wait_seconds}s")
42
+ else
22
43
  debug('provider: no open tasks')
23
44
  @stderr.puts("letsdo: no open tasks for #{@name}, retrying in #{@wait_seconds}s")
24
- else
25
- debug("provider: #{tasks.length} open task(s)")
26
- @stderr.puts("letsdo: #{@name} has #{tasks.length} open task(s)")
27
45
  end
28
46
  end
29
47
 
48
+ def report_backoff(skipped)
49
+ return unless skipped.positive?
50
+
51
+ @stderr.puts("letsdo: #{skipped} open task(s) in retry backoff")
52
+ end
53
+
54
+ def provider_unavailable
55
+ debug('provider: backlog unavailable')
56
+ @stderr.puts(unavailable_message)
57
+ @metrics&.provider_result(nil)
58
+ nil
59
+ end
60
+
30
61
  def wrapped_run
31
62
  lambda do |task|
32
63
  wait_while_paused
@@ -42,21 +73,68 @@ module Letsdo
42
73
  code = @run_one.call(task)
43
74
  debug("agent run exit #{code}")
44
75
  @stderr.puts("letsdo: #{@name} exited with code #{code}") if code != 0
76
+ @last_attempted[task_key(task)] = code
45
77
  ensure
46
- @metrics&.run_finished
78
+ @metrics&.run_finished(code)
47
79
  end
48
80
 
49
81
  def task_label(task)
50
- id = task.respond_to?(:[]) ? task['id'] : nil
82
+ id = task.respond_to?(:id) ? task.id : nil
51
83
  return id.to_s unless id.nil? || id.to_s.empty?
52
84
 
53
85
  task.to_s
54
86
  end
55
87
 
88
+ def task_key(task)
89
+ task_label(task)
90
+ end
91
+
56
92
  def wait_while_paused
57
93
  return unless @pause_gate
58
94
 
59
95
  @sleeper.call(AgentLoop::PAUSE_POLL_SECONDS) while @pause_gate.paused?
60
96
  end
97
+
98
+ # Compare tasks from the previous run batch with the fresh provider
99
+ # result: tasks still open are failures, tasks gone are successes.
100
+ # Tasks already in retry cooldown (or given up) were skipped and are
101
+ # NOT re-recorded as failed.
102
+ def reconcile_attempts(tasks)
103
+ return unless @last_attempted
104
+
105
+ fresh = tasks.to_h { |t| [task_key(t), true] }
106
+ @last_attempted.each_key { |key| reconcile_key(key, fresh) }
107
+ @last_attempted.clear
108
+ end
109
+
110
+ def reconcile_key(key, fresh)
111
+ if fresh.key?(key)
112
+ @retry_policy.record_failure(key)
113
+ log_give_up(key) if @retry_policy.failures(key) >= @retry_policy.max_retries
114
+ else
115
+ @retry_policy.record_success(key)
116
+ end
117
+ end
118
+
119
+ # Filters tasks that must not be attempted this batch (retry cooldown
120
+ # or give-up). A filtered task is also dropped from the attempted-set so
121
+ # the next reconciliation does not re-record it as failed.
122
+ def reject_cooled_down(tasks)
123
+ return tasks if tasks.nil? || tasks.empty?
124
+
125
+ tasks.reject do |task|
126
+ key = task_key(task)
127
+ cooled = @retry_policy.cooldown?(key) || @retry_policy.gave_up?(key)
128
+ @last_attempted.delete(key) if cooled
129
+ cooled
130
+ end
131
+ end
132
+
133
+ def log_give_up(task_key)
134
+ message = "letsdo: giving up on #{task_key} after " \
135
+ "#{@retry_policy.failures(task_key)} failed runs - " \
136
+ 'task stays open, next session will retry it'
137
+ @stderr.puts(message)
138
+ end
61
139
  end
62
140
  end
@@ -18,8 +18,7 @@ module Letsdo
18
18
  end
19
19
 
20
20
  def run
21
- @loop = build_loop
22
- install_signal_handlers
21
+ prepare_run
23
22
  debug("loop start (agent=#{@name}, handle=#{@handle}, wait=#{@wait_seconds}s)")
24
23
  run_until_stopped
25
24
  debug('loop stopped')
@@ -30,6 +29,15 @@ module Letsdo
30
29
  restore_signal_handlers
31
30
  end
32
31
 
32
+ # The `letsdo doctor` hint rides on the first unavailable backlog call
33
+ # of a run only: later nils repeat the same diagnosis, and a wall of
34
+ # hints would just add noise (TASK-93).
35
+ def unavailable_message
36
+ hint = @unavailable_hinted ? '' : ' - run `letsdo doctor` to diagnose'
37
+ @unavailable_hinted = true
38
+ "letsdo: backlog unavailable, retrying in #{@wait_seconds}s#{hint}"
39
+ end
40
+
33
41
  def close_watcher
34
42
  @watcher&.close
35
43
  end
@@ -38,8 +46,14 @@ module Letsdo
38
46
  warn("[letsdo] loop: #{message}") if @debug
39
47
  end
40
48
 
49
+ # Stop the running backend immediately. The signal handler is
50
+ # invoked from the main thread (Ruby 4.0) or from a dedicated signal
51
+ # thread (3.x fallback via AGENT_SIGNAL_THREAD=1). Thread.raise
52
+ # works reliably on CRuby 4.0: M:N fibers make threads interruptible
53
+ # everywhere, so a trap handler may safely raise Letsdo::Stopped on
54
+ # the main thread to unwind the current run.
41
55
  def on_signal(_signum)
42
- @agent&.runner&.terminate_now
56
+ @agent&.backend&.terminate_now
43
57
  raise Letsdo::Stopped
44
58
  end
45
59
 
@@ -50,9 +64,28 @@ module Letsdo
50
64
  @run_one = opts[:run_one] || ->(_task) { @agent.run }
51
65
  @wait_seconds = opts.fetch(:wait_seconds, 10.0)
52
66
  @stderr = opts.fetch(:stderr, $stderr)
67
+ @retry_policy = build_retry_policy(opts)
68
+ @last_attempted = {}
53
69
  assign_control_opts(opts)
54
70
  end
55
71
 
72
+ def prepare_run
73
+ @loop = build_loop
74
+ install_signal_handlers
75
+ @unavailable_hinted = false
76
+ end
77
+
78
+ def build_retry_policy(opts)
79
+ return opts[:retry_policy] if opts[:retry_policy]
80
+
81
+ Letsdo::RetryPolicy.new(
82
+ base: opts.fetch(:retry_base, @wait_seconds),
83
+ cap: opts.fetch(:retry_cap, Letsdo::RetryPolicy::DEFAULT_CAP),
84
+ max_retries: opts.fetch(:max_retries, Letsdo::RetryPolicy::DEFAULT_MAX_RETRIES),
85
+ clock: opts[:clock]
86
+ )
87
+ end
88
+
56
89
  def assign_control_opts(opts)
57
90
  @metrics = opts[:metrics]
58
91
  @pause_gate = opts[:pause_gate]