letsdo 0.3.0 → 0.6.1

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (52) hide show
  1. checksums.yaml +4 -4
  2. data/CHANGELOG.md +232 -0
  3. data/README.md +156 -23
  4. data/bin/letsdo +8 -1
  5. data/docs/config.md +176 -0
  6. data/docs/prompts.md +210 -0
  7. data/docs/task-selection.md +244 -0
  8. data/docs/usage.md +296 -0
  9. data/letsdo.gemspec +63 -0
  10. data/lib/letsdo/agent.rb +30 -17
  11. data/lib/letsdo/agent_identity.rb +46 -0
  12. data/lib/letsdo/agent_loop/assignee_hints.rb +34 -0
  13. data/lib/letsdo/agent_loop/tasks.rb +88 -13
  14. data/lib/letsdo/agent_loop.rb +46 -4
  15. data/lib/letsdo/backends/backend.rb +137 -0
  16. data/lib/letsdo/backends/pi/events.rb +77 -0
  17. data/lib/letsdo/backends/pi.rb +88 -0
  18. data/lib/letsdo/cli/builder.rb +102 -0
  19. data/lib/letsdo/cli/builder_assembly.rb +154 -0
  20. data/lib/letsdo/cli/builder_metrics.rb +82 -0
  21. data/lib/letsdo/cli/doctor.rb +16 -0
  22. data/lib/letsdo/cli.rb +17 -56
  23. data/lib/letsdo/config.rb +152 -0
  24. data/lib/letsdo/control/reader.rb +131 -0
  25. data/lib/letsdo/control.rb +6 -3
  26. data/lib/letsdo/doctor/checks.rb +144 -0
  27. data/lib/letsdo/doctor.rb +42 -0
  28. data/lib/letsdo/duration.rb +23 -0
  29. data/lib/letsdo/errors.rb +7 -0
  30. data/lib/letsdo/metrics/fanout.rb +47 -0
  31. data/lib/letsdo/prompt_store.rb +48 -1
  32. data/lib/letsdo/providers/backlog.rb +180 -0
  33. data/lib/letsdo/providers/task.rb +48 -0
  34. data/lib/letsdo/retry_policy.rb +98 -0
  35. data/lib/letsdo/session_recorder/jsonl_writer.rb +73 -0
  36. data/lib/letsdo/session_recorder.rb +188 -0
  37. data/lib/letsdo/task_time_writeback.rb +106 -0
  38. data/lib/letsdo/tui/metrics.rb +10 -3
  39. data/lib/letsdo/tui/renderer.rb +9 -1
  40. data/lib/letsdo/tui/session/terminal.rb +47 -0
  41. data/lib/letsdo/tui/session/view.rb +10 -2
  42. data/lib/letsdo/tui/session.rb +27 -22
  43. data/lib/letsdo/tui/window_title.rb +133 -0
  44. data/lib/letsdo/tui.rb +4 -0
  45. data/lib/letsdo/version.rb +1 -1
  46. data/lib/letsdo.rb +17 -5
  47. metadata +39 -13
  48. data/lib/letsdo/backlog_tasks.rb +0 -59
  49. data/lib/letsdo/cli/launch.rb +0 -87
  50. data/lib/letsdo/pi_runner/events.rb +0 -75
  51. data/lib/letsdo/pi_runner/process.rb +0 -68
  52. data/lib/letsdo/pi_runner.rb +0 -101
data/docs/config.md ADDED
@@ -0,0 +1,176 @@
1
+ # letsdo — configuration reference
2
+
3
+ Configuration comes from two places: environment variables (everything in
4
+ this reference) and an optional YAML front-matter block at the top of an
5
+ agent's prompt file, `agents/<name>.md` — see
6
+ [Per-agent configuration](#per-agent-configuration-yaml-front-matter).
7
+ Nothing else is configured in files.
8
+
9
+ ## Project layout
10
+
11
+ ```
12
+ LETSDO_ROOT (default: the current working directory)
13
+ ├── backlog/ # the Backlog.md tasks — the task provider reads them
14
+ ├── agents/<name>.md # one prompt file per agent
15
+ └── bin/letsdo # the executable (or the installed letsdo command)
16
+ ```
17
+
18
+ - `LETSDO_ROOT` — project root: `agents/` lives there, and the `backlog`
19
+ CLI resolves `backlog/` there. Default: the folder letsdo was started
20
+ from.
21
+ - The agent's assignee is `<name>` by default — the agent works on tasks
22
+ assigned to that name. The tracker stores bare names; `@<name>` is
23
+ prose-only notation (`letsdo doctor` warns about legacy `@`-prefixed
24
+ data).
25
+
26
+ ## Environment variables
27
+
28
+ | Variable | Default | Meaning |
29
+ | --- | --- | --- |
30
+ | `LETSDO_ROOT` | current directory | Project root where `agents/` lives (and where the `backlog` CLI finds `backlog/`). |
31
+ | `LETSDO_PI_FLAGS` | unset (no flags) | Extra pi flags, split on whitespace, e.g. `--model anthropic/claude-sonnet-4-5`. A `--model` here overrides the agent's front-matter `model:` (see [Per-agent configuration](#per-agent-configuration-yaml-front-matter)). |
32
+ | `AGENT_PI_FLAGS` | unset | Fallback for `LETSDO_PI_FLAGS` when it is blank (compatibility with the old `bin/agent`). |
33
+ | `LETSDO_PI_COMMAND` | `pi` | The pi command used to run agents; overridable for tests / fake pi. |
34
+ | `AGENT_ASSIGNEE_HANDLE` | `<name>` | The agent's backlog assignee (used verbatim when set; the bare name is canonical, a legacy `@`-prefixed value still matches but `letsdo doctor` warns). Also injected as the agent's identity in its prompt. |
35
+ | `LETSDO_WAIT_SECONDS` | `10` | Retry interval (seconds) when there are no open tasks. |
36
+ | `AGENT_WAIT_SECONDS` | `10` (via fallback) | Fallback for `LETSDO_WAIT_SECONDS` when it is blank (`bin/agent-loop` compatibility). |
37
+ | `LETSDO_MAX_RETRIES` | `3` | Max consecutive failed runs of the same task before giving up for the session. |
38
+ | `LETSDO_RETRY_BASE` | = `LETSDO_WAIT_SECONDS` | Base backoff seconds; doubles per failure, capped by `LETSDO_RETRY_CAP`. |
39
+ | `LETSDO_RETRY_CAP` | `300` | Maximum backoff seconds between attempts. |
40
+ | `LETSDO_BACKLOG_COMMAND` | `backlog` | The Backlog.md CLI command used as the task provider. |
41
+ | `LETSDO_PROVIDER` | `backlog` | Task provider name used by the loop (currently only `backlog`). |
42
+ | `LETSDO_BACKEND` | `pi` | AI backend that runs each agent (only `pi` today; `LETSDO_PI_COMMAND`/`LETSDO_PI_FLAGS` keep working as before). |
43
+ | `LETSDO_DEBUG` | unset | Set to `1` to trace loop and runner decisions (`[letsdo] loop: ...`) on stderr. |
44
+ | `LETSDO_TASK_TIME_COMMENT` | unset (off) | Set to `1` to append a `letsdo: completed in <time>` comment to each completed task's backlog record at session stop. Off by default: no task file is modified and no extra backlog subprocess runs. |
45
+
46
+ ### Precedence rules
47
+
48
+ - `LETSDO_PI_FLAGS` → `AGENT_PI_FLAGS`: the flags are read from
49
+ `LETSDO_PI_FLAGS`; when it is blank/absent, `AGENT_PI_FLAGS` is used. A
50
+ blank result means no flags.
51
+ - `LETSDO_WAIT_SECONDS` → `AGENT_WAIT_SECONDS`: same pattern, falling back
52
+ to the default `10`. A non-numeric value also falls back to `10`.
53
+ - `LETSDO_RETRY_BASE` → `LETSDO_WAIT_SECONDS`: when
54
+ `LETSDO_RETRY_BASE` is blank or non-numeric, the backoff base falls back
55
+ to the effective `LETSDO_WAIT_SECONDS` value.
56
+ - `LETSDO_MAX_RETRIES`: an invalid (non-integer) value falls back to `3`.
57
+ - `LETSDO_RETRY_CAP`: an invalid (non-numeric) value falls back to `300`.
58
+ - `AGENT_ASSIGNEE_HANDLE`: used as-is when set and non-blank; otherwise the
59
+ one rule: assignee = `<name>` (the bare name). The resolved value is both
60
+ the assignee the loop matches backlog tasks by and the assignee letsdo
61
+ injects into the agent's prompt identity. The match tolerates the legacy
62
+ `@` prefix, whitespace and case, so old data keeps working — but the
63
+ canonical stored value is the bare name, and deviations trigger a
64
+ once-per-run warning plus a `letsdo doctor` WARN.
65
+ - Agent `model` → `LETSDO_PI_FLAGS`/`AGENT_PI_FLAGS`: a `--model` in the
66
+ flags wins; the agent's front-matter `model:` is used only when the flags
67
+ carry no `--model`. See
68
+ [Per-agent configuration](#per-agent-configuration-yaml-front-matter).
69
+
70
+ ### Examples
71
+
72
+ ```sh
73
+ # Point letsdo at a project from anywhere
74
+ export LETSDO_ROOT=/srv/projects/acme-backlog
75
+
76
+ # Choose the pi model
77
+ export LETSDO_PI_FLAGS="--model anthropic/claude-sonnet-4-5"
78
+
79
+ # Less chatty backlog polling
80
+ export LETSDO_WAIT_SECONDS=30
81
+ export LETSDO_DEBUG=1 # trace loop decisions when diagnosing
82
+
83
+ # Non-standard installations
84
+ export LETSDO_PI_COMMAND=/opt/pi/bin/pi
85
+ export LETSDO_BACKLOG_COMMAND=~/.local/bin/backlog
86
+ ```
87
+
88
+ ## Per-agent configuration (YAML front matter)
89
+
90
+ A prompt file may start with a YAML front-matter block that carries launch
91
+ settings for that one agent. The block is optional: a file without it
92
+ behaves exactly as before.
93
+
94
+ ```markdown
95
+ ---
96
+ model: anthropic/claude-sonnet-4-5
97
+ ---
98
+
99
+ # Developer agent (developer)
100
+ You are a developer agent named developer.
101
+ ...
102
+ ```
103
+
104
+ Rules:
105
+
106
+ - The block must be the *very first* thing in `agents/<name>.md`: a `---`
107
+ line, the YAML keys, a closing `---` line. The rest of the file stays the
108
+ prompt.
109
+ - The front matter is stripped before the file is handed to the agent, so it
110
+ never appears in the system prompt.
111
+ - `model` is the only key consumed today:
112
+ - `model: <name>` is passed to pi as `--model <name>` for that agent.
113
+ - Absent → no `--model` is added and pi uses its own default model.
114
+ - A missing file, no front matter, or invalid YAML is treated the same as
115
+ absent: the block is ignored and the run continues.
116
+ - Precedence: a `--model` coming from `LETSDO_PI_FLAGS`/`AGENT_PI_FLAGS`
117
+ wins over the file's `model:`. If you set a global `--model`, every agent
118
+ uses it and the front matter is ignored — leave `--model` out of the
119
+ global flags to select the model per agent.
120
+ - The block is extensible: unknown keys (e.g. `tags:`) are parsed and
121
+ ignored, so new parameters can be added later without breaking existing
122
+ prompt files.
123
+
124
+ Read by `Letsdo::PromptStore#config` (`lib/letsdo/prompt_store.rb`), passed
125
+ to the backend by `Letsdo::Agent#run` (`lib/letsdo/agent.rb`) and applied in
126
+ `Letsdo::Backends::Pi#initialize` (`lib/letsdo/backends/pi.rb`).
127
+
128
+ ## TERM and TUI selection
129
+
130
+ The interactive TUI is engaged only when **all** of these hold:
131
+
132
+ 1. stdout is a TTY, and
133
+ 2. stdin is a TTY, and
134
+ 3. `TERM` is not `dumb`.
135
+
136
+ Otherwise letsdo prints the plain line-stream output — byte-identical to
137
+ the pre-TUI behavior, with no escape codes. Set `TERM=dumb` (or pipe stdout
138
+ through something) to force the plain mode explicitly, e.g. in scripts or
139
+ CI.
140
+
141
+ ## Where each variable is read
142
+
143
+ - `LETSDO_ROOT` — `Letsdo::CLI#initialize` (`lib/letsdo/cli.rb`).
144
+ - `LETSDO_PI_FLAGS` / `AGENT_PI_FLAGS` — `Letsdo::CLI#parse_pi_flags`
145
+ (`lib/letsdo/cli.rb`).
146
+ - `LETSDO_PI_COMMAND` — `Letsdo::CLI#pi_command` (`lib/letsdo/cli.rb`),
147
+ default `PiRunner::COMMAND` (`lib/letsdo/pi_runner.rb`).
148
+ - `AGENT_ASSIGNEE_HANDLE` — `Letsdo::CLI#assignee_handle`
149
+ (`lib/letsdo/cli.rb`).
150
+ - `LETSDO_WAIT_SECONDS` / `AGENT_WAIT_SECONDS` — `Letsdo::CLI#wait_seconds`
151
+ (`lib/letsdo/cli.rb`).
152
+ - `LETSDO_BACKLOG_COMMAND` — `Letsdo::CLI#backlog_command`
153
+ (`lib/letsdo/cli.rb`).
154
+ - `LETSDO_PROVIDER` — `Letsdo::Config#provider` and
155
+ `Letsdo::CLI::Builder#resolve_provider!` (unknown values fail fast with
156
+ `letsdo: unknown task provider: <name>`, exit code 1).
157
+ - `LETSDO_BACKEND` — `Letsdo::Config#backend` and
158
+ `Letsdo::CLI::Builder#resolve_backend!` (unknown values fail fast with
159
+ `letsdo: unknown AI backend: <name>`, exit code 1).
160
+ - `LETSDO_DEBUG` — `Letsdo::AgentLoop#initialize` (`lib/letsdo/agent_loop.rb`)
161
+ and `Letsdo::PiRunner#initialize` (`lib/letsdo/pi_runner.rb`); enabled when
162
+ the value is exactly `"1"`.
163
+ - `LETSDO_TASK_TIME_COMMENT` — `Letsdo::Config#task_time_comment?`
164
+ (`lib/letsdo/config.rb`); enabled only when the value is exactly `"1"`.
165
+ Consumed by `Letsdo::CLI::Builder#finish_session`, which runs
166
+ `Letsdo::TaskTimeWriteback` at stop.
167
+ - `TERM` — `Letsdo::CLI#tui?` (`lib/letsdo/cli.rb`).
168
+
169
+ ## External Requirements (not configurable)
170
+
171
+ - The **pi CLI** is a runtime requirement of letsdo — an external binary,
172
+ not a rubygem dependency. `LETSDO_PI_COMMAND` only replaces the command
173
+ name. See the [usage guide](usage.md) prerequisites.
174
+ - Test-only environment variables (`FAKE_PI_*`, `FAKE_BACKLOG_*`) belong to
175
+ the test fixtures (`test/fixtures/`) and are not part of the runtime
176
+ configuration.
data/docs/prompts.md ADDED
@@ -0,0 +1,210 @@
1
+ # letsdo — prompt-authoring guide
2
+
3
+ An agent in letsdo is exactly one file: `agents/<name>.md`. The file is the
4
+ agent's whole identity — role, rules, workflow. The assignee is
5
+ derived from the name (`agents/developer.md` ⇄ assignee `developer`; in
6
+ prose you write `@developer`), so the agent
7
+ automatically works on the tasks already assigned to it. Adding a file
8
+ creates an agent; no code, no schemas, no setup.
9
+
10
+ This guide covers what a good agent prompt contains, what to avoid, and two
11
+ worked examples from this repository's own agents.
12
+
13
+ - [How letsdo uses the prompt](#how-letsdo-uses-the-prompt)
14
+ - [Optional: per-agent configuration](#optional-per-agent-configuration)
15
+ - [Must-haves](#must-haves)
16
+ - [Anti-patterns](#anti-patterns)
17
+ - [Worked examples](#worked-examples)
18
+ - [Checklist](#checklist)
19
+
20
+ ## How letsdo uses the prompt
21
+
22
+ - The prompt is read on every run from `<LETSDO_ROOT>/agents/<name>.md` and
23
+ handed to pi as the agent's system prompt (`pi --mode json`).
24
+ - Before the prompt is handed over, letsdo prepends the agent's identity —
25
+ its name and backlog assignee (`Config#assignee_handle`, default
26
+ `<name>`, the canonical bare name) — to it. The injection always happens, whatever the template
27
+ contains, so a template never has to state the name or assignee.
28
+ - When the file is missing, the agent runs on the built-in default prompt
29
+ (process-only instructions); letsdo announces where it looked and how to
30
+ create a prompt (`letsdo <name> --init`).
31
+ - The pipeline is deterministic per run: one run = one agent invocation =
32
+ one task. The prompt defines HOW the agent executes that task.
33
+
34
+ ## Optional: per-agent configuration
35
+
36
+ A prompt file can open with a YAML front-matter block that sets launch
37
+ parameters for that one agent. The block is optional — a file without it
38
+ behaves exactly as before.
39
+
40
+ ```markdown
41
+ ---
42
+ model: anthropic/claude-sonnet-4-5
43
+ ---
44
+
45
+ # Developer agent (developer)
46
+
47
+ You are a developer agent named developer.
48
+ ...
49
+ ```
50
+
51
+ What to know:
52
+
53
+ - The block must be the very first thing in the file (`---`, the keys, a
54
+ closing `---`). It is stripped before the prompt reaches the agent, so it
55
+ never shows up in the system prompt.
56
+ - `model` is the only key used today: it is passed to pi as
57
+ `--model <name>` for this agent. Leave it out and pi uses its own default
58
+ model.
59
+ - A `--model` in `LETSDO_PI_FLAGS`/`AGENT_PI_FLAGS` wins over the file, so
60
+ keep `--model` out of the global flags when you want to choose the model
61
+ per agent.
62
+ - The block is extensible: unknown keys are parsed and ignored, so new
63
+ parameters can be added later without breaking existing prompts.
64
+
65
+ This is the only configuration a *prompt file* carries; everything else is
66
+ an environment variable (see the [configuration reference](config.md)).
67
+
68
+ ## Must-haves
69
+
70
+ A robust agent prompt has these parts:
71
+
72
+ 1. **Identity.** The agent's identity — its name and backlog assignee —
73
+ is injected by letsdo at the top of every prompt, so the
74
+ template does not have to state it. A hardcoded assignee is a bad idea:
75
+ it drifts from `AGENT_ASSIGNEE_HANDLE` and the backlog assignment. You
76
+ can still describe the role in prose; keep the name in the prompt
77
+ matching the file name.
78
+
79
+ 2. **Exactly one task per run.** The main rule. The agent picks up and
80
+ completes exactly one assigned task, then stops. This is what keeps an
81
+ orchestrator loop honest: no context-switching, no runaway churn.
82
+
83
+ ```
84
+ ## Main rule: exactly one task per run
85
+ In a single run you pick up and complete exactly one task assigned to
86
+ you, then stop. You take the next task only in the next run.
87
+ ```
88
+
89
+ 3. **What to do when there are no tasks.** The corollary of the one-task
90
+ rule: an agent must *not* invent work or create tasks itself.
91
+
92
+ ```
93
+ If there are no tasks assigned to you — do not invent work and do not
94
+ create tasks yourself. End the run with a message that there are no
95
+ tasks.
96
+ ```
97
+
98
+ 4. **How to choose a task.** Be explicit about the selection order, because
99
+ the agent will face multiple open tasks:
100
+
101
+ ```
102
+ Take the highest priority task assigned to you in order:
103
+ 0. if a task is already in progress
104
+ 1. priority
105
+ 2. order (ordinal), if priorities are equal
106
+ If the chosen task is currently blocked, take the task that blocks it.
107
+ ```
108
+
109
+ 5. **An execution protocol.** Step-by-step instruction for the lifecycle of
110
+ a task: read the instructions, start (status + assign to yourself), plan,
111
+ record progress as you go, verify and finish, commit. Reference the
112
+ project's actual commands and guides (`backlog instructions
113
+ task-execution`, `backlog instructions task-finalization`, ...) — the
114
+ agent executes them literally.
115
+
116
+ 6. **Deliverables.** What the final result must look like — concrete
117
+ artifacts: implementation tasks for the developer with acceptance
118
+ criteria, documentation files, recorded decisions. Name file paths and
119
+ commands.
120
+
121
+ 7. **Style and language rules.** Project conventions the agent must follow
122
+ in every artifact it writes (for example English-only task tracking).
123
+ State them as a hard rule, not a suggestion.
124
+
125
+ 8. **Prohibitions.** What the agent must never do — explicitly. The most
126
+ valuable one in a multi-agent system: *do not take work that is not
127
+ assigned to you* and *do not complete several tasks in one run*.
128
+
129
+ ## Anti-patterns
130
+
131
+ - **Multi-task runs.** "Work through all open tasks" defeats the loop
132
+ design: one task per run is the contract that makes progress observable
133
+ and stops possible.
134
+ - **Inventing work.** No task → done, not "let me also fix wiring/". A
135
+ worker without an explicit no-invent rule will drift off the backlog.
136
+ - **Vague instructions.** "Improve the code" — improve what, where, how
137
+ will we know it worked? Name files, components, commands and exit
138
+ criteria. The agent's job is execution, not guessing intent.
139
+ - **No stop condition.** A prompt without a finish ("end the run with a
140
+ message when there are no tasks") produces an agent that never knows it
141
+ is done.
142
+ - **Keeping state only in the head.** Long tasks fail when intermediate
143
+ results live only in the conversation. The prompt must require recording
144
+ (task notes, comments) as it goes.
145
+ - **Untestable acceptance criteria.** "Works well" cannot be verified.
146
+ Prefer testable, objective criteria ("`rake test` green, 0 failures",
147
+ "the created task is assigned to developer").
148
+ - **Missing identity/role.** Without a role the agent cannot tell what it
149
+ owns. letsdo injects the name and handle, but the role and rules are the
150
+ template's job.
151
+ - **Mixing languages in tracked artifacts.** In this project the convention
152
+ is English for tasks, notes, decisions and docs (TASK-35/48) even though
153
+ conversations may be in any language. State the convention; do not rely
154
+ on osmosis.
155
+
156
+ ## Worked examples
157
+
158
+ Both examples are real agents of this repository: `agents/developer.md` and
159
+ `agents/analyst.md`.
160
+
161
+ ### agents/developer.md — the implementer
162
+
163
+ Structure: identity → main rule → task selection → execution protocol →
164
+ prohibitions.
165
+
166
+ - **Identity + main rule:** "You are a developer agent named developer",
167
+ then exactly-one-task-per-run with the no-tasks rule.
168
+ - **Task selection:** the priority → ordinal ladder, with the blocked-task
169
+ rule (take the blocker into work).
170
+ - **Execution protocol:** read `backlog instructions task-execution` first;
171
+ move the task to In Progress and assign it to yourself; draft a plan and
172
+ record it; implement in short iterations with progress in task notes;
173
+ on completion verify every acceptance criterion with objective evidence,
174
+ write a final summary and move the task to Done; commit the changes
175
+ including the backlog folder.
176
+ - **Prohibitions:** no unassigned work, no several tasks in one run.
177
+
178
+ ### agents/analyst.md — the designer
179
+
180
+ Same skeleton, different deliverables: identity, main rule, task selection,
181
+ then an analysis protocol (start → plan → research → record rationale in
182
+ task notes/comments → produce deliverables) and a finalization protocol
183
+ (verify each acceptance criterion with objective evidence, mark completed
184
+ items, final summary, Done, commit). The analyst's "deliverables" step is
185
+ the interesting part — it says documentation deliverables are written into
186
+ the repo (`backlog doc/decision`), reference them from the task, and
187
+ implementation work is handed to the developer as new tasks: clear title,
188
+ description, testable acceptance criteria, affected components. Both agents
189
+ end with the same prohibitions and an explicit language/style section.
190
+
191
+ The pattern to copy: **identity → one task per run → selection → protocol
192
+ with concrete commands → deliverables → style rule → prohibitions.**
193
+
194
+ ## Checklist
195
+
196
+ Before writing `agents/<name>.md`, check:
197
+
198
+ - [ ] Role described in prose (the name/handle identity block is injected
199
+ by letsdo).
200
+ - [ ] Main rule: exactly one task per run.
201
+ - [ ] No-tasks rule: do not invent work, do not create tasks.
202
+ - [ ] Task selection: in progress → priority → ordinal; blocked → take the
203
+ blocker.
204
+ - [ ] Execution protocol with the project's real commands and guides.
205
+ - [ ] Deliverables: concrete artifacts, paths, acceptance criteria.
206
+ - [ ] Style/language conventions for tracked artifacts.
207
+ - [ ] Prohibitions (unassigned work, several tasks per run, ...).
208
+ - [ ] Prompt matches the file name (`<name>.md` ⇄ assignee `<name>`).
209
+ - [ ] (Optional) Front-matter `model:` is set only when you want a
210
+ per-agent model — a global `--model` flag overrides it.
@@ -0,0 +1,244 @@
1
+ # letsdo — task selection: feasibility and requirements
2
+
3
+ This document is the outcome of TASK-86: can the next task an agent should
4
+ pick up be selected programmatically and algorithmically? It records how
5
+ selection works today, which criteria an algorithm would use, what data is
6
+ missing, the risks and edge cases, a sketch of a possible approach, and the
7
+ recommendation.
8
+
9
+ - [Summary](#summary)
10
+ - [How selection works today](#how-selection-works-today)
11
+ - [Selection criteria](#selection-criteria)
12
+ - [Missing data](#missing-data)
13
+ - [Risks, limitations, edge cases](#risks-limitations-edge-cases)
14
+ - [Possible approaches](#possible-approaches)
15
+ - [Recommendation](#recommendation)
16
+ - [Out of scope](#out-of-scope)
17
+
18
+ ## Summary
19
+
20
+ Selection today is **LLM-driven, not algorithmic**. `Letsdo::Loop` receives a
21
+ batch of open tasks and runs the agent once per batch element, but the task
22
+ object it holds is discarded before the run: the agent re-queries the backlog
23
+ inside its own prompt and decides which task to work. letsdo never enforces
24
+ the batch order it just read.
25
+
26
+ The backlog CLI already exposes the primitives needed for deterministic
27
+ selection (`--sort priority`, `--ready`, `--type`, `--labels`, `--limit`), and
28
+ `Letsdo::Providers::Backlog` uses none of them. The recommendation is to build
29
+ a **small, scoped deterministic-ordering change** (scope A below): request a
30
+ priority-sorted, ready-only batch and keep the fields the selector needs. Do
31
+ **not** build a cross-agent scheduler (scope C); task injection into the
32
+ prompt (scope B) is a documented follow-up that needs its own decision because
33
+ it changes the agent contract.
34
+
35
+ **Status: scope A is implemented (TASK-95).** The provider now requests
36
+ `--ready --sort priority` and sorts the normalized batch in the adapter (In
37
+ Progress first, then priority, then ordinal, then id), so the run order it
38
+ hands to the loop is the authoritative one. The analysis below is kept as the
39
+ record; scopes B and C remain out of scope.
40
+
41
+ ## How selection works today
42
+
43
+ End-to-end, one `letsdo <name>` session:
44
+
45
+ 1. `Letsdo::Providers::Backlog` runs
46
+ `backlog task list --exclude-status Done --ready
47
+ --sort priority --json` (`lib/letsdo/providers/backlog.rb`,
48
+ `#command_line`) — the CLI line carries no `--assignee` (backlog CLI
49
+ matches it by exact string, which made notation drift invisible) — and
50
+ matches the agent's assignee on the returned tasks in Ruby
51
+ (`#normalized` comparison of the stored values against the resolved
52
+ handle). `--ready` drops tasks whose dependencies are not all
53
+ done; `--sort priority` orders by priority then ordinal as a first pass.
54
+ 2. The adapter projects each raw task onto `TASK_FIELDS =
55
+ %w[id title status priority assignees ordinal type labels milestone]` and
56
+ drops `reporter`, `parentTaskId`, `createdAt`, `updatedAt` and everything
57
+ else. It then sorts the batch itself, because the CLI does not put In
58
+ Progress first: In Progress first, then priority High > Medium > Low, then
59
+ ordinal ascending, then id ascending. The batch the provider returns is
60
+ therefore the authoritative run order (TASK-95).
61
+ 3. `Letsdo::Loop#run_batch` (`lib/letsdo/loop.rb`) iterates the batch once and
62
+ calls `@run_task.call(task)` for each element. The provider is **not**
63
+ re-polled between the runs of one batch.
64
+ 4. `Letsdo::AgentLoop#run_one_task` (`lib/letsdo/agent_loop/tasks.rb`) logs
65
+ the task label and calls the runner. Its default runner is
66
+ `->(_task) { @agent.run }` (`lib/letsdo/agent_loop.rb`, `#assign_opts`) —
67
+ **the task object is thrown away**.
68
+ 5. `Letsdo::Agent#run` (`lib/letsdo/agent.rb`) takes no task argument. It
69
+ prepends the injected identity and hands the prompt to pi. The prompt is
70
+ what chooses the task: `agents/analyst.md` and `agents/developer.md`
71
+ instruct the agent to run
72
+ `backlog task list --assignee <name> --exclude-status Done --sort priority --plain`
73
+ and take the first task, with the "already In Progress first" and "if
74
+ blocked, take the blocker" exceptions.
75
+ 6. Retry state lives in `Letsdo::RetryPolicy` and is keyed on the **batch
76
+ element** passed to `run_one_task` (`@last_attempted[task_key(task)]`),
77
+ then reconciled against the next provider batch. It is not keyed on the
78
+ task the agent actually worked.
79
+
80
+ So the authority over "which task next" sits in the prompt, and letsdo's own
81
+ batch is only a run counter. The order letsdo read does not govern the work.
82
+
83
+ ## Selection criteria
84
+
85
+ A deterministic selector needs these inputs, in this order:
86
+
87
+ 1. **Eligibility**
88
+ - assignee contains the agent's assignee (`Config#assignee_handle`,
89
+ default `<name>`, the canonical bare name); other agents' and `human`
90
+ tasks are never auto-run.
91
+ - `status != Done`.
92
+ - not in retry cooldown and not given up for this session
93
+ (`RetryPolicy#cooldown?` / `#gave_up?`).
94
+ - dependencies satisfied (all blocking tasks done).
95
+ 2. **Ordering**
96
+ - `In Progress` first: a task already started by this agent resumes before
97
+ a new task starts (the prompt's existing rule 3).
98
+ - `priority`: High > Medium > Low.
99
+ - **tie-break: `ordinal` ascending, then `id` ascending.** Equal priority
100
+ is the common case (all three open tasks in this repo are Medium), so the
101
+ tie-break is not optional — without it the order is unstable. The backlog
102
+ CLI's `--sort priority` already orders same-priority tasks by ordinal
103
+ ascending (verified: TASK-86 ordinal 75000 before TASK-89 ordinal 78000).
104
+ 3. **Optional routing inputs** (only if per-agent routing is wanted later)
105
+ - `type` (`spike`, `bug`, `feature`, `chore`, `docs`, `task`) — e.g. the
106
+ analyst takes spikes, the developer takes bugs/features.
107
+ - `labels` / `milestone` — team-defined tags usable as capability filters.
108
+ - `parentTaskId` — subtask ordering, if subtasks ever become selectable
109
+ units.
110
+ 4. **Load**
111
+ - number of open tasks runnable by this agent, or per-agent concurrency.
112
+ This criterion is **not implementable in the current design**: each agent
113
+ runs as its own process and the backlog is the only shared state, so no
114
+ process can see another's queue or in-flight work.
115
+
116
+ ## Missing data
117
+
118
+ | Gap | Where | Impact |
119
+ | --- | --- | --- |
120
+ | No sort requested | `Providers::Backlog#command_line` | Resolved in TASK-95: the command requests `--sort priority` and the adapter applies the full In-Progress-first tie-break. |
121
+ | No readiness filter | same | Resolved in TASK-95: the command requests `--ready`, so blocked tasks never reach the loop. |
122
+ | `type`, `ordinal`, `labels`, `milestone` dropped | `TASK_FIELDS` | Resolved in TASK-95: all four are normalized onto `Task`. |
123
+ | `dependencies` absent from list JSON | backlog `task list --json` | Blocked state is not a field; the provider relies on the `--ready` filter instead of N per-task `backlog task <id> --json` calls. |
124
+ | Task object discarded before the run | `AgentLoop#assign_opts` + `Agent#run` | letsdo cannot force the task it selected; the LLM re-selects and can diverge from the batch element (see risks). |
125
+ | No machine-readable agent capabilities | `agents/<name>.md` front matter (only `model` is consumed, `lib/letsdo/agent.rb`) | Type/skill routing exists only as prose in the prompt, not as data. The front-matter block is the natural place to add `types:` / `skills:`. |
126
+ | No claim/lock | backlog only | Two agents (or two processes of the same agent) can race for the same unassigned task; a status update is not atomic. |
127
+ | No selection telemetry | `Metrics` records provider counts and run outcomes | There is no record of which task was selected, on which basis, or why a task was skipped. |
128
+
129
+ ## Risks, limitations, edge cases
130
+
131
+ - **Equal priority.** Must be resolved deterministically. Use `ordinal`, then
132
+ `id`; do not rely on the raw CLI order.
133
+ - **Blocked top task.** If the highest-priority task is blocked, the selector
134
+ must skip it. The prompt's current rule ("take the blocker") is ambiguous
135
+ for an algorithm: if the blocker is assigned to this agent, selecting the
136
+ blocker is correct; if it is another agent's task, that would violate the
137
+ "do not take work that is not assigned to you" rule, so the selector must
138
+ fall through to the next runnable task and surface the skip.
139
+ - **Transitive and stale dependencies.** `--ready` is documented as
140
+ "unblocked tasks with all dependencies completed" — rely on the CLI rather
141
+ than reimplementing graph traversal locally. A dangling dependency id
142
+ (blocker deleted/archived) may make a task permanently "not ready"; the
143
+ session should make that visible (a `--ready`/blocked count in the log)
144
+ instead of silently never running it.
145
+ - **Batch order vs. prompt order (the real defect).** The loop runs the agent
146
+ once per batch element without re-polling. If the batch is
147
+ `[B(Low), A(High)]` and run 1 does not close `A`, run 2 runs for batch
148
+ element `B` while the agent again selects `A`. `B` starves and the retry
149
+ accounting misattributes a failure to `B` (which was never worked) while
150
+ `A` is not cooled down. Aligning the provider order with the prompt order
151
+ fixes the common case; only task injection (scope B) fixes it absolutely.
152
+ - **New work during a batch.** Because the provider is polled once per batch,
153
+ a High task created mid-batch cannot preempt it; it is picked up after the
154
+ current batch drains. Acceptable for short runs; note it if runs get long.
155
+ - **Race conditions.** Unassigned tasks are a coordination gap: two agents can
156
+ pick the same task. The design's answer is assignment by a human
157
+ (`developer`, `analyst`, `human`), not a lock. A lock/claim mechanism
158
+ would require shared mutable state and belongs to scope C, not here.
159
+ - **Retry interaction.** Any selector must compose with `RetryPolicy`: a task
160
+ in cooldown or given up must not be re-offered, and the outcome must be
161
+ recorded against the task that was actually selected.
162
+ - **Scope creep.** "Algorithmic selection" can silently grow into a central
163
+ scheduler with a queue, locks, load balancing and preemption. That is a
164
+ different architecture (shared coordinator) from letsdo's "one process per
165
+ agent, backlog is the only shared state" model and should not be smuggled
166
+ into this change.
167
+
168
+ ## Possible approaches
169
+
170
+ ### A. Provider-side deterministic ordering (small)
171
+
172
+ Make the provider return the batch in authoritative order and keep the fields
173
+ a selector needs:
174
+
175
+ - request `--ready --sort priority` in
176
+ `Providers::Backlog#command_line`;
177
+ - extend `TASK_FIELDS` with `type`, `ordinal`, `labels`, `milestone`;
178
+ - order in the selector/provider as: In Progress first, then priority, then
179
+ ordinal, then id (the CLI sorts by priority+ordinal but does not put In
180
+ Progress first — verified: `--ready --sort priority` returned the In
181
+ Progress task first here only because its ordinal is lower);
182
+ - `Letsdo::Loop` stays generic; no prompt or `Agent` signature change.
183
+
184
+ Pros: small, no agent-contract change, testable against the fake backlog, uses
185
+ documented CLI features. Cons: the agent still re-selects inside its prompt,
186
+ so letsdo's batch order is only advisory and the failure/accounting defect in
187
+ the risks section narrows but is not eliminated.
188
+
189
+ ### B. Selector + task injection (medium)
190
+
191
+ Build on A: a pure `Letsdo::Selection` object computes the next runnable task,
192
+ and letsdo passes it into the run so the agent does not re-select:
193
+
194
+ - `Agent#run(task:)` injects "your task for this run is `<id>` — `<title>`"
195
+ into the prompt;
196
+ - the prompts' "Choosing a task" section is replaced by "work the task letsdo
197
+ handed you" (the blocker/readiness logic moves into the selector);
198
+ - `AgentLoop` records retry outcomes against the injected task.
199
+
200
+ Pros: one run = one task becomes enforced, not advisory; retry accounting and
201
+ starvation are fixed; per-agent `types:`/`skills:` front matter routing becomes
202
+ possible. Cons: changes the agent contract and both prompt files; needs its own
203
+ decision, prompt tests and a transition for custom prompts that still
204
+ self-select.
205
+
206
+ ### C. Central scheduler (large, not recommended)
207
+
208
+ A coordinator process with a shared queue, atomic claim/lock, per-agent load
209
+ balancing and preemption.
210
+
211
+ Pros: true multi-agent scheduling. Cons: introduces shared mutable state,
212
+ concurrency, a new failure domain and an operational component; contradicts the
213
+ project's stated design ("multiple agents run as separate processes ... they
214
+ coordinate through the shared backlog — nothing else in common", README). Not
215
+ justified by any observed failure.
216
+
217
+ ## Recommendation
218
+
219
+ **Build it — to scope A.** Add deterministic ordering and readiness to the
220
+ provider: `--ready --sort priority`, the missing normalized fields, and the
221
+ explicit In-Progress→priority→ordinal→id order. It is small, low-risk, uses
222
+ existing backlog CLI features, needs no prompt or agent-contract change, and
223
+ removes the common divergence between the order letsdo reads and the order the
224
+ agent picks. **Implemented in TASK-95** (`Providers::Backlog#command_line` and
225
+ `#sort`, `Providers::Task` fields).
226
+
227
+ Record scopes **B and C as out of scope for now**. B (task injection into the
228
+ prompt) is the correct way to fully enforce the batch and fix retry
229
+ misattribution, but it changes the agent contract and both prompt files and
230
+ deserves a separate decision once A is in use. C (central scheduler) is
231
+ explicitly rejected: it conflicts with the backlog-only, one-process-per-agent
232
+ architecture and is not justified by current pain.
233
+
234
+ The follow-up implementation task for scope A is tracked as TASK-95
235
+ (`Deterministic task selection: priority-sorted, ready-only provider batch`).
236
+
237
+ ## Out of scope
238
+
239
+ - Task injection into the agent prompt and the prompt rewrite (scope B).
240
+ - Per-agent `types:`/`skills:` front-matter routing (needs B's injection to be
241
+ enforceable).
242
+ - Cross-agent load balancing, shared queue, atomic claim/lock, preemption
243
+ (scope C).
244
+ - Subtask-level selection and milestone-driven scheduling.