letsdo 0.3.0 → 0.5.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- checksums.yaml +4 -4
- data/CHANGELOG.md +179 -0
- data/README.md +147 -17
- data/bin/letsdo +5 -0
- data/docs/config.md +171 -0
- data/docs/prompts.md +209 -0
- data/docs/task-selection.md +239 -0
- data/docs/usage.md +295 -0
- data/letsdo.gemspec +63 -0
- data/lib/letsdo/agent.rb +29 -17
- data/lib/letsdo/agent_identity.rb +42 -0
- data/lib/letsdo/agent_loop/tasks.rb +91 -13
- data/lib/letsdo/agent_loop.rb +43 -4
- data/lib/letsdo/backends/backend.rb +137 -0
- data/lib/letsdo/backends/pi/events.rb +77 -0
- data/lib/letsdo/backends/pi.rb +88 -0
- data/lib/letsdo/cli/builder.rb +102 -0
- data/lib/letsdo/cli/builder_assembly.rb +151 -0
- data/lib/letsdo/cli/builder_metrics.rb +82 -0
- data/lib/letsdo/cli/doctor.rb +16 -0
- data/lib/letsdo/cli.rb +17 -56
- data/lib/letsdo/config.rb +133 -0
- data/lib/letsdo/control/reader.rb +131 -0
- data/lib/letsdo/control.rb +6 -3
- data/lib/letsdo/doctor/checks.rb +98 -0
- data/lib/letsdo/doctor.rb +42 -0
- data/lib/letsdo/duration.rb +23 -0
- data/lib/letsdo/errors.rb +7 -0
- data/lib/letsdo/metrics/fanout.rb +47 -0
- data/lib/letsdo/prompt_store.rb +48 -1
- data/lib/letsdo/providers/backlog.rb +124 -0
- data/lib/letsdo/providers/task.rb +48 -0
- data/lib/letsdo/retry_policy.rb +98 -0
- data/lib/letsdo/session_recorder/jsonl_writer.rb +73 -0
- data/lib/letsdo/session_recorder.rb +188 -0
- data/lib/letsdo/task_time_writeback.rb +106 -0
- data/lib/letsdo/tui/metrics.rb +7 -2
- data/lib/letsdo/tui/session/terminal.rb +47 -0
- data/lib/letsdo/tui/session/view.rb +10 -2
- data/lib/letsdo/tui/session.rb +27 -22
- data/lib/letsdo/tui/window_title.rb +133 -0
- data/lib/letsdo/tui.rb +4 -0
- data/lib/letsdo/version.rb +1 -1
- data/lib/letsdo.rb +17 -5
- metadata +37 -12
- data/lib/letsdo/backlog_tasks.rb +0 -59
- data/lib/letsdo/cli/launch.rb +0 -87
- data/lib/letsdo/pi_runner/events.rb +0 -75
- data/lib/letsdo/pi_runner/process.rb +0 -68
- data/lib/letsdo/pi_runner.rb +0 -101
data/docs/prompts.md
ADDED
|
@@ -0,0 +1,209 @@
|
|
|
1
|
+
# letsdo — prompt-authoring guide
|
|
2
|
+
|
|
3
|
+
An agent in letsdo is exactly one file: `agents/<name>.md`. The file is the
|
|
4
|
+
agent's whole identity — role, rules, workflow. The assignee handle is
|
|
5
|
+
derived from the name (`agents/developer.md` ⇄ `@developer`), so the agent
|
|
6
|
+
automatically works on the tasks already assigned to it. Adding a file
|
|
7
|
+
creates an agent; no code, no schemas, no setup.
|
|
8
|
+
|
|
9
|
+
This guide covers what a good agent prompt contains, what to avoid, and two
|
|
10
|
+
worked examples from this repository's own agents.
|
|
11
|
+
|
|
12
|
+
- [How letsdo uses the prompt](#how-letsdo-uses-the-prompt)
|
|
13
|
+
- [Optional: per-agent configuration](#optional-per-agent-configuration)
|
|
14
|
+
- [Must-haves](#must-haves)
|
|
15
|
+
- [Anti-patterns](#anti-patterns)
|
|
16
|
+
- [Worked examples](#worked-examples)
|
|
17
|
+
- [Checklist](#checklist)
|
|
18
|
+
|
|
19
|
+
## How letsdo uses the prompt
|
|
20
|
+
|
|
21
|
+
- The prompt is read on every run from `<LETSDO_ROOT>/agents/<name>.md` and
|
|
22
|
+
handed to pi as the agent's system prompt (`pi --mode json`).
|
|
23
|
+
- Before the prompt is handed over, letsdo prepends the agent's identity —
|
|
24
|
+
its name and backlog assignee handle (`Config#assignee_handle`, default
|
|
25
|
+
`@<name>`) — to it. The injection always happens, whatever the template
|
|
26
|
+
contains, so a template never has to state the name or handle.
|
|
27
|
+
- When the file is missing, the agent runs on the built-in default prompt
|
|
28
|
+
(process-only instructions); letsdo announces where it looked and how to
|
|
29
|
+
create a prompt (`letsdo <name> --init`).
|
|
30
|
+
- The pipeline is deterministic per run: one run = one agent invocation =
|
|
31
|
+
one task. The prompt defines HOW the agent executes that task.
|
|
32
|
+
|
|
33
|
+
## Optional: per-agent configuration
|
|
34
|
+
|
|
35
|
+
A prompt file can open with a YAML front-matter block that sets launch
|
|
36
|
+
parameters for that one agent. The block is optional — a file without it
|
|
37
|
+
behaves exactly as before.
|
|
38
|
+
|
|
39
|
+
```markdown
|
|
40
|
+
---
|
|
41
|
+
model: anthropic/claude-sonnet-4-5
|
|
42
|
+
---
|
|
43
|
+
|
|
44
|
+
# Developer agent (developer)
|
|
45
|
+
|
|
46
|
+
You are a developer agent named developer.
|
|
47
|
+
...
|
|
48
|
+
```
|
|
49
|
+
|
|
50
|
+
What to know:
|
|
51
|
+
|
|
52
|
+
- The block must be the very first thing in the file (`---`, the keys, a
|
|
53
|
+
closing `---`). It is stripped before the prompt reaches the agent, so it
|
|
54
|
+
never shows up in the system prompt.
|
|
55
|
+
- `model` is the only key used today: it is passed to pi as
|
|
56
|
+
`--model <name>` for this agent. Leave it out and pi uses its own default
|
|
57
|
+
model.
|
|
58
|
+
- A `--model` in `LETSDO_PI_FLAGS`/`AGENT_PI_FLAGS` wins over the file, so
|
|
59
|
+
keep `--model` out of the global flags when you want to choose the model
|
|
60
|
+
per agent.
|
|
61
|
+
- The block is extensible: unknown keys are parsed and ignored, so new
|
|
62
|
+
parameters can be added later without breaking existing prompts.
|
|
63
|
+
|
|
64
|
+
This is the only configuration a *prompt file* carries; everything else is
|
|
65
|
+
an environment variable (see the [configuration reference](config.md)).
|
|
66
|
+
|
|
67
|
+
## Must-haves
|
|
68
|
+
|
|
69
|
+
A robust agent prompt has these parts:
|
|
70
|
+
|
|
71
|
+
1. **Identity.** The agent's identity — its name and backlog assignee
|
|
72
|
+
handle — is injected by letsdo at the top of every prompt, so the
|
|
73
|
+
template does not have to state it. A hardcoded handle is a bad idea:
|
|
74
|
+
it drifts from `AGENT_ASSIGNEE_HANDLE` and the backlog assignment. You
|
|
75
|
+
can still describe the role in prose; keep the name in the prompt
|
|
76
|
+
matching the file name.
|
|
77
|
+
|
|
78
|
+
2. **Exactly one task per run.** The main rule. The agent picks up and
|
|
79
|
+
completes exactly one assigned task, then stops. This is what keeps an
|
|
80
|
+
orchestrator loop honest: no context-switching, no runaway churn.
|
|
81
|
+
|
|
82
|
+
```
|
|
83
|
+
## Main rule: exactly one task per run
|
|
84
|
+
In a single run you pick up and complete exactly one task assigned to
|
|
85
|
+
you, then stop. You take the next task only in the next run.
|
|
86
|
+
```
|
|
87
|
+
|
|
88
|
+
3. **What to do when there are no tasks.** The corollary of the one-task
|
|
89
|
+
rule: an agent must *not* invent work or create tasks itself.
|
|
90
|
+
|
|
91
|
+
```
|
|
92
|
+
If there are no tasks assigned to you — do not invent work and do not
|
|
93
|
+
create tasks yourself. End the run with a message that there are no
|
|
94
|
+
tasks.
|
|
95
|
+
```
|
|
96
|
+
|
|
97
|
+
4. **How to choose a task.** Be explicit about the selection order, because
|
|
98
|
+
the agent will face multiple open tasks:
|
|
99
|
+
|
|
100
|
+
```
|
|
101
|
+
Take the highest priority task assigned to you in order:
|
|
102
|
+
0. if a task is already in progress
|
|
103
|
+
1. priority
|
|
104
|
+
2. order (ordinal), if priorities are equal
|
|
105
|
+
If the chosen task is currently blocked, take the task that blocks it.
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
5. **An execution protocol.** Step-by-step instruction for the lifecycle of
|
|
109
|
+
a task: read the instructions, start (status + assign to yourself), plan,
|
|
110
|
+
record progress as you go, verify and finish, commit. Reference the
|
|
111
|
+
project's actual commands and guides (`backlog instructions
|
|
112
|
+
task-execution`, `backlog instructions task-finalization`, ...) — the
|
|
113
|
+
agent executes them literally.
|
|
114
|
+
|
|
115
|
+
6. **Deliverables.** What the final result must look like — concrete
|
|
116
|
+
artifacts: implementation tasks for the developer with acceptance
|
|
117
|
+
criteria, documentation files, recorded decisions. Name file paths and
|
|
118
|
+
commands.
|
|
119
|
+
|
|
120
|
+
7. **Style and language rules.** Project conventions the agent must follow
|
|
121
|
+
in every artifact it writes (for example English-only task tracking).
|
|
122
|
+
State them as a hard rule, not a suggestion.
|
|
123
|
+
|
|
124
|
+
8. **Prohibitions.** What the agent must never do — explicitly. The most
|
|
125
|
+
valuable one in a multi-agent system: *do not take work that is not
|
|
126
|
+
assigned to you* and *do not complete several tasks in one run*.
|
|
127
|
+
|
|
128
|
+
## Anti-patterns
|
|
129
|
+
|
|
130
|
+
- **Multi-task runs.** "Work through all open tasks" defeats the loop
|
|
131
|
+
design: one task per run is the contract that makes progress observable
|
|
132
|
+
and stops possible.
|
|
133
|
+
- **Inventing work.** No task → done, not "let me also fix wiring/". A
|
|
134
|
+
worker without an explicit no-invent rule will drift off the backlog.
|
|
135
|
+
- **Vague instructions.** "Improve the code" — improve what, where, how
|
|
136
|
+
will we know it worked? Name files, components, commands and exit
|
|
137
|
+
criteria. The agent's job is execution, not guessing intent.
|
|
138
|
+
- **No stop condition.** A prompt without a finish ("end the run with a
|
|
139
|
+
message when there are no tasks") produces an agent that never knows it
|
|
140
|
+
is done.
|
|
141
|
+
- **Keeping state only in the head.** Long tasks fail when intermediate
|
|
142
|
+
results live only in the conversation. The prompt must require recording
|
|
143
|
+
(task notes, comments) as it goes.
|
|
144
|
+
- **Untestable acceptance criteria.** "Works well" cannot be verified.
|
|
145
|
+
Prefer testable, objective criteria ("`rake test` green, 0 failures",
|
|
146
|
+
"the created task is assigned to @developer").
|
|
147
|
+
- **Missing identity/role.** Without a role the agent cannot tell what it
|
|
148
|
+
owns. letsdo injects the name and handle, but the role and rules are the
|
|
149
|
+
template's job.
|
|
150
|
+
- **Mixing languages in tracked artifacts.** In this project the convention
|
|
151
|
+
is English for tasks, notes, decisions and docs (TASK-35/48) even though
|
|
152
|
+
conversations may be in any language. State the convention; do not rely
|
|
153
|
+
on osmosis.
|
|
154
|
+
|
|
155
|
+
## Worked examples
|
|
156
|
+
|
|
157
|
+
Both examples are real agents of this repository: `agents/developer.md` and
|
|
158
|
+
`agents/analyst.md`.
|
|
159
|
+
|
|
160
|
+
### agents/developer.md — the implementer
|
|
161
|
+
|
|
162
|
+
Structure: identity → main rule → task selection → execution protocol →
|
|
163
|
+
prohibitions.
|
|
164
|
+
|
|
165
|
+
- **Identity + main rule:** "You are a developer agent named developer",
|
|
166
|
+
then exactly-one-task-per-run with the no-tasks rule.
|
|
167
|
+
- **Task selection:** the priority → ordinal ladder, with the blocked-task
|
|
168
|
+
rule (take the blocker into work).
|
|
169
|
+
- **Execution protocol:** read `backlog instructions task-execution` first;
|
|
170
|
+
move the task to In Progress and assign it to yourself; draft a plan and
|
|
171
|
+
record it; implement in short iterations with progress in task notes;
|
|
172
|
+
on completion verify every acceptance criterion with objective evidence,
|
|
173
|
+
write a final summary and move the task to Done; commit the changes
|
|
174
|
+
including the backlog folder.
|
|
175
|
+
- **Prohibitions:** no unassigned work, no several tasks in one run.
|
|
176
|
+
|
|
177
|
+
### agents/analyst.md — the designer
|
|
178
|
+
|
|
179
|
+
Same skeleton, different deliverables: identity, main rule, task selection,
|
|
180
|
+
then an analysis protocol (start → plan → research → record rationale in
|
|
181
|
+
task notes/comments → produce deliverables) and a finalization protocol
|
|
182
|
+
(verify each acceptance criterion with objective evidence, mark completed
|
|
183
|
+
items, final summary, Done, commit). The analyst's "deliverables" step is
|
|
184
|
+
the interesting part — it says documentation deliverables are written into
|
|
185
|
+
the repo (`backlog doc/decision`), reference them from the task, and
|
|
186
|
+
implementation work is handed to the developer as new tasks: clear title,
|
|
187
|
+
description, testable acceptance criteria, affected components. Both agents
|
|
188
|
+
end with the same prohibitions and an explicit language/style section.
|
|
189
|
+
|
|
190
|
+
The pattern to copy: **identity → one task per run → selection → protocol
|
|
191
|
+
with concrete commands → deliverables → style rule → prohibitions.**
|
|
192
|
+
|
|
193
|
+
## Checklist
|
|
194
|
+
|
|
195
|
+
Before writing `agents/<name>.md`, check:
|
|
196
|
+
|
|
197
|
+
- [ ] Role described in prose (the name/handle identity block is injected
|
|
198
|
+
by letsdo).
|
|
199
|
+
- [ ] Main rule: exactly one task per run.
|
|
200
|
+
- [ ] No-tasks rule: do not invent work, do not create tasks.
|
|
201
|
+
- [ ] Task selection: in progress → priority → ordinal; blocked → take the
|
|
202
|
+
blocker.
|
|
203
|
+
- [ ] Execution protocol with the project's real commands and guides.
|
|
204
|
+
- [ ] Deliverables: concrete artifacts, paths, acceptance criteria.
|
|
205
|
+
- [ ] Style/language conventions for tracked artifacts.
|
|
206
|
+
- [ ] Prohibitions (unassigned work, several tasks per run, ...).
|
|
207
|
+
- [ ] Prompt matches the file name (`<name>.md` ⇄ `@<name>`).
|
|
208
|
+
- [ ] (Optional) Front-matter `model:` is set only when you want a
|
|
209
|
+
per-agent model — a global `--model` flag overrides it.
|
|
@@ -0,0 +1,239 @@
|
|
|
1
|
+
# letsdo — task selection: feasibility and requirements
|
|
2
|
+
|
|
3
|
+
This document is the outcome of TASK-86: can the next task an agent should
|
|
4
|
+
pick up be selected programmatically and algorithmically? It records how
|
|
5
|
+
selection works today, which criteria an algorithm would use, what data is
|
|
6
|
+
missing, the risks and edge cases, a sketch of a possible approach, and the
|
|
7
|
+
recommendation.
|
|
8
|
+
|
|
9
|
+
- [Summary](#summary)
|
|
10
|
+
- [How selection works today](#how-selection-works-today)
|
|
11
|
+
- [Selection criteria](#selection-criteria)
|
|
12
|
+
- [Missing data](#missing-data)
|
|
13
|
+
- [Risks, limitations, edge cases](#risks-limitations-edge-cases)
|
|
14
|
+
- [Possible approaches](#possible-approaches)
|
|
15
|
+
- [Recommendation](#recommendation)
|
|
16
|
+
- [Out of scope](#out-of-scope)
|
|
17
|
+
|
|
18
|
+
## Summary
|
|
19
|
+
|
|
20
|
+
Selection today is **LLM-driven, not algorithmic**. `Letsdo::Loop` receives a
|
|
21
|
+
batch of open tasks and runs the agent once per batch element, but the task
|
|
22
|
+
object it holds is discarded before the run: the agent re-queries the backlog
|
|
23
|
+
inside its own prompt and decides which task to work. letsdo never enforces
|
|
24
|
+
the batch order it just read.
|
|
25
|
+
|
|
26
|
+
The backlog CLI already exposes the primitives needed for deterministic
|
|
27
|
+
selection (`--sort priority`, `--ready`, `--type`, `--labels`, `--limit`), and
|
|
28
|
+
`Letsdo::Providers::Backlog` uses none of them. The recommendation is to build
|
|
29
|
+
a **small, scoped deterministic-ordering change** (scope A below): request a
|
|
30
|
+
priority-sorted, ready-only batch and keep the fields the selector needs. Do
|
|
31
|
+
**not** build a cross-agent scheduler (scope C); task injection into the
|
|
32
|
+
prompt (scope B) is a documented follow-up that needs its own decision because
|
|
33
|
+
it changes the agent contract.
|
|
34
|
+
|
|
35
|
+
**Status: scope A is implemented (TASK-95).** The provider now requests
|
|
36
|
+
`--ready --sort priority` and sorts the normalized batch in the adapter (In
|
|
37
|
+
Progress first, then priority, then ordinal, then id), so the run order it
|
|
38
|
+
hands to the loop is the authoritative one. The analysis below is kept as the
|
|
39
|
+
record; scopes B and C remain out of scope.
|
|
40
|
+
|
|
41
|
+
## How selection works today
|
|
42
|
+
|
|
43
|
+
End-to-end, one `letsdo <name>` session:
|
|
44
|
+
|
|
45
|
+
1. `Letsdo::Providers::Backlog` runs
|
|
46
|
+
`backlog task list --assignee <handle> --exclude-status Done --ready
|
|
47
|
+
--sort priority --json` (`lib/letsdo/providers/backlog.rb`,
|
|
48
|
+
`#command_line`). `--ready` drops tasks whose dependencies are not all
|
|
49
|
+
done; `--sort priority` orders by priority then ordinal as a first pass.
|
|
50
|
+
2. The adapter projects each raw task onto `TASK_FIELDS =
|
|
51
|
+
%w[id title status priority assignees ordinal type labels milestone]` and
|
|
52
|
+
drops `reporter`, `parentTaskId`, `createdAt`, `updatedAt` and everything
|
|
53
|
+
else. It then sorts the batch itself, because the CLI does not put In
|
|
54
|
+
Progress first: In Progress first, then priority High > Medium > Low, then
|
|
55
|
+
ordinal ascending, then id ascending. The batch the provider returns is
|
|
56
|
+
therefore the authoritative run order (TASK-95).
|
|
57
|
+
3. `Letsdo::Loop#run_batch` (`lib/letsdo/loop.rb`) iterates the batch once and
|
|
58
|
+
calls `@run_task.call(task)` for each element. The provider is **not**
|
|
59
|
+
re-polled between the runs of one batch.
|
|
60
|
+
4. `Letsdo::AgentLoop#run_one_task` (`lib/letsdo/agent_loop/tasks.rb`) logs
|
|
61
|
+
the task label and calls the runner. Its default runner is
|
|
62
|
+
`->(_task) { @agent.run }` (`lib/letsdo/agent_loop.rb`, `#assign_opts`) —
|
|
63
|
+
**the task object is thrown away**.
|
|
64
|
+
5. `Letsdo::Agent#run` (`lib/letsdo/agent.rb`) takes no task argument. It
|
|
65
|
+
prepends the injected identity and hands the prompt to pi. The prompt is
|
|
66
|
+
what chooses the task: `agents/analyst.md` and `agents/developer.md`
|
|
67
|
+
instruct the agent to run
|
|
68
|
+
`backlog task list --assignee @<name> --exclude-status Done --sort priority --plain`
|
|
69
|
+
and take the first task, with the "already In Progress first" and "if
|
|
70
|
+
blocked, take the blocker" exceptions.
|
|
71
|
+
6. Retry state lives in `Letsdo::RetryPolicy` and is keyed on the **batch
|
|
72
|
+
element** passed to `run_one_task` (`@last_attempted[task_key(task)]`),
|
|
73
|
+
then reconciled against the next provider batch. It is not keyed on the
|
|
74
|
+
task the agent actually worked.
|
|
75
|
+
|
|
76
|
+
So the authority over "which task next" sits in the prompt, and letsdo's own
|
|
77
|
+
batch is only a run counter. The order letsdo read does not govern the work.
|
|
78
|
+
|
|
79
|
+
## Selection criteria
|
|
80
|
+
|
|
81
|
+
A deterministic selector needs these inputs, in this order:
|
|
82
|
+
|
|
83
|
+
1. **Eligibility**
|
|
84
|
+
- assignee contains the agent's handle (`Config#assignee_handle`, default
|
|
85
|
+
`@<name>`); other agents' and `@human` tasks are never auto-run.
|
|
86
|
+
- `status != Done`.
|
|
87
|
+
- not in retry cooldown and not given up for this session
|
|
88
|
+
(`RetryPolicy#cooldown?` / `#gave_up?`).
|
|
89
|
+
- dependencies satisfied (all blocking tasks done).
|
|
90
|
+
2. **Ordering**
|
|
91
|
+
- `In Progress` first: a task already started by this agent resumes before
|
|
92
|
+
a new task starts (the prompt's existing rule 3).
|
|
93
|
+
- `priority`: High > Medium > Low.
|
|
94
|
+
- **tie-break: `ordinal` ascending, then `id` ascending.** Equal priority
|
|
95
|
+
is the common case (all three open tasks in this repo are Medium), so the
|
|
96
|
+
tie-break is not optional — without it the order is unstable. The backlog
|
|
97
|
+
CLI's `--sort priority` already orders same-priority tasks by ordinal
|
|
98
|
+
ascending (verified: TASK-86 ordinal 75000 before TASK-89 ordinal 78000).
|
|
99
|
+
3. **Optional routing inputs** (only if per-agent routing is wanted later)
|
|
100
|
+
- `type` (`spike`, `bug`, `feature`, `chore`, `docs`, `task`) — e.g. the
|
|
101
|
+
analyst takes spikes, the developer takes bugs/features.
|
|
102
|
+
- `labels` / `milestone` — team-defined tags usable as capability filters.
|
|
103
|
+
- `parentTaskId` — subtask ordering, if subtasks ever become selectable
|
|
104
|
+
units.
|
|
105
|
+
4. **Load**
|
|
106
|
+
- number of open tasks runnable by this agent, or per-agent concurrency.
|
|
107
|
+
This criterion is **not implementable in the current design**: each agent
|
|
108
|
+
runs as its own process and the backlog is the only shared state, so no
|
|
109
|
+
process can see another's queue or in-flight work.
|
|
110
|
+
|
|
111
|
+
## Missing data
|
|
112
|
+
|
|
113
|
+
| Gap | Where | Impact |
|
|
114
|
+
| --- | --- | --- |
|
|
115
|
+
| No sort requested | `Providers::Backlog#command_line` | Resolved in TASK-95: the command requests `--sort priority` and the adapter applies the full In-Progress-first tie-break. |
|
|
116
|
+
| No readiness filter | same | Resolved in TASK-95: the command requests `--ready`, so blocked tasks never reach the loop. |
|
|
117
|
+
| `type`, `ordinal`, `labels`, `milestone` dropped | `TASK_FIELDS` | Resolved in TASK-95: all four are normalized onto `Task`. |
|
|
118
|
+
| `dependencies` absent from list JSON | backlog `task list --json` | Blocked state is not a field; the provider relies on the `--ready` filter instead of N per-task `backlog task <id> --json` calls. |
|
|
119
|
+
| Task object discarded before the run | `AgentLoop#assign_opts` + `Agent#run` | letsdo cannot force the task it selected; the LLM re-selects and can diverge from the batch element (see risks). |
|
|
120
|
+
| No machine-readable agent capabilities | `agents/<name>.md` front matter (only `model` is consumed, `lib/letsdo/agent.rb`) | Type/skill routing exists only as prose in the prompt, not as data. The front-matter block is the natural place to add `types:` / `skills:`. |
|
|
121
|
+
| No claim/lock | backlog only | Two agents (or two processes of the same agent) can race for the same unassigned task; a status update is not atomic. |
|
|
122
|
+
| No selection telemetry | `Metrics` records provider counts and run outcomes | There is no record of which task was selected, on which basis, or why a task was skipped. |
|
|
123
|
+
|
|
124
|
+
## Risks, limitations, edge cases
|
|
125
|
+
|
|
126
|
+
- **Equal priority.** Must be resolved deterministically. Use `ordinal`, then
|
|
127
|
+
`id`; do not rely on the raw CLI order.
|
|
128
|
+
- **Blocked top task.** If the highest-priority task is blocked, the selector
|
|
129
|
+
must skip it. The prompt's current rule ("take the blocker") is ambiguous
|
|
130
|
+
for an algorithm: if the blocker is assigned to this agent, selecting the
|
|
131
|
+
blocker is correct; if it is another agent's task, that would violate the
|
|
132
|
+
"do not take work that is not assigned to you" rule, so the selector must
|
|
133
|
+
fall through to the next runnable task and surface the skip.
|
|
134
|
+
- **Transitive and stale dependencies.** `--ready` is documented as
|
|
135
|
+
"unblocked tasks with all dependencies completed" — rely on the CLI rather
|
|
136
|
+
than reimplementing graph traversal locally. A dangling dependency id
|
|
137
|
+
(blocker deleted/archived) may make a task permanently "not ready"; the
|
|
138
|
+
session should make that visible (a `--ready`/blocked count in the log)
|
|
139
|
+
instead of silently never running it.
|
|
140
|
+
- **Batch order vs. prompt order (the real defect).** The loop runs the agent
|
|
141
|
+
once per batch element without re-polling. If the batch is
|
|
142
|
+
`[B(Low), A(High)]` and run 1 does not close `A`, run 2 runs for batch
|
|
143
|
+
element `B` while the agent again selects `A`. `B` starves and the retry
|
|
144
|
+
accounting misattributes a failure to `B` (which was never worked) while
|
|
145
|
+
`A` is not cooled down. Aligning the provider order with the prompt order
|
|
146
|
+
fixes the common case; only task injection (scope B) fixes it absolutely.
|
|
147
|
+
- **New work during a batch.** Because the provider is polled once per batch,
|
|
148
|
+
a High task created mid-batch cannot preempt it; it is picked up after the
|
|
149
|
+
current batch drains. Acceptable for short runs; note it if runs get long.
|
|
150
|
+
- **Race conditions.** Unassigned tasks are a coordination gap: two agents can
|
|
151
|
+
pick the same task. The design's answer is assignment by a human
|
|
152
|
+
(`@developer`, `@analyst`, `@human`), not a lock. A lock/claim mechanism
|
|
153
|
+
would require shared mutable state and belongs to scope C, not here.
|
|
154
|
+
- **Retry interaction.** Any selector must compose with `RetryPolicy`: a task
|
|
155
|
+
in cooldown or given up must not be re-offered, and the outcome must be
|
|
156
|
+
recorded against the task that was actually selected.
|
|
157
|
+
- **Scope creep.** "Algorithmic selection" can silently grow into a central
|
|
158
|
+
scheduler with a queue, locks, load balancing and preemption. That is a
|
|
159
|
+
different architecture (shared coordinator) from letsdo's "one process per
|
|
160
|
+
agent, backlog is the only shared state" model and should not be smuggled
|
|
161
|
+
into this change.
|
|
162
|
+
|
|
163
|
+
## Possible approaches
|
|
164
|
+
|
|
165
|
+
### A. Provider-side deterministic ordering (small)
|
|
166
|
+
|
|
167
|
+
Make the provider return the batch in authoritative order and keep the fields
|
|
168
|
+
a selector needs:
|
|
169
|
+
|
|
170
|
+
- request `--ready --sort priority` in
|
|
171
|
+
`Providers::Backlog#command_line`;
|
|
172
|
+
- extend `TASK_FIELDS` with `type`, `ordinal`, `labels`, `milestone`;
|
|
173
|
+
- order in the selector/provider as: In Progress first, then priority, then
|
|
174
|
+
ordinal, then id (the CLI sorts by priority+ordinal but does not put In
|
|
175
|
+
Progress first — verified: `--ready --sort priority` returned the In
|
|
176
|
+
Progress task first here only because its ordinal is lower);
|
|
177
|
+
- `Letsdo::Loop` stays generic; no prompt or `Agent` signature change.
|
|
178
|
+
|
|
179
|
+
Pros: small, no agent-contract change, testable against the fake backlog, uses
|
|
180
|
+
documented CLI features. Cons: the agent still re-selects inside its prompt,
|
|
181
|
+
so letsdo's batch order is only advisory and the failure/accounting defect in
|
|
182
|
+
the risks section narrows but is not eliminated.
|
|
183
|
+
|
|
184
|
+
### B. Selector + task injection (medium)
|
|
185
|
+
|
|
186
|
+
Build on A: a pure `Letsdo::Selection` object computes the next runnable task,
|
|
187
|
+
and letsdo passes it into the run so the agent does not re-select:
|
|
188
|
+
|
|
189
|
+
- `Agent#run(task:)` injects "your task for this run is `<id>` — `<title>`"
|
|
190
|
+
into the prompt;
|
|
191
|
+
- the prompts' "Choosing a task" section is replaced by "work the task letsdo
|
|
192
|
+
handed you" (the blocker/readiness logic moves into the selector);
|
|
193
|
+
- `AgentLoop` records retry outcomes against the injected task.
|
|
194
|
+
|
|
195
|
+
Pros: one run = one task becomes enforced, not advisory; retry accounting and
|
|
196
|
+
starvation are fixed; per-agent `types:`/`skills:` front matter routing becomes
|
|
197
|
+
possible. Cons: changes the agent contract and both prompt files; needs its own
|
|
198
|
+
decision, prompt tests and a transition for custom prompts that still
|
|
199
|
+
self-select.
|
|
200
|
+
|
|
201
|
+
### C. Central scheduler (large, not recommended)
|
|
202
|
+
|
|
203
|
+
A coordinator process with a shared queue, atomic claim/lock, per-agent load
|
|
204
|
+
balancing and preemption.
|
|
205
|
+
|
|
206
|
+
Pros: true multi-agent scheduling. Cons: introduces shared mutable state,
|
|
207
|
+
concurrency, a new failure domain and an operational component; contradicts the
|
|
208
|
+
project's stated design ("multiple agents run as separate processes ... they
|
|
209
|
+
coordinate through the shared backlog — nothing else in common", README). Not
|
|
210
|
+
justified by any observed failure.
|
|
211
|
+
|
|
212
|
+
## Recommendation
|
|
213
|
+
|
|
214
|
+
**Build it — to scope A.** Add deterministic ordering and readiness to the
|
|
215
|
+
provider: `--ready --sort priority`, the missing normalized fields, and the
|
|
216
|
+
explicit In-Progress→priority→ordinal→id order. It is small, low-risk, uses
|
|
217
|
+
existing backlog CLI features, needs no prompt or agent-contract change, and
|
|
218
|
+
removes the common divergence between the order letsdo reads and the order the
|
|
219
|
+
agent picks. **Implemented in TASK-95** (`Providers::Backlog#command_line` and
|
|
220
|
+
`#sort`, `Providers::Task` fields).
|
|
221
|
+
|
|
222
|
+
Record scopes **B and C as out of scope for now**. B (task injection into the
|
|
223
|
+
prompt) is the correct way to fully enforce the batch and fix retry
|
|
224
|
+
misattribution, but it changes the agent contract and both prompt files and
|
|
225
|
+
deserves a separate decision once A is in use. C (central scheduler) is
|
|
226
|
+
explicitly rejected: it conflicts with the backlog-only, one-process-per-agent
|
|
227
|
+
architecture and is not justified by current pain.
|
|
228
|
+
|
|
229
|
+
The follow-up implementation task for scope A is tracked as TASK-95
|
|
230
|
+
(`Deterministic task selection: priority-sorted, ready-only provider batch`).
|
|
231
|
+
|
|
232
|
+
## Out of scope
|
|
233
|
+
|
|
234
|
+
- Task injection into the agent prompt and the prompt rewrite (scope B).
|
|
235
|
+
- Per-agent `types:`/`skills:` front-matter routing (needs B's injection to be
|
|
236
|
+
enforceable).
|
|
237
|
+
- Cross-agent load balancing, shared queue, atomic claim/lock, preemption
|
|
238
|
+
(scope C).
|
|
239
|
+
- Subtask-level selection and milestone-driven scheduling.
|