loki-mode 9.16.0 → 9.17.0
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/SKILL.md +2 -2
- package/VERSION +1 -1
- package/autonomy/issue-providers.sh +12 -9
- package/autonomy/lib/agent_readiness.py +79 -1
- package/autonomy/lib/gate_policy.py +166 -0
- package/autonomy/lib/outcome_ledger.py +122 -0
- package/autonomy/lib/proof-generator.py +71 -4
- package/autonomy/loki +83 -12
- package/autonomy/notify.sh +70 -1
- package/autonomy/queue-consumer.sh +290 -18
- package/autonomy/run.sh +11 -1
- package/completions/_loki +1 -0
- package/completions/loki.bash +2 -2
- package/dashboard/__init__.py +1 -1
- package/dashboard/server.py +168 -1
- package/dashboard/static/index.html +248 -55
- package/docs/QUEUE-OPERATIONS.md +107 -0
- package/docs/VERIFICATION-COST.md +170 -0
- package/loki-ts/dist/loki.js +287 -293
- package/mcp/__init__.py +1 -1
- package/package.json +1 -1
- package/plugins/loki-mode/.claude-plugin/plugin.json +1 -1
|
@@ -0,0 +1,107 @@
|
|
|
1
|
+
# Queue operations: what happens when a worker dies
|
|
2
|
+
|
|
3
|
+
Written for the person on call when builds stop coming out.
|
|
4
|
+
|
|
5
|
+
## The short version
|
|
6
|
+
|
|
7
|
+
A build that vanished is recoverable. Run:
|
|
8
|
+
|
|
9
|
+
```bash
|
|
10
|
+
bash autonomy/queue-consumer.sh --reap
|
|
11
|
+
```
|
|
12
|
+
|
|
13
|
+
It requeues in-flight items whose worker died. It never touches a build that is
|
|
14
|
+
still running, and it never deletes anything.
|
|
15
|
+
|
|
16
|
+
## What the queue guarantees, and what it does not
|
|
17
|
+
|
|
18
|
+
| | redis backend | file backend |
|
|
19
|
+
|---|---|---|
|
|
20
|
+
| Delivery | at-least-once (default) | at-least-once |
|
|
21
|
+
| Crash leaves the item | in `<key>:processing` | in `processing/` |
|
|
22
|
+
| Automatic recovery | `--reap` | see below |
|
|
23
|
+
| Claim time recorded | `<key>:claims` hash | file mtime |
|
|
24
|
+
|
|
25
|
+
**A crashed worker does not lose the build.** It used to: the redis backend was
|
|
26
|
+
at-most-once (`LPOP`-then-run), so a worker killed between the pop and the
|
|
27
|
+
finish took the item with it and nothing anywhere recorded that it existed.
|
|
28
|
+
Proven against a real redis 8.6.3, then fixed with `LMOVE`, which pops and
|
|
29
|
+
records in-flight in one atomic step.
|
|
30
|
+
|
|
31
|
+
`LOKI_QUEUE_ACK=0` restores the old lossy behaviour if you are running a broker
|
|
32
|
+
that already guarantees delivery and do not want a second in-flight record.
|
|
33
|
+
|
|
34
|
+
## Recovering a stalled fleet
|
|
35
|
+
|
|
36
|
+
Symptom: items are queued, workers are up, and nothing completes.
|
|
37
|
+
|
|
38
|
+
```bash
|
|
39
|
+
# 1. What is stuck?
|
|
40
|
+
redis-cli -u "$LOKI_QUEUE_URL" LLEN loki-builds # waiting
|
|
41
|
+
redis-cli -u "$LOKI_QUEUE_URL" LLEN loki-builds:processing # claimed, in flight
|
|
42
|
+
|
|
43
|
+
# 2. Requeue anything whose worker is gone.
|
|
44
|
+
bash autonomy/queue-consumer.sh --reap
|
|
45
|
+
# -> reap: requeued 2, left 1 in flight (timeout 7200s)
|
|
46
|
+
```
|
|
47
|
+
|
|
48
|
+
The reaper reports what it did. "left 1 in flight" means one item's claim is
|
|
49
|
+
still inside the visibility timeout, so it is treated as a running build and not
|
|
50
|
+
touched.
|
|
51
|
+
|
|
52
|
+
### The timeout is the safety knob
|
|
53
|
+
|
|
54
|
+
`LOKI_QUEUE_VISIBILITY_SEC` (default 7200) is how long an item may be in flight
|
|
55
|
+
before the reaper assumes its worker died.
|
|
56
|
+
|
|
57
|
+
**It must exceed your longest legitimate build.** Set it too low and the reaper
|
|
58
|
+
requeues work that is still running, and the user gets the same build twice.
|
|
59
|
+
That is the only way this tool can hurt you, so it errs long by default.
|
|
60
|
+
|
|
61
|
+
### An item with no claim is reaped immediately
|
|
62
|
+
|
|
63
|
+
If a worker dies between claiming the item and stamping the claim time, the
|
|
64
|
+
item sits in `processing` with no entry in `<key>:claims`. The reaper treats
|
|
65
|
+
that as infinitely old, not infinitely young -- it is exactly the case that must
|
|
66
|
+
be recoverable, and erring the other way would strand the items this exists to
|
|
67
|
+
rescue.
|
|
68
|
+
|
|
69
|
+
## Running the reaper unattended
|
|
70
|
+
|
|
71
|
+
A CronJob is the intended shape. It is a separate mode rather than part of the
|
|
72
|
+
consume loop on purpose: a consumer that reaped on every poll would race its own
|
|
73
|
+
peers on a busy queue.
|
|
74
|
+
|
|
75
|
+
```yaml
|
|
76
|
+
schedule: "*/15 * * * *"
|
|
77
|
+
command: ["bash", "autonomy/queue-consumer.sh", "--reap"]
|
|
78
|
+
```
|
|
79
|
+
|
|
80
|
+
## What is NOT claimed
|
|
81
|
+
|
|
82
|
+
- **The file backend has no automatic reaper.** `--reap` is redis-only today; a
|
|
83
|
+
crashed worker on the file backend leaves its item in `processing/` for a
|
|
84
|
+
human to move back to `pending/`. The item is not lost, but nothing recovers
|
|
85
|
+
it on its own.
|
|
86
|
+
- **No dead-letter queue.** An item that fails repeatedly is requeued
|
|
87
|
+
repeatedly. There is no automatic give-up-and-park.
|
|
88
|
+
- **`LREM` clears every equal copy.** Two identical specs queued twice both
|
|
89
|
+
clear on the first ack. Accepted deliberately: the alternative needs a
|
|
90
|
+
per-item token from the producer, and the failure mode here is one duplicate
|
|
91
|
+
re-drive rather than a lost build.
|
|
92
|
+
- **Only redis and file ship.** SQS, Pub/Sub, RabbitMQ and Kafka are
|
|
93
|
+
bring-your-own: override `queue.command` with your own consumer.
|
|
94
|
+
|
|
95
|
+
## Check any of this yourself
|
|
96
|
+
|
|
97
|
+
```bash
|
|
98
|
+
bash tests/test-queue-at-least-once.sh # 16 assertions; live half needs a broker
|
|
99
|
+
redis-server --port 6399 --save '' --appendonly no &
|
|
100
|
+
LOKI_TEST_REDIS_PORT=6399 bash tests/test-queue-at-least-once.sh
|
|
101
|
+
```
|
|
102
|
+
|
|
103
|
+
The live checks SKIP with a named reason when no broker is present rather than
|
|
104
|
+
passing on a mock. Both bugs found building this were invisible to a fake:
|
|
105
|
+
`redis-cli --no-raw` escapes the inner quotes of a JSON payload, which mangled
|
|
106
|
+
the item and made the ack a silent no-op. A green mock would have certified both
|
|
107
|
+
as working.
|
|
@@ -101,3 +101,173 @@ bash tests/test-competitor-verify-surface.sh # the competitor CLI measurement
|
|
|
101
101
|
|
|
102
102
|
If a number here does not reproduce on your machine, that is a defect and we
|
|
103
103
|
want the report.
|
|
104
|
+
|
|
105
|
+
## How we compare, and what we cannot measure
|
|
106
|
+
|
|
107
|
+
A goal was set to be "2-10x better than factory.ai, cognition devin, 8090.ai
|
|
108
|
+
and replit". This section reports what is measurable and refuses the rest.
|
|
109
|
+
|
|
110
|
+
### What is NOT benchmarked, and why
|
|
111
|
+
|
|
112
|
+
None of those four products has a runnable local arm. Checked on this machine:
|
|
113
|
+
|
|
114
|
+
```
|
|
115
|
+
droid NOT installed devin NOT installed
|
|
116
|
+
replit NOT installed
|
|
117
|
+
claude on PATH aider on PATH
|
|
118
|
+
codex on PATH loki on PATH
|
|
119
|
+
```
|
|
120
|
+
|
|
121
|
+
Devin and 8090 are hosted services with no CLI. Factory's droid and Replit run
|
|
122
|
+
cloud-side. `benchmarks/bench/adapters/` can drive a competing arm as a
|
|
123
|
+
subprocess (`claude_code.py` invokes `claude -p` live), but it cannot drive a
|
|
124
|
+
product that has no local binary.
|
|
125
|
+
|
|
126
|
+
**So there is no "2-10x vs Factory/Devin/8090/Replit" number here, and any such
|
|
127
|
+
figure elsewhere should be treated as unearned.** Publishing one would be the
|
|
128
|
+
same error as sourcing a colour token from a frontend that does not ship: a
|
|
129
|
+
number that looks authoritative and measures something else.
|
|
130
|
+
|
|
131
|
+
### What IS measured: the verification surface
|
|
132
|
+
|
|
133
|
+
`tests/test-competitor-verify-surface.sh`, run on this machine:
|
|
134
|
+
|
|
135
|
+
> **5 installed competitor CLIs. 0 expose an output-verification command.**
|
|
136
|
+
|
|
137
|
+
That is not a multiplier, it is a category. The comparison is not "our
|
|
138
|
+
verification is faster" but "there is nothing on the other side to compare
|
|
139
|
+
against". Re-run it yourself; it names each CLI it checked and SKIPs the ones
|
|
140
|
+
it could not find rather than counting them as absent.
|
|
141
|
+
|
|
142
|
+
### What the competitors say about verification, in their own words
|
|
143
|
+
|
|
144
|
+
Verbatim, with the file each came from, so every line is checkable against the
|
|
145
|
+
scraped corpora:
|
|
146
|
+
|
|
147
|
+
| Source | Their words |
|
|
148
|
+
|---|---|
|
|
149
|
+
| `factory_ai/docs.factory.ai_missions_overview.md` | "**How do you maximize correctness?** Long-running plans accumulate errors." -- published as an OPEN QUESTION |
|
|
150
|
+
| `factory_ai/docs.factory.ai_missions_overview.md` | "Without it, the mission **cannot reliably verify its own work**" (Missions require repo readiness Level 4+) |
|
|
151
|
+
| `devin_cognition_ai/docs.devin.ai_admin_security.md.md` | "it can still experience **hallucinations, introduce bugs into code**, or suggest insecure code" |
|
|
152
|
+
| `8090_ai/www.8090.ai_terms-of-service.md` | "**HUMAN REVIEW AND VERIFICATION OF ALL OUTPUT**" required, while the same ToS caps liability at "FIFTY US DOLLARS" and disclaims "ACCURACY" |
|
|
153
|
+
|
|
154
|
+
Factory's is the most honest of the four: they name verification as an open
|
|
155
|
+
research question rather than a solved feature, and they state that
|
|
156
|
+
self-verification is a property of the ENVIRONMENT, not of the agent. We agree,
|
|
157
|
+
which is why `loki readiness` measures the repo and not the model.
|
|
158
|
+
|
|
159
|
+
### The claim we will actually defend
|
|
160
|
+
|
|
161
|
+
Not a multiplier. A receipt you can recompute:
|
|
162
|
+
|
|
163
|
+
```
|
|
164
|
+
loki proof verify <id> # re-hash the receipt; exit 1 on tamper
|
|
165
|
+
loki outcomes --json # anchored, or UNKNOWN with a named reason
|
|
166
|
+
```
|
|
167
|
+
|
|
168
|
+
Against "human review and verification of all output" and "cannot reliably
|
|
169
|
+
verify its own work", an anchored receipt is a categorical difference. It is
|
|
170
|
+
also falsifiable: if `loki outcomes` reports UNKNOWN, we say UNKNOWN. On this
|
|
171
|
+
repo it reported ANCHORED 0 of 9 for weeks, and the fix
|
|
172
|
+
(`facts.git.base_sha` was empty) is in the history.
|
|
173
|
+
|
|
174
|
+
### Honest limits on this section
|
|
175
|
+
|
|
176
|
+
- The verify-surface count is a check for a COMMAND, not for internal
|
|
177
|
+
verification a product may do without exposing it. A hosted product could
|
|
178
|
+
verify server-side and expose nothing to a CLI.
|
|
179
|
+
- It measures what is installed HERE. A CLI absent from this machine is
|
|
180
|
+
reported as SKIP, never as a competitor that lacks the feature.
|
|
181
|
+
- Nothing here measures build quality, speed, or cost against those four.
|
|
182
|
+
Those comparisons are not available to us and are not claimed.
|
|
183
|
+
|
|
184
|
+
### Measured: what the harness is worth, model held constant
|
|
185
|
+
|
|
186
|
+
The one comparison we CAN run. `benchmarks/bench/matrix.sh` defines two
|
|
187
|
+
configs against the same task and the same model:
|
|
188
|
+
|
|
189
|
+
- `baseline` -- raw model, minimal orchestration (in-repo comment: "the
|
|
190
|
+
Replit/Cursor mode")
|
|
191
|
+
- `full` -- the harness: council, code review, self-heal, auto-tune
|
|
192
|
+
|
|
193
|
+
Paired on identical tasks, both arms on `haiku`, from
|
|
194
|
+
`benchmarks/bench/results/`:
|
|
195
|
+
|
|
196
|
+
| Task | harness (`full`) | raw model (`baseline`) | verdict |
|
|
197
|
+
|---|---|---|---|
|
|
198
|
+
| `hard-2-ledger` | **4/4**, $0.89 | **0/4**, $0.19 | harness wins outright |
|
|
199
|
+
| `hard-1-order-api` | 3 runs, 1.00, $0.58 | 1 run, 1.00, $0.20 | harness cost 2.8x for nothing |
|
|
200
|
+
| `multifail-1-two-modules` | 2 runs, 1.00, $0.27 | 1 run, 1.00, $0.14 | harness cost 1.9x for nothing |
|
|
201
|
+
|
|
202
|
+
**Two of the three paired tasks are UNFAVOURABLE to the harness.** That is the
|
|
203
|
+
honest headline, and it is narrower than "the harness is worth more than the
|
|
204
|
+
model":
|
|
205
|
+
|
|
206
|
+
**The harness matters on hard tasks and is pure overhead on easy ones.**
|
|
207
|
+
|
|
208
|
+
On `hard-2-ledger` the raw model never finished the task across 4 trials and
|
|
209
|
+
the harness finished it every time. That is not a percentage improvement; the
|
|
210
|
+
baseline success rate is zero. It cost 4.7x more per run and produced a working
|
|
211
|
+
result instead of nothing.
|
|
212
|
+
|
|
213
|
+
On the other two, both arms succeed and the harness simply costs 1.9x-2.8x
|
|
214
|
+
more. We publish those rows because the direction is unfavourable to us, and a
|
|
215
|
+
benchmark table that only survives its favourable rows is an advertisement.
|
|
216
|
+
|
|
217
|
+
**Two further baseline cells were attempted and produced NO data.**
|
|
218
|
+
`tokenheavy-1-crm` (2 trials) and a wider `hard-1-order-api` (3 trials) both hit
|
|
219
|
+
the 1200s cell timeout, which caps all trials of a cell together. The harness
|
|
220
|
+
wrote no result file and the runner said so:
|
|
221
|
+
|
|
222
|
+
> WARNING: cell haiku-baseline / tokenheavy-1-crm wrote NO new result (likely
|
|
223
|
+
> timed out at 1200s). This cell is MISSING from the report.
|
|
224
|
+
|
|
225
|
+
They are absent from the table rather than counted as failures. A timeout is
|
|
226
|
+
not evidence the raw model cannot do the task; it is evidence we did not
|
|
227
|
+
measure it. Reporting them as baseline losses would have made the harness look
|
|
228
|
+
better on data that does not exist.
|
|
229
|
+
|
|
230
|
+
An earlier version of this section reported only the first two tasks and read
|
|
231
|
+
as a stronger claim than the data supported. The `multifail` baseline arm was
|
|
232
|
+
run afterwards specifically to test whether the claim would survive more data.
|
|
233
|
+
It did not survive intact, and the table was corrected rather than the
|
|
234
|
+
measurement dropped.
|
|
235
|
+
|
|
236
|
+
Aggregate across all recorded cells:
|
|
237
|
+
|
|
238
|
+
| Cell | n | success (median) | cost (median) |
|
|
239
|
+
|---|---|---|---|
|
|
240
|
+
| `haiku` + harness | 11 | 1.00 | $0.54 |
|
|
241
|
+
| `opus` + baseline | 5 | 1.00 | $0.83 |
|
|
242
|
+
| `haiku` + baseline | 8 | 0.50 | $0.19 |
|
|
243
|
+
|
|
244
|
+
The cheap model WITH the harness matches the expensive model without it, at
|
|
245
|
+
35% lower cost. Treat that as directional, not as a headline: the task sets
|
|
246
|
+
differ between those three cells, which is exactly why the paired table above
|
|
247
|
+
is the one that carries the argument.
|
|
248
|
+
|
|
249
|
+
**Limits of this measurement, stated so it cannot be over-read:**
|
|
250
|
+
|
|
251
|
+
- Two paired tasks. `hard-1-order-api`'s baseline arm is n=1.
|
|
252
|
+
- It measures OUR harness against OUR baseline config. `baseline` is a
|
|
253
|
+
documented stand-in for "raw model, minimal orchestration", NOT a
|
|
254
|
+
measurement of Replit, Cursor, or any other product.
|
|
255
|
+
- Success is a held-out acceptance exit code, not a judgement of code quality.
|
|
256
|
+
- Reproduce it: `LOKI_BENCH_SPEND_APPROVED=1 bash benchmarks/bench/matrix.sh pilot`.
|
|
257
|
+
The spend interlock is default-deny on purpose; a benchmark that starts a
|
|
258
|
+
paid tool without an explicit opt-in is how a surprise bill happens.
|
|
259
|
+
|
|
260
|
+
**The exclusion rule, observed on a live run rather than asserted.** A fresh
|
|
261
|
+
`haiku-full` trial on `hard-1-order-api` (2026-08-07) hit the 1200s adapter
|
|
262
|
+
timeout. The held-out grader inspected the workdir and returned
|
|
263
|
+
`success: true` -- there WAS a passing artifact. The harness still recorded
|
|
264
|
+
`measured: false`, `unmeasured_reasons: ["adapter exit_status=timeout"]`, and
|
|
265
|
+
excluded the trial from k/N, with this note in the result file:
|
|
266
|
+
|
|
267
|
+
> "the held-out grader's own verdict on whatever was in the workdir. Real
|
|
268
|
+
> evidence about the ARTIFACT; NOT evidence that a run happened."
|
|
269
|
+
|
|
270
|
+
A naive harness scores that 1/1. Ours scores it 0/0 and says why, so the
|
|
271
|
+
`hard-1-order-api` row above stayed at n=3 instead of being inflated to n=4 by
|
|
272
|
+
a run that never finished. The number in this document is smaller because of
|
|
273
|
+
that rule, which is the point of having it.
|