loki-mode 9.17.0 → 9.18.2

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
@@ -0,0 +1,174 @@
1
+ # Stale-State Audit
2
+
3
+ Audit for siblings of the `loki.pgid` session-killer (fixed in 4792b521).
4
+
5
+ **The shape being hunted:** a file recording a PID, PGID, port, lock, or session
6
+ id, written by one run, NOT removed on abnormal exit, later TRUSTED by another
7
+ run after the OS recycled the identifier.
8
+
9
+ **Method:** grep for the write, find its `rm -f`, check whether a trap covers
10
+ INT/TERM/crash, then check whether the reader proves the record is still
11
+ current. Every candidate was measured on this host rather than judged by
12
+ reading, because the pgid finding was credible on evidence (155h/202h orphans),
13
+ not on the code looking wrong.
14
+
15
+ Scope: read-only across the repo; fix applied to one file (`autonomy/run.sh`).
16
+
17
+ ---
18
+
19
+ ## Measured state of this host
20
+
21
+ Run from the repo root on 2026-08-08:
22
+
23
+ | File | Contents | Live? | Age |
24
+ |---|---|---|---|
25
+ | `~/.loki/dashboard/dashboard.pid` | absent | n/a | n/a |
26
+ | `.loki/dashboard/dashboard.pid` | `87992` | **DEAD** | **mtime Jul 31 (8 days)** |
27
+ | `$TMPDIR/loki-local-ci.lock` | `61876` | LIVE (real local-ci) | current |
28
+ | `.loki/app-runner/app.pid` | absent | n/a | n/a |
29
+
30
+ The dashboard pid file is the same measured shape as the pgid orphans: a dead
31
+ identifier, days old, still on disk, still trusted by a code path that kills.
32
+
33
+ ---
34
+
35
+ ## Findings, ranked by blast radius
36
+
37
+ ### FINDING 1 (HIGH -- fixed here): shared dashboard pid is killed unverified
38
+
39
+ **Evidence.** `autonomy/run.sh:1597-1600` reads a pid from
40
+ `~/.loki/dashboard/dashboard.pid` and sends it `kill` then `kill -9` with no
41
+ liveness and no identity check.
42
+
43
+ Removal of that file happens ONLY on explicit stop paths -- `run.sh:16608`,
44
+ `:16658`, `:17399`, and `autonomy/loki:7078-7079`. **No trap covers a crash or
45
+ Ctrl+C.** That is precisely why the measured copy in this checkout is 8 days
46
+ stale with a dead pid.
47
+
48
+ **Reachable:** yes. Called from `cleanup` on both stop paths (`run.sh:25311`,
49
+ `:25384`).
50
+
51
+ **Blast radius: the worst class.** The file lives under `~/.loki`, so it is
52
+ machine-global -- the victim need not belong to this project. A recycled pid
53
+ names an unrelated live process and receives `kill -9`.
54
+
55
+ **Why `kill -0` would not have fixed it:** a recycled pid IS alive, so a
56
+ liveness check passes. This is the identical insufficiency the `loki.pgid`
57
+ self-check had. The guard has to check IDENTITY.
58
+
59
+ **Fix applied.** `_loki_pid_looks_like_dashboard()` at `autonomy/run.sh:1526`,
60
+ gating the kill at `:1597`. It mirrors `_app_runner_pid_is_ours`
61
+ (`app-runner.sh:242`) and **fails OPEN**: when `ps` reports nothing we signal
62
+ exactly as before, so a legitimate dashboard is never left running by this
63
+ check. The only behavior change is refusing to kill a process positively
64
+ identified as not a dashboard.
65
+
66
+ Deliberately NOT done: adding a trap to remove the pid file on crash. That is a
67
+ larger change across the dashboard lifecycle, and the identity guard already
68
+ makes a stale file harmless at the point where it does damage. The stale file
69
+ still being present is untidy, not dangerous, once the killer verifies.
70
+
71
+ **Test:** `tests/test-stale-dashboard-pid.sh`, 6/6, mutation-verified below.
72
+
73
+ ### FINDING 2 (MEDIUM -- reported, not fixed): `app_runner_stop` group-kills unverified
74
+
75
+ **Evidence.** `app-runner.sh:1407` falls back to reading `app.pid` from disk,
76
+ then `:1456` sends `kill -TERM "-$_APP_RUNNER_PID"` -- a **process-group**
77
+ signal -- without an identity check.
78
+
79
+ The repo already has the right guard: `_app_runner_pid_is_ours`
80
+ (`app-runner.sh:242`). It is called at `:1657` and `:1829` but **not** on the
81
+ stop path. `app-runner.sh` has **no trap at all** (`grep -n "trap " ` returns
82
+ nothing), so `app.pid` survives a crash exactly like the pgid file did.
83
+
84
+ **Blast radius:** higher per-hit than Finding 1 (a group signal reaches a whole
85
+ tree) but narrower reach: the file is project-local (`.loki/app-runner/`), not
86
+ machine-global, and was absent on this host. Not fixed because it is a third
87
+ file and the lead scoped this task to one; it needs its own change adding
88
+ `_app_runner_pid_is_ours` to the stop path.
89
+
90
+ ### FINDING 3 (LOW -- reported, not fixed): local-ci lock guard is a no-op under `/bin/bash`
91
+
92
+ **Evidence.** `scripts/local-ci.sh:104`:
93
+
94
+ ```
95
+ trap '[ "$$" = "$_lci_owner" ] && rm -f "$_lci_lock" 2>/dev/null || true' EXIT
96
+ ```
97
+
98
+ The scar comment above it (`:99-103`) correctly diagnoses that a bare EXIT trap
99
+ fires in every subshell and that a finishing child deleted the parent's lock.
100
+ **But `$$` does not change in a bash subshell, so this guard does not
101
+ discriminate.** Verified:
102
+
103
+ ```
104
+ $ bash -c 'p=$$; ( [ "$$" = "$p" ] && echo same )'
105
+ same
106
+ ```
107
+
108
+ The correct discriminators are `BASHPID` (bash 4+) or `BASH_SUBSHELL` (works in
109
+ 3.2). Measured on this host: the script is `#!/usr/bin/env bash` which resolves
110
+ to **bash 5.3**, where `BASHPID` is available. But `/bin/bash` here is **3.2**,
111
+ where `BASHPID` is unset and `BASH_SUBSHELL` is the version-safe choice.
112
+
113
+ **Blast radius: benign** by the lead's own criterion. Worst case is two
114
+ concurrent local-ci runs starving each other into phantom failures -- costly in
115
+ time and trust, but it kills nothing. Left unfixed as out of scope; the one-line
116
+ change is `[ "${BASH_SUBSHELL:-0}" = 0 ]`.
117
+
118
+ Note on the sibling fix already shipped: `_loki_remove_pgid_file`
119
+ (`run.sh:25189`) uses `${BASHPID:-$$}`. Under bash 3.2 that collapses to `$$`
120
+ and the guard degrades to the no-op. **The degradation direction is safe** -- a
121
+ subshell deletes the pgid file early, the reap then finds no file and skips, so
122
+ an orphan survives and nobody gets killed. `run.sh` is also `#!/usr/bin/env
123
+ bash` (5.3 here), so this is a portability note, not a live defect. Worth
124
+ knowing: the mutation test for that guard asserts *source text* contains
125
+ `BASHPID`, so it proves the code is present, not that it behaves correctly on a
126
+ 3.2 host.
127
+
128
+ ---
129
+
130
+ ## Already guarded -- do not re-audit
131
+
132
+ Checked and found correct. Listed explicitly so this ground is not covered
133
+ twice.
134
+
135
+ - **`autonomy/lib/lock.sh:56-67`** -- `_loki_lock_is_stale` requires the
136
+ sentinel PID to be dead **AND** mtime > 30s. Both conditions, not either.
137
+ Correct.
138
+ - **`cleanup_orphan_pids` (`run.sh:2300+`)** -- reaps only on liveness AND
139
+ parent death AND (for wrappers) idle-past-budget AND no live engine child.
140
+ Also self-skips `$$`. Correct, and notably stricter than the pgid reap was.
141
+ - **`_app_runner_pid_is_ours` (`app-runner.sh:242`)** -- a real identity token
142
+ captured post-exec, failing open on a missing token so a live app is never
143
+ falsely killed. Correct where it is called; see Finding 2 for where it is not.
144
+ - **`_app_runner_collect_descendants` (`app-runner.sh:289`)** -- refuses pid
145
+ 0/1 and walks parent-child links from our own pid only, so it structurally
146
+ cannot signal outside our subtree. Correct.
147
+ - **`status.ts:266-270` and the bash status reader (`loki:5092`)** -- both do
148
+ `os.kill(pid, 0)` before reporting. A wrong answer only mislabels a URL in
149
+ status output. Benign even when stale.
150
+ - **The `CLEAR`/`KEEP` registry check (`run.sh:1545-1560`)** -- correctly
151
+ refuses to tear the shared dashboard down while any other project holds a
152
+ live pid. This gates Finding 1's call site; the defect was the missing check
153
+ on the pid itself, not this decision.
154
+
155
+ ---
156
+
157
+ ## Mutation verification (Finding 1's fix)
158
+
159
+ Each guard was reverted individually; the test must go red, and only on its own
160
+ assertion.
161
+
162
+ | Mutation | Before | After | Assertion killed |
163
+ |---|---|---|---|
164
+ | Identity check downgraded to bare `kill -0` | 6 pass / 0 fail | 5 / 1 | "a live non-dashboard process is refused" |
165
+ | Call site ungated (guard present but unused) | 6 / 0 | 5 / 1 | "the kill is gated on the identity guard" |
166
+ | `pid > 1` check removed | 6 / 0 | 5 / 1 | "pid 0/1, empty and malformed inputs refused" |
167
+
168
+ Restored after each. Final: `bash -n autonomy/run.sh` clean,
169
+ `tests/test-stale-dashboard-pid.sh` 6/6, `tests/test-pgid-stale-reap.sh` 8/8
170
+ (unaffected).
171
+
172
+ The first mutation is the important one: it replaces the identity check with
173
+ exactly the insufficient guard (`kill -0`) that a reviewer would most likely
174
+ propose, and the test catches it.
@@ -25,6 +25,8 @@ different answer on your hardware.
25
25
  | FULL local gate | 23 to 26 minutes | `LOCAL_CI_TIER=full bash scripts/local-ci.sh`, 166 checks, five runs on an M-series Mac |
26
26
  | FAST local gate | about 1 minute | `bash scripts/local-ci.sh`, the documented pre-push tier |
27
27
  | Shell suite, sharded | 352 seconds | 323 suites, `LOCAL_CI_SHARDS=4`; 1440 seconds serial, so 4.1x |
28
+ | `loki verify` | 77 seconds | `loki verify HEAD~1` on this repository, 5 changed files; dominated by the project's own test suite, not by our gates |
29
+ | of which, shipped-vs-dev CVE split | 418 ms | one extra `npm audit --omit=dev`; 0.5% of the run, and the reason the audit finding can say whether a CVE reaches users |
28
30
  | `loki outcomes` | under 1 second per receipt | `git blame` and one `git log` per changed file |
29
31
  | `loki intent status` | milliseconds | hash comparison against `.loki/spec/spec.lock` |
30
32
  | Agent readiness | milliseconds | filesystem checks only, no network |
@@ -50,6 +52,24 @@ non-forgeability requires the signed path
50
52
  (`LOKI_PROOF_GPG_KEY`, see [SIGNED-RECEIPTS.md](SIGNED-RECEIPTS.md)). We removed
51
53
  our own "non-forgeable" claim in v7.111.0 after finding it false on that path.
52
54
 
55
+ **The tests gate depends on correctly identifying YOUR test runner, and it has
56
+ been wrong before.** Until 2026-08-07 the runner was chosen by grepping
57
+ `package.json` for `"jest"` / `"vitest"` / `"mocha"`, which matches a
58
+ **devDependency**. This repository is the case that exposed it: jest is a
59
+ devDependency with no jest config while `scripts.test` runs `bash -n` plus
60
+ `node --test`, so verify ran jest, jest globbed 895 files that are not jest
61
+ tests, and `loki verify` returned BLOCKED on a clean tree -- permanently, for a
62
+ defect that did not exist. It now reads `scripts.test` with a JSON parser and
63
+ runs what the project declares.
64
+
65
+ We record this rather than quietly fixing it because a false BLOCK is the more
66
+ damaging direction of that error: a gate that cries wolf on every run trains
67
+ you to ignore the verdict, which costs more than the gate ever earned. If the
68
+ tests gate reports a runner you do not use, that is a bug in our detection, not
69
+ a finding about your code -- the runner is named in
70
+ `.loki/verify/evidence.json` under `deterministic_gates[].runner` so you can
71
+ check which one it picked.
72
+
53
73
  **Only four of the eight quality gates are agent-independent.** Static analysis,
54
74
  mock-integrity, test-mutation and documentation coverage do not ask a model
55
75
  anything. The other four involve model judgment and are labelled ASSESSMENTS
@@ -95,179 +115,39 @@ above less trustworthy:
95
115
  LOCAL_CI_TIER=full bash scripts/local-ci.sh # the full gate, timed
96
116
  loki outcomes --json # post-merge outcomes, or UNKNOWN with reasons
97
117
  loki intent status --json # spec-vs-intent drift
98
- loki proof verify <id> # re-hash a receipt, exit 1 on tamper
118
+ loki proof verify <id> # re-hash a receipt (see below)
99
119
  bash tests/test-competitor-verify-surface.sh # the competitor CLI measurement
100
120
  ```
101
121
 
102
- If a number here does not reproduce on your machine, that is a defect and we
103
- want the report.
104
-
105
- ## How we compare, and what we cannot measure
106
-
107
- A goal was set to be "2-10x better than factory.ai, cognition devin, 8090.ai
108
- and replit". This section reports what is measurable and refuses the rest.
109
-
110
- ### What is NOT benchmarked, and why
111
-
112
- None of those four products has a runnable local arm. Checked on this machine:
113
-
114
- ```
115
- droid NOT installed devin NOT installed
116
- replit NOT installed
117
- claude on PATH aider on PATH
118
- codex on PATH loki on PATH
119
- ```
120
-
121
- Devin and 8090 are hosted services with no CLI. Factory's droid and Replit run
122
- cloud-side. `benchmarks/bench/adapters/` can drive a competing arm as a
123
- subprocess (`claude_code.py` invokes `claude -p` live), but it cannot drive a
124
- product that has no local binary.
125
-
126
- **So there is no "2-10x vs Factory/Devin/8090/Replit" number here, and any such
127
- figure elsewhere should be treated as unearned.** Publishing one would be the
128
- same error as sourcing a colour token from a frontend that does not ship: a
129
- number that looks authoritative and measures something else.
122
+ ### Prove the tamper detection yourself, in three commands
130
123
 
131
- ### What IS measured: the verification surface
124
+ The claim is narrow and worth stating exactly: editing a receipt's recorded
125
+ facts is DETECTED. Run this against any receipt in `.loki/proofs/`:
132
126
 
133
- `tests/test-competitor-verify-surface.sh`, run on this machine:
134
-
135
- > **5 installed competitor CLIs. 0 expose an output-verification command.**
136
-
137
- That is not a multiplier, it is a category. The comparison is not "our
138
- verification is faster" but "there is nothing on the other side to compare
139
- against". Re-run it yourself; it names each CLI it checked and SKIPs the ones
140
- it could not find rather than counting them as absent.
141
-
142
- ### What the competitors say about verification, in their own words
143
-
144
- Verbatim, with the file each came from, so every line is checkable against the
145
- scraped corpora:
146
-
147
- | Source | Their words |
148
- |---|---|
149
- | `factory_ai/docs.factory.ai_missions_overview.md` | "**How do you maximize correctness?** Long-running plans accumulate errors." -- published as an OPEN QUESTION |
150
- | `factory_ai/docs.factory.ai_missions_overview.md` | "Without it, the mission **cannot reliably verify its own work**" (Missions require repo readiness Level 4+) |
151
- | `devin_cognition_ai/docs.devin.ai_admin_security.md.md` | "it can still experience **hallucinations, introduce bugs into code**, or suggest insecure code" |
152
- | `8090_ai/www.8090.ai_terms-of-service.md` | "**HUMAN REVIEW AND VERIFICATION OF ALL OUTPUT**" required, while the same ToS caps liability at "FIFTY US DOLLARS" and disclaims "ACCURACY" |
153
-
154
- Factory's is the most honest of the four: they name verification as an open
155
- research question rather than a solved feature, and they state that
156
- self-verification is a property of the ENVIRONMENT, not of the agent. We agree,
157
- which is why `loki readiness` measures the repo and not the model.
158
-
159
- ### The claim we will actually defend
160
-
161
- Not a multiplier. A receipt you can recompute:
162
-
163
- ```
164
- loki proof verify <id> # re-hash the receipt; exit 1 on tamper
165
- loki outcomes --json # anchored, or UNKNOWN with a named reason
127
+ ```bash
128
+ ID=$(ls .loki/proofs | head -1)
129
+ V() { loki proof verify "$ID" --json | python3 -c 'import json,sys;print(json.load(sys.stdin)["hash_ok"])'; }
130
+ V # True
131
+ python3 -c "import json;p='.loki/proofs/$ID/proof.json';d=json.load(open(p));d['files_changed']={'count':999999};json.dump(d,open(p,'w'))"
132
+ V # False
166
133
  ```
167
134
 
168
- Against "human review and verification of all output" and "cannot reliably
169
- verify its own work", an anchored receipt is a categorical difference. It is
170
- also falsifiable: if `loki outcomes` reports UNKNOWN, we say UNKNOWN. On this
171
- repo it reported ANCHORED 0 of 9 for weeks, and the fix
172
- (`facts.git.base_sha` was empty) is in the history.
135
+ Measured on this repository: `True` -> `False` -> `True` after restoring.
173
136
 
174
- ### Honest limits on this section
137
+ **Read `hash_ok`, not `ok`.** They answer different questions and conflating
138
+ them produces a false alarm. `hash_ok` is integrity: do the recorded facts still
139
+ hash to the recorded digest. `ok` also folds in `tree_drift`, which is true
140
+ whenever the working tree has moved since the receipt was written -- so an
141
+ untampered receipt from last month correctly reports `ok: false` with
142
+ `hash_ok: true`. An earlier version of this document said "exit 1 on tamper",
143
+ which is wrong in exactly that way: a drifted-but-intact receipt also exits 1.
175
144
 
176
- - The verify-surface count is a check for a COMMAND, not for internal
177
- verification a product may do without exposing it. A hosted product could
178
- verify server-side and expose nothing to a CLI.
179
- - It measures what is installed HERE. A CLI absent from this machine is
180
- reported as SKIP, never as a competitor that lacks the feature.
181
- - Nothing here measures build quality, speed, or cost against those four.
182
- Those comparisons are not available to us and are not claimed.
145
+ **What this does NOT establish.** Integrity is not provenance. On the unsigned
146
+ path a party who rewrites the facts AND recomputes the digest passes this check
147
+ -- see the forgeability limit above. Provenance requires the signed path
148
+ (`LOKI_PROOF_GPG_KEY`, [SIGNED-RECEIPTS.md](SIGNED-RECEIPTS.md)), and the remote
149
+ client reports the two separately for that reason: VERIFIED, UNSIGNED,
150
+ UNCHECKED and TAMPERED are four distinct verdicts, never collapsed.
183
151
 
184
- ### Measured: what the harness is worth, model held constant
185
-
186
- The one comparison we CAN run. `benchmarks/bench/matrix.sh` defines two
187
- configs against the same task and the same model:
188
-
189
- - `baseline` -- raw model, minimal orchestration (in-repo comment: "the
190
- Replit/Cursor mode")
191
- - `full` -- the harness: council, code review, self-heal, auto-tune
192
-
193
- Paired on identical tasks, both arms on `haiku`, from
194
- `benchmarks/bench/results/`:
195
-
196
- | Task | harness (`full`) | raw model (`baseline`) | verdict |
197
- |---|---|---|---|
198
- | `hard-2-ledger` | **4/4**, $0.89 | **0/4**, $0.19 | harness wins outright |
199
- | `hard-1-order-api` | 3 runs, 1.00, $0.58 | 1 run, 1.00, $0.20 | harness cost 2.8x for nothing |
200
- | `multifail-1-two-modules` | 2 runs, 1.00, $0.27 | 1 run, 1.00, $0.14 | harness cost 1.9x for nothing |
201
-
202
- **Two of the three paired tasks are UNFAVOURABLE to the harness.** That is the
203
- honest headline, and it is narrower than "the harness is worth more than the
204
- model":
205
-
206
- **The harness matters on hard tasks and is pure overhead on easy ones.**
207
-
208
- On `hard-2-ledger` the raw model never finished the task across 4 trials and
209
- the harness finished it every time. That is not a percentage improvement; the
210
- baseline success rate is zero. It cost 4.7x more per run and produced a working
211
- result instead of nothing.
212
-
213
- On the other two, both arms succeed and the harness simply costs 1.9x-2.8x
214
- more. We publish those rows because the direction is unfavourable to us, and a
215
- benchmark table that only survives its favourable rows is an advertisement.
216
-
217
- **Two further baseline cells were attempted and produced NO data.**
218
- `tokenheavy-1-crm` (2 trials) and a wider `hard-1-order-api` (3 trials) both hit
219
- the 1200s cell timeout, which caps all trials of a cell together. The harness
220
- wrote no result file and the runner said so:
221
-
222
- > WARNING: cell haiku-baseline / tokenheavy-1-crm wrote NO new result (likely
223
- > timed out at 1200s). This cell is MISSING from the report.
224
-
225
- They are absent from the table rather than counted as failures. A timeout is
226
- not evidence the raw model cannot do the task; it is evidence we did not
227
- measure it. Reporting them as baseline losses would have made the harness look
228
- better on data that does not exist.
229
-
230
- An earlier version of this section reported only the first two tasks and read
231
- as a stronger claim than the data supported. The `multifail` baseline arm was
232
- run afterwards specifically to test whether the claim would survive more data.
233
- It did not survive intact, and the table was corrected rather than the
234
- measurement dropped.
235
-
236
- Aggregate across all recorded cells:
237
-
238
- | Cell | n | success (median) | cost (median) |
239
- |---|---|---|---|
240
- | `haiku` + harness | 11 | 1.00 | $0.54 |
241
- | `opus` + baseline | 5 | 1.00 | $0.83 |
242
- | `haiku` + baseline | 8 | 0.50 | $0.19 |
243
-
244
- The cheap model WITH the harness matches the expensive model without it, at
245
- 35% lower cost. Treat that as directional, not as a headline: the task sets
246
- differ between those three cells, which is exactly why the paired table above
247
- is the one that carries the argument.
248
-
249
- **Limits of this measurement, stated so it cannot be over-read:**
250
-
251
- - Two paired tasks. `hard-1-order-api`'s baseline arm is n=1.
252
- - It measures OUR harness against OUR baseline config. `baseline` is a
253
- documented stand-in for "raw model, minimal orchestration", NOT a
254
- measurement of Replit, Cursor, or any other product.
255
- - Success is a held-out acceptance exit code, not a judgement of code quality.
256
- - Reproduce it: `LOKI_BENCH_SPEND_APPROVED=1 bash benchmarks/bench/matrix.sh pilot`.
257
- The spend interlock is default-deny on purpose; a benchmark that starts a
258
- paid tool without an explicit opt-in is how a surprise bill happens.
259
-
260
- **The exclusion rule, observed on a live run rather than asserted.** A fresh
261
- `haiku-full` trial on `hard-1-order-api` (2026-08-07) hit the 1200s adapter
262
- timeout. The held-out grader inspected the workdir and returned
263
- `success: true` -- there WAS a passing artifact. The harness still recorded
264
- `measured: false`, `unmeasured_reasons: ["adapter exit_status=timeout"]`, and
265
- excluded the trial from k/N, with this note in the result file:
266
-
267
- > "the held-out grader's own verdict on whatever was in the workdir. Real
268
- > evidence about the ARTIFACT; NOT evidence that a run happened."
269
-
270
- A naive harness scores that 1/1. Ours scores it 0/0 and says why, so the
271
- `hard-1-order-api` row above stayed at n=3 instead of being inflated to n=4 by
272
- a run that never finished. The number in this document is smaller because of
273
- that rule, which is the point of having it.
152
+ If a number here does not reproduce on your machine, that is a defect and we
153
+ want the report.