loki-mode 9.17.0 → 9.18.2
This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
- package/SKILL.md +2 -2
- package/VERSION +1 -1
- package/autonomy/issue-providers.sh +0 -24
- package/autonomy/lib/agent_readiness.py +1 -79
- package/autonomy/lib/outcome_ledger.py +0 -122
- package/autonomy/lib/proof-generator.py +4 -71
- package/autonomy/lib/verdict.py +4 -25
- package/autonomy/loki +444 -264
- package/autonomy/notify.sh +1 -70
- package/autonomy/queue-consumer.sh +18 -290
- package/autonomy/run.sh +280 -189
- package/autonomy/trigger-server.py +518 -17
- package/autonomy/verify.sh +100 -12
- package/bin/loki +7 -1
- package/completions/_loki +0 -3
- package/completions/loki.bash +2 -2
- package/dashboard/__init__.py +1 -1
- package/dashboard/server.py +1 -168
- package/dashboard/static/index.html +222 -249
- package/docs/COMPETITIVE-SCORECARD.md +38 -0
- package/docs/COMPETITOR-DEPLOYMENT-MODELS.md +475 -0
- package/docs/DEPLOYMENT.md +542 -0
- package/docs/STALE-STATE-AUDIT.md +174 -0
- package/docs/VERIFICATION-COST.md +46 -166
- package/loki-ts/dist/loki.js +311 -305
- package/mcp/__init__.py +1 -1
- package/mcp/_sdk_loader.py +0 -25
- package/package.json +1 -1
- package/plugins/loki-mode/.claude-plugin/plugin.json +1 -1
- package/autonomy/lib/gate_policy.py +0 -166
- package/docs/QUEUE-OPERATIONS.md +0 -107
|
@@ -0,0 +1,174 @@
|
|
|
1
|
+
# Stale-State Audit
|
|
2
|
+
|
|
3
|
+
Audit for siblings of the `loki.pgid` session-killer (fixed in 4792b521).
|
|
4
|
+
|
|
5
|
+
**The shape being hunted:** a file recording a PID, PGID, port, lock, or session
|
|
6
|
+
id, written by one run, NOT removed on abnormal exit, later TRUSTED by another
|
|
7
|
+
run after the OS recycled the identifier.
|
|
8
|
+
|
|
9
|
+
**Method:** grep for the write, find its `rm -f`, check whether a trap covers
|
|
10
|
+
INT/TERM/crash, then check whether the reader proves the record is still
|
|
11
|
+
current. Every candidate was measured on this host rather than judged by
|
|
12
|
+
reading, because the pgid finding was credible on evidence (155h/202h orphans),
|
|
13
|
+
not on the code looking wrong.
|
|
14
|
+
|
|
15
|
+
Scope: read-only across the repo; fix applied to one file (`autonomy/run.sh`).
|
|
16
|
+
|
|
17
|
+
---
|
|
18
|
+
|
|
19
|
+
## Measured state of this host
|
|
20
|
+
|
|
21
|
+
Run from the repo root on 2026-08-08:
|
|
22
|
+
|
|
23
|
+
| File | Contents | Live? | Age |
|
|
24
|
+
|---|---|---|---|
|
|
25
|
+
| `~/.loki/dashboard/dashboard.pid` | absent | n/a | n/a |
|
|
26
|
+
| `.loki/dashboard/dashboard.pid` | `87992` | **DEAD** | **mtime Jul 31 (8 days)** |
|
|
27
|
+
| `$TMPDIR/loki-local-ci.lock` | `61876` | LIVE (real local-ci) | current |
|
|
28
|
+
| `.loki/app-runner/app.pid` | absent | n/a | n/a |
|
|
29
|
+
|
|
30
|
+
The dashboard pid file is the same measured shape as the pgid orphans: a dead
|
|
31
|
+
identifier, days old, still on disk, still trusted by a code path that kills.
|
|
32
|
+
|
|
33
|
+
---
|
|
34
|
+
|
|
35
|
+
## Findings, ranked by blast radius
|
|
36
|
+
|
|
37
|
+
### FINDING 1 (HIGH -- fixed here): shared dashboard pid is killed unverified
|
|
38
|
+
|
|
39
|
+
**Evidence.** `autonomy/run.sh:1597-1600` reads a pid from
|
|
40
|
+
`~/.loki/dashboard/dashboard.pid` and sends it `kill` then `kill -9` with no
|
|
41
|
+
liveness and no identity check.
|
|
42
|
+
|
|
43
|
+
Removal of that file happens ONLY on explicit stop paths -- `run.sh:16608`,
|
|
44
|
+
`:16658`, `:17399`, and `autonomy/loki:7078-7079`. **No trap covers a crash or
|
|
45
|
+
Ctrl+C.** That is precisely why the measured copy in this checkout is 8 days
|
|
46
|
+
stale with a dead pid.
|
|
47
|
+
|
|
48
|
+
**Reachable:** yes. Called from `cleanup` on both stop paths (`run.sh:25311`,
|
|
49
|
+
`:25384`).
|
|
50
|
+
|
|
51
|
+
**Blast radius: the worst class.** The file lives under `~/.loki`, so it is
|
|
52
|
+
machine-global -- the victim need not belong to this project. A recycled pid
|
|
53
|
+
names an unrelated live process and receives `kill -9`.
|
|
54
|
+
|
|
55
|
+
**Why `kill -0` would not have fixed it:** a recycled pid IS alive, so a
|
|
56
|
+
liveness check passes. This is the identical insufficiency the `loki.pgid`
|
|
57
|
+
self-check had. The guard has to check IDENTITY.
|
|
58
|
+
|
|
59
|
+
**Fix applied.** `_loki_pid_looks_like_dashboard()` at `autonomy/run.sh:1526`,
|
|
60
|
+
gating the kill at `:1597`. It mirrors `_app_runner_pid_is_ours`
|
|
61
|
+
(`app-runner.sh:242`) and **fails OPEN**: when `ps` reports nothing we signal
|
|
62
|
+
exactly as before, so a legitimate dashboard is never left running by this
|
|
63
|
+
check. The only behavior change is refusing to kill a process positively
|
|
64
|
+
identified as not a dashboard.
|
|
65
|
+
|
|
66
|
+
Deliberately NOT done: adding a trap to remove the pid file on crash. That is a
|
|
67
|
+
larger change across the dashboard lifecycle, and the identity guard already
|
|
68
|
+
makes a stale file harmless at the point where it does damage. The stale file
|
|
69
|
+
still being present is untidy, not dangerous, once the killer verifies.
|
|
70
|
+
|
|
71
|
+
**Test:** `tests/test-stale-dashboard-pid.sh`, 6/6, mutation-verified below.
|
|
72
|
+
|
|
73
|
+
### FINDING 2 (MEDIUM -- reported, not fixed): `app_runner_stop` group-kills unverified
|
|
74
|
+
|
|
75
|
+
**Evidence.** `app-runner.sh:1407` falls back to reading `app.pid` from disk,
|
|
76
|
+
then `:1456` sends `kill -TERM "-$_APP_RUNNER_PID"` -- a **process-group**
|
|
77
|
+
signal -- without an identity check.
|
|
78
|
+
|
|
79
|
+
The repo already has the right guard: `_app_runner_pid_is_ours`
|
|
80
|
+
(`app-runner.sh:242`). It is called at `:1657` and `:1829` but **not** on the
|
|
81
|
+
stop path. `app-runner.sh` has **no trap at all** (`grep -n "trap " ` returns
|
|
82
|
+
nothing), so `app.pid` survives a crash exactly like the pgid file did.
|
|
83
|
+
|
|
84
|
+
**Blast radius:** higher per-hit than Finding 1 (a group signal reaches a whole
|
|
85
|
+
tree) but narrower reach: the file is project-local (`.loki/app-runner/`), not
|
|
86
|
+
machine-global, and was absent on this host. Not fixed because it is a third
|
|
87
|
+
file and the lead scoped this task to one; it needs its own change adding
|
|
88
|
+
`_app_runner_pid_is_ours` to the stop path.
|
|
89
|
+
|
|
90
|
+
### FINDING 3 (LOW -- reported, not fixed): local-ci lock guard is a no-op under `/bin/bash`
|
|
91
|
+
|
|
92
|
+
**Evidence.** `scripts/local-ci.sh:104`:
|
|
93
|
+
|
|
94
|
+
```
|
|
95
|
+
trap '[ "$$" = "$_lci_owner" ] && rm -f "$_lci_lock" 2>/dev/null || true' EXIT
|
|
96
|
+
```
|
|
97
|
+
|
|
98
|
+
The scar comment above it (`:99-103`) correctly diagnoses that a bare EXIT trap
|
|
99
|
+
fires in every subshell and that a finishing child deleted the parent's lock.
|
|
100
|
+
**But `$$` does not change in a bash subshell, so this guard does not
|
|
101
|
+
discriminate.** Verified:
|
|
102
|
+
|
|
103
|
+
```
|
|
104
|
+
$ bash -c 'p=$$; ( [ "$$" = "$p" ] && echo same )'
|
|
105
|
+
same
|
|
106
|
+
```
|
|
107
|
+
|
|
108
|
+
The correct discriminators are `BASHPID` (bash 4+) or `BASH_SUBSHELL` (works in
|
|
109
|
+
3.2). Measured on this host: the script is `#!/usr/bin/env bash` which resolves
|
|
110
|
+
to **bash 5.3**, where `BASHPID` is available. But `/bin/bash` here is **3.2**,
|
|
111
|
+
where `BASHPID` is unset and `BASH_SUBSHELL` is the version-safe choice.
|
|
112
|
+
|
|
113
|
+
**Blast radius: benign** by the lead's own criterion. Worst case is two
|
|
114
|
+
concurrent local-ci runs starving each other into phantom failures -- costly in
|
|
115
|
+
time and trust, but it kills nothing. Left unfixed as out of scope; the one-line
|
|
116
|
+
change is `[ "${BASH_SUBSHELL:-0}" = 0 ]`.
|
|
117
|
+
|
|
118
|
+
Note on the sibling fix already shipped: `_loki_remove_pgid_file`
|
|
119
|
+
(`run.sh:25189`) uses `${BASHPID:-$$}`. Under bash 3.2 that collapses to `$$`
|
|
120
|
+
and the guard degrades to the no-op. **The degradation direction is safe** -- a
|
|
121
|
+
subshell deletes the pgid file early, the reap then finds no file and skips, so
|
|
122
|
+
an orphan survives and nobody gets killed. `run.sh` is also `#!/usr/bin/env
|
|
123
|
+
bash` (5.3 here), so this is a portability note, not a live defect. Worth
|
|
124
|
+
knowing: the mutation test for that guard asserts *source text* contains
|
|
125
|
+
`BASHPID`, so it proves the code is present, not that it behaves correctly on a
|
|
126
|
+
3.2 host.
|
|
127
|
+
|
|
128
|
+
---
|
|
129
|
+
|
|
130
|
+
## Already guarded -- do not re-audit
|
|
131
|
+
|
|
132
|
+
Checked and found correct. Listed explicitly so this ground is not covered
|
|
133
|
+
twice.
|
|
134
|
+
|
|
135
|
+
- **`autonomy/lib/lock.sh:56-67`** -- `_loki_lock_is_stale` requires the
|
|
136
|
+
sentinel PID to be dead **AND** mtime > 30s. Both conditions, not either.
|
|
137
|
+
Correct.
|
|
138
|
+
- **`cleanup_orphan_pids` (`run.sh:2300+`)** -- reaps only on liveness AND
|
|
139
|
+
parent death AND (for wrappers) idle-past-budget AND no live engine child.
|
|
140
|
+
Also self-skips `$$`. Correct, and notably stricter than the pgid reap was.
|
|
141
|
+
- **`_app_runner_pid_is_ours` (`app-runner.sh:242`)** -- a real identity token
|
|
142
|
+
captured post-exec, failing open on a missing token so a live app is never
|
|
143
|
+
falsely killed. Correct where it is called; see Finding 2 for where it is not.
|
|
144
|
+
- **`_app_runner_collect_descendants` (`app-runner.sh:289`)** -- refuses pid
|
|
145
|
+
0/1 and walks parent-child links from our own pid only, so it structurally
|
|
146
|
+
cannot signal outside our subtree. Correct.
|
|
147
|
+
- **`status.ts:266-270` and the bash status reader (`loki:5092`)** -- both do
|
|
148
|
+
`os.kill(pid, 0)` before reporting. A wrong answer only mislabels a URL in
|
|
149
|
+
status output. Benign even when stale.
|
|
150
|
+
- **The `CLEAR`/`KEEP` registry check (`run.sh:1545-1560`)** -- correctly
|
|
151
|
+
refuses to tear the shared dashboard down while any other project holds a
|
|
152
|
+
live pid. This gates Finding 1's call site; the defect was the missing check
|
|
153
|
+
on the pid itself, not this decision.
|
|
154
|
+
|
|
155
|
+
---
|
|
156
|
+
|
|
157
|
+
## Mutation verification (Finding 1's fix)
|
|
158
|
+
|
|
159
|
+
Each guard was reverted individually; the test must go red, and only on its own
|
|
160
|
+
assertion.
|
|
161
|
+
|
|
162
|
+
| Mutation | Before | After | Assertion killed |
|
|
163
|
+
|---|---|---|---|
|
|
164
|
+
| Identity check downgraded to bare `kill -0` | 6 pass / 0 fail | 5 / 1 | "a live non-dashboard process is refused" |
|
|
165
|
+
| Call site ungated (guard present but unused) | 6 / 0 | 5 / 1 | "the kill is gated on the identity guard" |
|
|
166
|
+
| `pid > 1` check removed | 6 / 0 | 5 / 1 | "pid 0/1, empty and malformed inputs refused" |
|
|
167
|
+
|
|
168
|
+
Restored after each. Final: `bash -n autonomy/run.sh` clean,
|
|
169
|
+
`tests/test-stale-dashboard-pid.sh` 6/6, `tests/test-pgid-stale-reap.sh` 8/8
|
|
170
|
+
(unaffected).
|
|
171
|
+
|
|
172
|
+
The first mutation is the important one: it replaces the identity check with
|
|
173
|
+
exactly the insufficient guard (`kill -0`) that a reviewer would most likely
|
|
174
|
+
propose, and the test catches it.
|
|
@@ -25,6 +25,8 @@ different answer on your hardware.
|
|
|
25
25
|
| FULL local gate | 23 to 26 minutes | `LOCAL_CI_TIER=full bash scripts/local-ci.sh`, 166 checks, five runs on an M-series Mac |
|
|
26
26
|
| FAST local gate | about 1 minute | `bash scripts/local-ci.sh`, the documented pre-push tier |
|
|
27
27
|
| Shell suite, sharded | 352 seconds | 323 suites, `LOCAL_CI_SHARDS=4`; 1440 seconds serial, so 4.1x |
|
|
28
|
+
| `loki verify` | 77 seconds | `loki verify HEAD~1` on this repository, 5 changed files; dominated by the project's own test suite, not by our gates |
|
|
29
|
+
| of which, shipped-vs-dev CVE split | 418 ms | one extra `npm audit --omit=dev`; 0.5% of the run, and the reason the audit finding can say whether a CVE reaches users |
|
|
28
30
|
| `loki outcomes` | under 1 second per receipt | `git blame` and one `git log` per changed file |
|
|
29
31
|
| `loki intent status` | milliseconds | hash comparison against `.loki/spec/spec.lock` |
|
|
30
32
|
| Agent readiness | milliseconds | filesystem checks only, no network |
|
|
@@ -50,6 +52,24 @@ non-forgeability requires the signed path
|
|
|
50
52
|
(`LOKI_PROOF_GPG_KEY`, see [SIGNED-RECEIPTS.md](SIGNED-RECEIPTS.md)). We removed
|
|
51
53
|
our own "non-forgeable" claim in v7.111.0 after finding it false on that path.
|
|
52
54
|
|
|
55
|
+
**The tests gate depends on correctly identifying YOUR test runner, and it has
|
|
56
|
+
been wrong before.** Until 2026-08-07 the runner was chosen by grepping
|
|
57
|
+
`package.json` for `"jest"` / `"vitest"` / `"mocha"`, which matches a
|
|
58
|
+
**devDependency**. This repository is the case that exposed it: jest is a
|
|
59
|
+
devDependency with no jest config while `scripts.test` runs `bash -n` plus
|
|
60
|
+
`node --test`, so verify ran jest, jest globbed 895 files that are not jest
|
|
61
|
+
tests, and `loki verify` returned BLOCKED on a clean tree -- permanently, for a
|
|
62
|
+
defect that did not exist. It now reads `scripts.test` with a JSON parser and
|
|
63
|
+
runs what the project declares.
|
|
64
|
+
|
|
65
|
+
We record this rather than quietly fixing it because a false BLOCK is the more
|
|
66
|
+
damaging direction of that error: a gate that cries wolf on every run trains
|
|
67
|
+
you to ignore the verdict, which costs more than the gate ever earned. If the
|
|
68
|
+
tests gate reports a runner you do not use, that is a bug in our detection, not
|
|
69
|
+
a finding about your code -- the runner is named in
|
|
70
|
+
`.loki/verify/evidence.json` under `deterministic_gates[].runner` so you can
|
|
71
|
+
check which one it picked.
|
|
72
|
+
|
|
53
73
|
**Only four of the eight quality gates are agent-independent.** Static analysis,
|
|
54
74
|
mock-integrity, test-mutation and documentation coverage do not ask a model
|
|
55
75
|
anything. The other four involve model judgment and are labelled ASSESSMENTS
|
|
@@ -95,179 +115,39 @@ above less trustworthy:
|
|
|
95
115
|
LOCAL_CI_TIER=full bash scripts/local-ci.sh # the full gate, timed
|
|
96
116
|
loki outcomes --json # post-merge outcomes, or UNKNOWN with reasons
|
|
97
117
|
loki intent status --json # spec-vs-intent drift
|
|
98
|
-
loki proof verify <id> # re-hash a receipt
|
|
118
|
+
loki proof verify <id> # re-hash a receipt (see below)
|
|
99
119
|
bash tests/test-competitor-verify-surface.sh # the competitor CLI measurement
|
|
100
120
|
```
|
|
101
121
|
|
|
102
|
-
|
|
103
|
-
want the report.
|
|
104
|
-
|
|
105
|
-
## How we compare, and what we cannot measure
|
|
106
|
-
|
|
107
|
-
A goal was set to be "2-10x better than factory.ai, cognition devin, 8090.ai
|
|
108
|
-
and replit". This section reports what is measurable and refuses the rest.
|
|
109
|
-
|
|
110
|
-
### What is NOT benchmarked, and why
|
|
111
|
-
|
|
112
|
-
None of those four products has a runnable local arm. Checked on this machine:
|
|
113
|
-
|
|
114
|
-
```
|
|
115
|
-
droid NOT installed devin NOT installed
|
|
116
|
-
replit NOT installed
|
|
117
|
-
claude on PATH aider on PATH
|
|
118
|
-
codex on PATH loki on PATH
|
|
119
|
-
```
|
|
120
|
-
|
|
121
|
-
Devin and 8090 are hosted services with no CLI. Factory's droid and Replit run
|
|
122
|
-
cloud-side. `benchmarks/bench/adapters/` can drive a competing arm as a
|
|
123
|
-
subprocess (`claude_code.py` invokes `claude -p` live), but it cannot drive a
|
|
124
|
-
product that has no local binary.
|
|
125
|
-
|
|
126
|
-
**So there is no "2-10x vs Factory/Devin/8090/Replit" number here, and any such
|
|
127
|
-
figure elsewhere should be treated as unearned.** Publishing one would be the
|
|
128
|
-
same error as sourcing a colour token from a frontend that does not ship: a
|
|
129
|
-
number that looks authoritative and measures something else.
|
|
122
|
+
### Prove the tamper detection yourself, in three commands
|
|
130
123
|
|
|
131
|
-
|
|
124
|
+
The claim is narrow and worth stating exactly: editing a receipt's recorded
|
|
125
|
+
facts is DETECTED. Run this against any receipt in `.loki/proofs/`:
|
|
132
126
|
|
|
133
|
-
|
|
134
|
-
|
|
135
|
-
|
|
136
|
-
|
|
137
|
-
|
|
138
|
-
|
|
139
|
-
against". Re-run it yourself; it names each CLI it checked and SKIPs the ones
|
|
140
|
-
it could not find rather than counting them as absent.
|
|
141
|
-
|
|
142
|
-
### What the competitors say about verification, in their own words
|
|
143
|
-
|
|
144
|
-
Verbatim, with the file each came from, so every line is checkable against the
|
|
145
|
-
scraped corpora:
|
|
146
|
-
|
|
147
|
-
| Source | Their words |
|
|
148
|
-
|---|---|
|
|
149
|
-
| `factory_ai/docs.factory.ai_missions_overview.md` | "**How do you maximize correctness?** Long-running plans accumulate errors." -- published as an OPEN QUESTION |
|
|
150
|
-
| `factory_ai/docs.factory.ai_missions_overview.md` | "Without it, the mission **cannot reliably verify its own work**" (Missions require repo readiness Level 4+) |
|
|
151
|
-
| `devin_cognition_ai/docs.devin.ai_admin_security.md.md` | "it can still experience **hallucinations, introduce bugs into code**, or suggest insecure code" |
|
|
152
|
-
| `8090_ai/www.8090.ai_terms-of-service.md` | "**HUMAN REVIEW AND VERIFICATION OF ALL OUTPUT**" required, while the same ToS caps liability at "FIFTY US DOLLARS" and disclaims "ACCURACY" |
|
|
153
|
-
|
|
154
|
-
Factory's is the most honest of the four: they name verification as an open
|
|
155
|
-
research question rather than a solved feature, and they state that
|
|
156
|
-
self-verification is a property of the ENVIRONMENT, not of the agent. We agree,
|
|
157
|
-
which is why `loki readiness` measures the repo and not the model.
|
|
158
|
-
|
|
159
|
-
### The claim we will actually defend
|
|
160
|
-
|
|
161
|
-
Not a multiplier. A receipt you can recompute:
|
|
162
|
-
|
|
163
|
-
```
|
|
164
|
-
loki proof verify <id> # re-hash the receipt; exit 1 on tamper
|
|
165
|
-
loki outcomes --json # anchored, or UNKNOWN with a named reason
|
|
127
|
+
```bash
|
|
128
|
+
ID=$(ls .loki/proofs | head -1)
|
|
129
|
+
V() { loki proof verify "$ID" --json | python3 -c 'import json,sys;print(json.load(sys.stdin)["hash_ok"])'; }
|
|
130
|
+
V # True
|
|
131
|
+
python3 -c "import json;p='.loki/proofs/$ID/proof.json';d=json.load(open(p));d['files_changed']={'count':999999};json.dump(d,open(p,'w'))"
|
|
132
|
+
V # False
|
|
166
133
|
```
|
|
167
134
|
|
|
168
|
-
|
|
169
|
-
verify its own work", an anchored receipt is a categorical difference. It is
|
|
170
|
-
also falsifiable: if `loki outcomes` reports UNKNOWN, we say UNKNOWN. On this
|
|
171
|
-
repo it reported ANCHORED 0 of 9 for weeks, and the fix
|
|
172
|
-
(`facts.git.base_sha` was empty) is in the history.
|
|
135
|
+
Measured on this repository: `True` -> `False` -> `True` after restoring.
|
|
173
136
|
|
|
174
|
-
|
|
137
|
+
**Read `hash_ok`, not `ok`.** They answer different questions and conflating
|
|
138
|
+
them produces a false alarm. `hash_ok` is integrity: do the recorded facts still
|
|
139
|
+
hash to the recorded digest. `ok` also folds in `tree_drift`, which is true
|
|
140
|
+
whenever the working tree has moved since the receipt was written -- so an
|
|
141
|
+
untampered receipt from last month correctly reports `ok: false` with
|
|
142
|
+
`hash_ok: true`. An earlier version of this document said "exit 1 on tamper",
|
|
143
|
+
which is wrong in exactly that way: a drifted-but-intact receipt also exits 1.
|
|
175
144
|
|
|
176
|
-
|
|
177
|
-
|
|
178
|
-
|
|
179
|
-
-
|
|
180
|
-
|
|
181
|
-
|
|
182
|
-
Those comparisons are not available to us and are not claimed.
|
|
145
|
+
**What this does NOT establish.** Integrity is not provenance. On the unsigned
|
|
146
|
+
path a party who rewrites the facts AND recomputes the digest passes this check
|
|
147
|
+
-- see the forgeability limit above. Provenance requires the signed path
|
|
148
|
+
(`LOKI_PROOF_GPG_KEY`, [SIGNED-RECEIPTS.md](SIGNED-RECEIPTS.md)), and the remote
|
|
149
|
+
client reports the two separately for that reason: VERIFIED, UNSIGNED,
|
|
150
|
+
UNCHECKED and TAMPERED are four distinct verdicts, never collapsed.
|
|
183
151
|
|
|
184
|
-
|
|
185
|
-
|
|
186
|
-
The one comparison we CAN run. `benchmarks/bench/matrix.sh` defines two
|
|
187
|
-
configs against the same task and the same model:
|
|
188
|
-
|
|
189
|
-
- `baseline` -- raw model, minimal orchestration (in-repo comment: "the
|
|
190
|
-
Replit/Cursor mode")
|
|
191
|
-
- `full` -- the harness: council, code review, self-heal, auto-tune
|
|
192
|
-
|
|
193
|
-
Paired on identical tasks, both arms on `haiku`, from
|
|
194
|
-
`benchmarks/bench/results/`:
|
|
195
|
-
|
|
196
|
-
| Task | harness (`full`) | raw model (`baseline`) | verdict |
|
|
197
|
-
|---|---|---|---|
|
|
198
|
-
| `hard-2-ledger` | **4/4**, $0.89 | **0/4**, $0.19 | harness wins outright |
|
|
199
|
-
| `hard-1-order-api` | 3 runs, 1.00, $0.58 | 1 run, 1.00, $0.20 | harness cost 2.8x for nothing |
|
|
200
|
-
| `multifail-1-two-modules` | 2 runs, 1.00, $0.27 | 1 run, 1.00, $0.14 | harness cost 1.9x for nothing |
|
|
201
|
-
|
|
202
|
-
**Two of the three paired tasks are UNFAVOURABLE to the harness.** That is the
|
|
203
|
-
honest headline, and it is narrower than "the harness is worth more than the
|
|
204
|
-
model":
|
|
205
|
-
|
|
206
|
-
**The harness matters on hard tasks and is pure overhead on easy ones.**
|
|
207
|
-
|
|
208
|
-
On `hard-2-ledger` the raw model never finished the task across 4 trials and
|
|
209
|
-
the harness finished it every time. That is not a percentage improvement; the
|
|
210
|
-
baseline success rate is zero. It cost 4.7x more per run and produced a working
|
|
211
|
-
result instead of nothing.
|
|
212
|
-
|
|
213
|
-
On the other two, both arms succeed and the harness simply costs 1.9x-2.8x
|
|
214
|
-
more. We publish those rows because the direction is unfavourable to us, and a
|
|
215
|
-
benchmark table that only survives its favourable rows is an advertisement.
|
|
216
|
-
|
|
217
|
-
**Two further baseline cells were attempted and produced NO data.**
|
|
218
|
-
`tokenheavy-1-crm` (2 trials) and a wider `hard-1-order-api` (3 trials) both hit
|
|
219
|
-
the 1200s cell timeout, which caps all trials of a cell together. The harness
|
|
220
|
-
wrote no result file and the runner said so:
|
|
221
|
-
|
|
222
|
-
> WARNING: cell haiku-baseline / tokenheavy-1-crm wrote NO new result (likely
|
|
223
|
-
> timed out at 1200s). This cell is MISSING from the report.
|
|
224
|
-
|
|
225
|
-
They are absent from the table rather than counted as failures. A timeout is
|
|
226
|
-
not evidence the raw model cannot do the task; it is evidence we did not
|
|
227
|
-
measure it. Reporting them as baseline losses would have made the harness look
|
|
228
|
-
better on data that does not exist.
|
|
229
|
-
|
|
230
|
-
An earlier version of this section reported only the first two tasks and read
|
|
231
|
-
as a stronger claim than the data supported. The `multifail` baseline arm was
|
|
232
|
-
run afterwards specifically to test whether the claim would survive more data.
|
|
233
|
-
It did not survive intact, and the table was corrected rather than the
|
|
234
|
-
measurement dropped.
|
|
235
|
-
|
|
236
|
-
Aggregate across all recorded cells:
|
|
237
|
-
|
|
238
|
-
| Cell | n | success (median) | cost (median) |
|
|
239
|
-
|---|---|---|---|
|
|
240
|
-
| `haiku` + harness | 11 | 1.00 | $0.54 |
|
|
241
|
-
| `opus` + baseline | 5 | 1.00 | $0.83 |
|
|
242
|
-
| `haiku` + baseline | 8 | 0.50 | $0.19 |
|
|
243
|
-
|
|
244
|
-
The cheap model WITH the harness matches the expensive model without it, at
|
|
245
|
-
35% lower cost. Treat that as directional, not as a headline: the task sets
|
|
246
|
-
differ between those three cells, which is exactly why the paired table above
|
|
247
|
-
is the one that carries the argument.
|
|
248
|
-
|
|
249
|
-
**Limits of this measurement, stated so it cannot be over-read:**
|
|
250
|
-
|
|
251
|
-
- Two paired tasks. `hard-1-order-api`'s baseline arm is n=1.
|
|
252
|
-
- It measures OUR harness against OUR baseline config. `baseline` is a
|
|
253
|
-
documented stand-in for "raw model, minimal orchestration", NOT a
|
|
254
|
-
measurement of Replit, Cursor, or any other product.
|
|
255
|
-
- Success is a held-out acceptance exit code, not a judgement of code quality.
|
|
256
|
-
- Reproduce it: `LOKI_BENCH_SPEND_APPROVED=1 bash benchmarks/bench/matrix.sh pilot`.
|
|
257
|
-
The spend interlock is default-deny on purpose; a benchmark that starts a
|
|
258
|
-
paid tool without an explicit opt-in is how a surprise bill happens.
|
|
259
|
-
|
|
260
|
-
**The exclusion rule, observed on a live run rather than asserted.** A fresh
|
|
261
|
-
`haiku-full` trial on `hard-1-order-api` (2026-08-07) hit the 1200s adapter
|
|
262
|
-
timeout. The held-out grader inspected the workdir and returned
|
|
263
|
-
`success: true` -- there WAS a passing artifact. The harness still recorded
|
|
264
|
-
`measured: false`, `unmeasured_reasons: ["adapter exit_status=timeout"]`, and
|
|
265
|
-
excluded the trial from k/N, with this note in the result file:
|
|
266
|
-
|
|
267
|
-
> "the held-out grader's own verdict on whatever was in the workdir. Real
|
|
268
|
-
> evidence about the ARTIFACT; NOT evidence that a run happened."
|
|
269
|
-
|
|
270
|
-
A naive harness scores that 1/1. Ours scores it 0/0 and says why, so the
|
|
271
|
-
`hard-1-order-api` row above stayed at n=3 instead of being inflated to n=4 by
|
|
272
|
-
a run that never finished. The number in this document is smaller because of
|
|
273
|
-
that rule, which is the point of having it.
|
|
152
|
+
If a number here does not reproduce on your machine, that is a defect and we
|
|
153
|
+
want the report.
|