loki-mode 9.8.0 → 9.12.0

This diff represents the content of publicly available package versions that have been released to one of the supported registries. The information contained in this diff is provided for informational purposes only and reflects changes between package versions as they appear in their respective public registries.
Files changed (79) hide show
  1. package/README.md +19 -14
  2. package/SKILL.md +3 -2
  3. package/VERSION +1 -1
  4. package/autonomy/loki +122 -1
  5. package/autonomy/run.sh +49 -2
  6. package/dashboard/__init__.py +1 -1
  7. package/dashboard/api_evidence.py +411 -0
  8. package/dashboard/api_operator.py +283 -0
  9. package/dashboard/api_phases.py +262 -0
  10. package/dashboard/api_releases.py +242 -0
  11. package/dashboard/api_runs.py +477 -0
  12. package/dashboard/api_tests.py +444 -0
  13. package/dashboard/api_v2.py +47 -1
  14. package/dashboard/server.py +54 -0
  15. package/dashboard/static/index.html +246 -135
  16. package/docs/ARCHITECTURE-OVERVIEW.md +5 -3
  17. package/docs/CAPABILITY-BACKLOG.md +53 -0
  18. package/docs/COMPARISON.md +2 -2
  19. package/docs/COMPETITIVE-ANALYSIS.md +1 -1
  20. package/docs/COMPETITIVE-SCORECARD.md +422 -0
  21. package/docs/DASHBOARD-9.12-EVIDENCE.md +97 -0
  22. package/docs/DASHBOARD-ARCHITECTURE.md +423 -0
  23. package/docs/DEMOS.md +21 -23
  24. package/docs/HANDOFF-2026-08-03.md +439 -0
  25. package/docs/INSTALLATION.md +17 -10
  26. package/docs/OUTCOME-FRONTIER.md +536 -0
  27. package/docs/PROMPT-ABLATION-RESULT.md +97 -0
  28. package/docs/TOOLS.md +800 -0
  29. package/docs/alternative-installations.md +2 -3
  30. package/docs/audit-logging.md +44 -35
  31. package/docs/authentication.md +13 -2
  32. package/docs/authorization.md +87 -81
  33. package/docs/git-workflow.md +6 -3
  34. package/docs/metrics.md +15 -16
  35. package/docs/network-security.md +16 -13
  36. package/docs/openclaw-integration.md +36 -556
  37. package/docs/show-hn-post.md +2 -2
  38. package/docs/siem-integration.md +39 -36
  39. package/loki-ts/dist/loki.js +18 -18
  40. package/mcp/__init__.py +1 -1
  41. package/package.json +2 -2
  42. package/plugins/loki-mode/.claude-plugin/plugin.json +1 -1
  43. package/references/confidence-routing.md +18 -1
  44. package/references/invariant-checks.md +13 -8
  45. package/references/magic-rarv-integration.md +0 -1
  46. package/references/multi-provider.md +27 -5
  47. package/skills/healing.md +4 -2
  48. package/tools/audit-docs.py +488 -0
  49. package/tools/baseline-pin.py +19 -1
  50. package/tools/calibration-audit.py +523 -0
  51. package/tools/ci-gate.py +19 -1
  52. package/tools/cost-forecast.py +344 -0
  53. package/tools/cost-guard.py +19 -1
  54. package/tools/cost-history.py +19 -1
  55. package/tools/cost-per-outcome.py +394 -0
  56. package/tools/estimate-run.py +19 -1
  57. package/tools/evidence-freshness.py +307 -0
  58. package/tools/gate-init.py +19 -1
  59. package/tools/gate-report.py +19 -1
  60. package/tools/gate-simulate.py +570 -0
  61. package/tools/gate-trend.py +354 -0
  62. package/tools/model-advisor.py +52 -1
  63. package/tools/policy-load.py +19 -1
  64. package/tools/prompt-cost.py +363 -0
  65. package/tools/prompt-diff.py +448 -0
  66. package/tools/prompt-lint.py +448 -0
  67. package/tools/receipt-bundle.py +72 -2
  68. package/tools/receipt-diff.py +19 -1
  69. package/tools/receipt-find.py +19 -1
  70. package/tools/receipt-stats.py +380 -0
  71. package/tools/receipt-timeline.py +478 -0
  72. package/tools/receipt-verify-batch.py +291 -0
  73. package/tools/run-replay.py +19 -1
  74. package/tools/signing-status.py +19 -1
  75. package/tools/token-guard.py +19 -1
  76. package/tools/token-tax.py +375 -0
  77. package/tools/tool-index.py +19 -1
  78. package/tools/verification-tax.py +277 -0
  79. package/tools/verify-chain.py +361 -0
@@ -0,0 +1,439 @@
1
+ # Handoff, 2026-08-03
2
+
3
+ ## 0. DECISION-GRADE RC SUMMARY (read this first)
4
+
5
+ Pinned: **`worktree-pre-push-scoped-pytest` @ `16ba5c5e`**.
6
+
7
+ | Fact | Value | Command |
8
+ |---|---|---|
9
+ | RC HEAD | `16ba5c5e` at the time of writing; this document's own commit
10
+ advances it, so ALWAYS trust the command over this cell | `git rev-parse HEAD` |
11
+ | vs `origin/main` (`dd692561`) | **27 ahead, 0 behind** | `git rev-list --count origin/main..HEAD` |
12
+ | vs RC remote (`3808f63f`) | **17 ahead, 0 behind** | `git rev-list --count origin/worktree-...` |
13
+
14
+ WHY THE SHA CELL IS ALWAYS ONE BEHIND. Committing this file moves HEAD, so
15
+ a handoff can never pin its own commit. The cell records the SHA the
16
+ receipts below were taken on; the live value is one commit later. Re-run
17
+ `git rev-parse --short HEAD` and `git rev-list --count origin/main..HEAD`
18
+ rather than trusting any number typed here -- that habit is what caught
19
+ three stale pins in this document already.
20
+
21
+ LOCAL EVIDENCE AND CI EVIDENCE ARE DIFFERENT THINGS, and this document keeps
22
+ them apart deliberately. Everything under "Local receipts" ran on this machine
23
+ against this exact SHA. NO CI has ever run on `16ba5c5e` or on any commit in this
24
+ branch: a branch push triggers no workflow (see gate 2), so there is no CI
25
+ verdict to report for the RC. The CI table below is for `origin/main` only.
26
+
27
+ **A BRANCH PUSH ALONE RUNS NO CI.** Every workflow triggers only on
28
+ push-to-main or a pull_request targeting main, and `test.yml` has no
29
+ `workflow_dispatch`. Pushing this branch produces zero runs, as observed.
30
+ A CI verdict on this SHA requires a branch push AND a PR to main.
31
+
32
+ ### CI evidence, exact
33
+
34
+ `origin/main` = `dd692561`:
35
+
36
+ | Workflow | Conclusion | URL |
37
+ |---|---|---|
38
+ | Post-Release Soak Monitor | success | actions/runs/30805237104 |
39
+ | Security Audit | success | actions/runs/30796434331 |
40
+ | **Tests** | **failure** | actions/runs/30794017151 |
41
+
42
+ The Tests failure is ONE job, ONE cause: `Shell tests (shard 3/4)` ->
43
+ `ShellCheck Linting FAILED`. Shards 0, 1, 2 pass. Held commit `a4d8b839`
44
+ takes shellcheck from 7 warning sites to **361 passed / 0 failed**.
45
+
46
+ **RC branch CI: 0 runs, and it will stay 0.** Every workflow triggers only on
47
+ push-to-main or a PR targeting main, and `test.yml` has no `workflow_dispatch`.
48
+ A branch push cannot produce a CI verdict. PR #181 is MERGED 2026-07-28 on a
49
+ different diff and is NOT evidence about current main.
50
+
51
+ ### Local receipts on the exact RC HEAD `16ba5c5e`
52
+
53
+ Measured on this machine. NOT CI: no workflow has run on this SHA.
54
+
55
+ | Check | Result |
56
+ |---|---|
57
+ | `bash scripts/local-ci.sh` | **exit 0**, Failed 0 |
58
+ | `bash tests/run-shellcheck.sh` | **361 passed, 0 failed** |
59
+ | `bash tests/test-plan-command.sh` | **27 passed, 0 failed** (was 25/2; two stale June cost pins fixed) |
60
+ | `python3 -m pytest tests/ -q -k "bench or schema"` | **149 passed**, 1 skipped |
61
+ | `python3 -m pytest tests/` (full) | **2970 passed, 1 failed**, 19 skipped |
62
+
63
+ THE ONE FULL-SUITE FAILURE, characterised rather than hidden:
64
+ `tests/dashboard/test_build_supervisor.py::test_confined_claude_auth_uses_exact_login_keychain_capability`.
65
+
66
+ - PRE-EXISTING, not caused by anything here: a pristine `git archive HEAD`
67
+ extract with no working-tree residue reproduces it.
68
+ - ENVIRONMENT-SENSITIVE, not a code defect: the same file passes **37/37 in
69
+ isolation in this worktree, twice in a row**, and the failing assertion is
70
+ about a real login-keychain capability, which differs between this worktree
71
+ and a bare extract.
72
+ - It is NOT in the fast tier, which is why `local-ci.sh` is green while the
73
+ full suite is not. Both numbers are reported rather than the flattering one.
74
+
75
+ ### Working-tree residue, classified. NOTHING deleted, reverted or committed.
76
+
77
+ | Path | Class | Note |
78
+ |---|---|---|
79
+ | `coverage/clover.xml` | generated residue | rewritten by test runs |
80
+ | `coverage/lcov-report/index.html` | generated residue | rewritten by test runs |
81
+ | `loki-ts/dist/loki.js.map` | generated residue | rewritten by `bun run build` |
82
+ | `f.txt` (deleted) | **user-owned** | deleted in working tree only; **still present in HEAD**, fully recoverable via `git checkout -- f.txt` |
83
+ | `benchmarks/results/prompt-ablation.jsonl` | measurement output | untracked; the 6 real ablation trials |
84
+
85
+
86
+ ### Every held commit (27, unpushed to main)
87
+
88
+ | SHA | Subject |
89
+ |---|---|
90
+ | `16ba5c5e` | feat(bench): `loki bench oracles` -- reachable where the spend decision is made |
91
+ | `5cc5ffd2` | feat(bench): attest the artifact hashes so an edited answer key cannot pass |
92
+ | `6dbef513` | docs: both remaining tiers now verified against a near-miss, gate 3 precondition MET |
93
+ | `89babec7` | docs: gate 3's precondition is PARTIALLY met, with the hashes that make it checkable |
94
+ | `43e4d530` | test(bench): pin the private probe's own guarantees, and refuse to promote a tier on private evidence |
95
+ | `f9895095` | feat(bench): hash-bound private oracle probes -- the replayable version of what I retracted |
96
+ | `ed9f7a6c` | fix(bench,docs): retract the medium/high VERIFIED label -- the evidence cannot be replayed |
97
+ | `a54c190c` | feat(bench): medium and high oracles verified against a plausible wrong answer |
98
+ | `1fd65420` | docs: make HEAD decision-ready -- exact-SHA receipts, local evidence kept apart from CI |
99
+ | `5e0ed5b5` | fix(tests): two demo assertions pinned a June cost the estimator has since corrected |
100
+ | `49ebefa2` | docs: reconcile the handoff to live topology and separate the three founder gates |
101
+ | `7b25a772` | docs(bench): record that adversarially_unverified is unreachable today, and why it stays |
102
+ | `ea930bbb` | fix(bench,ci): fail on model mismatch, wire the integrity check, add adversarial oracle probes |
103
+ | `8980313e` | docs: decision-grade RC handoff reconciled to current topology |
104
+ | `2c031a14` | feat(safety): default-deny release interlock, and a detector for the recurring core.bare fault |
105
+ | `6ffa5080` | feat(bench): founder-gated spend interlock, and keep grader evidence beside the measured flag |
106
+ | `35add6d7` | fix(bench): flag acceptance commands that cannot fail -- a timeout scored as success |
107
+ | `3808f63f` | feat(bench): cross-tier oracle receipt -- verifies the graders, runs no agent |
108
+ | `1591f880` | docs: re-verify ref topology after an audit, and name the single CI blocker |
109
+ | `ef0a6d49` | docs: put the spend decision's actual numbers in the handoff |
110
+ | `12669cd3` | docs: the no-spend outcome frontier -- 3 held-out tasks, deterministic oracles, nothing run |
111
+ | `fdba7023` | docs: handoff -- reconciliation, receipts, risks, founder-gated next action |
112
+ | `19c2b5ad` | fix(tests): bound each completion probe, so a hang is a NAMED failure not rc143 |
113
+ | `90468210` | feat(advisor): point at the calibration signal, and pin that it can never rank |
114
+ | `0f4da110` | docs: freeze the T1/T2 benchmark designs, unrun and unpaid |
115
+ | `8c33d123` | feat(tools): offline calibration audit, and it discloses that it cannot measure accuracy |
116
+ | `a4d8b839` | fix(lint): the last red CI job was three pre-existing shellcheck warnings |
117
+
118
+ ### THE THREE FOUNDER GATES, deliberately separate
119
+
120
+ Ordered by dependency, not by value. Gate 1 first because gates 2 and 3 both
121
+ move code toward main, and today nothing non-bypassable stands between main
122
+ and a publish.
123
+
124
+ **GATE 1 -- a durable, non-bypassable publish control.** Zero spend, zero
125
+ code. Either an `environment:` key on release.yml's publish jobs (GitHub then
126
+ requires a named reviewer before those jobs start), or a ruleset on main.
127
+ Needs repository settings access, which is why it is not done here.
128
+
129
+ Everything I built is operator-side and BYPASSABLE by not running it:
130
+ `scripts/release-approval-gate.sh` refuses by default when run, and is a
131
+ no-op when skipped. Treat a green run as "the operator checked", never as
132
+ "publishing is gated". Measured today: release.yml triggers on any
133
+ VERSION-path push to main and runs `gh release create`, `npm publish` twice
134
+ and a Docker push, with no `environment:` key, and
135
+ `gh api .../branches/main/protection` returns 404.
136
+
137
+ **GATE 2 -- branch push PLUS a PR to main, for CI on this exact SHA.** Zero
138
+ spend. These are ONE gate because either alone yields nothing: a branch push
139
+ runs no workflow, and there is no PR without the push. Only this produces a
140
+ CI verdict on `43e4d530`. It is also the only path by which held `a4d8b839`
141
+ reaches main and clears the single remaining Tests failure.
142
+
143
+ **GATE 3 -- benchmark spend, roughly $25-$180 and 4-12 hours.** Independent
144
+ of gates 1 and 2. Buys the first OUTCOME evidence rather than affordance
145
+ evidence (`docs/OUTCOME-FRONTIER.md`).
146
+
147
+ PRECONDITION ON GATE 3: **MET, with one qualifier that decides who can
148
+ check it.** Small is verified by a receipt ANY reader can reproduce.
149
+ Medium and high are verified by a PRIVATE hash-bound receipt (below):
150
+ replayable by whoever holds the artifacts, not by a reader on a fresh
151
+ clone.
152
+
153
+ | Tier | Positive | Adversarial | Receipt status |
154
+ |---|---|---|---|
155
+ | small `simple-1-contact-form` | exit 0 | exit 1 | **VERIFIED, reproducible** -- fixtures are committed, `tier_receipt.py` re-runs both probes on demand |
156
+ | medium `multifail-1-two-modules` | -- | -- | **NOT ATTEMPTED** in the shared receipt; VERIFIED privately, see below |
157
+ | high `hard-1-order-api` | -- | -- | **NOT ATTEMPTED** in the shared receipt; VERIFIED privately, see below |
158
+
159
+ NON-REPRODUCIBLE OBSERVATION, recorded as a NOTE and explicitly NOT a
160
+ receipt. On 2026-08-03 the medium and high graders were exercised by hand
161
+ against artifacts built in a temp directory and destroyed. Observed exits
162
+ were 0 for a correct implementation and 1 for a near-miss (high: omitting
163
+ the subtotal>=100 discount; medium: fixing only cluster A). The high-tier
164
+ near-miss passed all ten validation cases and both non-discount totals.
165
+
166
+ That observation must not be promoted to VERIFIED, and the earlier draft of
167
+ this section wrongly did. The artifacts are gone, no hash binds them to the
168
+ run, and nothing in this repo can replay it -- so it fails the same standard
169
+ this codebase applies to every other claim: an unreproducible observation is
170
+ not evidence. It is recorded because it is a useful prior for whoever builds
171
+ the real probes, and for no stronger purpose.
172
+
173
+ THE MECHANISM THAT MEETS THIS PRECONDITION NOW EXISTS, and one of the two
174
+ tiers is verified against it. `benchmarks/bench/private_probe.py` runs the
175
+ in-repo graders against artifacts held OUTSIDE this repository (default
176
+ `~/loki-bench-private`), content-hashes each set over its sorted
177
+ (filename, bytes) pairs, and records the hash beside the exit code. The
178
+ artifacts stay out of the repo, so the held-out design is uncontaminated;
179
+ the hash makes the claim replayable, which the earlier by-hand run was not.
180
+
181
+ Measured 2026-08-03:
182
+
183
+ | Task | Probe | Exit | Expected | Artifact sha256 (first 16) |
184
+ |---|---|---|---|---|
185
+ | `hard-1-order-api` | positive | 0 | 0 | `884d555b6e53833d` |
186
+ | `hard-1-order-api` | adversarial | 1 | 1 | `486a62137c64167a` |
187
+ | `multifail-1-two-modules` | positive | 0 | 0 | `a7134c9c1ae36ae9` |
188
+ | `multifail-1-two-modules` | adversarial | 1 | 1 | `f4daff9d346c6116` |
189
+
190
+ Both tiers read `verified`. Each adversarial artifact is the near-miss the
191
+ grader's own docstring names: the high one omits the subtotal>=100 discount
192
+ while passing all ten validation cases and both non-discount totals; the
193
+ medium one fixes cluster A and leaves cluster B absent, and the grader
194
+ rejects it naming `cluster B_roman`.
195
+
196
+ So every tier now rejects a PLAUSIBLE WRONG answer, not merely an absent
197
+ one, which is the property that makes a paid run's number mean something.
198
+
199
+ HOW TO CHECK IT YOURSELF, and what stops it rotting.
200
+
201
+ loki bench oracles # or: python3 benchmarks/bench/private_probe.py
202
+
203
+ Reachable from the CLI deliberately, and placed beside `bench run`: the
204
+ question in front of a paid run is whether a green cell would mean
205
+ anything, so the answer sits one command from the spend decision. Costs
206
+ nothing -- local graders against artifacts already on disk, never a
207
+ provider.
208
+
209
+ `benchmarks/bench/private_attestation.json` commits the HASHES (never the
210
+ artifacts) and the probe compares against them. Without that the probe
211
+ would verify whatever bytes it found, so an artifact edited afterwards
212
+ would be re-blessed under a new hash while the receipt still read
213
+ `verified` -- which is how an edited answer key launders itself green.
214
+
215
+ Proven on the case that matters: appending a COSMETIC comment leaves
216
+ behaviour identical, the grader still returns its expected exit code, and
217
+ the run still fails with `attestation=DRIFTED` naming both hashes. Only
218
+ the hash comparison can catch that.
219
+
220
+ An unattested probe reads `unattested`, not drift. Absence of a record is
221
+ not evidence of tampering.
222
+
223
+ TWO QUALIFIERS THAT MUST TRAVEL WITH THAT TABLE.
224
+
225
+ First, this is a PRIVATE receipt. It is checkable by whoever holds the
226
+ artifacts, and not by a reader on a fresh clone. `tier_receipt.py`
227
+ therefore still reports medium and high as `not_attempted`, deliberately:
228
+ promoting them there would make a shared receipt's verdict depend on who
229
+ ran it, which re-creates the retracted defect one level up. The two
230
+ receipts answer different questions and a human combines them.
231
+
232
+ Second, the hashes above identify one operator's artifacts on one machine.
233
+ They prove the claim is REPLAYABLE, not that anyone else can replay it
234
+ today. Making it shareable (a second repository, an attested bundle) is a
235
+ distribution decision, not a code change.
236
+
237
+ ### RELEASE EXPOSURE, measured today
238
+
239
+ `.github/workflows/release.yml` triggers on `push: paths: ['VERSION'],
240
+ branches: [main]` and runs `gh release create`, `npm publish --access public`
241
+ (twice) and a Docker push. It has **no `environment:` approval gate**, and
242
+ `gh api .../branches/main/protection` returns **404 Branch not protected**.
243
+
244
+ So a single VERSION-editing commit reaching main publishes with no human in
245
+ the loop. `scripts/release-approval-gate.sh` is a default-DENY local interlock
246
+ against exactly that, verified on a throwaway 9.11.0 -> 9.99.99 bump. The
247
+ durable fix (an `environment:` gate or a ruleset) changes the publishing path
248
+ or repo settings and is **founder-gated** -- it is listed as an approval still
249
+ needed, not silently applied.
250
+
251
+
252
+ State at handoff. Everything below is verified by a command whose output was
253
+ read, or is marked UNKNOWN. Nothing has been pushed, published, merged, or
254
+ released, and nothing was spent.
255
+
256
+ ## 1. Reconciliation
257
+
258
+ SUPERSEDED SNAPSHOT. The HEAD and commit count below are from an earlier
259
+ point in the session and are kept as history, NOT as current state. Section 0
260
+ is the live topology: `7b25a772`, 16 ahead of origin/main, 6 ahead of the RC
261
+ remote. The VERSION / npm / tag rows are still accurate.
262
+
263
+ | Fact | Value (as of `19c2b5ad`) | How verified |
264
+ |---|---|---|
265
+ | Local HEAD | `19c2b5ad` (SUPERSEDED, now `7b25a772`) | `git rev-parse HEAD` |
266
+ | `origin/main` | `dd692561` (unchanged) | `git rev-parse origin/main` after fetch |
267
+ | Commits ahead | 5 (SUPERSEDED, now 16) | `git rev-list --count origin/main..HEAD` |
268
+ | repo VERSION | 9.11.0 | `cat VERSION` |
269
+ | npm latest | 9.8.1 | `npm view loki-mode version` |
270
+ | newest git tag | v9.8.1 | `git tag --sort=-v:refname \| head -1` |
271
+
272
+ **npm and the tag agree at 9.8.1.** The repo being at 9.11.0 is not drift: it
273
+ is an unreleased repo under a no-publish directive. Three minor versions of
274
+ work are committed and unpublished by instruction. Reconciled, not a defect.
275
+
276
+ ## 1b. Ref topology, re-verified after an independent audit (09:04Z)
277
+
278
+ An audit reported the session ref `loki/session-1785167314-9337` at
279
+ `09138e26` as 0 ahead / 176 behind, fully contained in `origin/main`, and
280
+ concluded the "commits held/unpushed" claim was stale. **Both statements are
281
+ correct, about DIFFERENT refs.** Re-measured here:
282
+
283
+ ```
284
+ git rev-list --left-right --count origin/main...loki/session-1785167314-9337
285
+ -> 176 0 (behind 176, ahead 0)
286
+ git merge-base --is-ancestor loki/session-1785167314-9337 origin/main
287
+ -> true origin/main CONTAINS it; nothing to push from that ref
288
+
289
+ git rev-list --left-right --count origin/main...HEAD # worktree-pre-push-scoped-pytest
290
+ -> 0 8 (behind 0, ahead 8)
291
+ git merge-base --is-ancestor HEAD origin/main
292
+ -> false NOT contained; the 8 are unique
293
+ ```
294
+
295
+ Each of the 8 was additionally tested individually with
296
+ `git merge-base --is-ancestor <sha> origin/main`; none is in `origin/main`.
297
+
298
+ The audited session ref is not one I have worked on. The held work is on
299
+ `worktree-pre-push-scoped-pytest`. The push was NOT cancelled as stale,
300
+ because it does not target the audited ref -- but it remains UNPUSHED by
301
+ directive, pending explicit approval.
302
+
303
+ ## 1c. origin/main is ONE JOB from CI-green
304
+
305
+ Receipts for `origin/main` = `dd692561`:
306
+
307
+ | Workflow | Conclusion | URL |
308
+ |---|---|---|
309
+ | Security Audit | success | actions/runs/30796434331 |
310
+ | SBOM (CycloneDX) | success | actions/runs/30794017044 |
311
+ | Bun Parity | success | actions/runs/30794017132 |
312
+ | Coverage (baseline) | success | actions/runs/30794017182 |
313
+ | pages build | success | actions/runs/30794016322 |
314
+ | **Tests** | **failure** | actions/runs/30794017151 |
315
+
316
+ The Tests failure is ONE job with ONE cause: `Shell tests (shard 3/4)` ->
317
+ `ShellCheck Linting FAILED`. Shards 0, 1 and 2 pass.
318
+
319
+ Measured on both sides of the fix:
320
+
321
+ - `origin/main`'s copies of the three files carry **7 shellcheck warning
322
+ sites** (test-iteration-grace 5, test-go-cargo-gate-timeout 1,
323
+ cleanup-test-processes 1).
324
+ - The held commit `a4d8b839` takes those same three files to **0**, and
325
+ `tests/run-shellcheck.sh` reports **360 passed / 0 failed**.
326
+
327
+ So the single blocker to a CI-green release candidate is already fixed and
328
+ held. No new work is required to clear it; only the push gate.
329
+
330
+ ## 2. Held commits (8, unpushed)
331
+
332
+ | SHA | What it is |
333
+ |---|---|
334
+ | `a4d8b839` | Last red CI job: three PRE-EXISTING shellcheck warnings (not mine; last touched Jul 31 / Aug 1). One was a real hazard -- an unguarded `cd` that could make a gate examine the wrong tree and report a pass for a repo it never looked at. |
335
+ | `8c33d123` | Offline calibration audit. Discloses, before any number, that it measures agreement with the council majority and NOT accuracy. |
336
+ | `0f4da110` | T1/T2 benchmark designs frozen with held-out discipline. Unrun, unpaid. |
337
+ | `90468210` | Calibration wired into model-advisor as a ONE-WAY caveat. Pinned by test so it can never become a ranking input. |
338
+ | `19c2b5ad` | Per-command timeout on the completion-coverage probe. Root cause of the weekly audit's rc143, with the culprit named. |
339
+ | `fdba7023` | This handoff. |
340
+ | `12669cd3` | The no-spend outcome frontier: 3 held-out tasks, deterministic non-self-grading oracles, nothing run. |
341
+ | `ef0a6d49` | Spend-decision figures moved into the handoff so the decision needs one file, not two. |
342
+
343
+ ## 3. Receipts
344
+
345
+ - **shellcheck: 360 passed, 0 failed.** Was 357/3.
346
+ - **Weekly-integrity rc143: root-caused and fixed.**
347
+ `test-completion-coverage.sh` probes ~274 command candidates as
348
+ `loki <c> --help`. `help` is among them (verified by re-running the
349
+ candidate extraction standalone), so it executed `loki help --help` -- the
350
+ unbounded self-delegation fixed in `ccf8dcbb`, which spawned a process per
351
+ level until the fork table was exhausted and the runner was SIGTERM'd.
352
+ Fixed with a PER-COMMAND 15s bound, never a blanket skip: skipping `help`
353
+ would have hidden the bug that needed fixing.
354
+ Proven non-vacuous by injecting a 60s hang into `doctor`:
355
+ `rc=1`, 34s, `PROBE HUNG: 'loki doctor --help' did not return within 15s`.
356
+ Without the hang: 6 passed, 0 failed, 20s.
357
+ - **CI on `dd692561` (last pushed SHA), terminal:** shards 0, 1, 2 GREEN --
358
+ the first shell shards to pass on a runner this session. Shard 3 red on
359
+ ShellCheck only, which `a4d8b839` fixes locally.
360
+ - **Local gate:** `scripts/local-ci.sh` exit 0, Failed 0, on every commit above.
361
+ - **Fork bomb:** contained 4,172 -> 375 processes. Root cause and counterfactual
362
+ in `ccf8dcbb` (already pushed).
363
+
364
+ ## 4. The calibration audit has NEVER run on real data
365
+
366
+ Searched this machine for council transcripts:
367
+
368
+ ```
369
+ find ~/git ~/loki-bench-v7 -type d -name transcripts -path "*council*" -> empty
370
+ find ~ -maxdepth 4 -type d -name council -> empty
371
+ ```
372
+
373
+ **No real council transcripts exist here.** Every calibration number produced
374
+ so far is fixture-derived, and the fixtures were built to have a hand-checkable
375
+ answer rather than to resemble production. The audit is correct against
376
+ hand-derived arithmetic; it is unvalidated against reality.
377
+
378
+ This is stated because "the tool works" and "the tool has told us something
379
+ about our system" are different claims, and only the first is supported.
380
+
381
+ ## 5. Risks
382
+
383
+ - **The calibration signal is not an accuracy signal, and it never becomes
384
+ one.** `tools/calibration-audit.py` scores agreement with the council
385
+ MAJORITY. The council outcome is derived from the votes, so a voter partly
386
+ causes its own label. A low Brier score there means "voted with the pack",
387
+ not "was correct". No artifact in this repo records whether the council was
388
+ right. This must stay a caveat; `90468210` pins by test that it cannot enter
389
+ model-advisor's machine-readable contract.
390
+ - **Local green is not CI green.** `a4d8b839` and `19c2b5ad` are verified only
391
+ on this machine. They are unpushed by instruction, so no runner has executed
392
+ them. Do not read the local 360/0 as a CI verdict.
393
+ - **Shard 3 has never been fully measured locally.** Its run was stopped during
394
+ the fork-bomb containment. Shards 0/1/2 were measured, partly during a window
395
+ when the process table was exhausted -- and in that state `grep` and `echo`
396
+ return EMPTY WITH NO ERROR, indistinguishable from "no matches". Readings
397
+ taken then are absent measurements, not results.
398
+ - **Three worktree agents' output was landed after review, and review found
399
+ real defects** (an unregistered guard that would never have run; a 48-vs-36
400
+ test-count misreport). Self-reports are not receipts.
401
+ - **Version skew is deliberate but load-bearing.** Anyone installing from npm
402
+ gets 9.8.1 and none of this work. If a user reports a bug fixed in 9.9-9.11,
403
+ that is why.
404
+
405
+ ## 6. Founder-gated next action
406
+
407
+ Exactly one decision unblocks the rest. All five commits are gate-green
408
+ locally and held only by the standing directive.
409
+
410
+ **Option A -- release the gate.** Push the 5 commits. CI then runs them for the
411
+ first time, which is the only way to confirm shard 3 goes green and the
412
+ rc143 fix holds on a runner. No publish is implied by a push.
413
+
414
+ **Option B -- keep the gate, name the next target.** Work continues and
415
+ continues to be held. The next unblocked item is the outcome frontier design
416
+ (in flight, no-spend).
417
+
418
+ **Option C -- authorize spend.** Only this unlocks measuring OUTCOME quality
419
+ against competitors. Every axis in the scorecard that reads UNKNOWN reads that
420
+ way because no paid run has happened.
421
+
422
+ The design is ready and unrun: `docs/OUTCOME-FRONTIER.md`, three held-out
423
+ tasks with deterministic non-self-grading oracles, a frozen three-dimension
424
+ rubric, and seven dimensions deliberately left UNSCORED rather than
425
+ approximated with an LLM judge.
426
+
427
+ **Estimated cost: roughly $25 to $180 in provider spend and 4 to 12 hours of
428
+ wall clock**, for a 4-tool, 2-task matrix. That is an ESTIMATE with its
429
+ assumptions listed in the document, not a measurement, and the two runnable
430
+ tasks (brownfield, multi-file migration) are the ones priced. The scientific
431
+ task is PLAN ONLY -- it needs a paper chosen and its number transcribed before
432
+ it can run at all, and its results must never be reported as equally rigorous
433
+ to the other two.
434
+
435
+ What approving this buys: the first evidence about OUTCOME rather than CLI
436
+ affordances. What it does not buy: the seven unscored dimensions, which no
437
+ amount of spend makes deterministic.
438
+
439
+ Nothing in this handoff requires spend, and nothing in it has been published.
@@ -2,7 +2,7 @@
2
2
 
3
3
  The flagship product of [Autonomi](https://www.autonomi.dev/). Loki Mode is a spec-driven autonomous builder with a built-in trust layer that takes any spec to a deployed product and verifies completion with evidence (quality gates plus a completion council), not just a "done" claim. Complete installation instructions for all platforms and use cases.
4
4
 
5
- **Version:** v8.0.0
5
+ **Version:** v9.8.1
6
6
 
7
7
  ---
8
8
 
@@ -36,7 +36,7 @@ Claude Code skill that compresses the model's OUTPUT tokens only (keeping all
36
36
  technical substance). It activates on free-form generation (the main RARV dev
37
37
  loop) and is HARD-SUPPRESSED on every trust-gate subcall (council votes, code
38
38
  review verdict, evidence-related parses) so determinism is never affected.
39
- - Claude-provider-only; runs are byte-identical on Codex / Cline / Aider.
39
+ - Claude-provider-only; runs are byte-identical on Cline / Codex / Aider / opencode.
40
40
  - Default on; opt out with `LOKI_CAVEMAN=0`.
41
41
  - Level: `LOKI_CAVEMAN_LEVEL` (default `full`; also `lite`, `ultra`, `wenyan*`).
42
42
  - Pinned + vendor-less: `LOKI_CAVEMAN_VERSION` (default `1.9.0`); Loki bootstraps
@@ -46,7 +46,7 @@ review verdict, evidence-related parses) so determinism is never affected.
46
46
 
47
47
  ### Earlier highlights still in scope
48
48
  - Bash-to-Bun runtime migration in progress (see `UPGRADING.md`)
49
- - Provider-agnostic runtime: Claude (full), Codex, Cline, Aider (no vendor lock-in)
49
+ - Provider-agnostic runtime: Claude (full), Cline, Codex, Aider, opencode (no vendor lock-in)
50
50
  - Memory system (episodic / semantic / procedural)
51
51
  - ChromaDB semantic code search via MCP
52
52
 
@@ -193,8 +193,9 @@ loki start ./spec.md
193
193
  loki start ./spec.md --provider codex
194
194
  loki start ./spec.md --provider cline
195
195
 
196
- # Spec as a GitHub issue
197
- loki start --github-issue https://github.com/owner/repo/issues/42
196
+ # Spec as a GitHub issue -- pass it positionally, it is auto-detected
197
+ loki start https://github.com/owner/repo/issues/42
198
+ loki start owner/repo#42
198
199
 
199
200
  # Spec as a YAML feature description
200
201
  loki start ./feature.yaml
@@ -298,7 +299,7 @@ By default, prompt injection is **disabled** for enterprise safety:
298
299
  ./autonomy/run.sh ./my-spec.md
299
300
 
300
301
  # Opt-in to enable prompt injection
301
- LOKI_PROMPT_INJECTION_ENABLED=true ./autonomy/run.sh ./my-spec.md
302
+ LOKI_PROMPT_INJECTION=true ./autonomy/run.sh ./my-spec.md
302
303
  ```
303
304
 
304
305
  #### Human Input Security
@@ -329,7 +330,13 @@ Loki Mode supports four active providers across three tiers, plus historical/upc
329
330
  | `cline` | Active | Tier 2 | Full feature set; small models (<13B) may fail tool-use. |
330
331
  | `codex` | Active | Tier 3 (degraded) | Sequential only, no Task tool; aligned with `@openai/codex` v0.125+. |
331
332
  | `aider` | Active | Tier 3 (degraded) | Sequential only; `ollama_chat/<model>` works for local models. |
332
- | `gemini` | DEPRECATED v7.5.18 | -- | Upstream Gemini CLI deprecated by Google. Runtime removed; `LOKI_PROVIDER=gemini` exits with migration message. |
333
+ | `opencode` | Active | Sequential | Model-agnostic route (`providers/opencode.sh`); autonomous flag `--auto`. Install: `npm install -g opencode-ai`. |
334
+ | `gemini` | REMOVED v7.5.18 | -- | Upstream Gemini CLI deprecated by Google. Runtime removed; `LOKI_PROVIDER=gemini` exits with a migration message. |
335
+
336
+ When `LOKI_PROVIDER` is unset, Loki auto-detects the first installed provider in
337
+ this order: `claude`, `cline`, `codex`, `aider`, `opencode`
338
+ (`providers/loader.sh:191`). An explicit `LOKI_PROVIDER` always wins and is
339
+ never silently substituted.
333
340
 
334
341
  ### Configuration
335
342
 
@@ -440,7 +447,7 @@ one-line export command and the compose volume to uncomment.
440
447
 
441
448
  ##### Other providers in Docker
442
449
 
443
- The image ships only the Claude Code CLI. Codex, Cline, and Aider are
450
+ The image ships only the Claude Code CLI. Cline, Codex, Aider, and opencode are
444
451
  bring-your-own-CLI: install the provider CLI in a derived image (or mount it),
445
452
  then select it with `-e LOKI_PROVIDER=<name>`. See
446
453
  [DOCKER_README.md](../DOCKER_README.md) for details.
@@ -498,7 +505,7 @@ Claude Code skill that compresses the model's OUTPUT tokens only (keeping all
498
505
  technical substance). It activates on free-form generation (the main RARV dev
499
506
  loop) and is HARD-SUPPRESSED on every trust-gate subcall (council votes, code
500
507
  review verdicts, evidence-related parses) so determinism is never affected. It
501
- is Claude-provider-only; runs are byte-identical on Codex / Cline / Aider. These
508
+ is Claude-provider-only; runs are byte-identical on Cline / Codex / Aider / opencode. These
502
509
  variables are read in `autonomy/lib/claude-flags.sh`.
503
510
 
504
511
  - `LOKI_CAVEMAN` (default on) -- set `LOKI_CAVEMAN=0` to disable the compressor.
@@ -843,7 +850,7 @@ The completion scripts support:
843
850
 
844
851
  * **Smart Context**
845
852
 
846
- * `loki start --provider <TAB>` shows only installed providers (`claude`, `codex`, `cline`, `aider`).
853
+ * `loki start --provider <TAB>` completes the supported providers (`claude`, `cline`, `codex`, `aider`, `opencode`; see `completions/loki.bash:24`).
847
854
  * `loki start <TAB>` defaults to file completion for spec files (PRD templates, YAML).
848
855
 
849
856
  * **Nested Commands**